Method and apparatus for using inter prediction information

By generating combined inter prediction information and using the combination of inter prediction information combination of adjacent blocks, the problem of insufficient inter prediction efficiency and accuracy in high-resolution and high-definition image encoding/decoding is solved, and the image encoding/decoding effect is improved.

CN120378629APending Publication Date: 2025-07-25ELECTRONICS & TELECOMM RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510499174.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-06-20
Filing Date
2018-10-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, in high resolution and high definition image encoding/decoding, inter prediction efficiency and accuracy need to be improved, especially in the derivation and use of prediction information of target blocks.

Method used

By generating combined inter prediction information, using two or more inter prediction information in neighboring blocks to combine, inter prediction information is derived and used to perform inter prediction, including motion vector weighted average of inter prediction information, a combination of block size and picture order counting, or a merged list or advanced motion vector prediction list is configured.

Benefits of technology

Improves the efficiency and accuracy of inter-frame prediction, and enhances the encoding/decoding effect of high-resolution and high-definition images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378629A_ABST
    Figure CN120378629A_ABST
Patent Text Reader

Abstract

A method and apparatus for using inter prediction information are disclosed. When performing video encoding or decoding, inter prediction information of a block to be encoded or decoded may be derived, and inter prediction of the block to be encoded or decoded may be performed using the derived inter prediction information. Combined inter prediction information may be generated by combining a plurality of pieces of inter prediction information, and the generated combined inter prediction information may be added as a candidate to a list for inter prediction. One of the candidates in the list may be selected for inter prediction of a block to be encoded or decoded, and inter prediction may be performed using the selected candidate.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a patent application for an invention named "Method and Apparatus for Using Inter-Frame Prediction Information" with an application date of October 10, 2018, an application number of 201880076042.X. Technical Field

[0002] The following embodiments generally relate to a video decoding method and apparatus and a video encoding method and apparatus, and more particularly, to a method and apparatus for using inter-frame prediction information in video encoding and decoding. Background Art

[0003] With the continuous development of the information and communication industry, broadcast services supporting high definition (HD) resolution have become widespread throughout the world. Through this widespread use, a large number of users have become accustomed to high-resolution and high-definition images and / or videos.

[0004] To meet the users' demand for high definition, a large number of institutions have accelerated the development of next-generation imaging devices. In addition to high-definition TVs (HDTVs) and full high-definition (FHD) TVs, users' interest in UHD TVs has also increased, where the resolution of UHD TVs is more than four times that of full high-definition (FHD) TVs. With the increase in their interest, there is a continuous need for image encoding / decoding techniques for images with higher resolution and higher clarity.

[0005] Image encoding / decoding devices and methods can use inter-frame prediction techniques, intra-frame prediction techniques, entropy encoding techniques, etc., in order to perform encoding / decoding on high-resolution and high-definition images. The inter-frame prediction technique can be a technique for predicting the values of pixels included in a target picture using a picture that is temporally previous and / or a picture that is temporally subsequent. The intra-frame prediction technique can be a technique for predicting the values of pixels included in a target picture using information about the pixels in the target picture. The entropy encoding technique can be a technique for assigning short codewords to frequently occurring symbols and long codewords to rarely occurring symbols.

[0006] Various prediction methods have been developed to improve the efficiency and accuracy of intra-frame prediction and / or inter-frame prediction. The prediction efficiency can vary greatly depending on which of the various applicable prediction methods is used for encoding and / or decoding a block. Summary of the Invention

[0007] Technical Problem

[0008] Embodiments are directed to providing an encoding device and method and a decoding device and method for performing inter-frame prediction on a target block.

[0009] An embodiment aims to provide an encoding device and method for deriving combined inter-prediction information for a target block and performing inter-prediction using the derived combined inter-prediction information, and a decoding device and method.

[0010] Solution

[0011] According to one aspect, there is provided an encoding device including: a processing unit configured to derive inter-prediction information for a target block and perform inter-prediction for the target block using the derived inter-prediction information, wherein the processing unit configures a list for the target block using combined inter-prediction information, and wherein the processing unit generates the combined inter-prediction information by combining two or more pieces of inter-prediction information of neighboring blocks of the target block.

[0012] According to another aspect, there is provided a decoding device including: a processing unit configured to derive inter-prediction information for a target block and perform inter-prediction for the target block using the derived inter-prediction information, wherein the processing unit configures a list for the target block using combined inter-prediction information, and wherein the processing unit generates the combined inter-prediction information by combining two or more pieces of inter-prediction information of neighboring blocks of the target block.

[0013] According to another aspect, there is provided a decoding method including: deriving inter-prediction information for a target block; and performing inter-prediction for the target block using the derived inter-prediction information, wherein a list for the target block is configured using combined inter-prediction information, and wherein the combined inter-prediction information is generated by combining two or more pieces of inter-prediction information of neighboring blocks of the target block.

[0014] The inter-prediction information may include at least one of an illumination compensation (IC) flag and an overlapping block motion compensation (OBMC) flag.

[0015] The list may be a merge list or an advanced motion vector prediction (AMVP) list.

[0016] The neighboring blocks may include spatial neighboring blocks and temporal neighboring blocks of the target block.

[0017] When the inter-prediction information of one of the neighboring blocks is not available, combined inter-prediction information for the one neighboring block may be derived.

[0018] When the inter-prediction information of one of the neighboring blocks is not added to the list, the combined inter-prediction information derived for the one neighboring block may be added to the list.

[0019] The motion vector of the combined inter-prediction information may be a result of a formula using the motion vectors of the neighboring blocks.

[0020] The motion vector for combining the inter-frame prediction information may be a weighted average of the motion vectors of neighboring blocks.

[0021] The motion vector for combining the inter-frame prediction information may be the result of a weighted combination of the motion vectors of neighboring blocks based on the block size.

[0022] The motion vector for combining the inter-frame prediction information may be the result of a weighted combination of the motion vectors of neighboring blocks based on the picture order count (POC).

[0023] The motion vector for combining the inter-frame prediction information may be the result of a combination of the motion vectors of neighboring blocks based on extrapolation.

[0024] First inter-frame prediction information related to the left position of a specific block and second inter-frame prediction information related to the right position of the specific block may be derived, and the combined inter-frame prediction information may be generated by combining the first inter-frame prediction information and the second inter-frame prediction information.

[0025] The scheme for configuring the list may be determined based on the shape of the target block.

[0026] The scheme for configuring the list may be determined based on the partitioning state of the target block.

[0027] The scheme for configuring the list may be determined based on the position of the target block.

[0028] The combined inter-frame prediction information may be added to the list with a lower priority than the priorities of multiple inter-frame prediction information of neighboring blocks.

[0029] The combined inter-frame prediction information may be added to a position in the list after the inter-frame prediction information of spatially neighboring blocks and before the inter-frame prediction information of temporally neighboring blocks.

[0030] The scheme for configuring the list may be determined based on the depth of the target block.

[0031] Advantageous Effects

[0032] There are provided an encoding device and method and a decoding device and method for performing inter-frame prediction on a target block.

[0033] There are provided an encoding device and method and a decoding device and method for deriving combined inter-frame prediction information for a target block and performing inter-frame prediction using the derived combined inter-frame prediction information. Description of the Drawings

[0034] Figure 1 is a block diagram showing the configuration of an embodiment of an encoding device to which the present disclosure is applied;

[0035] Figure 2 is a block diagram showing the configuration of an embodiment of a decoding device to which the present disclosure is applied;

[0036] Figure 3 is a diagram schematically showing the partitioning structure of an image when the image is encoded and decoded;

[0037] Figure 4 is a diagram showing the forms of prediction units (PUs) that a coding unit (CU) can include;

[0038] Figure 5 is a diagram showing the forms of transform units (TUs) that can be included in a CU;

[0039] Figure 6 is a diagram for explaining an embodiment of the intra prediction process;

[0040] Figure 7 is a diagram for explaining the positions of reference sample points used in the intra prediction process;

[0041] Figure 8 is a diagram for explaining an embodiment of the inter prediction process;

[0042] Figure 9 shows spatial candidates according to an embodiment;

[0043] Figure 10 shows the order of adding the motion information of spatial candidates to the merge list according to an embodiment;

[0044] Figure 11 shows transform and quantization processing according to an example;

[0045] Figure 12 is a configuration diagram of an encoding device according to an embodiment;

[0046] Figure 13 is a configuration diagram of a decoding device according to an embodiment;

[0047] Figure 14 is a flowchart showing an inter prediction method according to an embodiment;

[0048] Figure 15 shows spatial neighboring blocks of a target block according to an example;

[0049] Figure 16 shows temporal neighboring blocks of a target block according to an example;

[0050] Figure 17 shows the generation of combined inter prediction information for upper-right neighboring blocks according to an example;

[0051] Figure 18Shows the generation of combined inter-prediction information for upper neighboring blocks according to an example;

[0052] Figure 19 Shows the generation of combined inter-prediction information for neighboring blocks according to an example;

[0053] Figure 20 Shows the generation of inter-prediction information for block AL according to an example;

[0054] Figure 21 Shows the generation of inter-prediction information for block AR according to an example;

[0055] Figure 22 Shows the generation of inter-prediction information for a target CU according to an example;

[0056] Figure 23 Shows the case where a CU with the same width and height is vertically divided;

[0057] Figure 24 Shows the case where a CU with the same width and height is horizontally divided;

[0058] Figure 25 Shows the case where a CU with a width greater than its height is vertically divided;

[0059] Figure 26 Shows the case where a CU with a height greater than its width is horizontally divided;

[0060] Figure 27 Shows sub-blocks of a temporal neighboring block and sub-blocks of a target block according to an example;

[0061] Figure 28 Shows spatial neighboring blocks of a target block and sub-blocks of the target block according to an example;

[0062] Figure 29 Shows the derivation of inter-prediction information using bilateral matching according to an example;

[0063] Figure 30 Shows the derivation of inter-prediction information using a template matching mode according to an example;

[0064] Figure 31 Shows the application of OBMC according to an example;

[0065] Figure 32 Shows sub-PUs in the ATMVP mode according to an example;

[0066] Figure 33 Is a flowchart showing a target block prediction method and a bitstream generation method according to an embodiment; and

[0067] Figure 34It is a flowchart showing a target block prediction method using a bitstream according to an embodiment. Embodiment

[0069] The present invention can be variously changed and can have various embodiments. Specific embodiments will be described in detail below with reference to the accompanying drawings. However, it should be understood that these embodiments are not intended to limit the present invention to a specific disclosed form, and they include all changes, equivalent forms, or modifications included within the spirit and scope of the present invention.

[0070] The following exemplary embodiments will be described in detail with reference to the accompanying drawings showing specific embodiments. These embodiments are described so that those of ordinary skill in the art to which the present disclosure pertains can easily practice these embodiments. It should be noted that the various embodiments are different from each other but do not need to be mutually exclusive. For example, the specific shapes, structures, and characteristics described herein can be implemented as other embodiments without departing from the spirit and scope of the multiple embodiments related to one embodiment. In addition, it should be understood that the positions or arrangements of the respective components in each disclosed embodiment can be changed without departing from the spirit and scope of the embodiment. Therefore, the following detailed description is not intended to limit the scope of the present disclosure, and the scope of the exemplary embodiments is only defined by the appended claims and their equivalents (as long as they are properly described).

[0071] In the drawings, like reference numerals are used to designate the same or similar functions in various aspects. The shapes, sizes, etc. of the components in the drawings may be exaggerated to make the description clear.

[0072] Terms such as "first" and "second" may be used to describe various components, but the components are not limited by these terms. These terms are only used to distinguish one component from another. For example, without departing from the scope of this specification, the first component may be referred to as the second component. Similarly, the second component may be referred to as the first component. The term "and / or" may include combinations of multiple related description items or any one of the multiple related description items.

[0073] It will be understood that when a component is referred to as being "connected" or "coupled" to another component, the two components may be directly connected or coupled to each other, or there may be an intermediate component between the two components. It will be understood that when a component is referred to as being "directly connected or coupled", there is no intermediate component between the two components.

[0074] In addition, the components described in the embodiments are shown independently to represent different characteristic functions, but this does not mean that each component is formed by a single piece of hardware or software. That is, for convenience of description, multiple components are arranged and included separately. For example, at least two of the multiple components may be integrated into a single component. Conversely, one component may be divided into multiple components. As long as it does not deviate from the essence of this specification, embodiments in which multiple components are integrated or embodiments in which some components are separated are included within the scope of this specification.

[0075] In addition, it should be noted that in the exemplary embodiments, the expression of describing a component as "including" a specific component means that additional components may be included within the scope of the practice or technical spirit of the exemplary embodiments, but it does not exclude the existence of components other than the said specific component.

[0076] The terms used in this specification are only for describing specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless specifically stated to the contrary in the context. In this specification, it should be understood that terms such as "including" or "having" are only intended to indicate the existence of features, numbers, steps, operations, components, parts, or combinations thereof, and are not intended to exclude the possibility of the existence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0077] Embodiments will be described in detail below with reference to the accompanying drawings, so that those of ordinary skill in the art to which the embodiments belong can easily practice the embodiments. In the following description of the embodiments, a detailed description of well-known functions or configurations that are considered to obscure the gist of this specification will be omitted. In addition, the same reference numerals are used throughout the drawings to designate the same components, and repeated descriptions of the same components will be omitted.

[0078] Hereinafter, an "image" may represent a single frame constituting a video, or may represent the video itself. For example, "encoding and / or decoding of an image" may represent "encoding and / or decoding of a video", and may also represent "encoding and / or decoding of any one of the multiple images constituting the video".

[0079] Hereinafter, the terms "video" and "moving picture" may be used with the same meaning and may be used interchangeably with each other.

[0080] Hereinafter, a target image may be an encoding target image that is a target to be encoded and / or a decoding target image that is a target to be decoded. In addition, the target image may be an input image input to an encoding device or an input image input to a decoding device.

[0081] Hereinafter, the terms "image", "picture", "frame", and "screen" may be used with the same meaning and may be used interchangeably with each other.

[0082] Hereinafter, a target block may be an encoding target block (i.e., the target to be encoded) and / or a decoding target block (i.e., the target to be decoded). In addition, the target block may be the current block, i.e., the current target to be encoded and / or decoded. Here, the terms "target block" and "current block" may be used with the same meaning and may be used interchangeably with each other.

[0083] Hereinafter, the terms "block" and "unit" may be used with the same meaning and may be used interchangeably with each other. Optionally, a "block" may represent a specific unit.

[0084] Hereinafter, the terms "region" and "section" may be used interchangeably with each other.

[0085] Hereinafter, a specific signal may be a signal indicating a specific block. For example, an original signal may be a signal indicating a target block. A prediction signal may be a signal indicating a prediction block. A residual signal may be a signal indicating a residual block.

[0086] In the following embodiments, specific information, data, flags, elements, and attributes may have their respective values. The value "0" corresponding to each of the information, data, flags, elements, and attributes may indicate logical false or a first predefined value. In other words, the values "0", false, logical false, and the first predefined value may be used interchangeably with each other. The value "1" corresponding to each of the information, data, flags, elements, and attributes may indicate logical true or a second predefined value. In other words, the values "1", true, logical true, and the second predefined value may be used interchangeably with each other.

[0087] When variables such as i or j are used to indicate a row, column, or index, the value of i may be the integer 0 or an integer greater than 0, or may be the integer 1 or an integer greater than 1. In other words, in an embodiment, each of the row, column, and index may be counted starting from 0, or may be counted starting from 1.

[0088] Next, terms to be used in the embodiments will be described.

[0089] Encoder: An encoder represents a device for performing encoding.

[0090] Decoder: A decoder represents a device for performing decoding.

[0091] Unit: A "unit" may represent a unit of image encoding and decoding. The terms "unit" and "block" may be used with the same meaning and may be used interchangeably with each other.

[0092] – A “unit” can be an M×N sample array. M and N can be positive integers respectively. The term “unit” generally can represent a two-dimensional (2D) sample array.

[0093] – During the encoding and decoding processes of an image, a “unit” can be a region generated by partitioning an image. A single image can be partitioned into multiple units. Optionally, an image can be partitioned into sub-parts, and a unit can represent each partitioned sub-part when encoding or decoding is performed on the partitioned sub-parts.

[0094] – During the encoding and decoding processes of an image, predefined processing can be performed on each unit according to the type of the unit.

[0095] – According to functions, unit types can be classified as macro units, coding units (CUs), prediction units (PUs), residual units, transform units (TUs), etc. Optionally, according to functions, a unit can represent a block, a macro block, a coding tree unit (CTU), a coding tree block, a coding unit, a coding block, a prediction unit, a prediction block, a residual unit, a residual block, a transform unit, a transform block, etc.

[0096] – The term “unit” can represent information including a luma component block, a chroma component block corresponding to the luma component block, and syntax elements for each block, such that the unit is designated to be distinct from a block.

[0097] – The size and shape of a unit can be implemented differently. In addition, a unit can have any one of various sizes and shapes. Specifically, the shape of a unit can include not only a square but also geometric shapes (such as rectangles, trapezoids, triangles, and pentagons) that can be represented in two dimensions (2D).

[0098] – In addition, unit information can include one or more of the type of the unit (indicating a coding unit, a prediction unit, a residual unit, or a transform unit), the size of the unit, the depth of the unit, the order of encoding and decoding of the unit, etc.

[0099] – A unit can be partitioned into sub-units, each sub-unit having a size smaller than that of the relevant unit.

[0100] – Depth: Depth can represent the degree to which a unit is partitioned. In addition, depth can indicate the level at which a corresponding unit exists when the unit is represented in a tree structure.

[0101] – Unit partition information can include a depth indicating the depth of the unit. The depth can indicate the number of times a unit is partitioned and / or the degree to which a unit is partitioned.

[0102] – In a tree structure, it can be considered that the depth of the root node is the smallest and the depth of the leaf node is the largest.

[0103] – A single unit can be hierarchically partitioned into multiple sub-units, while the single unit has depth information based on a tree structure. In other words, the unit and the sub-units generated by partitioning the unit can respectively correspond to a node and the child nodes of the node. Each partitioned sub-unit can have a depth. Since the depth indicates the number of times the unit is partitioned and / or the degree of partitioning of the unit, the partitioning information of the sub-unit can include information about the size of the sub-unit.

[0104] – In a tree structure, the top node can correspond to the initial node before partitioning. The top node can be referred to as the "root node". In addition, the root node can have the minimum depth value. Here, the depth of the top node can be level "0".

[0105] – A node with a depth of level "1" can represent the unit generated when the initial unit is partitioned once. A node with a depth of level "2" can represent the unit generated when the initial unit is partitioned twice.

[0106] – A leaf node with a depth of level "n" can represent the unit generated when the initial unit is partitioned n times.

[0107] – A leaf node can be the bottom node that cannot be further partitioned. The depth of the leaf node can be the maximum level. For example, the predefined value for the maximum level can be 3.

[0108] – QT depth can represent the depth of quadtree partitioning. BT depth can represent the depth of binary tree partitioning. TT depth can represent the depth of ternary tree partitioning.

[0109] – Sample point: A sample point can be the basic unit that constitutes a block. The sample point can be represented by values from 0 to 2 Bd- ^1 according to the bit depth (Bd).

[0110] – A sample point can be a pixel or a pixel value.

[0111] – Hereinafter, the terms "pixel" and "sample point" can be used with the same meaning and can be used interchangeably with each other.

[0112] Coding tree unit (CTU): A CTU can be composed of a single luma component (Y) coding tree block and two chroma component (Cb, Cr) coding tree blocks related to the luma component coding tree block. In addition, a CTU can represent the information including the above blocks and the syntax elements for each block.

[0113] – Each coding tree unit (CTU) can be partitioned using one or more partitioning methods (such as quadtree (QT), binary tree (BT), and ternary tree (TT)) to configure sub-units such as coding units, prediction units, and transform units.

[0114] – "CTU" can be used as a term to specify a pixel block that serves as a processing unit in image decoding and encoding processes (such as in the case of partitioning an input image).

[0115] Coding Tree Block (CTB): "CTB" can be used as a term to specify any one of a Y coding tree block, a Cb coding tree block, and a Cr coding tree block.

[0116] – Neighboring (adjacent) block: A neighboring block represents a block adjacent to a target block. The term "neighboring block" can also refer to a reconstructed neighboring block.

[0117] In the following text, the terms "neighboring block" and "adjacent block" can be used with the same meaning and can be used interchangeably with each other.

[0118] – Spatial neighboring block: A spatial neighboring block can be a block that is spatially adjacent to a target block. A neighboring block can include a spatial neighboring block.

[0119] – A target block and a spatial neighboring block can be included in a target picture.

[0120] – A spatial neighboring block can be a block that touches the boundary of a target block or a block located at a predetermined distance from the target block.

[0121] – A spatial neighboring block can be a block adjacent to a vertex of a target block. Here, a block adjacent to a vertex of a target block can be a block that is vertically adjacent to a neighboring block horizontally adjacent to the target block or a block that is horizontally adjacent to a neighboring block vertically adjacent to the target block.

[0122] – Temporal neighboring block: A temporal neighboring block can be a block that is temporally adjacent to a target block. A neighboring block can include a temporal neighboring block.

[0123] – A temporal neighboring block can include a col block.

[0124] – A col block can be a block in a previously reconstructed col picture. The position of the col block in the col picture can correspond to the position of the target block in the target picture. The col picture can be a picture included in a reference picture list.

[0125] – A temporal neighboring block can be a spatial neighboring block of a target block.

[0126] Prediction unit: A prediction unit can be a basic unit for prediction (such as inter-frame prediction, intra-frame prediction, inter-frame compensation, intra-frame compensation, and motion compensation).

[0127] – A single prediction unit can be divided into multiple partitions or sub-prediction units with smaller sizes. The multiple partitions can also be basic units during prediction or compensation. Partitions generated by dividing a prediction unit can also be prediction units.

[0128] Prediction unit partition: The prediction unit partition can be the shape into which the prediction unit is divided.

[0129] Reconstructed neighboring unit: The reconstructed neighboring unit can be a unit that has been decoded and reconstructed around the target unit.

[0130] – The reconstructed neighboring unit can be a unit that is spatially adjacent to the target unit or temporally adjacent to the target unit.

[0131] – The reconstructed spatial neighboring unit can be a unit that has been reconstructed through encoding and / or decoding and is included in the target picture.

[0132] – The reconstructed temporal neighboring unit can be a unit that has been reconstructed through encoding and / or decoding and is included in the reference picture. The position of the reconstructed temporal neighboring unit in the reference picture can be the same as the position of the target unit in the target picture, or can correspond to the position of the target unit in the target picture.

[0133] Parameter set: The parameter set can be the header information in the structure of the bitstream. For example, the parameter set can include a sequence parameter set, a picture parameter set, an adaptive parameter set, etc.

[0134] Rate distortion optimization: The encoding device can use rate distortion optimization to provide high encoding efficiency by using a combination of the following: the size of the coding unit (CU), the prediction mode, the size of the prediction unit (PU), the motion information, and the size of the transform unit (TU).

[0135] – The rate distortion optimization scheme can calculate the rate distortion cost for each combination to select the optimal combination from these combinations. Equation 1 below can be used to calculate the rate distortion cost. Generally, the combination that minimizes the rate distortion cost can be selected as the optimal combination under the rate distortion optimization scheme.

[0136] [Equation 1]

[0137] D + λ * R

[0138] – D can represent the distortion. D can be the average of the squares of the differences between the original transform coefficients and the reconstructed transform coefficients in the transform unit (i.e., the mean squared error).

[0139] – R can represent the rate, which can use relevant context information to represent the bitrate.

[0140] – λ represents the Lagrange multiplier. R can include not only the coding parameter information (such as the prediction mode, the motion information, and the coding block flag), but also the bits generated due to encoding the transform coefficients.

[0141] – The encoding device may perform processes such as inter - frame prediction and / or intra - frame prediction, transformation, quantization, entropy coding, inverse quantization (de - quantization), and inverse transformation in order to calculate accurate D and R. These processes greatly increase the complexity of the encoding device.

[0142] – Bitstream: A bitstream may represent a stream of bits including encoded image information.

[0143] – Parameter set: A parameter set may be header information in the structure of a bitstream.

[0144] The parameter set may include at least one of a video parameter set, a sequence parameter set, a picture parameter set, and an adaptive parameter set. In addition, the parameter set may include information about slice headers and information about parallel block headers.

[0145] Parsing: Parsing may be a determination of the value of a syntax element made by performing entropy decoding on a bitstream. Optionally, the term "parsing" may represent this entropy decoding itself.

[0146] Symbol: A symbol may be at least one of a syntax element of an encoded target unit and / or a decoded target unit, an encoding parameter, and a transform coefficient. In addition, a symbol may be the target of entropy coding or the result of entropy decoding.

[0147] Reference picture: A reference picture may be an image that is referenced by a unit to perform inter - frame prediction or motion compensation. Optionally, a reference picture may be an image including reference units that are referenced by a target unit to perform inter - frame prediction or motion compensation.

[0148] Hereinafter, the terms "reference picture" and "reference image" may be used with the same meaning and may be used interchangeably with each other.

[0149] Reference picture list: A reference picture list may be a list including one or more reference images used for inter - frame prediction or motion compensation.

[0150] – The types of reference picture lists may include a merged list (LC), list 0 (L0), list 1 (L1), list 2 (L3), list 3 (L3), etc.

[0151] – For inter - frame prediction, one or more reference picture lists may be used.

[0152] Inter - frame prediction indicator: An inter - frame prediction indicator may indicate the inter - frame prediction direction of a target unit. Inter - frame prediction may be one of unidirectional prediction and bidirectional prediction. Optionally, the inter - frame prediction indicator may represent the number of reference images used to generate a prediction unit of a target unit. Optionally, the inter - frame prediction indicator may represent the number of prediction blocks used for inter - frame prediction or motion compensation of a target unit.

[0153] Reference picture index: The reference picture index can be an index indicating a specific reference image in the reference picture list.

[0154] Motion vector (MV): The motion vector can be a 2D vector used for inter-frame prediction or motion compensation. The motion vector can represent the offset between the encoded / decoded target image and the reference image.

[0155] – For example, the MV can be represented in a form such as (mv x , mv y ). mv x can indicate the horizontal component, and mv y can indicate the vertical component.

[0156] – Search range: The search range can be a 2D area where the search for the MV is performed during inter-frame prediction. For example, the size of the search range can be M×N. M and N can be positive integers respectively.

[0157] Motion vector candidate: The motion vector candidate can be a block that is a prediction candidate when the motion vector is predicted or the motion vector of the block that is a prediction candidate.

[0158] – The motion vector candidate can be included in the motion vector candidate list.

[0159] Motion vector candidate list: The motion vector candidate list can be a list configured using one or more motion vector candidates.

[0160] Motion vector candidate index: The motion vector candidate index can be an indicator used to indicate the motion vector candidate in the motion vector candidate list. Optionally, the motion vector candidate index can be an index of the motion vector predictor.

[0161] Motion information: The motion information can be information including at least one of the reference picture list, reference image, motion vector candidate, motion vector candidate index, merge candidate, and merge index, as well as the motion vector, reference picture index, and inter-frame prediction indicator.

[0162] Merge candidate list: The merge candidate list can be a list configured using merge candidates.

[0163] Merge candidate: The merge candidate can be a spatial merge candidate, temporal merge candidate, combined merge candidate, combined bi-prediction merge candidate, zero merge candidate, etc. The merge candidate can include motion information such as the inter-frame prediction indicator, reference picture index for each list, and motion vector.

[0164] Merge index: The merge index can be an indicator used to indicate the merge candidate in the merge candidate list.

[0165] – The merge index may indicate a reconstruction unit for deriving a merge candidate between a reconstruction unit adjacent to a target unit spatially and a reconstruction unit adjacent to the target unit temporally.

[0166] – The merge index may indicate at least one of multiple pieces of motion information of a merge candidate.

[0167] Transform unit: A transform unit may be a basic unit for residual signal encoding and / or residual signal decoding (such as transform, inverse transform, quantization, dequantization, transform coefficient encoding, and transform coefficient decoding). A single transform unit may be partitioned into multiple transform units with smaller sizes.

[0168] Scaling: Scaling may represent a process of multiplying a factor by a transform coefficient level.

[0169] – As a result of scaling the transform coefficient level, transform coefficients may be generated. Scaling may also be referred to as “dequantization”.

[0170] Quantization parameter (QP): A quantization parameter may be a value used to generate a transform coefficient level for transform coefficients in quantization. Optionally, the quantization parameter may also be a value used to generate transform coefficients by scaling a transform coefficient level in dequantization. Optionally, the quantization parameter may be a value mapped to a quantization step size.

[0171] Variable delta (Delta) quantization parameter: A variable delta quantization parameter is the difference between the quantization parameter of an encoded / decoded target unit and a predicted quantization parameter.

[0172] Scanning: Scanning may represent a method of arranging the coefficient order in a unit, block, or matrix. For example, a method for arranging a 2D array in the form of a one-dimensional (1D) array may be referred to as “scanning”. Optionally, a method for arranging a 1D array in the form of a 2D array may also be referred to as “scanning” or “inverse scanning”.

[0173] Transform coefficient: A transform coefficient may be a coefficient value generated when an encoding device performs a transform. Optionally, a transform coefficient may be a coefficient value generated when a decoding device performs at least one of entropy decoding and dequantization.

[0174] – When quantization is applied to a transform coefficient or the quantization level of a residual signal quantization or the quantized transform coefficient level may also be included in the meaning of the term “transform coefficient”.

[0175] Quantization level: A quantization level may be a value generated when an encoding device performs quantization on a transform coefficient or a residual signal. Optionally, a quantization level may be a value targeted for dequantization when a decoding device performs dequantization.

[0176] – The quantized transform coefficient levels resulting from transformation and quantization may also be included in the meaning of the quantized levels.

[0177] Non-zero transform coefficient: A non-zero transform coefficient may be a transform coefficient having a value other than 0, or may be a transform coefficient level having a value other than 0. Optionally, a non-zero transform coefficient may be a transform coefficient whose value magnitude is not 0, or may be a transform coefficient level whose value magnitude is not 0.

[0178] Quantization matrix: A quantization matrix may be a matrix used in the quantization or inverse quantization process to improve the subjective or objective image quality of an image. The quantization matrix may also be referred to as a "scaling list".

[0179] Quantization matrix coefficient: A quantization matrix coefficient may be each element in the quantization matrix. The quantization matrix coefficient may also be referred to as a "matrix coefficient".

[0180] Default matrix: A default matrix may be a quantization matrix predefined by an encoding device and a decoding device.

[0181] Non-default matrix: A non-default matrix may be a quantization matrix not predefined by an encoding device and a decoding device. The non-default matrix may be signaled by the encoding device to the decoding device.

[0182] Signaling: Signaling may indicate that information is sent from an encoding device to a decoding device. Optionally, signaling may represent that information is included in a bitstream or a storage medium. The information signaled by the encoding device may be used by the decoding device.

[0183] Figure 1 is a block diagram showing the configuration of an embodiment of an encoding device to which the present disclosure is applied.

[0184] Encoding device 100 may be an encoder, a video encoding device, or an image encoding device. The video may include one or more images (frames). Encoding device 100 may sequentially encode one or more images of the video.

[0185] Refer to Figure 1 , encoding device 100 includes an inter prediction unit 110, an intra prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.

[0186] Encoding device 100 may perform encoding on a target image using an intra mode and / or an inter mode.

[0187] In addition, the encoding device 100 can generate a bitstream including information about the encoding by encoding a target image, and can output the generated bitstream. The generated bitstream can be stored in a computer-readable storage medium and can be streamed via a wireless / wired transmission medium.

[0188] When the intra mode is used as the prediction mode, the switch 115 can switch to the intra mode. When the inter mode is used as the prediction mode, the switch 115 can switch to the inter mode.

[0189] The encoding device 100 can generate a prediction block of a target block. In addition, after the prediction block has been generated, the encoding device 100 can encode the residual between the target block and the prediction block.

[0190] When the prediction mode is the intra mode, the intra prediction unit 120 can use the pixels of previously encoded / decoded neighboring blocks around the target block as reference sample points. The intra prediction unit 120 can perform spatial prediction on the target block using the reference sample points and can generate prediction sample points for the target block via spatial prediction.

[0191] The inter prediction unit 110 can include a motion prediction unit and a motion compensation unit.

[0192] When the prediction mode is the inter mode, the motion prediction unit can search for the region in the reference image that best matches the target block during the motion prediction process, and can derive a motion vector for the target block and the found region based on the found region.

[0193] The reference image can be stored in the reference picture buffer 190. More specifically, when the encoding and / or decoding of the reference image has been processed, the reference image can be stored in the reference picture buffer 190.

[0194] The motion compensation unit can generate a prediction block for the target block by performing motion compensation using the motion vector. Here, the motion vector can be a two-dimensional (2D) vector for inter prediction. In addition, the motion vector can represent the offset between the target image and the reference image.

[0195] When the motion vector has a value other than an integer, the motion prediction unit and the motion compensation unit can generate a prediction block by applying an interpolation filter to a partial region of the reference image. To perform inter prediction or motion compensation, it can be determined which of the skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to the method for predicting and compensating the motion of a PU included in a CU based on the CU, and inter prediction or motion compensation can be performed according to the mode.

[0196] The subtractor 125 may generate a residual block, where the residual block is the difference between a target block and a prediction block. The residual block may also be referred to as a "residual signal".

[0197] The residual signal may be the difference between an original signal and a prediction signal. Optionally, the residual signal may be a signal generated by transforming or quantizing the difference between the original signal and the prediction signal or a signal generated by transforming and quantizing the difference. The residual block may be a residual signal for a block unit.

[0198] The transform unit 130 may generate transform coefficients by transforming the residual block and may output the generated transform coefficients. Here, the transform coefficients may be coefficient values generated by transforming the residual block.

[0199] When using the transform skip mode, the transform unit 130 may omit the operation of transforming the residual block.

[0200] By performing quantization on the transform coefficients, quantized transform coefficient levels or quantized levels may be generated. Hereinafter, in the embodiments, each of the quantized transform coefficient levels and the quantized levels may also be referred to as "transform coefficients".

[0201] The quantization unit 140 may generate quantized transform coefficient levels or quantized levels by quantizing the transform coefficients according to quantization parameters. The quantization unit 140 may output the generated quantized transform coefficient levels or quantized levels. In this case, the quantization unit 140 may use a quantization matrix to quantize the transform coefficients.

[0202] The entropy coding unit 150 may generate a bitstream by performing entropy coding based on probability distribution on the values calculated by the quantization unit 140 and / or the coding parameter values calculated during the coding process. The entropy coding unit 150 may output the generated bitstream.

[0203] The entropy coding unit 150 may perform entropy coding on information about pixels of an image and information required for decoding the image. For example, the information required for decoding the image may include syntax elements and the like.

[0204] The coding parameters may be information required for encoding and / or decoding. The coding parameters may include information encoded by the encoding device 100 and transmitted from the encoding device 100 to the decoding device, and may also include information derived during the encoding or decoding process. For example, the information transmitted to the decoding device may include syntax elements.

[0205] For example, the coding parameters may include values or statistical information, such as a prediction mode, a motion vector, a reference picture index, a coded block style, the presence or absence of a residual signal, transform coefficients, quantized transform coefficients, quantization parameters, block size, and block partitioning information. The prediction mode may be an intra prediction mode or an inter prediction mode.

[0206] The residual signal may represent the difference between the original signal and the predicted signal. Optionally, the residual signal may be a signal generated by transforming the difference between the original signal and the predicted signal. Optionally, the residual signal may be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal.

[0207] When entropy coding is applied, fewer bits may be allocated to more frequently occurring symbols, and more bits may be allocated to less frequently occurring symbols. Since the symbols are represented by this allocation, the size of the bit string for the target symbols to be coded can be reduced. Thus, the compression performance of video coding can be improved by entropy coding.

[0208] In addition, for entropy coding, the entropy coding unit 150 may use coding methods such as exponential Golomb, context adaptive variable length coding (CAVLC), or context adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 may use a variable length coding / code (VLC) table to perform entropy coding. For example, the entropy coding unit 150 may derive a binarization method for the target symbol. In addition, the entropy coding unit 150 may derive a probability model for the target symbol / bits. The entropy coding unit 150 may use the derived binarization method, probability model, and context model to perform arithmetic coding.

[0209] The entropy coding unit 150 may transform the coefficients in 2D block form into 1D vector form by a transform coefficient scanning method in order to code the transform coefficient levels.

[0210] Coding coefficients may include not only information such as syntax elements (or flags or indices) that are coded by an encoding device and signaled by the encoding device to a decoding device, but may also include information derived during the encoding or decoding process. Additionally, coding parameters may include information required to encode or decode an image.For example, the coding parameters may include at least one of the following items or a combination of the following items: the size of a unit / block, the depth of a unit / block, the partition information of a unit / block, the partition structure of a unit / block, information indicating whether a unit / block is partitioned in a quadtree structure, information indicating whether a unit / block is partitioned in a binary tree (BT) structure, the partition direction (horizontal direction or vertical direction) of the binary tree structure, the partition form (symmetric partition or asymmetric partition) of the binary tree structure, information indicating whether a unit / block is partitioned in a ternary tree structure, the partition direction (horizontal direction or vertical direction) of the ternary tree structure, a prediction scheme (intra prediction or inter prediction), an intra prediction mode / direction, a reference sample filtering method, a prediction block filtering method, a prediction block boundary filtering method, filter taps for filtering, filter coefficients for filtering, an inter prediction mode, motion information, a motion vector, a reference picture index, an inter prediction direction, an inter prediction indicator, a reference picture list, a reference image, a motion vector prediction factor, a motion vector prediction candidate, a motion vector candidate list, information indicating whether a merge mode is used, a merge candidate, a merge candidate list, information indicating whether a skip mode is used, the type of an interpolation filter, the taps of an interpolation filter, the filter coefficients of an interpolation filter, the size of a motion vector, the precision of the motion vector representation, a transform type, a transform size, information indicating whether a first transform is used, information indicating whether an additional (second) transform is used, a first transform index, a second transform index, information indicating the presence or absence of a residual signal, a coded block style, a coded block flag, a quantization parameter, a quantization matrix, information about an in-loop filter, information indicating whether an in-loop filter is applied, the coefficients of an in-loop filter, the taps of an in-loop filter, the shape / form of an in-loop filter, information indicating whether a deblocking filter is applied, the coefficients of a deblocking filter, the taps of a deblocking filter, the deblocking filter strength, the shape / form of a deblocking filter, information indicating whether an adaptive sample offset is applied, the value of an adaptive sample offset, the category of an adaptive sample offset, the type of an adaptive sample offset, information indicating whether an adaptive loop filter is applied, the coefficients of an adaptive loop filter, the taps of an adaptive loop filter, the shape / form of an adaptive loop filter, a binarization / de-binarization method, a context model, a context model determination method, a context model update method, information indicating whether a normal mode is executed, information indicating whether a bypass mode is executed, context bits, bypass bits, transform coefficients, transform coefficient levels, a transform coefficient level scanning method, an image display / output order, slice identification information, slice type, slice partition information, parallel block identification information, parallel block type, parallel block partition information, picture type, bit depth, information about a luminance signal, and information about a chrominance signal.

[0211] Here, it can be indicated by a signaling flag or index that the encoding device 100 includes an entropy-coded flag or an entropy-coded index generated by performing entropy coding on the flag or index in the bitstream, and it can be indicated that the decoding device 200 obtains the flag or index by performing entropy decoding on the entropy-coded flag or the entropy-coded index extracted from the bitstream.

[0212] Since the encoding device 100 performs encoding via inter-frame prediction, the encoded target image can be used as a reference image for another image to be subsequently processed. Therefore, the encoding device 100 can reconstruct or decode the encoded target image and store the reconstructed or decoded image in the reference picture buffer 190 as a reference image. For decoding, inverse quantization and inverse transformation of the encoded target image can be performed.

[0213] The quantization levels can be inverse quantized by the inverse quantization unit 160 and can be inverse transformed by the inverse transformation unit 170. The coefficients that have been inverse quantized and / or inverse transformed can be added to the prediction block by the adder 175. By adding the inverse quantized and / or inverse transformed coefficients and the prediction block, a reconstructed block can then be generated. Here, the inverse quantized and / or inverse transformed coefficients can represent the coefficients on which one or more of inverse quantization and inverse transformation have been performed, and can also represent the reconstructed residual block.

[0214] The reconstructed block can be filtered by the filter unit 180. The filter unit 180 can apply one or more of a deblocking filter, a sample adaptive offset (SAO) filter, and an adaptive loop filter (ALF) to the reconstructed block or the reconstructed picture. The filter unit 180 can also be referred to as a "loop filter".

[0215] The deblocking filter can eliminate block distortion that appears at the boundaries between blocks. To determine whether to apply the deblocking filter, it can be decided the number of columns or rows of pixels included in the block and that include the basis for determining whether to apply the deblocking filter to the target block. When the deblocking filter is applied to the target block, the applied filter can vary according to the required strength of the deblocking filtering. In other words, among different filters, the filter determined considering the strength of the deblocking filtering can be applied to the target block.

[0216] SAO can add an appropriate offset to the pixel value to compensate for the coding error. SAO can perform correction on the image to which deblocking has been applied based on pixels, where the correction uses the offset of the difference between the original image and the image to which deblocking has been applied. A method for dividing the pixels included in the image into a specific number of regions, determining the regions to which the offset will be applied among the divided regions, and applying the offset to the determined regions can be used, and a method for applying the offset considering the edge information of each pixel can also be used.

[0217] The ALF can perform filtering based on values obtained by comparing a reconstructed image with an original image. After pixels included in an image have been divided into a predetermined number of groups, a filter to be applied to the groups can be determined, and filtering can be performed differently for each group. Information related to whether to apply an adaptive loop filter can be signaled for each CU. The shape and filter coefficients of the ALF to be applied to each block can be different for each block.

[0218] The reconstructed block or reconstructed image filtered by the filter unit 180 can be stored in the reference picture buffer 190. The reconstructed block filtered by the filter unit 180 can be part of a reference picture. In other words, the reference picture can be a reconstructed picture composed of reconstructed blocks filtered by the filter unit 180. The stored reference picture can then be used for inter prediction.

[0219] Figure 2 is a block diagram showing the configuration of an embodiment of a decoding device to which the present disclosure is applied.

[0220] The decoding device 200 can be a decoder, a video decoding device, or an image decoding device.

[0221] Referring to Figure 2 , the decoding device 200 may include an entropy decoding unit 210, an inverse quantization (dequantization) unit 220, an inverse transform unit 230, an intra prediction unit 240, an inter prediction unit 250, an adder 255, a filter unit 260, and a reference picture buffer 270.

[0222] The decoding device 200 can receive a bitstream output from the encoding device 100. The decoding device 200 can receive a bitstream stored in a computer-readable storage medium, and can receive a bitstream streamed through a wired / wireless transmission medium.

[0223] The decoding device 200 can perform decoding on the bitstream in the intra mode and / or the inter mode. In addition, the decoding device 200 can generate a reconstructed image or a decoded image via decoding, and can output the reconstructed image or the decoded image.

[0224] For example, the operation of switching to the intra mode or the inter mode based on a prediction mode for decoding can be performed by a switcher. When the prediction mode for decoding is the intra mode, the switcher can be operated to switch to the intra mode. When the prediction mode for decoding is the inter mode, the switcher can be operated to switch to the inter mode.

[0225] The decoding device 200 can obtain the reconstructed residual block by decoding the input bitstream and can generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block, which is the target of decoding, by adding the reconstructed residual block and the prediction block.

[0226] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream based on the probability distribution of the bitstream. The generated symbols can include quantized level format symbols. Here, the entropy decoding method can be similar to the entropy encoding method described above. That is to say, the entropy decoding method can be the inverse process of the entropy encoding method described above.

[0227] The quantized coefficients can be dequantized by the dequantization unit 220. The dequantization unit 220 can generate dequantized coefficients by performing dequantization on the quantized coefficients. In addition, the dequantized coefficients can be inverse-transformed by the inverse transform unit 230. The inverse transform unit 230 can generate a reconstructed residual block by performing an inverse transform on the dequantized coefficients. As a result of performing dequantization and inverse transform on the quantized coefficients, a reconstructed residual block can be generated. Here, when generating the reconstructed residual block, the dequantization unit 220 can apply a quantization matrix to the quantized coefficients.

[0228] When using the intra mode, the intra prediction unit 240 can generate a prediction block by performing spatial prediction, where the spatial prediction uses the pixel values of previously decoded neighboring blocks around the target block.

[0229] The inter prediction unit 250 can include a motion compensation unit. Optionally, the inter prediction unit 250 can be designated as the "motion compensation unit".

[0230] When using the inter mode, the motion compensation unit can generate a prediction block by performing motion compensation, where the motion compensation uses a motion vector and a reference image stored in the reference picture buffer 270.

[0231] The motion compensation unit can apply an interpolation filter to a partial region of the reference image when the motion vector has a value other than an integer, and can use the reference image to which the interpolation filter has been applied to generate a prediction block. To perform motion compensation, the motion compensation unit can determine, based on the CU, which of the skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to the motion compensation method for the PU included in the CU, and can perform motion compensation according to the determined mode.

[0232] The reconstructed residual block and the prediction block can be added to each other by the adder 255. The adder 255 can generate a reconstructed block by adding the reconstructed residual block and the prediction block.

[0233] The reconstructed block can be filtered by the filter unit 260. The filter unit 260 can apply at least one of a deblocking filter, an SAO filter, and an ALF to the reconstructed block or the reconstructed picture.

[0234] The reconstructed block filtered by the filter unit 260 can be stored in the reference picture buffer 270. The reconstructed block filtered by the filter unit 260 can be part of a reference picture. In other words, the reference image can be an image composed of the reconstructed blocks filtered by the filter unit 260. The stored reference image can then be used for inter prediction.

[0235] Figure 3 is a diagram schematically showing the partitioning structure of an image when the image is encoded and decoded.

[0236] Figure 3 An example in which a single unit is partitioned into multiple sub-units can be schematically shown.

[0237] To effectively partition an image, coding units (CUs) can be used in encoding and decoding. The term "unit" can be used to commonly specify 1) a block including image samples and 2) syntax elements. For example, "partitioning of a unit" can mean "partitioning of the block corresponding to the unit".

[0238] A CU can be used as a basic unit for image encoding / decoding. A CU can be used as a unit to which one mode selected from an intra mode and an inter mode is applied in image encoding / decoding. In other words, in image encoding / decoding, it can be determined which of the intra mode and the inter mode will be applied to each CU.

[0239] In addition, a CU can be a basic unit for predicting, transforming, quantizing, inverse-transforming, dequantizing, and encoding / decoding transform coefficients.

[0240] Referring to Figure 3 , the image 300 can be sequentially partitioned into units corresponding to the largest coding unit (LCU), and the partitioning structure of the image 300 can be determined according to the LCU. Here, the LCU can be used to have the same meaning as the coding tree unit (CTU).

[0241] Partitioning a unit can mean partitioning the block corresponding to the unit. The block partitioning information can include depth information about the depth of the unit. The depth information can indicate the number of times the unit is partitioned and / or the degree to which the unit is partitioned. A single unit can be hierarchically partitioned into sub-units while the single unit has depth information based on a tree structure. Each partitioned sub-unit can have depth information. The depth information can be information indicating the size of the CU. The depth information can be stored for each CU. Each CU can have depth information.

[0242] The partitioning structure can represent the distribution of coding units (CUs) in the LCU 310 for efficiently encoding an image. Such a distribution can be determined based on whether a single CU will be partitioned into multiple CUs. The number of CUs generated by partitioning can be a positive integer of 2 or greater, including 2, 4, 8, 16, etc. According to the number of CUs generated by partitioning, the horizontal size and vertical size of each CU generated by partitioning can be smaller than the horizontal size and vertical size of the CU before partitioning.

[0243] Each partitioned CU can be recursively partitioned into four CUs in the same way. Compared with at least one of the horizontal size and vertical size of the CU before partitioning, at least one of the horizontal size and vertical size of each partitioned CU can be reduced via recursive partitioning.

[0244] The partitioning of CUs can be recursively executed until a predefined depth or a predefined size. For example, the depth of the LCU can be 0, and the depth of the smallest coding unit (SCU) can be the predefined maximum depth. Here, as described above, the LCU can be a CU with the maximum coding unit size, and the SCU can be a CU with the smallest coding unit size.

[0245] The partitioning can start at the LCU 310, and whenever the horizontal size and / or vertical size of a CU is reduced by partitioning, the depth of the CU can be incremented by 1.

[0246] For example, for each depth, an unpartitioned CU can have a size of 2N×2N. Additionally, in the case where a CU is partitioned, a CU with a size of 2N×2N can be partitioned into four CUs each with a size of N×N. Whenever the depth is incremented by 1, the value of N can be halved.

[0247] Referring to Figure 3 , the LCU with a depth of 0 can have 64×64 pixels or a 64×64 block. 0 can be the minimum depth. The SCU with a depth of 3 can have 8×8 pixels or an 8×8 block. 3 can be the maximum depth. Here, a CU with a 64×64 block as the LCU can be represented by a depth of 0. A CU with a 32×32 block can be represented by a depth of 1. A CU with a 16×16 block can be represented by a depth of 2. A CU with an 8×8 block as the SCU can be represented by a depth of 3.

[0248] Information on whether a corresponding CU is partitioned can be represented by the partitioning information of the CU. The partitioning information can be 1-bit information. All CUs except the SCU can include the partitioning information. For example, the value of the partitioning information of an unpartitioned CU can be 0. The value of the partitioning information of a partitioned CU can be 1.

[0249] For example, when a single CU is partitioned into four CUs, the horizontal size and vertical size of each of the four CUs generated by the partitioning can be half of the horizontal size and vertical size of the CU before partitioning. When a CU with a size of 32×32 is partitioned into four CUs, the size of each of the four partitioned CUs can be 16×16. When a single CU is partitioned into four CUs, it can be considered that the CU has been partitioned in a quadtree structure.

[0250] For example, when a single CU is partitioned into two CUs, the horizontal size or vertical size of each of the two CUs generated by the partitioning can be half of the horizontal size or vertical size of the CU before partitioning. When a CU with a size of 32×32 is vertically partitioned into two CUs, the size of each of the two partitioned CUs can be 16×32. When a single CU is partitioned into two CUs, it can be considered that the CU has been partitioned in a binary tree structure.

[0251] Both quadtree partitioning and binary tree partitioning can be applied to Figure 3 the LCU 310.

[0252] Figure 4 is a diagram showing the forms of prediction units (PUs) that a coding unit (CU) can include.

[0253] Among the CUs partitioned from the LCU, the CUs that are no longer partitioned can be divided into one or more prediction units (PUs). This division is also referred to as "partitioning".

[0254] A PU can be a basic unit for prediction. A PU can be encoded and decoded in any one of the skip mode, inter-frame mode, and intra-frame mode. A PU can be partitioned into various shapes according to each mode. For example, the target blocks described above with reference to Figure 1 and the target blocks described above with reference to Figure 2 can both be PUs.

[0255] In the skip mode, there may be no partitioning in the CU. In the skip mode, the 2N×2N mode 410 can be supported without partitioning, where in the 2N×2N mode 410, the size of the PU and the size of the CU are the same.

[0256] In the inter-frame mode, there may be 8 types of partitioning shapes in the CU. For example, in the inter-frame mode, the 2N×2N mode 410, 2N×N mode 415, N×2N mode 420, N×N mode 425, 2N×nU mode 430, 2N×nD mode 435, nL×2N mode 440, and nR×2N mode 445 can be supported.

[0257] In the intra mode, the 2N×2N mode 410 and the N×N mode 425 are supported.

[0258] In the 2N×2N mode 410, a PU with a size of 2N×2N can be encoded. The PU with a size of 2N×2N can represent a PU having the same size as the CU. For example, the PU with a size of 2N×2N can have a size of 64×64, 32×32, 16×16, or 8×8.

[0259] In the N×N mode 425, a PU with a size of N×N can be encoded.

[0260] For example, in intra prediction, when the size of the PU is 8×8, the PUs divided into four partitions can be encoded. The size of each partitioned PU can be 4×4.

[0261] When encoding a PU in the intra mode, any one of multiple intra prediction modes can be used to encode the PU. For example, the HEVC technology can provide 35 intra prediction modes, and the PU can be encoded under any one of the 35 intra prediction modes.

[0262] It is possible to determine which of the 2N×2N mode 410 and the N×N mode 425 will be used to encode the PU based on the rate - distortion cost.

[0263] The encoding device 100 can perform an encoding operation on a PU with a size of 2N×2N. Here, the encoding operation can be an operation of encoding the PU under each of multiple intra prediction modes that can be used by the encoding device 100. Through the encoding operation, the best intra prediction mode for the PU with a size of 2N×2N can be derived. The best intra prediction mode can be the intra prediction mode that incurs the minimum rate - distortion cost when encoding the PU with a size of 2N×2N among the multiple intra prediction modes that can be used by the encoding device 100.

[0264] In addition, the encoding device 100 can sequentially perform an encoding operation on each PU obtained by performing N×N partitioning. Here, the encoding operation can be an operation of encoding the PU under each of multiple intra prediction modes that can be used by the encoding device 100. Through the encoding operation, the best intra prediction mode for the PU with a size of N×N can be derived. The best intra prediction mode can be the intra prediction mode that incurs the minimum rate - distortion cost when encoding the PU with a size of N×N among the multiple intra prediction modes that can be used by the encoding device 100.

[0265] The encoding device 100 may determine which one of the PUs with a size of 2N×2N and the PUs with a size of N×N will be encoded based on a comparison between the rate-distortion cost of the PUs with a size of 2N×2N and the rate-distortion cost of the PUs with a size of N×N.

[0266] Figure 5 is a diagram showing the form of transform units (TUs) that can be included in a CU.

[0267] The transform unit (TU) may be a basic unit in the CU that is used for processes such as transformation, quantization, inverse transformation, dequantization, entropy coding, and entropy decoding. The TU may have a square or rectangular shape.

[0268] In a CU partitioned from an LCU, a CU that is no longer partitioned into CUs may be partitioned into one or more TUs. Here, the partitioning structure of the TUs may be a quadtree structure. For example, as Figure 5 shown, a single CU 510 may be partitioned one or more times according to the quadtree structure. Through this partitioning, a single CU 510 may be composed of TUs with various sizes.

[0269] In the encoding device 100, a coding tree unit (CTU) with a size of 64×64 may be partitioned into multiple smaller CUs according to a recursive quadtree structure. A single CU may be partitioned into four CUs with the same size. Each CU may be recursively divided and may have a quadtree structure.

[0270] A CU may have a given depth. When a CU is partitioned, the CUs generated by the partitioning may have a depth increased by 1 from the depth of the partitioned CU.

[0271] For example, the depth of a CU may have a value ranging from 0 to 3. According to the depth of the CU, the size range of the CU may range from a size of 64×64 to a size of 8×8.

[0272] Through the recursive partitioning of the CU, the best partitioning method that generates the minimum rate-distortion cost may be selected.

[0273] Figure 6 is a diagram for explaining an embodiment of the intra prediction process.

[0274] From Figure 6 the arrows radially extending from the center of the diagram in represent the prediction directions of the intra prediction modes. In addition, the numbers appearing near the arrows may represent examples of the mode values assigned to the intra prediction modes or the prediction directions of the intra prediction modes.

[0275] Reference sample points of blocks adjacent to a target block may be used to perform intra coding and / or decoding. The adjacent blocks may be adjacent reconstructed blocks. For example, intra coding and / or decoding may be performed using values of reference sample points included in each adjacent reconstructed block or coding parameters of the adjacent reconstructed blocks.

[0276] Encoding device 100 and / or decoding device 200 may generate a prediction block by performing intra prediction on a target block based on information about sample points in a target image. When intra prediction is performed, encoding device 100 and / or decoding device 200 may generate a prediction block for the target block by performing intra prediction based on information about sample points in the target image. When intra prediction is performed, encoding device 100 and / or decoding device 200 may perform directional prediction and / or non-directional prediction based on at least one reconstructed reference sample point.

[0277] The prediction block may be a block generated as a result of performing intra prediction. The prediction block may correspond to at least one of a CU, a PU, and a TU.

[0278] The unit of the prediction block may have a size corresponding to at least one of a CU, a PU, and a TU. The prediction block may have a square shape with a size of 2N×2N or N×N. The size N×N may include sizes such as 4×4, 8×8, 16×16, 32×32, 64×64, etc.

[0279] Optionally, the prediction block may be a square block with a size of 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, etc. or a rectangular block with a size of 2×8, 4×8, 2×16, 4×16, 8×16, etc.

[0280] Intra prediction may be performed considering an intra prediction mode for a target block. The number of intra prediction modes that the target block may have may be a predefined fixed value and may be a value determined differently according to attributes of the prediction block. For example, attributes of the prediction block may include the size of the prediction block, the type of the prediction block, etc.

[0281] For example, regardless of the size of the prediction block, the number of intra prediction modes may be fixed to 35. Optionally, the number of intra prediction modes may be, for example, 3, 5, 9, 17, 34, 35, or 36.

[0282] The intra prediction mode may be a non-directional mode or a directional mode. For example, as Figure 6 shown, the intra prediction mode may include two non-directional modes and 33 directional modes.

[0283] The two non-directional modes may include a DC mode and a planar mode.

[0284] The directional mode may be a mode with a specific direction or a specific angle.

[0285] Each available mode number, mode value, and mode angle in the intra prediction mode represents at least one of them. The number of intra prediction modes may be M. The value of M may be 1 or greater. In other words, the number of intra prediction modes may be M, where M includes the number of non-directional modes and the number of directional modes.

[0286] The number of intra prediction modes may be fixed to M regardless of the size of the block. Optionally, the number of intra prediction modes may vary according to the size of the block and / or the type of color component. For example, the number of prediction modes may vary according to whether the color component is a luminance signal or a chrominance signal. For example, the larger the size of the block, the larger the number of intra prediction modes. Optionally, the number of intra prediction modes corresponding to the luminance component block may be greater than the number of intra prediction modes corresponding to the chrominance component block.

[0287] For example, in the vertical mode with a mode value of 26, prediction may be performed along the vertical direction based on the pixel values of the reference samples. For example, in the horizontal mode with a mode value of 10, prediction may be performed along the horizontal direction based on the pixel values of the reference samples.

[0288] Even in the directional modes other than the above-mentioned modes, the encoding device 100 and the decoding device 200 may still perform intra prediction on the target unit using the reference samples according to the angle corresponding to the directional mode.

[0289] The intra prediction mode located to the right of the vertical mode may be referred to as the "vertical - right mode". The intra prediction mode located below the horizontal mode may be referred to as the "horizontal - below mode". For example, in Figure 6 Among them, the intra prediction mode with a mode value being one of 27, 28, 29, 30, 31, 32, 33, and 34 may be the vertical - right mode 613. The intra prediction mode with a mode value being one of 2, 3, 4, 5, 6, 7, 8, and 9 may be the horizontal - below mode 616.

[0290] The non - directional modes may include the DC mode and the plane mode. For example, the mode value of the DC mode may be 1. The mode value of the plane mode may be 0.

[0291] The directional modes may include the angle mode. Among the multiple intra prediction modes, the remaining modes except the DC mode and the plane mode may be the directional modes.

[0292] When the intra prediction mode is the DC mode, a prediction block may be generated based on the average value of the pixel values of multiple reference pixels. For example, the pixel values of the prediction block may be determined based on the average value of the pixel values of multiple reference pixels.

[0293] The number of the intra prediction modes described above and the mode values of the respective intra prediction modes are merely exemplary. The number of the intra prediction modes described above and the mode values of the respective intra prediction modes may be defined differently according to embodiments, implementations, and / or requirements.

[0294] To perform intra prediction on a target block, a step of checking whether samples included in reconstructed neighboring blocks can be used as reference samples for the target block may be performed. When there are samples among the samples in the neighboring blocks that cannot be used as reference samples for the target block, values generated by interpolation and / or copying of at least one of the sample values among the samples included in the reconstructed neighboring blocks may replace the sample values of the samples that cannot be used as reference samples. When the values generated by copying and / or interpolation replace the sample values of existing samples, the samples may be used as reference samples for the target block.

[0295] In intra prediction, a filter may be applied to at least one of reference samples and prediction samples based on at least one of an intra prediction mode and the size of a target block.

[0296] When the intra prediction mode is a planar mode, when generating a prediction block of a target block, sample values of the prediction target block may be generated using a weighted sum of an upper reference sample of the target block, a left reference sample of the target block, an upper right reference sample of the target block, and a lower left reference sample of the target block according to the position of the prediction target sample in the prediction block.

[0297] When the intra prediction mode is a DC mode, when generating a prediction block of a target block, an average value of a reference sample above the target block and a reference sample to the left of the target block may be used.

[0298] When the intra prediction mode is a direction mode, an upper reference sample, a left reference sample, an upper right reference sample, and / or a lower left reference sample of the target block may be used to generate a prediction block.

[0299] To generate the above prediction samples, interpolation based on real numbers may be performed.

[0300] The intra prediction mode of a target block may perform prediction from the intra prediction of neighboring blocks adjacent to the target block, and information used for prediction may be entropy encoded / decoded.

[0301] For example, when the intra prediction modes of a target block and a neighboring block are the same, a predefined flag may be used to signal that the intra prediction modes of the target block and the neighboring block are the same.

[0302] For example, an indicator for indicating an intra prediction mode that is the same as the intra prediction mode of a target block among the intra prediction modes of a plurality of neighboring blocks may be signaled.

[0303] When the intra prediction modes of the target block and the neighboring blocks are different from each other, the intra prediction mode information of the target block can be entropy-coded / decoded based on the intra prediction mode of the neighboring block.

[0304] Figure 7 It is a diagram for explaining the positions of the reference samples used in the intra prediction process.

[0305] Figure 7 It shows the positions of the reference samples for intra prediction of the target block. Refer to Figure 7 , the reconstructed reference samples for intra prediction of the target block may include a lower left reference sample 731, a left reference sample 733, an upper left reference sample 735, an upper reference sample 737, and an upper right reference sample 739.

[0306] For example, the left reference sample 733 may represent a reconstructed reference pixel adjacent to the left side of the target block. The upper reference sample 737 may represent a reconstructed reference pixel adjacent to the top of the target block. The upper left reference sample 735 may represent a reconstructed reference pixel located at the upper left corner of the target block. The lower left reference sample 731 may represent a reference sample located below the left side sample line among the samples on the same line as the left side sample line composed of the left reference sample 733. The upper right reference sample 739 may represent a reference sample located to the right of the upper side sample line among the samples on the same line as the upper side sample line composed of the upper reference sample 737.

[0307] When the size of the target block is N×N, the numbers of the lower left reference sample 731, the left reference sample 733, the upper reference sample 737, and the upper right reference sample 739 may all be N.

[0308] By performing intra prediction on the target block, a prediction block can be generated. The process of generating the prediction block may include determining the values of the pixels in the prediction block. The sizes of the target block and the prediction block may be the same.

[0309] The reference samples for intra prediction of the target block may change according to the intra prediction mode of the target block. The direction of the intra prediction mode may represent the dependency relationship between the reference sample and the pixels in the prediction block. For example, the value of the specified reference sample may be used as the value of one or more specified pixels in the prediction block. In this case, the specified reference sample and the one or more specified pixels in the prediction block may be samples and pixels located on a straight line along the direction of the intra prediction mode. In other words, the value of the specified reference sample may be copied as the value of the pixels located in the direction opposite to the direction of the intra prediction mode. Optionally, the value of the pixel in the prediction block may be the value of the reference sample located in the direction of the intra prediction mode with respect to the position of the pixel.

[0310] In the example, when the intra prediction mode of the target block is the vertical mode with a mode value of 26, the upper reference sample 737 can be used for intra prediction. When the intra prediction mode is the vertical mode, the value of a pixel in the prediction block can be the value of the reference sample vertically located above the position of the pixel. Therefore, the upper reference sample 737 adjacent to the top of the target block can be used for intra prediction. In addition, the values of the pixels in a row of the prediction block can be the same as the values of the pixels of the upper reference sample 737.

[0311] In the example, when the intra prediction mode of the target block is the horizontal mode with a mode value of 10, the left reference sample 733 can be used for intra prediction. When the intra prediction mode is the horizontal mode, the value of a pixel in the prediction block can be the value of the reference sample horizontally located to the left of the position of the pixel. Therefore, the left reference sample 733 adjacent to the left side of the target block can be used for intra prediction. In addition, the values of the pixels in a column of the prediction block can be the same as the values of the pixels of the left reference sample 733.

[0312] In the example, when the mode value of the intra prediction mode of the current block is 18, at least some of the left reference sample 733, the upper left reference sample 735, and at least some of the upper reference sample 737 can be used for intra prediction. When the mode value of the intra prediction mode is 18, the value of a pixel in the prediction block can be the value of the reference sample diagonally located at the upper left corner of the pixel.

[0313] In addition, when an intra prediction mode with a mode value of 27, 28, 29, 30, 31, 32, 33, or 34 is used, at least a part of the upper right reference sample 739 can be used for intra prediction.

[0314] In addition, when an intra prediction mode with a mode value of 2, 3, 4, 5, 6, 7, 8, or 9 is used, at least a part of the lower left reference sample 731 can be used for intra prediction.

[0315] In addition, in the case of an intra prediction mode with a mode value in the range from 11 to 25, the upper left reference sample 735 can be used for intra prediction.

[0316] The number of reference samples used to determine the pixel value of a pixel in the prediction block can be 1 or 2 or more.

[0317] As described above, the pixel value of a pixel in the prediction block can be determined according to the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode. When the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode are integer positions, the value of one reference sample indicated by the integer position can be used to determine the pixel value of the pixel in the prediction block.

[0318] When the position of a pixel and the position of a reference sample indicated by the direction of an intra prediction mode are not integer positions, an interpolated reference sample can be generated based on two reference samples closest to the position of the reference sample. The value of the interpolated reference sample can be used to determine the pixel value of a pixel in a prediction block. In other words, when the position of a pixel in a prediction block and the position of a reference sample indicated by the direction of an intra prediction mode indicate a position between two reference samples, an interpolated value based on the values of these two samples can be generated.

[0319] A prediction block generated through prediction can be different from an original target block. In other words, there may be a prediction error, which is the difference between the target block and the prediction block, and there may also be a prediction error between the pixels of the target block and the pixels of the prediction block.

[0320] Hereinafter, the terms "difference", "error", and "residual" can be used to have the same meaning and can be used interchangeably with each other.

[0321] For example, in the case of directional intra prediction, the longer the distance between the pixels of a prediction block and a reference sample, the greater the possible prediction error. Such a prediction error can lead to discontinuity between the generated prediction block and neighboring blocks.

[0322] To reduce the prediction error, a filtering operation for a prediction block can be used. The filtering operation can be configured to adaptively apply a filter to a region in the prediction block that is considered to have a relatively large prediction error. For example, a region considered to have a relatively large prediction error can be the boundary of the prediction block. In addition, the region in the prediction block considered to have a relatively large prediction error can vary according to the intra prediction mode, and the characteristics of the filter can also vary according to the intra prediction mode.

[0323] Figure 8 is a diagram for explaining an embodiment of an inter prediction process.

[0324] Figure 8 The rectangle shown in can represent an image (or picture). In addition, in Figure 8 The arrow can represent a prediction direction. That is, each image can be encoded and / or decoded according to the prediction direction.

[0325] An image can be classified into an intra picture (I picture), a unidirectional prediction picture or a predictive coded picture (P picture), and a bidirectional prediction picture or a bidirectional predictive coded picture (B picture) according to the coding type. Each picture can be encoded according to the coding type of each picture.

[0326] When the target image to be encoded is an I picture, the target image can be encoded using the data contained in the image itself without performing inter-frame prediction with reference to other images. For example, an I picture can be encoded only via intra-frame prediction.

[0327] When the target image is a P picture, the target image can be encoded via inter-frame prediction using a reference picture existing in one direction. Here, the one direction can be the forward direction or the backward direction.

[0328] When the target image is a B picture, the image can be encoded via inter-frame prediction using reference pictures existing in two directions, or can be encoded via inter-frame prediction using a reference picture existing in one of the forward direction and the backward direction. Here, the two directions can be the forward direction and the backward direction.

[0329] P pictures and B pictures encoded and / or decoded using reference pictures can be regarded as images using inter-frame prediction.

[0330] Hereinafter, inter-frame prediction in the inter-frame mode according to an embodiment will be described in detail.

[0331] Motion information can be used to perform inter-frame prediction.

[0332] In the inter-frame mode, the encoding device 100 can perform inter-frame prediction and / or motion compensation on a target block. The decoding device 200 can perform inter-frame prediction and / or motion compensation corresponding to the inter-frame prediction and / or motion compensation performed by the encoding device 100 on the target block.

[0333] The motion information of the target block can be separately derived by the encoding device 100 and the decoding device 200 during inter-frame prediction. The motion information of the reconstructed neighboring block, the motion information of the col block, and / or the motion information of the block adjacent to the col block can be used to derive the motion information.

[0334] For example, the encoding device 100 or the decoding device 200 can perform prediction and / or motion compensation by using the motion information of the spatial candidate and / or the temporal candidate as the motion information of the target block. The target block can represent a PU and / or a PU partition.

[0335] The spatial candidate can be a reconstructed block spatially adjacent to the target block.

[0336] The temporal candidate can be a reconstructed block corresponding to the target block in a previously reconstructed co-located picture (col picture).

[0337] In inter - frame prediction, the encoding device 100 and the decoding device 200 can improve the encoding efficiency and decoding efficiency by using the motion information of spatial candidates and / or temporal candidates. The motion information of spatial candidates can be referred to as "spatial motion information". The motion information of temporal candidates can be referred to as "temporal motion information".

[0338] Hereinafter, the motion information of spatial candidates can be the motion information of a PU including the spatial candidates. The motion information of temporal candidates can be the motion information of a PU including the temporal candidates. The motion information of a candidate block can be the motion information of a PU including the candidate block.

[0339] Inter - frame prediction can be performed using reference pictures.

[0340] The reference picture can be at least one of a picture before the target picture and a picture after the target picture. The reference picture can be an image for predicting the target block.

[0341] In inter - frame prediction, a reference picture index (or refIdx) for indicating a reference picture, a motion vector to be described later, etc. can be used to specify a region in the reference picture. Here, the region specified in the reference picture can indicate a reference block.

[0342] Inter - frame prediction can select a reference picture and can also select a reference block corresponding to the target block from the reference picture. In addition, inter - frame prediction can use the selected reference block to generate a prediction block for the target block.

[0343] Each of the encoding device 100 and the decoding device 200 can derive motion information during inter - frame prediction.

[0344] Spatial candidates can be 1) blocks that exist in the target picture, 2) have been previously reconstructed via encoding and / or decoding, and 3) are adjacent to the target block or are located at the corners of the target block. Here, "a block located at the corner of the target block" can be a block that is vertically adjacent to a neighboring block horizontally adjacent to the target block, or a block that is horizontally adjacent to a neighboring block vertically adjacent to the target block. In addition, "a block located at the corner of the target block" can have the same meaning as "a block adjacent to the corner of the target block". The meaning of "a block located at the corner of the target block" can be included in the meaning of "a block adjacent to the target block".

[0345] For example, spatial candidates can be a reconstructed block located to the left of the target block, a reconstructed block located above the target block, a reconstructed block located at the lower - left corner of the target block, a reconstructed block located at the upper - right corner of the target block, or the target block located at the upper - left corner of the target block.

[0346] Each of the encoding device 100 and the decoding device 200 can identify a block that exists in the col picture and is spatially corresponding to the target block. The position of the target block in the target picture and the position of the identified block in the col picture can correspond to each other.

[0347] Each of the encoding device 100 and the decoding device 200 can determine a col block existing at a predefined relative position for the identified block as a temporal candidate. The predefined relative position can be a position existing inside and / or outside the identified block.

[0348] For example, the col block can include a first col block and a second col block. When the coordinates of the identified block are (xP, yP) and the size of the identified block is represented by (nPSW, nPSH), the first col block can be the block located at the coordinates (xP + nPSW, yP + nPSH). The second col block can be the block located at the coordinates (xP + (nPSW >> 1), yP + (nPSH >> 1)). When the first col block is not available, the second col block can be selectively used.

[0349] The motion vector of the target block can be determined based on the motion vector of the col block. Each of the encoding device 100 and the decoding device 200 can scale the motion vector of the col block. The scaled motion vector of the col block can be used as the motion vector of the target block. In addition, the motion vector of the running information of the temporal candidate stored in the list can be the scaled motion vector.

[0350] The ratio of the motion vector of the target block to the motion vector of the col block can be the same as the ratio of the first distance to the second distance. The first distance can be the distance between the reference picture and the target picture of the target block. The second distance can be the distance between the reference picture and the col picture of the col block.

[0351] The scheme for deriving motion information can be changed according to the inter-frame prediction mode of the target block. For example, as the inter-frame prediction mode applied to inter-frame prediction, there can be an Advanced Motion Vector Prediction (AMVP) mode, a merge mode, a skip mode, a current picture reference mode, etc. The merge mode can also be referred to as the "motion merge mode". Each mode will be described in detail below.

[0352] 1) AMVP Mode

[0353] When using the AMVP mode, the encoding device 100 can search for a similar block in the neighboring area of the target block. The encoding device 100 can perform prediction on the target block by using the motion information of the found similar block to obtain a predicted block. The encoding device 100 can encode the residual block that is the difference between the target block and the predicted block.

[0354] 1-1) Create a list of predicted motion vector candidates

[0355] When the AMVP mode is used as a prediction mode, each of the encoding device 100 and the decoding device 200 may create a list of predicted motion vector candidates by using the motion vectors of spatial candidates, the motion vectors of temporal candidates, and the zero vector. The list of predicted motion vector candidates may include one or more predicted motion vector candidates. At least one of the motion vectors of spatial candidates, the motion vectors of temporal candidates, and the zero vector may be determined and used as a predicted motion vector candidate.

[0356] Hereinafter, the terms "predicted motion vector (candidate)" and "motion vector (candidate)" may be used with the same meaning and may be used interchangeably with each other.

[0357] Hereinafter, the terms "predicted motion vector candidate" and "AMVP candidate" may be used with the same meaning and may be used interchangeably with each other.

[0358] Hereinafter, the terms "list of predicted motion vector candidates" and "list of AMVP candidates" may be used with the same meaning and may be used interchangeably with each other.

[0359] Spatial motion candidates may include reconstructed spatially neighboring blocks. In other words, the motion vectors of the reconstructed neighboring blocks may be referred to as "spatial predicted motion vector candidates".

[0360] Temporal motion candidates may include col blocks and blocks adjacent to the col blocks. In other words, the motion vectors of the col blocks or the motion vectors of the blocks adjacent to the col blocks may be referred to as "temporal predicted motion vector candidates".

[0361] The zero vector may be a (0, 0) motion vector.

[0362] Predicted motion vector candidates may be motion vector predictors for predicting motion vectors. In addition, in the encoding device 100, each predicted motion vector candidate may be an initial search position for a motion vector.

[0363] 1-2) Search for motion vectors using the list of predicted motion vector candidates

[0364] The encoding device 100 may determine, within a search range, the motion vector to be used for encoding a target block by using the list of predicted motion vector candidates. In addition, the encoding device 100 may determine, among the predicted motion vector candidates present in the list of predicted motion vector candidates, the predicted motion vector candidate to be used as the predicted motion vector of the target block.

[0365] The motion vector to be used for encoding a target block may be the motion vector that can be encoded at the minimum cost.

[0366] In addition, the encoding device 100 can determine whether to use the AMVP mode to encode the target block.

[0367] 1-3) Transmission of inter-frame prediction information

[0368] The encoding device 100 can generate a bitstream including the inter-frame prediction information required for inter-frame prediction. The decoding device 200 can perform inter-frame prediction on the target block using the inter-frame prediction information of the bitstream.

[0369] The inter-frame prediction information may include 1) mode information indicating whether AMVP is used, 2) a predicted motion vector index, 3) a motion vector difference (MVD), 4) a reference direction, and 5) a reference picture index.

[0370] Hereinafter, the terms "predicted motion vector index" and "AMVP index" may be used with the same meaning and may be used interchangeably with each other. In addition, the inter-frame prediction information may include a residual signal.

[0371] When the mode information indicates that the AMVP mode is used, the decoding device 200 can obtain the predicted motion vector index, MVD, reference direction, and reference picture index from the bitstream through entropy decoding.

[0372] The predicted motion vector index can indicate the predicted motion vector candidate among the predicted motion vector candidates included in the predicted motion vector candidate list that will be used to predict the target block.

[0373] 1-4) Inter-frame prediction in AMVP mode using inter-frame prediction information

[0374] The decoding device 200 can use the predicted motion vector candidate list to derive the predicted motion vector candidate, and can determine the motion information of the target block based on the derived predicted motion vector candidate.

[0375] The decoding device 200 can use the predicted motion vector index to determine the motion vector candidate for the target block among the predicted motion vector candidates included in the predicted motion vector candidate list. The decoding device 200 can select the predicted motion vector candidate indicated by the predicted motion vector index as the predicted motion vector of the target block from among the predicted motion vector candidates included in the predicted motion vector candidate list.

[0376] The motion vector that will actually be used for inter-frame prediction of the target block may not match the predicted motion vector. In order to indicate the difference between the motion vector that will actually be used for inter-frame prediction of the target block and the predicted motion vector, the MVD can be used. The encoding device 100 can derive a predicted motion vector similar to the motion vector that will actually be used for inter-frame prediction of the target block so as to use the smallest possible MVD.

[0377] The MVD may be the difference between the motion vector of the target block and the predicted motion vector. The encoding device 100 may calculate the MVD and may perform entropy coding on the MVD.

[0378] The MVD may be sent from the encoding device 100 to the decoding device 200 via a bitstream. The decoding device 200 may decode the received MVD. The decoding device 200 may derive the motion vector of the target block by summing the decoded MVD and the predicted motion vector. In other words, the motion vector of the target block derived by the decoding device 200 may be the sum of the entropy-decoded MVD and the motion vector candidate.

[0379] The reference direction may indicate a list of reference pictures to be used for predicting the target block. For example, the reference direction may indicate one of the reference picture list L0 and the reference picture list L1.

[0380] The reference direction only indicates the list of reference pictures to be used for predicting the target block and does not necessarily mean that the direction of the reference picture is limited to the forward direction or the backward direction. In other words, each of the reference picture list L0 and the reference picture list L1 may include pictures in the forward direction and / or the backward direction.

[0381] The reference direction being unidirectional may mean using a single reference picture list. The reference direction being bidirectional may mean using two reference picture lists. In other words, the reference direction may indicate one of the following cases: the case of using only the reference picture list L0, the case of using only the reference picture list L1, and the case of using two reference picture lists.

[0382] The reference picture index may indicate the reference picture among the reference pictures in the reference picture list to be used for predicting the target block. The encoding device 100 may perform entropy coding on the reference picture index. The entropy-coded reference picture index may be signaled by the encoding device 100 to the decoding device 200 via a bitstream.

[0383] When two reference picture lists are used for predicting the target block, a single reference picture index and a single motion vector may be used for each of the reference picture lists. In addition, when two reference picture lists are used for predicting the target block, two prediction blocks may be specified for the target block. For example, the average value or weighted sum of the two prediction blocks for the target block may be used to generate the (final) prediction block of the target block.

[0384] The motion vector of the target block may be derived by the prediction motion vector index, the MVD, the reference direction, and the reference picture index.

[0385] The decoding device 200 can generate a predicted block for a target block based on the derived motion vector and reference picture index. For example, the predicted block can be a reference block indicated by the derived motion vector in the reference picture indicated by the reference picture index.

[0386] Since the predicted motion vector index and MVD are encoded while the motion vector of the target block itself is not encoded, the number of bits sent from the encoding device 100 to the decoding device 200 can be reduced, and the encoding efficiency can be improved.

[0387] The motion information of the reconstructed neighboring blocks can be used for the target block. In a specific inter prediction mode, the encoding device 100 may not encode the actual motion information of the target block separately. Instead of encoding the motion information of the target block, additional information can be encoded, where the additional information enables the derivation of the motion information of the target block using the motion information of the reconstructed neighboring blocks. Since the additional information is encoded, the number of bits sent to the decoding device 200 can be reduced, and the encoding efficiency can be improved.

[0388] For example, as inter prediction modes in which the motion information of the target block is not directly encoded, there may be a skip mode and / or a merge mode. Here, each of the encoding device 100 and the decoding device 200 can use an indicator and / or an index of a unit indicating that its motion information among the reconstructed neighboring units will be used as the motion information of the target unit.

[0389] 2) Merge Mode

[0390] As a scheme for deriving the motion information of the target block, there is merge. The term "merge" can mean merging the motions of multiple blocks. "Merge" can also mean that the motion information of one block is also applied to other blocks. In other words, the merge mode can be a mode of deriving the motion information of the target block from the motion information of neighboring blocks.

[0391] When using the merge mode, the encoding device 100 can use the motion information of spatial candidates and / or temporal candidates to predict the motion information of the target block. Spatial candidates can include reconstructed spatial neighboring blocks that are spatially adjacent to the target block. Spatially adjacent blocks can include the left adjacent block and the upper adjacent block. Temporal candidates can include col blocks. The terms "spatial candidate" and "spatial merge candidate" can be used with the same meaning and can be used interchangeably with each other. The terms "temporal candidate" and "temporal merge candidate" can be used with the same meaning and can be used interchangeably with each other.

[0392] The encoding device 100 can obtain a predicted block through prediction. The encoding device 100 can encode the residual block that is the difference between the target block and the predicted block.

[0393] 2-1) Create a merge candidate list

[0394] When using the merge mode, each of the encoding device 100 and the decoding device 200 can create a merge candidate list using the motion information of the spatial candidate and / or the motion information of the temporal candidate. The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction can be unidirectional or bidirectional.

[0395] The merge candidate list may include merge candidates. A merge candidate can be motion information. In other words, the merge candidate list can be a list storing multiple pieces of motion information.

[0396] A merge candidate can be the motion information of multiple temporal candidates and / or spatial candidates. In addition, the merge candidate list may include new merge candidates generated by combining the merge candidates already existing in the merge candidate list. In other words, the merge candidate list may include new motion information generated by combining multiple pieces of motion information previously existing in the merge candidate list.

[0397] A merge candidate can be a specific mode for deriving inter-frame prediction information. A merge candidate can be information indicating a specific mode for deriving inter-frame prediction information. The inter-frame prediction information of the target block can be derived according to the specific mode indicated by the merge candidate. In addition, the specific mode may include a process for deriving a series of inter-frame prediction information. Such a specific mode can be an inter-frame prediction information derivation mode or a motion information derivation mode.

[0398] The inter-frame prediction information of the target block can be derived according to the mode indicated by the merge candidate selected from the merge candidates in the merge candidate list through a merge index.

[0399] For example, the motion information derivation mode in the merge candidate list can be at least one of the following modes: 1) a motion information derivation mode for a sub-block unit; 2) an affine motion information derivation mode. In addition, the merge candidate list may include the motion information of a zero vector. The zero vector may also be referred to as a "zero merge candidate".

[0400] In other words, multiple pieces of motion information in the merge candidate list can be at least one of the following information: 1) the motion information of a spatial candidate, 2) the motion information of a temporal candidate, 3) the motion information generated by combining multiple pieces of motion information previously existing in the merge candidate list, and 4) a zero vector.

[0401] The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may also be referred to as an "inter-frame prediction indicator". The reference direction can be unidirectional or bidirectional. The unidirectional reference direction may indicate L0 prediction or L1 prediction.

[0402] A merge candidate list may be created before performing prediction in the merge mode.

[0403] The number of merge candidates in the merge candidate list may be predefined. Each of the encoding device 100 and the decoding device 200 may add merge candidates to the merge candidate list according to a predefined scheme and a predefined priority, such that the merge candidate list has a predefined number of merge candidates. The merge candidate list of the encoding device 100 and the merge candidate list of the decoding device 200 may be made the same as each other using the predefined scheme and the predefined priority.

[0404] Merging may be applied based on a CU or a PU. When performing merging based on a CU or a PU, the encoding device 100 may send a bitstream including predefined information to the decoding device 200. For example, the predefined information may include 1) information indicating whether merging is performed for each block partition, and 2) information about the block among the blocks that are spatial candidates and / or temporal candidates for the target block and for which merging is to be performed.

[0405] 2-2) Search for motion vectors using the merge candidate list

[0406] The encoding device 100 may determine a merge candidate to be used for encoding a target block. For example, the encoding device 100 may perform prediction on the target block using the merge candidates in the merge candidate list, and may generate a residual block for the merge candidate. The encoding device 100 may use the merge candidate that results in the minimum cost in the prediction and the encoding of the residual block to encode the target block.

[0407] In addition, the encoding device 100 may determine whether to encode the target block using the merge mode.

[0408] 2-3) Transmission of inter-frame prediction information

[0409] The encoding device 100 may generate a bitstream including inter prediction information required for inter prediction. The encoding device 100 may generate entropy-coded inter prediction information by performing entropy coding on the inter prediction information, and may send the bitstream including the entropy-coded inter prediction information to the decoding device 200. The entropy-coded inter prediction information may be signaled by the encoding device 100 to the decoding device 200 through the bitstream.

[0410] The decoding device 200 may perform inter prediction on the target block using the inter prediction information of the bitstream.

[0411] The inter prediction information may include 1) mode information indicating whether the merge mode is used and 2) a merge index.

[0412] In addition, the inter prediction information may include a residual signal.

[0413] The decoding device 200 may obtain the merge index from the bitstream only when the mode information indicates that the merge mode is used.

[0414] The mode information may be a merge flag. The unit of the mode information may be a block. Information about the block may include the mode information, and the mode information may indicate whether the merge mode is applied to the block.

[0415] The merge index may indicate the merge candidate among the merge candidates included in the merge candidate list that will be used to predict the target block. Optionally, the merge index may indicate the block among the neighboring blocks that are spatially or temporally adjacent to the target block and that will be merged with the target block.

[0416] The encoding device 100 may select the merge candidate with the highest encoding performance from the merge candidates included in the merge candidate list, and may set the value of the merge index such that the merge index indicates the selected merge candidate.

[0417] 2-4) Inter-frame prediction in merge mode using inter-frame prediction information

[0418] The decoding device 200 may perform prediction on the target block using the merge candidate indicated by the merge index among the merge candidates included in the merge candidate list.

[0419] The motion vector of the target block may be specified by the motion vector, reference picture index, and reference direction of the merge candidate indicated by the merge index.

[0420] 3) Skip Mode

[0421] The skip mode may be a mode that applies the motion information of a spatial candidate or a temporal candidate to the target block without change. In addition, the skip mode may be a mode that does not use the residual signal. In other words, when the skip mode is used, the reconstructed block may be a predicted block.

[0422] The difference between the merge mode and the skip mode lies in whether the residual signal is sent or used. That is, the skip mode may be similar to the merge mode except that the residual signal is not sent or used.

[0423] When the skip mode is used, the encoding device 100 may send information related to the block whose motion information will be used as the motion information of the target block among the blocks as spatial candidates or temporal candidates to the decoding device 200 through the bitstream. The encoding device 100 may generate entropy-encoded information by performing entropy encoding on this information, and may signal the entropy-encoded information to the decoding device 200 through the bitstream.

[0424] In addition, when using the skip mode, the encoding device 100 may not send other syntax information (such as MVD) to the decoding device 200. For example, when using the skip mode, the encoding device 100 may not signal to the decoding device 200 the syntax elements related to at least one of MVC, coding block flag, and transform coefficient level.

[0425] 3-1) Create a merge candidate list

[0426] The skip mode may also use a merge candidate list. In other words, the merge candidate list may be used in both the merge mode and the skip mode. In this regard, the merge candidate list may also be referred to as a "skip candidate list" or a "merge / skip candidate list".

[0427] Optionally, the skip mode may use an additional candidate list different from the candidate list of the merge mode. In this case, in the following description, the merge candidate list and the merge candidate may be replaced with the skip candidate list and the skip candidate, respectively.

[0428] The merge candidate list may be created before performing prediction in the skip mode.

[0429] 3-2) Search for motion vectors using the merge candidate list

[0430] The encoding device 100 may determine the merge candidate to be used for encoding the target block. For example, the encoding device 100 may perform prediction on the target block using the merge candidate in the merge candidate list. The encoding device 100 may encode the target block using the merge candidate that produces the minimum cost in the prediction.

[0431] In addition, the encoding device 100 may determine whether to use the skip mode to encode the target block.

[0432] 3-3) Transmission of inter-frame prediction information

[0433] The encoding device 100 may generate a bitstream including the inter-frame prediction information required for inter-frame prediction. The decoding device 200 may perform inter-frame prediction on the target block using the inter-frame prediction information of the bitstream.

[0434] The inter-frame prediction information may include 1) mode information indicating whether the skip mode is used and 2) a skip index.

[0435] The skip index may be the same as the merge index described above.

[0436] When using the skip mode, the target block may be encoded without using a residual signal. The inter-frame prediction information may not include a residual signal. Optionally, the bitstream may not include a residual signal.

[0437] The decoding device 200 may obtain a skip index from the bitstream only when the mode information indicates that the skip mode is used. As described above, the merge index and the skip index may be the same as each other. The decoding device 200 may obtain a skip index from the bitstream only when the mode information indicates that the merge mode or the skip mode is used.

[0438] The skip index may indicate a merge candidate among the merge candidates included in the merge candidate list that will be used to perform prediction on the target block.

[0439] 3-4) Inter-frame prediction in skip mode using inter-frame prediction information

[0440] The decoding device 200 may perform prediction on the target block using the merge candidate indicated by the skip index among the merge candidates included in the merge candidate list.

[0441] The motion vector of the target block may be specified by the motion vector, reference picture index, and reference direction of the merge candidate indicated by the skip index.

[0442] 4) Current Picture Reference Mode

[0443] The current picture reference mode may represent a prediction mode that uses a previously reconstructed region in the target picture to which the target block belongs.

[0444] A vector for specifying the previously reconstructed region may be defined. The reference picture index of the target block may be used to determine whether the target block has been encoded in the current picture reference mode.

[0445] A flag or index indicating whether the target block is a block encoded in the current picture reference mode may be signaled from the encoding device 100 to the decoding device 200. Optionally, it may be inferred whether the target block is a block encoded in the current picture reference mode by the reference picture index of the target block.

[0446] When the target block is encoded in the current picture reference mode, the current picture may be added to a fixed position or an arbitrary position in the reference picture list for the target block.

[0447] For example, the fixed position may be the position where the reference picture index is 0 or the last position.

[0448] When the target picture is added to an arbitrary position in the reference picture list, an additional reference picture index indicating such an arbitrary position may be signaled from the encoding device 100 to the decoding device 200.

[0449] In the AMVP mode, merge mode, and skip mode described above, an index of a list may be used to specify the motion information among multiple pieces of motion information in the list that will be used to perform prediction on the target block.

[0450] To improve the coding efficiency, the coding device 100 may transmit only the index of the element that generates the minimum cost in the inter prediction of the target block among the elements in the list. The coding device 100 may encode the index and transmit the encoded index as a signal.

[0451] Therefore, it must be possible for the coding device 100 and the decoding device 200 to derive the above-described lists (i.e., the predicted motion vector candidate list and the merge candidate list) based on the same data using the same scheme. Here, the same data may include the reconstructed picture and the reconstructed block. In addition, in order to specify an element using an index, the order of the elements in the list must be fixed.

[0452] Figure 9 Shows spatial candidates according to an embodiment.

[0453] In Figure 9 the positions of the spatial candidates are shown.

[0454] The large block at the center of the figure may represent the target block. The five small blocks may represent spatial candidates.

[0455] The coordinates of the target block may be (xP, yP), and the size of the target block may be represented by (nPSW, nPSH).

[0456] Spatial candidate A0 may be a block adjacent to the lower left corner of the target block. A0 may be a block occupying the pixels located at coordinates (xP - 1, yP + nPSH + 1).

[0457] Spatial coordinate A1 may be a block adjacent to the left side of the target block. A1 may be the lowermost block among the blocks adjacent to the left side of the target block. Alternatively, A1 may be a block adjacent to the top of A0. A1 may be a block occupying the pixels located at coordinates (xP - 1, yP + nPSH).

[0458] Spatial candidate B0 may be a block adjacent to the upper right corner of the target block. B0 may be a block occupying the pixels located at coordinates (xP + nPSW + 1, yP - 1).

[0459] Spatial candidate B1 may be a block adjacent to the top of the target block. B1 may be the rightmost block among the blocks adjacent to the top of the target block. Alternatively, B1 may be a block adjacent to the left of B0. B1 may be a block occupying the pixels located at coordinates (xP + nPSW, yP - 1).

[0460] Spatial candidate B2 may be a block adjacent to the upper left corner of the target block. B2 may be a block occupying the pixels located at coordinates (xP - 1, yP - 1).

[0461] Determination of the availability of spatial and temporal candidates

[0462] In order to include spatial candidate motion information or temporal candidate motion information in a list, it is necessary to determine whether the spatial candidate motion information or the temporal candidate motion information is available.

[0463] Hereinafter, a candidate block may include a spatial candidate and a temporal candidate.

[0464] For example, the determination may be performed by sequentially applying the following steps 1) to 4).

[0465] Step 1) When the PU including the candidate block is outside the boundary of the picture, the availability of the candidate block may be set to "false". The expression "the availability is set to false" may have the same meaning as "set to unavailable".

[0466] Step 2) When the PU including the candidate block is outside the boundary of the slice, the availability of the candidate block may be set to "false". When the target block and the candidate block are in different slices, the availability of the candidate block may be set to "false".

[0467] Step 3) When the PU including the candidate block is outside the boundary of the parallel block, the availability of the candidate block may be set to "false". When the target block and the candidate block are in different parallel blocks, the availability of the candidate block may be set to "false".

[0468] Step 4) When the prediction mode of the PU including the candidate block is the intra prediction mode, the availability of the candidate block may be set to "false". When the PU including the candidate block does not use inter prediction, the availability of the candidate block may be set to "false".

[0469] Figure 10 Shows the order of adding spatial candidate motion information to the merge list according to an embodiment.

[0470] As Figure 10 shown, when multiple pieces of motion information of spatial candidates are added to the merge list, the order of A1, B1, B0, A0, and B2 may be used. That is, multiple pieces of motion information of available spatial candidates may be added to the merge list in the order of A1, B1, B0, A0, and B2.

[0471] Method for deriving a merge list in merge mode and skip mode

[0472] As described above, the maximum number of merge candidates in the merge list may be set. The set maximum number may be indicated by "N". The set number may be sent from the encoding device 100 to the decoding device 200. The slice header of the slice may include N. In other words, the maximum number of merge candidates in the merge list for the target block of the slice may be set through the slice header. For example, the value of N may be substantially 5.

[0473] Multiple pieces of motion information (i.e., merge candidates) can be added to the merge list in the order of the following steps 1) to 4).

[0474] Step 1) Among the spatial candidates, available spatial candidates can be added to the merge list. Multiple pieces of motion information of the available spatial candidates can be added to the merge list in the order shown in Figure 10 Here, when the motion information of the available spatial candidate overlaps with other motion information already existing in the merge list, the motion information of the available spatial candidate may not be added to the merge list. The operation of checking whether the corresponding motion information overlaps with other motion information existing in the list can be simply referred to as "overlap check".

[0475] The maximum number of pieces of motion information to be added can be N.

[0476] Step 2) When the number of pieces of motion information in the merge list is less than N and time candidates are available, the motion information of the time candidates can be added to the merge list. Here, when the motion information of the available time candidate overlaps with other motion information already existing in the merge list, the motion information of the time candidate may not be added to the merge list.

[0477] Step 3) When the number of pieces of motion information in the merge list is less than N and the type of the target strip is "B", the combined motion information generated by combined bidirectional prediction (bi-prediction) can be added to the merge list.

[0478] The target strip can be a strip including the target block.

[0479] The combined motion information can be a combination of L0 motion information and L1 motion information. The L0 motion information can be motion information that only refers to the reference picture list L0. The L1 motion information can be motion information that only refers to the reference picture list L1.

[0480] In the merge list, there can be one or more pieces of L0 motion information. In addition, in the merge list, there can be one or more pieces of L1 motion information.

[0481] The combined motion information can include one or more pieces of combined motion information. When generating the combined motion information, the L0 motion information and the L1 motion information among the one or more pieces of L0 motion information and the one or more pieces of L1 motion information that will be used in the steps for generating the combined motion information can be predefined. One or more pieces of combined motion information can be generated in a predefined order through combined bidirectional prediction using a combination of a pair of different motion information in the merge list. One piece of motion information in the pair of different motion information can be L0 motion information, and the other piece of motion information in the pair of different motion information can be L1 motion information.

[0482] For example, the combined motion information with the highest priority added thereto may be a combination of L0 motion information having a merge index 0 and L1 motion information having a merge index 1. When the motion information having a merge index 0 is not L0 motion information or when the motion information having a merge index 1 is not L1 motion information, the combined motion information may neither be generated nor added. Next, the combined motion information with the next highest priority added thereto may be a combination of L0 motion information having a merge index 1 and L1 motion information having a merge index 0. Subsequent detailed combinations may conform to other combinations in the field of video encoding / decoding.

[0483] Here, when the combined motion information overlaps with other motion information already existing in the merge list, the combined motion information may not be added to the merge list.

[0484] Step 4) When the number of motion information in the merge list is less than N, motion information of a zero vector may be added to the merge list.

[0485] The motion information of a zero vector may be motion information whose motion vector is a zero vector.

[0486] The number of motion information of a zero vector may be one or more. The reference picture indices of one or more pieces of motion information of a zero vector may be different from each other. For example, the value of the reference picture index of the first motion information of a zero vector may be 0. The value of the reference picture index of the second motion information of a zero vector may be 1.

[0487] The number of motion information of a zero vector may be the same as the number of reference pictures in the reference picture list.

[0488] The reference direction of the motion information of a zero vector may be bidirectional. Two motion vectors may be zero vectors. The number of motion information of a zero vector may be the smaller one of the number of reference pictures in the reference picture list L0 and the number of reference pictures in the reference picture list L1. Optionally, when the number of reference pictures in the reference picture list L0 and the number of reference pictures in the reference picture list L1 are different from each other, a unidirectional reference direction may be used for the reference picture index that can be applied only to a single reference picture list.

[0489] The encoding device 100 and / or the decoding device 200 may then add the motion information of a zero vector to the merge list while changing the reference picture index.

[0490] When the motion information of a zero vector overlaps with other motion information already existing in the merge list, the motion information of a zero vector may not be added to the merge list.

[0491] The order of the above steps 1) to 4) is merely exemplary and can be changed. In addition, some of the above steps can be omitted according to predefined conditions.

[0492] Method for deriving a list of predicted motion vector candidates in AMVP mode

[0493] The maximum number of predicted motion vector candidates in the predicted motion vector candidate list can be predefined. The predefined maximum number can be indicated by N. For example, the predefined maximum number can be 2.

[0494] Multiple pieces of motion information (i.e., predicted motion vector candidates) can be added to the predicted motion vector candidate list in the order of the following steps 1) to 3).

[0495] Step 1) Available spatial candidates among the spatial candidates can be added to the predicted motion vector candidate list. The spatial candidates can include a first spatial candidate and a second spatial candidate.

[0496] The first spatial candidate can be one of A0, A1, scaled A0, and scaled A1. The second spatial candidate can be one of B0, B1, B2, scaled B0, scaled B1, and scaled B2.

[0497] Multiple pieces of motion information of the available spatial candidates can be added to the predicted motion vector candidate list in the order of the first spatial candidate and the second spatial candidate. In this case, when the motion information of the available spatial candidates overlaps with other motion information already existing in the predicted motion vector candidate list, the motion information of the available spatial candidates may not be added to the predicted motion vector candidate list. In other words, when the value of N is 2, if the motion information of the second spatial candidate is the same as the motion information of the first spatial candidate, the motion information of the second spatial candidate may not be added to the predicted motion vector candidate list.

[0498] The maximum number of added motion information can be N.

[0499] Step 2) When the number of motion information in the predicted motion vector candidate list is less than N and a temporal candidate is available, the motion information of the temporal candidate can be added to the predicted motion vector candidate list. In this case, when the motion information of the available temporal candidate overlaps with other motion information already existing in the predicted motion vector candidate list, the motion information of the available temporal candidate may not be added to the predicted motion vector candidate list.

[0500] Step 3) When the number of motion information in the predicted motion vector candidate list is less than N, zero vector motion information can be added to the predicted motion vector candidate list.

[0501] The zero vector motion information may include one or more pieces of zero vector motion information. The reference picture indices of the one or more pieces of zero vector motion information may be different from each other.

[0502] The encoding device 100 and / or the decoding device 200 may sequentially add multiple pieces of zero vector motion information to the prediction motion vector candidate list while changing the reference picture index.

[0503] When the zero vector motion information overlaps with other motion information already existing in the prediction motion vector candidate list, the zero vector motion information may not be added to the prediction motion vector candidate list.

[0504] The description of the zero vector motion information made in combination with the merge list above may also be applied to the zero vector motion information. The repetitive description thereof will be omitted.

[0505] The order of steps 1) to 3) described above is merely exemplary and may be changed. In addition, some steps may be omitted according to predefined conditions.

[0506] Figure 11 Illustrates the transform and quantization processing according to an example.

[0507] As Figure 11 shown, quantization levels may be generated by performing transform and / or quantization processing on the residual signal.

[0508] The residual signal may be generated as the difference between the original block and the prediction block. Here, the prediction block may be a block generated via intra prediction or inter prediction.

[0509] The transform may include at least one of a first transform and a second transform. Transform coefficients may be generated by performing the first transform on the residual signal, and second transform coefficients may be generated by performing the second transform on the transform coefficients.

[0510] At least one of a plurality of predefined transform methods may be used to perform the first transform. For example, the plurality of predefined transform methods may include discrete cosine transform (DCT), discrete sine transform (DST), Karhunen-Loève transform (KLT), etc.

[0511] The second transform may be performed on the transform coefficients generated by performing the first transform.

[0512] The transform method applied to the first transform and / or the second transform may be determined based on at least one of the encoding parameters for the target block and / or neighboring blocks. Optionally, transform information indicating the transform method may be signaled from the encoding device to the decoding device 200.

[0513] A quantized level can be generated by performing quantization on the result generated by performing a first transformation and / or a second transformation or on the residual signal.

[0514] The quantized level can be scanned based on at least one of a right upper diagonal scan, a vertical scan, and a horizontal scan according to at least one of an intra prediction mode, a block size, and a block form.

[0515] For example, the coefficients can be changed to a 1D vector form by scanning the coefficients of a block using a right upper diagonal scan. Optionally, according to the size of the transform block and / or the intra prediction mode, a vertical scan that scans the 2D block format coefficients in the column direction or a horizontal scan that scans the 2D block format coefficients in the row direction can be used instead of the right upper diagonal scan.

[0516] The scanned quantized level can be entropy encoded, and the bitstream can include the entropy encoded quantized level.

[0517] The decoding device 200 can generate a quantized level by performing entropy decoding on the bitstream. The quantized level can be arranged in the form of a 2D block via inverse scanning. Here, as a method of inverse scanning, at least one of a right upper diagonal scan, a vertical scan, and a horizontal scan can be performed.

[0518] Inverse quantization can be performed on the quantized level. According to whether a second inverse transformation is performed, a second inverse transformation can be performed on the result generated by performing inverse quantization. In addition, according to whether a first inverse transformation will be performed, a first inverse transformation can be performed on the result generated by performing the second inverse transformation. A reconstructed residual signal can be generated by performing a first inverse transformation on the result generated by performing the second inverse transformation.

[0519] Figure 12 is a configuration diagram of an encoding device according to an embodiment.

[0520] The encoding device 1200 can correspond to the encoding device 100 described above.

[0521] The encoding device 1200 can include a processing unit 1210, a memory 1230, a user interface (UI) input device 1250, a UI output device 1260, and a storage 1240 that communicate with each other via a bus 1290. The electronic device 1200 can also include a communication unit 1220 connected to a network 1299.

[0522] The processing unit 1210 can be a central processing unit (CPU) or a semiconductor device for running processing instructions stored in the memory 1230 or the storage 1240. The processing unit 1210 can be at least one hardware processor.

[0523] The processing unit 1210 can generate and process signals, data, or information that are input to, output from, or used in the encoding device 1200, and can perform inspections, comparisons, determinations, etc. related to the signals, data, or information. In other words, in an embodiment, the generation and processing of data or information and the inspections, comparisons, and determinations related to the data or information can be performed by the processing unit 1210.

[0524] The processing unit 1210 may include an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transformation unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transformation unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.

[0525] At least some of the inter-frame prediction unit 110, the intra-frame prediction unit 120, the switch 115, the subtractor 125, the transformation unit 130, the quantization unit 140, the entropy encoding unit 150, the inverse quantization unit 160, the inverse transformation unit 170, the adder 175, the filter unit 180, and the reference picture buffer 190 may be program modules and may communicate with an external device or system. The program modules may be included in the encoding device 1200 in the form of an operating system, an application program module, or other program modules.

[0526] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device capable of communicating with the encoding device 1200.

[0527] The program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.

[0528] The program modules may be implemented using instructions or code run by at least one processor of the encoding device 1200.

[0529] The processing unit 1210 can run the instructions or code in the inter-frame prediction unit 110, the intra-frame prediction unit 120, the switch 115, the subtractor 125, the transformation unit 130, the quantization unit 140, the entropy encoding unit 150, the inverse quantization unit 160, the inverse transformation unit 170, the adder 175, the filter unit 180, and the reference picture buffer 190.

[0530] The storage unit may represent memory 1230 and / or storage 1240. Each of memory 1230 and storage 1240 may be any one of various types of volatile or non-volatile storage media. For example, memory 1230 may include at least one of read-only memory (ROM) 1231 and random access memory (RAM) 1232.

[0531] The storage unit may store data or information for encoding the operations of device 1200. In an embodiment, the data or information of encoding device 1200 may be stored in the storage unit.

[0532] For example, the storage unit may store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.

[0533] Encoding device 1200 may be implemented in a computer system including a computer-readable storage medium.

[0534] The storage medium may store at least one module required for the operations of encoding device 1200. Memory 1230 may store at least one module and may be configured such that the at least one module is run by processing unit 1210.

[0535] Functions related to the communication of data or information of encoding device 1200 may be performed by communication unit 1220.

[0536] For example, communication unit 1220 may send a bitstream to decoding device 1300 which will be described later.

[0537] Figure 13 is a configuration diagram of a decoding device according to an embodiment.

[0538] Decoding device 1300 may correspond to decoding device 200 described above.

[0539] Decoding device 1300 may include a processing unit 1310, a memory 1330, a user interface (UI) input device 1350, a UI output device 1360, and a storage 1340 that communicate with each other via a bus 1390. Decoding device 1300 may further include a communication unit 1320 connected to a network 1399.

[0540] Processing unit 1310 may be a central processing unit (CPU) or a semiconductor device for running processing instructions stored in memory 1330 or storage 1340. Processing unit 1310 may be at least one hardware processor.

[0541] The processing unit 1310 can generate and process signals, data, or information that is input to, output from, or used in the decoding device 1300, and can perform checks, comparisons, determinations, etc. related to the signals, data, or information. In other words, in an embodiment, the generation and processing of data or information and the checks, comparisons, and determinations related to the data or information can be performed by the processing unit 1310.

[0542] The processing unit 1310 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transformation unit 230, an intra prediction unit 240, an inter prediction unit 250, an adder 255, a filter unit 260, and a reference picture buffer 270.

[0543] At least some of the entropy decoding unit 210, the inverse quantization unit 220, the inverse transformation unit 230, the intra prediction unit 240, the inter prediction unit 250, the adder 255, the filter unit 260, and the reference picture buffer 270 of the decoding device 1300 may be program modules and may communicate with an external device or system. The program modules may be included in the decoding device 1300 in the form of an operating system, an application program module, or other program modules.

[0544] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device capable of communicating with the decoding device 1300.

[0545] The program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.

[0546] The program modules can be implemented using instructions or code run by at least one processor of the decoding device 1300.

[0547] The processing unit 1310 can run instructions or code in the entropy decoding unit 210, the inverse quantization unit 220, the inverse transformation unit 230, the intra prediction unit 240, the inter prediction unit 250, the adder 255, the filter unit 260, and the reference picture buffer 270.

[0548] The storage unit may represent the memory 1330 and / or the storage 1340. Each of the memory 1330 and the storage 1340 can be any of various types of volatile or non-volatile storage media. For example, the memory 1330 may include at least one of a ROM 1331 and a RAM 1332.

[0549] The storage unit can store data or information for the operation of the decoding device 1300. In an embodiment, the data or information of the decoding device 1300 can be stored in the storage unit.

[0550] For example, the storage unit can store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.

[0551] The decoding device 1300 can be implemented in a computer system including a computer-readable storage medium.

[0552] The storage medium can store at least one module required for the operation of the decoding device 1300. The memory 1330 can store at least one module and can be configured such that the at least one module is run by the processing unit 1310.

[0553] Functions related to the communication of data or information of the decoding device 1300 can be performed through the communication unit 1320.

[0554] For example, the communication unit 1320 can receive a bitstream from the encoding device 1200.

[0555] In the following embodiments, when encoding and decoding using inter-frame prediction, a method for deriving the inter-frame prediction information of a target block using the inter-frame prediction information of neighboring blocks will be described.

[0556] When encoding and decoding using inter-frame prediction, the inter-frame prediction information of a target block in a target picture can be searched for in previously encoded and / or previously decoded pictures in order to remove temporal redundancy from the video.

[0557] Since the inter-frame prediction information of neighboring blocks is used to derive the inter-frame prediction information of the target block, the amount of information required for inter-frame prediction can be reduced. Here, the amount of information can be the number of bits.

[0558] As a method for deriving the inter-frame prediction information of a target block using the inter-frame prediction information of neighboring blocks, the AMVP mode and the merge mode can be used. In the AMVP mode and the merge mode, temporally adjacent neighboring blocks and spatially adjacent neighboring blocks can be used to configure the AMVP candidate list and the merge candidate list, respectively. Each candidate in such a list can be inter-frame prediction information or a part of the inter-frame prediction information.

[0559] In the configuration of the list, the inter-frame prediction information of available neighboring blocks can be used as candidates to fill the list.

[0560] When there is no inter-frame prediction information of neighboring blocks, or when the inter-frame prediction information of neighboring blocks cannot be used, the inter-frame prediction information of neighboring blocks cannot be used as candidates.

[0561] The maximum number of candidates in each list can be predefined. When the multiple inter - prediction information of available neighboring blocks cannot fill the predefined maximum number of candidates in the list, zero - vector motion information can be added to the list.

[0562] Since the correlation between the candidates in the list (i.e., inter - prediction information) and the inter - prediction information of the target block is high, the coding performance can be improved.

[0563] Conversely, for candidates with low correlation with the inter - prediction information of the target block (e.g., zero - vector motion information), the number of bits to be signaled may increase when deriving the inter - prediction information. Since the number of bits to be signaled increases, the coding performance may decrease.

[0564] In an embodiment, for a target block without inter - prediction information, the encoding device 1200 and the decoding device 1300 can use the inter - prediction information of neighboring blocks to generate the inter - prediction information for the target block, and can add the generated inter - prediction information as a candidate to the corresponding list.

[0565] In an embodiment, the inter - prediction information with high correlation rather than the inter - prediction information with low correlation can be added as a candidate to the list. Since the inter - prediction information with high correlation is added as a candidate to the list, the coding efficiency can be improved.

[0566] In an embodiment, each of the encoding device 1200 and the decoding device 1300 can use the multiple inter - prediction information of multiple neighboring blocks to configure a list for the target block. The multiple neighboring blocks may include temporal neighboring blocks and spatial neighboring blocks. Each of the encoding device 1200 and the decoding device 1300 can use the multiple inter - prediction information of multiple neighboring blocks to add the inter - prediction information with higher correlation with the inter - prediction information of the target block to the list. By using and adding the inter - prediction information, the number of bits for indicating indexes of the inter - prediction information, etc., can be reduced, and the coding performance can be improved.

[0567] In an embodiment, when there is no inter - prediction information of the target block, each of the encoding device 1200 and the decoding device 1300 can generate the inter - prediction information of the target block by combining the multiple inter - prediction information of the neighboring blocks of the target block. Each of the encoding device 1200 and the decoding device 1300 can add the generated inter - prediction information as a candidate to the list. Instead of the inter - prediction information with low correlation with the inter - prediction information of the target block (e.g., zero - vector motion information), the inter - prediction information with high correlation with the inter - prediction information of the target block is added as a candidate to the list, thus enabling more efficient coding in the derivation of the inter - prediction information.

[0568] In the process of recursively partitioning a CU, the CU can be partitioned into four square blocks of the same size or two blocks of the same size. When the CU is partitioned into two sub-blocks, the CU can be partitioned horizontally or vertically. Optionally, the CU can be partitioned into three sub-blocks and can be partitioned horizontally or vertically. For example, when the CU is partitioned vertically, the ratio of the widths of the sub-blocks generated from the partition can be 1:2:1. Similarly, when the CU is partitioned horizontally, the ratio of the heights of the sub-blocks can be 1:2:1.

[0569] The partitioning of a block in inter prediction can indicate that the coding efficiency obtained when performing inter prediction using the motion information of each sub-block generated from the partition is higher than the coding efficiency obtained when performing inter prediction using the motion information of a single non-partitioned block. In other words, when a block is partitioned, it is very likely that two or four sub-blocks will have different motion information.

[0570] When the target block is a block partitioned from an upper layer block, when deriving the motion information of the spatially neighboring blocks in the merge mode for the target block, the motion information of another sub-block in the upper layer block can be used. In other words, the motion information of the target block can be derived in the same way as the motion information of other sub-blocks. In this case, (even though the sub-blocks are generated by partitioning the upper layer block,) the two sub-blocks have the same motion information, so the coding performance may be reduced.

[0571] Each of the encoding device 1200 and the decoding device 1300 can allocate a smaller number of bits to the candidate with a higher priority among the candidates in the list. The allocated bit can be a value indicating the index of the corresponding candidate. In other words, a smaller number of bits than the number of bits indicating the index of the candidate with a lower priority can be used to signal the index indicating the candidate with a higher priority.

[0572] Each of the encoding device 1200 and the decoding device 1300 can be configured to assign a higher priority to the candidate expected or estimated to have higher coding performance when configuring the list.

[0573] Here, assigning a higher priority can mean 1) allocating a smaller number of bits, 2) allocating a smaller index, or 3) the candidate is included in the list with a higher priority to take precedence over other candidates in the list.

[0574] In addition, each of the encoding device 1200 and the decoding device 1300 can be configured to assign a lower priority to the candidate expected or estimated to have lower coding performance when configuring the list.

[0575] Here, assigning a lower priority may indicate 1) assigning a larger number of bits, 2) assigning a larger index, 3) a candidate being included in the list with a lower priority after other candidates in the list, or 4) a candidate not being included in the list. With such a list configuration, the coding performance can be improved.

[0576] When configuring the list using the motion information of spatially neighboring blocks for a partitioned CU, if the spatially neighboring block is a block divided from an upper CU including the partitioned CU, each of the coding device 1200 and the decoding device 1300 may include the motion information of the spatially neighboring block in the list with a lower priority. Optionally, if the spatially neighboring block is a block divided from an upper CU including the partitioned CU, the coding device 1200 may not include the motion information of the spatially neighboring block in the list.

[0577] In other words, each of the coding device 1200 and the decoding device 1300 may not include the motion information that is very likely to have low coding performance in the list, and may assign a lower priority to the motion information that is very likely to have low coding performance. Through such exclusion and assignment, each of the coding device 1200 and the decoding device 1300 prevents the motion information that is expected or estimated to have low coding performance from being selected, or reduces the possibility of such motion information being selected, thereby improving the coding performance.

[0578] Figure 14 is a flowchart of an inter-frame prediction method according to an embodiment.

[0579] The inter-frame prediction method may be executed by the coding device 1200 and / or the decoding device 1300.

[0580] For example, the coding device 1200 may execute the inter-frame prediction method according to an embodiment to compare the efficiencies of multiple prediction methods for a target block, and may execute the inter-frame prediction method according to an embodiment to generate a reconstructed block for the target block.

[0581] The target block may be any one of the various blocks described above. For example, the target block may be a coding unit (CU), a prediction unit (PU), or a transform unit (TU).

[0582] For example, the decoding device 1300 may execute the inter-frame prediction method according to an embodiment to generate a reconstructed block for the target block.

[0583] Hereinafter, the processing unit may be the processing unit 1210 of the coding device 1200 and / or the processing unit 1310 of the decoding device 1300.

[0584] In step 1410, the processing unit may derive the inter-frame prediction information for the target block.

[0585] Inter-frame prediction information may include 1) motion vectors, 2) reference picture lists, 3) reference picture indices, 4) merge flags, 5) merge indices, 6) advanced motion vector prediction (AMVP) indices, 7) illumination compensation (IC) flags, and 8) overlapping block motion compensation (OBMC) flags.

[0586] The IC flag may be a flag indicating whether IC will be applied.

[0587] The OBMC flag may be a flag indicating whether OBMC will be applied.

[0588] The processing unit may use at least one method to derive inter-frame prediction information.

[0589] The at least one method may include 1) merge mode, 2) AMVP mode, 3) a method for deriving inter-frame prediction information based on sub-blocks, and 4) a method for deriving inter-frame prediction information in decoding device 1300.

[0590] The processing unit may use at least one piece of information to derive inter-frame prediction information.

[0591] The at least one piece of information may include 1) inter-frame prediction information of spatially neighboring blocks, 2) inter-frame prediction information of temporally neighboring blocks, 3) combined inter-frame prediction information, 4) unified candidate lists, 5) adaptive candidate lists depending on block shapes, and 6) adaptive candidate lists depending on block partitioning states.

[0592] In step 1420, the processing unit may perform inter-frame prediction for a target block using the derived inter-frame prediction information.

[0593] Inter-frame prediction may include motion compensation and / or motion correction.

[0594] The processing unit may perform inter-frame prediction using at least one of compensation and / or correction.

[0595] The at least one of the compensation and / or correction may include 1) motion compensation, 2) IC, 3) OBMC, 4) bidirectional optical flow (BIO), 5) affine spatial motion compensation, and 6) motion vector correction in decoding device 1300.

[0596] Using merge mode to derive inter-frame prediction information

[0597] The processing unit may use the merge mode to derive inter-frame prediction information. In an embodiment, the merge mode may be replaced with the AMVP mode or a specific inter-frame prediction mode such as using a list. In other words, the use of the merge mode to derive inter-frame prediction information described in the embodiment may also be applied to deriving inter-frame prediction information using the AMVP mode or a specific inter-frame prediction mode.

[0598] The processing unit can configure a merge candidate list. The number of merge candidates in the merge candidate list can be N. N can be a positive integer. For example, the merge candidate can be inter prediction information and can include a motion vector and a reference picture list.

[0599] The processing unit can use one or more of the inter prediction information of spatially neighboring blocks, the inter prediction information of temporally neighboring blocks, and combined inter prediction information to configure the merge candidate list. Here, the processing unit can add multiple pieces of inter prediction information to the merge candidate list in a specific order.

[0600] When configuring the merge candidate list, the processing unit can add multiple pieces of inter prediction information of neighboring blocks to the merge candidate list as merge candidates. Here, the processing unit can add multiple pieces of inter prediction information of neighboring blocks to the merge candidate list in a specific order of the neighboring blocks. 1) When the inter prediction information of a neighboring block does not exist or 2) when the inter prediction information of a neighboring block is the same as the inter prediction information existing in the merge candidate list (i.e., when the inter prediction information of the neighboring block has already been included in the merge candidate list), the processing unit may not add the inter prediction information of the neighboring block to the merge candidate list. In other words, when the multiple pieces of inter prediction information of two neighboring blocks are the same as each other, the inter prediction information of the neighboring block with a lower priority may not be added to the merge candidate list.

[0601] When the inter prediction information of one of the neighboring blocks is not added to the merge candidate list, the processing unit can add the combined inter prediction information instead of the unadded inter prediction information to the merge candidate list. For example, when the inter prediction information of a specific neighboring block does not exist, or when the inter prediction information of a specific neighboring block is the same as the inter prediction information in the merge candidate list, the processing unit can derive the combined inter prediction information for the specific neighboring block and add the derived combined inter prediction information to the merge candidate list.

[0602] The processing unit can generate combined inter prediction information by combining two or more pieces of inter prediction information of neighboring blocks of a target block with each other.

[0603] The processing unit can configure an inter prediction information palette. The inter prediction information palette can be a list having N pieces of inter prediction information. N can be a positive integer. Here, the processing unit can 1) add the inter prediction mode of a target block to the inter prediction information palette, and 2) manage the inter prediction modes in the inter prediction information palette according to a specific order and method. For example, when the inter prediction information palette is filled with information, the processing unit can manage the inter prediction information palette in a first-in first-out (FIFO) manner.

[0604] When the inter-frame prediction information of the target block is the same as the inter-frame prediction information existing in the inter-frame prediction information palette (i.e., when the inter-frame prediction information of the target block has been included in the inter-frame prediction information palette), the processing unit may not add the inter-frame prediction information of the target block to the inter-frame prediction information palette.

[0605] When the inter-frame prediction information of the target block is the same as the inter-frame prediction information existing in the inter-frame prediction information palette, the processing unit may move the inter-frame prediction information in the inter-frame prediction information palette that is the same as the inter-frame prediction information of the target block to the position of the first inter-frame prediction information in the inter-frame prediction information palette. In other words, the processing unit may assign a specific priority (such as the highest priority) to the inter-frame prediction information in the inter-frame prediction information palette that is the same as the inter-frame prediction information of the target block, and may adjust the positions of the multiple pieces of inter-frame prediction information existing in the inter-frame prediction information palette based on the assigned priority.

[0606] The processing unit may initialize the inter-frame prediction information palette for all blocks in the target picture as a unit for each picture. In other words, the blocks in the target picture may share a single inter-frame prediction information palette with each other.

[0607] The processing unit may use the inter-frame prediction information existing in the inter-frame prediction information palette as the inter-frame prediction information for temporally adjacent blocks.

[0608] Figure 15 Shows the spatial adjacent blocks of the target block according to the example.

[0609] In Figure 15 A to K may indicate respective spatial adjacent blocks.

[0610] The inter-frame prediction information of the spatial adjacent blocks may be the inter-frame prediction information of the block existing at any one of the positions corresponding to A to K of Figure 15 .

[0611] Hereinafter, the term "inter-frame prediction information of block X" may be understood as "inter-frame prediction information corresponding to position X".

[0612] For example, the size of the adjacent block may be M×N. M and N may each be at least one of 2, 4, 8, 16, 32, 64, and 128.

[0613] The left adjacent block may be an adjacent block adjacent to the left side of the target block, and may be one or more of block A, block B, block C, block D, and block E.

[0614] The upper adjacent block may be an adjacent block adjacent to the upper side of the target block, and may be one or more of block G, block H, block I, block J, and block K.

[0615] The upper-left neighboring block may be a neighboring block adjacent to the upper-left corner of the target block, and may be block F.

[0616] Such a spatially neighboring block may be a block adjacent to the boundary of the target block or a block not adjacent to the boundary of the target block.

[0617] Figure 16 Shows the temporal neighboring blocks of the target block according to the example.

[0618] In Figure 16 L to W may represent respective temporal neighboring blocks.

[0619] The temporal neighboring block may be a block in a previous picture. The previous picture may be a previously reconstructed col picture. The previous picture may be a picture that has been encoded or decoded before the target picture is encoded or decoded.

[0620] The previous picture may be a picture having a picture order count (POC) larger than the picture order count (POC) of the target picture.

[0621] The position of the temporal neighboring block in the previous picture may be the same as the position of the target block in the target picture. Optionally, the position of the temporal neighboring block in the previous picture may correspond to the position of the target block in the target picture. Optionally, the position of the temporal neighboring block in the previous picture may correspond to at least one of the position of the lower-right part of the target block, the central part of the target block, and a specific position of the target block.

[0622] Optionally, the temporal neighboring block may be a block adjacent to the col block. For example, the temporal neighboring block may be a block adjacent to the lower-right vertex of the col block.

[0623] Optionally, the temporal neighboring block may be a block in the target picture that is temporally previous. The block that is temporally previous may be a block that has been encoded or decoded before the target block is encoded or decoded.

[0624] The temporal neighboring block may be a specific neighboring block referred to in the process of configuring the merge candidate list. Here, the specific neighboring block may be a neighboring block corresponding to the inter prediction information included in the merge candidate list.

[0625] The inter prediction information of the temporal neighboring block may be the inter prediction information of a block arranged at a specific position in the previous picture. Here, the specific position may be the position of the target block in the target picture.

[0626] The inter prediction information of the temporal neighboring block may be the inter prediction information of a block arranged at a specific position in the target picture. Here, the specific position may be the position of the spatial neighboring block of the target block in the target picture.

[0627] Using combined inter-frame prediction information to configure the merge candidate list

[0628] Figure 17 Shows the generation of combined inter-frame prediction information for the upper-right neighboring block according to an example.

[0629] Figure 18 Shows the generation of combined inter-frame prediction information for the upper neighboring block according to an example.

[0630] The processing unit may use the combined inter-frame prediction information to configure the merge candidate list. The combined inter-frame prediction information may replace the inter-frame prediction information or motion information of the neighboring block, and may then be added to the merge candidate list as a new merge candidate.

[0631] The processing unit may generate combined inter-frame prediction information by combining multiple pieces of inter-frame prediction information related to the target block. For example, the inter-frame prediction information related to the target block may be the inter-frame prediction information of the neighboring block of the target block. The inter-frame prediction information related to the target block may be neighboring inter-frame prediction information. The neighboring inter-frame prediction information may be the inter-frame prediction information of the neighboring block.

[0632] The combined inter-frame prediction information may include motion information. The motion information of the combined inter-frame prediction information may include at least one of a reference picture list, a reference picture index, an inter-frame prediction indicator, a motion vector, a motion vector candidate, a motion vector candidate list, and a picture order count (POC).

[0633] In an embodiment, the term "inter-frame prediction information" may be replaced with "motion information" and "motion vector", and "combined inter-frame prediction information" may be replaced with "combined motion information" and "combined motion vector".

[0634] When generating partial information of the combined inter-frame prediction information, the processing unit may select the neighboring inter-frame prediction information to be used from multiple pieces of neighboring inter-frame prediction information.

[0635] For example, the processing unit may use partial information of the selected neighboring inter-frame prediction information as partial information of the combined inter-frame prediction information. In other words, the processing unit may assign the value of the partial information of the selected neighboring inter-frame prediction information to the partial information of the combined inter-frame prediction information.

[0636] For example, the partial information may be an IC flag or an OBMC flag.

[0637] For example, the processing unit may use the neighboring inter-frame prediction information selected in relation to the combination of motion vectors from multiple pieces of neighboring inter-frame prediction information to generate partial information of the combined inter-frame prediction information.

[0638] For example, the processing unit may use the inter-frame prediction information for the combination of motion vectors in multiple neighboring inter-frame prediction information (scaled reference) to generate partial information of the combined inter-frame prediction information.

[0639] For example, referring to Figure 15 , when using the inter-frame prediction information of blocks B and J to generate the combined inter-frame prediction information for block F, the IC flag and / or OBMC flag of block B or block J may be used as the IC flag and / or OBMC flag of block F.

[0640] The processing unit may generate the motion vector of the combined inter-frame prediction information by combining the motion vectors of multiple neighboring motion information. The neighboring motion information may be the motion information of neighboring blocks. In addition, the neighboring motion vectors may be the motion vectors of neighboring blocks.

[0641] For example, the neighboring motion information may be the motion information of blocks (such as Figure 15 the spatially neighboring blocks A to K shown in Figure 16 and the temporally neighboring blocks L to W shown in

[0642] The neighboring motion information may be the motion information of each block that is not adjacent to the target block. The block that is not adjacent to the target block may be a block adjacent to the neighboring block of the target block.

[0643] The neighboring motion information may be the motion information of a block having the above specific relationship with the target block, and the block having the specific relationship may also be a block not adjacent to the target block. For example, the block having the specific relationship may be a block adjacent to the neighboring block of the target block. The neighboring block may be inserted between the block having the specific relationship and the target block.

[0644] The combined inter-frame prediction information may be a result obtained by selecting one from the motion information of multiple neighboring blocks. Here, one of the multiple motion information may be selected as the combined inter-frame prediction information according to specified conditions. For example, the combined inter-frame prediction information may be a result of calculation, selection, combination, and transformation using the motion information of multiple neighboring blocks.

[0645] For example, the combined inter-frame prediction information may be the motion information of the following neighboring block: for this neighboring block, the difference between the POC of the target picture and the POC of the reference picture for this neighboring block is the smallest.

[0646] For example, the combined inter-frame prediction information may be specific motion information existing in the merge candidate list.

[0647] The combined inter-frame prediction information may be a result obtained by selecting and combining one or more neighboring motion information from multiple neighboring motion information. Here, the combined inter-frame prediction information may be selected according to specified conditions.

[0648] For example, the motion vector for combining inter-frame prediction information may be the motion vector of a specific neighboring block among a plurality of neighboring blocks. Here, the specific neighboring block may be the neighboring block with the smallest difference between the POC of the target picture among the plurality of neighboring blocks and the POC of the reference picture for the neighboring block. The motion vector for combining inter-frame prediction information may be the result of a formula using a plurality of neighboring motion vectors. The neighboring motion vectors may be the motion vectors of neighboring blocks. The neighboring motion vectors of neighboring blocks may include a plurality of motion vectors.

[0649] When the combined inter-frame prediction information is generated, the processing unit may generate unidirectional combined inter-frame prediction information or bidirectional combined inter-frame prediction information. Here, the unidirectional combined inter-frame prediction information may be forward (L0) inter-frame prediction information or backward (L1) inter-frame prediction information, and the bidirectional combined inter-frame prediction information may be forward inter-frame prediction information and backward inter-frame prediction information.

[0650] The unidirectional combined inter-frame prediction information may be a combination of: 1) multiple bidirectional inter-frame prediction information of neighboring blocks, 2) multiple unidirectional prediction information of neighboring blocks, and 3) multiple L0 inter-frame prediction information or L1 inter-frame prediction information among multiple combined inter-frame prediction information.

[0651] For example, the processing unit may generate L0 (L1) direction inter-frame prediction information by combining two or more temporally neighboring blocks, and may add the generated unidirectional prediction information to the merge candidate list.

[0652] For example, the combined inter-frame prediction information may be the result of a combination of multiple L0 (L1) direction inter-frame prediction information of two previous blocks and L0 (L1) direction inter-frame prediction information of one temporally neighboring block, and may subsequently be added to the merge candidate list.

[0653] For example, the combined inter-frame prediction information may be the result of a combination of multiple L0 (L1) direction inter-frame prediction information of two specific neighboring blocks referred to during the merge candidate list configuration process, and may be added to the merge candidate list.

[0654] The bidirectional combined inter-frame prediction information may be a combination of the above forward inter-frame prediction information and backward inter-frame prediction information.

[0655] The motion vector of the combined inter-frame prediction information may be a combination of neighboring motion vectors. For example, the neighboring motion information may be a combination of multiple motion information of a plurality of neighboring blocks A to W. For example, the motion vector of the combined inter-frame prediction information may be the average value, maximum value, minimum value, or median value of a plurality of neighboring motion vectors, and may be a combination of one or more of the average value, maximum value, minimum value, and median value.

[0656] The average value can be obtained by dividing the sum of the combined motion vectors by the number of combined motion vectors. For example, the average value of motion vector (4, 6) and motion vector (6, 10) can be (5, 8).

[0657] For example, the motion vector of the combined inter-frame prediction information can be a weighted average of multiple neighboring motion vectors, or can be a combination using the variations between multiple neighboring motion vectors.

[0658] For example, as Figure 17 shown, for target block 1710, there may be an upper-left neighboring block 1720, an upper neighboring block 1730, an upper-right neighboring block 1740, a left neighboring block 1750, and a lower-left neighboring block 1760. By combining the motion information 1721 of the upper-left neighboring block 1720 and the motion information 1731 of the upper neighboring block 1730, combined inter-frame prediction information corresponding to the motion information 1741 of the upper-right neighboring block 1740 can be generated. Such combined inter-frame prediction information can be used as the motion information 1711 of the target block 1710.

[0659] For example, as Figure 18 shown, by combining the motion information 1721 of the upper-left neighboring block 1720 and the motion information 1741 of the upper-right neighboring block 1740, combined inter-frame prediction information corresponding to the motion information 1731 of the upper neighboring block 1730 can be generated. Such combined inter-frame prediction information can be used as the motion information 1711 of the target block 1710.

[0660] As described above, the motion vector of the combined inter-frame prediction information can be the result of a weighted combination of neighboring motion vectors. A higher weight can be assigned to the motion vector of the neighboring block that has a higher correlation with the target block.

[0661] The motion vector of the combined inter-frame prediction information can be the result of a weighted combination of neighboring motion vectors based on the block size (i.e., a weighted combination based on the block size). The weighted combination based on the block size can be represented by Equation 2 below:

[0662] [Equation 2]

[0663]

[0664] MV Left can be the left neighboring motion vector of the target block. The left neighboring motion vector of the target block can be the motion vector of the neighboring block adjacent to the left side of the target block.

[0665] “width” can represent the width of the target block and can be the weight for the left neighboring motion vector of the target block.

[0666] MV AboveIt can be the upper neighboring motion vector of the target block. The upper neighboring motion vector of the target block can be the motion vector of the neighboring block adjacent to the upper side of the target block.

[0667] "height" can represent the height of the target block and can be the weight for the upper neighboring motion vector of the target block.

[0668] The motion vector for combining the inter-frame prediction information can be the result of a weighted combination of POC-based neighboring motion vectors (i.e., a weighted combination based on POC).

[0669] For example, as the POC of the reference picture for the neighboring motion information gets closer to the POC of the target picture, the weight for the neighboring motion vector can be larger.

[0670] Using a varying combination can be using the variation between two or more motion vectors to generate the inter-frame prediction information for the block before the combined block or the block after the combined block.

[0671] For example, referring to Figure 15 , the processing unit can derive the motion vector of block K by using a combination of the variation between the motion vector of block I and the motion vector of block J. The motion vector of block K can be obtained by adding the difference between the motion vector of block J and the motion vector of block I to the motion vector of block J.

[0672] Optionally, the motion vector for combining the inter-frame prediction information can be the result of an interpolation-based combination of neighboring motion vectors.

[0673] For example, the interpolation-based combination of two neighboring motion vectors can be represented by the following Equation 3:

[0674] [Equation 3]

[0675] 2×MV0 - MV1

[0676] MV0 can be the first neighboring motion vector. MV1 can be the second neighboring motion vector.

[0677] Scaling can be applied to the inter-frame prediction information used to generate the combined inter-frame prediction information.

[0678] For example, when the neighboring motion information indicates bi-directional prediction, the processing unit can scale the motion vector of L0 based on the motion vector of L1. The processing unit can generate the combined inter-frame prediction information by combining the scaled motion vector of L0 with the motion vector of L1.

[0679] For example, when the neighboring motion information indicates bi - directional prediction, the processing unit may scale the motion vector of L1 based on the motion vector of L0. The processing unit may generate combined inter - prediction information by combining the scaled motion vector of L1 with the motion vector of L0, and may add the generated combined inter - prediction information to the merge candidate list.

[0680] When the reference pictures for multiple neighboring motion information used to generate combined inter - prediction information are different from each other, the processing unit may change the inter - prediction information by applying scaling to the inter - prediction information.

[0681] For example, referring to Figure 15 , when the neighboring blocks for combination are block B and block J, and the reference pictures for block B and block J are different from each other, the processing unit may scale the motion vector of block J according to the temporal distance between block B and the reference picture for block B.

[0682] For example, referring to Figure 15 , when the neighboring blocks for combination are block B and block J, and the reference pictures for block B and block J are different from each other, the processing unit may scale the motion vector of block B according to the temporal distance between block J and the reference picture for block J.

[0683] When scaling, the processing unit may select the neighboring motion information used as a reference for scaling.

[0684] The processing unit may select the neighboring motion information used as a reference for scaling based on the POC. The processing unit may select the following motion information from multiple neighboring motion information as the reference motion information: for this motion information, the POC of the reference picture for this motion information is closer to the POC of the target picture.

[0685] When the combined inter - prediction information is generated, the processing unit may determine whether to perform combination based on specified conditions.

[0686] When the combined inter - prediction information is generated, the processing unit may determine whether to perform combination based on the similarity between multiple pieces of inter - prediction information or multiple pieces of motion information used for combination. For example, when the similarity is less than a predefined threshold, the processing unit may not perform combination. Optionally, when the similarity is greater than a predefined threshold, the processing unit may not perform combination.

[0687] Here, the similarity may indicate the value or result of a formula using multiple pieces of motion information.

[0688] The processing unit may generate combined inter - prediction information by combining multiple neighboring motion information based on the direction of the reference picture for the neighboring motion information. For example, the processing unit may generate combined inter - prediction information by combining neighboring motion vectors in the same direction.

[0689] As described above, the processing unit may generate combined inter prediction information by combining a plurality of pieces of inter prediction information of a plurality of blocks. Here, each of the plurality of blocks may be a block that satisfies a specified condition.

[0690] exist Figure 15 In each block shown in , the inter-frame prediction information of the specific block may be replaced by the combined inter-frame prediction information. The processing unit may derive the first inter-frame prediction information for the position on the left side of the specific block, and derive the second inter-frame prediction information for the position on the right side of the specific block, and may generate the combined inter-frame prediction information by combining the first inter-frame prediction information with the second inter-frame prediction information. The generated combined inter-frame prediction information may replace the inter-frame prediction information of the specific block. Here, the inter-frame prediction information for the position on the left side of the specific block may be the inter-frame prediction information of the block located on the left side of the specific block. Here, the inter-frame prediction information for the position on the right side of the specific block may be the inter-frame prediction information of the block located on the right side of the specific block. In addition, such derivation, combination and generation may also be applied to a part of the inter-frame prediction information, such as motion information and motion vector.

[0691] For example, when first inter prediction information for a left position is derived, if inter prediction information of only one block located on the left side of a specific block is available, the processing unit may perform combination using the available inter prediction information.

[0692] For example, when second inter prediction information for a right position is derived, if inter prediction information of only one block located on the right side of a specific block is available, the processing unit may perform combination using the available inter prediction information.

[0693] For example, when pieces of inter-frame prediction information of a plurality of blocks located on the left side of a specific block are available, the processing unit may derive first inter-frame prediction information by combining the available pieces of inter-frame prediction information.

[0694] For example, when a plurality of pieces of inter-frame prediction information of a plurality of blocks located on the right side of a specific block are available, the processing unit may derive second inter-frame prediction information by combining the available plurality of pieces of inter-frame prediction information.

[0695] For example, when the first inter-frame prediction information for the left position is derived, if multiple inter-frame prediction information of multiple blocks located on the left side of a specific block are available, the processing unit may select specific inter-frame prediction information from the multiple inter-frame prediction information as the first inter-frame prediction information, and the selected inter-frame prediction information may be used for the generation of combined inter-frame prediction information.

[0696] For example, if multiple pieces of inter-frame prediction information of multiple blocks located to the left of a specific block are available, the processing unit may select, from the multiple blocks, the block having the shortest temporal distance from the target picture. The processing unit may select the inter-frame prediction information of the selected block as the first inter-frame prediction information. The temporal distance between pictures may be the difference between the sequential positions of the displayed pictures.

[0697] For example, when the second inter-frame prediction information for the right position is derived, if multiple pieces of inter-frame prediction information of multiple blocks located to the right of a specific block are available, the processing unit may select specific inter-frame prediction information from the multiple pieces of inter-frame prediction information as the second inter-frame prediction information, and may use the selected inter-frame prediction information for generating combined inter-frame prediction information.

[0698] For example, if multiple pieces of inter-frame prediction information of multiple blocks located to the right of a specific block are available, the processing unit may select, from the multiple blocks, the block having the shortest temporal distance from the target picture. The processing unit may select the inter-frame prediction information of the selected block as the second inter-frame prediction information.

[0699] For example, in Figure 15 in each of the blocks shown, when multiple pieces of inter-frame prediction information of multiple blocks located to the left of a specific block are available, the processing unit may generate combined inter-frame prediction information by combining the multiple pieces of inter-frame prediction information. The generated combined inter-frame prediction information may replace the inter-frame prediction information of the specific block. In addition, when multiple pieces of inter-frame prediction information of multiple blocks located to the right of a specific block are available, the processing unit may generate combined inter-frame prediction information by combining the multiple pieces of inter-frame prediction information.

[0700] For example, in Figure 16 in a specific block among blocks N, O, P, and Q, when multiple pieces of inter-frame prediction information of multiple co-located blocks in different previous pictures are available, the processing unit may generate combined inter-frame prediction information by combining the multiple available pieces of inter-frame prediction information. The generated combined inter-frame prediction information may replace the inter-frame prediction information of the specific block. Such co-located blocks may be the col blocks for the specific block. In other words, the positions of the co-located blocks in different previous pictures may be the same as the position of the specific block in the previous pictures.

[0701] For example, the combined inter-frame prediction information for block N may be a combination of multiple pieces of inter-frame prediction information for co-located blocks L and M. For example, the combined inter-frame prediction information for block O may be a combination of multiple pieces of inter-frame prediction information for co-located blocks U and R. For example, the combined inter-frame prediction information for block P may be a combination of multiple pieces of inter-frame prediction information for co-located blocks V and S. For example, the combined inter-frame prediction information for block Q may be a combination of multiple pieces of inter-frame prediction information for co-located blocks W and T.

[0702] For example, the processing unit may generate combined inter prediction information by combining one or more pieces of inter prediction information of one or more spatially adjacent blocks with one or more pieces of inter prediction information of one or more temporally adjacent blocks.

[0703] Figure 19 Illustrates the generation of combined inter prediction information of adjacent blocks according to an example.

[0704] In Figure 19 the target CU may indicate the target block.

[0705] In Figure 19 block AL, block A, block AR, block L, and block LB may be the upper left adjacent block, the upper adjacent block, the upper right adjacent block, the left adjacent block, and the lower left adjacent block of the target block, respectively.

[0706] The upper adjacent block may refer to the rightmost (or leftmost) block among multiple adjacent blocks above the target block. The left adjacent block may refer to the lowermost (or uppermost) block among multiple adjacent blocks to the left of the target block.

[0707] Block A and block AR may be understood as multiple blocks adjacent to or above the target block. Block L and block LB may be understood as multiple blocks adjacent to or to the left of the target block.

[0708] The processing unit may generate first adjacent inter prediction information by combining multiple pieces of inter prediction information of block A and block AR, where the multiple pieces of inter prediction information may be understood as upper inter prediction information or upper motion vectors. The processing unit may generate second adjacent inter prediction information by combining multiple pieces of inter prediction information of block L and block LB. These multiple pieces of inter prediction information may be understood as left inter prediction information or left motion vectors. The processing unit may generate combined inter prediction information by combining the first adjacent inter prediction information with the second adjacent inter prediction information.

[0709] When the inter prediction information of block AL is unavailable, the processing unit may use the generated combined inter prediction information to replace the inter prediction information of block AL. In addition, the processing unit may add the combined inter prediction information generated for block AL as a new merge candidate to the merge candidate list.

[0710] When the above first adjacent inter prediction information, second adjacent inter prediction information, and combined inter prediction information are obtained, the foregoing combination method and the combination method to be described below may be used.

[0711] When only one of the multiple inter - prediction information of block A and block AR is available, the processing unit may use the available inter - prediction information among the multiple inter - prediction information as the first neighboring inter - prediction information. In addition, when only one of the multiple inter - prediction information of block L and block LB is available, the processing unit may use the available inter - prediction information among the multiple inter - prediction information as the second neighboring inter - prediction information.

[0712] The processing unit may select any one of the multiple inter - prediction information of block A and block AR, and may use the selected inter - prediction information as the first neighboring inter - prediction information. The processing unit may select any one of the multiple inter - prediction information of block L and block LB, and may use the selected inter - prediction information as the second neighboring inter - prediction information.

[0713] The processing unit may select the inter - prediction information of block A and block AR whose POC is closer to the POC of the target picture as the first neighboring inter - prediction information. Here, the POC of the inter - prediction information may be the POC of the motion information for the inter - prediction information.

[0714] The processing unit may select the inter - prediction information of block L and block LB whose POC is closer to the POC of the target picture as the second neighboring inter - prediction information.

[0715] Figure 20 Illustrates the generation of the inter - prediction information of block AL according to an example.

[0716] MV0 and MV1 may refer to the respective motion vectors used for combining the inter - prediction information.

[0717] The processing unit may generate combined inter - prediction information by combining the multiple inter - prediction information of block L and block LB. The combined inter - prediction information may replace the inter - prediction information of block AL, and may be added as a merge candidate to the merge candidate list for the target block.

[0718] Figure 21 Illustrates the generation of the inter - prediction information of block AR according to an example.

[0719] The processing unit may generate combined inter - prediction information by combining the multiple inter - prediction information of block A and block L. The combined inter - prediction information may replace the inter - prediction information of block AR, and may be added as a merge candidate to the merge candidate list for the target block. For example, the aforementioned combination may be a combination based on extrapolation.

[0720] Figure 22 Illustrates the generation of the inter - prediction information of the target CU according to an example.

[0721] The processing unit may generate combined inter - prediction information by combining multiple pieces of inter - prediction information of block L and non - adjacent blocks. The combined inter - prediction information may be added as a merge candidate to the merge candidate list for the target block. For example, the aforementioned combination may be an interpolation - based combination.

[0722] Using the inter-frame prediction information in the merge candidate list to generate combined inter-frame prediction information

[0723] The processing unit may generate combined inter - prediction information by combining M pieces of inter - prediction information in the merge candidate list, and may use the generated combined inter - prediction information for inter - prediction, or may add the generated combined inter - prediction information to the merge candidate list.

[0724] Here, the directions of the motion vectors of the multiple pieces of inter - prediction information to be combined may be the same as each other.

[0725] Here, M may be an integer of 2 or greater, and may be less than or equal to the number of multiple pieces of inter - prediction information in the merge candidate list.

[0726] For example, when there are three pieces of inter - prediction information in the merge candidate list and the value of M is 2, the combinations that can be used may be (the first piece of inter - prediction information, the second piece of inter - prediction information), (the first piece of inter - prediction information, the third piece of inter - prediction information), and (the second piece of inter - prediction information, the third piece of inter - prediction information), and the processing unit may generate combined inter - prediction information by combining multiple pieces of inter - prediction information according to these combinations.

[0727] For example, when there are four pieces of inter - prediction information in the merge candidate list and the value of M is 4, the combination that can be used may be (the first piece of inter - prediction information, the second piece of inter - prediction information, the third piece of inter - prediction information, and the fourth piece of inter - prediction information), and the processing unit may generate combined inter - prediction information by combining multiple pieces of inter - prediction information according to this combination.

[0728] Configuration of the merge candidate list

[0729] As described above, the processing unit may add multiple pieces of inter - prediction information of adjacent blocks as merge candidates to the merge candidate list when configuring the merge candidate list. Here, the processing unit may add multiple pieces of inter - prediction information of adjacent blocks to the merge candidate list in a specific order of adjacent blocks.

[0730] Return reference Figure 15 When configuring the merge candidate list, the processing unit may use multiple pieces of inter - prediction information of specific spatial adjacent blocks as merge candidates. The specific spatial adjacent blocks may be block A, block B, block F, block J, and block K.

[0731] The processing unit may add 1) combined inter-frame prediction information, 2) a mode derived based on sub-block motion information (e.g., an optional temporal motion vector prediction (ATMVP) mode, a spatio-temporal motion vector prediction (STMVP) mode, etc.), and 3) an affine spatial motion information derivation mode to the merge candidate list in a specific order.

[0732] For example, the processing unit may configure the merge candidate list in the order (B, J, K, A, ATMVP, F, combined inter-frame prediction information).

[0733] For example, the processing unit may configure the merge candidate list in the order (B, J, K, A, first combined inter-frame prediction information, ATMVP, F, second combined inter-frame prediction information). Here, the first combined inter-frame prediction information and the second combined inter-frame prediction information may differ from each other in terms of neighboring blocks to be referred to for generating the first combined inter-frame prediction information and the second combined inter-frame prediction information, the number of neighboring blocks, and the directionality of the combined inter-frame prediction information.

[0734] Hereinafter, "(α, β, γ, δ, ε)" may indicate the order of blocks, and may represent that the block with a symbol earlier in the parentheses is processed earlier than the block with a symbol later. The expression of configuring the merge candidate list in the order of "(α, β, γ, δ, ε)" may indicate that in the configuration of the merge candidate list, the task of configuring the merge candidate list is performed in the order of block α, block β, block γ, block δ, and block ε, and may also represent that the blocks are processed in the order of the listed blocks.

[0735] Here, the processing unit may perform the following tasks 1) to 5) for each block in the order of the blocks.

[0736] 1) The processing unit may determine whether to add the inter-frame prediction information of the corresponding block to the merge candidate list.

[0737] 2) If it is determined to add the inter-frame prediction information of the block to the merge candidate list, the processing unit may add the inter-frame prediction information of the block to the merge candidate list.

[0738] 3) (If it is determined not to add the inter-frame prediction information of the block to the merge candidate list), the processing unit may determine whether to derive combined inter-frame prediction information for the corresponding block.

[0739] 4) If it is determined to derive combined inter-frame prediction information, the processing unit may derive combined inter-frame prediction information.

[0740] 5) The processing unit may determine whether to add the combined inter-frame prediction information to the merge candidate list.

[0741] 6) If it is determined to add the combined inter-frame prediction information to the merge candidate list, the processing unit may add the combined inter-frame prediction information to the merge candidate list.

[0742] When tasks 1) to 5) are performed on a block, task 1) can be performed on a subsequent block.

[0743] For example, the processing unit can configure the merge candidate list in the order of (B, J, K, A, F).

[0744] For example, the processing unit can configure the merge candidate list in the order of (J, B, A, K, F).

[0745] Figure 23 Illustrates the case where a CU with the same width and height is vertically divided.

[0746] Figure 24 Illustrates the case where a CU with the same width and height is horizontally divided.

[0747] Figure 25 Illustrates the case where a CU with a width greater than its height is vertically divided.

[0748] Figure 26 Illustrates the case where a CU with a height greater than its width is horizontally divided.

[0749] The comparison between the width and height and the direction of division can be used to determine the order of neighboring blocks.

[0750] The processing unit can determine a scheme for configuring the merge candidate list based on the shape of the target block.

[0751] The scheme for configuring the merge candidate list can include the order of neighboring blocks required for configuring the merge list. The order of neighboring blocks can be the order of availability tests for multiple inter-frame prediction information for neighboring blocks and the addition of the multiple inter-frame prediction information.

[0752] When the height of the target block is greater than the width of the target block, the processing unit can configure the merge candidate list in the order of (J, B, K, A, F), and when the height of the target block is less than or equal to the width of the target block, the processing unit can configure the merge candidate list in the order of (B, J, K, A, F).

[0753] When the height of the target block is greater than the width of the target block, the processing unit can configure the merge candidate list in the order of (J, B, A, K, F), and when the height of the target block is less than or equal to the width of the target block, the processing unit can configure the merge candidate list in the order of (B, J, K, A, F).

[0754] For example, when the height of the target block is greater than the width of the target block, the processing unit may configure the merge candidate list in the order of (J, K, B, A, F). When the height of the target block is less than the width of the target block, the processing unit may configure the merge candidate list in the order of (B, A, J, K, F). And when the height and width of the target block are equal to each other, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0755] For example, when the height of the target block is greater than the width of the target block, the processing unit may use multiple inter-frame prediction information of spatial neighboring blocks located above the target block to configure the merge candidate list. When the height of the target block is greater than the width of the target block, the processing unit may configure the merge candidate list in the order of (F, G, H, I, J, K).

[0756] For example, when the height of the target block is less than the width of the target block, the processing unit may use multiple inter-frame prediction information of spatial neighboring blocks located to the left of the target block to configure the merge candidate list. When the height of the target block is less than the width of the target block, the processing unit may configure the merge candidate list in the order of (A, B, C, D, E, F).

[0757] The processing unit may determine a scheme for configuring the merge candidate list based on the partitioning state of the target block.

[0758] The partitioning state may represent the partitioning direction. The partitioning state of the target block may be the type or direction of the partitioning used to generate the partitioning of the target block. Optionally, the partitioning state of the target block may be the type or direction of the partitioning applied to the upper-level block of the target block.

[0759] For example, when the target block is obtained as a result of vertical partitioning, the processing unit may configure the merge candidate list in the order of (J, B, K, A, F), while when the target block is not obtained as a result of vertical partitioning, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0760] For example, when the target block is obtained as a result of vertical partitioning, the processing unit may configure the merge candidate list in the order of (J, B, A, K, F), while when the target block is not obtained as a result of vertical partitioning, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0761] For example, according to which one of vertical partitioning, horizontal partitioning, and quad partitioning has been used to obtain the target block, the processing unit may select one order from different orders of neighboring blocks and may configure the merge candidate list in the selected order.

[0762] For example, when the target block is obtained as a result of vertical partitioning, the processing unit may configure the merge candidate list in the order of (J, K, B, A, F). When the target block is obtained as a result of horizontal partitioning, the processing unit may configure the merge candidate list in the order of (B, A, J, K, F). When the target block is obtained as a result of quad partitioning, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0763] For example, when the target block is obtained as a result of vertical partitioning, the processing unit may use multiple pieces of inter-frame prediction information of spatially adjacent blocks located above the target block to configure the merge candidate list. When the target block is obtained as a result of vertical partitioning, the processing unit may configure the merge candidate list in the order of (F, G, H, I, J, K).

[0764] For example, when the target block is obtained as a result of horizontal partitioning, the processing unit may use multiple pieces of inter-frame prediction information of spatially adjacent blocks located to the left of the target block to configure the merge candidate list. When the target block is obtained as a result of horizontal partitioning, the processing unit may configure the merge candidate list in the order of (A, B, C, D, E, F).

[0765] The processing unit may determine a scheme for configuring the merge candidate list based on both the shape of the target block and the partitioning state of the target block.

[0766] For example, when the height of the target block is greater than the width of the target block or when the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical partitioning, the processing unit may configure the merge candidate list in the order of (J, B, K, A, F). In other cases, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0767] For example, when the height of the target block is greater than the width of the target block or the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical partitioning, the processing unit may configure the merge candidate list in the order of (J, B, A, K, F). In other cases, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0768] For example, when the height of the target block is greater than the width of the target block or when the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical partitioning, the processing unit may configure the merge candidate list in the order of (J, K, B, A, F). When the height of the target block is less than the width of the target block or when the target block is obtained as a result of horizontal partitioning, the processing unit may configure the merge candidate list in the order of (B, A, J, K, F). When the height and width of the target block are equal to each other and when the target block is obtained as a result of quad partitioning, the processing unit may configure the merge candidate list in the order of (B, J, K, A, F).

[0769] For example, when the height of the target block is greater than the width of the target block, or when the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical partitioning, the processing unit may use multiple pieces of inter prediction information of spatially adjacent blocks located above the target block to configure the merge candidate list. When the height of the target block is greater than the width of the target block, or when the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical partitioning, the processing unit may configure the merge candidate list in the order of (F, G, H, I, J, K).

[0770] For example, when the height of the target block is less than the width of the target block or when the target block is obtained as a result of horizontal partitioning, the processing unit may use multiple pieces of inter prediction information of spatially adjacent blocks located to the left of the target block to configure the merge candidate list. When the height of the target block is less than the width of the target block or when the target block is obtained as a result of horizontal partitioning, the processing unit may configure the merge candidate list in the order of (A, B, C, D, E, F).

[0771] The processing unit may determine a scheme for configuring the merge candidate list based on the position of the target block.

[0772] The position of the target block may be the relative position of the target block in the upper layer block. By partitioning the upper layer block, multiple sub-blocks may be generated, and the target block may be one of the multiple sub-blocks. The position of the target block may be the position of the target block in the upper layer block, or the position of the target block among the multiple sub-blocks. The partitioning may be binary partitioning or quad partitioning.

[0773] For example, in Figure 23 the processing unit may apply a unified configuration method to the merge candidate list for the first partition CU and apply an adaptive configuration method to the merge candidate list for the second partition CU.

[0774] The processing unit may determine a scheme for configuring the merge candidate list based on the presence or absence of combined inter-frame prediction information. In the configuration of the merge candidate list, when there is combined inter-frame prediction information, the processing unit may adjust the priority of the combined inter-frame prediction information. Here, the priority may refer to the position of the combined inter-frame prediction information in the merge candidate list, the index of the combined inter-frame prediction information, or the order of adding the combined inter-frame prediction information to the merge candidate list.

[0775] For example, when configuring the merge candidate list in the order of (B, J, K, A, F), if block B has combined inter-frame prediction information, the processing unit may configure the merge candidate list in the order of (J, K, A, F, B). In other words, the processing unit may assign the lowest priority to the combined inter-frame prediction information. Optionally, in the configuration of the merge candidate list, the processing unit may add the combined inter-frame prediction information to a position after multiple pieces of inter-frame prediction information in the merge candidate list. In other words, the processing unit may add the combined inter-frame prediction information to the merge candidate list with a lower priority, after multiple pieces of inter-frame prediction information of neighboring blocks.

[0776] For example, when configuring the merge candidate list in the order of (B, J, K, N, F) and block N has temporally neighboring blocks, if block F has combined inter-frame prediction information, the processing unit may configure the merge candidate list in the order of (B, J, K, F, N). In other words, the processing unit may assign a priority lower than that of multiple pieces of inter-frame prediction information of spatially neighboring blocks and higher than that of the inter-frame prediction information of temporally neighboring blocks to the combined inter-frame prediction information.

[0777] Optionally, in the configuration of the merge candidate list, the processing unit may add the combined inter-frame prediction information to a position after the inter-frame prediction information of spatially neighboring blocks and before the inter-frame prediction information of temporally neighboring blocks in the merge candidate list. Optionally, the processing unit may assign a priority higher than that of the inter-frame prediction information of temporally neighboring blocks to the combined inter-frame prediction information. Optionally, in the configuration of the merge candidate list, the processing unit may add the combined inter-frame prediction information to a position before the inter-frame prediction information of temporally neighboring blocks in the merge candidate list.

[0778] For example, when there are multiple pieces of combined inter-frame prediction information, according to the above order determination scheme, the order of the multiple pieces of combined inter-frame prediction information may remain unchanged.

[0779] The processing unit may determine a method for configuring the merge candidate list based on the depth of the target block. The depth of the target block may be at least one of the QT depth based on quadtree (QT) partitioning and the BT depth based on binary tree (BT) partitioning.

[0780] For example, the processing unit may determine a scheme for configuring a merge candidate list based on whether the depth of the target block falls within a specific range.

[0781] For example, when the BT depth of the target block is less than or equal to n, the processing unit may use at least one of the above-mentioned schemes for configuring the merge candidate list based on the shape of the target block and the above-mentioned scheme for configuring the merge candidate list based on the partitioning state of the target block to configure the merge candidate list. For example, n may be 1.

[0782] In an embodiment, when the QT depth of the target block is equal to or greater than n and the BT depth of the target block is less than or equal to m, the processing unit may use at least one of the above-mentioned schemes for configuring the merge candidate list based on the shape of the target block and the above-mentioned scheme for configuring the merge candidate list based on the partitioning state of the target block to configure the merge candidate list. For example, n may be 3, and m may be 1.

[0783] The processing unit may determine a scheme for configuring a merge candidate list based on the position of the target block and the depth of the target block.

[0784] For example, when the BT depth of the target block is less than or equal to n and the target block is a lower block in the sub-blocks generated by horizontal partitioning, the processing unit may use at least one of the above-mentioned scheme for configuring the merge candidate list based on the shape of the target block and the above-mentioned scheme for configuring the merge candidate list based on the partitioning state of the target block to configure the merge candidate list. Here, n may be 1.

[0785] Configuration of the AMVP candidate list

[0786] The processing unit may use the AMVP mode to derive the inter-frame prediction information of the target block.

[0787] The processing unit may configure an AMVP candidate list. The number of AMVP candidates in the AMVP candidate list may be N. N may be a positive integer. For example, the AMVP candidate list may include two AMVP candidates.

[0788] For example, such an AMVP candidate may be inter-frame prediction information or motion information. Optionally, the AMVP candidate may include a motion vector or a reference picture list.

[0789] The processing unit may use one or more of the inter-frame prediction information of spatially adjacent blocks, the inter-frame prediction information of temporally adjacent blocks, and combined inter-frame prediction information to configure the AMVP candidate list.

[0790] The processing unit may derive one AMVP candidate from the inter - prediction information of neighboring blocks to the left of the target block and may derive one AMVP candidate from the inter - prediction information of neighboring blocks above the target block. When the AMVP candidate list is not filled with candidates, the processing unit may derive additional AMVP candidates from the inter - prediction information of temporally neighboring blocks.

[0791] The processing unit may use the inter - prediction information of multiple neighboring blocks in a specific order of neighboring blocks to derive AMVP candidates and may add the derived AMVP candidates to the AMVP candidate list. Here, the AMVP candidate list may be configured differently according to the order of neighboring blocks, and the prediction efficiency and coding efficiency of using the AMVP candidate list for encoding and decoding may vary according to the order of neighboring blocks.

[0792] The processing unit may derive AMVP candidates in a specific order of left - hand neighboring blocks. Here, deriving AMVP candidates in a specific order of left - hand neighboring blocks may mean using the inter - prediction information of multiple neighboring blocks selected in a specific order to derive AMVP candidates.

[0793] In the derivation of AMVP candidates in a specific order of neighboring blocks, when the inter - prediction information of the neighboring block at the current order position is available, the processing unit may use the inter - prediction information of the neighboring block at the current order position to derive AMVP candidates. When the inter - prediction information of the neighboring block at the current order position is not available, the processing unit may use the inter - prediction information of the neighboring block at a subsequent order position to derive AMVP candidates. In other words, the processing unit may use the inter - prediction information of the previous neighboring block that first appears and has available inter - prediction information among the neighboring blocks to derive AMVP candidates.

[0794] The fact that the inter - prediction information of the neighboring block is not available may indicate that at least one of the following cases 1) to 3) is satisfied.

[0795] 1) The case where there is no inter - prediction information of the neighboring block

[0796] 2) The case where the neighboring block and the target block are included in different strips, parallel blocks, or pictures

[0797] 3) The case where the AMVP candidate derived using the inter - prediction information is the same as another AMVP candidate already included in the AMVP list, that is, the case where the AMVP candidate derived using the inter - prediction information is a duplicate AMVP candidate

[0798] For example, when deriving AMVP candidates using the left neighboring block, the processing unit may derive the AMVP candidates in the order of (A, B). In other words, when the inter-frame prediction information of block A is available, the AMVP candidate derived from the inter-frame prediction information of block A may be included in the AMVP list, and when the inter-frame prediction information of block A is not available and the inter-frame prediction information of block B is available, the AMVP candidate derived from the inter-frame prediction information of block B may be included in the AMVP list.

[0799] For example, when deriving AMVP candidates using the left neighboring block, the processing unit may derive the AMVP candidates in the order of (B, A).

[0800] For example, when deriving AMVP candidates using the left neighboring block, the processing unit may derive the AMVP candidates in the order of (A, B, C, D, E).

[0801] The processing unit may derive AMVP candidates in a specific order of the upper neighboring blocks.

[0802] For example, when deriving AMVP candidates using the upper neighboring block, the processing unit may derive the AMVP candidates in the order of (K, J, F).

[0803] For example, when deriving AMVP candidates using the upper neighboring block, the processing unit may derive the AMVP candidates in the order of (K, F, J).

[0804] When the inter-frame prediction information of the neighboring block does not exist or is not available, in order to derive AMVP candidates, the processing unit may use the combined inter-frame prediction information instead of the inter-frame prediction information of the neighboring block.

[0805] When using the inter-frame prediction information of the spatial neighboring block to configure the AMVP candidate list, the processing unit may determine the scheme for configuring the AMVP candidate list based on the shape of the target block.

[0806] For example, when the height of the target block is greater than the width of the target block, the processing unit may use at least one of the following schemes 1) to 3) to configure the AMVP list.

[0807] 1) The processing unit may derive one AMVP candidate from the inter-frame prediction information of the upper neighboring block of the target block, and then may derive one AMVP candidate from the inter-frame prediction information of the left neighboring block of the target block.

[0808] 2) The processing unit may configure the AMVP candidate list only using the multiple inter-frame prediction information of the upper neighboring block of the target block.

[0809] 3) The processing unit may configure the AMVP candidate list in the order of (J, B, K, A, F) or (J, B, A, K, F).

[0810] For example, when the height of the target block is less than the width of the target block, the processing unit may use at least one of the following schemes 4) to 6) to configure the AMVP list.

[0811] 4) The processing unit may derive an AMVP candidate from the inter-frame prediction information of the left neighboring block of the target block, and subsequently may derive an AMVP candidate from the inter-frame prediction information of the upper neighboring block of the target block.

[0812] 5) The processing unit may configure the AMVP candidate list only using the multiple inter-frame prediction information of the left neighboring block of the target block.

[0813] 6) The processing unit may configure the AMVP candidate list in the order of (B, J, K, A, F) or in the order of (B, A, J, K, F).

[0814] For example, when the height of the target block is equal to the width of the target block, the processing unit may use at least one of the following schemes 7) and 8) to configure the AMVP candidate list.

[0815] 7) The processing unit may use at least one of the above schemes 1) to 6) to configure the AMVP candidate list.

[0816] 8) The processing unit may configure the AMVP candidate list using the multiple inter-frame prediction information of the left neighboring block of the target block and the multiple inter-frame prediction information of the upper neighboring block of the target block.

[0817] The processing unit may determine the scheme for configuring the AMVP candidate list based on the partitioning state of the target block.

[0818] For example, when the target block is obtained as a result of vertical partitioning, the processing unit may derive an AMVP candidate from the inter-frame prediction information of the upper neighboring block of the target block, and subsequently may derive an AMVP candidate from the inter-frame prediction information of the left neighboring block of the target block.

[0819] For example, when the target block is obtained as a result of horizontal partitioning, the processing unit may derive an AMVP candidate from the inter-frame prediction information of the left neighboring block of the target block, and subsequently may derive an AMVP candidate from the inter-frame prediction information of the upper neighboring block of the target block.

[0820] For example, the processing unit may use the above scheme for determining the merge candidate list based on the partitioning state of the target block to configure the AMVP candidate list.

[0821] The processing unit may determine the scheme for configuring the AMVP candidate list based on both the shape of the target block and the partitioning state of the target block.

[0822] For example, when the height of the target block is greater than the width of the target block or the height and width of the target block are equal to each other, and when the target block is obtained as a result of vertical partitioning, the processing unit may derive an AMVP candidate from the inter-frame prediction information for the upper neighboring block of the target block, and may then derive an AMVP candidate from the inter-frame prediction information for the left neighboring block or the upper neighboring block of the target block.

[0823] For example, when the height of the target block is less than the width of the target block or the height and width of the target block are equal to each other, and when the target block is obtained as a result of horizontal partitioning, the processing unit may derive an AMVP candidate from the inter-frame prediction information for the left neighboring block of the target block, and may then derive an AMVP candidate from the inter-frame prediction information for the left neighboring block or the upper neighboring block of the target block.

[0824] The processing unit may determine a scheme for configuring the AMVP candidate list based on the position of the target block.

[0825] The position of the target block may be the relative position of the target block in the upper-level block. By partitioning the upper-level block, a plurality of sub-blocks may be generated, and the target block may be one of the plurality of sub-blocks. The position of the target block may be the position of the target block in the upper-level block, or the position of the target block among the plurality of sub-blocks. The partitioning may be binary partitioning or quaternary partitioning.

[0826] For example, in Figure 23 , the processing unit may apply a unified configuration method to the AMVP candidate list for the first partition CU, and apply an adaptive configuration method to the AMVP candidate list for the second partition CU.

[0827] The processing unit may determine a method for configuring the AMVP candidate list based on the depth of the target block. The depth of the target block may be at least one of the QT depth based on quadtree (QT) partitioning and the BT depth based on binary tree (BT) partitioning.

[0828] For example, the processing unit may determine a scheme for configuring the AMVP candidate list based on whether the depth of the target block falls within a specific range.

[0829] For example, when the BT depth of the target block is less than or equal to n, the processing unit may use at least one of the above-mentioned schemes for configuring the AMVP candidate list based on the shape of the target block and the above-mentioned scheme for configuring the AMVP candidate list based on the partitioning state of the target block to configure the AMVP candidate list. For example, n may be 1.

[0830] For example, when the QT depth of the target block is equal to or greater than n and the BT depth of the target block is less than or equal to m, the processing unit may use at least one of the above-described schemes for configuring the AMVP candidate list based on the shape of the target block and the above-described scheme for configuring the AMVP candidate list based on the partitioning state of the target block to configure the AMVP candidate list. For example, n may be 3, and m may be 1.

[0831] The processing unit may determine a scheme for configuring the AMVP candidate list based on the position and depth of the target block.

[0832] For example, when the BT depth of the target block is less than or equal to n and the target block is the lower block in the sub-blocks generated by horizontal partitioning, the processing unit may use at least one of the above-described schemes for configuring the AMVP candidate list based on the shape of the target block and the above-described scheme for configuring the AMVP candidate list based on the partitioning state of the target block to configure the AMVP candidate list. Here, n may be 1.

[0833] The details described in the above configurations of the merge candidate list and the AMVP candidate list may be applied to each other. For example, the features described regarding the derivation and addition of one of the merge candidates and the AMVP candidates may also be applied to the derivation and addition of other candidates. Repeated descriptions will be omitted here.

[0834] Using sub-blocks to derive inter-frame prediction information for the target block

[0835] When deriving the inter-frame prediction information of the target block, the processing unit may use the inter-frame prediction information of the sub-blocks of the target block. In other words, the processing unit may use the inter-frame prediction information corresponding to the unit of the sub-blocks when deriving the inter-frame prediction information of the target block.

[0836] The processing unit may divide the target block into multiple sub-blocks and may derive the inter-frame prediction information of each of the multiple sub-blocks. For example, the processing unit may divide the target block into N sub-blocks and may derive N pieces of inter-frame prediction information for the N sub-blocks. N may be a positive integer.

[0837] Figure 27 Shows the sub-blocks of the temporally neighboring block and the sub-blocks of the target block according to the example.

[0838] When the inter-frame prediction information of the temporally neighboring block is available, the processing unit may partition the temporally neighboring block into multiple sub-blocks and may use the multiple pieces of inter-frame prediction information of the multiple sub-blocks of the temporally neighboring block to derive the multiple pieces of inter-frame prediction information of the multiple sub-blocks of the target block.

[0839] The processing unit may use multiple pieces of inter-frame prediction information of sub-blocks of a temporally neighboring block to derive multiple pieces of inter-frame prediction information of sub-blocks of a target block. Here, the positions of the sub-blocks of the target block within the target block and the positions of the sub-blocks of the temporally neighboring block within the temporally neighboring block may be the same as each other.

[0840] For example, the processing unit may partition Figure 16 the temporally neighboring block N shown in into 4×4 temporal sub-blocks, and may use the inter-frame prediction information of each temporal sub-block to derive multiple pieces of inter-frame prediction information of 4×4 sub-blocks of the target block.

[0841] For example, the processing unit may partition Figure 16 the temporally neighboring block N shown in into 2N×N temporal sub-blocks, and may use the inter-frame prediction information of each temporal sub-block to derive multiple pieces of inter-frame prediction information of 2N×N sub-blocks of the target block.

[0842] Figure 28 Shows the spatial neighboring blocks of a target block and the sub-blocks of the target block according to an example.

[0843] In Figure 28 the sub-blocks are indicated by capital letters "A" to "P", and the spatial neighboring blocks are indicated by lowercase letters "a" to "h".

[0844] The processing unit may partition the target block into multiple sub-blocks, and may use multiple pieces of inter-frame prediction information of the spatial neighboring blocks of the sub-blocks of the target block to derive multiple pieces of inter-frame prediction information of the sub-blocks of the target block.

[0845] Here, the spatial neighboring blocks of the sub-blocks of the target block may include 1) additional sub-blocks adjacent to the corresponding sub-blocks of the target block and 2) blocks adjacent to the corresponding sub-blocks of the target block and that are also spatial neighboring blocks of the target block. In addition, the spatial neighboring blocks may be blocks that have been encoded and / or decoded before the corresponding sub-blocks are encoded and / or decoded.

[0846] The processing unit may use multiple pieces of inter-frame prediction information of the spatial neighboring blocks of a sub-block to derive the inter-frame prediction information of the sub-block. The processing unit may use multiple pieces of inter-frame prediction information of multiple spatial neighboring blocks of a sub-block to derive the inter-frame prediction information of the sub-block. For example, the inter-frame prediction information of a sub-block may be the average of multiple pieces of inter-frame prediction information of multiple spatial neighboring blocks of the sub-block.

[0847] For example, the processing unit may use the following to derive the inter-frame prediction information of sub-block A: 1) the inter-frame prediction information of spatial neighboring block d, 2) the inter-frame prediction information of spatial neighboring block e, or 3) the average of multiple pieces of inter-frame prediction information of spatial neighboring block d and spatial neighboring block e.

[0848] For example, the processing unit may derive the inter-frame prediction information of sub-block K by using the average value of the inter-frame prediction information of one or more of blocks J, F, G, and H, which are spatially adjacent blocks of sub-block K.

[0849] When deriving the inter-frame prediction information of a sub-block of a target block, the processing unit may use both the temporal adjacent blocks of the sub-block of the target block and the spatial adjacent blocks of the sub-block of the target block.

[0850] Using bilateral matching to derive inter-frame prediction information

[0851] Figure 29 Shows the derivation of inter-frame prediction information using bilateral matching according to an example.

[0852] The processing unit may use bilateral matching to derive the inter-frame prediction information.

[0853] When performing bilateral matching, the processing unit may configure an initial motion vector candidate list for the target block and may use at least one of one or more initial motion vector candidates included in the configured initial motion vector candidate list as the initial motion vector.

[0854] For example, the processing unit may use the AMVP mode to configure the initial motion vector candidate list for the target block. The AMVP candidates in the AMVP candidate list in the AMVP mode may be one or more initial motion vector candidates in the initial motion vector candidate list. The processing unit may add the AMVP candidates included in the AMVP candidate list to the initial motion vector candidate list.

[0855] For example, the processing unit may use the merge mode to configure the initial motion vector candidate list for the target block. The merge candidates in the merge candidate list in the merge mode may be one or more initial motion vector candidates in the initial motion vector candidate list. The processing unit may add the merge candidates in the merge mode to the initial motion vector candidate list.

[0856] For example, the processing unit may configure the frame rate up-conversion (FRUC) unidirectional motion vector for the target block as the initial motion vector candidate list. The processing unit may add the FRUC unidirectional motion vector for the target block to the initial motion vector candidate list.

[0857] For example, the processing unit may configure the motion vectors of the adjacent blocks of the target block as the initial motion vector candidate list. The processing unit may add the motion vectors of the adjacent blocks of the target block to the initial motion vector candidate list.

[0858] For example, the processing unit may configure the combination of the above motion vectors as an initial motion vector candidate list. The number of combinations of motion vectors may be N or more. N may be a positive integer. The processing unit may add the combination of the above motion vectors to the initial motion vector candidate list.

[0859] For example, when configuring the initial motion vector candidate list, the processing unit may use the motion vectors for at least one of the direction of reference picture list L0 and the direction of reference picture list L1. The processing unit may add the motion vectors for at least one of the direction of reference picture list L0 and the direction of reference picture list L1 to the initial motion vector candidate list.

[0860] When performing bilateral matching, the processing unit may derive an initial motion vector for the target block. The processing unit may use the initial motion vector candidate list to derive the initial motion vector.

[0861] When performing bilateral matching, the processing unit may use the initial motion vector candidate list to derive a bidirectional motion vector that makes the initial motion vector indication block and the relative block best match each other.

[0862] The initial motion vector indication block may be the block indicated by the initial motion vector. The relative block may be a block that exists on the same trajectory as the initial motion vector indication block in the direction opposite to the direction of the initial motion vector indication block. In other words, the direction of the initial motion vector indication block and the direction of the relative block may be opposite to each other, and the trajectory of the initial motion vector indication block and the trajectory of the relative block may be the same as each other.

[0863] For example, as Figure 29 shown, the processing unit may perform bilateral matching on the target block 2911 in the target picture 2910. When the motion vector existing in the initial motion vector candidate list is MV0 in the reference picture reference 02920, the processing unit may derive a motion vector MV1, where 1) the motion vector MV1 exists in the reference picture reference 1 2930 that exists in the direction opposite to the direction of MV0, 2) the motion vector MV1 exists on the same trajectory as MV0, and 3) the motion vector MV1 indicates the block 2931 that best matches the block 2921 indicated by MV0.

[0864] In other words, when the processing unit uses the initial motion vector for the target block to derive the motion vector MV0 and determines the motion vector MV1 based on MV0, 1) the direction of MV1 may be opposite to the direction of MV0, and the motion trajectory of MV1 may be the same as the motion trajectory of MV0.

[0865] The processing unit may improve the initial motion vector.

[0866] For example, the processing unit may search for blocks adjacent to the block indicated by the derived MV0, and may also search for blocks adjacent to the block indicated by the derived MV1. The processing unit may refine the initial motion vector such that the block with the best match among the blocks adjacent to the block indicated by MV0 and the blocks adjacent to the block indicated by MV1 is indicated.

[0867] When performing bilateral matching, the processing unit may derive inter-frame prediction information based on sub-blocks. The inter-frame prediction information may include motion information and / or a motion vector.

[0868] When deriving motion information based on sub-blocks, the processing unit may use the above-described scheme for deriving an initial motion vector for a block to derive an initial motion vector for the sub-blocks.

[0869] When performing bilateral matching, the processing unit may define a degree of match between blocks. That is, the processing unit may use one of a specifically defined set of schemes when determining the degree of match between blocks.

[0870] For example, the processing unit may determine that two blocks best match each other when the sum of absolute differences (SAD) between the two blocks is minimized. That is, the processing unit may determine that the smaller the SAD between two blocks, the better the two blocks match each other.

[0871] For example, the processing unit may determine that two blocks best match each other when the sum of absolute transform differences (SATD) between the two blocks is minimized. That is, the processing unit may determine that the smaller the SATD between two blocks, the better the two blocks match each other.

[0872] Using template matching to derive inter-frame prediction information

[0873] Figure 30 Shows the derivation of inter-frame prediction information using a template matching mode according to an example.

[0874] In Figure 30 a target block 3011 in a target picture 3010 is shown.

[0875] The processing unit may use template matching to derive inter-frame prediction information.

[0876] When performing template matching, the processing unit may use neighboring blocks of the target block as templates. The size and position of the templates may be set based on a predefined scheme.

[0877] For example, the processing unit may use a neighboring block 3013 adjacent to the upper side of the target block 3011 as a template.

[0878] For example, the processing unit may use a neighboring block 3012 adjacent to the left side of the target block 3011 as a template.

[0879] For example, the processing unit may use the neighboring block 3013 adjacent to the upper side of the target block 3011 or the neighboring block 3012 adjacent to the left side of the target block 3011 as a template.

[0880] In an embodiment, when performing template matching, the processing unit may use the template to search for a block in the reference picture. In Figure 30 "Reference 0 (3020)" may be the reference picture.

[0881] When performing template matching, the processing unit may use the template to derive inter-frame prediction information.

[0882] The processing unit may search for a block corresponding to the template in the reference picture. The shape and size of the block corresponding to the template may be the same as the shape and size of the template. The block corresponding to the template may be the block in the reference picture that best matches the template.

[0883] The processing unit may derive a motion vector indicating the block corresponding to the template. In other words, the processing unit may derive a motion vector indicating the block in the reference picture that best matches the template.

[0884] For example, in Figure 30 a block 3022 corresponding to the template found for the neighboring block 3013 adjacent to the upper side of the target block and a block 3021 corresponding to the template found for the neighboring block 3012 adjacent to the left side of the target block are depicted.

[0885] The processing unit may use the motion vector of the block corresponding to the template to derive the motion vector of the target block. The processing unit may set the motion vector of the block corresponding to the template as the motion vector of the target block.

[0886] Here, the direction of the reference picture in which the block is found and the direction of the reference picture indicated by the motion vector of the target block may be opposite to each other.

[0887] The processing unit may refine the motion vector. For example, the processing unit may search for a block adjacent to the block indicated by the current template and the derived motion vector, and may refine the motion vector such that the motion vector indicates the block in the found neighboring blocks that best matches the template. In other words, the refined motion vector may be the motion vector for the block in the neighboring blocks adjacent to the block indicated by the derived motion vector that best matches the template.

[0888] When performing template matching, the processing unit may derive inter-frame prediction information based on sub-blocks. The inter-frame prediction information may include motion information and a motion vector.

[0889] When deriving motion information based on sub-blocks, the processing unit may use the above-described scheme for deriving the motion vector for a block to derive the motion vector for a sub-block.

[0890] When performing template matching, the processing unit may define the degree of match between blocks. That is, the processing unit may use one of the specifically defined schemes when determining the degree of match between blocks.

[0891] For example, the processing unit may determine that when the SAD between two blocks is the smallest, the two blocks best match each other. That is, the processing unit may determine that as the SAD between two blocks becomes smaller, the two blocks match each other better.

[0892] For example, the processing unit may determine that when the SATD between two blocks is the smallest, the two blocks best match each other. That is, the processing unit may determine that as the SATD between two blocks becomes smaller, the two blocks match each other better.

[0893] Motion compensation and motion correction for inter-frame prediction

[0894] In the inter - frame prediction for a target block, the processing unit may use at least one of motion compensation, IC, OBMC, BIO, affine - space motion prediction and motion compensation, and motion vector correction in the decoding device 1300 to perform inter - frame prediction.

[0895] The processing unit may perform motion compensation to perform inter - frame prediction for the target block.

[0896] The processing unit may use the derived inter - frame prediction information to generate a predicted block for the target block. Motion compensation may be uni - directional motion compensation or bi - directional motion compensation.

[0897] In an example, the processing unit may perform motion compensation using a block in a picture existing in a reference picture list L0.

[0898] In an example, the processing unit may perform motion compensation by combining multiple blocks in a picture existing in a reference picture list L0.

[0899] In an example, the processing unit may perform motion compensation by combining multiple blocks in multiple pictures existing in a reference picture list L0.

[0900] In an example, the processing unit may perform motion compensation by combining a block in a picture existing in a reference picture list L0 with a block in a picture existing in a reference picture list L1.

[0901] In an example, the processing unit may perform motion compensation by combining multiple blocks in a picture existing in a reference picture list L0 with multiple blocks in a picture existing in a reference picture list L1.

[0902] In an example, the processing unit may perform motion compensation by combining a plurality of blocks in a plurality of pictures present in a reference picture list L0 with a plurality of blocks in a plurality of pictures present in a reference picture list L1.

[0903] The processing unit may perform illumination compensation (IC) to perform inter-frame prediction for a target block.

[0904] When performing motion compensation, the processing unit may compensate for a luminance change and / or an illumination change between a reference picture including a reference block for motion compensation and a target picture including a target block.

[0905] For example, the processing unit may approximate a change between neighboring samples of a target block and neighboring samples of a reference block as N or more linear models, and may perform illumination compensation by applying the linear models to a motion compensation block. N may be a positive integer.

[0906] Figure 31 An application of OBMC according to an example is shown.

[0907] The processing unit may perform OBMC to perform inter-frame prediction for a target block.

[0908] The processing unit may generate a prediction block by combining a first block and a second block. The first block may be a block generated by compensation using inter-frame prediction information of a target block. The second block may be one or more blocks generated by compensation using one or more pieces of inter-frame prediction information of one or more neighboring blocks adjacent to the target block. Here, one or more neighboring blocks adjacent to the target block may include a left neighboring block adjacent to the left side of the target block, a right neighboring block adjacent to the right side of the target block, an upper neighboring block adjacent to the upper side of the target block, and a lower neighboring block adjacent to the lower side of the target block.

[0909] The processing unit may perform OBMC for each sub-block of a target block.

[0910] The processing unit may generate a prediction block by combining a first block and a second block. The first block may be a block generated by compensation using inter-frame prediction information of a sub-block of a target block. The second block may indicate one or more blocks generated by compensation using one or more pieces of inter-frame prediction information of one or more neighboring sub-blocks adjacent to the sub-block of the target block. The inter-frame prediction information may include a motion vector. Here, one or more neighboring sub-blocks adjacent to the corresponding sub-block of the target block may include a left neighboring sub-block adjacent to the left side of the sub-block of the target block, a right neighboring sub-block adjacent to the right side of the sub-block of the target block, an upper neighboring sub-block adjacent to the upper side of the sub-block of the target block, and a lower neighboring sub-block adjacent to the lower side of the sub-block of the target block.

[0911] The processing unit may perform OBMC only on specific sub - blocks among the sub - blocks of the target block.

[0912] For example, the processing unit may perform OBMC only on the sub - blocks adjacent to the internal boundary of the target block.

[0913] In the example, the processing unit may perform OBMC on all sub - blocks of the target block.

[0914] For example, the processing unit may perform OBMC only on the sub - blocks adjacent to the internal left boundary of the target block.

[0915] In the example, the processing unit may perform OBMC only on the sub - blocks adjacent to the internal right boundary of the target block.

[0916] In the example, the processing unit may perform OBMC only on the sub - blocks adjacent to the internal upper boundary of the target block.

[0917] In the example, the processing unit may perform OBMC only on the sub - blocks adjacent to the internal lower boundary of the target block.

[0918] In Figure 31 an example, it shows an example where a CU as a target block is divided into PU1 and PU2, and OBMC is applied to the following sub - blocks among the sub - blocks of the CU: 1) sub - blocks adjacent to the upper boundary, 2) sub - blocks adjacent to the left boundary, and 3) sub - blocks adjacent to the boundary between PUs.

[0919] Figure 32 Shows sub - PUs in the ATMVP mode according to the example.

[0920] The target block may be divided into multiple sub - blocks. The target block may be a target CU, and the sub - blocks may be sub - PUs. The processing unit may perform OBMC on all sub - blocks of the target CU.

[0921] The processing unit may generate a prediction block by combining a first block and a second block. The first block may be a block generated through compensation using the inter - frame prediction information of each sub - block of the target block. The second block may be a block generated through compensation using multiple inter - frame prediction information of neighboring sub - blocks adjacent to the sub - blocks of the target block. The inter - frame prediction information may be a motion vector. Here, the neighboring sub - blocks adjacent to the corresponding sub - blocks of the target block may include left - hand neighboring sub - blocks adjacent to the left side of the sub - blocks of the target block, right - hand neighboring sub - blocks adjacent to the right side of the sub - blocks of the target block, upper - hand neighboring sub - blocks adjacent to the upper side of the sub - blocks of the target block, and lower - hand neighboring sub - blocks adjacent to the lower side of the sub - blocks of the target block.

[0922] The processing unit may perform affine - space motion prediction and compensation to perform inter - frame prediction for the target block.

[0923] In an example, the processing unit may generate motion vectors for respective pixels in a target block by applying an affine transformation equation to a first motion vector at the upper left position of the target block and a second motion vector at the upper right position of the target block, and may perform motion compensation using the generated motion vectors.

[0924] In an example, the processing unit may generate motion vectors for respective sub - blocks in a target block by applying an affine transformation equation to a first motion vector at the upper left position of the target block and a second motion vector at the upper right position of the target block, and may perform motion compensation using the generated motion vectors.

[0925] To provide the first motion vector and the second motion vector, at least one of the following schemes 1) to 3) may be used.

[0926] 1) The first motion vector and the second motion vector may be sent from the encoding device 1200 to the decoding device 1300 (through a bitstream).

[0927] 2) For each of the first motion vector and the second motion vector, the difference between its corresponding motion vector and an adjacent motion vector may be sent from the encoding device 1200 to the decoding device 1300.

[0928] 3) The first motion vector and the second motion vector may be derived using the affine motion vectors of adjacent blocks of the target block without sending.

[0929] The processing unit may perform BIO to perform inter - frame prediction for the target block.

[0930] The processing unit may use the optical flow of unidirectional blocks to derive the motion vector of the target block.

[0931] The processing unit may use bidirectional optical flow to derive the motion vector of the target block, where the bidirectional optical flow uses blocks existing in the picture before the target picture and blocks existing in the picture after the target picture in the display order. Here, the blocks existing in the picture before the target picture and the blocks existing in the picture after the target picture may be blocks with opposite motions and similar to the target block.

[0932] The processing unit may perform motion vector correction on the decoding device 1300 to perform inter - frame prediction for the target block. The processing unit may use the motion vector sent to the decoding device 1300 to correct the motion vector of the target block.

[0933] Figure 33 is a flowchart showing a target block prediction method and a bitstream generation method according to an embodiment.

[0934] The target block prediction method and bitstream generation method according to the present embodiment may be executed by the encoding device 1200. This embodiment may be part of a target block encoding method or a video encoding method.

[0935] In step 3310, the processing unit 1210 may derive inter prediction information. Step 3310 may correspond to step 1410 described above with reference to Figure 14 description.

[0936] In step 3320, the processing unit 1210 may perform inter prediction using the derived inter prediction information. Step 3320 may correspond to step 1420 described above with reference to Figure 14 description.

[0937] In step 3330, the processing unit 1210 may generate a bitstream.

[0938] The bitstream may include information about the encoded target block. For example, the information about the encoded target block may include the transformed and quantized coefficients of the target block.

[0939] The bitstream may include information for deriving inter prediction information and information for inter prediction.

[0940] In an example, the bitstream may include an indicator indicating a method for deriving inter prediction information.

[0941] For example, the bitstream may include an index indicating one of the candidates present in the list of the encoding device 1200 or the decoding device 1300.

[0942] The processing unit 1210 may perform entropy coding on the information for deriving inter prediction information and the information for inter prediction, and may generate a bitstream including a plurality of entropy-coded information.

[0943] The processing unit 1210 may store the generated bitstream in the memory 1240. Optionally, the communication unit 1220 may send the bitstream to the decoding device 1300.

[0944] Figure 34 is a flowchart showing a target block prediction method using a bitstream according to an embodiment.

[0945] The target block prediction method using a bitstream according to the present embodiment may be executed by the decoding device 1300. This embodiment may be part of a target block decoding method or a video decoding method.

[0946] In step 3410, the communication unit 1320 may obtain a bitstream. The communication unit 1320 may receive the bitstream from the encoding device 1200.

[0947] The bitstream may include information about an encoded target block. For example, the information about the encoded target block may include the transformed and quantized coefficients of the target block.

[0948] The bitstream may include information for deriving inter-prediction information and information for inter-prediction.

[0949] For example, the bitstream may include an indicator indicating a method for deriving inter-prediction information.

[0950] For example, the bitstream may include an index indicating one candidate among candidates present in a list in each of the encoding device 1200 and the decoding device 1300.

[0951] The processing unit 1310 may store the acquired bitstream in the memory 1340.

[0952] The processing unit 1310 may obtain information for deriving inter-prediction information and information for inter-prediction by performing entropy decoding on a plurality of entropy-encoded information in the bitstream.

[0953] In step 3420, the processing unit 1310 may derive inter-prediction information. Step 3420 may correspond to step 1410 described above with reference to Figure 14 description.

[0954] In step 3430, the processing unit 1310 may perform inter-prediction using the derived inter-prediction information. Step 3430 may correspond to step 1420 described above with reference to Figure 14 description.

[0955] In the above-described embodiments, although the method has been described based on a flowchart as a series of steps or units, the present disclosure is not limited to the order of the steps, and some steps may be performed in an order different from the order of the steps described or may be performed simultaneously with other steps. Further, those skilled in the art will understand that the steps shown in the flowchart are not exclusive and may further include other steps, or one or more steps in the flowchart may be deleted without departing from the scope of the present disclosure.

[0956] Embodiments according to the present disclosure described above can be implemented as programs executable by various computer devices and can be recorded on a computer-readable storage medium. The computer-readable storage medium may include program instructions, data files, and data structures, either individually or in combination. The program instructions recorded on the storage medium may be specifically designed or configured for the present disclosure, or may be known or available to those of ordinary skill in the computer software field. Examples of computer-readable storage media may include all types of hardware devices specifically configured to record and execute program instructions, such as magnetic media (such as hard disks, floppy disks, and magnetic tapes), optical media (such as compact disc (CD)-ROMs and digital versatile discs (DVDs)), magneto-optical media (such as floppy optical discs, ROMs, RAMs, and flash memories). Examples of program instructions include machine code (such as code created by a compiler) and high-level language code that can be executed by a computer using an interpreter. The hardware device may be configured to operate as one or more software modules to perform the operations of the present disclosure, and vice versa.

[0957] As described above, although the present disclosure has been described based on specific details (such as detailed components and a limited number of embodiments and drawings), the specific details are provided only for easy understanding of the present disclosure. The present disclosure is not limited to these embodiments, and those skilled in the art will practice various changes and modifications based on the above description.

[0958] Therefore, it should be understood that the spirit of the present embodiment is not limited to the above embodiments, and the appended claims and their equivalents and modifications thereto fall within the scope of the present disclosure.

Claims

1. A video decoding method, comprising: Deriving prediction information for a target block; And Performing prediction on the target block using the prediction information.

2. The video decoding method according to claim 1, wherein A list including a plurality of candidates is configured for the target block, The list is used to perform the prediction, The plurality of candidates are generated based on a plurality of neighboring blocks adjacent to the target block, and The plurality of neighboring blocks include the leftmost block among the blocks adjacent to the upper side of the target block and the uppermost block among the blocks adjacent to the left side of the target block.

3. The video decoding method according to claim 1, wherein A list including a plurality of candidates is configured for the target block, The list is used to perform the prediction, The plurality of candidates are generated by adding the prediction information of a plurality of neighboring blocks adjacent to the target block to the list according to a specific order, According to the specific order, the prediction information of the upper neighboring block is added to the list first, and the prediction information of the left neighboring block is added to the list second, The upper neighboring block is the rightmost block among the blocks adjacent to the upper side of the target block, and The left neighboring block is the lowermost block among the blocks adjacent to the left side of the target block.

4. The video decoding method according to claim 1, wherein A list including a plurality of candidates is configured for the target block, The list is used to perform the prediction, One candidate in the list is generated based on the prediction information of a plurality of neighboring blocks adjacent to the target block.

5. The video decoding method according to claim 4, wherein The one candidate is the average of the prediction information of two neighboring blocks.

6. The video decoding method according to claim 4, wherein The one candidate is generated based on the prediction information of three neighboring blocks.

7. The video decoding method according to claim 1, wherein The prediction information includes a motion vector, and The motion vector is generated by improving an initial motion vector.

8. The video decoding method according to claim 7, wherein The improvement is performed using the sum of absolute differences SAD between two reference blocks.

9. According to the video decoding method of claim 1, wherein A first list including a plurality of first candidates is configured for the target block, The first list is used to perform the prediction, Candidates in a second list including a plurality of second candidates are used to add candidates to the first list.

10. A video encoding method, comprising: Deriving prediction information for a target block; And Performing prediction on the target block using the prediction information.

11. According to the video encoding method of claim 10, wherein A list including a plurality of candidates is configured for the target block, The list is used to perform the prediction, The plurality of candidates are generated based on a plurality of neighboring blocks adjacent to the target block, and The plurality of neighboring blocks include the leftmost block among the blocks adjacent to the upper side of the target block and the uppermost block among the blocks adjacent to the left side of the target block.

12. The video encoding method according to claim 10, wherein A list including a plurality of candidates is configured for the target block, and the prediction is performed using the list. The plurality of candidates are generated by adding prediction information of a plurality of neighboring blocks adjacent to the target block to the list according to a specific order. According to the specific order, the prediction information of the upper neighboring block is added to the list first, and the prediction information of the left neighboring block is added to the list second. The upper neighboring block is the rightmost block among the blocks adjacent to the upper side of the target block, and the left neighboring block is the lowermost block among the blocks adjacent to the left side of the target block.

13. The video coding method according to claim 10, wherein A list including a plurality of candidates is configured for the target block, and the prediction is performed using the list. One candidate in the list is generated based on prediction information of a plurality of neighboring blocks adjacent to the target block.

14. The video coding method according to claim 10, wherein the prediction information includes a motion vector, and the motion vector is generated by improving an initial motion vector.

15. The video coding method according to claim 10, wherein A first list including a plurality of first candidates is configured for the target block, and the prediction is performed using the first list. Candidates in a second list including a plurality of second candidates are used to add candidates to the first list.

16. A computer-readable medium for storing a bitstream, wherein, The bitstream is generated by an encoding device through the video coding method of claim 10.

17. A computer-readable medium for storing a bitstream, wherein, The bitstream includes prediction information for a target block; wherein the prediction for the target block is performed using the prediction information.

18. The computer-readable medium according to claim 17, wherein A list including a plurality of candidates is configured for the target block, and the prediction is performed using the list.

19. The computer-readable medium according to claim 17, wherein the prediction information includes a motion vector.

20. A computer-readable medium storing a bitstream generated by a video coding device executing a video coding method, the video coding method comprising: Deriving prediction information for a target block; Performing prediction for the target block using the prediction information; Storing the bitstream including the prediction information in a computer-readable medium.

21. A method for sending a bitstream, the method comprising: Sending a bitstream including prediction information for a target block, wherein the prediction information is information for performing prediction for the target block.

22. The method according to claim 21, wherein A list including a plurality of candidates is configured for the target block, and the prediction is performed using the list. The plurality of candidates are generated based on a plurality of neighboring blocks adjacent to the target block, and the plurality of neighboring blocks include the leftmost block among the blocks adjacent to the upper side of the target block and the uppermost block among the blocks adjacent to the left side of the target block.

23. The method according to claim 21, wherein A list including a plurality of candidates is configured for the target block, and the prediction is performed using the list. The plurality of candidates are generated by adding prediction information of a plurality of neighboring blocks adjacent to the target block to the list according to a specific order, According to the specific order, the prediction information of the upper neighboring block is added to the list first, and the prediction information of the left neighboring block is added to the list second, The upper neighboring block is the rightmost block among the blocks adjacent to the upper side of the target block, and The left neighboring block is the lowermost block among the blocks adjacent to the left side of the target block.

24. The method according to claim 21, wherein, A list including a plurality of candidates is configured for the target block, The prediction is performed using the list, One candidate in the list is generated based on prediction information of a plurality of neighboring blocks adjacent to the target block.

25. The method according to claim 21, wherein, The prediction information includes a motion vector, The motion vector is generated by improving an initial motion vector.

26. The method according to claim 21, wherein, A first list including a plurality of first candidates is configured for the target block, The prediction is performed using the first list, Candidates in a second list including a plurality of second candidates are used to add candidates to the first list.

27. A video decoding method, comprising: Deriving decoding information for a target block; And Performing decoding for the target block using the decoding information.

28. The video decoding method according to claim 27, wherein, A list including a plurality of candidates is configured for the target block, The decoding is performed using the list.

29. The video decoding method according to claim 27, wherein, The decoding information includes a motion vector.

30. A video encoding method, comprising: Deriving encoding information for a target block; And Performing encoding for the target block using the encoding information.

31. The video encoding method according to claim 30, wherein, A list including a plurality of candidates is configured for the target block, The prediction is performed using the list.

32. The video encoding method according to claim 30, wherein, The encoding information includes a motion vector.

33. A computer-readable medium storing a bitstream, wherein, The bitstream is generated by an encoding device through the video encoding method of claim 30.

34. A computer-readable medium storing a bitstream, the bitstream including decoding information for a target block, Among them, Performing decoding for the target block using the decoding information.

35. A computer-readable medium storing a bitstream generated by a video encoding device executing a video encoding method, the video encoding method comprising: Deriving encoding information for a target block; Performing encoding for the target block using the encoding information; And Storing the bitstream including the encoding information in a computer-readable medium.

36. A method for transmitting a bitstream, the method comprising: Transmitting a bitstream including decoding information for a target block, wherein the decoding information is information for performing decoding for the target block.

37. The method according to claim 36, wherein, A list including a plurality of candidates is configured for the target block, Use the said list to perform the said decoding.

38. The method according to claim 36, wherein the decoded information includes a motion vector.