Image decoding method and image encoding method

By generating a comprehensive merge candidate list that includes spatial, temporal, combined, and zero merge candidates, the method enhances the encoding/decoding efficiency of images and simplifies hardware logic, addressing the limitations of conventional motion compensation in the merge mode.

JP7690533B2Active Publication Date: 2025-06-10INTELLECTUAL DISCOVERY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023159358
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-07-12
Filing Date
2023-09-25
Publication Date
2025-06-10
Estimated Expiration
2037-07-12

AI Technical Summary

Technical Problem

Conventional motion compensation using the merge mode is limited by dependencies between temporal and bi-prediction merge candidate derivation processes, leading to reduced encoding efficiency and increased memory access bandwidth.

Method used

The method involves generating a merge candidate list for a current block that includes spatial and temporal merge candidates, combined merge candidates, and zero merge candidates, allowing for parallel derivation of motion information and reducing dependencies between processes.

Benefits of technology

This approach improves the encoding/decoding efficiency of images by increasing the processing amount of the merge mode and simplifying hardware logic, while also reducing memory access bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690533000010
    Figure 0007690533000010
  • Figure 0007690533000011
    Figure 0007690533000011
  • Figure 0007690533000012
    Figure 0007690533000012
Patent Text Reader

Abstract

To provide a technique capable of improving encoding / decoding efficiency of an image.SOLUTION: An image decoding method comprises the steps of: generating a candidate list of current blocks including at least one of space candidates derived from spatial neighboring blocks of the current block and time candidates derived from a corresponding position block of the current block; determining motion information by using the candidate list; generating a plurality of prediction blocks on the basis of the motion information; and generating a final prediction block of the current block by applying weight and offset to the plurality of prediction blocks. The weight is derived by weight index information that specifies the weight to be applied to the current block in a plurality of weights included in a weight set. The weight index information is acquired from a bit stream only when the size of the current block is equal to or greater than a pre-defined value. The weight index information is derived on the basis of index information on the neighboring blocks of the current block.SELECTED DRAWING: Figure 33
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and apparatus for encoding / decoding an image, and more particularly, to a method and apparatus for performing motion compensation using a merge mode.

Background Art

[0002] Recently, there has been an increasing demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, in various application fields. As the image data becomes higher in resolution and quality, the data volume relatively increases compared to conventional image data. Therefore, when transmitting image data using a medium such as a conventional wired or wireless broadband line or storing it using a conventional storage medium, the transmission and storage costs increase. In order to solve the problems arising from such high-resolution and high-quality image data, a high-efficiency image encoding / decoding technology for images with higher resolution and image quality is required.

[0003] As image compression technologies, there are various technologies such as an inter-prediction technology that predicts pixel values included in the current picture from a picture before or after the current picture, an intra-prediction technology that predicts pixel values included in the current picture using pixel information within the current picture, a transformation and quantization technology for compressing the energy of the residual signal, and an entropy encoding technology that assigns short codes to frequently occurring values and long codes to infrequently occurring values. Using such image compression technologies, image data can be effectively compressed and transmitted or stored.

[0004] In conventional motion compensation using the merge mode, only a spatial merge candidate, a temporal merge candidate, a bi-prediction merge candidate, and a zero merge candidate are added to and used in the merge candidate list. However, since only uni-directional prediction and bi-directional prediction are used, there is a limit to the improvement of encoding efficiency.

[0005] In motion compensation using the conventional merge mode, there is a dependency between the temporal merge candidate derivation process and the bi-prediction merge candidate derivation process, which limits the processing amount of the merge mode and makes it difficult to perform each merge candidate derivation process in parallel.

[0006] In motion compensation using the conventional merge mode, since bi-prediction merge candidates are generated using the bi-prediction merge candidate derivation process and used as motion information, there is a drawback that the memory access bandwidth during motion compensation increases compared to single-prediction merge candidates.

[0007] In motion compensation using the conventional merge mode, zero merge candidate derivation is performed differently according to the slice type, which has the drawback of complicating the hardware logic. Also, since bi-prediction zero merge candidates are generated using bi-prediction zero merge candidates and used for motion compensation, there is a drawback that the memory access bandwidth increases.

Summary of the Invention

Problems to be Solved by the Invention

[0008] The present invention can provide a method and apparatus for performing motion compensation using combined merge candidates in order to improve the encoding / decoding efficiency of an image.

[0009] The present invention can provide a method and apparatus for performing motion compensation using one-way prediction, bi-directional prediction, three-way prediction, and four-way prediction in order to improve the encoding / decoding efficiency of an image.

[0010] The present invention provides a method and apparatus for determining motion information by using parallelization of each merge candidate derivation process, removal of dependencies between merge candidate derivation processes, splitting of bi-prediction merge candidates, and derivation of single-prediction zero merge candidates in order to increase the processing amount of the merge mode and simplify the hardware logic.

Means for Solving the Problems

[0011] The image decoding method according to the present invention can include steps of generating a merge candidate list of a current block including at least one of merge candidates corresponding to each of a plurality of reference image lists, determining at least one motion information using the merge candidate list, and generating a predicted block of the current block using the determined at least one motion information.

[0012] In the image decoding method, the merge candidate list can include at least one of a spatial merge candidate derived from spatial neighboring blocks of the current block, a temporal merge candidate derived from corresponding position blocks of the current block, a modified spatial merge candidate derived by modifying the spatial merge candidate, a modified temporal merge candidate derived by modifying the temporal merge candidate, and a merge candidate having a predefined motion information value.

[0013] In the image decoding method, the merge candidate list can further include a combined merge candidate derived using two or more of the spatial merge candidate, the temporal merge candidate, the modified spatial merge candidate, and the modified temporal merge candidate.

[0014] In the image decoding method, the spatial merge candidate can be derived from sub-blocks of neighboring blocks adjacent to the current block, and the temporal merge candidate can be derived from sub-blocks of corresponding position blocks of the current block.

[0015] In the image decoding method, the step of generating a predicted block of the current block using the determined at least one motion information can generate a plurality of temporary predicted blocks according to an inter prediction indicator of the current block, and generate a predicted block of the current block by applying at least one of a weight and an offset to the generated plurality of temporary predicted blocks.

[0016] In the image decoding method, at least one of the weight and the offset can be shared among blocks smaller than a predetermined block size or deeper than a predetermined block depth.

[0017] In the image decoding method, the merge candidate list can be shared among blocks smaller than a predetermined block size or deeper than a predetermined block depth.

[0018] In the image decoding method, when the current block is smaller than a predetermined block size or deeper than a predetermined block depth, the merge candidate list of the current block can be generated based on a higher-level block of the current block having the predetermined block size or the predetermined block depth.

[0019] The image encoding method according to the present invention may include a step of generating a merge candidate list of a current block including at least one of merge candidates corresponding to each of a plurality of reference image lists, a step of determining at least one motion information using the merge candidate list, and a step of generating a predicted block of the current block using the determined at least one motion information.

[0020] In the image encoding method, the merge candidate list may include at least one of a spatial merge candidate derived from a spatially neighboring block of the current block, a temporal merge candidate derived from a corresponding position block of the current block, a modified spatial merge candidate derived by modifying the spatial merge candidate, a modified temporal merge candidate derived by modifying the temporal merge candidate, and a merge candidate having a predefined motion information value.

[0021] In the image encoding method, the merge candidate list may further include a combined merge candidate derived using two or more of the spatial merge candidate, the temporal merge candidate, the modified spatial merge candidate, and the modified temporal merge candidate.

[0022] In the image encoding method, the spatial merge candidate is derived from a lower block of a neighboring block adjacent to the current block, and the temporal merge candidate may be derived from a lower block of a corresponding position block of the current block.

[0023] In the image encoding method, the step of generating a prediction block of the current block using the determined at least one motion information includes generating a plurality of temporary prediction blocks according to an inter prediction indicator of the current block, and at least one of a weight and an offset may be applied to the generated plurality of temporary prediction blocks to generate a prediction block of the current block.

[0024] In the image encoding method, at least one of the weight and the offset can be shared by a block smaller than a predetermined block size or deeper than a predetermined block depth.

[0025] In the image encoding method, the merge candidate list can be shared by a block smaller than a predetermined block size or deeper than a predetermined block depth.

[0026] In the image encoding method, when the current block is smaller than a predetermined block size or deeper than a predetermined block depth, the merge candidate list can be generated based on an upper block of the current block having the predetermined block size or the predetermined block depth.

[0027] An image decoding apparatus according to the present invention may include an inter prediction unit that generates a merge candidate list of a current block including at least one of merge candidates corresponding to a plurality of reference image lists, determines at least one motion information using the merge candidate list, and generates a prediction block of the current block using the determined at least one motion information.

[0028] The image encoding apparatus according to the present invention can include an inter prediction unit that generates a merge candidate list for a current block including at least one of merge candidates corresponding to each of a plurality of reference image lists, determines at least one motion information using the merge candidate list, and generates a predicted block for the current block using the determined at least one motion information.

[0029] A recording medium storing a bitstream according to the present invention can store a bitstream generated by an image encoding method including a step of generating a merge candidate list for a current block including at least one of merge candidates corresponding to each of a plurality of reference image lists, a step of determining at least one motion information using the merge candidate list, and a step of generating a predicted block for the current block using the determined at least one motion information.

Effects of the Invention

[0030] In the present invention, a method and an apparatus for performing motion compensation using combined merge candidates are provided to improve the encoding / decoding efficiency of an image.

[0031] In the present invention, a method and an apparatus for performing motion compensation using uni-directional prediction, bi-directional prediction, tri-directional prediction, and quad-directional prediction are provided to improve the encoding / decoding efficiency of an image.

[0032] In the present invention, a method and an apparatus for performing motion compensation using parallelization of each merge candidate derivation process, removal of dependency between merge candidate derivation processes, splitting of bi-prediction merge candidates, and derivation of single-prediction zero merge candidates are provided for increasing the processing amount of the merge mode and simplifying the hardware logic.

Brief Description of the Drawings

[0033]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

[0034] The present invention can be modified in various ways and can have various embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, and it should be understood that the present invention includes all modifications, equivalents, and alternatives included in the spirit and technical scope of the present invention. Similar reference numerals in the drawings indicate the same or similar functions across various aspects. The shape and size of elements in the drawings may be exaggerated for clearer explanation. The detailed description of the exemplary embodiments described below refers to the accompanying drawings that illustrate specific embodiments. These embodiments are described in sufficient detail for those skilled in the art to implement the embodiments. It should be understood that the various embodiments are different from each other but do not necessarily exclude each other. For example, the specific shapes, structures, and characteristics described herein can be implemented in various embodiments without departing from the spirit and scope of the present invention in relation to one embodiment. Also, it should be understood that the position or arrangement of individual components within each disclosed embodiment can be changed without departing from the spirit and scope of the embodiment. Therefore, the detailed description below is not taken in a limiting sense, and the scope of the exemplary embodiments is limited only by all ranges equivalent to those claimed by their claims and the appended claims if appropriately described.

[0035] In the present invention, terms such as "first" and "second" can be used to describe various components, but these components should not be limited by the above terms. These terms are only used for the purpose of distinguishing one component from another. For example, unless it exceeds the scope of the rights of the present invention, the first component can be named the second component, and similarly, the second component can also be named the first component. The term "and / or" includes a combination of a plurality of related described items or any one of a plurality of related described items.

[0036] When a certain component of the present invention is described as "connected to" or "attached to" another component, it should be understood that it may be directly connected to or attached to the other component, but there may also be another component intervening between them. In contrast, when a certain component is described as "directly connected to" or "directly attached to" another component, it should be understood that there is no other component intervening between them.

[0037] The components shown in the embodiments of the present invention are independently illustrated to show different characteristic functions from each other, and it does not mean that each component consists of separate hardware or one software component unit. That is, each component is included by listing each component for the convenience of explanation. At least two of each component can be combined to form one component, or each component can be divided into a plurality of components to perform functions. Such integrated embodiments and separated embodiments of each component are included in the scope of the rights of the present invention unless they deviate from the essence of the present invention.

[0038] The terms used in the present invention are merely for describing specific embodiments and do not limit the present invention. Singular expressions include plural expressions unless the context clearly indicates a different meaning. In the present invention, terms such as "including" or "having" are used to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the presence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof is not precluded in advance. That is, in the present invention, the description of a specific configuration as "including" does not exclude configurations other than the corresponding configuration, but means that additional configurations may be included within the scope of the implementation of the present invention or the technical idea of the present invention.

[0039] Some components of the present invention may not be essential components for performing essential functions in the present invention, but may be optional components merely for improving performance. The present invention can be implemented by including only the essential components that are indispensable for realizing the essence of the present invention, excluding the components used merely for performance improvement, and a structure including only the essential components excluding the optional components used merely for performance improvement is also included in the scope of the rights of the present invention.

[0040] Hereinafter, embodiments of the present invention will be specifically described with reference to the drawings. When it is determined that a specific description of a related known configuration or function may obscure the gist of this specification in the description of the embodiments of this specification, the detailed description thereof will be omitted, the same reference numerals will be used for the same components in the drawings, and duplicate descriptions of the same components will be omitted.

[0041] Also, hereinafter, an image may sometimes mean one picture constituting a video, and may sometimes indicate the video itself. For example, "encoding and / or decoding of an image" can mean "encoding and / or decoding of a video", and can also mean "encoding and / or decoding of one image among the images constituting a video". Here, a picture can have the same meaning as an image.

[0042] <Term Explanation> Encoder: It can mean a device that performs encoding.

[0043] Decoder: It can mean a device that performs decoding.

[0044] Parsing: It can mean entropy decoding to determine the value of a syntax element, or it can mean entropy decoding itself.

[0045] Block: It is an MxN array of samples, where M and N represent positive integer values, and a block can generally mean a 2D-shaped sample array.

[0046] Sample: It is the basic unit that makes up a block and can represent values from 0 to 2 d -1 according to the bit depth (bit depth, B Bd . In the present invention, pixel and picture element can be used in the same meaning as sample.

[0047] Unit: It can mean the unit of image encoding and decoding. In image encoding and decoding, a unit can be a region generated by dividing an image. Also, when an image is divided into subdivided units for encoding or decoding, the divided unit can be meant by the unit. In image encoding and decoding, for each unit, a predefined process can be executed. One unit can be further divided into sub-units with a smaller size than the unit. Depending on the function, the unit can mean a block, a macroblock, a coding tree unit, a coding tree block, a coding unit, a coding block, a prediction unit, a prediction block, a transform unit, a transform block, etc. Also, in order to indicate separately from a block, the unit can mean including a luminance component block, a corresponding chroma component block, and syntax elements for each block. The unit can have various sizes and shapes. In particular, the shape of the unit can include geometric figures that can be two-dimensionally represented, such as not only rectangles but also squares, trapezoids, triangles, pentagons, etc. Also, the unit information can include at least one of the type of the unit indicating a coding unit, a prediction unit, a transform unit, etc., the size of the unit, the depth of the unit, the encoding and decoding order of the unit, etc.

[0048] Reconstructed Neighbor Unit: It can mean a unit that has already been encoded or decoded and spatially / temporally reconstructed around the unit to be encoded / decoded. At this time, the reconstructed neighbor unit can mean a reconstructed neighbor block.

[0049] Neighbor block: It can mean a block adjacent to the block to be encoded / decoded. The block adjacent to the block to be encoded / decoded can mean a block whose boundary is adjacent to the block to be encoded / decoded. The neighbor block can mean a block located at the adjacent vertex of the block to be encoded / decoded. The neighbor block can also mean the reconstructed neighbor block.

[0050] Depth: It means the degree to which a unit is divided. In a tree structure, it can be said that the root node has the shallowest depth and the leaf node has the deepest depth.

[0051] Symbol: It can mean the syntax element of the unit to be encoded / decoded, the coding parameter, the value of the transform coefficient, etc.

[0052] Parameter Set: It can correspond to the header information in the structure within the bitstream. At least one of the video parameter set, sequence parameter set, picture parameter set, and adaptation parameter set can be included in the parameter set. Also, the parameter set can be meant to include slice header and tile header information.

[0053] Bitstream: It can mean a sequence of bits containing the encoded image information.

[0054] Prediction Unit: It is the basic unit when performing inter prediction or intra prediction and compensation therefor. One prediction unit can also be divided into a plurality of partitions with small sizes. In this case, each of the plurality of partitions is used as the basic unit during the prediction and compensation execution, and the partition into which the prediction unit is divided can also be regarded as a prediction unit. The prediction unit can have various sizes and shapes. In particular, the shape of the prediction unit can include geometric figures that can be two-dimensionally represented, such as not only rectangles, but also squares, trapezoids, triangles, pentagons, etc.

[0055] Prediction Unit Partition: It can mean the shape into which the prediction unit is divided.

[0056] Reference Picture List: It can mean a list containing one or more reference pictures used for inter prediction or motion compensation. The types of reference picture lists can be LC (List Combined), L0 (List 0), L1 (List 1), L2 (List 2), L3 (List 3), etc. One or more reference picture lists can be used for inter prediction.

[0057] Inter Prediction Indicator: It can mean the inter prediction direction (such as one-direction prediction, bi-direction prediction, etc.) of the block to be encoded / decoded during inter prediction, the number of reference pictures used when the block to be encoded / decoded generates a prediction block, and the number of prediction blocks used when the block to be encoded / decoded performs inter prediction or motion compensation.

[0058] Reference Picture Index: It can mean the index for a specific reference picture from the reference picture list. Here, the index can mean an index.

[0059] Reference Picture: It can mean an image that a specific unit refers to for inter prediction or motion compensation, and the reference picture can also be called a reference picture.

[0060] Motion Vector: It is a two-dimensional vector used for inter prediction or motion compensation, and can mean the offset between the picture to be coded / decoded and the reference picture. For example, (mvX, mvY) can indicate a motion vector, where mvX can indicate the horizontal component and mvY can indicate the vertical component.

[0061] Motion Vector Candidate: It can mean a unit that becomes a prediction candidate when predicting a motion vector, or the motion vector of that unit.

[0062] Motion Vector Candidate List: It can mean a list composed of motion vector candidates.

[0063] Motion Vector Candidate Index: It is an indicator that indicates a motion vector candidate in the motion vector candidate list, and can also be called the index of the motion vector predictor.

[0064] Motion Information: It can mean information that includes at least one of a motion vector, a reference picture index, an inter prediction indicator, as well as reference picture list information, a reference picture, a motion vector candidate, a motion vector candidate index, etc.

[0065] Merge Candidate List: It can mean a list configured using merge candidates.

[0066] Merge Candidate: It can include a spatial merge candidate, a temporal merge candidate, a combined merge candidate, a combined dual-prediction merge candidate, a zero merge candidate, etc. A merge candidate can include motion information such as prediction type information, a reference picture index for each list, and a motion vector.

[0067] Merge Index: It can mean information indicating a merge candidate in a merge candidate list. Also, the merge index can indicate the block from which the merge candidate was derived among the blocks reconstructed to be spatially / temporally adjacent to the current block. Also, the merge index can indicate at least one of the motion information that the merge candidate has.

[0068] Transform Unit: It can mean the basic unit when performing residual signal encoding / decoding such as transformation, inverse transformation, quantization, inverse quantization, and transformation coefficient encoding / decoding. One transform unit can be divided into a plurality of smaller-sized transform units. Transform units can have various sizes and shapes. In particular, the shape of a transform unit can include geometric figures that can be two-dimensionally represented, such as not only rectangles but also squares, trapezoids, triangles, pentagons, etc.

[0069] Scaling: It can mean the process of multiplying a factor to the transform coefficient level, and as a result, generate transform coefficients. Scaling can also be called dequantization.

[0070] Quantization Parameter: It can be meant as a value used when scaling the transform coefficient level in quantization and inverse quantization. At this time, the quantization parameter can be a value mapped to the quantization step size.

[0071] Delta Quantization Parameter: It can be meant as the difference value between the predicted quantization parameter and the quantization parameter of the unit to be encoded / decoded.

[0072] Scan: It can be meant as a method of sorting the order of coefficients within a block or matrix. For example, sorting a two-dimensional array into a one-dimensional array is called scan, and sorting a one-dimensional array into a two-dimensional array can also be called scan or inverse scan.

[0073] Transform Coefficient: The coefficient value generated after performing the transform. In the present invention, the quantization transform coefficient level to which quantization is applied to the transform coefficient can also be included in the meaning of the transform coefficient.

[0074] Non-zero Transform Coefficient: It can be meant as a transform coefficient whose transform coefficient value magnitude is not 0, or a transform coefficient level whose value magnitude is not 0.

[0075] Quantization Matrix: It can be meant as a matrix used in quantization or inverse quantization processing to improve the subjective or objective image quality. The quantization matrix can also be called a scaling list.

[0076] Quantization Matrix Coefficient: It can mean each element within the quantization matrix. The quantization matrix coefficient can also be called the matrix coefficient.

[0077] Default Matrix: It can mean a predetermined quantization matrix that is predefined in the encoder and decoder.

[0078] Non-default Matrix: It can mean a quantization matrix that is not predefined in the encoder and decoder and is signaled by the user.

[0079] Coding Tree Unit: It can be composed of two chrominance component (Cb, Cr) coding tree blocks related to one luminance component (Y) coding tree block. Each coding tree unit can be divided using one or more splitting methods such as a quadtree, binary tree, etc. to form lower-level units such as coding units, prediction units, and transform units. It can be used as a term to indicate a pixel block that is a processing unit in the decoding / encoding process of an image, like the splitting of an input image.

[0080] Coding Tree Block: It can be used as a term to indicate any one of a Y coding tree block, a Cb coding tree block, and a Cr coding tree block.

[0081] FIG. 1 is a block diagram showing a configuration according to an embodiment of an encoding apparatus to which the present invention is applied.

[0082] Encoding apparatus 100 can be a video encoding apparatus or an image encoding apparatus. The video can include one or more images. Encoding apparatus 100 can sequentially encode one or more images of the video according to time.

[0083] Referring to FIG. 1, the encoding apparatus 100 can include a motion prediction unit 111, a motion compensation unit 112, an intra prediction unit 120, a switch 115, a subtractor 125, a conversion unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse conversion unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.

[0084] The encoding apparatus 100 can perform encoding on the input image in the intra mode and / or the inter mode. Also, the encoding apparatus 100 can generate a bitstream through the encoding of the input image and output the generated bitstream. When the intra mode is used as the prediction mode, the switch 115 may be switched to intra. When the inter mode is used as the prediction mode, the switch 115 may be switched to inter. Here, the intra mode can mean the intra prediction mode, and the inter mode can mean the inter prediction mode. The encoding apparatus 100 can generate a prediction block for the input block of the input image. Also, after the prediction block is generated, the encoding apparatus 100 can encode the difference (residual) between the input block and the prediction block. The input image may also be referred to as the current image that is currently the encoding target. The input block may also be referred to as the current block or the encoding target block that is currently the encoding target.

[0085] When the prediction mode is the intra mode, the intra prediction unit 120 can also use the pixel values of the blocks that have already been encoded in the vicinity of the current block as reference pixels. The intra prediction unit 120 can perform spatial prediction using the reference pixels and generate prediction samples for the input block through the spatial prediction. Here, the intra prediction can mean the intra prediction.

[0086] When the prediction mode is the inter mode, the motion prediction unit 111 can search for a region in the reference image that best matches the input block in the motion prediction process, and can derive a motion vector using the searched region. The reference image can be stored in the reference picture buffer 190.

[0087] The motion compensation unit 112 can generate a prediction block by performing motion compensation using the motion vector. Here, the motion vector can be a two-dimensional vector used for inter prediction. Also, the motion vector can indicate an offset between the current image and the reference image. Here, inter prediction can mean inter prediction.

[0088] When the value of the motion vector does not have an integer value, the motion prediction unit 111 and the motion compensation unit 112 can generate a prediction block by applying an interpolation filter to a partial region in the reference image. In order to perform inter prediction or motion compensation, based on the coding unit, it is possible to determine which method among the skip mode, merge mode, AMVP mode, and current picture reference mode is the motion prediction and motion compensation method of the prediction unit included in the corresponding coding unit, and inter prediction or motion compensation can be performed according to each mode. Here, the current picture reference mode can mean a prediction mode using a previously reconstructed region in the current picture to which the coding target block belongs. In order to specify the previously reconstructed region, a motion vector for the current picture reference mode can be defined. Whether the coding target block is coded in the current picture reference mode can be coded using the reference image index of the coding target block.

[0089] The subtractor 125 can generate a residual block using the difference between the input block and the prediction block. The residual block is also referred to as a residual signal.

[0090] The conversion unit 130 can perform a transform on the residual block to generate transform coefficients and output the transform coefficients. Here, the transform coefficients can be the coefficient values generated by performing a transform on the residual block. When the transform skip mode is applied, the conversion unit 130 can also omit the transform on the residual block.

[0091] By applying quantization to the transform coefficients, quantized transform coefficient levels can be generated. Hereinafter, in the embodiments, the quantized transform coefficient levels may also be referred to as transform coefficients.

[0092] The quantization unit 140 can generate quantized transform coefficient levels by quantizing the transform coefficients based on quantization parameters and output the quantized transform coefficient levels. At this time, in the quantization unit 140, the transform coefficients can be quantized using a quantization matrix.

[0093] The entropy encoding unit 150 can generate a bitstream by performing entropy encoding based on a probability distribution on the value calculated by the quantization unit 140, or the coding parameter value calculated in the encoding process, etc., and output the bitstream. The entropy encoding unit 150 can perform entropy encoding on information for decoding the image in addition to the information of the pixels of the image. For example, the information for decoding the image can include syntax elements and the like.

[0094] When entropy coding is applied, symbols with a high occurrence probability are assigned a small number of bits, and symbols with a low occurrence probability are assigned a large number of bits to represent the symbols, so that the size of the bit string for the symbol to be coded can be reduced. Therefore, the compression efficiency of image coding can be increased through entropy coding. The entropy coding unit 150 can use coding methods such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding) for entropy coding. For example, the entropy coding unit 150 can perform entropy coding using a variable length coding (VLC) table. Also, after deriving the binarization method of the target symbol and the probability model of the target symbol / bin, the entropy coding unit 150 can perform arithmetic coding using the derived binarization method or probability model.

[0095] The entropy coding unit 150 can change the two-dimensional block-shaped coefficients into a one-dimensional vector through a transform coefficient scanning method to code the transform coefficient levels. For example, by scanning the coefficients of the block using upright scanning, it can be changed into a one-dimensional vector. Depending on the size of the transform unit and the intra prediction mode, instead of upright scanning, vertical scanning that scans the two-dimensional block-shaped coefficients in the column direction or horizontal scanning that scans the two-dimensional block-shaped coefficients in the row direction may be used. That is, it is possible to determine which of the upright scan, vertical scan, and horizontal scan methods is used according to the size of the transform unit and the intra prediction mode.

[0096] Coding parameters can include not only information that is encoded by an encoder and signaled to a decoder like syntax elements, but also information derived in the encoding or decoding process, and can mean information necessary when encoding or decoding an image. For example, block size, block depth, block partitioning information, unit size, unit depth, unit partitioning information, quadtree partitioning flag, binary tree partitioning flag, binary tree partitioning direction, intra prediction mode, intra prediction direction, reference sample filtering method, prediction block boundary filtering method, filter taps, filter coefficients, inter prediction mode, motion information, motion vectors, reference image index, inter prediction direction, inter prediction indicator, reference image list, motion vector prediction, motion vector candidate list, presence or absence of use of motion merge mode, motion merge candidate, motion merge candidate list, presence or absence of use of skip mode, type of interpolation filter, size of motion vectors, accuracy of motion vector representation, transform type, transform size, information on presence or absence of use of additional (secondary) transform, information on presence or absence of residual signal, Coded Block Pattern, Coded Block Flag, quantization parameter, quantization matrix, in-loop filter information, information on whether to apply in-loop filter, in-loop filter coefficients, binarization / inverse binarization method, context model, context bin, bypass bin, transform coefficients, transform coefficient levels, scanning method of transform coefficient levels, image display / output order, slice identification information, slice type, slice partitioning information, tile identification information, tile type, tile partitioning information, picture type, bit depth, at least one value or combination form of information for luminance signal or chrominance signal can be included in the coding parameters.

[0097] The residual signal can mean the difference between the original signal and the predicted signal. Or, the residual signal can be a signal generated by transforming the difference between the original signal and the predicted signal. Or, the residual signal can be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. The residual block can be a residual signal in block units.

[0098] When the encoding device 100 performs encoding using inter-prediction, the encoded current image can be used as a reference image for other images to be processed later. Therefore, the encoding device 100 can further decode the encoded current image and save the decoded image as a reference image. For decoding, inverse quantization and inverse transformation for the encoded current image can be processed.

[0099] The quantized coefficients can be inverse quantized by the inverse quantization unit 160 and inverse transformed by the inverse transformation unit 170. The inverse quantized and inverse transformed coefficients can be combined with the prediction block via the adder 175. By combining the inverse quantized and inverse transformed coefficients with the prediction block, a reconstructed block can be generated.

[0100] The reconstructed block can pass through the filter unit 180. The filter unit 180 can apply at least one of a deblocking filter, a Sample Adaptive Offset (SAO), and an Adaptive Loop Filter (ALF) to the reconstructed block or the reconstructed image. The filter unit 180 is also referred to as an in-loop filter.

[0101] The deblocking filter can remove block distortion that occurs at the boundaries between blocks. To determine whether to perform the deblocking filter, it is possible to determine whether to apply the deblocking filter to the current block based on the pixels included in several columns or rows included in the block. When applying the deblocking filter to a block, a strong filter or a weak filter can be applied according to the necessary deblocking filtering strength. Also, when applying the deblocking filter, horizontal filtering and vertical filtering can be processed in parallel during vertical filtering and horizontal filtering.

[0102] The sample adaptive offset can add an appropriate offset value to the pixel value to compensate for coding errors. The sample adaptive offset can correct the offset from the original image for each pixel in the image that has undergone deblocking. To perform offset correction for a specific picture, after dividing the pixels included in the image into a certain number of regions, a method of determining the region to which the offset should be applied and applying the offset to the corresponding region, or a method of applying the offset in consideration of the edge information of each pixel can be used.

[0103] The adaptive loop filter can perform filtering based on the value obtained by comparing the reconstructed image and the original image. After dividing the pixels included in the image into predetermined groups, one filter to be applied to the group can be determined and differential filtering can be performed for each group. Information related to whether to apply the adaptive loop filter can be signaled for each coding unit (CU) of the luminance signal, and the shape and filter coefficients of the adaptive loop filter to be applied can be different according to each block. Also, regardless of the characteristics of the block to be applied, it is also possible to apply an adaptive loop filter of the same form (fixed form).

[0104] The reconstructed block that has passed through the filter unit 180 can be stored in the reference picture buffer 190.

[0105] FIG. 2 is a block diagram showing a configuration according to an embodiment of a decoding apparatus to which the present invention is applied. The decoding apparatus 200 can be a video decoding apparatus or an image decoding apparatus.

[0106] Referring to FIG. 2, the decoding apparatus 200 can include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, a motion compensation unit 250, an adder 255, a filter unit 260, and a reference picture buffer 270.

[0107] The decoding apparatus 200 can receive the bit stream output from the encoding apparatus 100. The decoding apparatus 200 can perform decoding on the bit stream in an intra mode or an inter mode. Further, the decoding apparatus 200 can generate a reconstructed image through decoding and output the reconstructed image.

[0108] When the prediction mode used for decoding is the intra mode, the switch may be switched to intra. When the prediction mode used for decoding is the inter mode, the switch may be switched to inter.

[0109] The decoding apparatus 200 can obtain a reconstructed residual block from the input bit stream and generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding apparatus 200 can generate a reconstructed block, which is a block to be decoded, by adding the reconstructed residual block and the prediction block. The block to be decoded may also be referred to as the current block.

[0110] The entropy decoding unit 210 can generate symbols by performing entropy decoding based on the probability distribution for the bit stream. The generated symbols can include symbols in the form of quantized transform coefficient levels. Here, the entropy decoding method can be similar to the entropy encoding method described above. For example, the entropy decoding method can be the inverse process of the entropy encoding method described above.

[0111] In order to decode the transform coefficient levels, the entropy decoding unit 210 can change the one-dimensional vector form coefficients into a two-dimensional block form by a transform coefficient scanning method. For example, by scanning the coefficients of the block using up right scanning, it can be changed into a two-dimensional block form. Depending on the size of the transform unit and the intra prediction mode, vertical scanning or horizontal scanning may be used instead of up right scanning. That is, it is possible to determine which of the up right scanning, vertical scanning, and horizontal scanning methods is used according to the size of the transform unit and the intra prediction mode.

[0112] The quantized transform coefficient levels can be inverse quantized by the inverse quantization unit 220 and inverse transformed by the inverse transform unit 230. As a result of inverse quantization and inverse transformation of the quantized transform coefficient levels, a reconstructed residual block can be generated. At this time, the inverse quantization unit 220 can apply a quantization matrix to the quantized transform coefficient levels.

[0113] When the intra mode is used, the intra prediction unit 240 can generate a prediction block by performing spatial prediction using the pixel values of the already decoded blocks in the vicinity of the block to be decoded.

[0114] When inter mode is used, the motion compensation unit 250 can generate a prediction block by performing motion compensation using the motion vector and the reference image stored in the reference picture buffer 270. When the value of the motion vector does not have an integer value, the motion compensation unit 250 can generate a prediction block by applying an interpolation filter to a partial region in the reference image. In order to perform motion compensation, based on the coding unit, it is possible to determine which method among the skip mode, merge mode, AMVP mode, and current picture reference mode is the motion compensation method of the prediction unit included in the corresponding coding unit, and motion compensation can be performed according to each mode. Here, the current picture reference mode can mean a prediction mode using a reconstructed region in the current picture to which the block to be decoded belongs. In order to specify the already reconstructed region, a motion vector for the current picture reference mode can be used. A flag or index indicating whether the block to be decoded is a block encoded in the current picture reference mode can also be signaled, or can be inferred using the reference picture index of the block to be decoded. The current picture for the current picture reference mode can exist at a fixed position (for example, a position where the reference picture index is 0 or the last position) in the reference picture list for the block to be decoded. Or, it may be variably positioned in the reference picture list, and for this purpose, a separate reference picture index indicating the position of the current picture may be signaled. Here, signaling a flag or index can mean that in the encoder, the flag or index is entropy encoded and included in the bitstream, and in the decoder, the flag or index is entropy decoded from the bitstream.

[0115] The reconstructed residual block and the prediction block can be added via the adder 255. The block generated by the addition of the reconstructed residual block and the prediction block can pass through the filter unit 260. The filter unit 260 can apply at least one of a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the reconstructed block or the reconstructed image. The filter unit 260 can output the reconstructed image. The reconstructed image can be stored in the reference picture buffer 270 and used for inter prediction.

[0116] FIG. 3 is a schematic diagram showing a division structure of an image when encoding and decoding the image. FIG. 3 schematically shows an embodiment in which one unit is divided into a plurality of lower units.

[0117] In order to efficiently divide an image, in encoding and decoding, a coding unit (CU) can be used. Here, the coding unit can mean a coding unit. The unit can be a term that collectively refers to 1) a syntax element and 2) a block including image samples. For example, "division of a unit" can mean "division of a block corresponding to the unit". The block division information may include information regarding the depth of the unit. The depth information can indicate the number of times and / or the degree to which the unit is divided.

[0118] Referring to FIG. 3, the image 300 is sequentially divided in units of the largest coding unit (LCU), and the division structure is determined in units of the LCU. Here, the LCU can be used with the same meaning as a coding tree unit (CTU). One unit can be hierarchically divided with depth information based on a tree structure. Each divided lower unit can have depth information. Since the depth information indicates the number of times and / or the degree to which the unit is divided, it can also include information regarding the size of the lower unit.

[0119] The splitting structure can mean the distribution of coding units (CUs) within the LCU310. A CU can be a unit for efficiently encoding / decoding an image. Such a distribution can be determined by whether to split a single CU into multiple (two or more positive integers including 2, 4, 8, 16, etc.) CUs. The horizontal width and vertical height of the CUs generated by the splitting can be, respectively, half of the horizontal width and half of the vertical height of the CU before splitting, or can have sizes smaller than the horizontal width and smaller than the vertical height of the CU before splitting according to the number of CUs into which it is split. The split CUs can be recursively split into multiple CUs with reduced horizontal width and vertical height in a similar manner.

[0120] At this time, the splitting of the CU can be recursively performed up to a predetermined depth. The depth information is information indicating the size of the CU and can be saved for each CU. For example, the depth of the LCU is 0, and the depth of the smallest coding unit (SCU) can be a predefined maximum depth. Here, as described above, the LCU is a coding unit having the size of the largest coding unit, and the SCU can be a coding unit having the size of the smallest coding unit.

[0121] The splitting starts from the LCU310, and each time the horizontal width and vertical height of the CU decrease due to the splitting, the depth of the CU increases by 1. For each depth, the non-split CUs can have a size of 2Nx2N. In the case of the CUs to be split, the 2N×2N-sized CUs can be split into multiple NxN-sized CUs. The size of N decreases by half each time the depth increases by 1.

[0122] For example, when one encoding unit is divided into four encoding units, the horizontal and vertical widths of the four divided encoding units can each have half the size compared to the horizontal and vertical widths of the encoding unit before division. As an example, when a 32x32-sized encoding unit is divided into four encoding units, the four divided encoding units can each have a size of 16x16. When one encoding unit is divided into four encoding units, it can be said that the encoding unit is divided in a quad-tree shape.

[0123] For example, when one encoding unit is divided into two encoding units, the horizontal or vertical width of the two divided encoding units can have half the size compared to the horizontal or vertical width of the encoding unit before division. As an example, when a 32x32-sized encoding unit is vertically divided into two encoding units, the two divided encoding units can each have a size of 16x32. As an example, when a 32x32-sized encoding unit is horizontally divided into two encoding units, the two divided encoding units can each have a size of 32x16. When one encoding unit is divided into two encoding units, it can be said that the encoding unit is divided in a binary-tree shape.

[0124] Referring to FIG. 3, the LCU with a depth of 0 can be 64x64 pixels. 0 can be the minimum depth. The SCU with a depth of 3 can be 8x8 pixels. 3 can be the maximum depth. At this time, the CU of 64x64 pixels, which is an LCU, can be represented at depth 0. The CU of 32x32 pixels can be represented at depth 1. The CU of 16x16 pixels can be represented at depth 2. The CU of 8x8 pixels, which is an SCU, can be represented at depth 3.

[0125] In addition, information on whether a CU is divided or not can be expressed using CU division information. The division information can be 1-bit information. All CUs except the SCU can include the division information. For example, if the value of the division information is 0, the CU does not have to be divided, and if the value of the division information is 1, the CU may be divided.

[0126] Figure 4 is a diagram showing the forms of prediction units (PUs) that a coding unit (CU) can include.

[0127] Among the CUs divided from the LCU, the CUs that are not further divided can be divided into one or more prediction units (PUs). Such processing may also be referred to as division.

[0128] A PU can be a basic unit for prediction. A PU can be encoded and decoded in any of the skip mode, inter mode, and intra mode. A PU can be divided into various forms according to the mode.

[0129] In addition, the coding unit may not be divided into prediction units, and the coding unit and the prediction unit can have the same size.

[0130] As shown in Figure 4, in the skip mode, there may be no division within the CU. In the skip mode, a 2Nx2N mode 410 having the same size as the CU without division can be supported.

[0131] In the inter mode, a form divided into eight within the CU can be supported. For example, in the inter mode, a 2Nx2N mode 410, a 2NxN mode 415, an Nx2N mode 420, an NxN mode 425, a 2NxnU mode 430, a 2NxnD mode 435, an nLx2N mode 440, and an nRx2N mode 445 can be supported. In the intra mode, a 2Nx2N mode 410 and an NxN mode 425 can be supported.

[0132] One coding unit can be divided into one or more prediction units, and one prediction unit can also be divided into one or more prediction units.

[0133] For example, when one prediction unit is divided into four prediction units, the horizontal and vertical widths of the four divided prediction units can each have half the size compared to the horizontal and vertical widths of the prediction unit before division. As an example, when a 32x32-sized prediction unit is divided into four prediction units, the four divided prediction units can each have a size of 16x16. When one prediction unit is divided into four prediction units, it can be said that the prediction unit is divided in a quad-tree shape.

[0134] For example, when one prediction unit is divided into two prediction units, the horizontal or vertical width of the two divided prediction units can have half the size compared to the horizontal or vertical width of the prediction unit before division. As an example, when a 32x32-sized prediction unit is vertically divided into two prediction units, the two divided prediction units can each have a size of 16x32. As an example, when a 32x32-sized prediction unit is horizontally divided into two prediction units, the two divided prediction units can each have a size of 32x16. When one prediction unit is divided into two prediction units, it can be said that the prediction unit is divided in a binary-tree shape.

[0135] FIG. 5 is a diagram showing the form of a transform unit (TU) that a coding unit (CU) can include.

[0136] A transform unit (TU) can be a basic unit used for processing such as transformation, quantization, inverse transformation, and inverse quantization within a CU. The TU can have a shape such as a square or a rectangle. The TU may also be determined dependently on the size and / or shape of the CU.

[0137] Of the CUs split from the LCU, a CU that is not split into further CUs can be split into one or more TUs. At this time, the split structure of the TUs can be a quad-tree structure. For example, as shown in FIG. 5, one CU 510 can be split once or more by a quad-tree structure. When a CU is split more than once, it can be said that it is recursively split. Through splitting, one CU 510 can be composed of TUs of various sizes. Alternatively, it is also possible to split into one or more TUs based on the number of vertical lines and / or horizontal lines for splitting the CU. The CU may be split into symmetric TUs or asymmetric TUs. For splitting into an asymmetric TU, information regarding the size / shape of the TU may be signaled or derived from information regarding the size / shape of the CU.

[0138] Also, the coding unit may not be split into transform units, and the coding unit and the transform unit can have the same size.

[0139] One coding unit can be split into one or more transform units, and one transform unit can also be split into one or more transform units.

[0140] For example, when one transform unit is split into four transform units, the horizontal and vertical widths of the four split transform units can each have a size that is half of the horizontal and vertical widths of the transform unit before splitting. As an example, when a transform unit of size 32x32 is split into four transform units, the four split transform units can each have a size of 16x16. When one transform unit is split into four transform units, it can be said that the transform unit is split in a quad-tree manner.

[0141] For example, when one conversion unit is divided into two conversion units, the horizontal or vertical width of the two divided conversion units can have a size that is half of the horizontal or vertical width of the conversion unit before division. As an example, when a 32x32-sized conversion unit is vertically divided into two conversion units, the two divided conversion units can each have a size of 16x32. As an example, when a 32x32-sized conversion unit is horizontally divided into two conversion units, the two divided conversion units can each have a size of 32x16. When one conversion unit is divided into two conversion units, it can be said that the conversion unit is divided in a binary-tree shape.

[0142] When performing conversion, the residual block can be converted using at least one of a plurality of predefined conversion methods. As an example, as a plurality of predefined conversion methods, DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), or KLT, etc. can be used. Which conversion method is applied to convert the residual block may be determined using at least one of the inter-prediction mode information, intra-prediction mode information, and size / shape of the conversion block of the prediction unit. In a certain case, information indicating the conversion method may also be signaled.

[0143] FIG. 6 is a diagram for explaining an embodiment of intra-prediction processing.

[0144] The intra-prediction mode can be a non-directional mode or a directional mode. The non-directional mode is a DC mode or a Planar mode, and the directional mode is a prediction mode with a specific direction or angle, and the number can be M of one or more. The directional mode can be represented by at least one of a mode number, a mode value, a mode number, and a mode angle.

[0145] The number of intra-prediction modes can be N of one or more including the non-directional and directional modes.

[0146] The number of intra prediction modes may vary depending on the block size. For example, when the block size is 4x4 or 8x8, it may be 67, when it is 16x16, it may be 35, when it is 32x32, it may be 19, and when it is 64x64, it may be 7.

[0147] The number of intra prediction modes can be fixed to N regardless of the block size. For example, it can be fixed to at least one of 35 or 67 regardless of the block size.

[0148] The number of intra prediction modes may vary depending on the type of color component. For example, the number of prediction modes may vary depending on whether the color component is a luma signal or a chroma signal.

[0149] Intra encoding and / or decoding can be performed using sample values or encoding parameters included in neighboring reconstructed blocks.

[0150] A step of checking whether samples included in neighboring reconstructed blocks can be used as reference samples for the block to be encoded / decoded by intra prediction can be performed. If there are samples that cannot be used as reference samples for the block to be encoded / decoded, the sample values can be copied and / or interpolated to at least one of the samples included in neighboring reconstructed blocks and used as reference samples for the block to be encoded / decoded.

[0151] When performing intra prediction, a filter can be applied to at least one of the reference sample or the prediction sample based on at least one of the intra prediction mode and the size of the encoding / decoding target block. At this time, the encoding / decoding target block can mean the current block, and can mean at least one of the encoding block, the prediction block, and the transform block. The type of filter applied to the reference sample or the prediction sample may vary depending on at least one of the intra prediction mode or the size / shape of the current block. The type of the filter may vary depending on at least one of the filter tap number, the filter coefficient value, or the filter strength.

[0152] Among the intra prediction modes, in the non-directional Planar mode, when generating the prediction block of the target encoding / decoding block, the sample value in the prediction block can be generated as a weighted sum of the upper reference sample of the current sample, the left reference sample of the current sample, the upper right reference sample of the current block, and the lower left reference sample of the current block according to the sample position.

[0153] Among the intra prediction modes, in the non-directional DC mode, when generating the prediction block of the target encoding / decoding block, it can be generated as the average value of the upper reference sample of the current block and the left reference sample of the current block. Also, filtering can be performed using the reference sample values for one or more upper rows and one or more left columns adjacent to the reference samples in the encoding / decoding block.

[0154] Among the intra prediction modes, in the case of a plurality of angular modes, the prediction block can be generated using the upper right and / or lower left reference samples, and the angular modes can have different directions from each other. To generate the prediction sample value, interpolation in real number units can also be performed.

[0155] To perform an intra prediction method, the intra prediction mode of a current prediction block can be predicted from the intra prediction modes of prediction blocks existing in the vicinity of the current prediction block. When predicting the intra prediction mode of the current prediction block using the mode information predicted from the intra prediction modes in the vicinity, if the intra prediction modes of the current prediction block and the neighboring prediction blocks are the same, information indicating that the intra prediction modes of the current prediction block and the neighboring prediction blocks are the same can be signaled using predetermined flag information. If the intra prediction modes of the current prediction block and the neighboring prediction blocks are different from each other, entropy coding can be performed to encode the intra prediction mode information of the block to be encoded / decoded.

[0156] FIG. 7 is a diagram for explaining an embodiment of an inter prediction process.

[0157] The rectangles in FIG. 7 can indicate an image (or picture). Also, the arrows in FIG. 7 can indicate a prediction direction. That is, an image can be encoded and / or decoded according to the prediction direction. Each image can be classified into an I picture (Intra Picture), a P picture (Uni-predictive Picture), a B picture (Bi-predictive Picture), etc. according to the encoding type. Each picture can be encoded and decoded according to the encoding type of each picture.

[0158] When the image to be symbolized is an I picture, the image can be intra-coded for the image itself without inter-prediction. When the image to be coded is a P picture, the image can be coded through inter-prediction or motion compensation that uses a reference image only in the forward direction. When the image to be coded is a B picture, it can be coded through inter-prediction or motion compensation that uses reference pictures on both the forward and backward sides, or it can be coded through inter-prediction or motion compensation that uses a reference picture in either the forward or backward direction. Here, when the inter-prediction mode is used, the encoder can perform inter-prediction or motion compensation, and the decoder can perform corresponding motion compensation. The images of P pictures and B pictures that are coded and / or decoded using a reference image can be regarded as images for which inter-prediction is used.

[0159] Next, the inter-prediction according to the embodiment will be specifically described.

[0160] Inter-prediction or motion compensation can be performed using a reference picture and motion information. Also, the above-described skip mode can be used for inter-prediction.

[0161] The reference picture can be at least one of a picture before the current picture or a picture after the current picture. At this time, inter-prediction can perform prediction for a block of the current picture based on the reference picture. Here, the reference picture can mean an image used for predicting the block. At this time, the area in the reference picture can be specified by using a reference picture index (refIdx) indicating the reference picture and a motion vector described later.

[0162] Inter prediction can select a reference picture and a reference block corresponding to the current block within the reference picture, and can generate a predicted block for the current block using the selected reference block. The current block can be a block to be currently encoded or decoded among the blocks of the current picture.

[0163] Motion information can be derived from the inter prediction process by each of the encoding device 100 and the decoding device 200. Also, the derived motion information can be used for inter prediction. At this time, the encoding device 100 and the decoding device 200 can improve the encoding and / or decoding efficiency by using the motion information of the reconstructed neighboring block and / or the motion information of the collocated block. The collocated block can be a block corresponding to the spatial position of the block to be encoded / decoded within the already reconstructed collocated picture. The reconstructed neighboring block can be a block within the current picture and can be a block that has already been reconstructed through encoding and / or decoding. Also, the reconstructed block can be an adjacent block adjacent to the block to be encoded / decoded and / or a block located at the outer corner of the block to be encoded / decoded. Here, the block located at the outer corner of the block to be encoded / decoded refers to a block that is vertically adjacent to an adjacent block horizontally adjacent to the block to be encoded / decoded, or a block that is horizontally adjacent to an adjacent block vertically adjacent to the block to be encoded / decoded.

[0164] Each of the symbolization device 100 and the decoding device 200 can determine a block existing at a position spatially corresponding to an encoding / decoding target block within a collocated picture, and can determine a predetermined relative position based on the determined block. The predetermined relative position can be a position inside and / or outside the block existing at the position spatially corresponding to the encoding / decoding target block. Also, each of the symbolization device 100 and the decoding device 200 can derive a collocated block based on the determined predetermined relative position. Here, the collocated picture can be any one of at least one reference picture included in the reference picture list.

[0165] The method for deriving motion information can vary according to the prediction mode of the encoding / decoding target block. For example, as prediction modes applied for inter prediction, there can be Advanced Motion Vector Prediction (AMVP) and merge mode, etc. Here, the merge mode can be called the motion merge mode.

[0166] For example, when AMVP is applied as the prediction mode, each of the symbolization device 100 and the decoding device 200 can generate a motion vector candidate list using the motion vector of the reconstructed neighboring block and / or the motion vector of the collocated block. The motion vector of the reconstructed neighboring block and / or the motion vector of the collocated block can be used as a motion vector candidate. Here, the motion vector of the collocated block can be called a temporal motion vector candidate, and the motion vector of the reconstructed neighboring block can be called a spatial motion vector candidate.

[0167] The bitstream generated by the symbolization device 100 can include a motion vector candidate index. That is, the symbolization device 100 can perform entropy coding on the motion vector candidate index to generate a bitstream. The motion vector candidate index can indicate the optimal motion vector candidate selected from among the motion vector candidates included in the motion vector candidate list. The motion vector candidate index can be signaled from the symbolization device 100 to the decoding device 200 via the bitstream.

[0168] The decoding device 200 can perform entropy decoding on the motion vector candidate index from the bitstream, and use the entropy-decoded motion vector candidate index to select the motion vector candidate of the block to be decoded from among the motion vector candidates included in the motion vector candidate list.

[0169] The symbolization device 100 can calculate the motion vector difference (MVD: Motion Vector Difference) between the motion vector of the block to be symbolized and the motion vector candidate, and can perform entropy coding on the MVD. The bitstream can include the entropy-coded MVD. The MVD can be signaled from the symbolization device 100 to the decoding device 200 via the bitstream. At this time, the decoding device 200 can perform entropy decoding on the received MVD from the bitstream. The decoding device 200 can derive the motion vector of the block to be decoded by adding the decoded MVD and the motion vector candidate.

[0170] The bitstream can include a reference picture index indicating a reference picture, etc. The reference picture index can be entropy-coded and signaled from the encoding device 100 to the decoding device 200 via the bitstream. The decoding device 200 can predict the motion vector of the block to be decoded using the motion information of neighboring blocks, and can derive the motion vector of the block to be decoded using the predicted motion vector and the motion vector difference. The decoding device 200 can generate a predicted block for the block to be decoded based on the derived motion vector and the reference picture index information.

[0171] As another example of the method for deriving motion information, there is a merge mode. The merge mode can mean the merging of motions for a plurality of blocks. The merge mode can mean applying the motion information of one block to other blocks together. When the merge mode is applied, each of the encoding device 100 and the decoding device 200 can generate a merge candidate list using the motion information of the reconstructed neighboring blocks and / or the motion information of collocated blocks. The motion information can include at least one of 1) a motion vector, 2) a reference picture index, and 3) an inter prediction indicator. The prediction indicator can be unidirectional (L0 prediction, L1 prediction) or bidirectional.

[0172] At this time, the merge mode can be applied in CU units or PU units. When the merge mode is performed in CU units or PU units, the encoding device 100 can entropy-encode predefined information to generate a bitstream and then signal it to the decoding device 200. The bitstream can include predefined information. The predefined information can include 1) a merge flag which is information indicating whether to perform the merge mode for each block partition, and 2) a merge index which is information about which block among the neighboring blocks adjacent to the block to be encoded is merged with. For example, the neighboring blocks of the block to be encoded can include the left adjacent block of the block to be encoded, the upper adjacent block of the block to be encoded, and the temporal adjacent blocks of the block to be encoded.

[0173] The merge candidate list can indicate a list in which motion information is stored. Also, the merge candidate list can be generated before the merge mode is performed. The motion information stored in the merge candidate list can be at least one of the motion information of the neighboring blocks adjacent to the block to be encoded / decoded, the motion information of the block collocated with the block to be encoded / decoded in the reference picture, the new motion information generated by the combination of the motion information already existing in the merge candidate list, and the zero merge candidate. Here, the motion information of the neighboring blocks adjacent to the block to be encoded / decoded can be called a spatial merge candidate, and the motion information of the block collocated with the block to be encoded / decoded in the reference picture can be called a temporal merge candidate.

[0174] The skip mode may be a mode that directly applies the motion information of neighboring blocks to the block to be coded / decoded. The skip mode may be any of the modes used for inter prediction. When the skip mode is used, the encoding device 100 can entropy-encode information about which block's motion information is to be used as the motion information of the block to be encoded and signal it to the decoding device 200 via a bitstream. The encoding device 100 may not need to signal other information to the decoding device 200. For example, the other information may be syntax element information. The syntax element information can include at least one of motion vector difference information, coding block flag, and transform coefficient level.

[0175] The residual signal generated after intra or inter prediction can be converted to the frequency domain through conversion processing as part of the quantization process. At this time, in addition to DCT type 2 (DCT-II), various DCT and DST kernels can be used for the primary conversion to be performed. These conversion kernels can be converted by a separable transform that performs a one-dimensional transform (1D transform) in the horizontal and / or vertical directions on the residual signal respectively, or by a two-dimensional non-separable transform.

[0176] As an example, for the DCT and DST types used for conversion, as shown in the following table, in addition to DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII can be adaptively used during 1D transform. For example, as in the examples of Table 1 and Table 2, a transform set can be configured to derive the DCT or DST type used for conversion.

Table 1

Table 2

[0177] For example, as shown in FIG. 8, after defining different transform sets for the horizontal and vertical directions according to the intra prediction mode, in the encoder / decoder, the intra prediction mode of the currently encoded / decoded block and the transform included in the corresponding transform set can be used to perform the transform and / or inverse transform. In this case, the transform set can be defined based on the same rule in the encoder / decoder instead of being entropy encoded / decoded. In this case, the information indicating which transform among the transforms belonging to the transform set is used is entropy encoding / decoding. For example, when the size of the block is 64x64 or less, according to the intra prediction mode, a total of three transform sets are configured as in the example of Table 2, and after combining a total of nine multiple transform methods using three transforms each for the horizontal transform and the vertical transform, the encoding efficiency can be improved by encoding / decoding the residual signal with the optimal transform method. At this time, in order to entropy encode / decide the information about which of the three transforms belonging to one transform set is used, truncated unary binarization can also be used. At this time, for at least one of the vertical transform and the horizontal transform, the information indicating which transform among the transforms belonging to the transform set is used can be entropy encoded / decoded.

[0178] After the above-described primary transformation is completed, the symbolizer can perform a secondary transform to increase the energy concentration for the transformed coefficients, as in the example of FIG. 9. The secondary transform can also perform a separable transform that respectively executes one-dimensional transforms for the horizontal and / or vertical directions, or can perform a two-dimensional non-separable transform, and the transform information used can be signaled or implicitly derived by the encoder / decoder according to the current and neighboring coding information. For example, a set of transforms for the secondary transform can be defined like the primary transform, and the set of transforms can be defined by the encoder / decoder based on the same rules instead of being entropy-coded / decoded. In this case, information indicating which of the transforms belonging to the set of transforms is used can be signaled and applied to at least one of the residual signals by intra or inter prediction.

[0179] At least one of the number or type of transform candidates is different for each set of transforms, and at least one of the number or type of transform candidates can be variably determined in consideration of at least one of the position, size, division form, prediction mode (intra / inter mode), or the directionality / non-directionality of the intra prediction mode of the block (such as CU, PU, TU, etc.).

[0180] The decoder can perform the secondary inverse transform according to whether to perform the secondary inverse transform, and can perform the primary inverse transform according to whether to perform the primary inverse transform from the execution result of the secondary inverse transform.

[0181] The above-described primary transform and secondary transform are applied to at least one signal component of the luminance / chrominance components, or can be applied according to the size / shape of any coding block, and whether to use it in any coding block, and the index indicating the used primary transform / secondary transform can be entropy-coded / decoded, or implicitly derived by the encoder / decoder based on at least one of the current / neighboring coding information.

[0182] The residual signal generated after intra or inter prediction undergoes quantization processing after the first and / or second transformation is completed. The quantized transform coefficients undergo entropy encoding processing. At this time, as shown in FIG. 10, the quantized transform coefficients can be scanned along the diagonal, vertical, and horizontal directions based on at least one of the intra prediction mode or the minimum block size / shape.

[0183] Also, the entropy-decoded, quantized transform coefficients can be inverse-scanned, aligned in block form, and at least one of inverse quantization or inverse transformation may be performed on the block. At this time, as a method of inverse scanning, at least one of diagonal scan, horizontal scan, and vertical scan can be performed.

[0184] As an example, when the size of the current encoding block is 8x8, the residual signal for the 8x8 block, after the first transformation, second transformation, and quantization, for each of the four 4x4 sub-blocks, can be entropy-coded while scanning the quantized transform coefficients according to at least one of the three scanning order methods shown in FIG. 10. Also, the quantized transform coefficients can be entropy-decoded while being inverse-scanned. The inverse-scanned, quantized transform coefficients become the transform coefficients after inverse quantization, and a reconstructed residual signal can be generated by performing at least one of the second inverse transformation or the first inverse transformation.

[0185] In video encoding processing, as shown in FIG. 11, a block may be divided, and an indicator corresponding to the division information may be signaled. At this time, the division information may be at least one of a split flag, a quad / binary tree flag (QB_flag), a quadtree division flag (quadtree_flag), a binary tree division flag (binarytree_flag), and a binary tree division type flag (Btype_flag). Here, the split_flag is a flag indicating whether the block is divided; the QB_flag is a flag indicating whether the block is divided in a quadtree shape or a binary tree shape; the quadtree_flag is a flag indicating whether the block is divided in a quadtree shape; the binarytree_flag is a flag indicating whether the block is divided in a binary tree shape; and the Btype_flag may be a flag indicating whether the block is divided vertically or horizontally when it is divided in a binary tree shape.

[0186] When the split flag is 1, it can indicate that the block is divided, and when the split flag is 0, it can indicate that the block is not divided. In the case of the quad / binary tree flag, if it is 0, it indicates quadtree division, and if it is 1, it indicates binary tree division. Conversely, if it is 0, it indicates binary tree division, and if it is 1, it indicates quadtree division. In the case of the binary tree division type flag, if it is 0, it indicates horizontal division, and if it is 1, it indicates vertical division. Conversely, if it is 0, it indicates vertical division, and if it is 1, it indicates horizontal division.

[0187] For example, regarding FIG. 11, the division information can be derived by signaling at least one of quadtree_flag, binarytree_flag, and Btype_flag as shown in Table 3 below.

Table 3

[0188] For example, for the upper part of the split in FIG. 11, at least one of split_flag, QB_flag, and Btype_flag can be signaled and derived as shown in Table 4 below.

Table 4

[0189] The above splitting method can only be split in a quadtree shape or only in a binary tree shape according to the size / shape of the block. In such a case, the split_flag can mean a flag indicating whether it is a quadtree split or a binary tree split. The size / shape of the block can be derived according to the depth information of the block, and the depth information can be signaled.

[0190] When the size of the block belongs to a predetermined range, it may be possible to split only in a quadtree shape. Here, the predetermined range can be defined as at least one of the size of the largest block or the size of the smallest block that can be split only in a quadtree shape. Information indicating the size of the largest / smallest block allowed for the quadtree split can be signaled via a bitstream, and this information can be signaled in at least one unit of a sequence, a picture parameter, or a slice (segment). Or, the size of the largest / smallest block may be a predetermined fixed size in the encoder / decoder. For example, when the size of the block corresponds to 256x256 to 64x64, it is only possible to split in a quadtree shape. In such a case, the split_flag can be a flag indicating whether it is a quadtree split or not.

[0191] When the size of the block belongs to a predetermined range, it can only be divided in a binary tree shape. Here, the predetermined range can be defined as at least one of the size of the largest block or the size of the smallest block that can only be divided in a binary tree shape. Information indicating the size of the largest / smallest block for which the binary tree division is allowed can be signaled via a bitstream, and the corresponding information can be signaled in at least one unit of a sequence, a picture parameter, or a slice (segment). Or, the size of the largest / smallest block may be a predetermined fixed size for the encoder / decoder. For example, when the size of the block corresponds to 16x16 to 8x8, it can only be divided in a binary tree shape. In such a case, the split_flag can be a flag indicating whether it is a binary tree division.

[0192] After the one block is divided in a binary tree shape, if the divided block is further divided, it can only be divided in a binary tree shape.

[0193] When the horizontal width or vertical height of the divided block is a size that cannot be further divided, the one or more indicators may not be signaled.

[0194] In addition to the binary tree division of the quadtree base, after the binary tree division, the quadtree base can be divided.

[0195] Based on the above matters, the image encoding / decoding method according to the present invention will be described in detail.

[0196] FIG. 12 is a flowchart showing an image encoding method using the merge mode according to the present invention, and FIG. 13 is a flowchart showing an image decoding method using the merge mode according to the present invention.

[0197] Referring to FIG. 12, the encoding device can derive merge candidates (S1201) and generate a merge candidate list based on the derived merge candidates. When the merge candidate list is generated, motion information is determined using the generated merge candidate list (S1202), and motion compensation for the current block can be performed using the determined motion information (S1203). Thereafter, the encoding device can entropy-encode information related to motion compensation (S1204).

[0198] Referring to FIG. 13, the decoding device entropy-decodes information related to motion compensation received from the encoding device (S1301), derives merge candidates (S1302), and can generate a merge candidate list based on the derived merge candidates. When the merge candidate list is generated, the motion information of the current block can be determined using the generated merge candidate list (S1303). Thereafter, the decoding device can perform motion compensation using the motion information (S1304).

[0199] Hereinafter, each step shown in FIGS. 12 and 13 will be described in detail.

[0200] First, the step of deriving merge candidates will be specifically described (S1201, S1302).

[0201] The merge candidates for the current block can include at least one of a spatial merge candidate, a temporal merge candidate, or an additional merge candidate.

[0202] The spatial merge candidate of the current block can be derived from the reconstructed blocks in the vicinity of the current block. As an example, the motion information of the reconstructed blocks in the vicinity of the current block can be determined as the spatial merge candidate for the current block. Here, the motion information can include at least one of a motion vector, a reference picture index, or a prediction list utilization flag.

[0203] In this case, the motion information of the spatial merge candidate can have not only the motion information corresponding to L0 and L1, but also the motion information corresponding to L0, L1, ..., LX. Here, X can be a positive integer including 0. Therefore, the reference picture list can include at least one of L0, L1, ..., LX, etc.

[0204] FIG. 14 is a diagram for explaining an example of deriving the spatial merge candidate of the current block. Here, deriving the spatial merge candidate can mean deriving the spatial merge candidate and adding it to the merge candidate list.

[0205] Referring to FIG. 14, the spatial merge candidate of the current block can be derived from the neighboring blocks adjacent to the current block X. The neighboring blocks adjacent to the current block can include at least one of the block B1 adjacent to the upper side of the current block, the block A1 adjacent to the left side of the current block, the block B0 adjacent to the upper right corner of the current block, the block B2 adjacent to the upper left corner of the current block, and the block A0 adjacent to the lower left corner of the current block. On the other hand, the neighboring blocks adjacent to the current block can be square or non-square.

[0206] In order to derive the spatial merge candidate of the current block, it is possible to determine whether the neighboring block adjacent to the current block can be used for deriving the spatial merge candidate of the current block. At this time, whether the neighboring block adjacent to the current block can be used for deriving the spatial merge candidate of the current block can be determined based on a predetermined priority order. As an example, in the example shown in FIG. 14, the availability of deriving the spatial merge candidate can be determined in the order of the blocks at the positions of A1, B1, B0, A0, and B2. The spatial merge candidate determined based on the availability determination order can be sequentially added to the merge candidate list of the current block. Hereinafter, an example of a neighboring block that cannot be used for deriving the spatial merge candidate of the current block is shown. 1) When the neighboring block is the block at the B2 position and all spatial merge candidates have been derived from the blocks at the A0, A1, B0, and B1 positions 2) When there is no neighboring block (when the current block exists at a picture boundary, slice boundary, or tile boundary, etc.) 3) When the neighboring block is intra-coded 4) When the previously derived spatial merge candidate of the neighboring block is the same as at least one of the motion vector, reference picture index, and reference picture 5) When the motion vector of the neighboring block refers to an external region of at least one of the pictures, slices, or tiles containing the current block

[0207] FIG. 15 is a diagram for explaining an example in which a spatial merge candidate is added to a merge candidate list.

[0208] Referring to FIG. 15, when four spatial merge candidates are derived from neighboring blocks at positions A1, B0, A0, and B2, the derived spatial merge candidates can be sequentially added to the merge candidate list.

[0209] maxNumSpatialMergeCand means the maximum number of spatial merge candidates that can be included in the merge candidate list, and numMergeCand can mean the number of merge candidates included in the merge candidate list. maxNumSpatialMergeCand can be a positive integer including 0. maxNumSpatialMergeCand can be preset to use the same value in the encoding device and the decoding device. Alternatively, the encoding device can also encode the maximum number of merge candidates that can be included in the merge candidate list of the current block and signal it to the decoding device via a bitstream.

[0210] As described above, when at least one spatial merge candidate is derived from the neighboring blocks A1, B1, B0, A0, and B2, spatial merge candidate flag information (spatialCand) indicating whether each derived merge candidate is a spatial merge candidate can be set. As an example, when a spatial merge candidate is derived, spatialCand can be set to a predetermined value of 1, and when not, it can be set to a predetermined value of 0. Also, each time a spatial merge candidate is derived, the spatial merge candidate count (spatialCandCnt) can be incremented by 1.

[0211] The spatial merge candidate can be derived based on at least one of the encoding parameters of the current block or the neighboring blocks.

[0212] Based on the size or depth of the block for which the information related to motion compensation is entropy-encoded / decoded, the spatial merge candidate can be shared in a block with a size or depth smaller than the size or depth of the block for which the information related to motion compensation is entropy-encoded / decoded. Here, the information related to motion compensation can be at least one of skip mode usage information, merge mode usage information, or merge index information.

[0213] The block for which the information related to motion compensation is entropy-encoded / decoded can be a CTU or a lower unit of the CTU, a CU, or a PU.

[0214] Hereinafter, for convenience of explanation, the size of the block for which the information related to motion compensation is entropy-encoded / decoded is defined as the first block size, and the depth of the block for which the information related to motion compensation is entropy-encoded / decoded is defined as the depth of the first block.

[0215] Specifically, when the size of the current block is smaller than the first block size, a spatial merge candidate for the current block can be derived from at least one of the blocks reconstructed in the vicinity of the upper block having the first block size. And the blocks included inside the upper block can share the derived spatial merge candidate. Here, a block having the first block size can be used as the upper block of the current block.

[0216] FIG. 16 is a diagram for explaining an embodiment of deriving and sharing a spatial merge candidate in a CTU. Referring to FIG. 16, when the first block size is 32x32, blocks 1601, 1602, 1603, and 1604 having a block size smaller than 32x32 can derive a spatial merge candidate from at least one of the adjacent neighboring blocks of the upper block 1600 having the first block size and share the derived spatial merge candidate.

[0217] As an example, when the first block size is 32x32 and the block size of the encoded block is 32x32, a prediction block having a block size smaller than 32x32 can derive a spatial merge candidate for the prediction block from at least one of the motion information of the neighboring blocks of the encoded block, and the prediction blocks within the encoded block can share the derived spatial merge candidate. Here, the encoded block and the prediction block can mean a block in a more generalized expression.

[0218] When the block depth of the current block is deeper than the first block depth, a spatial merge candidate can be derived from at least one of the blocks reconstructed in the vicinity of the upper block having the first block depth. And the blocks included inside the upper block can share the derived spatial merge candidate. Here, a block having the first block depth can be used as the upper block of the current block.

[0219] As an example, when the depth of the first block is 2 and the depth of the encoded block is 2, a prediction block having a depth deeper than the block depth of 2 can derive a spatial merge candidate of the prediction block based on at least one of the motion information of the neighboring blocks of the encoded block, and the prediction block within the encoded block can share the derived spatial merge candidate.

[0220] Here, sharing the spatial merge candidate can mean that, based on the same spatial merge candidate, a merge candidate list for each of the blocks to be shared can be generated.

[0221] Also, sharing the spatial merge candidate can mean that the blocks to be shared can perform motion compensation using one merge candidate list. Here, the shared merge candidate list can include at least one of the spatial merge candidates derived based on the upper block for which the information related to motion compensation is entropy encoded / decoded.

[0222] The neighboring block adjacent to the current block or the current block can have a square or non-square shape.

[0223] Then, the neighboring blocks adjacent to the current block can be divided into sub-block units. In this case, the motion information of any one of the sub-blocks of the neighboring blocks adjacent to the current block can be determined as the spatial merge candidate of the current block. Also, based on at least one of the motion information of the sub-blocks of the neighboring blocks adjacent to the current block, the spatial merge candidate of the current block can be determined. Here too, it is possible to determine whether the sub-blocks of the neighboring blocks can be used for deriving the spatial merge candidate, and determine it as the spatial merge candidate of the current block. Whether it can be used for deriving the spatial merge candidate can include at least one of whether the motion information of the sub-blocks of the neighboring blocks exists and whether the motion information of the sub-blocks of the neighboring blocks is available as the spatial merge candidate of the current block.

[0224] Also, any one of the median value, average value, minimum value, maximum value, weighted average value or mode value of at least one (for example, motion vector) of the motion information of the sub-blocks of the neighboring blocks can be determined as the spatial merge candidate of the current block.

[0225] Next, a method for deriving the temporal merge candidate of the current block will be described.

[0226] The temporal merge candidate of the current block can be derived from the reconstructed block included in the co-located picture of the current image. Here, the co-located picture is an image for which encoding / decoding has been completed before the current image, and can be an image having a different temporal order from the current image.

[0227] FIG. 17 is a diagram for explaining an example of deriving the temporal merge candidate of the current block. Here, deriving the temporal merge candidate can mean deriving the temporal merge candidate and adding it to the merge candidate list.

[0228] Referring to FIG. 17, in the collocated picture of the current image, the temporal merge candidate of the current block can be derived from a block including an external position of a block corresponding to a position spatially identical to the current block X, or a block including an internal position of a block corresponding to a position spatially identical to the current block X. Here, the temporal merge candidate can mean the motion information of the corresponding position block. As an example, the temporal merge candidate of the current block X can be derived from block H adjacent to the lower right corner of block C corresponding to a position spatially identical to the current block, or block C3 including the center point of block C. Blocks such as block H or block C3 used to derive the temporal merge candidate of the current block can be referred to as "collocated blocks".

[0229] On the other hand, the corresponding position block of the current block or the current block can be in a square or non-square shape.

[0230] When the temporal merge candidate of the current block can be derived from block H including the external position of block C, block H can be set as the corresponding position block of the current block. In this case, the temporal merge candidate of the current block can be derived based on the motion information of block H. On the contrary, when the temporal merge candidate of the current block cannot be derived from block H, block C3 including the internal position of block C can be set as the corresponding position block of the current block. In this case, the temporal merge candidate of the current block can be derived based on the motion information of block C3. If the temporal merge of the current block cannot be derived from both block H and block C3 (for example, when both block H and block C3 are intra-coded), the temporal merge candidate for the current block may not be derived, or it may be derived from a block at a position different from block H and block C3.

[0231] As another example, the time merge candidates for the current block can also be derived from a plurality of blocks within the corresponding position image. As an example, a plurality of time merge candidates for the current block can also be derived from block H and block C3.

[0232] FIG. 18 is a diagram for explaining an example in which a time merge candidate is added to a merge candidate list.

[0233] Referring to FIG. 18, when one time merge candidate is derived from the corresponding position block at the H1 position, the derived time merge candidate can be added to the merge candidate list.

[0234] The corresponding position block of the current block can be divided into sub-block units. In this case, among the sub-blocks of the corresponding position block of the current block, the motion information of any one of the sub-blocks can be determined as the time merge candidate of the current block. Also, the time merge candidate of the current block can be determined based on at least one of the motion information of the sub-blocks of the corresponding position block of the current block.

[0235] Here too, it is possible to determine whether there is motion information of the sub-blocks of the corresponding position block, or whether the motion information of the sub-blocks of the corresponding position block is available as the time merge candidate of the current block, and determine it as the time merge candidate of the current block.

[0236] Also, any one of the median value, average value, minimum value, maximum value, weighted average value, or most frequent value of at least one (for example, motion vector) of the motion information of the sub-blocks of the corresponding position block can be determined as the time merge candidate of the current block.

[0237] In FIG. 17, it is illustrated that the temporal merge candidate of the current block can be derived from a block adjacent to the lower right corner of the corresponding position block or a block including the center point of the corresponding position block. However, the position of the block for deriving the temporal merge candidate of the current block is not limited to the example shown in FIG. 17. As an example, the temporal merge candidate of the current block may be derived from a block adjacent to the upper / lower boundary, left / right boundary or one corner of the corresponding position block, or may be derived from a block including a specific position within the corresponding position block (for example, a block adjacent to the corner boundary of the corresponding position block).

[0238] The temporal merge candidate of the current block may be determined in consideration of the reference picture lists (or prediction directions) of the current block and the corresponding position block. On the other hand, the motion information of the temporal merge candidate can have motion information corresponding not only to L0 and L1, but also to L0, L1, …, LX. Here, X can be a positive integer including 0.

[0239] As an example, when the reference picture list available for the current block is L0 (that is, when the inter prediction indicator indicates PRED_L0), the motion information corresponding to L0 in the corresponding position block can be derived as the temporal merge candidate of the current block. That is, when the reference picture list available for the current block is LX (where X is an integer indicating the index of the reference picture list, such as 0, 1, 2 or 3), the motion information corresponding to LX in the corresponding position block (hereinafter referred to as "LX motion information") can be derived as the temporal merge candidate of the current block.

[0240] Even when the current block uses a plurality of reference picture lists, the temporal merge candidate of the current block can be determined in consideration of the reference picture lists of the current block and the corresponding position block.

[0241] As an example, when the current block performs bidirectional prediction (i.e., when the inter prediction indicator is PRED_BI), at least two of the L0 motion information, L1 motion information, L2 motion information, …, LX motion information of the corresponding position block can be derived as temporal merge candidates. When the current block performs three-directional prediction (i.e., when the inter prediction indicator is PRED_TRI), at least three of the L0 motion information, L1 motion information, L2 motion information, …, LX motion information of the corresponding position block can be derived as temporal merge candidates. When the current block performs four-directional prediction (i.e., when the inter prediction indicator is PRED_QUAD), at least four of the L0 motion information, L1 motion information, L2 motion information, …, LX motion information of the corresponding position block can be derived as temporal merge candidates.

[0242] Also, at least one of the temporal merge candidates, the corresponding position image, the corresponding position block, the prediction list utilization flag, and the reference image index can be derived based on at least one of the encoding parameters of the current block, the neighboring block, or the corresponding position block.

[0243] The temporal merge candidates can be preliminarily derived when the number of derived spatial merge candidates is smaller than the maximum number of merge candidates. Thereby, when the number of derived spatial merge candidates reaches the maximum number of merge candidates, the process of deriving the temporal merge candidates can be omitted.

[0244] As an example, when the maximum number of merge candidates is two and the two derived spatial merge candidates have different values from each other, the process of deriving the temporal merge candidates can be omitted.

[0245] As another example, the temporal merge candidates of the current block can also be derived based on the maximum number of temporal merge candidates. Here, the maximum number of temporal merge candidates can be preset so that the same value is used in the encoding device and the decoding device. Alternatively, information indicating the maximum number of temporal merge candidates of the current block can be encoded via the bitstream and signaled to the decoding device. As an example, the encoding device can encode "maxNumTemporalMergeCand" indicating the maximum number of temporal merge candidates of the current block and signal it to the decoding device via the bitstream. At this time, "maxNumTemporalMergeCand" can be set as a positive integer including 0. For example, "maxNumTemporalMergeCand" can be set to 1. The value of maxNumTemporalMergeCand may be variably derived based on information regarding the number of signaled temporal merge candidates, or may be a fixed value preset in the encoder / decoder.

[0246] When the distance between the current image containing the current block and the reference image of the current block is different from the distance between the corresponding position image containing the corresponding position block and the reference image of the corresponding position block, the motion vector of the temporal merge candidate of the current block can be obtained by scaling the motion vector of the corresponding position block. Here, the scaling can be performed based on at least one of the distance between the current image and the reference image referred to by the current block and the distance between the corresponding position image and the reference image referred to by the corresponding position block. As an example, the motion vector of the temporal merge candidate of the current block can be derived by scaling the motion vector of the corresponding position block according to the ratio of the distance between the current image and the reference image referred to by the current block and the distance between the corresponding position image and the reference image referred to by the corresponding position block.

[0247] Based on the size of the block (the first block size) or the depth of the block (the first block depth) for which information related to motion compensation is entropy-coded / decoded, time merge candidates can be shared among blocks with a size smaller than or a depth deeper than the block size or block depth for which information related to motion compensation is entropy-coded / decoded. Here, the information related to motion compensation can be at least one of information on whether to use the skip mode, information on whether to use the merge mode, or merge index information.

[0248] The block for which information related to motion compensation is entropy-coded / decoded can be a CTU or a lower unit of a CTU, a CU, or a PU.

[0249] Specifically, when the size of the current block is smaller than the first block size, the time merge candidates of the current block can be derived from the corresponding position block of the upper block having the first block size. And the blocks included inside the upper block can share the derived time merge candidates.

[0250] Also, when the block depth of the current block is deeper than the first block depth, the time merge candidates can be derived from the corresponding position block of the upper block having the first block depth. And the blocks included inside the upper block can share the derived time merge candidates.

[0251] Here, sharing time merge candidates can mean that, based on the same time merge candidates, merge candidate lists for each of the blocks to be shared can be generated.

[0252] Also, sharing time merge candidates can mean that the blocks to be shared can perform motion compensation using one merge candidate list. Here, the shared merge candidate list can include time merge candidates derived based on the upper block for which information related to motion compensation is entropy-coded / decoded.

[0253] FIG. 19 shows an example of scaling a motion vector in the motion information of a corresponding position block in order to derive a time merge candidate for a current block.

[0254] The motion vector of the corresponding position block can be scaled based on at least one of a difference value (td) between a POC (Picture order count) indicating the display order of the corresponding position image and the POC of the reference image of the corresponding position block, and a difference value (tb) between the POC of the current image and the POC of the reference image of the current block.

[0255] Prior to performing the scaling, td or tb can be adjusted so that td or tb exists within a predetermined range. As an example, when the predetermined range is -128 to 127, if td or tb is smaller than -128, td or tb can be adjusted to -128. If td or tb is larger than 127, td or tb can be adjusted to 127. If td or tb is within the range of -128 to 127, td or tb is not adjusted.

[0256] A scaling factor DistScaleFactor can be calculated based on td or tb. At this time, the scaling factor can be calculated based on the following Equation 1.

Equation

[0257] In the formula, Abs() represents an absolute value function, and the output value of the corresponding function is the absolute value of the input value.

[0258] The value of the scaling factor DistScaleFactor calculated based on Equation 1 can be adjusted to a predetermined range. As an example, DistScaleFactor can be adjusted to exist within the range of -1024 to 1023.

[0259] By scaling the motion vector of the corresponding position block using a scaling factor, the motion vector of the temporal merge candidate for the current block can be determined. As an example, the motion vector of the temporal merge candidate for the current block can be determined by the following Equation 2.

Equation

[0260] In the formula, Sign() is a function that outputs the sign information of the value contained in (). As an example, if it is Sign(-1), - is output. In Equation 2 above, mvCol can mean the motion vector of the corresponding position block.

[0261] Next, a method for deriving additional merge candidates for the current block will be described.

[0262] The additional merge candidate can mean at least one of a modified spatial merge candidate, a modified temporal merge candidate, a combined merge candidate, and a merge candidate having a predetermined motion information value. Here, deriving an additional merge candidate can mean deriving an additional merge candidate and adding it to the merge candidate list.

[0263] The modified spatial merge candidate can mean a merge candidate that changes at least one of the motion information of the derived spatial merge candidate.

[0264] The modified temporal merge candidate can mean a merge candidate that changes at least one of the motion information of the derived temporal merge candidate.

[0265] The combined merge candidate can mean a merge candidate derived by combining at least one of the motion information of the spatial merge candidate, the temporal merge candidate, the changed spatial merge candidate, the changed temporal merge candidate, the combined merge candidate, and the merge candidate having a predetermined motion information value existing in the merge candidate list.

[0266] Alternatively, the combined merge candidate can mean a merge candidate derived by combining at least one of the motion information of the spatial merge candidate and the temporal merge candidate that do not exist in the merge candidate list but can be derived from a block capable of deriving at least one of them, the derived spatial merge candidate and the derived temporal merge candidate, the changed spatial merge candidate, the changed temporal merge candidate, the combined merge candidate, and the merge candidate generated based on this.

[0267] Alternatively, the combined merge candidate can be derived using the motion information entropy-decoded from the bitstream by the decoder. At this time, in the encoder, the motion information used for deriving the combined merge candidate can be entropy-coded into the bitstream.

[0268] The combined merge candidate can mean a combined bi-prediction merge candidate. The combined bi-prediction merge candidate is a merge candidate using bi-prediction and can mean a merge candidate having L0 motion information and L1 motion information.

[0269] Also, the combined merge candidate can mean a merge candidate having at least N of L0 motion information, L1 motion information, L2 motion information, and L3 motion information. Here, N can mean a positive integer of 2 or more.

[0270] The merge candidate having a predetermined motion information value can mean a zero merge candidate whose motion vector is (0,0). On the other hand, the merge candidate having a predetermined motion information value may be preset so that the same value is used in the encoding device and the decoding device.

[0271] Based on at least one of the encoding parameters of the current block, neighboring blocks, or corresponding position blocks, at least one of a modified spatial merge candidate, a modified temporal merge candidate, a combined merge candidate, and a merge candidate having a predetermined motion information value can be derived or generated. Also, at least one of a modified spatial merge candidate, a modified temporal merge candidate, a combined merge candidate, and a merge candidate having a predetermined motion information value can be added to a merge candidate list based on at least one of the encoding parameters of the current block, neighboring blocks, or corresponding position blocks.

[0272] Additional merge candidates can be derived for each sub-block of the current block, neighboring blocks, or corresponding position blocks, and the merge candidates derived for each sub-block can be added to the merge candidate list of the current block.

[0273] Additional merge candidates can be derived only in the case of a B slice / B picture or a slice / picture using M or more reference picture lists. Here, M may be 3 or 4, and can mean a positive integer of 3 or more.

[0274] Additional merge candidates can be derived up to a maximum of N. At this time, N is a positive integer including 0. N may be a variable value derived based on information regarding the maximum number of merge candidates included in the merge candidate list. Or, it may be a fixed value preset in the encoder / decoder. Here, N may vary according to the size, shape, depth, or position of the block encoded / decoded in the merge mode.

[0275] The merge candidate list has a preset size and can increase by the number of additional merge candidates generated after adding spatial merge candidates or temporal merge candidates. In this case, any of the generated additional merge candidates can be included in the merge candidate list. On the other hand, the size of the merge candidate list can also increase to a size smaller than the number of additional merge candidates (for example, the number of additional merge candidates - N, where N is a positive integer). In this case, only some of the generated additional merge candidates can be included in the merge candidate list.

[0276] Also, the size of the merge candidate list can be determined based on the encoding parameters of the current block, neighboring blocks, or corresponding position blocks, and the size can be changed based on the encoding parameters.

[0277] To increase the throughput of the merge mode in the encoder and decoder, combined merge candidates are not derived, and only the derivation of spatial merge candidates, the derivation of temporal merge candidates, and the derivation of zero merge candidates are performed to perform motion compensation using the merge mode. If the combined merge candidate derivation process is performed after the execution of the temporal merge candidate derivation process that requires a relatively large cycle time, if the combined merge candidate derivation process is not performed, the hardware complexity of the merge mode can, in the worst case, be the temporal merge candidate derivation process instead of the combined merge candidate derivation process after the temporal merge candidate derivation process. Therefore, in the merge mode, the cycle time required to derive each merge candidate can be reduced. Also, the merge mode that does not derive combined merge candidates has the advantage that there is no dependency between the derivations of each merge candidate, so the derivation of spatial merge candidates, the derivation of temporal merge candidates, and the derivation of zero merge mode candidates can be performed in parallel.

[0278] FIG. 21 (FIGS. 21a and 21b) is a diagram for explaining an embodiment of a method for deriving a combined merge candidate. The method for deriving the combined merge candidate of FIGS. 21a and 21b can be performed when there is one or more merge candidates in the merge candidate list, or when the number of merge candidates (numOrigMergeCand) in the merge candidate list is smaller than the maximum number of merge candidates (MaxNumMergeCand) before deriving the combined merge candidate.

[0279] Referring to FIGS. 21a and 21b, the encoder / decoder can set the input number of merge candidates (numInputMergeCand) as the number of merge candidates (numMergeCand) in the current merge candidate list, and set the combination index (combIdx) to 0. The k (numMergeCand numInputMergeCand)-th combined merge candidate can be derived.

[0280] The encoder / decoder can derive at least one of the L0 candidate index (l0CandIdx), L1 candidate index (l1CandIdx), L2 candidate index (l2CandIdx), and L3 candidate index (l3CandIdx) using the combination index as shown in FIG. 20 (S2101).

[0281] Each candidate index indicates a merge candidate in the merge candidate list, and the motion information of the candidate index according to L0, L1, L2, L3 can become the motion information for L0, L1, L2, L3 of the combined merge candidate.

[0282] The symbolizer / decoder can derive the L0 candidate (l0Cand) as the merge candidate corresponding to the L0 candidate index in the merge candidate list (mergeCandList[l0CandIdx]), the L1 candidate (l1Cand) as the merge candidate corresponding to the L1 candidate index in the merge candidate list (mergeCandList[l1CandIdx]), the L2 candidate (l2Cand) as the merge candidate corresponding to the L2 candidate index in the merge candidate list (mergeCandList[l2CandIdx]), and the L3 candidate (l3Cand) as the merge candidate corresponding to the L3 candidate index in the merge candidate list (mergeCandList[l3CandIdx]) (S2102).

[0283] If the symbolizer / decoder satisfies at least one of the following cases, it can perform step S2104; otherwise, it can perform step S2105 (S2103). 1) When the L0 candidate uses L0 single prediction (predFlagL0l0Cand == 1) 2) When the L1 candidate uses L1 single prediction (predFlagL1l1Cand == 1) 3) When the L2 candidate uses L2 single prediction (predFlagL2l2Cand == 1) 4) When the L3 candidate uses L3 single prediction (predFlagL3l3Cand == 1) 5) When at least one reference picture among the L0, L1, L2, and L3 candidates is different from the reference pictures of other candidates, and at least one motion vector among the L0, L1, L2, and L3 candidates is different from the motion vectors of other candidates

[0284] If at least one of these five cases is satisfied (S2103 - Yes), the encoder / decoder determines the L0 motion information of the L0 candidate as the combined candidate's L0 motion information, determines the L1 motion information of the L1 candidate as the combined candidate's L1 motion information, determines the L2 motion information of the L2 candidate as the combined candidate's L2 motion information, determines the L3 motion information of the L3 candidate as the combined candidate's L3 motion information, and can add the combined merge candidate (combCandk) to the merge candidate list (S2104).

[0285] For example, the information about the combined merge candidate is as follows. The L0 reference picture index (refIdxL0combCandk) of the K - th combined merge candidate = the L0 reference picture index (refIdxL0l0Cand) of the L0 candidate The L1 reference picture index (refIdxL1combCandk) of the K - th combined merge candidate = the L1 reference picture index (refIdxL1l1Cand) of the L1 candidate The L2 reference picture index (refIdxL2combCandk) of the K - th combined merge candidate = the L2 reference picture index (refIdxL2l2Cand) of the L2 candidate The L3 reference picture index (refIdxL3combCandk) of the K - th combined merge candidate = the L3 reference picture index (refIdxL3l3Cand) of the L3 candidate The L0 prediction list utilization flag (predFlagL0combCandk) of the K - th combined merge candidate = 1 The L1 prediction list utilization flag (predFlagL1combCandk) of the K - th combined merge candidate = 1 The L2 prediction list utilization flag (predFlagL2combCandk) of the K - th combined merge candidate = 1 The L3 prediction list utilization flag (predFlagL3combCandk) of the K - th combined merge candidate = 1 The x component of the L0 motion vector of the K-th combination merge candidate (mvL0combCandk[0]) = the x component of the L0 motion vector of the L0 candidate (mvL0l0Cand[0]) The y component of the L0 motion vector of the K-th combination merge candidate (mvL0combCandk[1]) = the y component of the L0 motion vector of the L0 candidate (mvL0l0Cand[1]) The x component of the L1 motion vector of the K-th combination merge candidate (mvL1combCandk[0]) = the x component of the L1 motion vector of the L1 candidate (mvL1l1Cand[0]) The y component of the L1 motion vector of the K-th combination merge candidate (mvL1combCandk[1]) = the y component of the L1 motion vector of the L1 candidate (mvL1l1Cand[1]) The x component of the L2 motion vector of the K-th combination merge candidate (mvL2combCandk[0]) = the x component of the L2 motion vector of the L2 candidate (mvL2l2Cand[0]) The y component of the L2 motion vector of the K-th combined merge candidate (mvL2combCandk[1]) = the y component of the L2 motion vector of the L2 candidate (mvL2l2Cand[1]) The x component of the L3 motion vector of the K-th combination merge candidate (mvL3combCandk[0]) = the x component of the L3 motion vector of the L3 candidate (mvL3l3Cand[0]) The y component of the L3 motion vector of the K-th combination merge candidate (mvL3combCandk[1]) = the y component of the L3 motion vector of the L3 candidate (mvL3l3Cand[1]) numMergeCand = numMergeCand + 1

[0286] Also, the encoder / decoder can increment the combination index by 1 (S2105).

[0287] Also, when the combination index is the same as (numOrigMergeCand*(numOrigMergeCand - 1)) or the number of merge candidates (numMergeCand) in the current merge candidate list is the same as the maximum number of merge candidates (MaxNumMergeCand), the combination merge candidate derivation step ends. Otherwise, it can return to step S2101 (S2106).

[0288] When the combination merge candidate derivation method of FIGS. 21a and 21b is performed, as shown in FIG. 22, the combination merge candidates derived to the merge candidate list can be added.

[0289] On the other hand, when there are two or more spatial merge candidates in the merge candidate list or when the number of merge candidates (numOrigMergeCand) in the merge candidate list is smaller than the maximum number of merge candidates (MaxNumMergeCand) before the combination merge candidate derivation, a method of deriving combination merge candidates using only spatial merge candidates can be performed. Also in this case, it can be performed using the combination merge candidate derivation method of FIGS. 21a and 21b.

[0290] However, the L0 candidate index, L1 candidate index, L2 candidate index, and L3 candidate index derived in step S2101 of FIG. 21a can indicate only the merge candidates for which the spatial merge candidate flag information (spatialCand) is 1.

[0291] Therefore, the L0 candidate, L1 candidate, L2 candidate, and L3 candidate derived in step S2102 can be derived using only the merge candidates for which the spatial merge candidate flag information (spatialCand) is 1 in the merge candidate list, that is, the spatial merge candidates.

[0292] Then, in step S2106 of FIG. 21b, by comparing the combination index with the value of (spatialCandCnt*(spatialCandCnt - 1)) instead of the value of (numOrigMergeCand*(numOrigMergeCand - 1)), if the combination index is the same as (spatialCandCnt*(spatialCandCnt - 1)), or if the number of merge candidates (numMergeCand) in the current merge candidate list is the same as "MaxNumMergeCand", the combined merge candidate derivation step is terminated; otherwise, it is possible to return to step S2101.

[0293] When a method of deriving combined merge candidates using only spatial merge candidates is performed, as shown in FIG. 23, combined merge candidates combined only from spatial merge candidates can be added to the merge candidate list.

[0294] FIGS. 22 and 23 show an example of deriving combined merge candidates using at least one of spatial merge candidates, temporal merge candidates, and zero merge candidates and adding them to the merge candidate list.

[0295] Here, the merge candidate list may include merge candidates having at least one piece of motion information among L0 motion information, L1 motion information, L2 motion information, and L3 motion information. On the other hand, although examples of L0, L1, L2, and L3 reference image lists have been described, the present invention is not limited thereto, and merge candidates having motion information for L0 to LX reference image lists (X is a positive integer) may be included in the merge candidate list.

[0296] Each piece of motion information may include at least one of a motion vector, a reference image index, and a prediction list utilization flag.

[0297] As shown in FIGS. 22 and 23, at least one of the merge candidates can be determined as the final merge candidate. The determined final merge candidate can be used as the motion information of the current block. The motion information can be used for inter prediction or motion compensation of the current block. Also, at least one value of the information corresponding to the motion information of the current block can be changed to use the motion information for inter prediction or motion compensation of the current block. Here, among the information corresponding to the motion information, the value to be changed can be at least one of the x component of the motion vector, the y component of the motion vector, and the reference picture index. Also, when changing at least one value of the information corresponding to the motion information, at least one value of the information corresponding to the motion information can be changed so as to show the minimum distortion using a distortion calculation method (SAD, SSE, MSE, etc.).

[0298] Based on the motion information of the merge candidate, at least one of the L0 motion information, L1 motion information, L2 motion information, and L3 motion information is used to generate a predicted block for the current block, and the generated predicted block can be used for inter prediction or motion compensation of the current block.

[0299] The inter prediction indicator can be expressed as PRED_LX, which is a one-way prediction indicating PRED_L0 or PRED_L1, or PRED_BI_LX, which is a bidirectional prediction for the reference picture list X, when at least one of the L0 motion information, L1 motion information, L2 motion information, and L3 motion information is used for generating the predicted block. Here, X means a positive integer including 0, and can be 0, 1, 2, 3, etc.

[0300] Also, the inter prediction indicator can be expressed as PRED_TRI, which is a three-way prediction, when at least three of the L0 motion information, L1 motion information, L2 motion information, and L3 motion information are used. Also, the inter prediction indicator can be expressed as PRED_QUAD, which is a four-way prediction, when at least four of the L0 motion information, L1 motion information, L2 motion information, and L3 motion information are used.

[0301] For example, when the inter prediction indicator for reference picture list L0 is PRED_L0 and the inter prediction indicator for reference picture list L1 is PRED_BI_L1, the inter prediction indicator of the current block can be PRED_TRI. That is, the sum of the number of predicted blocks indicated by the inter prediction indicator for each reference picture list can be the inter prediction indicator of the current block.

[0302] Also, the reference picture list can be at least one such as L0, L1, L2, L3, etc., and a merge candidate list as shown in FIGS. 22 and 23 can be generated for each reference picture list. Therefore, when generating a predicted block for the current block, at least 1 to a maximum of N predicted blocks can be generated and used for inter prediction or motion compensation for the current block. Here, N means a positive integer of 1 or more and can be 1, 2, 3, 4, etc.

[0303] In order to reduce the memory bandwidth and improve the processing speed, it can be used to derive a combined merge candidate only when at least one of the reference picture indexes or motion vector values of the merge candidates is the same as other merge candidates or is included within a predetermined range.

[0304] As an example, a combined merge candidate can be derived using a merge candidate whose reference picture index is equal to a predetermined value among the merge candidates included in the merge candidate list. At this time, the predetermined value can be a positive integer including 0.

[0305] As another example, a combined merge candidate can be derived using a merge candidate whose reference picture index is included within a predetermined range among the merge candidates included in the merge candidate list. At this time, the predetermined range can be a range of positive integer values including 0.

[0306] As another example, among the merge candidates included in the merge candidate list, a combined merge candidate can be derived using a merge candidate whose motion vector value is within a predetermined range. At this time, the predetermined range can be a range of positive integer values including 0.

[0307] As another example, among the merge candidates included in the merge candidate list, a combined merge candidate can be derived using a merge candidate whose value of the difference between the motion vectors between the merge candidates is within a predetermined range. At this time, the predetermined range can be a range of positive integer values including 0.

[0308] Here, at least one of the predetermined value and the predetermined range can be determined based on a value commonly set in the encoder / decoder. Also, at least one of the predetermined value and the predetermined range can be determined based on an entropy-encoded / decoded value.

[0309] Also, when deriving a modified spatial merge candidate, a modified temporal merge candidate, or a merge candidate having a predetermined motion information value, it can be derived and added to the merge candidate list only when at least one of the reference picture index or the motion vector value of the merge candidate is the same as or within a predetermined range of another merge candidate.

[0310] FIG. 24 is a diagram for explaining the advantage of deriving a combined merge candidate using only a spatial merge candidate during motion compensation using the merge mode.

[0311] Referring to FIG. 24, in order to increase the throughput of the merge mode in the encoder and decoder, the combined merge candidates can be derived using only the spatial merge candidates without using the temporal merge candidates. The temporal merge candidate derivation process requires relatively more cycle time due to the execution of motion vector scaling than the spatial merge candidate derivation process. Therefore, when the combined merge candidate derivation process is performed after the execution of the temporal merge candidate derivation process, a large amount of cycle time is required when determining motion information using the merge mode.

[0312] However, when deriving the combined merge candidates using only the spatial merge candidates without using the temporal merge candidates, the combined merge candidate derivation process can be performed immediately after the derivation of the spatial merge candidates, which takes relatively less cycle time than the temporal merge candidate derivation process. Therefore, the cycle time required when determining motion information using the merge mode can be reduced compared to the method including the temporal merge candidate derivation process.

[0313] That is, by removing the dependency between the temporal merge candidates and the derivation of the combined merge candidates, the throughput of the merge mode can be increased. Also, when an error occurs in the reference picture due to a transmission error or the like, the error resiliency of the decoder can also be increased by deriving the combined merge candidates using only the spatial merge candidates instead of the temporal merge candidates.

[0314] Also, when using the method of deriving the combined merge candidates using only the spatial merge candidates without using the temporal merge candidates, the method of deriving the combined merge candidates using the temporal merge candidates and the method of deriving the combined merge candidates without using the temporal merge candidates can operate similarly to each other, and each method can be realized similarly to each other. Therefore, there is an advantage that the hardware logic is integrated.

[0315] FIG. 25 is a diagram for explaining an embodiment of a combined bi-predictive merge candidate splitting method. Here, the combined bi-predictive merge candidate can be a combined merge candidate including two pieces of motion information among the L0 motion information, ..., LX motion information. Hereinafter, FIG. 25 will be described assuming that the combined bi-predictive merge candidate includes the L0 motion information and the L1 motion information.

[0316] Referring to FIG. 25, the encoder / decoder can add each piece of motion information generated by splitting the information of the combined bi-predictive merge candidate in the merge candidate list into the L0 motion information and the L1 motion information to the merge candidate list as a new merge candidate.

[0317] Specifically, the encoder / decoder can determine a combined bi-predictive merge candidate to be split using a split index (splitIdx) from the merge candidate list (S2501). Here, the split index (splitIdx) can be index information indicating a combined bi-predictive merge candidate to be split.

[0318] The encoder / decoder can set the L0 motion information of the combined bi-predictive merge candidate as the motion information of the L0 split candidate and add it to the merge candidate list, and can increase numMergeCand by 1 (S2502).

[0319] The encoder / decoder can determine whether the number of merge candidates (numMergeCand) in the current merge candidate list is the same as the maximum number of merge candidates (MaxNumMergeCand). If they are the same (S2503 - Yes), the splitting process can be terminated. On the contrary, if they are not the same (S2503 - No), the encoder / decoder can set the L1 motion information of the combined bi-predictive merge candidate as the motion information of the L1 split candidate and add it to the merge candidate list, and can increase numMergeCand by 1 (S2504). The split index (splitIdx) can be increased by 1 (S2505).

[0320] Also, when the number of merge candidates (numMergeCand) in the current merge candidate list is the same as the maximum number of merge candidate cores (MaxNumMergeCand), the combined merge candidate splitting process ends; otherwise, it can return to step S2501 (S2506).

[0321] As shown in FIG. 25, the combined bi-prediction merge candidate splitting method for splitting bi-prediction merge candidates can only be executed when it is a B slice / B picture, or when the reference picture list uses M or more slices / pictures. Here, M can be 3 or 4, and can represent a positive integer of 3 or more.

[0322] Combined bi-prediction merge candidate splitting can be performed using at least one of the following methods: 1) When there are combined bi-prediction merge candidates, split them into single-prediction merge candidates; 2) When there are combined bi-prediction merge candidates and the L0 reference picture and the L1 reference picture in the combined bi-prediction merge candidates are different from each other, split them into single-prediction merge candidates; 3) When there are combined bi-prediction merge candidates and the L0 reference picture and the L1 reference picture in the combined bi-prediction merge candidates are the same as each other, split them into single-prediction merge candidates.

[0323] Since combined bi-prediction merge candidates use bi-prediction, motion compensation is performed using the reconstructed pixel data in a maximum of two different reference pictures. Therefore, the memory access bandwidth during motion compensation is larger than that of single-prediction using the reconstructed pixel data in one reference picture. Thus, by using combined bi-prediction merge candidate splitting, since the combined bi-prediction merge candidates are split into single-prediction merge candidates, if the split single-prediction merge candidates are determined as the motion information of the current block, the memory access bandwidth during motion compensation can be reduced.

[0324] The symbolizer / decoder can derive a zero merge candidate having a zero motion vector whose motion vector is (0,0).

[0325] The zero merge candidate can mean a merge candidate in which the motion vector in at least one of the L0 motion information, L1 motion information, L2 motion information, and L3 motion information has a motion vector of (0,0).

[0326] Also, the zero merge candidate can be at least one of two types. The first zero merge candidate can mean a merge candidate whose motion vector is (0,0) and whose reference picture index can have a value of 0 or more. And the second zero merge candidate can mean a merge candidate whose motion vector is (0,0) and whose reference picture index can have only a value of 0.

[0327] When the number of merge candidates (numMergeCand) in the current merge candidate list is not the same as the maximum number of merge candidates (MaxNumMergeCand) (that is, when the merge candidate list is not full of merge candidates), at least one of the first zero merge candidate and the second zero merge candidate can be repeatedly added to the merge candidate list until the number of merge candidates (numMergeCand) becomes the same as the maximum number of merge candidates (MaxNumMergeCand).

[0328] Also, the first zero merge candidate can be derived and added to the merge candidate list, and the second zero merge candidate can be derived and added to the merge candidate list when the merge candidate list is not full of merge candidates.

[0329] FIG. 26 is a diagram for explaining an embodiment of a zero merge candidate derivation method. When the number of merge candidates (numMergeCand) in the current merge candidate list is smaller than the maximum number of merge candidates (MaxNumMergeCand), the derivation of the zero merge candidate can be executed in the order as shown in FIG. 26.

[0330] First, the encoder / decoder can set the number of input merge candidates (numInputMergeCand) as the number of merge candidates (numMergeCand) in the current merge candidate list. Also, the reference picture index (zeroIdx) of the zero merge candidate can be set to 0. At this time, the m (numMergeCand numInputMergeCand) -th zero merge candidate can be derived.

[0331] The encoder / decoder can determine whether the slice type is a P - slice (S2601).

[0332] If the slice type (slice_type) is a P - slice (S2601 - yes), the encoder / decoder can set the number of reference pictures (numRefIdx) to the number of available reference pictures (num_ref_idx_l0_active_minus1 + 1) in the L0 list.

[0333] Also, the encoder / decoder can derive zero merge candidates as follows and increase numMergeCand by 1 (S2602). The L0 reference picture index of the m - th zero merge candidate (refIdxL0zeroCandm) = the reference picture index of the zero merge candidate (zeroIdx) The L1 reference picture index of the m - th zero merge candidate (refIdxL1zeroCandm) = 1 The L0 prediction list utilization flag of the m - th zero merge candidate (predFlagL0zeroCandm) = 1 The L1 prediction list utilization flag of the m - th zero merge candidate (predFlagL1zeroCandm) = 0 The x - component of the L0 motion vector of the m - th zero merge candidate (mvL0zeroCandm[0]) = 0 The y - component of the L0 motion vector of the m - th zero merge candidate (mvL0zeroCandm[1]) = 0 The x-component of the L1 motion vector of the m-th zero merge candidate (mvL1zeroCandm[0]) = 0 The y-component of the L1 motion vector of the m-th zero merge candidate (mvL1zeroCandm[1]) = 0

[0334] On the other hand, when the slice type is not a P slice (when it is a B slice or other slices) (S2601 - No), the number of reference images (numRefIdx) can be set to a value less than at least one of the number of available reference images in the L0 list (num_ref_idx_l0_active_minus1 + 1), the number of available reference images in the L1 list (num_ref_idx_l1_active_minus1 + 1), the number of available reference images in the L2 list (num_ref_idx_l2_active_minus1 + 1), and the number of available reference images in the L3 list (num_ref_idx_l3_active_minus1 + 1).

[0335] Also, the encoder / decoder can derive zero merge candidates as follows and increment numMergeCand by 1 (S2603). refIdxL0zeroCandm = zeroIdx refIdxL1zeroCandm = zeroIdx refIdxL2zeroCandm = zeroIdx refIdxL3zeroCandm = zeroIdx predFlagL0zeroCandm = 1 predFlagL1zeroCandm = 1 predFlagL2zeroCandm = 1 predFlagL3zeroCandm = 1 mvL0zeroCandm[0] = 0 mvL0zeroCandm[1] = 0 mvL1zeroCandm[0] = 0 mvL1zeroCandm[1] = 0 mvL2zeroCandm[0] = 0 mvL2 zero candidate m[1] = 0 mvL3 zero candidate m[0] = 0 mvL3 zero candidate m[1] = 0

[0336] After the S2602 step or the S2603 step is performed, if the reference picture count (refCnt) is the same as the number of reference pictures (numRefIdx) - 1, the encoder / decoder can set the reference picture index (zeroIdx) of the zero merge candidate to 0, and if not, it can increment refCnt and zeroIdx by 1 respectively (S2604).

[0337] Also, if numMergeCand is the same as MaxNumMergeCand, the encoder / decoder can end the zero merge candidate derivation process, and if not, it can return to the S2601 step (S2605).

[0338] When the zero merge candidate derivation method in FIG. 26 is performed, the zero merge candidate derived into the merge candidate list can be added as shown in FIG. 27.

[0339] FIG. 28 is a diagram for explaining another embodiment of the zero merge candidate derivation method. When the number of merge candidates (numMergeCand) in the current merge candidate list is smaller than the maximum number of merge candidates (MaxNumMergeCand), the derivation of the L0 single prediction zero merge candidate can be performed in the order as shown in FIG. 28.

[0340] First, the encoder / decoder can set the number of input merge candidates (numInputMergeCand) as the number of merge candidates (numMergeCand) in the current merge candidate list. Also, the reference picture index of zero merge candidates (zeroIdx) can be set to 0. At this time, the m (numMergeCand numInputMergeCand) -th zero merge candidate can be derived. Then, the number of reference pictures (numRefIdx) can be set to the number of available reference pictures in the L0 list (num_ref_idx_l0_active_minus1 + 1).

[0341] The encoder / decoder can derive zero merge candidates as follows and increment numMergeCand by 1 (S2801). The L0 reference picture index of the m -th zero merge candidate (refIdxL0zeroCandm) = the reference picture index of zero merge candidates (zeroIdx) The L1 reference picture index of the m -th zero merge candidate (refIdxL1zeroCandm) = 1 The L0 prediction list utilization flag of the m -th zero merge candidate (predFlagL0zeroCandm) = 1 The L1 prediction list utilization flag of the m -th zero merge candidate (predFlagL1zeroCandm) = 0 The x - component of the L0 motion vector of the m -th zero merge candidate (mvL0zeroCandm[0]) = 0 The y - component of the L0 motion vector of the m -th zero merge candidate (mvL0zeroCandm[1]) = 0 The x - component of the L1 motion vector of the m -th zero merge candidate (mvL1zeroCandm[0]) = 0 The y - component of the L1 motion vector of the m -th zero merge candidate (mvL1zeroCandm[1]) = 0

[0342] When the reference picture count (refCnt) is the same as numRefIdx - 1, the symbolizer / decoder can set zeroIdx to 0; otherwise, it can increment refCnt and zeroIdx by 1 respectively (S2802).

[0343] Also, when numMergeCand is the same as MaxNumMergeCand (S2803 - Yes), the encoder / decoder can end the zero merge candidate derivation step; otherwise, it can return to step S2801 (S2803 - No).

[0344] Since the zero merge candidate derivation method in FIG. 26 performs the derivation of dual - prediction zero merge candidates or the derivation of L0 single - prediction zero merge candidates according to the slice type, two implementation methods are required according to the slice type.

[0345] The zero merge candidate derivation method in FIG. 28 does not perform the derivation of dual - prediction zero merge candidates or the derivation of L0 single - prediction zero merge candidates according to the slice type. Instead, it derives L0 single - prediction zero merge candidates regardless of the slice type, which simplifies the hardware logic and can also reduce the cycle time required during the execution of the zero merge candidate derivation step. Also, when the L0 single - prediction zero merge candidate is determined as the motion information of the current block instead of the dual - prediction zero merge candidate, single - prediction motion compensation is performed instead of dual - prediction motion compensation, so the memory access bandwidth during motion compensation can be reduced.

[0346] As an example, when it is not a P - slice, an L0 single - prediction zero merge candidate can be derived and added to the merge candidate list.

[0347] After the symbolizer / decoder adds other merge candidates except zero merge candidates to the merge candidate list, it can add L0 single prediction zero merge candidates to the merge candidate list. Also, after the symbolizer / decoder initializes the merge candidate list with L0 single prediction zero merge candidates, it can also add spatial merge candidates, temporal merge candidates, combined merge candidates, zero merge candidates, additional merge candidates, etc. to the initialized merge candidate list.

[0348] FIG. 29 is a diagram for explaining an embodiment of deriving and sharing a merge candidate list in a CTU. The merge candidate list can be shared for blocks with a size or depth smaller than a predetermined block size or a predetermined block depth. Here, the predetermined block size or the predetermined block depth can be the size or block depth of a block for which information related to motion compensation is entropy encoded / decoded. Also, the predetermined block size or the predetermined block depth may be information that is entropy encoded by an encoder and entropy decoded by a decoder, or may be a value preset in common for the encoder / decoder.

[0349] Referring to FIG. 29, when the predetermined block size is 128x128, blocks having a size smaller than 128x128 (the hatched blocks in FIG. 29) can share the merge candidate list.

[0350] Next, the step of determining the motion information of the current block using the generated merge candidate list will be specifically described (S1202, S1303).

[0351] The encoder can determine, through motion estimation, among the merge candidates in the merge candidate list, the merge candidates to be used for motion compensation, and encode the merge candidate index (merge_idx) indicating the determined merge candidates into the bitstream.

[0352] On the one hand, the symbolizer can select a merge candidate from the merge candidate list based on the above-described merge candidate index to generate a prediction block and determine the motion information of the current block. Here, based on the determined motion information, motion compensation can be performed to generate a prediction block for the current block.

[0353] As an example, when the merge candidate index is selected as 3, the merge candidate indicated by the merge candidate index 3 in the merge candidate list is determined as the motion information and can be used for motion compensation of the block to be coded.

[0354] The decoder can decode the merge candidate index in the bitstream to determine the merge candidate in the merge candidate list indicated by the merge candidate index. The determined merge candidate can be determined as the motion information of the current block. The determined motion information is used for motion compensation of the current block. At this time, the motion compensation can be the same as the meaning of inter prediction.

[0355] As an example, when the merge candidate index is 2, the merge candidate indicated by the merge candidate index 2 in the merge candidate list is determined as the motion information and can be used for motion compensation of the block to be decoded.

[0356] In addition, at least one value of the information corresponding to the motion information of the current block can be changed, and the motion information can be used for inter prediction or motion compensation of the current block. Here, among the information corresponding to the motion information, the value to be changed can be at least one of the x component of the motion vector, the y component of the motion vector, and the reference picture index. Also, when changing at least one value of the information corresponding to the motion information, at least one value of the information corresponding to the motion information can be changed so as to show the minimum distortion using a distortion calculation method (such as SAD, SSE, MSE, etc.).

[0357] Next, the step of performing motion compensation on the current block using the determined motion information will be specifically described (S1203, S1304).

[0358] In the encoder and decoder, inter prediction or motion compensation can be performed using the determined motion information of the merge candidates. Here, the current block (the block to be encoded / decoded) can have the motion information of the determined merge candidates.

[0359] The current block can have a minimum of 1 to a maximum of N pieces of motion information according to the prediction direction. Using the motion information, a minimum of 1 to a maximum of N prediction blocks can be generated to derive the final prediction block of the current block.

[0360] As an example, when the current block has one piece of motion information, the prediction block generated using the motion information can be determined as the final prediction block of the current block.

[0361] On the other hand, when the current block has a plurality of pieces of motion information, a plurality of prediction blocks can be generated using the plurality of pieces of motion information, and the final prediction block of the current block can be determined based on the weighted sum of the plurality of prediction blocks. The reference images each including a plurality of prediction blocks indicated by the plurality of pieces of motion information may be included in different reference image lists from each other, or may be included in the same reference image list. Also, when the current block has a plurality of pieces of motion information, among the plurality of pieces of motion information, the plurality of reference images may indicate the same reference image as each other.

[0362] As an example, a plurality of prediction blocks can be generated based on at least one of a spatial merge candidate, a temporal merge candidate, a modified spatial merge candidate, a modified temporal merge candidate, a merge candidate having a predetermined motion information value or a combined merge candidate, and an additional merge candidate, and the final prediction block of the current block can be determined based on the weighted sum of the plurality of prediction blocks.

[0363] As another example, based on merge candidates indicated by a preset merge candidate index, a plurality of prediction blocks can be generated, and based on the weighted sum of the plurality of prediction blocks, the final prediction block of the current block can be determined. Also, based on merge candidates existing within a preset merge candidate index range, a plurality of prediction blocks can be generated, and based on the weighted sum of the plurality of prediction blocks, the final prediction block of the current block can be determined.

[0364] The weight applied to each prediction block can be 1 / N (where N is the number of generated prediction blocks) and can have an equal value. As an example, when two prediction blocks are generated, the weight applied to each prediction block is 1 / 2, when three prediction blocks are generated, the weight applied to each prediction block is 1 / 3, and when four prediction blocks are generated, the weight applied to each prediction block can be 1 / 4. Alternatively, different weights can be assigned to each prediction block to determine the final prediction block of the current block.

[0365] The weights do not have to have fixed values for each prediction block, but can also have variable values for each prediction block. At this time, the weights applied to each prediction block may be the same as each other or different from each other. As an example, when two prediction blocks are generated, the weights applied to the two prediction blocks can be variable values for each block such as (1 / 2, 1 / 2), (1 / 3, 2 / 3), (1 / 4, 3 / 4), (2 / 5, 3 / 5), (3 / 8, 5 / 8), etc. On the other hand, the weights can be positive real values and negative real values. As an example, it can include negative real values such as (-1 / 2, 3 / 2), (-1 / 3, 4 / 3), (-1 / 4, 5 / 4), etc.

[0366] On the one hand, in order to apply variable weights, one or more weight information for the current block may be signaled via a bitstream. The weight information may be signaled separately for each prediction block or separately for each reference picture. It is also possible for multiple prediction blocks to share one weight information.

[0367] The encoder and decoder can determine whether to use the motion information of the merge candidate based on the prediction block list utilization flag. As an example, when the prediction block list utilization flag indicates 1, which is the first value, for each reference picture list, the encoder and decoder can use the motion information of the merge candidate of the current block to perform inter prediction or motion compensation, and when it indicates 0, which is the second value, the encoder and decoder can indicate that they do not perform inter prediction or motion compensation using the motion information of the merge candidate of the current block. On the other hand, the first value of the prediction block list utilization flag can also be set to 0 and the second value to 1 respectively.

[0368] The following Equations 3 to 5 respectively show examples of generating the final prediction block of the current block when the inter prediction indicator of the current block is PRED_BI (or when the current block can use two motion information), PRED_TRI (or when the current block can use three motion information), and PRED_QUAD (or when the current block can use four motion information), and the prediction direction for each reference picture list is unidirectional.

Equation

Equation

Equation

[0369] In the above formulas (3) to (5), P_BI, P_TRI, and P_QUAD indicate the final predicted block of the current block, and LX (X = 0, 1, 2, 3) can mean a reference picture list. WF_LX indicates the weight value of the predicted block generated using LX, and OFFSET_LX can indicate the offset value for the predicted block generated using LX. P_LX means the predicted block generated using the motion information for LX of the current block. RF means a rounding factor and can be set to 0, a positive number, or a negative number. The LX reference picture list can include at least one of a long-term reference picture, a reference picture without a deblocking filter, a reference picture without a sample adaptive offset, a reference picture without an adaptive loop filter, a reference picture with only a deblocking filter and a sample adaptive offset, a reference picture with only a deblocking filter and an adaptive loop filter, a reference picture with only a sample adaptive offset and an adaptive loop filter, and a reference picture with all of a deblocking filter, a sample adaptive offset, and an adaptive loop filter. In this case, the LX reference picture list can be at least one of an L0 reference picture list, an L1 reference picture list, an L2 reference picture list, and an L3 reference picture list.

[0370] Even when there are a plurality of prediction directions for a predetermined reference picture list, the final predicted block for the current block can be obtained based on the weighted sum of the predicted blocks. At this time, the weights applied to the predicted blocks derived from the same reference picture list may have the same value or may have different values from each other.

[0371] At least one of the weights (WF_LX) and offsets (OFFSET_LX) for a plurality of predicted blocks can be an encoded parameter that is entropy-coded / decoded.

[0372] As another example, the weights and offsets may be derived from the encoded / decoded neighboring blocks in the vicinity of the current block. Here, the neighboring blocks of the current block may include at least one of the blocks used to derive the spatial merge candidates of the current block or the blocks used to derive the temporal merge candidates of the current block.

[0373] As another example, the weights and offsets may be determined based on the presentation order (POC) of the current picture and each reference picture. In this case, the farther the distance between the current picture and the reference picture, the smaller the value of the weight or offset can be set, and the closer the distance between the current picture and the reference picture, the larger the value of the weight or offset can be set. As an example, when the difference in POC between the current picture and the L0 reference picture is 2, the weight value applied to the predicted block generated with reference to the L0 reference picture is set to 1 / 3, whereas when the difference in POC between the current picture and the L0 reference picture is 1, the weight value applied to the predicted block generated with reference to the L0 reference picture can be set to 2 / 3. As illustrated above, the weight or offset value can have an inverse proportional relationship with the difference in presentation order between the current picture and the reference picture. As another example, it is also possible to make the weight or offset value have a proportional relationship with the difference in presentation order between the current picture and the reference picture.

[0374] As another example, at least one of the weights or offsets can be entropy encoded / decoded based on at least one of the encoding parameters of the current block, neighboring blocks, or corresponding position blocks. Also, the weighted sum of the predicted blocks can be calculated based on at least one of the encoding parameters of the current block, neighboring blocks, or corresponding position blocks.

[0375] The weighted sum of a plurality of prediction blocks can be applied only to a partial region within the prediction block. Here, the partial region may be a region corresponding to the boundary within the prediction block. As described above, in order to apply the weighted sum only to a partial region, the weighted sum can be performed in units of sub-blocks of the prediction block.

[0376] Next, a process of entropy encoding / decoding information related to motion compensation will be described in detail (S1204, S1301).

[0377] FIGS. 30 and 31 are diagrams illustrating the syntax of information related to motion compensation. FIG. 30 shows an embodiment of the syntax of information related to motion compensation in a coding unit, and FIG. 31 shows an embodiment of the syntax of information related to motion compensation in a prediction unit.

[0378] The encoding device can entropy-encode information related to motion compensation via a bitstream, and the decoding device can entropy-decode information related to motion compensation included in the bitstream. Here, the information related to motion compensation to be entropy-encoded / decoded can include at least one of skip mode usage information (cu_skip_flag), merge mode usage information (merge_flag), merge index information (merge_index), inter prediction indicator (inter_pred_idc), weight values (wf_l0, wf_l1, wf_l2, wf_l3), and offset values (offset_l0, offset_l1, offset_l2, offset_l3). The information related to motion compensation can be entropy-encoded / decoded in units of at least one of CTU, coding block, and prediction block.

[0379] The skip mode usage presence / absence information (cu_skip_flag) can indicate the use of the skip mode when it has a first value of 1, and can indicate the non-use of the skip mode when it has a second value of 0. Based on the skip mode usage presence / absence information, motion compensation for the current block can be performed using the skip mode.

[0380] The merge mode usage presence / absence information (merge_flag) can indicate the use of the merge mode when it has a first value of 1, and can indicate the non-use of the merge mode when it has a second value of 0. Based on the merge mode usage presence / absence information, motion compensation for the current block can be performed using the skip mode.

[0381] The merge index information (merge_index) can mean information indicating a merge candidate within a merge candidate list.

[0382] Also, the merge index information can mean information regarding a merge index.

[0383] Also, the merge index information can indicate a block from which a merge candidate among the blocks reconstructed to be spatially / temporally adjacent to the current block was derived.

[0384] In addition, the merge index information can indicate at least one of the motion information that the merge candidates have. For example, when the merge index information has a first value of 0, it can indicate the first merge candidate in the merge candidate list; when it has a second value of 1, it can indicate the second merge candidate in the merge candidate list; when it has a third value of 2, it can indicate the third merge candidate in the merge candidate list. Similarly, when it has a fourth to Nth value, it can indicate the merge candidate corresponding to the value based on the order in the merge candidate list. Here, N can represent a positive integer including 0.

[0385] Based on the merge mode index information, motion compensation of the current block can be performed using the merge mode.

[0386] The inter prediction indicator can mean at least one of the inter prediction direction of the current block or the number of prediction directions when the current block is encoded / decoded by inter prediction. As an example, the inter prediction indicator can indicate one-direction prediction, or can indicate multi-direction prediction such as bi-directional prediction, tri-directional prediction, or quadri-directional prediction. The inter prediction indicator can mean the number of reference images used when the current block generates a prediction block. Or, one reference image may be used for multiple direction predictions. In this case, N (N>M) direction predictions can be performed using M reference images. The inter prediction indicator can also mean the number of prediction blocks used when performing inter prediction or motion compensation for the current block.

[0387] In this way, according to the inter-prediction indicator, the number of reference images used when generating the predicted block of the current block, the number of predicted blocks used when performing inter-prediction or motion compensation of the current block, or the number of available reference image lists of the current block can be determined. Here, the number N of reference image lists is a positive integer and can have a value of 1, 2, 3, 4, or more. For example, the reference image lists can include L0, L1, L2, and L3, etc. The current block can perform motion compensation using one or more reference image lists.

[0388] As an example, the current block can generate at least one predicted block using at least one reference image list to perform motion compensation of the current block. As an example, one or more predicted blocks can be generated using the reference image list L0 to perform motion compensation, or one or more predicted blocks can be generated using the reference image lists L0 and L1 to perform motion compensation. Or, one, one or more predicted blocks or a maximum of N predicted blocks (where N is a positive integer of 3 or more) can be generated using the reference image lists L0, L1, and L2 to perform motion compensation, or, using the reference image lists L0, L1, L2, and L3, one, one or more predicted blocks or a maximum of N predicted blocks (where N is a positive integer of 4 or more) can be used to perform motion compensation for the current block.

[0389] The reference image indicator can indicate one-direction (PRED_LX), bi-direction (PRED_BI), tri-direction (PRED_TRI), quad-direction (PRED_QUAD), or more directions according to the number of prediction directions of the current block.

[0390] As an example, assuming that one-direction prediction is performed for each reference picture list, the inter prediction indicator PRED_LX can mean generating one prediction block using the reference picture list LX (X is an integer such as 0, 1, 2, or 3), and performing inter prediction or motion compensation using the generated one prediction block. The inter prediction indicator PRED_BI can mean generating two prediction blocks using the L0, L1, L2, and L3 reference picture lists, and performing inter prediction or motion compensation using the generated two prediction blocks. The inter prediction indicator PRED_TRI can mean generating three prediction blocks using at least one of the L0, L1, L2, and L3 reference picture lists, and performing inter prediction or motion compensation using the generated three prediction blocks. The inter prediction indicator PRED_QUAD can mean generating four prediction blocks using at least one of the L0, L1, L2, and L3 reference picture lists, and performing inter prediction or motion compensation using the generated four prediction blocks. That is, the total number of prediction blocks used for executing the inter prediction of the current block can be set in the inter prediction indicator.

[0391] When multi-direction prediction is performed for the reference picture list, the inter prediction indicator PRED_BI means that bidirectional prediction is performed for the L0 reference picture list, and the inter prediction indicator PRED_TRI means that three-direction prediction is performed for the L0 reference picture list, one-direction prediction is performed for the L0 reference picture list, bidirectional prediction is performed for the L1 reference picture list, or bidirectional prediction is performed for the L0 reference picture list, and one-direction prediction is performed for the L1 reference picture list.

[0392] Thus, the inter prediction indicator means generating at least one to a maximum of N prediction blocks (where N is the number of prediction directions indicated by the inter prediction indicator) from at least one reference picture list for motion compensation, or alternatively, it can mean generating at least one to a maximum of N prediction blocks from N reference pictures and performing motion compensation for the current block using the generated prediction blocks.

[0393] For example, the inter prediction indicator PRED_TRI means generating three prediction blocks using at least one of the L0, L1, L2, L3 reference picture lists for inter prediction or motion compensation of the current block, or alternatively, it can mean generating three prediction blocks using at least three of the L0, L1, L2, L3 reference picture lists for inter prediction or motion compensation of the current block. Also, PRED_QUAD means generating four prediction blocks using at least one of the L0, L1, L2, L3 reference picture lists for inter prediction or motion compensation of the current block, or alternatively, it can mean an inter prediction indicator that generates four prediction blocks using at least four of the L0, L1, L2, L3 reference picture lists for inter prediction or motion compensation of the current block.

[0394] The available inter prediction directions can be determined according to the inter prediction indicator, and all or part of the available inter prediction directions may be selectively used based on the size and / or shape of the current block.

[0395] The prediction list utilization flag indicates whether to generate a prediction block using the corresponding reference picture list.

[0396] For example, when the prediction list utilization flag indicates the first value, which is 1, it indicates that a prediction block can be generated using the reference image list. When it indicates the second value, which is 0, it can indicate that a prediction block is not generated using the reference image list. Here, the first value of the prediction list utilization flag may be set to 0, and the second value may be set to 1.

[0397] That is, when the prediction list utilization flag indicates the first value, a prediction block of the current block can be generated using the motion information corresponding to the reference image list.

[0398] On the other hand, the prediction list utilization flag can be set based on the inter prediction indicator. For example, when the inter prediction indicator indicates PRED_LX, PRED_BI, PRED_TRI, or PRED_QUAD, the prediction list utilization flag predFlagLX can be set to the first value, which is 1. If the inter prediction indicator is PRED_LN (where N is a positive integer other than X), the prediction list utilization flag predFlagLX can be set to the second value, which is 0.

[0399] Also, the inter prediction indicator can be set based on the prediction list utilization flag. For example, when the prediction list utilization flags predFlagL0 and predFlagL1 indicate the first value, which is 1, the inter prediction indicator can be set to PRED_BI. For example, when only the prediction list utilization flag predFlagL0 indicates the first value, which is 1, the inter prediction indicator can be set to PRED_L0.

[0400] When two or more prediction blocks are generated during motion compensation for a current block, a final prediction block for the current block can be generated through a weighted sum for each prediction block. During the weighted sum operation, at least one of a weight and an offset can be applied to each prediction block. The weighted sum factor used in the weighted sum operation, such as a weight (weighted factor) or an offset, can be entropy-coded / decoded by at least the number of at least one of a reference image list, a reference image, a motion vector candidate index, a motion vector difference, a motion vector, skip mode usage information, merge mode usage information, and merge index information. Also, the weighted sum factor for each prediction block can be entropy-coded / decoded based on an inter prediction indicator. Here, the weighted sum factor can include at least one of a weight and an offset.

[0401] The weighted sum factor can also be derived by index information specifying any one of a set predefined in an encoding device and a decoding device. In this case, the index information for specifying at least one of a weight and an offset can be entropy-coded / decoded. The set predefined in the encoder and the decoder can be defined for the weight and the offset, respectively. The predefined set can include one or more weight candidates or offset candidates. Also, a table defining a mapping relationship between the weight and the offset may be used. In this case, a weight value and an offset value for a prediction block can be obtained from the table using one index information. For each index information for the weight to be entropy-coded / decoded, the index information for the offset mapped thereto can also be entropy-coded / decoded.

[0402] Information related to the weighting and factors may be entropy-coded / decoded in block units or at a higher level. As an example, the weight or offset may be entropy-coded / decoded in block units such as CTU, CU, or PU, or at a higher level such as in a Video Parameter Set, Sequence Parameter Set, Picture Parameter Set, Adaptation Parameter Set, or Slice Header.

[0403] The weighting and factors can also be entropy-coded / decoded based on a weighting and factor difference value that indicates the difference value between the weighting and factor and the predicted value of the weighting and factor. As an example, the predicted weight value and the weight difference value can be entropy-coded / decoded, or the predicted offset value and the offset difference value can be entropy-coded / decoded. Here, the weight difference value indicates the difference value between the weight and the predicted weight value, and the offset difference value can indicate the difference value between the offset and the predicted offset value.

[0404] At this time, the weighting and factor difference value may be entropy-coded / decoded in block units, and the predicted value of the weighting and factor may be entropy-coded / decoded at a higher level. When the predicted value of the weighting and factor such as the predicted weight value or the predicted offset value is entropy-coded / decoded in picture or slice units, the blocks included in the picture or slice can use a common predicted value of the weighting and factor.

[0405] The weighted sum factor prediction value can also be derived through a specific area within an image, slice, or tile, or a specific area within a CTU or CU. As an example, the weight value or offset value of a specific area within an image, slice, tile, CTU, or CU can be used as the weight prediction value or offset prediction value. In this case, the entropy coding / decoding of the weighted sum factor prediction value can be omitted, and only the weighted sum factor difference value can be entropy coded / decoded.

[0406] Alternatively, the weighted sum factor prediction value can also be derived from neighboring blocks that have been encoded / decoded in the vicinity of the current block. As an example, the weight value or offset value of a neighboring block that has been encoded / decoded in the vicinity of the current block can be set as the weight prediction value or offset prediction value of the current block. Here, the neighboring blocks of the current block can include at least one of the blocks used for deriving spatial merge candidates and the blocks used for deriving temporal merge candidates.

[0407] When using the weight prediction value and the weight difference value, the decoding device can calculate the weight value for the predicted block by combining the weight prediction value and the weight difference value. Also, when using the offset prediction value and the offset difference value, the decoding device can calculate the offset value for the predicted block by combining the offset prediction value and the offset difference value.

[0408] The weighted sum factor or the difference value of the weighted sum factor can be entropy coded / decoded based on at least one of the coding parameters of the current block, neighboring blocks, or corresponding position blocks.

[0409] Based on at least one of the coding parameters of the current block, neighboring blocks, or corresponding position blocks, the weighted sum factor, the prediction value of the weighted sum factor, or the difference value of the weighted sum factor can be derived as the weighted sum factor, the prediction value of the weighted sum factor, or the difference value of the weighted sum factor of the current block.

[0410] Instead of entropy encoding / decoding information on the weighted sum factor of the current block, it is also possible to use the weighted sum factor of the blocks encoded / decoded in the vicinity of the current block as the weighted sum factor of the current block. As an example, the weight or offset of the current block can be set to the same value as the weight or offset of the neighboring blocks encoded / decoded in the vicinity of the current block.

[0411] The current block can perform motion compensation using at least one of the weighted sum factors, or can perform motion compensation using at least one of the derived weighted sum factors.

[0412] The weighted sum factor may be included in the information related to motion compensation.

[0413] At least one of the above-described information related to motion compensation can be entropy encoded / decoded in at least one unit of a CTU or a sub-unit (sub-CTU) of a CTU. Here, the sub-unit of a CTU can include at least one unit of a CU and a PU. The blocks of the sub-unit of a CTU can be in a square or non-square form. The information related to motion compensation described later can, for the sake of convenience, mean at least one of the information related to motion compensation.

[0414] When the information related to the motion compensation of a CTU is entropy encoded / decoded, depending on the value of the information related to the motion compensation, motion compensation can be performed using the information related to the motion compensation in all or some of the blocks existing in the CTU.

[0415] When the information related to motion compensation is entropy encoded / decoded in a CTU or a sub-unit of a CTU, the information related to motion compensation can be entropy encoded / decoded based on at least one of a predetermined block size or a predetermined block depth.

[0416] Here, information regarding a predetermined block size or a predetermined block depth can be further entropy-coded / decoded. Alternatively, information regarding a predetermined block size or a predetermined block depth may be determined based on at least one of preset values, encoding parameters, or at least one of values of other syntax elements in an encoder and a decoder.

[0417] Information regarding motion compensation is entropy-coded / decoded only for blocks having a block size equal to or larger than the predetermined block size. For blocks having a block size smaller than the predetermined block size, information regarding motion compensation may not be entropy-coded / decoded. In this case, sub-blocks within a block having a block size equal to or larger than the predetermined block size can perform motion compensation based on information regarding motion compensation entropy-coded / decoded in a block having a block size equal to or larger than the predetermined block size. That is, sub-blocks within a block having a block size equal to or larger than the predetermined block size can share information regarding motion compensation including motion vector candidates, motion vector candidate lists, merge candidates, merge candidate lists, and the like.

[0418] Information regarding motion compensation is entropy-coded / decoded only for blocks having a block depth equal to or shallower than the predetermined block depth. For blocks having a block depth deeper than the predetermined block depth, information regarding motion compensation may not be entropy-coded / decoded. In this case, sub-blocks within a block having a block depth equal to or shallower than the predetermined block depth can perform motion compensation based on information regarding motion compensation entropy-coded / decoded in a block having a block depth equal to or shallower than the predetermined block depth. That is, sub-blocks within a block having a block depth equal to or shallower than the predetermined block depth can share information regarding motion compensation including motion vector candidates, motion vector candidate lists, merge candidates, merge candidate lists, and the like.

[0419] As an example, when information related to motion compensation is entropy encoded / decoded in a CTU sub-unit of size 32x32 where the block size of the CTU is 64x64, for blocks that belong to the 32x32 block and are smaller in size than the 32x32 block unit, motion compensation can be performed based on the information related to motion compensation that is entropy encoded / decoded in the 32x32 block unit.

[0420] As another example, when information related to motion compensation is entropy encoded / decoded in a CTU sub-unit of size 16x16 where the block size of the CTU is 128x128, for blocks that belong to the 16x16 block and are of the same size as or smaller than the 16x16 block unit, motion compensation can be performed based on the information related to motion compensation that is entropy encoded / decoded in the 16x16 block unit.

[0421] As another example, when information related to motion compensation is entropy encoded / decoded in CTU sub-units where the block depth of the CTU is 0 and the block depth is 1, for blocks that belong to block depth 1 and have a block depth deeper than block depth 1, motion compensation can be performed based on the information related to motion compensation that is entropy encoded / decoded at block depth 1.

[0422] For example, when at least one of the information related to motion compensation is entropy encoded / decoded in CTU sub-units where the block depth of the CTU is 0 and the block depth is 2, for blocks that belong to block depth 2 and have a block depth equal to or deeper than block depth 2, motion compensation can be performed based on the information related to motion compensation that is entropy encoded / decoded at block depth 2.

[0423] Here, the value of the block depth can have a positive integer including 0. It can be meant that the greater the value of the block depth, the deeper the depth, and the smaller the value of the block depth, the shallower the depth. Therefore, the greater the value of the block depth, the smaller the block size can be, and the smaller the value of the block depth, the larger the block size can be. Also, the lower level of a predetermined block depth can mean a depth deeper than the predetermined block depth, and the lower level of a predetermined block depth can mean a deeper depth within the block corresponding to the predetermined block depth.

[0424] The information related to motion compensation may be entropy encoded / decoded in block units or at a higher level. As an example, the information related to motion compensation can be entropy encoded / decoded in block units such as CTU, CU, or PU, or at a higher level such as a Video Parameter Set, Sequence Parameter Set, Picture Parameter Set, Adaptation Parameter Set, or Slice Header of the video.

[0425] The information related to motion compensation can also be entropy - encoded / decoded based on the information - related difference value for motion compensation, which indicates the difference between the information related to motion compensation and the predicted value of the information related to motion compensation. Taking the inter - prediction indicator, which is one of the information related to motion compensation, as an example, the predicted value of the inter - prediction indicator and the inter - prediction indicator difference value can be entropy - encoded / decoded. At this time, the inter - prediction indicator difference value is entropy - encoded / decoded in block units, and the predicted value of the inter - prediction indicator can be entropy - encoded / decoded at a higher level. When the predicted value of the information related to motion compensation, such as the predicted value of the inter - prediction indicator, is entropy - encoded / decoded in picture or slice units, the blocks included in the picture or slice can use a common predicted value of the information related to motion compensation.

[0426] The predicted value of the information related to motion compensation may also be derived through a specific area within an image, slice or tile, or a specific area within a CTU or CU. As an example, the inter - prediction indicator of a specific area within an image, slice, tile, CTU or CU can be used as the predicted value of the inter - prediction indicator. In this case, the entropy - encoding / decoding of the predicted value of the information related to motion compensation can be omitted, and only the information - related difference value for motion compensation can be entropy - encoded / decoded.

[0427] Alternatively, the predicted value of the information related to motion compensation may be derived from neighboring blocks encoded / decoded in the vicinity of the current block. As an example, the inter-prediction indicator of a neighboring block encoded / decoded in the vicinity of the current block can be set as the predicted value of the inter-prediction indicator of the current block. Here, the neighboring blocks of the current block can include at least one of the blocks used to derive spatial merge candidates and the blocks used to derive temporal merge candidates. Also, the neighboring blocks may have a depth equal to or smaller than the depth of the current block. When there are multiple neighboring blocks, any one of them can be selectively used based on a predetermined priority order. The neighboring blocks used to predict the information related to motion compensation may have a fixed position based on the current block, or may have a variable position according to the position of the current block. Here, the position of the current block may be a position based on the picture or slice to which the current block belongs, or may be a position based on the position of the CTU, CU, or PU to which the current block belongs.

[0428] The merge index information can be calculated using index information within a set predetermined in the encoder and decoder.

[0429] When using the predicted value of the information related to motion compensation and the difference value of the information related to motion compensation, the decoding device can calculate the value of the information related to motion compensation for the predicted block by combining the predicted value of the information related to motion compensation and the difference value of the information related to motion compensation.

[0430] The information related to motion compensation or the difference value of the information related to motion compensation can be entropy-encoded / decoded based on at least one of the encoding parameters of the current block, neighboring blocks, or corresponding position blocks.

[0431] Based on at least one of the encoding parameters of the current block, neighboring block, or corresponding position block, information related to motion compensation, a predicted value of the information related to motion compensation, or a difference value of the information related to motion compensation can be derived as the information related to motion compensation of the current block, the predicted value of the information related to motion compensation, or the difference value of the information related to motion compensation.

[0432] Instead of entropy encoding / decoding the information related to motion compensation of the current block, it is also possible to use the information related to motion compensation of the block encoded / decoded in the neighborhood of the current block as the information related to motion compensation of the current block. As an example, the inter-prediction indicator of the current block can be set to the same value as the inter-prediction indicator of the neighboring block encoded / decoded in the neighborhood of the current block.

[0433] Also, at least one of the information related to motion compensation can have a fixed value preset by the encoder and decoder. The preset fixed value can be determined as a value for at least one value of the information related to motion compensation, and in a block having a smaller block size within a specific block size, at least one of the information related to motion compensation having the preset fixed value can be shared. Similarly, in a block having a deeper block depth below a specific block size, at least one of the information related to motion compensation having the preset fixed value can be shared. Here, the fixed value can be a positive integer value including 0 or an integer vector value including (0, 0).

[0434] Here, sharing at least one of the information related to motion compensation means that in a block, at least one of the information related to motion compensation can have the same value for each other, or motion compensation can be performed using at least one of the information related to motion compensation having the same value for each other.

[0435] The information related to motion compensation may further include at least one of a motion vector, a motion vector candidate, a motion vector candidate index, a motion vector difference value, a motion vector prediction value, skip mode usage information (skip_flag), merge mode usage information (merge_flag), merge index information (merge_index), motion vector resolution information, overlapped block motion compensation information, local illumination compensation information, affine motion compensation information, decoder-side motion vector derivation information, and bi-directional optical flow information.

[0436] The motion vector resolution information may be information indicating whether to use a specific resolution for at least one of the motion vector and the motion vector difference value. Here, the resolution can mean precision. Also, the specific resolution can be set to at least one of an integer-pel unit, a 1 / 2-pel unit, a 1 / 4-pel unit, a 1 / 8-pel unit, a 1 / 16-pel unit, a 1 / 32-pel unit, and a 1 / 64-pel unit.

[0437] The overlapped block motion compensation information may be information indicating whether to further use the motion vectors of neighboring blocks that are spatially adjacent to the current block to calculate the weighted sum of the predicted blocks of the current block during motion compensation of the current block.

[0438] The regional illumination compensation information may be information indicating whether to apply at least one of a weight value and an offset value when generating a predicted block of the current block. Here, at least one of the weight value and the offset value may be a value calculated based on a reference block.

[0439] The affine motion compensation information may be information indicating whether to use an affine motion model during motion compensation for the current block. Here, the affine motion model may be a model that divides one block into a plurality of sub-blocks using a plurality of parameters and calculates the motion vectors of the divided sub-blocks using representative motion vectors.

[0440] The decoder motion vector derivation information may be information indicating whether to derive and use the motion vectors required for motion compensation by a decoder. Based on the decoder motion vector derivation information, information regarding the motion vectors may not be entropy-coded / decoded. And when the decoder motion vector derivation information indicates that the decoder derives and uses the motion vectors, the information regarding the merge mode can be entropy-coded / decoded. That is, the decoder motion vector derivation information can indicate whether to use the merge mode in the decoder.

[0441] The bidirectional optical flow information may be information indicating whether to correct the motion vectors in units of pixels or sub-blocks to perform motion compensation. Based on the bidirectional optical flow information, the motion vectors in units of pixels or sub-blocks may not be entropy-coded / decoded. Here, the motion vector correction may be to change the motion vector in units of blocks to the motion vector values in units of pixels or sub-blocks.

[0442] The current block can perform motion compensation using at least one of the information regarding motion compensation and can entropy-code / decrypt at least one of the information regarding motion compensation.

[0443] FIG. 32 is a diagram for explaining an embodiment in which a merge mode is used for blocks smaller than a predetermined block size.

[0444] Referring to FIG. 32, when the predetermined block size is 8x8, blocks smaller than 8x8 (hatched blocks) can use the merge mode.

[0445] On the other hand, when comparing the sizes between blocks, being smaller than the predetermined block size can mean that the sum of the samples existing in the block is small. As an example, a 32x16 block has 512 samples, so it is smaller in size than a 32x32 block having 1024 samples, and a 4x16 block has 64 samples, so it can be said to be the same size as an 8x8 block.

[0446] When entropy encoding / decoding information related to motion compensation, binarization methods such as Truncated Rice binarization method, K-th order Exp_Golomb binarization method, Restricted K-th order Exp_Golomb binarization method, Fixed-length binarization method, Unary binarization method, or Truncated Unary binarization method can be used.

[0447] When entropy encoding / decoding information related to motion compensation, a context model can be determined using at least one of information related to motion compensation of neighboring blocks in the vicinity of the current block or region information of neighboring blocks, information related to motion compensation encoded / decoded previously or region information encoded / decoded previously, information related to the depth of the current block, and information related to the size of the current block.

[0448] Also, when entropy-encoding / decoding information related to motion compensation, at least one of information related to motion compensation of neighboring blocks, information related to motion compensation that has been previously encoded / decoded, information related to the depth of the current block, and information related to the size of the current block can be used as a predicted value for information related to motion compensation of the current block to perform entropy-encoding / decoding.

[0449] The encoding / decoding process can be performed for each of the luminance and chrominance signals. For example, in the encoding / decoding process, at least one method among obtaining an inter-prediction indicator, generating a merge candidate list, deriving motion information, and executing motion compensation can be applied differently to the luminance signal and the chrominance signal.

[0450] The encoding / decoding process for the luminance and chrominance signals can be performed in the same manner. For example, at least one of an inter-prediction indicator, a merge candidate list, a merge candidate, a reference image, and a reference image list applied in the encoding / decoding process for the luminance signal can be similarly applied to the chrominance signal.

[0451] These methods can be performed in a similar manner by an encoder and a decoder. For example, in the encoding / decoding process, at least one method among obtaining an inter-prediction indicator, generating a merge candidate list, deriving motion information, and executing motion compensation can be similarly applied by the encoder and the decoder. Also, the order of application of these methods may be different between the encoder and the decoder.

[0452] The above-described embodiments of the present invention can be applied according to the size of at least one of an encoding block, a prediction block, a block, and a unit. The size here may be defined as a minimum size and / or a maximum size for these embodiments to be applied, or may be defined as a fixed size for which the above embodiments are applied. Also, in these embodiments, the first embodiment can be applied for the first size, and the second embodiment can be applied for the second size. That is, these embodiments can be applied in a composite manner according to the size. Further, the above-described embodiments of the present invention can be applied only when the size is equal to or greater than the minimum size and equal to or less than the maximum size. That is, these embodiments can be applied only when the size of the block is within a certain range.

[0453] For example, the above embodiments can be applied only when the size of the encoding / decoding target block is 8x8 or more. For example, the above embodiments can be applied only when the size of the encoding / decoding target block is 16x16 or more. For example, the above embodiments can be applied only when the size of the encoding / decoding target block is 32x32 or more. For example, the above embodiments can be applied only when the size of the encoding / decoding target block is 64x64 or more. For example, the above embodiments can be applied only when the size of the encoding / decoding target block is 128x128 or more. For example, the above embodiments can be applied only when the size of the encoding / decoding target block is 4x4. For example, the above embodiments can be applied only when the size of the encoding / decoding target block is 8x8 or less. For example, the above embodiments can be applied only when the size of the encoding / decoding target block is 16x16 or less. For example, the above embodiments can be applied only when the size of the encoding / decoding target block is 8x8 or more and 16x16 or less. For example, the above embodiments can be applied only when the size of the encoding / decoding target block is 16x16 or more and 64x64 or less.

[0454] The embodiments of the present invention described above can be applied according to a temporal layer. In order to identify the temporal layer to which these embodiments are applicable, a separate identifier is signaled, and these embodiments can be applied to the temporal layer specified by the corresponding identifier. The identifier here may be defined as the minimum layer and / or the maximum layer applicable to the above embodiments, or may be defined as indicating a specific layer to which the above embodiments are applied.

[0455] For example, the above embodiments can be applied only when the temporal layer of the current image is the lowest layer. For example, the above embodiments can be applied only when the temporal layer identifier of the current image is 0. For example, the above embodiments can be applied only when the temporal layer identifier of the current image is 1 or more. For example, the above embodiments can be applied only when the temporal layer of the current image is the highest layer.

[0456] As in the embodiments of the present invention described above, the reference picture set used in the reference picture list construction and reference picture list modification processes can use at least one of the reference picture lists L0, L1, L2, and L3.

[0457] When the deblocking filter calculates the boundary strength according to the embodiments of the present invention, one or more motion vectors of the coding / decoding target block can be used up to a maximum of N. Here, N is a positive integer of 1 or more, and indicates 2, 3, 4, etc.

[0458] When the motion vector during motion vector prediction has at least one of 16 - pixel (16 - pel) unit, 8 - pixel (8 - pel) unit, 4 - pixel (4 - pel) unit, integer - pixel (integer - pel) unit, 1 / 2 - pixel (1 / 2 - pel) unit, 1 / 4 - pixel (1 / 4 - pel) unit, 1 / 8 - pixel (1 / 8 - pel) unit, 1 / 16 - pixel (1 / 16 - pel) unit, 1 / 32 - pixel (1 / 32 - pel) unit, and 1 / 64 - pixel (1 / 64 - pel) unit, the above - described embodiments of the present invention are applicable. Also, when executing the merge mode, the motion vector can be selectively used for each of the above - mentioned pixel units.

[0459] A slice type to which the above - described embodiments of the present invention are applied is defined, and the embodiments of the present invention can be applied according to the slice type.

[0460] For example, when the slice type is a T (Tri - predictive) - slice, at least three motion vectors are used to generate a prediction block, and a weighted sum of at least three prediction blocks is calculated and used as the final prediction block of the block to be coded / decoded. For example, when the slice type is a Q (Quad - predictive) - slice, at least four motion vectors are used to generate a prediction block, and a weighted sum of at least four prediction blocks is calculated and used as the final prediction block of the block to be coded / decoded.

[0461] The above - described embodiments of the present invention are applicable not only to the inter - prediction and motion compensation method using the merge mode, but also to the inter - prediction and motion compensation method using motion vector prediction, and the inter - prediction and motion compensation method using the skip mode and the like.

[0462] The shape of the block to which the above - described embodiments of the present invention are applied can have a square or non - square shape.

[0463] Above, with reference to FIGS. 12 to 32, the image encoding and decoding method using the merge mode according to the present invention has been described. Hereinafter, with reference to FIGS. 33 and 34, the image decoding method, image encoding method, image decoder, image encoder, and bit stream according to the present invention will be specifically described.

[0464] FIG. 33 is a diagram for explaining the image decoding method according to the present invention.

[0465] Referring to FIG. 33, a merge candidate list of a current block including at least one of merge candidates corresponding to each of a plurality of reference image lists can be generated (S3301).

[0466] Here, the merge candidates corresponding to each of the plurality of reference image lists can mean merge candidates having corresponding LX motion information of the reference image list LX. As an example, there may be an L0 merge candidate having L0 motion information, an L1 merge candidate having L1 motion information, an L2 merge candidate having L2 motion information, an L3 merge candidate having L3 motion information, and the like.

[0467] On the other hand, the merge candidate list can include at least one of a spatial merge candidate derived from a spatial neighboring block of the current block, a temporal merge candidate derived from a corresponding position block of the current block, a modified spatial merge candidate derived by modifying the spatial merge candidate, a modified temporal merge candidate derived by modifying the temporal merge candidate, and a merge candidate having a predefined motion information value. Here, the merge candidate having a predefined motion information value can be a zero merge candidate.

[0468] In this case, the spatial merge candidate can be derived from a lower block of a neighboring block adjacent to the current block. And the temporal merge candidate can be derived from a lower block of the corresponding position block of the current block.

[0469] On the other hand, the merge candidate list can further include a combined merge candidate derived using two or more of a spatial merge candidate, a temporal merge candidate, a modified spatial merge candidate, and a modified temporal merge candidate.

[0470] At least one motion information can be determined using the generated merge candidate list (S3302).

[0471] Using the determined at least one motion information, a predicted block of the current block can be generated (S3303).

[0472] Here, the step of generating the predicted block of the current block (S3303) can generate a plurality of temporary predicted blocks according to the inter prediction indicator of the current block, and apply at least one of a weight and an offset to the generated plurality of temporary predicted blocks to generate the predicted block of the current block.

[0473] In this case, at least one of the weight and the offset can be shared by a block smaller than a predetermined block size or deeper than a predetermined block depth.

[0474] On the other hand, the merge candidate list can be shared by a block smaller than a predetermined block size or deeper than a predetermined block depth.

[0475] And, when the current block is smaller than a predetermined block size or deeper than a predetermined block depth, the merge candidate list can be generated based on the upper block of the current block having the predetermined block size or the predetermined block depth.

[0476] The predicted block of the current block can be generated by applying information on the weighted sum to a plurality of predicted blocks generated based on a plurality of merge candidates or a plurality of merge candidate lists.

[0477] FIG. 34 is a diagram for explaining the image encoding method according to the present invention.

[0478] Referring to FIG. 34, a merge candidate list for the current block including at least one of the merge candidates corresponding to each of the plurality of reference image lists can be generated (S3401).

[0479] Using the generated merge candidate list, at least one motion information can be determined (S3402).

[0480] Then, using the determined at least one motion information, a predicted block of the current block can be generated (S3403).

[0481] The image decoder according to the present invention may be an image decoder including an inter prediction unit that generates a merge candidate list for the current block including at least one of the merge candidates corresponding to each of the plurality of reference image lists, determines at least one motion information using the merge candidate list, and generates a predicted block of the current block using the determined at least one motion information.

[0482] The image encoder according to the present invention may be an image encoder including an inter prediction unit that generates a merge candidate list for the current block including at least one of the merge candidates corresponding to each of the plurality of reference image lists, determines at least one motion information using the merge candidate list, and generates a predicted block of the current block using the determined at least one motion information.

[0483] The bitstream according to the present invention may be a bitstream generated by an image encoding method including a step of generating a merge candidate list for the current block including at least one of the merge candidates corresponding to each of the plurality of reference image lists, a step of determining at least one motion information using the merge candidate list, and a step of generating a predicted block of the current block using the determined at least one motion information.

[0484] In the above-described embodiments, these methods are described based on a flowchart with a series of steps or units. However, the present invention is not limited to the order of these steps, and a certain step may be performed in a different step and in a different order or simultaneously from those described above. Also, those having ordinary knowledge in the relevant technical field will be able to understand that the steps shown in the flowchart are not exclusive, and that other steps may be included, or one or more steps of the flowchart may be deleted without affecting the scope of the present invention.

[0485] The above-described embodiments include examples of various aspects. Although it is not possible to describe all possible combinations for showing various aspects, those having ordinary knowledge in the relevant technical field will be able to recognize that other combinations are possible. Therefore, it can be said that the present invention includes all various alternatives, modifications, and changes that fall within the scope of the following claims.

[0486] The embodiments of the present invention described above can be realized in the form of program instructions executable via various computer components and can be recorded on a computer-readable recording medium. The computer-readable recording medium can include program instructions, data files, data structures, etc. alone or in combination. The program instructions recorded on the computer-readable recording medium are those specially designed and configured for the present invention or those known and usable by those skilled in the field of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks (registered trademark), and magnetic tapes, optical recording media such as CD-ROMs, DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions such as ROMs, RAMs, and flash memories. Examples of program instructions include not only machine language codes created by compilers but also high-level language codes executable by a computer using an interpreter or the like. The hardware device can be configured to operate as one or more software modules for performing the processing according to the present invention, and vice versa.

[0487] As described above, the present invention has been described with reference to specific matters such as specific components, limited embodiments, and drawings, but this is only provided to assist in a more general understanding of the present invention, and the present invention is not limited to the above embodiments. Those with ordinary knowledge in the technical field to which the present invention pertains can make various modifications and variations from such descriptions.

[0488] Therefore, the idea of the present invention should not be defined as being limited to the above-described embodiments, and not only the following claims but also all those equivalently or equivalently modified from these claims belong to the scope of the idea of the present invention.

Industrial Applicability

[0489] The present invention can be used in an apparatus for encoding / decoding an image.

Claims

1. Generating a candidate list for the current block, the candidate list including at least one of a spatial candidate derived from spatially neighboring blocks of the current block and a temporal candidate derived from corresponding position blocks of the current block; Determining motion information using the candidate list; Generating a plurality of predicted blocks based on the determined motion information; Generating a final predicted block for the current block by applying weights and offsets to the generated plurality of predicted blocks; comprising wherein the weight is derived by weight index information that specifies the weight applied to the current block among a plurality of weights included in a weight set; the weight index information is obtained from a bit stream only when the size of the current block is equal to or greater than a predefined value; the weight index information is derived based on index information of neighboring blocks of the current block for specifying the weight applied to the current block; An image decoding method comprising.

2. The weight set includes positive weights and negative weights, The image decoding method according to claim 1.

3. The index information of the neighboring blocks is obtained from the bit stream, The image decoding method according to claim 1.

4. Generating a candidate list for the current block, the candidate list including at least one of a spatial candidate derived from spatially neighboring blocks of the current block and a temporal candidate derived from corresponding position blocks of the current block; Determining motion information using the candidate list; Generating a plurality of predicted blocks based on the determined motion information; Generating a final predicted block for the current block by applying weights and offsets to the generated plurality of predicted blocks; comprising weight index information for specifying the weight applied to the current block among a plurality of weights included in a weight set is encoded; the weight index information is encoded in a bit stream only when the size of the current block is equal to or greater than a predefined value; the weight index information is derived based on index information of neighboring blocks of the current block for specifying the weight applied to the current block; Image symbolization method.

5. Generating a candidate list for the current block, the candidate list including at least one of a spatial candidate derived from spatial neighboring blocks of the current block and a temporal candidate derived from corresponding position blocks of the current block; Determining motion information using the candidate list; Generating a plurality of prediction blocks based on the determined motion information; Applying weights and offsets to the generated plurality of prediction blocks to generate a final prediction block for the current block; Transmitting a bitstream obtained based on the final prediction block; comprising: Weight index information that specifies the weight applied to the current block among a plurality of weights included in a weight set is encoded; The weight index information is encoded in the bitstream only when the size of the current block is equal to or greater than a predefined value; The weight index information is derived based on index information of neighboring blocks of the current block for specifying the weight applied to the current block; Method for transmitting a bitstream.

Citation Information

Patent Citations

  • Method for encoding moving image and method for decoding moving image

    JP2004007379A

  • Implicit weighting of reference pictures in video encoder

    JP2010220265A