Image decoding method and image encoding method

By incorporating combined motion vector candidates and multiple prediction directions, the method enhances the encoding/decoding efficiency for high-resolution and high-quality images, addressing limitations in conventional motion compensation techniques.

JP2026020352APending Publication Date: 2026-02-06ELECTRONICS & TELECOMM RES INST
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025207070
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2016-05-24
Filing Date
2025-11-27
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Conventional motion compensation methods in image encoding/decoding are limited by the use of unidirectional and bidirectional prediction, which restricts the improvement in coding efficiency for high-resolution and high-quality images.

Method used

The method employs combined motion vector candidates, including spatial, temporal, and predefined values, along with unidirectional, bidirectional, tridirectional, and quaternary predictions to enhance encoding/decoding efficiency.

Benefits of technology

This approach improves image encoding/decoding efficiency by utilizing a broader range of motion vector candidates and prediction directions, optimizing the compression and transmission of high-resolution and high-quality images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026020352000001_ABST
    Figure 2026020352000001_ABST
Patent Text Reader

Abstract

To provide a technique capable of improving encoding / decoding efficiency of an image.SOLUTION: Deriving a plurality of motion vectors of a current block according to an inter prediction direction of the current block, determining a plurality of prediction blocks of the current block by using the plurality of motion vectors, obtaining a final prediction block of the current block based on a weighted sum of the plurality of prediction blocks, and obtaining a residual block of the current block by inverse transformation; The current block may be reconstructed based on the final prediction block and the residual block, the inverse transform may be performed using one of predefined transform sets, a weight of the current block for the weighted sum may be derived from weight index information of the current block specifying one of weights included in a predefined weight set, and the weight index information of the current block may be derived based on weight index information of a neighboring block of the current block.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image encoding / decoding method and apparatus, and more particularly to a method and apparatus for performing motion compensation using motion vector prediction. [Background technology]

[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various application fields. As image data becomes higher in resolution and quality, the amount of data increases relatively compared to conventional image data. Therefore, when image data is transmitted over conventional media such as wired or wireless broadband lines or stored using conventional storage media, transmission and storage costs increase. To solve the problems that arise with the increase in resolution and quality of image data, highly efficient image encoding / decoding technologies for images with higher resolution and quality are required.

[0003] There are various image compression techniques, such as inter-prediction techniques that predict pixel values ​​contained in a current picture from pictures before or after the current picture, intra-prediction techniques that predict pixel values ​​contained in a current picture using pixel information within the current picture, transformation and quantization techniques that compress the energy of residual signals, and entropy coding techniques that assign short codes to values ​​that occur frequently and long codes to values ​​that occur less frequently.By using such image compression techniques, image data can be effectively compressed and transmitted or stored.

[0004] Conventional motion compensation adds only spatial motion vector candidates, temporal motion vector candidates, and zero motion vector candidates to a motion vector candidate list and uses only unidirectional prediction and bidirectional prediction, which limits the improvement in coding efficiency. Summary of the Invention [Problem to be solved by the invention]

[0005] The present invention can provide a method and apparatus for performing motion compensation using combined motion vector candidates to improve image encoding / decoding efficiency.

[0006] The present invention can provide a method and apparatus for performing motion compensation using unidirectional prediction, bidirectional prediction, tridirectional prediction, and quaternary prediction to improve image encoding / decoding efficiency. [Means for solving the problem]

[0007] The image decoding method of the present invention may include the steps of generating a plurality of motion vector candidate lists according to the inter-prediction direction of a current block, deriving a plurality of motion vectors for the current block using the plurality of motion vector candidate lists, determining a plurality of predictive blocks for the current block using the plurality of motion vectors, and obtaining a final predictive block for the current block based on the plurality of predictive blocks.

[0008] The image encoding method of the present invention may include the steps of generating a plurality of motion vector candidate lists according to the inter-prediction direction of a current block, deriving a plurality of motion vectors for the current block using the plurality of motion vector candidate lists, determining a plurality of predictive blocks for the current block using the plurality of motion vectors, and obtaining a final predictive block for the current block based on the plurality of predictive blocks.

[0009] In the image decoding / encoding method, the inter-prediction direction may indicate one-way or multi-way prediction, and the multi-way prediction may include prediction in three or more directions.

[0010] In the image decoding / encoding method, the motion vector candidate list can be generated for each reference image list.

[0011] In the image decoding / encoding method, the motion vector candidate list may include at least one of spatial motion vector candidates derived from spatially neighboring blocks of the current block, temporal motion vector candidates derived from corresponding position blocks of the current block, and motion vector candidates of predefined values.

[0012] In the image decoding / encoding method, the motion vector candidate list may include a combined motion vector candidate generated by combining two or more of the spatial motion vector candidate, the temporal motion vector candidate, and the predefined value motion vector candidate.

[0013] In the image decoding / encoding method, the final predicted block can be determined based on a weighted sum of the plurality of predicted blocks.

[0014] In the image decoding / encoding method, weights applied to the plurality of predicted blocks may be determined based on weighted prediction values ​​and weighted differential values. [Effects of the Invention]

[0015] The present invention provides a method and apparatus for performing motion compensation using combined motion vector candidates to improve image encoding / decoding efficiency.

[0016] The present invention provides a method and apparatus for performing motion compensation using unidirectional prediction, bidirectional prediction, tridirectional prediction, and quaternary prediction to improve image encoding / decoding efficiency. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a block diagram showing a configuration of an embodiment of an encoding device to which the present invention is applied. [Figure 2] 1 is a block diagram showing a configuration of an embodiment of a decoding device to which the present invention is applied. [Figure 3]FIG. 2 is a schematic diagram showing the division structure of an image when encoding and decoding the image. [Figure 4] FIG. 1 is a diagram showing the form of a prediction unit (PU) that can be included in a coding unit (CU). [Figure 5] FIG. 2 is a diagram showing the form of a transform unit (TU) that can be included in a coding unit (CU). [Figure 6] FIG. 1 is a diagram illustrating an embodiment of intra-prediction processing. [Figure 7] FIG. 1 is a diagram illustrating an embodiment of inter prediction processing. [Figure 8] FIG. 10 is a diagram illustrating a transform set according to an intra prediction mode. [Figure 9] FIG. 10 is a diagram illustrating a conversion process. [Figure 10] FIG. 10 is a diagram illustrating scanning of quantized transform coefficients. [Figure 11] FIG. 10 is a diagram for explaining block division. [Figure 12] 1 is a flowchart illustrating an image coding method according to the present invention. [Figure 13] 1 is a flowchart illustrating an image decoding method according to the present invention. [Figure 14] 10A and 10B are diagrams illustrating an example of deriving spatial motion vector candidates for a current block. [Figure 15] 10A and 10B are diagrams illustrating an example of deriving temporal motion vector candidates for a current block. [Figure 16] 10A and 10B are diagrams illustrating an example of scaling a motion vector of a corresponding position block to derive a temporal motion vector candidate for a current block. [Figure 17] FIG. 10 is a diagram showing an example in which a motion vector candidate list is generated. [Figure 18] FIG. 10 is a diagram showing an example of adding a motion vector having a predetermined value to a motion vector candidate list. [Figure 19] FIG. 10 is a diagram showing an example in which a motion vector candidate is removed from the motion vector candidate list. [Figure 20]FIG. 10 is a diagram showing an example of a motion vector candidate list. [Figure 21] FIG. 10 is a diagram illustrating an example of deriving a predicted motion vector candidate for a current block from a motion vector candidate list. [Figure 22A] FIG. 10 illustrates a syntax for information about motion compensation. [Figure 22B] FIG. 10 illustrates a syntax for information about motion compensation. DETAILED DESCRIPTION OF THE INVENTION

[0018] Because the present invention is susceptible to various modifications and can have various embodiments, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this does not limit the present invention to the specific embodiments, but rather should be understood to include all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention. In the drawings, like reference numerals indicate the same or similar functions throughout the various aspects. The shape and size of elements in the drawings may be exaggerated for clarity. The detailed description of exemplary embodiments below refers to the accompanying drawings, which show specific embodiments by way of example. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that various embodiments, although different from one another, are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein in connection with one embodiment can be implemented in various embodiments without departing from the spirit and scope of the present invention. It should also be understood that the location or arrangement of individual components within each disclosed embodiment can be changed without departing from the spirit and scope of the embodiment. Accordingly, the following detailed description is not to be taken in a limiting sense, and the scope of the exemplary embodiments is limited only by the appended claims, if properly recited, and to the full scope of equivalents to which those claims are entitled.

[0019] In the present invention, the terms "first," "second," etc. may be used to describe various elements, but these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element can be called a second element, and similarly, a second element can be called a first element, without departing from the scope of the present invention. The term "and / or" includes a combination of multiple related listed items or any of multiple related listed items.

[0020] When a component of the present invention is said to be "coupled" or "connected" to another component, it should be understood that the component may be directly coupled or connected to the other component, but that there may be other components between them. In contrast, when a component is said to be "directly coupled" or "directly connected" to another component, it should be understood that there are no other components between them.

[0021] The components shown in the embodiments of the present invention are illustrated independently to show different characteristic functions, and do not mean that each component is composed of separate hardware or a single software unit. That is, each component is included in the respective components for the convenience of explanation, and at least two of the components may be combined to form a single component, or each component may be divided into multiple components to perform a function. Such integrated and separated embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.

[0022] The terms used in the present invention are merely used to describe specific embodiments and do not limit the present invention. A singular expression includes a plural expression unless the context clearly dictates otherwise. In the present invention, terms such as "comprise" or "have" specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the presence or additional possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. In other words, in the present invention, a description that a specific configuration "comprises" does not exclude configurations other than the specified configuration, but means that additional configurations may be included within the scope of the implementation of the present invention or the technical idea of ​​the present invention.

[0023] Some components of the present invention may not be essential components that perform essential functions in the present invention, but may be optional components simply for improving performance. The present invention can be realized by including only components that are essential for achieving the essence of the present invention, excluding components used simply for improving performance, and a structure including only essential components excluding optional components used simply for improving performance is also included in the scope of the present invention.

[0024] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In describing the embodiments of this specification, if it is determined that a detailed description of related known configurations or functions may obscure the gist of this specification, the detailed description will be omitted, and the same reference numerals will be used for the same components in the drawings, and duplicate descriptions of the same components will be omitted.

[0025] In the following, an image may refer to a picture constituting a video, or may refer to the video itself. For example, "encoding and / or decoding an image" may mean "encoding and / or decoding a video" or "encoding and / or decoding one of the images constituting a video." Here, a picture may have the same meaning as an image.

[0026] Terminology

[0027] Encoder: Can refer to a device that performs encoding.

[0028] Decoder: Can refer to a device that performs decoding.

[0029] Parsing: This can refer to determining the value of a syntax element by entropy decoding, or it can refer to entropy decoding itself.

[0030] Block: An MxN array of samples, where M and N are positive integers, and a block can generally refer to a two-dimensional array of samples.

[0031] Sample: A basic unit constituting a block, which can represent values ​​from 0 to 2Bd-1 depending on the bit depth (Bd). In the present invention, the terms picture element and pixel can be used interchangeably with sample.

[0032] Unit: May refer to a unit of image encoding and decoding. In image encoding and decoding, a unit may be an area generated by dividing an image. Furthermore, when an image is divided into smaller units for encoding or decoding, a unit may refer to the divided units. In image encoding and decoding, a predefined process may be performed on each unit. A unit may be further divided into sub-units having a smaller size than the unit. Depending on the function, a unit may refer to a block, macroblock, coding tree unit, coding tree block, coding unit, coding block, prediction unit, prediction block, transform unit, transform block, etc. Furthermore, a unit may refer to including a luma component block, a corresponding chroma component block, and syntax elements for each block to distinguish it from a block. The unit may have various sizes and shapes, and in particular, the shape of the unit may include not only a rectangle but also geometric shapes that can be expressed two-dimensionally, such as a square, a trapezoid, a triangle, a pentagon, etc. Furthermore, the unit information may include at least one of a unit type indicating a coding unit, a prediction unit, a transform unit, etc., a unit size, a unit depth, an encoding and decoding order of the unit, etc.

[0033] Reconstructed Neighbor Unit: This can refer to a unit that has been spatially / temporally coded or decoded around a unit to be coded / decoded and reconstructed. In this case, the reconstructed neighbor unit can refer to a reconstructed neighbor block.

[0034] Neighbor block: This may refer to a block adjacent to the block to be coded / decoded. A block adjacent to the block to be coded / decoded may refer to a block whose boundary is adjacent to the block to be coded / decoded. A neighbor block may refer to a block located at an adjacent vertex of the block to be coded / decoded. A neighbor block may also refer to a reconstructed neighbor block.

[0035] Unit depth: This refers to the degree to which a unit is divided. In a tree structure, the root node has the shallowest depth, and the leaf node has the deepest depth.

[0036] Symbol: It can refer to a syntax element of a unit to be coded / decoded, a coding parameter, a value of a transform coefficient, etc.

[0037] Parameter Set: This may correspond to header information among structures in a bitstream, and may include at least one of a video parameter set, a sequence parameter set, a picture parameter set, and an adaptation parameter set. In addition, the parameter set may include slice header and tile header information.

[0038] Bitstream: Can mean a string of bits containing coded image information.

[0039] Prediction unit: A basic unit for performing inter-prediction or intra-prediction and compensation therefor. One prediction unit may be divided into multiple smaller partitions. In this case, each of the multiple partitions may be used as a basic unit for performing the prediction and compensation, and the partitions into which the prediction unit is divided may also be used as prediction units. Prediction units may have various sizes and shapes. In particular, the shape of a prediction unit may include not only a rectangle but also geometric shapes that can be expressed in two dimensions, such as a square, trapezoid, triangle, or pentagon.

[0040] Prediction Unit Partition: This may refer to the shape in which a prediction unit is divided.

[0041] Reference Picture List: This can refer to a list containing one or more reference pictures used for inter prediction or motion compensation. Types of reference picture lists can include LC (List Combined), L0 (List 0), L1 (List 1), L2 (List 2), and L3 (List 3). Inter prediction can use one or more reference picture lists.

[0042] Inter Prediction Indicator: This may refer to the inter prediction direction (unidirectional prediction, bidirectional prediction, etc.) of the block to be coded / decoded during inter prediction, may refer to the number of reference images used when the block to be coded / decoded generates a predicted block, or may refer to the number of predicted blocks used when the block to be coded / decoded performs inter prediction or motion compensation.

[0043] Reference Picture Index: This can refer to an index to a specific reference picture from a reference picture list.

[0044] Reference Picture: This can refer to a picture that a specific unit references for inter-prediction or motion compensation, and can also be called a reference picture.

[0045] Motion Vector: A two-dimensional vector used in inter-prediction or motion compensation, which may represent an offset between a current image to be coded / decoded and a reference image. For example, (mvX, mvY) may represent a motion vector, where mvX represents the horizontal component and mvY represents the vertical component.

[0046] Motion Vector Candidate: This can refer to a unit that is a prediction candidate when predicting a motion vector, or the motion vector of that unit.

[0047] Motion Vector Candidate List: This may refer to a list constructed using motion vector candidates.

[0048] Motion Vector Candidate Index: An indicator that indicates a motion vector candidate in a motion vector candidate list, and can also be called an index of a motion vector predictor.

[0049] Motion Information: This can refer to information including at least one of a motion vector, a reference image index, an inter prediction indicator, reference image list information, a reference image, a motion vector candidate, a motion vector candidate index, etc.

[0050] Merge Candidate List: This may refer to a list constructed using merge candidates.

[0051] Merge Candidate: May include spatial merge candidates, temporal merge candidates, combined merge candidates, combined bi-predictive merge candidates, zero merge candidates, etc., and merge candidates may include motion information such as prediction type information, reference picture index for each list, and motion vector.

[0052] Merge Index: This may refer to information indicating a merge candidate in a merge candidate list. The merge index may also indicate a block from which the merge candidate is derived among blocks reconstructed to be spatially / temporally adjacent to the current block. The merge index may also indicate at least one of the motion information items of the merge candidate.

[0053] Transform unit: A basic unit used in encoding / decoding residual signals, such as transform, inverse transform, quantization, inverse quantization, and transform coefficient encoding / decoding. A transform unit can be divided into multiple smaller transform units. Transform units can have various sizes and shapes, and in particular, the shape of a transform unit can include not only rectangles but also geometric shapes that can be expressed two-dimensionally, such as squares, trapezoids, triangles, and pentagons.

[0054] Scaling: can refer to the process of multiplying transform coefficient levels by a factor, resulting in transform coefficients. Scaling can also be called dequantization.

[0055] Quantization Parameter: This can refer to a value used when scaling a transform coefficient level in quantization and inverse quantization. In this case, the quantization parameter can be a value mapped to a quantization step size.

[0056] Delta Quantization Parameter: This can refer to a difference value between a predicted quantization parameter and a quantization parameter of a unit to be coded / decoded.

[0057] Scan: This can refer to a method of sorting the order of coefficients in a block or matrix. For example, sorting a two-dimensional array into a one-dimensional array is called a scan, and sorting a one-dimensional array into a two-dimensional array can also be called a scan or inverse scan.

[0058] Transform Coefficient: A coefficient value generated after transformation. In the present invention, the quantized transform coefficient level obtained by applying quantization to a transform coefficient may also be included in the meaning of the transform coefficient.

[0059] Non-zero Transform Coefficient: This can refer to a transform coefficient whose magnitude of the transform coefficient value is not 0, or a transform coefficient level whose magnitude of the value is not 0.

[0060] Quantization Matrix: A matrix used in quantization or inverse quantization to improve the subjective or objective image quality. A quantization matrix can also be called a scaling list.

[0061] Quantization Matrix Coefficient: This can refer to each element in a quantization matrix. The quantization matrix coefficient can also be called a matrix coefficient.

[0062] Default Matrix: This may refer to a predetermined quantization matrix that is predefined in the encoder and decoder.

[0063] Non-default Matrix: This may refer to a quantization matrix that is not predefined in the encoder and decoder and is transmitted / received by the user.

[0064] Coding tree unit: A coding tree unit can consist of one luminance (Y) coding tree block and two associated chrominance (Cb, Cr) coding tree blocks. Each coding tree unit can be divided using one or more partitioning methods, such as a quad tree or a binary tree, to form subunits such as coding units, prediction units, and transform units. This term can be used to refer to pixel blocks that are the processing units in the image decoding / encoding process, such as the division of an input image.

[0065] Coding Tree Block: A term that can be used to refer to any of the Y coding tree block, Cb coding tree block, and Cr coding tree block.

[0066] FIG. 1 is a block diagram showing the configuration of an embodiment of an encoding device to which the present invention is applied.

[0067] The encoding device 100 may be a video encoding device or an image encoding device. A video may include one or more images. The encoding device 100 may encode one or more images of the video in a time-sequential manner.

[0068] Referring to FIG. 1, the encoding device 100 may include a motion prediction unit 111, a motion compensation unit 112, an intra prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.

[0069] The encoding device 100 may perform encoding on an input image in intra mode and / or inter mode. The encoding device 100 may generate a bitstream through encoding of the input image and output the generated bitstream. When an intramode is used as a prediction mode, the switch 115 may be switched to intra. When an intermode is used as a prediction mode, the switch 115 may be switched to inter. Here, the intramode may refer to an intra prediction mode, and the intermode may refer to an inter prediction mode. The encoding device 100 may generate a prediction block for an input block of the input image. After the prediction block is generated, the encoding device 100 may encode a residual between the input block and the prediction block. The input image may also be referred to as a current image currently being encoded. The input block may also be referred to as a current block currently being encoded or a block to be encoded.

[0070] When the prediction mode is the intra mode, the intra prediction unit 120 may also use pixel values ​​of previously coded blocks adjacent to the current block as reference pixels. The intra prediction unit 120 may perform spatial prediction using the reference pixels and generate prediction samples for the input block through the spatial prediction. Here, the term "intra prediction" may refer to intra prediction.

[0071] When the prediction mode is inter mode, the motion prediction unit 111 can search for an area that best matches the input block in a reference image during motion prediction processing, and can derive a motion vector using the searched area. The reference image can be stored in the reference picture buffer 190.

[0072] The motion compensation unit 112 may generate a prediction block by performing motion compensation using a motion vector. Here, the motion vector may be a two-dimensional vector used for inter prediction. Also, the motion vector may indicate an offset between a current image and a reference image. Here, inter prediction may refer to inter-frame prediction.

[0073] When a motion vector value does not have an integer value, the motion prediction unit 111 and the motion compensation unit 112 may generate a prediction block by applying an interpolation filter to a portion of a reference image. To perform inter prediction or motion compensation, it is possible to determine whether the motion prediction and motion compensation method of a prediction unit included in the corresponding coding unit is skip mode, merge mode, AMVP mode, or current picture reference mode, based on the coding unit, and perform inter prediction or motion compensation according to each mode. Here, the current picture reference mode may refer to a prediction mode using an already reconstructed region in the current picture to which the current block belongs. A motion vector for the current picture reference mode may be defined to identify the already reconstructed region. Whether the current block is coded in the current picture reference mode may be coded using a reference image index of the current block.

[0074] The subtractor 125 can use the difference between the input block and the predicted block to generate a residual block, also referred to as a residual signal.

[0075] The transform unit 130 may perform a transform on the residual block to generate transform coefficients and output the transform coefficients. Here, the transform coefficients may be coefficient values ​​generated by performing a transform on the residual block. When a transform skip mode is applied, the transform unit 130 may skip transforming the residual block.

[0076] By applying quantization to the transform coefficients, quantized transform coefficient levels can be generated. Hereinafter, in the embodiments, the quantized transform coefficient levels may also be referred to as transform coefficients.

[0077] The quantization unit 140 may generate quantized transform coefficient levels by quantizing the transform coefficients based on the quantization parameter, and may output the quantized transform coefficient levels. In this case, the quantization unit 140 may quantize the transform coefficients using a quantization matrix.

[0078] The entropy coding unit 150 can generate a bitstream by performing entropy coding based on a probability distribution on values ​​calculated by the quantization unit 140 or coding parameter values ​​calculated in the coding process, and can output the bitstream. The entropy coding unit 150 can perform entropy coding on information for decoding the image in addition to information on pixels of the image. For example, the information for decoding the image can include syntax elements.

[0079] When entropy coding is applied, fewer bits are assigned to symbols with higher occurrence probabilities and more bits are assigned to symbols with lower occurrence probabilities to represent the symbols, thereby reducing the size of the bit string for the symbol to be coded. Therefore, entropy coding can improve the compression performance of image coding. The entropy coding unit 150 can use coding methods such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding) for entropy coding. For example, the entropy coding unit 150 can perform entropy coding using a variable length coding / code (VLC) table. The entropy coding unit 150 can also derive a binarization method for the target symbol and a probability model for the target symbol / bin, and then perform arithmetic coding using the derived binarization method or probability model.

[0080] To encode the transform coefficient levels, the entropy coding unit 150 may convert two-dimensional block shape coefficients into one-dimensional vectors through a transform coefficient scanning method. For example, the entropy coding unit 150 may convert the two-dimensional block shape coefficients into one-dimensional vectors by scanning the coefficients of the blocks using upright scanning. Instead of upright scanning, vertical scanning, in which the two-dimensional block shape coefficients are scanned in the column direction, or horizontal scanning, in which the two-dimensional block shape coefficients are scanned in the row direction, may be used depending on the size of the transform unit and the intra prediction mode. That is, which scanning method to use among upright scanning, vertical scanning, and horizontal scanning may be determined depending on the size of the transform unit and the intra prediction mode.

[0081] The coding parameters may include not only information encoded by an encoder and transmitted to a decoder like syntax elements, but also information derived in the encoding or decoding process, and may refer to information required for encoding or decoding an image. For example, the coding parameters may include block size, block depth, block partition information, unit size, unit depth, unit partition information, quadtree partition flag, binary tree partition flag, binary tree partition direction, intra prediction mode, intra prediction direction, reference sample filtering method, prediction block boundary filtering method, filter tap, filter coefficient, inter prediction mode, motion information, motion vector, reference image index, inter prediction direction, inter prediction indicator, reference image list, motion vector prediction, motion vector candidate list, whether motion merge mode is used, motion merge candidate, motion merge candidate list, whether skip mode is used, type of interpolation filter, motion vector size, precision of motion vector representation, transform type, transform size, information on whether additional (secondary) transform is used, information on whether residual signal is present, coded block pattern, coded block flag, etc. The coding parameters may include at least one value or a combination form of information on a luminance signal or a chrominance signal, such as a quantization flag, a quantization parameter, a quantization matrix, in-loop filter information, information on whether to apply an in-loop filter, in-loop filter coefficients, a binarization / de-binarization method, a context model, a context bin, a bypass bin, a transform coefficient, a transform coefficient level, a scanning method for the transform coefficient level, an image display / output order, slice identification information, a slice type, slice division information, a tile identification information, a tile type, a tile division information, a picture type, a bit depth, and information on a luminance signal or a chrominance signal.

[0082] The residual signal may refer to the difference between an original signal and a predicted signal. Alternatively, the residual signal may be a signal generated by transforming the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. The residual block may be a block-based residual signal.

[0083] When the encoding device 100 performs encoding using inter prediction, the encoded current image can be used as a reference image for other images to be processed later. Thus, the encoding device 100 can further decode the encoded current image and store the decoded image as a reference image. For decoding, the encoded current image can be subjected to inverse quantization and inverse transform.

[0084] The quantized coefficients can be dequantized in an inverse quantization unit 160 and inverse transformed in an inverse transform unit 170. The dequantized and inverse transformed coefficients can be combined with a prediction block via an adder 175. A reconstructed block can be generated by combining the dequantized and inverse transformed coefficients with the prediction block.

[0085] The reconstructed block may pass through a filter unit 180. The filter unit 180 may apply at least one of a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF) to the reconstructed block or image. The filter unit 180 is also referred to as an in-loop filter.

[0086] A deblocking filter can remove block artifacts that occur at boundaries between blocks. To determine whether to perform deblocking filtering, it can be determined whether to apply a deblocking filter to a current block based on pixels included in several columns or rows included in the block. When applying a deblocking filter to a block, a strong filter or a weak filter can be applied depending on the required deblocking filtering strength. In addition, when applying a deblocking filter, horizontal filtering and vertical filtering can be performed in parallel during vertical filtering and horizontal filtering.

[0087] The sample adaptive offset can add an appropriate offset value to a pixel value to compensate for encoding errors. The sample adaptive offset can correct the offset between a deblocked image and an original image on a pixel-by-pixel basis. To perform offset correction for a specific picture, the pixels included in the image can be divided into a certain number of regions, and then the region to be offset is determined and an offset is applied to the corresponding region. Alternatively, an offset can be applied taking into account edge information of each pixel.

[0088] The adaptive loop filter can perform filtering based on a comparison between the reconstructed image and the original image. After dividing the pixels included in the image into predetermined groups, a filter to be applied to each group can be determined, and differential filtering can be performed for each group. Information related to whether to apply an adaptive loop filter can be transmitted for each coding unit (CU) of the luminance signal, and the shape and filter coefficients of the adaptive loop filter applied can vary depending on each block. Alternatively, an adaptive loop filter of the same type (fixed type) can be applied regardless of the characteristics of the block to which it is applied.

[0089] The reconstructed blocks that have passed through the filter unit 180 can be stored in a reference picture buffer 190 .

[0090] FIG. 2 is a block diagram showing the configuration of an embodiment of a decoding device to which the present invention is applied.

[0091] The decoding device 200 may be a video decoding device or an image decoding device.

[0092] Referring to FIG. 2, the decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, a motion compensation unit 250, an adder 255, a filter unit 260, and a reference picture buffer 270.

[0093] The decoding device 200 can receive a bitstream output from the encoding device 100. The decoding device 200 can perform decoding on the bitstream in intra mode or inter mode. The decoding device 200 can also generate a reconstructed image through decoding and output the reconstructed image.

[0094] If the prediction mode used for decoding is an intra mode, the switch may be switched to intra. If the prediction mode used for decoding is an inter mode, the switch may be switched to inter.

[0095] The decoding device 200 can obtain a reconstructed residual block from an input bitstream and generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block, which is a block to be decoded, by adding the reconstructed residual block and the prediction block. The block to be decoded may also be referred to as a current block.

[0096] The entropy decoder 210 may generate symbols by performing entropy decoding based on a probability distribution on the bitstream. The generated symbols may include symbols in the form of quantized transform coefficient levels. Here, the entropy decoding method may be similar to the entropy coding method described above. For example, the entropy decoding method may be the inverse process of the entropy coding method described above.

[0097] To decode transform coefficient levels, the entropy decoding unit 210 may convert one-dimensional vector form coefficients into two-dimensional block form coefficients using a transform coefficient scanning method. For example, the two-dimensional block form coefficients may be converted by scanning the coefficients of the block using upright scanning. Depending on the size of the transform unit and the intra prediction mode, vertical scanning or horizontal scanning may be used instead of upright scanning. That is, it is possible to determine which scanning method to use, among upright scanning, vertical scanning, or horizontal scanning, depending on the size of the transform unit and the intra prediction mode.

[0098] The quantized transform coefficient levels can be inversely quantized by the inverse quantization unit 220 and inversely transformed by the inverse transform unit 230. A reconstructed residual block can be generated as a result of the inverse quantization and inverse transformation of the quantized transform coefficient levels. In this case, the inverse quantization unit 220 can apply a quantization matrix to the quantized transform coefficient levels.

[0099] When intra mode is used, the intra prediction unit 240 can generate a predicted block by performing spatial prediction using pixel values ​​of already decoded blocks that are neighboring the block to be decoded.

[0100] When the inter mode is used, the motion compensation unit 250 may generate a prediction block by performing motion compensation using a motion vector and a reference image stored in the reference picture buffer 270. If the value of the motion vector does not have an integer value, the motion compensation unit 250 may generate a prediction block by applying an interpolation filter to a portion of the reference image. To perform motion compensation, the motion compensation unit 250 may determine whether the motion compensation method of the prediction unit included in the corresponding coding unit is skip mode, merge mode, AMVP mode, or current picture reference mode, based on the coding unit, and perform motion compensation according to the determined mode. Here, the current picture reference mode may refer to a prediction mode using an already reconstructed region in the current picture to which the current block belongs. A motion vector for the current picture reference mode can be used to identify the already reconstructed region. A flag or index indicating whether the current block is coded in the current picture reference mode may be signaled or may be inferred using the reference image index of the current block. The current picture for the current picture reference mode may be located at a fixed position (e.g., the position where the reference picture index is 0 or the last position) in the reference picture list for the block to be decoded, or may be variably located in the reference picture list, and for this purpose, a separate reference picture index indicating the position of the current picture may be signaled.

[0101] The reconstructed residual block and the prediction block may be added via an adder 255. A block generated by adding the reconstructed residual block and the prediction block may be passed through a filter unit 260. The filter unit 260 may apply at least one of a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the reconstructed block or the reconstructed image. The filter unit 260 may output the reconstructed image. The reconstructed image may be stored in a reference picture buffer 270 and used for inter prediction.

[0102] 3 is a schematic diagram showing an image division structure when encoding and decoding an image, and shows an embodiment in which one unit is divided into multiple sub-units.

[0103] To efficiently divide an image, a coding unit (CU) can be used in encoding and decoding. Here, the coding unit may refer to a coding unit. A unit may be a term that collectively refers to 1) a syntax element and 2) a block including image samples. For example, "division of a unit" may mean "division of blocks corresponding to the unit." Block division information may include information about the depth of the unit. The depth information may indicate the number of times and / or the degree to which the unit is divided.

[0104] Referring to FIG. 3, an image 300 is sequentially divided into largest coding units (LCUs), and a division structure is determined for each LCU. Here, LCU can be used interchangeably with coding tree unit (CTU). One unit can be hierarchically divided using depth information based on a tree structure. Each divided sub-unit can have depth information. The depth information indicates the number and / or degree to which a unit is divided, and can therefore include information about the size of the sub-unit.

[0105] The partition structure may refer to the distribution of coding units (CUs) within the LCU 310. A CU may be a unit for efficiently encoding an image. This distribution may be determined by whether or not a single CU is partitioned into multiple CUs (a positive integer equal to or greater than 2, including 2, 4, 8, 16, etc.). The width and height of a CU generated by partitioning may be half the width and height of the CU before partitioning, respectively, or may be smaller than the width and height of the CU before partitioning depending on the number of partitions. A partitioned CU may be recursively partitioned into multiple CUs with reduced widths and heights in a similar manner.

[0106] In this case, the division of CUs may be performed recursively up to a predetermined depth. Depth information indicates the size of a CU and may be stored for each CU. For example, the depth of an LCU may be 0, and the depth of a smallest coding unit (SCU) may be a predefined maximum depth. Here, as described above, an LCU may be a coding unit having the largest coding unit size, and an SCU may be a coding unit having the smallest coding unit size.

[0107] The division starts from LCU 310, and each time the width and height of the CU decrease due to division, the depth of the CU increases by 1. For each depth, an undivided CU can have a size of 2Nx2N. For a divided CU, a 2Nx2N CU can be divided into multiple NxN CUs. The size of N is halved each time the depth increases by 1.

[0108] For example, when one coding unit is divided into four coding units, the width and height of the four divided coding units may be half the size of the width and height of the coding unit before division. As an example, when a coding unit of 32x32 size is divided into four coding units, each of the four divided coding units may have a size of 16x16. When one coding unit is divided into four coding units, it can be said that the coding unit is divided into a quad-tree.

[0109] For example, when one coding unit is divided into two coding units, the width or height of the two divided coding units may be half the width or height of the coding unit before division. As an example, when a 32x32 coding unit is divided vertically into two coding units, each of the two divided coding units may have a size of 16x32. As an example, when a 32x32 coding unit is divided horizontally into two coding units, each of the two divided coding units may have a size of 32x16. When one coding unit is divided into two coding units, the coding unit can be said to be divided into a binary tree.

[0110] Referring to FIG. 3, an LCU with a depth of 0 may be 64x64 pixels. 0 may be the minimum depth. An SCU with a depth of 3 may be 8x8 pixels. 3 may be the maximum depth. In this case, a CU of 64x64 pixels, which is an LCU, can be represented with a depth of 0. A CU of 32x32 pixels can be represented with a depth of 1. A CU of 16x16 pixels can be represented with a depth of 2. A CU of 8x8 pixels, which is an SCU, can be represented with a depth of 3.

[0111] In addition, information about whether a CU is split can be expressed using CU split information. The split information may be 1-bit information. All CUs except for SCUs may include split information. For example, if the value of the split information is 0, the CU may not be split, and if the value of the split information is 1, the CU may be split.

[0112] FIG. 4 is a diagram showing the form of a prediction unit PU that can be included in a coding unit CU.

[0113] Among the CUs divided from the LCU, CUs that are not further divided may be divided into one or more Prediction Units (PUs). This process may also be referred to as division.

[0114] A PU may be a basic unit for prediction. A PU can be coded and decoded in either skip mode, inter mode, or intra mode. A PU can be divided in various ways depending on the mode.

[0115] Also, the coding unit is not divided into prediction units, and the coding unit and prediction unit can have the same size.

[0116] In skip mode, there may be no partitioning within a CU, as shown in Figure 4. Skip mode can support a 2Nx2N mode 410 with the same size as the CU without partitioning.

[0117] In inter mode, eight division modes within a CU can be supported. For example, in inter mode, 2Nx2N mode 410, 2NxN mode 415, Nx2N mode 420, NxN mode 425, 2NxnU mode 430, 2NxnD mode 435, nLx2N mode 440, and nRx2N mode 445 can be supported. In intra mode, 2Nx2N mode 410 and NxN mode 425 can be supported.

[0118] A coding unit can be divided into one or more prediction units, and a prediction unit can also be divided into one or more prediction units.

[0119] For example, when one prediction unit is divided into four prediction units, the horizontal and vertical widths of the divided four prediction units may be half the size of the horizontal and vertical widths of the prediction unit before division. As an example, when a prediction unit of 32x32 size is divided into four prediction units, each of the divided four prediction units may have a size of 16x16. When one prediction unit is divided into four prediction units, it can be said that the prediction units are divided into a quad-tree shape.

[0120] For example, when one prediction unit is divided into two prediction units, the horizontal or vertical width of the two divided prediction units may be half the size of the horizontal or vertical width of the prediction unit before division. As an example, when a 32x32 prediction unit is vertically divided into two prediction units, each of the two divided prediction units may have a size of 16x32. As an example, when a 32x32 prediction unit is horizontally divided into two prediction units, each of the two divided prediction units may have a size of 32x16. When one prediction unit is divided into two prediction units, it can be said that the prediction unit is divided into a binary tree.

[0121] FIG. 5 shows the form of a transform unit TU that can be included in a coding unit CU.

[0122] A transform unit (TU) may be a basic unit used for transform, quantization, inverse transform, and inverse quantization within a CU. A TU may have a shape such as a square or a rectangle. The size of a TU may be determined depending on the size and / or shape of the CU.

[0123] Among CUs divided from an LCU, CUs that are not further divided into CUs can be divided into one or more TUs. In this case, the division structure of the TUs may be a quad-tree structure. For example, as shown in FIG. 5, one CU 510 can be divided one or more times using a quad-tree structure. When one CU is divided one or more times, it is said to be divided recursively. Through division, one CU 510 can be composed of TUs of various sizes. Alternatively, the CU can be divided into one or more TUs based on the number of vertical lines and / or horizontal lines dividing the CU. A CU may be divided into symmetric TUs or asymmetric TUs. For asymmetric division into TUs, information about the size / shape of the TUs may be signaled or derived from information about the size / shape of the CU.

[0124] Also, coding units are not divided into transform units, and coding units and transform units can have the same size.

[0125] A coding unit can be divided into one or more transform units, and a transform unit can also be divided into one or more transform units.

[0126] For example, when one transform unit is divided into four transform units, the width and height of the four divided transform units can be half the size of the width and height of the transform unit before division. As an example, when a 32x32 transform unit is divided into four transform units, each of the four divided transform units can be 16x16 in size. When one transform unit is divided into four transform units, the transform unit can be said to be divided into a quad-tree.

[0127] For example, when one transform unit is divided into two transform units, the width or height of the two divided transform units can be half the size of the width or height of the transform unit before being divided. As an example, when a 32x32 transform unit is divided vertically into two transform units, each of the two divided transform units can have a size of 16x32. As an example, when a 32x32 transform unit is divided horizontally into two transform units, each of the two divided transform units can have a size of 32x16. When one transform unit is divided into two transform units, the transform unit can be said to be divided into a binary tree.

[0128] When performing the transform, the residual block may be transformed using at least one of a plurality of predefined transform methods. For example, the predefined transform methods may include a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or a KLT. Which transform method is applied to transform the residual block may be determined using at least one of inter prediction mode information, intra prediction mode information, and the size / shape of the transform block of the prediction unit, and in certain cases, information indicating the transform method may be signaled.

[0129] FIG. 6 is a diagram illustrating an embodiment of the intra prediction process.

[0130] The intra prediction mode may be a non-directional mode or a directional mode. The non-directional mode may be a DC mode or a planar mode, and the directional mode may be a prediction mode having a specific direction or angle, the number of which may be one or more, M. The directional mode may be expressed by at least one of a mode number, a mode value, a mode number, and a mode angle.

[0131] The number of intra prediction modes may be one or more, N, including the non-directional and directional modes.

[0132] The number of intra prediction modes may vary depending on the block size, for example, 67 if the block size is 4x4 or 8x8, 35 if the block size is 16x16, 19 if the block size is 32x32, and 7 if the block size is 64x64.

[0133] The number of intra prediction modes may be fixed to N regardless of the block size, for example, at least one of 35 and 67 regardless of the block size.

[0134] The number of intra prediction modes may vary depending on the type of color component, for example, whether the color component is a luma signal or a chroma signal.

[0135] Intra coding and / or decoding may be performed using sample values ​​or coding parameters contained in neighboring reconstructed blocks.

[0136] In order to encode / decode a current block using intra prediction, a step of checking whether samples included in neighboring reconstructed blocks can be used as reference samples for the block to be encoded / decoded may be performed. If there is a sample that cannot be used as a reference sample for the block to be encoded / decoded, at least one of the samples included in the neighboring reconstructed blocks may be used as a reference sample for the block to be encoded / decoded by copying and / or interpolating the sample value for the sample that cannot be used as a reference sample.

[0137] During intra prediction, a filter may be applied to at least one of a reference sample or a predicted sample based on at least one of an intra prediction mode and a size of a block to be coded / decoded. In this case, the block to be coded / decoded may refer to a current block, or at least one of a coding block, a predicted block, and a transform block. The type of filter applied to the reference sample or the predicted sample may vary depending on at least one of the intra prediction mode and the size / shape of the current block. The type of filter may vary depending on at least one of the number of filter taps, filter coefficient values, and filter strength.

[0138] Among the intra prediction modes, the non-directional planar mode can generate sample values ​​within the prediction block as a weighted sum of the upper reference sample of the current sample, the left reference sample of the current sample, the upper right reference sample of the current block, and the lower left reference sample of the current block according to the sample position when generating a prediction block for the target encoding / decoding block.

[0139] Among the intra prediction modes, the non-directional DC mode may generate a predicted block of a target encoding / decoding block as an average value of an upper reference sample of a current block and a left reference sample of the current block, and may perform filtering using reference sample values ​​for one or more upper rows and one or more left columns adjacent to the reference sample in the encoding / decoding block.

[0140] In the case of a plurality of angular modes among the intra prediction modes, a prediction block may be generated using upper right and / or lower left reference samples, and the angular modes may have different directions. Real number-based interpolation may also be performed to generate the prediction sample values.

[0141] To perform the intra prediction method, the intra prediction mode of a current prediction block can be predicted from the intra prediction mode of a prediction block that exists in the neighborhood of the current prediction block. When predicting the intra prediction mode of the current prediction block using mode information predicted from the neighborhood intra prediction mode, if the intra prediction modes of the current prediction block and the neighborhood prediction block are the same, information that the intra prediction modes of the current prediction block and the neighborhood prediction block are the same can be transmitted using predetermined flag information. If the intra prediction modes of the current prediction block and the neighborhood prediction block are different from each other, entropy coding can be performed to encode the intra prediction mode information of the block to be coded / decoded.

[0142] FIG. 7 is a diagram illustrating an embodiment of the inter prediction process.

[0143] The rectangles in Figure 7 may represent images (or pictures). Also, the arrows in Figure 7 may represent prediction directions. That is, images can be coded and / or decoded according to their prediction directions. Each image can be classified into an I-picture (Intra Picture), a P-picture (Uni-predictive Picture), a B-picture (Bi-predictive Picture), etc., depending on its coding type. Each picture can be coded and decoded according to its coding type.

[0144] If the image to be coded is an I picture, the image can be intra-coded without inter-prediction. If the image to be coded is a P picture, the image can be coded via inter-prediction or motion compensation using a reference picture only in the forward direction. If the image to be coded is a B picture, the image can be coded via inter-prediction or motion compensation using a reference picture in both the forward and backward directions, or via inter-prediction or motion compensation using a reference picture in either the forward or backward direction. Here, when an inter-prediction mode is used, the encoder can perform inter-prediction or motion compensation, and the decoder can perform corresponding motion compensation. P and B picture images coded and / or decoded using a reference picture can be considered images using inter-prediction.

[0145] Next, inter prediction according to the embodiment will be described in detail.

[0146] Inter prediction or motion compensation can be performed using reference pictures and motion information, and inter prediction can also use skip mode as described above.

[0147] A reference picture may be at least one of a picture before the current picture or a picture after the current picture. In this case, inter prediction may perform prediction for a block of the current picture based on the reference picture. Here, the reference picture may refer to an image used for predicting a block. In this case, an area within the reference picture may be identified using a reference picture index (refIdx) indicating the reference picture and a motion vector (described later), etc.

[0148] Inter prediction may select a reference picture and a reference block corresponding to a current block in the reference picture, and generate a predicted block for the current block using the selected reference block. The current block may be a block currently being encoded or decoded among blocks in the current picture.

[0149] The motion information can be derived from an inter prediction process by each of the encoding device 100 and the decoding device 200. The derived motion information can be used to perform inter prediction. In this case, the encoding device 100 and the decoding device 200 can improve encoding and / or decoding efficiency by using motion information of a reconstructed neighboring block and / or motion information of a collocated block (col block). The collocated block may be a block corresponding to the spatial position of the current block to be encoded / decoded in an already reconstructed collocated picture (col picture). The reconstructed neighboring block may be a block in the current picture that has already been reconstructed through encoding and / or decoding. The reconstructed block may be a neighboring block adjacent to the current block to be encoded / decoded and / or a block located at an outer corner of the current block to be encoded / decoded. Here, the block located at an outer corner of the current block to be encoded / decoded may be a block vertically adjacent to a neighboring block horizontally adjacent to the current block to be encoded / decoded, or a block horizontally adjacent to a neighboring block vertically adjacent to the current block to be encoded / decoded.

[0150] Each of the encoding device 100 and the decoding device 200 may determine a block located at a spatially corresponding position within the co-located picture to the current block to be coded / decoded, and may determine a predetermined relative position based on the determined block. The predetermined relative position may be inside and / or outside the block located at a spatially corresponding position to the current block to be coded / decoded. Furthermore, each of the encoding device 100 and the decoding device 200 may derive the co-located block based on the determined predetermined relative position. Here, the co-located picture may be any one of at least one reference picture included in a reference picture list.

[0151] The method of deriving motion information may vary depending on the prediction mode of the block to be coded / decoded. For example, prediction modes applicable to inter prediction may include Advanced Motion Vector Prediction (AMVP) and merge mode. Here, the merge mode may be referred to as a motion merge mode.

[0152] For example, when AMVP is applied as a prediction mode, each of the encoding device 100 and the decoding device 200 can generate a motion vector candidate list using the motion vector of the reconstructed neighboring block and / or the motion vector of the co-located block. The reconstructed motion vector of the neighboring block and / or the motion vector of the co-located block can be used as a motion vector candidate. Here, the motion vector of the co-located block can be referred to as a temporal motion vector candidate, and the reconstructed motion vector of the neighboring block can be referred to as a spatial motion vector candidate.

[0153] The bitstream generated by the encoding device 100 may include a motion vector candidate index. That is, the encoding device 100 may generate a bitstream by entropy coding the motion vector candidate index. The motion vector candidate index may indicate an optimal motion vector candidate selected from among the motion vector candidates included in a motion vector candidate list. The motion vector candidate index may be transmitted from the encoding device 100 to the decoding device 200 via the bitstream.

[0154] The decoding device 200 entropy decodes the motion vector candidate index from the bitstream, and uses the entropy decoded motion vector candidate index to select a motion vector candidate for the block to be decoded from the motion vector candidates included in the motion vector candidate list.

[0155] The encoding device 100 can calculate a motion vector difference (MVD) between a motion vector of a current block to be encoded and a motion vector candidate, and can entropy code the MVD. The bitstream can include the entropy-coded MVD. The MVD can be transmitted from the encoding device 100 to the decoding device 200 via the bitstream. In this case, the decoding device 200 can entropy decode the received MVD from the bitstream. The decoding device 200 can derive a motion vector of a current block to be decoded by adding the decoded MVD and the motion vector candidate.

[0156] The bitstream may include a reference image index indicating a reference picture. The reference image index may be entropy coded and transmitted from the coding device 100 to the decoding device 200 via the bitstream. The decoding device 200 may predict a motion vector of a block to be decoded using motion information of neighboring blocks, and may derive a motion vector of the block to be decoded using the predicted motion vector and a motion vector differential. The decoding device 200 may generate a prediction block for the block to be decoded based on the derived motion vector and reference image index information.

[0157] Another example of a motion information derivation method is a merge mode. The merge mode may refer to the merging of motions for multiple blocks. The merge mode may refer to the joint application of motion information of one block to other blocks. When the merge mode is applied, each of the encoding device 100 and the decoding device 200 may generate a merge candidate list using the motion information of reconstructed neighboring blocks and / or the motion information of collocated blocks. The motion information may include at least one of 1) a motion vector, 2) a reference image index, and 3) an inter-prediction indicator. The prediction indicator may be unidirectional (L0 prediction, L1 prediction) or bidirectional.

[0158] In this case, the merge mode may be applied on a CU basis or a PU basis. When the merge mode is performed on a CU basis or a PU basis, the encoding device 100 may entropy encode predefined information to generate a bitstream and then transmit the bitstream to the decoding device 200. The bitstream may include predefined information. The predefined information may include 1) a merge flag indicating whether to perform the merge mode for each block partition, and 2) a merge index indicating which block among neighboring blocks adjacent to the current block to be encoded is to be merged with. For example, the neighboring blocks of the current block to be encoded may include the left neighboring block of the current block to be encoded, the upper neighboring block of the current block to be encoded, and the temporal neighboring block of the current block to be encoded.

[0159] The merge candidate list may indicate a list in which motion information is stored. The merge candidate list may also be generated before the merge mode is performed. The motion information stored in the merge candidate list may be at least one of motion information of neighboring blocks adjacent to the current block to be coded / decoded, motion information of a block collocated with the current block to be coded / decoded in a reference image, new motion information generated by combining motion information already present in the merge candidate list, and a zero merge candidate. Here, the motion information of neighboring blocks adjacent to the current block to be coded / decoded may be referred to as a spatial merge candidate, and the motion information of a block collocated with the current block to be coded / decoded in a reference image may be referred to as a temporal merge candidate.

[0160] The skip mode may be a mode in which motion information of a neighboring block is applied to the current block to be coded / decoded as is. The skip mode may be any of the modes used in inter prediction. When the skip mode is used, the coding apparatus 100 may entropy code information regarding which block's motion information is to be used as the motion information of the current block to be coded, and transmit the entropy code to the decoding apparatus 200 via a bitstream. The coding apparatus 100 may not transmit other information to the decoding apparatus 200. For example, the other information may be syntax element information. The syntax element information may include at least one of motion vector differential information, coded block flags, and transform coefficient levels.

[0161] The residual signal generated after intra- or inter-prediction may be transformed into the frequency domain through a transform process as part of the quantization process. In this case, the primary transform to be performed may be a DCT type 2 (DCT-II) or various DCT or DST kernels. These transform kernels may be separable transforms that perform 1D transforms on the residual signal in the horizontal and / or vertical directions, respectively, or may be 2D non-separable transforms.

[0162] For example, as shown in the table below, the DCT and DST types used for the transformation may adaptively use DCT-V, DCT-VIII, DST-I, and DST-VII in addition to DCT-II during 1D transformation. For example, as shown in the examples of Tables 1 and 2, a transform set may be constructed to derive the DCT or DST type used for the transformation.

[0163] [Table 1]

[0164] [Table 2]

[0165] For example, as shown in FIG. 8, different transform sets may be defined for the horizontal or vertical direction according to the intra prediction mode, and then an encoder / decoder may perform a transform and / or an inverse transform using the intra prediction mode of a block currently being encoded / decoded and the transforms included in the corresponding transform set. In this case, the transform sets may be defined in the encoder / decoder based on the same rule, rather than being entropy coded / decoded. In this case, entropy coding / decoding indicating which of the transforms belonging to the transform set has been used may be performed. For example, if the block size is 64x64 or less, a total of three transform sets may be configured according to the intra prediction mode, as shown in the example of Table 2, and a total of nine multiple transform methods may be combined using three transforms each for the horizontal and vertical directions, and the residual signal may be coded / decoded using the optimal transform method, thereby improving coding efficiency. In this case, truncated unary binarization may be used to entropy code / decode information regarding which of the three transforms belonging to one transform set has been used. In this case, for at least one of the vertical transform and the horizontal transform, information indicating which transform from the transform set has been used can be entropy coded / decoded.

[0166] After the encoder completes the primary transform, it may perform a secondary transform to increase the energy density of the transformed coefficients, as shown in the example of FIG. 9. The secondary transform may also be a separable transform that performs one-dimensional transforms in the horizontal and / or vertical directions, or a two-dimensional non-separable transform. The used transform information may be transmitted or may be implicitly derived in the encoder / decoder based on current and neighboring coding information. For example, a transform set for the secondary transform may be defined, as in the primary transform. The transform set may be defined in the encoder / decoder based on the same rules as the primary transform, rather than being entropy coded / decoded. In this case, information indicating which transforms from the transform set were used may be transmitted and may be applied to at least one of the residual signals obtained by intra- or inter-prediction.

[0167] At least one of the number or type of transform candidates varies depending on the transform set, and at least one of the number or type of transform candidates can be variably determined taking into account at least one of the position, size, division type, prediction mode (intra / inter mode), or directionality / non-directivity of the intra prediction mode of the block (CU, PU, ​​TU, etc.).

[0168] In the decoder, a secondary inverse transform can be performed depending on whether or not to perform a secondary inverse transform, and a primary inverse transform can be performed depending on whether or not to perform a primary inverse transform from the result of the secondary inverse transform.

[0169] The above-mentioned primary and secondary transforms can be applied to at least one signal component of the luminance / chrominance components, or can be applied depending on the size / shape of any coding block. Whether or not a given coding block is used and an index indicating the primary / secondary transform used can be entropy coded / decoded, or can be implicitly derived by the encoder / decoder based on at least one of the current / neighboring coding information.

[0170] The residual signal generated after intra or inter prediction undergoes a quantization process after primary and / or secondary transformation. The quantized transform coefficients undergo an entropy coding process. At this time, the quantized transform coefficients may be scanned diagonally, vertically, and horizontally based on at least one of an intra prediction mode or a minimum block size / shape, as shown in FIG. 10.

[0171] The entropy-decoded, quantized transform coefficients may be inverse scanned and arranged in a block format, and at least one of inverse quantization and inverse transform may be performed on the block. At this time, at least one of diagonal scan, horizontal scan, and vertical scan may be performed as the inverse scanning method.

[0172] For example, if the size of the current coding block is 8x8, the residual signal for the 8x8 block may be entropy coded by scanning quantized transform coefficients for each of four 4x4 sub-blocks according to at least one of the three scanning order methods shown in Fig. 10 after primary transformation, secondary transformation, and quantization. The quantized transform coefficients may also be entropy decoded by inverse scanning. The inverse-scanned and quantized transform coefficients become transform coefficients after inverse quantization, and a reconstructed residual signal may be generated by performing at least one of secondary inverse transform or primary inverse transform.

[0173] In a video encoding process, a block may be split, and an indicator corresponding to the split information may be signaled, as shown in Fig. 11. In this case, the split information may be at least one of a split flag (split_flag), a quad / binary tree flag (QB_flag), a quad tree split flag (quadtree_flag), a binary tree split flag (binarytree_flag), and a binary tree split type flag (Btype_flag). Here, split_flag may be a flag indicating whether the block has been split, QB_flag may be a flag indicating whether the block has been split in a quad tree or binary tree form, quadtree_flag may be a flag indicating whether the block has been split in a quad tree form, binarytree_flag may be a flag indicating whether the block has been split in a binary tree form, and Btype_flag may be a flag indicating whether the block has been split vertically or horizontally if the block has been split in a binary tree form.

[0174] When the split flag is 1, it indicates that the tree has been split, and when the split flag is 0, it indicates that the tree has not been split. In the case of the quad / binary tree flag, when it is 0, it indicates a quad tree split, and when it is 1, it indicates a binary tree split. Conversely, when it is 0, it indicates a binary tree split, and when it is 1, it indicates a quad tree split. In the case of the binary tree split type flag, when it is 0, it indicates a horizontal split, and when it is 1, it indicates a vertical split. Conversely, when it is 0, it indicates a vertical split, and when it is 1, it indicates a horizontal split.

[0175] For example, the partition information for FIG. 11 can be derived by signaling at least one of quadtree_flag, binarytree_flag, and Btype_flag as shown in Table 3 below.

[0176] [Table 3]

[0177] For example, the split flag for FIG. 11 can be derived by signaling at least one of split_flag, QB_flag, and Btype_flag, as shown in Table 4 below.

[0178] [Table 4]

[0179] The division method can be either quadtree division or binary tree division depending on the size / shape of the block. In this case, the split_flag can represent a flag indicating whether quadtree division or binary tree division is used. The size / shape of the block can be derived according to depth information of the block, and the depth information can be signaled.

[0180] If the size of the block falls within a predetermined range, it may be possible to perform quadtree division only. Here, the predetermined range may be defined as at least one of the maximum block size and the minimum block size that can be divided only into quadtrees. Information indicating the maximum / minimum block sizes that are permissible for quadtree division may be signaled via a bitstream, and the information may be signaled in units of at least one of a sequence, a picture parameter, or a slice (segment). Alternatively, the maximum / minimum block sizes may be fixed sizes that are predetermined for an encoder / decoder. For example, if the size of the block falls within the range of 256x256 to 64x64, it may be possible to perform quadtree division only. In this case, the split_flag may be a flag indicating whether or not quadtree division is permitted.

[0181] If the size of the block falls within a predetermined range, it can only be split into a binary tree. Here, the predetermined range may be defined as at least one of the maximum block size and the minimum block size that can only be split into a binary tree. Information indicating the maximum / minimum block size for which binary tree splitting is permitted may be signaled via a bitstream, and the information may be signaled in units of at least one of a sequence, a picture parameter, or a slice (segment). Alternatively, the maximum / minimum block size may be a fixed size that is predetermined for an encoder / decoder. For example, if the size of the block falls within a range of 16x16 to 8x8, it can only be split into a binary tree. In this case, the split_flag may be a flag indicating whether or not binary tree splitting is permitted.

[0182] After one block is divided into a binary tree, if the divided block is further divided, it can only be divided into a binary tree.

[0183] If the width or height of the divided block is such that it cannot be further divided, the one or more indicators may not be signaled.

[0184] In addition to the above-mentioned quadtree-based binary tree division, quadtree-based division is possible after binary tree division.

[0185] Based on the above, the image encoding / decoding method using motion vectors according to the present invention will be described in detail.

[0186] FIG. 12 is a flowchart showing an image coding method according to the present invention, and FIG. 13 is a flowchart showing an image decoding method according to the present invention.

[0187] 12, the encoding device derives motion vector candidates (S1201) and generates a motion vector candidate list based on the derived motion vector candidates (S1202). Once the motion vector candidate list is generated, the generated motion vector candidate list is used to determine a motion vector (S1203), and motion compensation can be performed using the motion vector (S1204). Thereafter, the encoding device can entropy code information related to motion compensation (S1205).

[0188] 13, the decoding device entropy decodes information related to motion compensation received from the encoding device (S1301) and can derive motion vector candidates (S1302). The decoding device then generates a motion vector candidate list based on the derived motion vector candidates (S1303) and can determine a motion vector using the generated motion vector candidate list (S1304). Thereafter, the decoding device can perform motion compensation using the motion vector (S1305).

[0189] Each step shown in FIGS. 12 and 13 will be described in detail below.

[0190] First, the steps of deriving motion vector candidates will be specifically described (S1201, S1302).

[0191] The motion vector candidates for the current block may include at least one of spatial motion vector candidates or temporal motion vector candidates.

[0192] The spatial motion vector of the current block can be derived from a reconstructed block neighboring the current block. For example, the motion vector of a reconstructed block neighboring the current block can be determined as a spatial motion vector candidate for the current block.

[0193] FIG. 14 is a diagram for explaining an example of deriving spatial motion vector candidates for the current block.

[0194] 14, spatial motion vector candidates for a current block can be derived from neighboring blocks adjacent to the current block X. Here, the neighboring blocks adjacent to the current block can include at least one of block B1 adjacent to the upper side of the current block, block A1 adjacent to the left side of the current block, block B0 adjacent to the upper right corner of the current block, block B2 adjacent to the upper left corner of the current block, and block A0 adjacent to the lower left corner of the current block.

[0195] If a motion vector exists in a neighboring block adjacent to the current block, the motion vector of the neighboring block may be determined as a spatial motion vector candidate for the current block. Whether a motion vector of a neighboring block exists or whether the motion vector of the neighboring block can be used as a spatial motion vector candidate for the current block may be determined based on whether a neighboring block exists or whether the neighboring block is coded through inter prediction. In this case, whether a motion vector of a neighboring block exists or whether the motion vector of the neighboring block can be used as a spatial motion vector candidate for the current block may be determined according to a predetermined priority. For example, in the example shown in FIG. 14, the availability of a motion vector may be determined in the order of blocks located at positions A0, A1, B0, B1, and B2.

[0196] When the reference block of the current block and the reference image of the neighboring block having the motion vector are different from each other, the motion vector of the neighboring block may be scaled and determined as a spatial motion vector candidate for the current block. Here, the scaling may be performed based on at least one of the distance between the current image and the reference image referenced by the current block and the distance between the current image and the reference image referenced by the neighboring block. For example, the spatial motion vector candidate for the current block may be derived by scaling the motion vector of the neighboring block by the difference between the distance between the current image and the reference image referenced by the current block and the distance between the current image and the reference image referenced by the neighboring block.

[0197] Even if the reference image list of the current block and the reference image list of the neighboring block are different, it can determine whether to scale the motion vector of the neighboring block depending on whether the reference images of the current block and the neighboring block are the same. Here, the reference image list can include at least one of List0 (L0), List1 (L1), List2 (L2), and List3 (L3).

[0198] In summary, spatial motion vector candidates can be derived taking into consideration at least one of the following: block availability, whether the block is coded in intra prediction mode, whether the reference image list is the same as that of the current block, or whether the reference image is the same as that of the current block. If the neighboring block is available and not coded in intra prediction mode, spatial motion vector candidates for the current block can be generated in the manner illustrated in Table 5 below.

[0199] [Table 5]

[0200] As shown in Table 5, even if the reference image lists of the current block and the neighboring block are different from each other, if the reference images of the current block and the neighboring block are the same, the motion vector of the neighboring block can be determined as the spatial motion vector candidate of the current block.

[0201] In contrast, if the reference images of the current block and the neighboring block are different from each other, the motion vector of the neighboring block can be determined as a spatial motion vector candidate for the current block by scaling the motion vector of the neighboring block, regardless of whether the reference image list of the current block and the reference image list of the neighboring block are the same.

[0202] When deriving spatial motion vector candidates for a current block from neighboring blocks, the order in which spatial motion vector candidates for the current block are derived may be determined taking into consideration whether the reference images of the current block and the neighboring blocks are the same. For example, spatial vector candidates may be derived preferentially from neighboring blocks having the same reference image as the current block, and if the number of derived spatial motion vector candidates (or the number of derived motion vector candidates) is equal to or less than a predetermined maximum value, spatial vector candidates may be derived from neighboring blocks having a reference image different from that of the current block.

[0203] Alternatively, spatial motion vector prediction candidates for the current block may be determined taking into consideration whether the current block and neighboring blocks have the same reference image and the positions of the neighboring blocks.

[0204] For example, spatial motion vector candidates for the current block may be derived from neighboring blocks (A0, A1) adjacent to the left of the current block depending on whether the reference images are identical, and then spatial motion vector candidates for the current block may be derived from neighboring blocks (B0, B1, B2) adjacent to the top of the current block depending on whether the reference images are identical. Table 6 shows an example of the order in which spatial motion vector candidates for the current block may be derived.

[0205] [Table 6]

[0206] The maximum number of spatial motion vector candidates for the current block may be preset so that the same value is used by the encoding device and the decoding device. Alternatively, the encoding device may encode information indicating the maximum number of spatial motion vector candidates for the current block and transmit it to the decoding device via a bitstream. For example, the encoding device may encode 'maxNumSpatialMVPCand' indicating the maximum number of spatial motion vector candidates for the current block and transmit it to the decoding device via a bitstream. In this case, 'maxNumSpatialMVPCand' may be set to a positive integer including 0. For example, 'maxNumSpatialMVPCand' may be set to 2.

[0207] The candidate temporal motion vectors for the current block can be derived from reconstructed blocks included in a co-located picture of the current image, where the co-located picture may be an image that has been coded / decoded before the current image and has a different temporal order from the current image.

[0208] FIG. 15 is a diagram for explaining an example of deriving temporal motion vector candidates for the current block.

[0209] 15, in a collocated picture of the current image, a temporal motion vector candidate for the current block can be derived from a block including an external position of a block corresponding to the spatially same position as the current block X, or a block including an internal position of a block corresponding to the spatially same position as the current block X. As an example, a temporal motion vector candidate for the current block X can be derived from block H adjacent to the lower left corner of block C corresponding to the spatially same position as the current block, or block C3 including the center point of block C. Block H or block C3 used to derive a temporal motion vector candidate for the current block can be referred to as a "collocated block."

[0210] If a temporal motion vector candidate for the current block can be derived from block H, which includes an external position of block C, block H can be set as the corresponding block of the current block. In this case, the temporal motion vector of the current block can be derived based on the motion vector of block H. On the other hand, if a temporal motion vector candidate for the current block cannot be derived from block H, block C3, which includes an internal position of block C, can be set as the corresponding block of the current block. In this case, the temporal motion vector of the current block can be derived based on the motion vector of block C3. If a temporal motion vector for the current block cannot be derived from blocks H and C3 (e.g., if blocks H and C3 are all intra-coded), a temporal motion vector candidate for the current block may not be derived, or may be derived from a block located different from blocks H and C3.

[0211] Alternatively, the candidate temporal motion vectors for the current block may be derived from multiple blocks in the corresponding position image. For example, multiple candidate temporal motion vectors for the current block may be derived from block H and block C3.

[0212] In Figure 15, it is illustrated that a temporal motion vector candidate for the current block can be derived from a block adjacent to the lower left corner of the corresponding position block or a block including the center point of the corresponding position block. However, the position of the block for deriving a temporal motion vector candidate for the current block is not limited to the example shown in Figure 15. For example, a temporal prediction candidate for the current block can be derived from a block adjacent to the upper / lower boundary, left / right boundary, or one corner of the corresponding position block, or can be derived from a block including a specific position within the corresponding position block (e.g., a block adjacent to a corner boundary of the corresponding position block).

[0213] The temporal motion vector candidates for the current block may be determined by taking into consideration the reference image lists (or prediction directions) of blocks located inside or outside the current block and the corresponding position block.

[0214] For example, when the reference image list available for the current block is L0 (i.e., when the inter prediction indicator indicates PRED_L0), a motion vector of a block located inside or outside the corresponding position block that uses L0 as a reference image can be derived as a temporal motion vector candidate for the current block. That is, when the reference image list available for the current block is LX (where X is an integer indicating an index of the reference image list, such as 0, 1, 2, or 3), a motion vector of a block located inside or outside the corresponding position block that uses LX as a reference image can be derived as a temporal motion vector candidate for the current block.

[0215] Even when the current block uses multiple reference image lists, a temporal motion vector candidate for the current block can be determined by considering whether the reference image lists of the current block and the blocks located inside or outside the corresponding position block are the same.

[0216] For example, if the current block performs bidirectional prediction (i.e., the inter prediction indicator is PRED_BI), a motion vector of a block located inside or outside the corresponding position block that uses L0 and L1 as reference images may be derived as a temporal motion vector candidate for the current block. If the current block performs tri-directional prediction (i.e., the inter prediction indicator is PRED_TRI), a motion vector of a block located inside or outside the corresponding position block that uses L0, L1, and L2 as reference images may be derived as a temporal motion vector candidate for the current block. If the current block performs four-directional prediction (i.e., the inter prediction indicator is PRED_QUAD), a motion vector of a block located inside or outside the corresponding position block that uses L0, L1, L2, and L3 as reference images may be derived as a temporal motion vector candidate for the current block.

[0217] Alternatively, if the current block is set to perform multi-directional prediction using any reference image, the temporal motion prediction vector candidate for the current block can be determined by considering whether the external block has the same reference picture list and the same prediction direction as the current block.

[0218] As an example, if the current block performs bidirectional prediction with respect to the L0 reference picture list (i.e., if the inter prediction indicator for list L0 is PRED_BI), the motion vector of a block located inside or outside the corresponding position block that performs bidirectional prediction with respect to L0 using L0 as a reference image can be derived as a candidate temporal motion vector for the current block.

[0219] Candidate temporal motion vectors may also be derived based on at least one of the coding parameters.

[0220] The temporal motion vector candidate can be preliminarily derived when the number of derived spatial motion vector candidates is smaller than the maximum number of motion vector candidates, thereby omitting the process of deriving the temporal motion vector candidate when the number of derived spatial motion vector candidates reaches the maximum number of motion vector candidates.

[0221] As an example, if the maximum number of motion vector candidates is two and the two derived spatial motion vector candidates have different values, the process of deriving a temporal motion vector candidate can be omitted.

[0222] As another example, the temporal motion vector candidates for the current block may be derived based on the maximum number of temporal motion vector candidates. Here, the maximum number of temporal motion vector candidates may be preset so that the same value is used in the encoding device and the decoding device. Alternatively, information indicating the maximum number of temporal motion vector candidates for the current block may be coded via a bitstream and transmitted to the decoding device. As an example, the encoding device may code 'maxNumTemporalMVPCand' indicating the maximum number of temporal motion vector candidates for the current block and transmit it to the decoding device via a bitstream. In this case, 'maxNumTemporalMVPCand' may be set as a positive integer including 0. For example, 'maxNumTemporalMVPCand' may be set to 1.

[0223] If the distance between the current image containing the current block and the reference image of the current block is different from the distance between the corresponding position picture containing the corresponding position block and the reference image of the corresponding position block, the candidate temporal motion vector for the current block can be obtained by scaling the motion vector of the corresponding position block.

[0224] FIG. 16 shows an example of scaling the motion vector of the corresponding position block to derive the temporal motion vector candidate for the current block.

[0225] The motion vector of the corresponding position vector can be scaled based on at least one of the difference value (td) between the POC (Picture order count) indicating the display order of the corresponding position image and the POC of the reference image of the corresponding position block, and the difference value (tb) between the POC of the current image and the POC of the reference image of the current block.

[0226] Prior to scaling, td or tb can be adjusted so that it is within a predetermined range. For example, if the predetermined range is -128 to 127, and td or tb is smaller than -128, td or tb can be adjusted to -128. If td or tb is larger than 127, td or tb can be adjusted to 127. If td or tb is within the range of -128 to 127, td or tb is not adjusted.

[0227] The scaling factor DistScaleFactor can be calculated based on td or tb. In this case, the scaling factor can be calculated based on the following Equation 1.

[0228]

number

[0229] In the formula, Abs() represents the absolute value function, and the output value of the function is the absolute value of the input value.

[0230] The value of the scaling factor DistScaleFactor calculated based on Equation 1 can be adjusted to a predetermined range. For example, DistScaleFactor can be adjusted to be within the range of -1024 to 1023.

[0231] A candidate temporal motion vector for the current block can be determined by scaling the motion vector of the corresponding block using the scaling factor. For example, the candidate temporal motion vector for the current block can be determined by the following Equation 2.

[0232]

number

[0233] In the formula, Sign() is a function that outputs the sign information of the value included in (). For example, if Sign(-1), it outputs -. In Equation 2, mvCol represents the motion vector of the corresponding position block, i.e., the temporal motion vector predictor before scaling.

[0234] Next, the steps of generating a motion vector candidate list based on the derived motion vector candidates will be described (S1202, S1303).

[0235] Generating the motion vector candidate list may include adding or removing motion vector candidates from the motion vector candidate list, and adding combined motion vector candidates to the motion vector candidate list.

[0236] Considering the step of adding or removing derived motion vector candidates to the motion vector candidate list, the encoding device and decoding device can add the derived motion vector candidates to the motion vector candidate list in the order in which the motion vector candidates are derived.

[0237] The generated motion vector candidate list may be determined according to the inter-prediction direction of the current block. For example, one motion vector candidate list may be generated for each reference image list, or one motion vector candidate list may be generated for each reference image. Multiple reference image lists or multiple reference images may share one motion vector candidate list.

[0238] In the embodiments described below, it is assumed that the motion vector candidate list mvpListLX refers to the motion vector candidate list corresponding to the reference image lists L0, L1, L2, and L3. For example, the motion vector candidate list corresponding to the reference image list L0 can be called mvpListL0.

[0239] The number of motion vector candidates included in the motion vector candidate list may be set so that the encoding device and the decoding device use the same preset value, or the maximum number of motion vector candidates included in the motion vector candidate list may be coded by the encoding device and transmitted to the decoding device via a bitstream.

[0240] For example, maxNumMVPCandList, the maximum number of motion vector candidates that the motion vector candidate list mvpListLX can include, may be a positive integer including 0. For example, maxNumMVPCandList may be an integer such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16. When maxNumMVPCandList is 2, it means that a maximum of two motion vector candidates can be included in mvpListLX. Thus, the index value of the motion vector candidate added first to mvpListLX may be set to 0, and the index value of the motion vector candidate added next may be set to 1. The maximum number of motion vector candidates may be defined for each motion vector candidate list, or may be commonly defined for all motion vector candidate lists. For example, the maximum motion vector candidates in mvpListL0 and mvpListL1 may have different values ​​or the same value.

[0241] FIG. 17 is a diagram showing an example of how a motion vector candidate list is generated.

[0242] Assume that a spatial motion vector candidate (1,0) that is not spatially scaled is derived from the block at position A1 shown in (b) of Figure 17, and a scaled temporal motion vector candidate (2,3) is derived from the block at position H shown in (a) of Figure 17. In this case, as in the example shown in (c) of Figure 17, the spatial motion vector candidate derived from the block at position A1 and the temporal motion vector candidate derived from the block at position H may be sequentially included in the motion vector candidate list.

[0243] The derived motion vector candidates may be included in the motion vector candidate list in a predetermined order. For example, after a spatial motion vector candidate is added to the motion vector candidate list, if the number of motion vector candidates included in the motion vector candidate list is less than the maximum number of motion vector candidates, a temporal motion vector candidate may be added to the motion vector candidate list. Conversely, a temporal motion vector candidate may be added to the motion vector candidate list with priority over a spatial motion vector candidate. In this case, a spatial motion vector candidate may be selectively added to the motion vector candidate list depending on whether it is identical to a temporal motion vector candidate.

[0244] Furthermore, the encoding device and the decoding device can assign an index to identify each motion vector candidate in the order in which they were added to the motion vector candidate list. (c) of Figure 17 shows that the index value of the motion vector candidate derived from the block at position A1 is set to 0, and the index value of the motion vector candidate derived from the block at position H is set to 1.

[0245] A motion vector having a predetermined value other than the spatial motion vector candidate and the temporal motion vector candidate may be added to the motion vector candidate list. For example, if the number of motion vector candidates included in the motion vector list is less than the maximum number of motion vector candidates, a motion vector with a value of 0 may be added to the motion vector candidate list.

[0246] FIG. 18 is a diagram showing an example of adding a motion vector having a predetermined value to the motion vector candidate list.

[0247] 18, "numMVPCandLX" indicates the number of motion vector candidates included in the motion vector candidate list mvpListLX. As an example, numMVPCandL0 may indicate the number of motion vector candidates included in the motion vector candidate list mvpListL0.

[0248] Also, maxNumMVPCand indicates the maximum number of motion vector candidates that can be included in the motion vector candidate list mvpListLX. numMVPCandLX and maxNumMVPCand can have integer values ​​including 0.

[0249] If numMVPCandLX is smaller than maxNumMVPCand, a motion vector having a predetermined value may be added to the motion vector candidate list, and the numMVPCandLX value may be increased by 1. In this case, the motion vector added to the motion vector candidate list has a fixed value and may be added to the end of the motion vector candidate list. For example, the motion vector having a predetermined value added to the motion vector candidate list may be a zero motion vector candidate having a value of (0,0).

[0250] As an example, as shown in (a) of Figure 18, if numMVPCandLX is 1 and maxNumMVPCand is 2, the numMVPCandLX value can be increased by 1 while adding one zero motion vector candidate with a (0,0) value to the motion vector candidate list.

[0251] If the difference between maxNumMVPCand and numMVPCandLX is 2 or more, a motion vector having a predetermined value can be included in the motion vector candidate list repeatedly by the difference value.

[0252] For example, if maxNumMVPCand is 2 and numMVPCandLX is 0, motion vectors having a predetermined value can be repeatedly added to the motion vector candidate list until numMVPCandLX becomes equal to maxNumMVPCand. In (b) of Figure 18, two zero motion vector candidates having a value of (0,0) are shown as having been added to the motion vector candidate list.

[0253] As another example, a motion vector of a predetermined value may be included in the motion vector candidate list only if the motion vector candidate list does not include any motion vector candidates that are the same as the motion vector of the predetermined value.

[0254] As an example, if numMVPCandLX is smaller than maxNumMVPCand and the motion vector candidate list does not include a motion vector with (0,0), a motion vector with (0,0) can be added to the motion vector candidate list and numMVPCandLX can be increased by 1, as shown in the example of (c) of Figure 18.

[0255] 18 illustrates an example in which the value of the predetermined motion vector added to the motion vector candidate list is (0,0), but the value of the predetermined motion vector added to the motion vector candidate list is not limited to the illustrated example. Also, when adding multiple predetermined motion vector candidates as in the example shown in (b) of FIG. 18, the multiple predetermined motion vectors added to the motion vector candidate list may have different values.

[0256] The encoding device and the decoding device can also adjust the size of the motion vector candidate list by removing a motion vector candidate included in the motion vector candidate list.

[0257] For example, the encoding device and the decoding device may check whether the same motion vector candidate exists in the motion vector candidate list. If the same motion vector candidate exists in the motion vector candidate list, the remaining motion vector candidates, except for the motion vector candidate with the smallest motion vector candidate index, may be removed from the motion vector candidate list.

[0258] The determination of identity between motion vector candidates may be applied only between spatial motion vectors or only between temporal motion vectors, or may be applied between spatial and temporal motion vectors.

[0259] When the number of motion vector candidates included in the motion vector candidate list is greater than the maximum number of motion vector candidates that the motion vector candidate list can contain, motion vector candidates equal to the difference between the number of motion vector candidates included in the motion vector candidate list and the maximum number of motion vector candidates can be removed from the motion vector candidate list.

[0260] FIG. 19 is a diagram showing an example in which a motion vector candidate is removed from the motion vector candidate list.

[0261] If numMVPCandLX is equal to or greater than maxNumMVPCand, motion vector candidates with index values ​​greater than maxNumMVPCand-1 can be removed from the motion vector candidate list.

[0262] As an example, in the example shown in Figure 19, when numMVPCandLX is 3 and maxNumMVPCand is 2, the motion vector candidate (4,-3) assigned an index 2, which is greater than maxNumMVPCand-1, is shown to be removed from the motion vector candidate list.

[0263] Next, the step of adding the combined motion vector candidate to the motion vector candidate list will be described.

[0264] If the number of motion vector candidates included in the motion vector candidate list is less than the maximum number of motion vector candidates, a combined motion vector may be added to the motion vector candidate list using at least one of the motion vector candidates included in the motion vector candidate list. For example, a combined motion vector candidate may be generated using at least one of the spatial motion vector candidate, the temporal motion vector candidate, and the zero motion vector candidate included in the motion vector candidate list, and the generated combined motion vector candidate may be included in the motion vector candidate list.

[0265] Alternatively, a combined motion vector candidate may be generated using a motion vector candidate that is not included in the motion vector candidate list. As an example, a combined motion vector candidate may be generated using a motion vector candidate derived from a block that is not included in the motion vector candidate list but can be used to derive at least one of the spatial motion vector candidate or the temporal motion vector candidate, or a motion vector candidate having a predetermined value that is not included in the motion vector candidate list (e.g., a zero motion vector).

[0266] Alternatively, the combined motion vector candidate may be generated based on at least one of the coding parameters, or the combined motion vector candidate may be added to the motion vector candidate list based on at least one of the coding parameters.

[0267] The maximum number of motion vector candidates that the motion vector candidate list can include may be increased by the number of combined motion vectors or by a smaller number after at least one of a spatial motion vector candidate, a temporal motion vector candidate, or a motion vector candidate of a predetermined value is added. For example, maxNumMVPCandList may have a first value for a spatial motion vector candidate or a temporal motion vector candidate, and may be increased to a second value greater than the first value to insert additional combined motion vector candidates after the spatial motion vector candidate or the temporal motion vector candidate is added.

[0268] FIG. 20 is a diagram showing an example of a motion vector candidate list.

[0269] The current block may be motion-compensated using motion vector candidates included in the motion vector candidate list. Motion compensation of the current block may be performed using one motion vector for one reference image list, or may be performed using multiple motion vectors for one reference image list. For example, if the inter-prediction direction of the current block is bidirectional, motion compensation of the current block may be performed by deriving one motion vector for each of reference picture lists L0 and L1, or by deriving two motion vectors for reference picture list L0.

[0270] The motion vector candidate list may include at least one of spatial motion vector candidates, temporal motion vector candidates, zero motion vector candidates, and combined motion vector candidates generated by combining two or more of them, each of which may be identified by a motion vector candidate index.

[0271] A motion vector candidate set including multiple motion vector candidates may be identified by a single motion vector candidate index based on the inter-prediction direction of the current block. Here, the motion vector candidate set may include N motion vector candidates, depending on the number N of inter-prediction directions of the current block. For example, the motion vector candidate set may include multiple motion vector candidates, such as a first motion vector candidate, a second motion vector candidate, a third motion vector candidate, and a fourth motion vector candidate.

[0272] A motion vector candidate set may be generated by combining at least two of a spatial motion vector candidate, a temporal motion vector candidate, and a zero motion vector candidate. As an example, in Figure 20, a motion vector candidate set including two motion vector candidates is illustrated as being assigned to motion vector candidate indexes 4 to 13. Also, each motion vector candidate set is illustrated as being generated by combining a spatial motion vector candidate (mxLXA, mxLXB), a temporal motion vector (mxLXCol), and a zero motion vector (mvZero).

[0273] One or more motion vectors can be derived from the reference image list LX based on the prediction direction for the reference image list LX. As an example, when unidirectional prediction is performed on the reference image list LX, a motion vector for the current block can be derived using one of the motion vector candidates assigned to motion vector indexes 0 to 3. On the other hand, when bidirectional prediction is performed on the reference image list LX, a motion vector for the current block can be derived using a set of motion vector candidates assigned to motion vector indexes 4 to 13. That is, in the encoding / decoding process, at least one motion vector can be derived based on the motion vector candidates included in the motion vector candidate list.

[0274] The motion vector of the current block can be derived by adding a motion vector differential value to a motion vector candidate. For example, in the example shown in Figure 20, when a motion vector candidate is selected from motion vector candidate indexes 0 to 3, the motion vector is derived by adding a motion vector differential value (MVD) to the selected motion vector candidate.

[0275] When a motion vector candidate set including multiple motion vector candidates is selected, multiple motion vectors for the current block can be derived based on the multiple motion vector candidates included in the motion vector candidate set. In this case, motion vector differential values ​​for each of the multiple motion vector candidates included in the motion vector candidate set can be coded / decoded. In this case, multiple motion vectors can be derived for the current block by adding up the motion vector differential values ​​corresponding to each motion vector candidate.

[0276] As another example, a motion vector differential value for some of the motion vector candidates included in the motion vector candidate set may be encoded / decoded. For example, one motion vector differential value may be encoded / decoded for a motion vector candidate set including multiple motion vector candidates. In this case, the current block may use a motion vector derived by adding a motion vector differential value to one of the motion vector candidates included in the motion vector candidate set, or a motion vector derived directly from the motion vector candidate. In the example shown in FIG. 20, for a motion vector candidate set including two motion vector candidates, the first motion vector or the second motion vector is derived by adding a motion vector differential value to one of the motion vector candidates, and the other is the same as the motion vector candidate.

[0277] As another example, multiple motion vector candidates included in a motion vector candidate set may share the same motion vector differential value.

[0278] The inter-prediction indicator may indicate unidirectional prediction or multi-directional prediction for a given reference image list. For example, the inter-prediction indicator may be expressed as PRED_LX indicating unidirectional prediction for reference image list LX, PRED_BI_LX indicating bidirectional prediction for reference image list LX, etc. Here, X may be an integer including 0 and may indicate an index of the reference image list, such as 0, 1, 2, 3, etc.

[0279] For example, when unidirectional prediction is performed on reference image list L0, the inter prediction indicator can be set to PRED_L0, and when unidirectional prediction is performed on reference image list L1, the inter prediction indicator can be set to PRED_L1.

[0280] On the other hand, when bidirectional prediction is performed on reference image list L1, the inter prediction indicator may be set to PRED_BI_L1. When the inter prediction indicator for reference image list L1 is PRED_BI_L1, the current block may be subjected to inter prediction by deriving two motion vectors using the motion vector candidate list and deriving two predictive blocks from reference images included in reference image list L1. In this case, each of the two predictive blocks may be derived from two different reference images included in reference image list L1, or may be derived from a single reference image included in reference image list L1.

[0281] The inter prediction indicator may be coded / decoded to indicate the total number of prediction directions for the current block, or may be coded / decoded to indicate the number of prediction directions for each of the reference image lists.

[0282] For example, an inter-prediction indicator (PRED_L0) indicating unidirectional prediction for reference image list L0 and an inter-prediction indicator (PRED_BI_L1) indicating bidirectional prediction for reference image list L1 may be coded for the current block. Alternatively, if unidirectional prediction is performed for reference image list L0 and bidirectional prediction is performed for reference image list L1, the inter-prediction indicator for the current block may indicate PRED_TRI.

[0283] The example shown in Figure 20 shows a motion vector candidate list mvpListLX for a specific reference image list LX. When multiple reference image lists, such as L0, L1, L2, and L3, exist, a motion vector candidate list can be generated for each reference image list. This allows for the generation of a minimum of 1 to a maximum of N prediction blocks to be used for inter prediction or motion compensation of the current block. Here, N is an integer equal to or greater than 1 and may represent 2, 3, 4, 5, 6, 7, 8, etc.

[0284] At least one of the motion vector candidates included in the motion vector candidate list can be determined as a predicted motion vector (or motion vector predictor) for the current block. The determined predicted motion vector can be used to calculate the motion vector of the current block, and the motion vector can be used for inter prediction or motion compensation of the current block.

[0285] When a motion vector candidate set including a plurality of motion vector candidates is selected for the current block, the plurality of motion vector candidates included in the motion vector candidate set and a motion vector of the current block calculated based on the plurality of motion vector candidates may be stored as information about motion compensation for the current block, and the stored information about motion compensation for the current block may be used later when generating a motion vector candidate list or performing motion compensation for neighboring blocks.

[0286] 20 illustrates an example in which a motion vector candidate list is generated for each reference image list. A motion vector candidate list may be generated for each reference image. For example, when bidirectional prediction is performed on a reference image list LX, a first motion vector candidate list may be generated for a first reference image used for bidirectional prediction among reference pictures included in the reference image list LX, and a second motion vector candidate list may be generated for a second reference image used for bidirectional prediction.

[0287] Next, consider the steps of determining a predicted motion vector from the motion vector candidate list (S1203, S1304).

[0288] Among the motion vector candidates included in the motion vector candidate list, the motion vector candidate indicated by the motion vector candidate index can be determined as the predicted motion vector for the current block.

[0289] FIG. 21 is a diagram showing an example of deriving a predicted motion vector candidate for a current block from a motion vector candidate list.

[0290] 21 shows that the maximum number of motion vector candidates that a motion vector candidate list can contain, maxNumMVPCand, is 2, and the number of motion vector candidates included in the motion vector candidate list is also 2. In this case, if the motion vector candidate index points to index 1, the second motion vector candidate included in the motion vector candidate list (i.e., the motion vector candidate assigned index 1), (2,3), can be determined as the predicted motion vector of the current block.

[0291] The encoding device can calculate the difference between a motion vector and a predicted motion vector to calculate a motion vector differential value, and the decoding device can calculate a motion vector by combining the predicted motion vector and the motion vector differential.

[0292] Although not shown, when a motion vector candidate index indicates a motion vector candidate set, multiple motion vectors can be derived from multiple motion vector candidates included in the motion vector candidate set. In this case, the motion vector of the current block may be a motion vector candidate plus a motion vector differential, or may have the same value as the motion vector candidate.

[0293] Next, the steps of performing motion compensation using the motion vectors will be described (S1204, S1305).

[0294] The encoding device and the decoding device can calculate a motion vector using the predicted motion vector and the motion vector difference value. Once the motion vector is calculated, inter prediction or motion compensation can be performed using the calculated motion vector. Alternatively, as shown in the example of FIG. 20, the motion vector prediction value can be determined as the motion vector itself.

[0295] The current block can have a minimum of 1 to a maximum of N motion vectors based on the prediction direction. Using the motion vectors, a minimum of 1 to a maximum of N prediction blocks can be generated to derive the final prediction block of the current block.

[0296] For example, if the current block has one motion vector, a predicted block generated using the motion vector may be determined as the final predicted block of the current block.

[0297] On the other hand, when the current block has multiple motion vectors, multiple prediction blocks may be generated using the multiple motion vectors, and the final prediction block of the current block may be determined based on a weighted sum of the multiple prediction blocks. Reference images including each of the multiple prediction blocks indicated by the multiple motion vectors may be included in different reference image lists or may be included in the same reference image list.

[0298] The weight applied to each prediction block may be 1 / N (where N is the number of generated prediction blocks) and may have an equal value. For example, if two prediction blocks are generated, the weight applied to each prediction block may be 1 / 2, if three prediction blocks are generated, the weight applied to each prediction block may be 1 / 3, and if four prediction blocks are generated, the weight applied to each prediction block may be 1 / 4. Alternatively, a different weight may be assigned to each prediction block to determine the final prediction block of the current block.

[0299] The weight does not have to have a fixed value for each prediction block, but may have a variable value for each prediction block. In this case, the weights applied to each prediction block may be the same or different from each other. To apply variable weights, one or more weight information for the current block may be signaled via a bitstream. The weight information may be signaled for each prediction block or for each reference image. It is also possible for multiple prediction blocks to share one weight information.

[0300] The following Equations 3 to 5 respectively show an example of generating a final predicted block of the current block when the inter prediction indicator of the current block is PRED_BI, PRED_TRI, and PRED_QUAD and the prediction direction for each reference image list is unidirectional.

[0301]

number

[0302]

number

[0303]

number

[0304] In Equations 3 to 5, P_BI, P_TRI, and P_QUAD indicate the final predicted block of the current block, and LX (X=0, 1, 2, 3) may indicate a reference image list. WF_LX indicates a weight value of a predicted block generated using LX, and OFFSET_LX may indicate an offset value for a predicted block generated using LX. P_LX indicates a predicted block generated using a motion vector for LX of the current block. RF indicates a rounding factor and may be set to 0, a positive number, or a negative number.

[0305] Even when there are multiple prediction directions for a given reference image list, a final predicted block for a current block can be obtained based on a weighted sum of predicted blocks. In this case, weights applied to predicted blocks derived from the same reference image list may have the same value or different values.

[0306] At least one of the weights (WF_LX) and offsets (OFFSET_LX) for the plurality of prediction blocks may be entropy coded / decoded coding parameters. As another example, the weights and offsets may be derived from coded / decoded neighboring blocks near the current block. Here, the neighboring blocks near the current block may include at least one of blocks used to derive spatial motion vector candidates for the current block or blocks used to derive temporal motion vector candidates for the current block.

[0307] As another example, the weight and offset may be determined based on the display order (POC) of the current image and each reference image. In this case, the greater the distance between the current image and the reference image, the smaller the weight or offset may be set, and the closer the distance between the current image and the reference image, the larger the weight or offset may be set. For example, if the difference in POC between the current image and the L0 reference image is 2, the weight value applied to the predicted block generated with reference to the L0 reference image may be set to 1 / 3, whereas if the difference in POC between the current image and the L0 reference image is 1, the weight value applied to the predicted block generated with reference to the L0 reference image may be set to 2 / 3. As illustrated above, the weight or offset value may be inversely proportional to the difference in display order between the current image and the reference image. As another example, the weight or offset value may be proportional to the difference in display order between the current image and the reference image.

[0308] As another example, at least one of the weights or offsets may be entropy coded / decoded based on at least one of the coding parameters, and a weighted sum of the prediction block may be calculated based on at least one of the coding parameters.

[0309] Next, the process of entropy encoding / decoding information related to motion compensation will be described in detail (S1205, S1301).

[0310] FIG. 22 (FIGS. 22A and 22B) is a diagram illustrating the syntax for information for motion compensation.

[0311] The encoding device may entropy encode information related to motion compensation via a bitstream, and the decoding device may entropy decode the information related to motion compensation included in the bitstream. Here, the information related to motion compensation to be entropy encoded / decoded may include at least one of an inter prediction indicator (inter_pred_idc), reference image indexes (ref_idx_l0, ref_idx_l1, ref_idx_l2, ref_idx_l3), motion vector candidate indexes (mvp_l0_idx, mvp_l1_idx, mvp_l2_idx, mvp_l3_idx), motion vector differentials, weight values ​​(wf_l0, wf_l1, wf_l2, wf_l3), and offset values ​​(offset_l0, offset_l1, offset_l2, offset_l3).

[0312] The inter prediction indicator may indicate the inter prediction direction of the current block when the current block is encoded / decoded using inter prediction. For example, the inter prediction indicator may indicate unidirectional prediction or multi-directional prediction such as bidirectional prediction, tridirectional prediction, or quaternary prediction. The inter prediction indicator may indicate the number of reference images used when the current block generates a predicted block. Alternatively, one reference image may be used for multi-directional prediction. In this case, N (N>M) directional predictions can be performed using M reference images. The inter prediction indicator may also indicate the number of predicted blocks used when performing inter prediction or motion compensation on the current block.

[0313] In this way, the number of reference images used when generating a predicted block of the current block, the number of predicted blocks used when performing inter prediction or motion compensation of the current block, or the number of reference image lists available to the current block can be determined according to the inter prediction indicator. Here, the number N of reference image lists is a positive integer and may have a value of 1, 2, 3, 4, or more. For example, the reference image lists may include L0, L1, L2, and L3. The current block may perform motion compensation using one or more reference image lists.

[0314] For example, the current block may be motion compensated by generating at least one predictive block using at least one reference image list. For example, motion compensation may be performed by generating one or more predictive blocks using reference image list L0, or by generating one or more predictive blocks using reference image lists L0 and L1. Alternatively, motion compensation may be performed by generating one or more predictive blocks or up to N predictive blocks (where N is a positive integer of 3 or greater) using reference image lists L0, L1, and L2, or by using one or more predictive blocks or up to N predictive blocks (where N is a positive integer of 4 or greater) using reference image lists L0, L1, L2, and L3.

[0315] The reference image indicator may indicate one-way (PRED_LX), two-way (PRED_BI), three-way (PRED_TRI), four-way (PRED_QUAD), or more directions depending on the number of prediction directions of the current block.

[0316] As an example, assuming that unidirectional prediction is performed for each reference image list, the inter prediction indicator PRED_LX may indicate that one prediction block is generated using the reference image list LX (X is an integer such as 0, 1, 2, or 3) and inter prediction or motion compensation is performed using the generated one prediction block. The inter prediction indicator PRED_BI may indicate that two prediction blocks are generated using the L0 and L1 reference image lists and inter prediction or motion compensation is performed using the generated two prediction blocks. The inter prediction indicator PRED_TRI may indicate that three prediction blocks are generated using the L0, L1, and L2 reference image lists and inter prediction or motion compensation is performed using the generated three prediction blocks. The inter prediction indicator PRED_QUAD may indicate that four prediction blocks are generated using the L0, L1, L2, and L3 reference image lists and inter prediction or motion compensation is performed using the generated four prediction blocks. That is, the total number of prediction blocks used to perform inter prediction of the current block may be set as the inter prediction indicator.

[0317] When multi-directional prediction is performed on the reference image list, the inter-prediction indicator PRED_BI may indicate that bidirectional prediction is performed on the L0 reference image list, and the inter-prediction indicator PRED_TRI may indicate that three-directional prediction is performed on the L0 reference image list, that one-directional prediction is performed on the L0 reference image list and bidirectional prediction is performed on the L1 reference image list, or that bidirectional prediction is performed on the L0 reference image list and one-directional prediction is performed on the L1 reference image list.

[0318] In this way, the inter prediction indicator can mean generating a minimum of 1 to a maximum of N (where N is the number of prediction directions indicated by the inter prediction indicator) predictive blocks from at least one reference image list and performing motion compensation, or it can mean generating a minimum of 1 to a maximum of N predictive blocks from N reference images and performing motion compensation for the current block using the generated predictive blocks.

[0319] For example, the inter prediction indicator PRED_TRI may mean that three predictive blocks are generated using at least one of the L0, L1, L2, and L3 reference image lists to perform inter prediction or motion compensation on the current block, or that three predictive blocks are generated using at least three of the L0, L1, L2, and L3 reference image lists to perform inter prediction or motion compensation on the current block. Also, PRED_QUAD may mean that four predictive blocks are generated using at least one of the L0, L1, L2, and L3 reference image lists to perform inter prediction or motion compensation on the current block, or that four predictive blocks are generated using at least four of the L0, L1, L2, and L3 reference image lists to perform inter prediction or motion compensation on the current block.

[0320] Available inter-prediction directions may be determined according to the inter-prediction indicator, and all or some of the available inter-prediction directions may be selectively used based on the size and / or shape of the current block.

[0321] The number of reference images included in each reference image list may be predefined or may be entropy coded in the coding device and transmitted to the decoding device. As an example, the syntax element "num_ref_idx_lX_active_minus1" (where X indicates the index of the reference image list, such as 0, 1, 2, 3, etc.) may indicate the number of reference images for a reference image list, such as L0, L1, L2, or L3.

[0322] The reference image index can identify the reference image from each reference image list that the current block refers to. For each reference image list, one or more reference image indexes can be entropy coded / decoded. The current block can perform motion compensation using one or more reference image indexes.

[0323] When N reference images are selected using N reference image indexes, at least one to N (or more than N) prediction blocks can be generated to perform motion compensation on the current block.

[0324] The motion vector candidate index indicates a motion vector candidate for the current block from a motion vector candidate list generated for each reference image list or each reference image index. At least one motion vector candidate index can be entropy coded / decoded for each motion vector candidate list. The current block can be motion compensated using at least one motion vector candidate index.

[0325] For example, based on N motion vector candidate indexes, at least one to N (or more than N) prediction blocks may be generated to perform motion compensation on the current block.

[0326] A motion vector differential indicates a difference value between a motion vector and a predicted motion vector. One or more motion vector differentials can be entropy coded / decoded for a reference image list for a current block or a motion vector candidate list generated for each reference image index. The current block can be motion compensated using one or more motion vector differentials.

[0327] For example, N motion vector differentials may be used to generate a minimum of 1 to a maximum of N (or more than N) prediction blocks, and motion compensation may be performed on the current block.

[0328] When two or more prediction blocks are generated during motion compensation for a current block, a final prediction block for the current block can be generated through a weighted sum of each prediction block. During the weighted sum operation, at least one of a weight and an offset can be applied to each prediction block. A weighted sum factor, such as a weight or an offset, used in the weighted sum operation can be entropy coded / decoded for at least one of a reference image list, a reference image, a motion vector candidate index, a motion vector differential, or a motion vector.

[0329] The weighted sum factor may be derived by index information specifying one of a set predefined in the encoding device and the decoding device, in which case the index information for specifying at least one of the weight and the offset may be entropy coded / decoded.

[0330] Information related to the weighted sum factors may be entropy coded / decoded on a block-by-block basis or at a higher level. For example, weights or offsets may be entropy coded / decoded on a block-by-block basis, such as a CTU, CU, or PU, or at a higher level, such as a video parameter set, a sequence parameter set, a picture parameter set, an adaptation parameter set, or a slice header.

[0331] The weighted sum factor may be entropy coded / decoded based on a weighted sum factor difference value indicating a difference between the weighted sum factor and the weighted sum factor predicted value. For example, the weight predicted value and the weight difference value may be entropy coded / decoded, or the offset predicted value and the offset difference value may be entropy coded / decoded. Here, the weight difference value may indicate a difference between the weight and the weight predicted value, and the offset difference value may indicate a difference between the offset and the offset predicted value.

[0332] In this case, the weighted sum factor differential value may be entropy coded / decoded on a block-by-block basis, and the weighted sum factor predicted value may be entropy coded / decoded at a higher level. When a weighted sum factor predicted value, such as a weighted predicted value or an offset predicted value, is entropy coded / decoded on a picture or slice-by-picture or slice basis, blocks included in the picture or slice may use a common weighted sum factor predicted value.

[0333] The weighted sum factor predicted value may also be derived via a specific region within an image, slice, or tile, or a specific region within a CTU or CU. As an example, the weight value or offset value of a specific region within an image, slice, tile, CTU, or CU may be used as the weight predicted value or offset predicted value. In this case, the entropy coding / decoding of the weighted sum factor predicted value may be omitted, and only the weighted sum factor difference value may be entropy coded / decoded.

[0334] Alternatively, the weighted sum factor prediction value may be derived from a neighboring block that is coded / decoded near the current block. For example, a weight value or an offset value of a neighboring block that is coded / decoded near the current block may be set as a weighted prediction value or an offset prediction value of the current block. Here, the neighboring block of the current block may include at least one of a block used to derive a spatial motion vector candidate and a block used to derive a temporal motion vector candidate.

[0335] When a weight prediction value and a weight difference value are used, the decoding device can calculate a weight value for a prediction block by combining the weight prediction value and the weight difference value. When an offset prediction value and an offset difference value are used, the decoding device can calculate an offset value for a prediction block by combining the offset prediction value and the offset difference value.

[0336] Instead of entropy encoding / decoding information about the weighted sum factor of the current block, the weighted sum factor of a block coded / decoded in the vicinity of the current block can be used as the weighted sum factor of the current block. For example, the weight or offset of the current block can be set to the same value as the weight or offset of a neighboring block coded / decoded in the vicinity of the current block.

[0337] At least one of the pieces of information relating to the motion compensation may be entropy coded / decoded via a bitstream using coding parameters, or at least one of the pieces of information relating to the motion compensation may be derived using at least one coding parameter.

[0338] When entropy encoding / decoding information related to motion compensation, binarization methods such as truncated Rice binarization, K-th order Exp_Golomb binarization, K-th order Exp_Golomb binarization, fixed-length binarization, unary binarization, or truncated unary binarization can be used.

[0339] When entropy encoding / decoding information about motion compensation, a context model can be determined using at least one of information about motion compensation of neighboring blocks near the current block, information about previously encoded / decoded motion compensation, information about the depth of the current block, and information about the size of the current block.

[0340] In addition, when entropy encoding / decoding information regarding motion compensation, entropy encoding / decoding can be performed using at least one of information regarding motion compensation of neighboring blocks, information regarding previously encoded / decoded motion compensation, information regarding the depth of the current block, and information regarding the size of the current block as a predicted value for information regarding motion compensation of the current block.

[0341] The inter encoding / decoding process can be performed for each of the luminance and chrominance signals. For example, in the inter encoding / decoding process, at least one of the methods of obtaining an inter prediction indicator, generating a motion vector candidate list, deriving a motion vector, and performing motion compensation can be applied differently to the luminance signal and the chrominance signal.

[0342] The inter encoding / decoding process for the luminance and chrominance signals can be performed in the same manner. For example, in the inter encoding / decoding process applied to the luminance signal, at least one of an inter prediction indicator, a motion vector candidate list, a motion vector candidate, a motion vector, and a reference image can be similarly applied to the chrominance signal.

[0343] These methods can be performed in the same way in the encoder and decoder. For example, in the inter encoding / decoding process, at least one of the methods of deriving a motion vector candidate list, deriving a motion vector candidate, deriving a motion vector, and motion compensation can be applied in the same way in the encoder and decoder. Also, the application order of these methods can be different between the encoder and the decoder.

[0344] The above-described embodiments of the present invention can be applied depending on the size of at least one of a coding block, a prediction block, a block, and a unit. The size here may be defined as a minimum size and / or a maximum size for applying these embodiments, or as a fixed size for applying the embodiments. Furthermore, these embodiments can be applied to a first size, and the second size, respectively. That is, these embodiments can be applied in combination depending on the size. Furthermore, the above-described embodiments of the present invention can be applied only when the block size is equal to or greater than the minimum size and equal to or less than the maximum size. That is, these embodiments can be applied only when the block size falls within a certain range.

[0345] For example, the above embodiment is applicable only when the size of the block to be coded / decoded is 8x8 or larger. For example, the above embodiment is applicable only when the size of the block to be coded / decoded is 16x16 or larger. For example, the above embodiment is applicable only when the size of the block to be coded / decoded is 32x32 or larger. For example, the above embodiment is applicable only when the size of the block to be coded / decoded is 64x64 or larger. For example, the above embodiment is applicable only when the size of the block to be coded / decoded is 128x128 or larger. For example, the above embodiment is applicable only when the size of the block to be coded / decoded is 4x4. For example, the above embodiment is applicable only when the size of the block to be coded / decoded is 8x8 or smaller. For example, the above embodiment is applicable only when the size of the block to be coded / decoded is 16x16 or smaller. For example, the above embodiment is applicable only when the size of the block to be coded / decoded is 8x8 or larger but not larger than 16x16. For example, the above embodiment is applicable only when the size of the block to be coded / decoded is 16x16 or more and 64x64 or less.

[0346] The above-described embodiments of the present invention can be applied according to a temporal layer. A separate identifier is signaled to identify a temporal layer to which the embodiments can be applied, and the embodiments can be applied to the temporal layer identified by the corresponding identifier. The identifier here may be defined as the minimum and / or maximum layer to which the embodiments can be applied, or may be defined as indicating a specific layer to which the embodiments can be applied.

[0347] For example, the above embodiment is applicable only when the temporal layer of the current image is the lowest layer. For example, the above embodiment is applicable only when the temporal layer identifier of the current image is 0. For example, the above embodiment is applicable only when the temporal layer identifier of the current image is 1 or greater. For example, the above embodiment is applicable only when the temporal layer of the current image is the highest layer.

[0348] As in the above-described embodiment of the present invention, the reference picture set used in the reference picture list construction and reference picture list modification processes can use at least one of the reference picture lists L0, L1, L2, and L3.

[0349] According to an embodiment of the present invention, when a deblocking filter calculates boundary strength, it can use one or more, up to a maximum of N, motion vectors of a block to be coded / decoded, where N is a positive integer equal to or greater than 1, such as 2, 3, 4, etc.

[0350] The above-described embodiments of the present invention can also be applied when the motion vector used in motion vector prediction has at least one of the following units: 16-pel, 8-pel, 4-pel, integer-pel, 1 / 2-pel, 1 / 4-pel, 1 / 8-pel, 1 / 16-pel, 1 / 32-pel, and 1 / 64-pel. Furthermore, the motion vector used in motion vector prediction can be selectively used in the above pixel units.

[0351] The slice types to which the above-described embodiments of the present invention are applied are defined, and the embodiments of the present invention can be applied according to the slice types.

[0352] For example, when the slice type is T (Tri-predictive)-slice, a prediction block is generated using at least three motion vectors, and a weighted sum of the at least three prediction blocks is calculated and used as the final prediction block of the block to be coded / decoded. For example, when the slice type is Q (Quad-predictive)-slice, a prediction block is generated using at least four motion vectors, and a weighted sum of the at least four prediction blocks is calculated and used as the final prediction block of the block to be coded / decoded.

[0353] The above-described embodiments of the present invention can be applied not only to inter prediction and motion compensation methods using motion vector prediction, but also to inter prediction and motion compensation methods using skip mode, merge mode, and the like.

[0354] The shape of the blocks to which the above-described embodiments of the present invention are applied may be square or non-square.

[0355] In the above-described embodiments, the methods are described based on flowcharts with a series of steps or units, but the present invention is not limited to the order of these steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included, or one or more steps in the flowcharts may be deleted without affecting the scope of the present invention.

[0356] The above-described embodiments include examples of various aspects. It is not possible to describe all possible combinations for illustrating the various aspects, but a person skilled in the art will recognize that other combinations are possible. Therefore, it can be said that the present invention includes all various alterations, modifications, and variations that fall within the scope of the following claims.

[0357] The above-described embodiments of the present invention may be embodied in the form of program instructions executable by various computer components and stored on a computer-readable storage medium. The computer-readable storage medium may include, alone or in combination, program instructions, data files, data structures, and the like. The program instructions stored on the computer-readable storage medium may be specially designed and constructed for the present invention, or may be well known and available to those skilled in the art of computer software. Examples of computer-readable storage media include magnetic media such as hard disks, floppy disks, and magnetic tape; optical media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specially configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include not only machine language code produced by a compiler, but also high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices may be configured to operate as one or more software modules to perform the processes of the present invention, or vice versa.

[0358] Although the present invention has been described above using specific details such as specific components, limited embodiments, and drawings, these are provided merely to facilitate a more comprehensive understanding of the present invention, and the present invention is not limited to the above embodiments. Those skilled in the art will be able to make various modifications and variations from these descriptions.

[0359] Therefore, the concept of the present invention should not be limited to the above-described embodiments, but should be considered to fall within the scope of the concept of the present invention, including not only the scope of the claims described below, but also all modifications equivalent to or equivalent to the scope of these claims. [Industrial Applicability]

[0360] The present invention can be used in devices for encoding / decoding images.

Claims

1. deriving a plurality of motion vectors for a current block according to an inter prediction direction of the current block; determining a plurality of prediction blocks for the current block using the plurality of motion vectors; obtaining a final predicted block of a current block based on a weighted sum of the plurality of predicted blocks; obtaining a residual block of the current block by inverse transform; reconstructing the current block based on the final predicted block and the residual block; Including, the inverse transformation is performed using one of a predefined set of transformations; a weight of the current block for the weighted sum is derived from weight index information of the current block that specifies one of weights included in a predefined weight set; The weight index information of the current block is derived based on weight index information of neighboring blocks of the current block. Image decoding method.

2. one of the predefined transform sets is determined based on transform type index information explicitly decoded from the bitstream. The image decoding method according to claim 1 .

3. determining a plurality of prediction blocks for a current block; determining a final predicted block of the current block based on a weighted sum of the plurality of predicted blocks; obtaining a residual block of the current block based on the final predicted block; performing a transform block on the residual block to obtain transform coefficients; Including, the transformation is performed using one of a predefined set of transformations; a weight of the current block for the weighted sum is derived from weight index information of the current block that specifies one of weights included in a predefined weight set; The weight index information of the current block is derived based on weight index information of neighboring blocks of the current block. Image encoding method.

4. transform type index information specifying one of the predefined transform sets is explicitly coded into the bitstream; The image encoding method according to claim 3 .

5. A method for transmitting a bitstream generated by the image coding method according to claim 3 or 4 and stored on a computer-readable recording medium.

Citation Information

Patent Citations

  • Motion picture coding / decoding method and apparatus

    JP2004007377A

  • Method for encoding moving image and method for decoding moving image

    JP2004007379A

  • Electronic apparatus and decoding method

    JP2013251752A

  • Image decoding apparatus, image decoding method and image encoding apparatus

    WO2013047805A1

  • Systems and methods for generalized multi-hypothesis prediction for video coding

    WO2017197146A1