Image encoding / decoding method and apparatus
By generating a reconfigured list of merging candidates and utilizing the distortion values of neighboring and reference blocks, the efficiency of inter-frame prediction is improved, the transmission and storage cost issues of high-resolution and high-quality image data are resolved, and the compression efficiency of image encoding/decoding methods is improved.
Patent Information
- Application Number
- CN202310382489.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-29
- Filing Date
- 2018-09-28
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2038-09-28
AI Technical Summary
The increased cost of transmitting and storing high-resolution and high-quality image data necessitates efficient image encoding/decoding technologies to improve compression efficiency.
By generating a reconfigured list of merge candidates, and using the distortion values of the current block's neighboring blocks and reference blocks, the list is rearranged to improve the efficiency of inter-frame prediction.
This improves the compression efficiency of image encoding/decoding methods and reduces transmission and storage costs.
Smart Images

Figure CN116489387B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of September 28, 2018, the application number of "201880063815.0", and the title of "Image encoding / decoding method and apparatus and recording medium for storing bitstream". TECHNICAL FIELD
[0002] The present invention relates to an image encoding / decoding method and apparatus. More particularly, the present invention relates to an image encoding / decoding method and apparatus using a reconfigured merge candidate list when performing inter prediction and a recording medium storing a bitstream generated by the image encoding method and apparatus of the present invention. BACKGROUND
[0003] Recently, in various application fields, the demand for high resolution and high quality images such as high definition (HD) images and ultra-high definition (UHD) images has increased. However, higher resolution and higher quality image data has an increased amount of data compared to conventional image data. Therefore, when image data is transmitted by using a medium such as a conventional wired and wireless broadband network, or when image data is stored by using a conventional storage medium, the cost of transmission and storage increases. In order to solve these problems that occur as the resolution and quality of image data increase, an efficient image encoding / decoding technique is required for higher resolution and higher quality images.
[0004] Image compression techniques include various techniques including an inter prediction technique that predicts pixel values included in a current picture from previous or subsequent pictures of the current picture, an intra prediction technique that predicts pixel values included in a current picture by using pixel information in the current picture, a transform and quantization technique for compressing the energy of a residual signal, an entropy encoding technique that allocates a short code to a value having a high frequency of occurrence and a long code to a value having a low frequency of occurrence, and the like. Image data can be efficiently compressed by using such image compression techniques, and can be transmitted or stored. SUMMARY
[0005] TECHNICAL PROBLEM
[0006] The present invention aims to provide an image encoding / decoding method and apparatus having improved compression efficiency and a recording medium storing a bitstream generated by the image encoding / decoding method and apparatus of the present invention.
[0007] In addition, another object of the present invention is to provide an image encoding / decoding method and apparatus using inter prediction having improved compression efficiency and a recording medium storing a bitstream generated by the image encoding / decoding method and apparatus of the present invention.
[0008] Furthermore, another object of the present application is to provide an image encoding / decoding method and apparatus that efficiently performs inter prediction by using a reconfigured merge candidate list and a recording medium storing a bitstream generated by the image encoding / decoding method and apparatus of the present application.
[0009] Technical Solution
[0010] A method of decoding an image of the present application can include deriving at least one merge candidate for a current block, generating an initial merge candidate list for the current block by using the derived at least one merge candidate, and generating a reconfigured merge candidate list by using the initial merge candidate list.
[0011] In the method of decoding an image of the present application, the step of generating a reconfigured merge candidate list can include calculating a distortion value for the at least one merge candidate by using a neighboring block of the current block and the at least one merge candidate, and reconfiguring the initial merge candidate list based on the distortion value.
[0012] In the method of decoding an image of the present application, the initial merge candidate list includes at least one of a spatial merge candidate, a temporal merge candidate, a subblock-based temporal merge candidate, a subblock-based spatio-temporal combination merge candidate, a combination merge candidate, and a zero merge candidate for the current block.
[0013] In the method of decoding an image of the present application, the distortion value is calculated by using at least one of a sum of absolute difference (SAD), a sum of absolute transformed difference (SATD), and a mean removed sum of absolute difference (MR-SAD) between a neighboring block of a reference block of the current block and a neighboring block of a reference block of the current block.
[0014] In the method of decoding an image of the present application, the distortion value is calculated based on at least one of a neighboring block of a reference block in an L0 direction indicated by L0 direction motion information of the at least one merge candidate and a neighboring block of a reference block in an L1 direction indicated by L1 direction motion information of the at least one merge candidate.
[0015] In the method of decoding an image of the present application, the distortion value is calculated based on a neighboring block of a reference block indicated by a motion vector obtained by applying a preset offset to a motion vector of the at least one merge candidate.
[0016] In the method of decoding an image of the present application, when the at least one merge candidate includes both L0 direction motion information and L1 direction motion information, the distortion value is calculated by a distortion value between a neighboring block of a reference block in the L0 direction and a neighboring block of a reference block in the L1 direction.
[0017] In the method of decoding an image of the present invention, when the at least one merge candidate includes any one of L0 direction motion information and L1 direction motion information, the distortion value is calculated from a distortion value between a neighboring block of a reference block indicated by motion information derived by applying mirroring to the any one of motion information and a neighboring block of the reference block according to the any one of motion information.
[0018] In the method of decoding an image of the present invention, when the at least one merge candidate includes both L0 direction motion information and L1 direction motion information, the distortion value is calculated from a distortion value between a neighboring block of a reference block indicated by motion information derived by applying mirroring to the L0 direction motion information and a neighboring block of the reference block indicated by motion information derived by applying mirroring to the L1 direction motion information.
[0019] In the method of decoding an image of the present invention, wherein the reconfigured merge candidate list is generated by rearranging distortion values of one or more merge candidates included in the initial merge candidate list in size.
[0020] In the method of decoding an image of the present invention, the merge candidate that becomes a target to be rearranged is determined according to an order of one or more merge candidates included in the initial merge candidate list.
[0021] In the method of decoding an image of the present invention, wherein the number of the merge candidate that becomes a target to be rearranged is a predefined value.
[0022] In the method of decoding an image of the present invention, further comprising entropy decoding motion estimation information of a current block, wherein the motion estimation information of the current block includes information indicating whether to reconfigure the initial merge candidate list.
[0023] In the method of decoding an image of the present invention, whether to reconfigure the initial merge candidate list is determined based on at least one of a size and a split shape of a current block.
[0024] In the method of encoding an image of the present invention, the method can include deriving at least one merge candidate of a current block, generating an initial merge candidate list of the current block by using the derived at least one merge candidate, and generating a reconfigured merge candidate list by using the initial merge candidate list.
[0025] In a non-transitory storage medium of the present invention, a bitstream is included, wherein the bitstream is generated by an image encoding method, the image encoding method can include deriving at least one merge candidate for a current block, generating an initial merge candidate list for the current block by using the derived at least one merge candidate, and generating a reconfigured merge candidate list by using the initial merge candidate list.
[0026] Advantageous Effects
[0027] According to the present invention, there are provided an image encoding / decoding method and apparatus having improved compression efficiency and a recording medium storing a bitstream generated by the image encoding / decoding method and apparatus of the present invention.
[0028] Further, according to the present invention, there are provided an image encoding / decoding method and apparatus using inter prediction having improved compression efficiency and a recording medium storing a bitstream generated by the image encoding / decoding method and apparatus of the present invention.
[0029] Further, according to the present invention, there are provided an image encoding / decoding method and apparatus effectively performing inter prediction by using a reconfigured merge candidate list and a recording medium storing a bitstream generated by the image encoding / decoding method and apparatus of the present invention. BRIEF DESCRIPTION OF DRAWINGS
[0030] Figure 1 is a block diagram illustrating a configuration of an encoding apparatus to which the present invention is applied according to an embodiment.
[0031] Figure 2 is a block diagram illustrating a configuration of a decoding apparatus according to an embodiment and to which the present invention is applied.
[0032] Figure 3 is a diagram schematically illustrating a partition structure of an image when the image is encoded and decoded.
[0033] Figure 4 is a diagram illustrating an embodiment of an inter picture prediction process.
[0034] Figure 5 is a diagram illustrating a flowchart of an image decoding method according to an embodiment of the present invention.
[0035] Figure 6 is a diagram illustrating a flowchart of an image decoding method according to an embodiment of the present invention.
[0036] Figure 7 is a diagram illustrating a method of deriving a spatial merge candidate.
[0037] Figure 8 is a diagram illustrating a method of deriving a temporal merge candidate.
[0038] Figure 9 is a diagram illustrating a method of deriving a sub-block based spatio-temporal combination merge candidate.
[0039] Figure 10 is a diagram illustrating a method of determining a merge candidate list according to an embodiment of the present application.
[0040] Figure 11 is a diagram illustrating a method of determining representative motion information according to an embodiment of the present application.
[0041] Figure 12 is a diagram illustrating a method of calculating a distortion value according to an embodiment of the present application.
[0042] Figure 13 is a diagram illustrating a method of calculating a distortion value according to another embodiment of the present application.
[0043] Figure 14 is a diagram illustrating a method of calculating a distortion value according to another embodiment of the present application.
[0044] Figure 15 is a diagram illustrating a flowchart of an image encoding method according to an embodiment of the present application. DETAILED DESCRIPTION
[0045] Various modifications can be made to the present application, and there are various embodiments of the present application, and examples of various embodiments of the present application will now be provided and described in detail with reference to the accompanying drawings. However, the present application is not limited thereto, and exemplary embodiments can be interpreted to include all modifications, equivalents, or substitutions within the technical concept and technical scope of the present application. Like reference numerals refer to the same or similar functions throughout. In the drawings, the shapes and sizes of elements can be exaggerated for clarity. In the following detailed description of the present application, reference is made to the accompanying drawings that illustrate a specific embodiment by which the present application can be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure. It is to be understood that the various embodiments of the present disclosure, although different, are not necessarily mutually exclusive. For example, specific features, structures, and characteristics described herein in connection with one embodiment can be implemented in other embodiments without departing from the spirit and scope of the disclosure. In addition, it should be understood that the position or arrangement of individual elements within each disclosed embodiment can be modified without departing from the spirit and scope of the disclosure. The following detailed description is, therefore, not to be taken in a limiting sense, and the scope of the present disclosure is defined only by the appended claims, along with the full range of equivalents to which such claims are entitled, in view of a proper interpretation in light of the specification.
[0046] The terms "first", "second", etc. used in the specification can be used to describe various components, but the components are not construed as being limited by the terms. The terms are used only to distinguish one component from another component. For example, a "first" component can be called a "second" component without departing from the scope of the present application, and a "second" component can also be similarly called a "first" component. The term "and / or" includes a combination of a plurality of items or any one of the plurality of items.
[0047] It will be understood that, in the specification, when an element is referred to as being "connected to" or "coupled to" another element, it can be "directly connected to" or "directly coupled to" the other element or connected to or coupled to the other element with other elements in between. On the contrary, it should be understood that when an element is referred to as being "directly coupled to" or "directly connected to" another element, there are no other elements interposed therebetween.
[0048] In addition, constituent components shown in the embodiments of the present application are independently shown in order to present different characteristic functions from each other. Therefore, this does not mean that each of the constituent components is composed of a separate hardware or software constituent unit. In other words, for convenience, each of the constituent components includes each of the enumerated constituent components. Therefore, at least two of the constituent components in each of the constituent components can be combined to form one constituent component, or one constituent component can be divided into a plurality of constituent components for performing each function. Embodiments in which each of the constituent components is combined and embodiments in which one constituent component is divided are also included in the scope of the present application without departing from the essence of the present application.
[0049] The terms used in the specification are used only to describe specific embodiments and are not intended to limit the present application. Expressions used in the singular include the plural, unless they have obviously a different meaning in the context. In the specification, it will be understood that the terms such as "include", "has", "comprise", and "have" are intended to indicate the existence of the features, numbers, steps, actions, elements, parts, or combinations thereof disclosed in the specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, actions, elements, parts, or combinations thereof can exist or can be added. In other words, when a certain element is referred to as "included", elements other than the corresponding element are not excluded, but additional elements can be included in the embodiments of the present application or can be within the scope of the present application.
[0050] Furthermore, some of the constituent elements can not be indispensable constituent elements to perform the essential functions of the present application, but optional constituent elements to improve the performance thereof. The present application can be implemented by including only the indispensable constituent elements to implement the essence of the present application, excluding the constituent elements used when improving the performance. A structure including only the indispensable constituent elements, excluding the optional constituent elements used when improving the performance, is also included in the scope of the present application.
[0051] Hereinafter, embodiments of the present application will be described in detail with reference to the accompanying drawings. In describing the exemplary embodiments of the present application, a well-known function or construction will not be described in detail in order not to unnecessarily obscure the understanding of the present application. The same or similar constituent elements are designated by the same reference numerals in the accompanying drawings, and repetitive description on the same elements will be omitted.
[0052] Hereinafter, an image can refer to a picture constituting a video, or can refer to a video itself. For example, "encoding or decoding or both encoding and decoding an image" can refer to "encoding or decoding or both encoding and decoding a moving picture", and can refer to "encoding or decoding or both encoding and decoding one of the pictures of a moving picture."
[0053] Hereinafter, the terms "moving picture" and "video" can be used as the same meaning and can be replaced with each other.
[0054] Hereinafter, a target image can be an encoding target image as an encoding target and / or a decoding target image as a decoding target. Furthermore, the target image can be an input image input to an encoding apparatus, and an input image input to a decoding apparatus. Here, the target image can have the same meaning as a current image.
[0055] Hereinafter, the terms "image", "picture", "frame", and "screen" can be used as the same meaning and can be replaced with each other.
[0056] Hereinafter, a target block can be an encoding target block as an encoding target and / or a decoding target block as a decoding target. Furthermore, the target block can be a current block as a target of current encoding and / or decoding. For example, the terms "target block" and "current block" can be used as the same meaning and can be replaced with each other.
[0057] Hereinafter, the terms "block" and "unit" can be used as the same meaning and can be replaced with each other. Or "block" can mean a specific unit.
[0058] Hereinafter, the terms "region" and "segment" can be replaced with each other.
[0059] Hereinafter, a certain signal can be a signal representing a certain block. For example, an original signal can be a signal representing a target block. A prediction signal can be a signal representing a prediction block. A residual signal can be a signal representing a residual block.
[0060] In an embodiment, each of certain information, data, flag, index, element, and attribute, etc. can have a value. A value of information, data, flag, index, element, and attribute equal to "0" can represent a logical false or a first predefined value. In other words, the value "0", false, logical false, and the first predefined value can be replaced with each other. A value of information, data, flag, index, element, and attribute equal to "1" can represent a logical true or a second predefined value. In other words, the value "1", true, logical true, and the second predefined value can be replaced with each other.
[0061] When a variable i or j is used to represent a column, a row, or an index, a value of i can be an integer equal to or greater than 0, or an integer equal to or greater than 1. That is, a column, a row, an index, etc. can be counted from 0, or can be counted from 1.
[0062] Term Description
[0063] Encoder: denotes a device performing encoding. That is, denotes an encoding device.
[0064] Decoder: denotes a device performing decoding. That is, denotes a decoding device.
[0065] Block: is an array of samples of MxN. Here, M and N can denote a positive integer, and the block can denote an array of samples in a two-dimensional form. The block can refer to a unit. A current block can denote an encoding target block which becomes a target at the time of encoding, or a decoding target block which becomes a target at the time of decoding. Further, the current block can be at least one of an encoding block, a prediction block, a residual block, and a transform block.
[0066] Sample: is a basic unit constituting a block. According to a bit depth (Bd), a sample can be represented as a value from 0 to 2 Bd -1. In the present invention, a sample can be used as a meaning of a pixel. That is, a sample, pel, pixel can have the same meaning as each other.
[0067] Unit: Can refer to a coding and decoding unit. When an image is coded and decoded, the unit can be a region generated by partitioning a single image. Also, when a single image is partitioned into sub-partitioned units during coding or decoding, the unit can mean a sub-partitioned unit. That is, an image can be partitioned into a plurality of units. When an image is coded and decoded, predetermined processing for each unit can be performed. A single unit can be partitioned into sub-units having sizes smaller than that of the unit. Depending on a function, the unit can mean a block, a macroblock, a coding tree unit, a coding tree block, a coding unit, a coding block, a prediction unit, a prediction block, a residual unit, a residual block, a transform unit, a transform block, etc. Also, to distinguish the unit from a block, the unit can include a luma component block, a chroma component block associated with the luma component block, and syntax elements of each color component block. The unit can have various sizes and shapes, and specifically, the shape of the unit can be a two-dimensional geometric figure such as a square, a rectangle, a trapezoid, a triangle, a pentagon, etc. Also, unit information can include at least one of a unit type indicating a coding unit, a prediction unit, a transform unit, etc., and a unit size, a unit depth, an order of coding and decoding of the unit, etc.
[0068] Coding tree unit: A single coding tree block configured with a luma component Y and two coding tree blocks related to chroma components Cb and Cr. Also, the coding tree unit can mean to include a block and syntax elements of each block. Each coding tree unit can be partitioned by using at least one of a quad-tree partitioning method, a binary-tree partitioning method, and a ternary-tree partitioning method to configure lower-level units such as coding units, prediction units, transform units, etc. The coding tree unit can be used as a term for designating a block of samples that becomes a processing unit when an image as an input image is coded / decoded. Here, the quad-tree can mean a quad-ary tree.
[0069] Coding tree block: Can be used as a term for designating any one of a Y coding tree block, a Cb coding tree block, and a Cr coding tree block.
[0070] Neighbor block: Can mean a block neighboring a current block. The block neighboring the current block can mean a block contacting a boundary of the current block, or a block located within a predetermined distance from the current block. The neighbor block can mean a block neighboring a vertex of the current block. Here, the block neighboring the vertex of the current block can mean a block vertically neighboring a neighbor block horizontally neighboring the current block, or a block horizontally neighboring a neighbor block vertically neighboring the current block.
[0071] Reconstructed neighboring block: can denote a neighboring block that is adjacent to the current block and has been encoded or decoded spatially / temporally. Here, the reconstructed neighboring block can denote a reconstructed neighboring unit. The reconstructed spatial neighboring block can be a block within the current picture and has been reconstructed by encoding or decoding or both. The reconstructed temporal neighboring block is a block or a neighboring block of the block at a position corresponding to the current block of the current picture within the reference picture.
[0072] Unit depth: can denote a partitioning degree of a unit. In a tree structure, the highest node (root node) can correspond to a first unit that is not partitioned. Also, the highest node can have a minimum depth value. In this case, the depth of the highest node can be level 0. A node with a depth of level 1 can denote a unit generated by partitioning the first unit once. A node with a depth of level 2 can denote a unit generated by partitioning the first unit twice. A node with a depth of level n can denote a unit generated by partitioning the first unit n times. A leaf node can be a lowest node and is a node that cannot be further partitioned. The depth of the leaf node can be a maximum level. For example, a predefined value of the maximum level can be 3. The depth of the root node can be the lowest, and the depth of the leaf node can be the deepest. Also, when a unit is denoted as a tree structure, a level in which the unit exists can denote a unit depth.
[0073] Bitstream: can denote a bitstream including encoded image information.
[0074] Parameter set: corresponds to header information among configurations within a bitstream. At least one of a video parameter set, a sequence parameter set, a picture parameter set, and an adaptation parameter set can be included in the parameter set. Also, the parameter set can include slice header and tile header information.
[0075] Parsing: can denote determining a value of a syntax element by performing entropy decoding, or can denote entropy decoding itself.
[0076] Symbol: can denote at least one of a syntax element, an encoding parameter, and a transform coefficient value of an encoding / decoding target unit. Also, the symbol can denote an entropy encoding target or an entropy decoding result.
[0077] Prediction mode: can be information indicating a mode encoded / decoded using intra prediction or a mode encoded / decoded using inter prediction.
[0078] Prediction unit: can mean a basic unit when performing prediction such as inter prediction, intra prediction, inter compensation, intra compensation, and motion compensation. A single prediction unit can be partitioned into multiple partitions having smaller sizes, or can be partitioned into multiple lower-level prediction units. The multiple partitions can be basic units when performing prediction or compensation. The partitions generated by partitioning the prediction unit can also be prediction units.
[0079] Prediction unit partition: can mean a shape obtained by partitioning a prediction unit.
[0080] Reference picture list: can mean a list including one or more reference pictures used for inter prediction or motion compensation. LC (List Combination), L0 (List 0), L1 (List 1), L2 (List 2), L3 (List 3), etc. are types of reference picture lists. One or more reference picture lists can be used for inter prediction.
[0081] Inter prediction indicator: can mean a direction of inter prediction (uni-prediction, bi-prediction, etc.) of a current block. Alternatively, the inter prediction indicator can mean the number of reference pictures used to generate a prediction block of the current block. Further alternatively, the inter prediction indicator can mean the number of prediction blocks used to perform inter prediction or motion compensation for the current block.
[0082] Prediction list utilization flag: can mean whether to use at least one reference picture included in a specific reference picture list to generate a prediction block. The prediction list utilization flag can be used to derive an inter prediction indicator, and conversely, the inter prediction indicator can be used to derive the prediction list utilization flag. For example, when the prediction list utilization flag indicates a first value "0", it means that a reference picture included in the corresponding reference picture list is not used to generate a prediction block. When the prediction list utilization flag indicates a second value "1", it means that a reference picture included in the corresponding reference picture list is used to generate a prediction block.
[0083] Reference picture index: can mean an index indicating a specific reference picture in a reference picture list.
[0084] Reference picture: can mean a reference picture referred to by a specific block for inter prediction or motion compensation. Alternatively, the reference picture can be a picture including a reference block referred to by a current block for inter prediction or motion compensation. Hereinafter, the terms "reference picture" and "reference image" can be used as the same meaning and used interchangeably.
[0085] Motion vector: is a two-dimensional vector used for inter prediction or motion compensation, and can represent an offset between a reference picture and a coding / decoding target picture. For example, (mvX, mvY) can represent a motion vector, mvX can represent a horizontal component, and mvY can represent a vertical component.
[0086] Search range: can be a two-dimensional region in which a search for a motion vector is performed during inter prediction. For example, the size of the search range can be MxN. M and N are integers, respectively.
[0087] Motion vector candidate: can represent a block or a motion vector of the block that becomes a prediction candidate when a motion vector is predicted. The motion vector candidate can be listed in a motion vector candidate list.
[0088] Motion vector candidate list: can represent a list configured using one or more motion vector candidates.
[0089] Motion vector candidate index: represents an indicator indicating a motion vector candidate in a motion vector candidate list. The motion vector candidate index is also referred to as an index of a motion vector predictor.
[0090] Motion information: can represent information including a motion vector, a reference picture index, an inter prediction indicator; and at least any one of reference picture list information, a reference picture, a motion vector candidate, a motion vector candidate index, a merge candidate, and a merge index.
[0091] Merge candidate list: can represent a list consisting of merge candidates.
[0092] Merge candidate: can represent a spatial merge candidate, a temporal merge candidate, a combined merge candidate, a combined bi-predictive merge candidate, a zero merge candidate, etc. The merge candidate can have an inter prediction indicator, a reference picture index for each list, and motion information such as a motion vector.
[0093] Merge index: can represent an indicator indicating a merge candidate within a merge candidate list. The merge index can indicate a block used to derive the merge candidate among reconstructed blocks that are spatially and / or temporally adjacent to a current block. The merge index can indicate at least one of motion information that the merge candidate has.
[0094] Transform unit: can represent a basic unit when encoding / decoding a residual signal such as transform, inverse transform, quantization, inverse quantization, transform coefficient encoding / decoding. A single transform unit can be partitioned into a plurality of lower-level transform units having smaller sizes. Here, the transform / inverse transform can include at least one of a first transform / first inverse transform and a second transform / second inverse transform.
[0095] Scaling: can denote a process of multiplying a quantized level by a factor. A transform coefficient can be generated by scaling a quantized level. Scaling can also be referred to as dequantization.
[0096] Quantization parameter: can denote a value used when a transform coefficient is used to generate a quantized level during quantization. The quantization parameter can also denote a value used when a transform coefficient is generated by scaling a quantized level during dequantization. The quantization parameter can be a value mapped with a quantization step.
[0097] Delta quantization parameter: can denote a difference value between a predicted quantization parameter and a quantization parameter of a coding / decoding target unit.
[0098] Scan: can denote a method of ordering coefficients within a unit, a block, or a matrix. For example, changing a two-dimensional matrix of coefficients into a one-dimensional matrix can be referred to as a scan, and changing a one-dimensional matrix of coefficients into a two-dimensional matrix can be referred to as a scan or an inverse scan.
[0099] Transform coefficient: can denote a coefficient value generated after a transform is performed in an encoder. The transform coefficient can denote a coefficient value generated after at least one of entropy decoding and dequantization is performed in a decoder. A quantized level or a quantized transform coefficient level obtained by quantizing a transform coefficient or a residual signal can also fall within the meaning of a transform coefficient.
[0100] Quantized level: can denote a value generated in an encoder by quantizing a transform coefficient or a residual signal. Alternatively, the quantized level can denote a value that is a dequantization target to be performed in a decoder. Similarly, a quantized transform coefficient level as a result of a transform and quantization can also fall within the meaning of a quantized level.
[0101] Non-zero transform coefficient: can denote a transform coefficient having a value other than zero, or a transform coefficient level or a quantized level having a value other than zero.
[0102] Quantization matrix: can denote a matrix used in a quantization process or a dequantization process performed to improve subjective image quality or objective image quality. The quantization matrix can also be referred to as a scaling list.
[0103] Quantization matrix coefficient: can denote each element within a quantization matrix. The quantization matrix coefficient can also be referred to as a matrix coefficient.
[0104] Default matrix: can denote a predetermined quantization matrix that is predefined in an encoder or a decoder.
[0105] Non-default matrix: can denote a quantization matrix that is not predefined but signaled by a user in an encoder or a decoder.
[0106] Statistical value: The statistical value for at least one of a variable having a calculable specific value, an encoding parameter, a constant value, etc., can be one or more of an average value, a weighted average value, a weighted sum value, a minimum value, a maximum value, a most frequently occurring value, a median value, an interpolation value of the respective specific value.
[0107] Figure 1 FIG. 1 is a block diagram illustrating a configuration of an encoding apparatus according to an embodiment of the present application.
[0108] The encoding apparatus 100 can be an encoder, a video encoding apparatus, or an image encoding apparatus. The video can include at least one image. The encoding apparatus 100 can sequentially encode the at least one image.
[0109] Referring to Figure 1 The encoding apparatus 100 can include a motion prediction unit 111, a motion compensation unit 112, an intra prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, a dequantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.
[0110] The encoding apparatus 100 can perform encoding on an input image by using an intra mode or an inter mode or both the intra mode and the inter mode. Further, the encoding apparatus 100 can generate a bitstream including encoded information by encoding the input image, and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium, or can be streamed through a wired / wireless transmission medium. When the intra mode is used as a prediction mode, the switch 115 can be switched to the intra. Alternatively, when the inter mode is used as the prediction mode, the switch 115 can be switched to the inter mode. Here, the intra mode can denote an intra prediction mode, and the inter mode can denote an inter prediction mode. The encoding apparatus 100 can generate a prediction block for an input block of an input image. Further, the encoding apparatus 100 can encode a residual block using a residual of the input block and the prediction block after generating the prediction block. The input image can be referred to as a current image as a current encoding target. The input block can be referred to as a current block as a current encoding target, or as an encoding target block.
[0111] When the prediction mode is the intra mode, the intra prediction unit 120 can use samples of a block that has been encoded / decoded and is adjacent to the current block as reference samples. The intra prediction unit 120 can perform spatial prediction on the current block by using the reference samples, or generate prediction samples of the input block by performing the spatial prediction. Here, the intra prediction can denote prediction within a frame.
[0112] When the prediction mode is the inter mode, the motion prediction unit 111 can retrieve a region that is most matched to the input block from a reference picture when performing motion prediction, and derive a motion vector by using the retrieved region. In this case, a search region can be used as the region. The reference picture can be stored in the reference picture buffer 190. Here, when encoding / decoding of the reference picture is performed, the reference picture can be stored in the reference picture buffer 190.
[0113] The motion compensation unit 112 can generate a prediction block by performing motion compensation on the current block by using the motion vector. Here, the inter prediction can mean prediction between frames or motion compensation.
[0114] When a value of the motion vector is not an integer, the motion prediction unit 111 and the motion compensation unit 112 can generate a prediction block by applying an interpolation filter to a partial region of a reference picture. In order to perform inter-picture prediction or motion compensation on a coding unit, it can be determined which mode among a skip mode, a merge mode, an advanced motion vector prediction (AMVP) mode, and a current picture reference mode is used for motion prediction and motion compensation of a prediction unit included in the corresponding coding unit. Then, depending on the determined mode, inter-picture prediction or motion compensation can be differently performed.
[0115] The subtractor 125 can generate a residual block by using a residual of the input block and the prediction block. The residual block can be referred to as a residual signal. The residual signal can mean a difference between an original signal and a prediction signal. Also, the residual signal can be a signal generated by transforming or quantizing or transforming and quantizing a difference between the original signal and the prediction signal. The residual block can be a residual signal of a block unit.
[0116] The transform unit 130 can generate transform coefficients by performing a transform on the residual block, and output the generated transform coefficients. Here, the transform coefficients can be coefficient values generated by performing a transform on the residual block. When a transform skip mode is applied, the transform unit 130 can skip the transform on the residual block.
[0117] A quantized level can be generated by applying quantization to the transform coefficients or to the residual signal. Hereinafter, the quantized level can also be referred to as a transform coefficient in an embodiment.
[0118] The quantization unit 140 can generate a quantized level by quantizing the transform coefficients or the residual signal according to a parameter, and output the generated quantized level. Here, the quantization unit 140 can quantize the transform coefficients by using a quantization matrix.
[0119] The entropy encoding unit 150 can generate a bitstream by performing entropy encoding on the values calculated by the quantization unit 140 or on the encoding parameter values calculated when encoding is performed according to a probability distribution, and output the generated bitstream. The entropy encoding unit 150 can perform entropy encoding on the sample information of the image and information used to decode the image. For example, the information used to decode the image can include syntax elements.
[0120] When entropy encoding is applied, symbols are represented such that a smaller number of bits is allocated to symbols having a high generation probability, and a larger number of bits is allocated to symbols having a low generation probability, and thus, the size of a bitstream of the symbols to be encoded can be reduced. The entropy encoding unit 150 can use an encoding method for entropy encoding such as exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), or the like. For example, the entropy encoding unit 150 can perform entropy encoding by using a variable length coding / code (VLC) table. Also, the entropy encoding unit 150 can derive a binarization method of a target symbol and a probability model of the target symbol / binary bit, and perform arithmetic encoding by using the derived binarization method and context model.
[0121] In order to encode the transform coefficient levels (quantized levels), the entropy encoding unit 150 can change the coefficients in a two-dimensional block form to a one-dimensional vector form by using a transform coefficient scanning method.
[0122] The coding parameters can include information such as syntax elements (flags, indices, etc.) that are coded in the encoder and signaled to the decoder, and information that is derived when encoding or decoding is performed. The coding parameters can represent information that is needed when encoding or decoding an image. For example, at least one value or combination of the following can be included in the coding parameters: unit / block size, unit / block depth, unit / block partition information, unit / block shape, unit / block partition structure, whether partitioning in a quad-tree form is performed, whether partitioning in a binary-tree form is performed, partitioning direction in a binary-tree form (horizontal direction or vertical direction), partitioning form in a binary-tree form (symmetric partitioning or asymmetric partitioning), whether the current coding unit is partitioned by triple-tree partitioning, triple-tree partitioning direction (horizontal direction or vertical direction), triple-tree partitioning type (symmetric type or asymmetric type), whether the current coding unit is partitioned by multi-type tree partitioning, multi-type tree partitioning direction (horizontal direction or vertical direction), multi-type tree partitioning type (symmetric type or asymmetric type), and multi-type tree partitioning tree (binary tree or triple tree) structure, prediction mode (intra prediction or inter prediction), luma intra prediction mode / direction, chroma intra prediction mode / direction, intra partition information, inter partition information, coding block partition flag, prediction block partition flag, transform block partition flag, reference sample filtering method, reference sample filter tap, reference sample filter coefficient, prediction block filtering method, prediction block filter tap, prediction block filter coefficient, prediction block boundary filtering method, prediction block boundary filter tap, prediction block boundary filter coefficient, intra prediction mode, inter prediction mode, motion information, motion vector, motion vector difference, reference picture index, inter prediction angle, inter prediction indicator, prediction list utilization flag, reference picture list, reference picture, motion vector predictor index, motion vector predictor candidate, motion vector candidate list, whether to use merge mode, merge index, merge candidate, merge candidate list, whether to use skip mode, interpolation filter type, interpolation filter tap, interpolation filter coefficient, motion vector size, representation precision of motion vector, transform type, transform size, information whether first (primary) transform is used, information whether secondary transform is used, first transform index, secondary transform index, information whether residual signal exists, coding block pattern, coding block flag (CBF), quantization parameter, quantization parameter residual, quantization matrix, whether to apply intra loop filter, intra loop filter coefficient, intra loop filter tap, intra loop filter shape / form, whether to apply deblocking filter, deblocking filter coefficient, deblocking filter tap, deblocking filter strength, deblocking filter shape / form, whether to apply adaptive sample offset, adaptive sample offset value, adaptive sample offset class, adaptive sample offset type, whether to apply adaptive loop filter, adaptive loop filter coefficient, adaptive loop filter tap, adaptive loop filter shape / form,binarization / de-binarization method, context model determination method, context model update method, whether to perform normal mode, whether to perform bypass mode, context bin, bypass bin, significant coefficient flag, last significant coefficient flag, coding flag for a unit of coefficient group, position of last significant coefficient, flag on whether value of coefficient is greater than 1, flag on whether value of coefficient is greater than 2, flag on whether value of coefficient is greater than 3, information on remaining coefficient value, sign information, reconstructed luma sample, reconstructed chroma sample, residual luma sample, residual chroma sample, luma transform coefficient, chroma transform coefficient, quantized luma level, quantized chroma level, transform coefficient level scanning method, size of motion vector search area at decoder side, shape of motion vector search area at decoder side, number of times of motion vector search at decoder side, information on CTU size, information on minimum block size, information on maximum block size, information on maximum block depth, information on minimum block depth, image display / output order, slice identification information, slice type, slice partition information, parallel block identification information, parallel block type, parallel block partition information, picture type, bit depth of input sample, bit depth of reconstructed sample, bit depth of residual sample, bit depth of transform coefficient, bit depth of quantized level, and information on luma signal or information on chroma signal.
[0123] Here, signaling a flag or an index can mean that the corresponding flag or index is entropy-encoded by the encoder and included in the bitstream, and can mean that the corresponding flag or index is entropy-decoded by the decoder from the bitstream.
[0124] When the encoding apparatus 100 performs encoding through inter prediction, the encoded current picture can be used as a reference picture for another picture which is processed later. Accordingly, the encoding apparatus 100 can reconstruct or decode the encoded current picture, or store the reconstructed or decoded picture as a reference picture in the reference picture buffer 190.
[0125] The quantized level can be dequantized in the dequantization unit 160, or can be inverse-transformed in the inverse transform unit 170. The coefficient which is dequantized or inverse-transformed or both can be added to the prediction block by the adder 175. By adding the coefficient which is dequantized or inverse-transformed or both to the prediction block, the reconstructed block can be generated. Here, the coefficient which is dequantized or inverse-transformed or both can mean a coefficient which has undergone at least one of dequantization and inverse transformation, and can mean a reconstructed residual block.
[0126] The reconstructed block can pass through a filter unit 180. The filter unit 180 can apply at least one of a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF) to the reconstructed sample, the reconstructed block, or the reconstructed picture. The filter unit 180 can be referred to as an in-loop filter.
[0127] The deblocking filter can remove block distortion generated in a boundary between blocks. In order to determine whether to apply the deblocking filter, it can be determined whether to apply the deblocking filter to the current block based on samples included in a number of rows or columns included in the block. When the deblocking filter is applied to the block, another filter can be applied according to a required deblocking filter strength.
[0128] In order to compensate for encoding errors, a suitable offset value can be added to a sample value by using a sample adaptive offset. The sample adaptive offset can correct an offset of a de-blocked picture from an original picture in units of samples. A method of applying an offset considering edge information about each sample can be used, or a method of partitioning samples of a picture into a predetermined number of regions, determining a region to which an offset is applied, and applying the offset to the determined region can be used.
[0129] The adaptive loop filter can perform filtering based on a comparison result of a filtered reconstructed picture and an original picture. Samples included in the picture can be partitioned into a predetermined group, a filter to be applied to each group can be determined, and differential filtering can be performed on each group. Information on whether to apply the ALF can be signaled through a coding unit (CU), and a form and coefficients of the ALF to be applied to each block can vary.
[0130] The reconstructed block or the reconstructed picture that has passed through the filter unit 180 can be stored in a reference picture buffer 190. The reconstructed block processed by the filter unit 180 can be a part of a reference picture. That is, the reference picture is a reconstructed picture composed of the reconstructed blocks processed by the filter unit 180. The stored reference picture can be used later at inter prediction or motion compensation.
[0131] Figure 2 FIG. 1 is a block diagram illustrating a configuration of a decoding apparatus according to an embodiment and to which the present application is applied.
[0132] The decoding apparatus 200 can be a decoder, a video decoding apparatus, or a picture decoding apparatus.
[0133] Referring to Figure 2 , the decoding apparatus 200 can include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, a motion compensation unit 250, an adder 225, a filter unit 260, and a reference picture buffer 270.
[0134] The decoding apparatus 200 can receive a bitstream output from the encoding apparatus 100. The decoding apparatus 200 can receive a bitstream stored in a computer-readable recording medium, or can receive a bitstream being streamed through a wired / wireless transmission medium. The decoding apparatus 200 can decode the bitstream by using an intra mode or an inter mode. Furthermore, the decoding apparatus 200 can generate a reconstructed image or a decoded image generated by decoding, and output the reconstructed image or the decoded image.
[0135] When the prediction mode used in decoding is the intra mode, the switch can be switched to the intra. Alternatively, when the prediction mode used in decoding is the inter mode, the switch can be switched to the inter mode.
[0136] The decoding apparatus 200 can obtain a reconstructed residual block by decoding an input bitstream, and generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding apparatus 200 can generate a reconstructed block that is a decoding target by adding the reconstructed residual block to the prediction block. The decoding target block can be referred to as a current block.
[0137] The entropy decoding unit 210 can generate a symbol by entropy-decoding a bitstream according to a probability distribution. The generated symbol can include a quantized level form of a symbol. Here, the entropy-decoding method can be an inverse process of the above-described entropy-encoding method.
[0138] In order to decode a transform coefficient level (quantized level), the entropy decoding unit 210 can change a coefficient in a one-dimensional vector form to a two-dimensional block form by using a transform coefficient scanning method.
[0139] The quantized level can be inverse-quantized in the inverse quantization unit 220, or can be inverse-transformed in the inverse transform unit 230. The quantized level can be a result of inverse-quantization or inverse-transformation, or both inverse-quantization and inverse-transformation, and can be generated as a reconstructed residual block. Here, the inverse quantization unit 220 can apply a quantization matrix to the quantized level.
[0140] When the intra mode is used, the intra prediction unit 240 can generate a prediction block by performing spatial prediction on the current block, in which the spatial prediction uses sample values of blocks adjacent to the decoding target block and already decoded.
[0141] When the inter mode is used, the motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block, in which the motion compensation uses a motion vector and a reference image stored in the reference picture buffer 270.
[0142] The adder 225 can generate a reconstructed block by adding the reconstructed residual block to the prediction block. The filter unit 260 can apply at least one of a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the reconstructed block or the reconstructed picture. The filter unit 260 can output the reconstructed picture. The reconstructed block or the reconstructed picture can be stored in the reference picture buffer 270 and used when performing inter prediction. The reconstructed block processed by the filter unit 260 can be a part of a reference picture. That is, the reference picture is a reconstructed picture composed of the reconstructed blocks processed by the filter unit 260. The stored reference picture can be used later when performing inter prediction or motion compensation.
[0143] Figure 3 FIG. 1 is a diagram schematically illustrating a partition structure of an image when encoding and decoding the image. Figure 3 FIG. 2 schematically illustrates an example of partitioning a single unit into a plurality of lower-level units.
[0144] To effectively partition an image, a coding unit (CU) can be used when encoding and decoding. The coding unit can be used as a basic unit when encoding / decoding an image. Also, the coding unit can be used as a unit for distinguishing an intra prediction mode from an inter prediction mode when encoding / decoding an image. The coding unit can be a basic unit for prediction, transform, quantization, inverse transform, dequantization, or encoding / decoding processing of transform coefficients.
[0145] Referring to Figure 3 , the image 300 is sequentially partitioned in a maximum coding unit (LCU) and the LCU unit is determined as a partition structure. Here, the LCU can be used in the same meaning as a coding tree unit (CTU). The unit partitioning can denote partitioning of a block associated with the unit. In the block partitioning information, information of a unit depth can be included. The depth information can denote either or both of a number or degree of partitioning of a unit or a number and degree of partitioning of a unit. A single unit can be partitioned into a plurality of lower-level units hierarchically associated with the depth information based on a tree structure. In other words, the unit and the lower-level units generated by partitioning the unit can correspond to a node and child nodes of the node, respectively. Each of the partitioned lower-level units can have the depth information. The depth information can be information denoting a size of a CU and can be stored in each CU. The unit depth denotes a number and / or degree related to partitioning of a unit. Accordingly, the partitioning information of the lower-level units can include information on sizes of the lower-level units.
[0146] The partition structure can represent a distribution of coding units (CUs) within the LCU 310. The distribution can be determined according to whether a single CU is partitioned into a plurality of CUs (a positive integer equal to or greater than 2, including 2, 4, 8, 16, etc.). The horizontal and vertical sizes of the CUs generated by the partitioning can be half of the horizontal and vertical sizes of the CU before the partitioning, respectively, or can have sizes smaller than the horizontal and vertical sizes before the partitioning, respectively, according to the number of times of partitioning. The CU can be recursively partitioned into a plurality of CUs. At least one of the height and width of the CU after the partitioning can be reduced compared to at least one of the height and width of the CU before the partitioning by the recursive partitioning. The partitioning of the CU can be recursively performed until a predefined depth or a predefined size is reached. For example, the depth of the LCU can be 0 and the depth of the smallest coding unit (SCU) can be a predefined maximum depth. Here, as described above, the LCU can be a coding unit having a maximum coding unit size and the SCU can be a coding unit having a minimum coding unit size. The partitioning starts from the LCU 310, and the CU depth is increased by 1 when the horizontal size or the vertical size, or both, of the CU is reduced by the partitioning. For example, the size of the CU that is not partitioned can be 2Nx2N for each depth. Further, in the case of the CU that is partitioned, the CU having a size of 2Nx2N can be partitioned into four CUs having a size of NxN. When the depth is increased by 1, the size of N can be halved.
[0147] Further, information on whether the CU is partitioned can be represented by using partitioning information of the CU. The partitioning information can be 1-bit information. All CUs except for the SCU can include the partitioning information. For example, when the value of the partitioning information is 1, the CU can not be partitioned, and when the value of the partitioning information is 2, the CU can be partitioned.
[0148] Referring to Figure 3 , the LCU having a depth of 0 can be a 64x64 block. 0 can be a minimum depth. The SCU having a depth of 3 can be an 8x8 block. 3 can be a maximum depth. The CUs of the 32x32 block and the 16x16 block can be represented as depths 1 and 2, respectively.
[0149] For example, when a single coding unit is partitioned into four coding units, the horizontal and vertical sizes of the partitioned four coding units can be half the size of the horizontal and vertical sizes of the CU before being partitioned. In one embodiment, when a coding unit having a size of 32x32 is partitioned into four coding units, each of the partitioned four coding units can have a size of 16x16. When a single coding unit is partitioned into four coding units, the coding unit can be said to be partitioned in a quad-tree form.
[0150] For example, when one coding unit is partitioned into two sub-coding units, each of the two sub-coding units can have a horizontal size or a vertical size (width or height) that is half of a horizontal size or a vertical size of the original coding unit. For example, when a coding unit having a size of 32x32 is vertically partitioned into two sub-coding units, each of the two sub-coding units can have a size of 16x32. For example, when a coding unit having a size of 8x32 is horizontally partitioned into two sub-coding units, each of the two sub-coding units can have a size of 8x16. When one coding unit is partitioned into two sub-coding units, the coding unit can be referred to as being bipartitioned, or partitioned according to a binary tree partitioning structure.
[0151] For example, when one coding unit is partitioned into three sub-coding units, a horizontal size or a vertical size of the coding unit can be partitioned in a ratio of 1:2:1, thereby resulting in three sub-coding units having a ratio of 1:2:1 in horizontal size or vertical size. For example, when a coding unit having a size of 16x32 is horizontally partitioned into three sub-coding units, the three sub-coding units can have sizes of 16x8, 16x16, and 16x8, in order from a topmost sub-coding unit to a bottommost sub-coding unit. For example, when a coding unit having a size of 32x32 is vertically partitioned into three sub-coding units, the three sub-coding units can have sizes of 8x32, 16x32, and 8x32, in order from a leftmost sub-coding unit to a rightmost sub-coding unit. When one coding unit is partitioned into three sub-coding units, the coding unit can be referred to as being tripartitioned, or partitioned according to a ternary tree partitioning structure.
[0152] In Figure 3 The coding tree unit (CTU) 320 is an example of a CTU to which all of the quadtree partitioning structure, the binary tree partitioning structure, and the ternary tree partitioning structure are applied.
[0153] As described above, in order to partition a CTU, at least one of the quadtree partitioning structure, the binary tree partitioning structure, and the ternary tree partitioning structure can be applied. The various tree partitioning structures can be applied to the CTU sequentially according to a predetermined priority order. For example, the quadtree partitioning structure can be preferentially applied to the CTU. A coding unit that cannot be further partitioned using the quadtree partitioning structure can correspond to a leaf node of the quadtree. The coding unit corresponding to the leaf node of the quadtree can be used as a root node of the binary tree and / or the ternary tree partitioning structure. That is, the coding unit corresponding to the leaf node of the quadtree can be further partitioned according to the binary tree partitioning structure or the ternary tree partitioning structure, or can not be further partitioned. Thus, by preventing binary tree partitioning or ternary tree partitioning of the coding unit corresponding to the leaf node of the quadtree, coding blocks resulting from further quadtree partitioning of the coding unit can be efficiently block-partitioned and / or signaling of partition information can be performed.
[0154] The fact that a coding unit corresponding to a node of a quadtree is partitioned can be signaled using quad-partition information. The quad-partition information having a first value (e.g., "1") can indicate that the current coding unit is partitioned according to a quadtree partition structure. The quad-partition information having a second value (e.g., "0") can indicate that the current coding unit is not partitioned according to a quadtree partition structure. The quad-partition information can be a flag having a predetermined length (e.g., one bit).
[0155] There can be no priority between binary tree partitioning and ternary tree partitioning. That is, a coding unit corresponding to a leaf node of a quadtree can be further partitioned by either of binary tree partitioning and ternary tree partitioning. Further, a coding unit generated by binary tree partitioning or ternary tree partitioning can be further partitioned by either of binary tree partitioning or ternary tree partitioning, or can not be further partitioned.
[0156] A tree structure in which there is no priority between binary tree partitioning and ternary tree partitioning is referred to as a multi-type tree structure. A coding unit corresponding to a leaf node of a quadtree can serve as a root node of a multi-type tree. Whether a coding unit corresponding to a node of a multi-type tree is partitioned can be signaled using at least one of multi-type tree partitioning indication information, partition direction information, and partition tree information. In order to partition a coding unit corresponding to a node of a multi-type tree, the multi-type tree partitioning indication information, the partition direction, and the partition tree information can be sequentially signaled.
[0157] The multi-type tree partitioning indication information having a first value (e.g., "1") can indicate that the current coding unit is to be partitioned by a multi-type tree partition. The multi-type tree partitioning indication information having a second value (e.g., "0") can indicate that the current coding unit is not to be partitioned by a multi-type tree partition.
[0158] When a coding unit corresponding to a node of a multi-type tree is further partitioned according to a multi-type tree partition structure, the coding unit can include partition direction information. The partition direction information can indicate in which direction the current coding unit is to be partitioned according to a multi-type tree partition. The partition direction information having a first value (e.g., "1") can indicate that the current coding unit is to be vertically partitioned. The partition direction information having a second value (e.g., "0") can indicate that the current coding unit is to be horizontally partitioned.
[0159] When a coding unit corresponding to a node of a multi-type tree is further partitioned according to a multi-type tree partition structure, the current coding unit can include partition tree information. The partition tree information can indicate a tree partition structure to be used for partitioning a node of a multi-type tree. The partition tree information having a first value (e.g., "1") can indicate that the current coding unit is to be partitioned according to a binary tree partition structure. The partition tree information having a second value (e.g., "0") can indicate that the current coding unit is to be partitioned according to a ternary tree partition structure.
[0160] The partition indication information, the partition tree information, and the partition direction information can each be a flag having a predetermined length (e.g., one bit).
[0161] At least any one of the quad-tree partition indication information, the multi-type tree partition indication information, the partition direction information, and the partition tree information can be entropy coded / entropy decoded. In order to entropy code / entropy decode those types of information, information about neighboring coding units adjacent to the current coding unit can be used. For example, it is highly likely that the partition type (partitioned or not partitioned, partition tree, and / or partition direction) of a left neighboring coding unit and / or an above neighboring coding unit of the current coding unit is similar to the partition type of the current coding unit. Thus, context information for entropy coding / entropy decoding information about the current coding unit can be derived from information about the neighboring coding units. The information about the neighboring coding units can include at least any one of the quad-tree partition information, the multi-type tree partition indication information, the partition direction information, and the partition tree information.
[0162] As another example, in the binary tree partition and the ternary tree partition, the binary tree partition can be preferentially performed. That is, the current coding unit can first be subjected to the binary tree partition, and then coding units corresponding to leaf nodes of the binary tree can be set as root nodes for the ternary tree partition. In this case, for coding units corresponding to nodes of the ternary tree, neither the quad-tree partition nor the binary tree partition can be performed.
[0163] A coding unit that cannot be partitioned according to the quad-tree partition structure, the binary tree partition structure, and / or the ternary tree partition structure becomes a basic unit for encoding, prediction, and / or transformation. That is, the coding unit cannot be further partitioned for prediction and / or transformation. Thus, in a bitstream, there can be no partition structure information and partition information for partitioning a coding unit into prediction units and / or transformation units.
[0164] However, when the size of a coding unit (i.e., a basic unit for partitioning) is larger than the size of the maximum transform block, the coding unit can be recursively partitioned until the size of the coding unit is reduced to be equal to or smaller than the size of the maximum transform block. For example, when the size of the coding unit is 64x64 and when the size of the maximum transform block is 32x32, the coding unit can be partitioned into four 32x32 blocks for transform. For example, when the size of the coding unit is 32x64 and the size of the maximum transform block is 32x32, the coding unit can be partitioned into two 32x32 blocks for transform. In this case, the partitioning of the coding unit for transform is not signaled separately and can be determined by a comparison between the horizontal size or the vertical size of the coding unit and the horizontal size or the vertical size of the maximum transform block. For example, when the horizontal size (width) of the coding unit is larger than the horizontal size (width) of the maximum transform block, the coding unit can be vertically bisected. For example, when the vertical size (length) of the coding unit is larger than the vertical size (length) of the maximum transform block, the coding unit can be horizontally bisected.
[0165] The information of the maximum size and / or the minimum size of the coding unit and the information of the maximum size and / or the minimum size of the transform block can be signaled or determined at a higher level of the coding unit. The higher level can be, for example, a sequence level, a picture level, a slice level, etc. For example, the minimum size of the coding unit can be determined to be 4x4. For example, the maximum size of the transform block can be determined to be 64x64. For example, the minimum size of the transform block can be determined to be 4x4.
[0166] The information of the minimum size of the coding unit corresponding to a leaf node of a quad tree (quad tree minimum size) and / or the information of the maximum depth of a multi-type tree from a root node to a leaf node (maximum tree depth of the multi-type tree) can be signaled or determined at a higher level of the coding unit. For example, the higher level can be a sequence level, a picture level, a slice level, etc. The information of the minimum size of the quad tree and / or the information of the maximum depth of the multi-type tree can be signaled or determined for each of an intra slice and an inter slice.
[0167] The difference information between the size of the CTU and the maximum size of the transform block can be signaled or determined at a higher level of the coding unit. For example, the higher level can be a sequence level, a picture level, a slice level, and the like. Information of the maximum size of the coding unit corresponding to each node of the binary tree (hereinafter, referred to as the maximum size of the binary tree) can be determined based on the size of the coding tree unit and the difference information. The maximum size of the coding unit corresponding to each node of the ternary tree (hereinafter, referred to as the maximum size of the ternary tree) can vary depending on the type of the slice. For example, the maximum size of the ternary tree can be 32x32 for an intra slice. For example, the maximum size of the ternary tree can be 128x128 for an inter slice. For example, the minimum size of the coding unit corresponding to each node of the binary tree (hereinafter, referred to as the minimum size of the binary tree) and / or the minimum size of the coding unit corresponding to each node of the ternary tree (hereinafter, referred to as the minimum size of the ternary tree) can be set to the minimum size of the coding block.
[0168] As another example, the maximum size of the binary tree and / or the maximum size of the ternary tree can be signaled or determined at the slice level. Alternatively, the minimum size of the binary tree and / or the minimum size of the ternary tree can be signaled or determined at the slice level.
[0169] According to the size information and the depth information of the various blocks described above, the quad partition information, the multi-type tree partitioning indication information, the partition tree information, and / or the partition direction information can or can not be included in the bitstream.
[0170] For example, when the size of the coding unit is not greater than the minimum size of the quad tree, the coding unit does not include the quad partition information. Accordingly, the quad partition information can be derived from the second value.
[0171] For example, when the size (horizontal size and vertical size) of the coding unit corresponding to the node of the multi-type tree is greater than the maximum size (horizontal size and vertical size) of the binary tree and / or the maximum size (horizontal size and vertical size) of the ternary tree, the coding unit can not be bi-partitioned or tri-partitioned. Accordingly, the multi-type tree partitioning indication information can not be signaled, but can be derived from the second value.
[0172] Optionally, the coding unit corresponding to the node of the multi-type tree can not be further bi-partitioned or tri-partitioned when the size (horizontal size and vertical size) of the coding unit is the same as the maximum size (horizontal size and vertical size) of the binary tree and / or twice as large as the maximum size (horizontal size and vertical size) of the ternary tree. Thus, the multi-type tree partitioning indication information can not be signaled, but can be derived from the second value. This is because, when the coding unit is partitioned according to the binary tree partitioning structure and / or the ternary tree partitioning structure, a coding unit smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree is generated.
[0173] Optionally, the coding unit corresponding to the node of the multi-type tree can not be further bi-partitioned and / or tri-partitioned when the depth of the coding unit is equal to the maximum depth of the multi-type tree. Thus, the multi-type tree partitioning indication information can not be signaled, but can be derived from the second value.
[0174] Optionally, the multi-type tree partitioning indication information can be signaled only when at least one of the vertical direction binary tree partitioning, the horizontal direction binary tree partitioning, the vertical direction ternary tree partitioning and the horizontal direction ternary tree partitioning is feasible for the coding unit corresponding to the node of the multi-type tree. Otherwise, the coding unit can not be bi-partitioned and / or tri-partitioned. Thus, the multi-type tree partitioning indication information can not be signaled, but can be derived from the second value.
[0175] Optionally, the partitioning direction information can be signaled only when both the vertical direction binary tree partitioning and the horizontal direction binary tree partitioning or both the vertical direction ternary tree partitioning and the horizontal direction ternary tree partitioning are feasible for the coding unit corresponding to the node of the multi-type tree. Otherwise, the partitioning direction information can not be signaled, but can be derived from the value indicating the possible partitioning directions.
[0176] Optionally, the partitioning tree information can be signaled only when both the vertical direction binary tree partitioning and the vertical direction ternary tree partitioning or both the horizontal direction binary tree partitioning and the horizontal direction ternary tree partitioning are feasible for the coding tree corresponding to the node of the multi-type tree. Otherwise, the partitioning tree information can not be signaled, but can be derived from the value indicating the possible partitioning tree structures.
[0177] Figure 4 is a diagram illustrating an embodiment of the inter picture prediction process.
[0178] In Figure 4 , a rectangle can represent a picture. In Figure 4In the middle, an arrow indicates a prediction direction. Pictures can be classified into an intra picture (I picture), a predicted picture (P picture), and a bi-predicted picture (B picture) according to an encoding type of the picture.
[0179] An I picture can be encoded by intra prediction without inter prediction. A P picture can be encoded by inter prediction using a reference picture existing in one direction (i.e., a forward direction or a backward direction) for a current block. A B picture can be encoded by inter prediction using reference pictures existing in two directions (i.e., a forward direction and a backward direction) for the current block. When inter prediction is used, an encoder can perform inter prediction or motion compensation, and a decoder can perform corresponding motion compensation.
[0180] Hereinafter, embodiments of inter prediction will be described in detail.
[0181] Inter prediction or motion compensation can be performed using a reference picture and motion information.
[0182] Motion information of a current block can be derived during inter prediction by each of the encoding apparatus 100 and the decoding apparatus 200. The motion information of the current block can be derived by using motion information of a reconstructed neighboring block, a collocated block (also referred to as a col block or a co-located block), and / or a block adjacent to the collocated block. The collocated block can indicate a block spatially located at the same position as the current block within a previously reconstructed co-located picture (also referred to as a col picture or a co-located picture). The co-located picture can be one picture among one or more reference pictures included in a reference picture list.
[0183] A method of deriving motion information of a current block can vary according to a prediction mode of the current block. For example, as a prediction mode for inter prediction, there can be an AMVP mode, a merge mode, a skip mode, a current picture reference mode, etc. The merge mode can be referred to as a motion merge mode.
[0184] For example, when the AMVP is used as the prediction mode, at least one of a motion vector of a reconstructed neighboring block, a motion vector of a collocated block, a motion vector of a block adjacent to the collocated block, and a (0, 0) motion vector can be determined as a motion vector candidate for the current block, and a motion vector candidate list is generated by using the motion vector candidate. A motion vector candidate of the current block can be derived by using the generated motion vector candidate list. Motion information of the current block can be determined based on the derived motion vector candidate. The motion vector of the collocated block or the motion vector of the block adjacent to the collocated block can be referred to as a temporal motion vector candidate, and the motion vector of the reconstructed neighboring block can be referred to as a spatial motion vector candidate.
[0185] The encoding apparatus 100 can calculate a motion vector difference (MVD) between a motion vector of the current block and a motion vector candidate, and can perform entropy encoding on the motion vector difference (MVD). Further, the encoding apparatus 100 can perform entropy encoding on a motion vector candidate index and generate a bitstream. The motion vector candidate index can indicate a best motion vector candidate among motion vector candidates included in a motion vector candidate list. The decoding apparatus can perform entropy decoding on the motion vector candidate index included in the bitstream, and can select a motion vector candidate of a decoding target block from among the motion vector candidates included in the motion vector candidate list by using the motion vector candidate index that is entropy-decoded. Further, the decoding apparatus 200 can add the entropy-decoded MVD to the motion vector candidate extracted through entropy decoding, thereby deriving a motion vector of the decoding target block.
[0186] The bitstream can include a reference picture index indicating a reference picture. The reference picture index can be entropy-encoded by the encoding apparatus 100 and then signaled to the decoding apparatus 200 as the bitstream. The decoding apparatus 200 can generate a prediction block of a decoding target block based on the derived motion vector and the reference picture index information.
[0187] Another example of a method of deriving motion information of a current block can be a merge mode. The merge mode can denote a method of merging motions of a plurality of blocks. The merge mode can denote a mode of deriving motion information of a current block from motion information of a neighboring block. When the merge mode is applied, a merge candidate list can be generated using motion information of a reconstructed neighboring block and / or motion information of a collocated block. The motion information can include at least one of a motion vector, a reference picture index, and an inter prediction indicator. The prediction indicator can indicate a uni-prediction (L0 prediction or L1 prediction) or bi-prediction (L0 prediction and L1 prediction).
[0188] The merge candidate list can be a list storing motion information. The motion information included in the merge candidate list can be at least one of a zero merge candidate, and a new motion information that is a combination of motion information of one neighboring block adjacent to the current block (spatial merge candidate), motion information of a collocated block of the current block included in a reference picture (temporal merge candidate), and motion information present in the merge candidate list.
[0189] The encoding apparatus 100 can generate a bitstream by performing entropy encoding on at least one of a merge flag and a merge index, and can signal the bitstream to the decoding apparatus 200. The merge flag can be information indicating whether or not the merge mode is performed for each block, and the merge index can be information indicating which of the neighboring blocks of the current block is a merge target block. For example, the neighboring blocks of the current block can include a left neighboring block disposed at a left side of the current block, an above neighboring block disposed above the current block, and a temporal neighboring block adjacent in time to the current block.
[0190] The skip mode can be a mode in which motion information of a neighboring block is applied as it is to the current block. When the skip mode is applied, the encoding apparatus 100 can entropy-encode information on which block's motion information is to be used as the motion information of the current block, thereby generating a bitstream, and can signal the bitstream to the decoding apparatus 200. The encoding apparatus 100 can not signal syntax elements on at least any one of motion vector difference information, a coded block flag, and transform coefficient levels to the decoding apparatus 200.
[0191] The current picture reference mode can denote a prediction mode in which a previously reconstructed region within a current picture to which the current block belongs is used for prediction. Here, a vector can be used to specify the previously reconstructed region. Information indicating whether the current block is to be encoded in the current picture reference mode can be encoded by using a reference picture index of the current block. A flag or index indicating whether the current block is encoded in the current picture reference mode can be signaled, and can be derived based on the reference picture index of the current block. In a case where the current block is encoded in the current picture reference mode, the current picture can be added to a reference picture list for the current block so as to be located at a fixed position or a random position in the reference picture list. The fixed position can be, for example, a position indicated by a reference picture index 0 or a last position in the list. In a case where the current picture is added to the reference picture list so as to be located at the random position, a reference picture index indicating the random position can be signaled.
[0192] Hereinafter, an image encoding / decoding method according to the present application will be described in detail with reference to the above description.
[0193] Figure 5 FIG. 4 is a diagram illustrating a flowchart of an image decoding method according to an embodiment of the present application.
[0194] Referring to Figure 5 At S510, the decoding apparatus can entropy-decode motion-estimated information based on the merge mode, and at S520, derive a merge candidate for the current block based on the decoded information. Subsequently, at S530, the decoding apparatus can generate an initial merge candidate list of the merge candidate by using the previously derived merge candidate. Subsequently, at S540, the decoding apparatus can generate a merge candidate list reconfigured by using the initial merge candidate list.
[0195] The merge candidate derived by the decoding apparatus can include at least one of a spatial merge candidate, a temporal merge candidate, a subblock-based temporal merge candidate, a subblock-based spatio-temporal combination merge candidate, and an additional merge candidate. However, the types of the merge candidates that can be merged by the decoding apparatus are not limited thereto, and various forms of the merge candidates that can be implemented by one of ordinary skill in the art can be applied to the present application.
[0196] Figure 6 FIG. 4 is a diagram illustrating a flowchart of an image decoding method according to an embodiment of the present application.
[0197] Referring to Figure 6 S540 in which the decoding device generates a reconstructed merge candidate list will be described in detail. In S610, the decoding device can calculate a distortion value of a derived merge candidate by using motion information of a neighboring block of the current block. Subsequently, in S620, the decoding device can generate a reconfigured merge candidate list by reconfiguring the initial merge candidate list based on the calculated distortion value.
[0198] Figure 7 FIG. 5 is a diagram illustrating a method of deriving a spatial merge candidate.
[0199] Referring to Figure 7 A method of deriving a spatial merge candidate of a current block to be decoded by the decoding device will be described in detail. The decoding device can derive a spatial merge candidate of the current block from reconstructed neighboring blocks that are spatially adjacent to the current block.
[0200] For example, motion information can be derived from blocks corresponding to a block Al located at the left side of the current block X, a block Bl located at the top of the current block X, a block B0 located at the top right corner of the current block X, a block A0 located at the bottom left corner of the current block X, and a block B2 located at the top left corner of the current block X, and the derived information can be used as a spatial merge candidate of the current block.
[0201] When deriving a spatial merge candidate of the current block from reconstructed neighboring blocks, when motion information of the neighboring blocks is decoded by an affine transform model mode (affine mode) or a current picture reference (CPR) mode, the decoding device does not use the corresponding neighboring blocks as spatial merge candidates. Here, the CPR mode can denote a prediction mode in which a current picture can be used as a reference picture when performing intra- or inter-prediction.
[0202] In addition, the neighboring blocks can include corrected motion information instead of initial motion information. Here, the decoding device can use the initial motion information of the neighboring blocks as a spatial merge candidate of the current block, instead of using the corrected motion information as a spatial merge candidate of the current block.
[0203] A spatial merge candidate can indicate motion information of a reconstructed neighboring block that is spatially adjacent to the current block, and can have a square shape or a non-square shape. In addition, the reconstructed neighboring block that is spatially adjacent to the current block can be divided into a lower level block (sub-block) unit. The decoding device can derive at least one spatial merge candidate for each lower level block.
[0204] In another example, the spatial merge candidate can include motion information of a reconstructed neighboring block that is spatially non-adjacent to the current block. Here, the reconstructed neighboring block that is spatially non-adjacent to the current block can be a block located within the same CTU as the current block.
[0205] Meanwhile, when the reconstructed neighboring block that is spatially non-adjacent is located in a different CTU from the current block, the reconstructed neighboring block that is spatially non-adjacent is not used as a spatial merge candidate of the current block. However, when the neighboring block that is spatially non-adjacent is located on the upper boundary or the left boundary or both the upper boundary and the left boundary of the CTU to which the current block belongs, the corresponding neighboring block can be used as a spatial merge candidate of the current block even though the neighboring block that is spatially non-adjacent is located within a different CTU.
[0206] Here, deriving a spatial merge candidate can mean deriving the spatial merge candidate and adding it to the merge candidate list. Here, the motion information of each of the merge candidates included in the merge candidate list can be different.
[0207] When the spatial merge candidate is added to the merge candidate list by the decoding device, the decoding device can determine whether the motion information of all the spatial merge candidates present in the merge candidate list is the same as that of the newly added spatial merge candidate, instead of directly determining whether the motion information of the candidate previously added to the merge candidate list is the same as that of the newly added spatial merge candidate. When the decoding device determines that there is no spatial merge candidate identical to the newly added spatial merge candidate in the merge candidate list, the decoding device can add the spatial merge candidate to the merge candidate list.
[0208] The decoding device can derive up to maxNumSpatialMergeCand spatial merge candidates. Here, maxNumSpatialMergeCand can be a positive integer including 0.
[0209] In an example, maxNumSpatialMVPCand can be 5. MaxNumMergeCand can be the maximum number of merge candidates that can be included in the merge candidate list, and can be a positive integer including 0. Further, numMergeCand can denote the number of merge candidates included in the actual merge candidate list within the preset MaxNumMergeCand. However, the use of numMergeCand and MaxNumMergeCand does not limit the scope of the present application. The decoding device can use the above information by using parameter values having the same meaning as numMergeCand and MaxNumMergeCand.
[0210] Figure 8 is a diagram showing a method of deriving a temporal merge candidate.
[0211] Referring to Figure 8 A method of deriving a temporal merge candidate for a current block to be decoded by a decoding apparatus will be described in detail. The decoding apparatus can derive the temporal merge candidate from a reconstructed block in a reference picture (reference picture) temporally neighboring the current block. The reference picture temporally neighboring the current block can represent a collocated picture (collocated picture). In addition, information of the collocated picture can be transmitted from an encoding apparatus to the decoding apparatus in units of at least one of a sequence, a picture, a slice, a parallel block, a CTU, and a coding block unit within a CTU.
[0212] Alternatively, the information of the collocated picture can be implicitly derived by using at least one piece of motion information of a block that has been encoded / decoded, which is currently or temporally-spatially or both currently and temporally-spatially neighboring a hierarchy according to an encoding / decoding order, and using inter prediction indicator or reference picture index information of the collocated picture at a sequence, a picture, a slice, and a parallel block level.
[0213] The information of the collocated picture can include at least one of an inter prediction indicator, a reference picture index, and motion vector information indicating a collocated block of the current block.
[0214] Here, when deriving the temporal merge candidate for the current block, a position of the collocated picture and a collocated block within the collocated picture can be determined by using at least one piece of motion information of a block that has been decoded, which is temporally-spatially neighboring or non-neighboring, by a block at the same position within the collocated picture based on a position of the current block.
[0215] Alternatively, by using at least one piece of motion vector information of a block that has been decoded, which is temporally-spatially neighboring or non-neighboring, a block positioned at a position spatially identical to the current block within a selected collocated picture can be defined as a collocated block of the current block by moving the corresponding motion vector.
[0216] Here, the motion information can include a motion vector, a reference picture index, an inter prediction indicator, a picture order count (POC), information of a collocated picture at a current coding picture (or slice) level.
[0217] Here, deriving the temporal merge candidate can mean deriving the temporal merge candidate and adding it to a merge candidate list. In addition, adding the temporal merge candidate to the merge candidate list can mean adding the corresponding temporal merge candidate to the merge candidate list when motion information existing in an existing merge candidate list and motion information of the newly derived temporal merge candidate are different.
[0218] When the decoding device adds a temporal merge candidate and there is a subblock-based temporal merge candidate to be described later in the existing merge candidate list, the decoding device can determine whether the motion information of the subblock-based temporal merge candidate is the same as that of the newly added temporal merge candidate. When the decoding device determines that there is no subblock-based temporal merge candidate having the same motion information as the newly added temporal merge candidate in the merge candidate list, the decoding device can add the corresponding temporal merge candidate to the merge candidate list.
[0219] In detail, the decoding device can compare the representative motion information of the subblock-based temporal merge candidate with the motion information of the newly added temporal merge candidate. Detailed embodiments of obtaining the representative motion information will be described below with reference to Figure 11
[0220] The decoding device can determine whether the inter prediction indicator of the representative motion information is the same as the value of the inter prediction indicator of the newly added temporal merge candidate. Here, when the values of the inter prediction indicators are not the same, the decoding device can add the temporal merge candidate to the merge candidate list. Meanwhile, when the values of the inter prediction indicators are the same, the decoding device can not add the newly added temporal merge candidate to the merge candidate list.
[0221] In another example, when the values of the inter prediction indicators are the same, the decoding device can additionally determine whether the motion vector or the reference picture index of the representative motion information is the same as that of the newly added temporal merge candidate. When at least one of the motion vector and the reference picture index is not the same, the decoding device can add the temporal merge candidate to the merge candidate list.
[0222] In another example, even when the values of the inter prediction indicators are the same, the decoding device can not add the temporal merge candidate to the merge candidate list when at least one piece of motion information of the L0 direction and the L1 direction is not the same.
[0223] On the other hand, when the temporal merge candidate is first included in the existing merge candidate list and the subblock-based temporal merge candidate is additionally added thereafter, the decoding device can determine whether to add the newly added subblock-based temporal merge candidate to the merge candidate list by using any one of the above-described methods. In other words, the decoding device can compare the motion information of the temporal merge candidate included in the existing merge candidate list with that of the newly added subblock-based temporal merge candidate and determine whether to add the newly added subblock-based temporal merge candidate to the merge candidate list.
[0224] The decoding device can derive up to maxNumTemporalMergeCand temporal merge candidates. Here, maxNumTemporalMergeCand can be a positive integer including 0.
[0225] In an example, maxNumTemporalMergeCand can be 1. However, the use of maxNumTemporalMergeCand does not limit the scope of the present invention. A decoding device can use the above information by using a parameter value having the same meaning as maxNumTemporalMergeCand.
[0226] Hereinafter, prediction through a temporal merge candidate is referred to as temporal motion vector prediction (TMVP) in the present specification.
[0227] Figure 8 is a diagram illustrating a method of deriving a temporal merge candidate. A decoding device can derive a temporal merge candidate by using a block at a position H, or by using a block at a position C3, wherein the position H is outside a collocated block C which is spatially identical to a current block X.
[0228] When a decoding device derives a temporal merge candidate from a block at a position H, the temporal merge candidate can be derived through a block at the position H, and when the decoding device does not derive a temporal merge candidate from a block at a position H, the temporal merge candidate can be derived through a block at a position C3.
[0229] Here, when a block at a position H or a block at a position C3 is encoded by using an intra prediction method, a decoding device cannot derive a temporal merge candidate. A collocated block can have a square shape or a non-square shape.
[0230] In another example, when a block associated with a block (a block at a position H or C3) is encoded by using an affine transform model mode (an affine mode) or a current picture reference mode (CPR mode), a decoding device cannot derive a temporal merge candidate of a current block from a corresponding collocated block.
[0231] When a distance between a picture including a current block and a reference picture of the current block is different from a distance between a picture including a collocated block and a reference picture of the collocated block, a decoding device can derive a temporal merge candidate by performing scaling on a motion vector of the collocated block. Scaling of the motion vector can be performed according to a ratio of td and tb (ratio = (tb / td)).
[0232] Here, td can denote a difference between a POC of a collocated picture and a POC of a reference picture of a collocated block, and tb can denote a difference between a POC of a picture to be decoded and a POC of a reference picture of a current block.
[0233] Hereinafter, a method of deriving a subblock-based temporal merge candidate by a decoding device will be described.
[0234] The decoding device can derive temporal merge candidates in units of sub-blocks from collocated sub-blocks. A sub-block is a block having a horizontal or vertical size smaller than that of the current block or having a depth deeper than that of the current block or having a shape smaller than that of the current block, and can be a block included in the current block.
[0235] The collocated sub-blocks of the sub-block to be decoded can have a square shape or a non-square shape. The decoding device can divide the collocated block of the current block in units of sub-blocks, and derive at least one temporal merge candidate for each sub-block.
[0236] When at least one temporal merge candidate is derived in units of sub-blocks, the temporal merge candidate can be derived from a collocated sub-block corresponding to H or C3 according to the shape or depth of the sub-block of the current block as shown in Figure 8 Alternatively, at least one temporal merge candidate can be derived from motion information stored in each sub-block unit of the collocated block associated with a position obtained by moving with motion information derived from neighboring blocks of the current block.
[0237] When deriving a temporal merge candidate of the current block or a sub-block of the current block, the decoding device can perform scaling on a motion vector of each reference picture list obtained from a collocated sub-block within a collocated block as a motion vector associated with an arbitrary reference picture of the current block.
[0238] The decoding device can obtain a plurality of motion vectors by performing scaling on a motion vector from a collocated sub-block to a motion vector associated with at least one reference picture among all reference pictures referable by a sub-block of the current block, obtain at least one prediction block using the scaled motion vector associated with each reference picture, and obtain a prediction block of the current block or the sub-block by a weighted sum of the at least one prediction block.
[0239] Hereinafter, prediction by a temporal merge candidate based on a sub-block is referred to as an alternative temporal motion vector prediction (ATMVP) in this specification.
[0240] Figure 9 is a diagram illustrating a method of deriving a spatio-temporal combination merge candidate based on a sub-block.
[0241] The decoding device can divide the current block into sub-blocks, and derive a merge candidate of the current block in units of sub-blocks by using at least one motion information of a neighboring spatial sub-block or a collocated sub-block within a collocated picture or both.
[0242] Figure 9 is a diagram illustrating a method of deriving a spatio-temporal combination merge candidate based on a sub-block by a decoding device. In Figure 9 a current block of which size is 8x8 in gray is divided into four 4x4 sub-blocks.
[0243] The decoding device can derive the subblock-based spatio-temporal combined merge candidate of the subblock being decoded by using the motion vector information of the spatio-temporal subblock of each subblock.
[0244] In Figure 9 , when the decoding device partitions the current block into subblocks and derives the residual signal according to the motion estimation, the decoding device can obtain the motion information by performing a scan from left to right based on the subblocks above the subblock A. For example, in Figure 9 , when the subblock above is encoded by an intra prediction method, the decoding device can sequentially scan the subblock b above.
[0245] The decoding device can perform a scan on the neighboring blocks above until the subblock above including available motion vector information is found. The decoding device can obtain the motion information of the available subblock above, and then obtain the motion information by performing a scan from top to bottom based on the subblock c left to the subblock A.
[0246] The decoding device can obtain the spatial neighboring motion information of at least one of the left subblock and the subblock above, and to derive the temporal motion information, the decoding device can obtain the motion information of at least one of the co-located subblock and the co-located block of the current subblock.
[0247] Here, the position of the co-located block or the subblock of the co-located block can be the motion information of the block at the position C3 or H described in Figure 8 , or can indicate the subblock of the co-located block at the position corrected by the motion vector derived adjacent to the current block or the co-located block at the corrected position.
[0248] By using the above method, the decoding device can obtain at least one of the motion information of at least one of the L0 spatial neighboring block and the L1 spatial neighboring block of the current block and the motion information of the temporal neighboring block, and derive the subblock-based spatio-temporal combined merge candidate of the subblock being decoded based on the at least one motion information.
[0249] In one embodiment, the decoding device can perform scaling on at least one motion vector derived in the spatio-temporal subblock describing the subblock of the current block for at least one of L0 and L1 to correspond to the first reference picture of the current block. Subsequently, the decoding device can derive the motion vector of the current subblock A or the spatio-temporal combined merge candidate of the subblock A by using at least one of the mean value, the maximum value, the minimum value, the median value, the weight value, the mode of up to three scaled motion vectors. By using the same method, the decoding device can derive the spatio-temporal combined merge candidate of the subblocks B, C, and D.
[0250] The decoding device can not partition the current block into as Figure 9The illustrated sub-blocks A, B, C, and D, and a merge candidate of the current block is derived by using at least one motion information of the neighboring spatial sub-blocks and the collocated sub-block within the collocated picture. For example, the decoding device can derive a spatio-temporal combination merge candidate of the current block by using the motion information of the upper sub-block b, the motion information of the left sub-block d, and the motion information of the collocated block.
[0251] Hereinafter, the prediction by the sub-block based spatio-temporal combination merge candidate is referred to as spatio-temporal motion vector prediction (STMVP) in the present specification.
[0252] Hereinafter, in addition to the above-described spatial merge candidate, temporal merge candidate, sub-block based temporal merge candidate, and sub-block based spatio-temporal combination merge candidate, additional merge candidates that can be applied to the present application will be described.
[0253] As the additional merge candidates that can be used in the present application, the decoding device can derive at least one of a modified spatial merge candidate, a modified temporal merge candidate, a combination merge candidate, and a merge candidate having a predetermined motion information value.
[0254] Here, deriving the additional merge candidate can mean that when there is a merge candidate having different motion information from the merge candidates existing in the existing merge candidate list, the corresponding merge candidate is added to the merge candidate list.
[0255] Here, the modified spatial merge candidate can mean a merge candidate obtained by modifying at least one motion information of the spatial merge candidate derived by using the above-described method.
[0256] The modified temporal merge candidate can mean a merge candidate obtained by modifying at least one motion information of the temporal merge candidate derived by using the above-described method.
[0257] Here, the merge candidate having a predetermined motion information value can mean a zero merge candidate whose motion vector is (0, 0). Hereinafter, the prediction by the zero merge candidate is referred to as zero motion prediction (ZMP) in the present specification.
[0258] The combination merge candidate can mean a merge candidate using motion information of at least one motion information of the spatial merge candidate, the temporal merge candidate, the modified spatial merge candidate, the modified temporal merge candidate, the combination merge candidate, and the merge candidate having a predetermined motion information value existing in the merge candidate list, and here, the combination merge candidate can mean a combination bi-prediction merge candidate.
[0259] Here, a combined merge candidate can be constructed for each list. For example, the decoding device can generate a new combined L0 list merge candidate by using an average of candidates present in the L0 list. In addition, the decoding device can generate a new combined L1 merge candidate by using an average of candidates present in the L1 list.
[0260] In addition, the decoding device can generate an L0 or L1 list merge candidate by using a candidate present in the L0 or L1 list.
[0261] For example, the decoding device can generate a new combined L0 list merge candidate by using an average of candidates generated by performing mirroring or scaling on the L0 merge candidate and the L1 merge candidate in the direction of the L0 list.
[0262] In another example, the decoding device can generate a new combined L1 list merge candidate by using an average of candidates generated by performing mirroring or scaling on the L1 merge candidate and the L0 merge candidate in the direction of the L1 list.
[0263] Hereinafter, prediction by a combined merge candidate will be referred to as combined motion prediction (CMP) in this specification.
[0264] The decoding device can derive at least one of the modified spatial merge candidate, the spatial merge candidate, the modified temporal merge candidate, the temporal merge candidate, the combined merge candidate, and the merge candidate having a predetermined motion information value for each sub-block, and add the merge candidate derived for each sub-block to the merge candidate list.
[0265] Figure 10 is a diagram illustrating a method of determining a merge candidate list according to an embodiment of the present application.
[0266] Hereinafter, a method of determining a merge candidate list according to an embodiment of the present application will be described. As described in S530 of FIG. 5, Figure 5 As described in S530 of FIG. 5, the decoding device can generate an initial merge candidate list of the current block. Subsequently, as described in S610 of FIG. 6, Figure 6 As described in S610 of FIG. 6, the decoding device can calculate distortion values of the merge candidates of the current block by using motion information of each of the merge candidates included in the initial merge candidate list.
[0267] The motion information can include at least one of an inter prediction indicator, an L0 reference or an L1 reference or both an L0 reference and an L1 reference image index, an L0 or an L1 or both an L0 and an L1 motion vector, a POC (Picture Order Count) and an LIC (Local Illumination Compensation) flag of a current coded picture or a reference picture or both a current coded picture and a reference picture, an affine flag, an OBMC (Overlapped Block Motion Compensation) flag, a reconstructed luma sample or a reconstructed chroma sample or both a reconstructed luma and chroma sample spatially neighboring the current block, and a luma sample or a chroma sample or both a luma and chroma sample of a reference picture indicated by the motion information of the merge candidate. However, the motion information of the present application is not limited thereto.
[0268] The initial merge candidate list can be configured with motion information of up to N merge candidates, and N can represent a positive integer greater than 0. Here, the spatio-temporal merge candidate can represent at least one of the above-described spatial merge candidate, the temporal merge candidate, the subblock-based temporal merge candidate, the subblock-based spatio-temporal combination merge candidate, the combination merge candidate, and the zero merge candidate.
[0269] To configure up to N merge candidates within the initial merge candidate list, the decoding device can fill the initial merge candidate list according to a preset order for the current block. Here, the decoding device can omit reconfiguring the additional merge candidate list, and determine the initial merge candidate list as the final merge candidate list of the current block.
[0270] When the decoding device adds a new merge candidate to the merge candidate list, when at least one piece of motion information of the newly added merge candidate is different from motion information of the merge candidate contained in the existing merge candidate list, the decoding device can add the new merge candidate to the merge candidate list.
[0271] In an example, assuming that up to seven spatio-temporal merge candidates are allowed to be added to the merge candidate list, up to seven initial merge candidate lists can be sequentially configured according to an arbitrary predetermined order as shown in Table 1 below. Here, the merge index can have a value from 0 to 6. The following example showing the order of adding the merge candidate list is an example of the present application, and the scope of the present application is not limited thereto.
[0272] [Table 1]
[0273]
[0274] Here, A1, B1, A0, B0, and B2 can represent the use of Figure 7 the spatial merge candidate described above.
[0275] For example, assuming that up to seven spatiotemporal merge candidates can be added to the merge candidate list, the decoding device can configure the merge candidate list using the number of merge candidates corresponding to the decoded merge index, thereby reducing computational load or complexity. Therefore, when merge indices are sent from 0 to 6, and the actual merge index decoded in the decoding device is 3, the decoding device can configure the initial merge candidate list by deriving up to four merge candidates.
[0276] In the example, within the module configuring the merge candidate list, the memory storing the motion information of the merge candidates on a sub-block basis can be initialized only before the actual derivation of merge candidates on a sub-block basis. Here, initialization can mean specifying the initial values for the motion vectors of the sub-block units, the inter-frame prediction indicators, and the reference image indices of L0 or L1, or both L0 and L1, on a sub-block basis.
[0277] When the decoding device configures an initial merge candidate list, and the number of spatial merge candidates included in the initial merge candidate list is less than a preset value K, the decoding device can omit calculating the distortion values of the merge candidates and determine the initial merge candidate list as the final merge candidate list for the current block. Here, K can represent any positive integer greater than 0.
[0278] The method for calculating distortion values by the decoding device will be described in detail below.
[0279] The decoding device can configure an initial list of merging candidates, and then calculate the distortion value between the current block and any merging candidate by using reconstructed luminance samples or reconstructed chrominance samples or both reconstructed luminance samples and reconstructed chrominance samples that are spatially adjacent to the current block (reconstructed samples adjacent to the current block), and luminance samples or chrominance samples or both luminance samples and chrominance samples that are spatially adjacent to a reference block of a reference image indicated by the motion information of each merging candidate (samples adjacent to the reference block).
[0280] The decoding device can calculate the distortion value by using at least one of SAD (sum of absolute differences), SATD (sum of absolute transform differences), and MR-SAD (sum of absolute differences with the mean removed) between the reconstructed samples adjacent to the current block and the samples adjacent to the reference block.
[0281] like Figure 10 As shown, at least one block having arbitrary size, shape, and depth and configured with at least one sample point adjacent to the current block can be defined as templates 1000 and 1005.
[0282] In addition, the decoding device can define at least one block having an arbitrary size, shape, and depth and configured with samples that are temporally motion estimated from a reference picture as a template 1010 and a template 1015 of a reference block by using at least one piece of motion information adjacent to an upper side of the current block or at least one piece of motion information adjacent to a left side of the current block or both of the at least one piece of motion information adjacent to the upper side of the current block and the at least one piece of motion information adjacent to the left side of the current block.
[0283] The decoding device can calculate a distortion value between the templates 1000 and 1005 of the current block and the templates 1010 and 1015 of the reference block indicated by the merge candidate. The decoding device can define the upper template 1000 (upper template) or the left template 1005 as a template for calculating a distortion value.
[0284] Figure 10 The width and the height described in the middle represent horizontal and vertical dimensions of the current block. Here, M and K can be positive integers greater than 0. M and K can have the same value or different values from each other. In addition, the width or the height or both of the width and the height can be set to have the same value as or different values from the width or the height or both of the width and the height of the current block.
[0285] Figure 10 is a diagram illustrating an embodiment in which the upper template has a value of width x M and the left template has a value of height x K.
[0286] In an example, a distortion value between the current block and a reference block within a reference picture indicated by motion information of an arbitrary merge candidate can be calculated by using at least one of SAD, SATD, and MRSAD between the template 1000 (template (current)) of the current block and the reference block template 1010 (template L0) or the reference block template 1015 (template L1) or both of the reference block template 1010 (template L0) and the reference block template 1015 (template L1).
[0287] The decoding device can use one of the template L0 and the template L1 as a template for calculating a distortion value, or can use both of the template L0 and the template L1. When both of the template L0 and the template L1 are used, the decoding device can calculate a distortion value by a weighted average of the template L0 and the template L1.
[0288] A distortion value between the current block templates 1000 and 1005 and the template L0 1010 can be defined as distortion (L0), a distortion value between the current block template and the template L1 can be defined as distortion (L1), and a distortion value between the weighted average of the current block template, the template L0, and the template L1 can be defined as distortion (Bi). Here, the distortion (L0) and the distortion (L1) can be defined as a first distortion value and a second distortion value. The template L0 and the template L1 can be defined as a first template and a second template, respectively.
[0289] When the arbitrary merge candidate includes only the L0 directional motion information, the decoding apparatus can calculate the distortion (L0), or can derive the L1 directional motion information by performing mirroring on the L0 directional motion information to calculate the distortion (L1) and the distortion (Bi). Here, the mirroring can be a symmetric operation performed on the value of the motion vector with respect to the origin.
[0290] For example, when the X and Y movement amounts of the L0 directional motion vector are (3, 5), the L1 vector value obtained by performing mirroring on the motion vector can be derived as (-3, -5).
[0291] In another example, after calculating all of the distortion (L0), the distortion (L1), and the distortion (Bi), the decoding apparatus can calculate a final distortion value of the current block by using at least one of a minimum value, a median value, and an average value.
[0292] In another example, when the decoding apparatus defines the distortion value between the current block and the merge candidate as a minimum value of the distortion (L0), the distortion (L1), and the distortion (Bi), the decoding apparatus can decode the current block by updating the motion information of the merge candidate list, wherein, when the distortion value in the distortion (L0) has the minimum value, the current merge candidate includes only the L0 motion information even though the merge candidate of the initial merge candidate list includes the bi-directional motion information.
[0293] When the initial merge candidate list is reconfigured, the decoding apparatus can reconfigure the initial merge candidate list based on information transmitted from the encoding apparatus without calculating the above-described distortion values. The information transmitted from the encoding apparatus can be indicator information indicating a list reconfiguration method preset in the encoding apparatus and the decoding apparatus, or can be an index indicating a preset list.
[0294] Figure 11 is a diagram illustrating a method of determining representative motion information according to an embodiment of the present application.
[0295] Referring to Figure 11 A case in which the arbitrary merge candidate includes at least one piece of motion information will be described. When the arbitrary merge candidate that constructs the initial merge candidate list is the ATMVP or the STMVP or both the ATMVP and the STMVP including at least one piece of motion information, the decoding apparatus can calculate the distortion value of the current block using the representative motion information by the same method as described above.
[0296] The representative motion information of the arbitrary merge candidate can be determined by the motion information at a preset position among the motion information of the sub-blocks having at least one piece of motion information different from each other, or can be derived by a weighted average value between the motion information of all of the sub-blocks.
[0297] In an example, when the size of the current block is a 32x32 block larger than 4x4 and the merge candidate is an ATMVP or an STMVP or both of ATMVP and STMVP in a 4x4 sub-block unit, to derive a template of a reference block for the current block, the decoding device can derive motion information of a first sub-block of the current block (e.g., a shaded area shown in Figure 11 (a)) as representative motion information of the merge candidate.
[0298] In another example, when the size of the current block is a 32x32 block larger than 4x4 and the merge candidate is an ATMVP or an STMVP or both of ATMVP and STMVP in a 4x4 sub-block unit, to derive a template of a reference block for the current block, the decoding device can derive motion information of a sub-block located at the center of the current block (e.g., a shaded area shown in Figure 11 (b)) as representative motion information of the merge candidate.
[0299] In another example, when the size of the current block is a 32x32 block larger than 4x4 and the merge candidate is an ATMVP or an STMVP or both of ATMVP and STMVP in a 4x4 sub-block unit, to derive a template of a reference block, the decoding device can derive representative motion information by using at least one of a mode, a median, and an average of sub-blocks different from each other.
[0300] When the decoding device calculates a distortion value, the decoding device can accurately calculate the distortion value by correcting (refining) a template of a reference block. As shown in Figure 10 , the decoding device can calculate a distortion (L0) by using an L0 direction reference block template, but can change a motion vector by applying an arbitrary offset to an L0 direction motion vector derived in a merge candidate, and then derive a template of a reference block.
[0301] In one example, when the offset is 1 and sizes of X and Y directions indicated by an L0 motion vector of an arbitrary merge candidate are (3, 4), in addition to a template of a reference block corresponding to (3, 4), the decoding device can derive at least one template of a reference block indicated by each motion vector in a cross form such as (2, 4), (4, 4), (3, 3), and (3, 5) by applying an offset of +1 and -1 to an X axis direction and applying an offset of +1 and -1 to a Y axis direction. Here, the decoding device can define a minimum value among values of distortion values calculated by using a plurality of templates as a distortion (L0). Here, when the distortion (L0) is determined at (3, 5), the L0 direction motion vector of the arbitrary merge candidate can be updated from (3, 4) to (3, 5) and be set in a merge candidate list.
[0302] Figure 12 is a diagram illustrating a method of calculating a distortion value according to an embodiment of the present application.
[0303] As shown in Figure 12 , when an arbitrary merge candidate of the current block includes bi-directional motion information, the decoding device can define a distortion value between L0 reference blocks and L1 reference blocks as a distortion value between the current block and the arbitrary merge candidate.
[0304] Figure 13 is a diagram illustrating a method of calculating a distortion value according to another embodiment of the present application.
[0305] As shown in Figure 13 , when an arbitrary merge candidate includes uni-directional (L0 or L1) motion information, the decoding device can define a distortion value between L0 reference blocks and L1 reference blocks obtained by performing mirroring on a motion vector of the uni-directional motion information as a distortion value between the current block and the arbitrary merge candidate.
[0306] Figure 14 is a diagram illustrating a method of calculating a distortion value according to another embodiment of the present application.
[0307] As shown in Figure 14 , when an arbitrary merge candidate includes bi-directional motion information, the decoding device can calculate a distortion (L0) between L0 reference blocks and L1 reference blocks after calculating reference blocks in the L1 direction by performing mirroring on a motion vector of L0. By using the same method, the decoding device can calculate a distortion (L1) between L0 reference blocks and L1 reference blocks after calculating reference blocks in the L0 direction by performing mirroring on a motion vector of L1. The decoding device can calculate a distortion value between the current block and the arbitrary merge candidate by using at least one of an average, a minimum value, and a maximum value of the distortion (L0) and the distortion (L1). For example, when the distortion (L0) has a minimum value, the decoding device can determine that the current merge candidate includes only L0 motion information, and perform decoding by updating motion information of the merge candidate list.
[0308] Referring to Figure 13 to Figure 15 , when calculating a distortion value between reference blocks, the decoding device can use motion vectors derived in a merge candidate as they are, or can change the motion vectors by applying an arbitrary offset to the initially derived motion vectors. The decoding device can define a minimum value among distortion values calculated by applying the offset as a final distortion value between the current block and the arbitrary merge candidate.
[0309] Hereinafter, a method of reconfiguring a merge candidate list by a decoding device will be described in detail.
[0310] As described in S620 of Figure 6 , the decoding device can reconfigure an initial merge candidate list by using distortion values calculated by using motion information of neighboring blocks.
[0311] The decoding device can calculate distortion values of all merge candidates of the initial merge candidate list, and then populate the merge candidate list from the merge candidates having small distortion values.
[0312] In an example, the decoding device can calculate distortion values for L arbitrary merge candidates of the initial merge candidate list, and then populate the merge candidate list from the merge candidates having small distortion values. When the maximum number of merge candidates included in the merge candidate list is P, L can be smaller than P.
[0313] The merge candidates used to calculate the distortion values can have merge indices from 0 to L-1 in the initial merge candidate list. For example, when L is 2, the merge candidate list can be reconfigured by calculating distortion values for two merge candidates populated in the initial merge candidate list.
[0314] For example, when the initial merge candidate list is determined to be configured in the order of (A1-B1-B0-A0-ATMVP-STMVP-B2), and L is 2 and the distortion value of the spatial merge candidate B1 is smaller than the distortion value of the spatial merge candidate A1, the decoding device can reconfigure the initial merge candidate list in the order of (B1-A1-B0-A0-ATMVP-STMVP-B2).
[0315] In another example, when reconfiguring the merge candidate list, the decoding device can reconfigure the merge candidate having the smallest calculated distortion value as the first merge candidate. When the order in which the initial merge candidate list is configured is determined to be (A1-B1-B0-A0-ATMVP-STMVP-B2), and in the initial merge candidate list, the distortion value of the spatial merge candidate B0 is smaller than the distortion value of the spatial merge candidate A1, the decoding device can reconfigure the initial merge candidate list to (B0-A1-B1-A0-ATMVP-STMVP-B2).
[0316] In another embodiment, the decoding device can reconfigure the merge candidate list of the current block using the order of the merge candidate list reconfigured in a neighboring block or in a higher level of the current block. Whether to use the order of the merge candidate list reconfigured in the neighboring block or in the higher level can be provided from the encoding device to the decoding device, or the decoding device can determine whether to use the order of the merge candidate list reconfigured in the neighboring block or in the higher level based on an encoding parameter.
[0317] The encoding device can determine whether to reconfigure the merge candidate list of the current block, and entropy encode information indicating whether to perform the reconfiguration. The encoding device can determine whether to perform the reconfiguration of the merge candidate list by comparing an RD cost before the merge candidate list reconfiguration method is applied and an RD cost after the merge candidate list reconfiguration method is performed.
[0318] The decoding apparatus can entropy-decode information indicating whether to perform reconfiguration of the merge candidate list from the bitstream, and reconfigure the merge candidate list according to the received information.
[0319] The encoding apparatus or the decoding apparatus can be set to perform entropy-encoding / decoding of the same information indicating whether to perform reconfiguration of the merge candidate list or omit the same information indicating whether to perform reconfiguration of the merge candidate list according to an encoding parameter of the current block.
[0320] For example, the encoding apparatus or the decoding apparatus can be set to perform or not to perform reconfiguration of the merge candidate list when a size of the current block is equal to or smaller than a predefined size, shape, and depth. On the other hand, the encoding apparatus or the decoding apparatus can be set to perform or not to perform reconfiguration of the merge candidate list when the size of the current block is equal to or greater than the predefined size, shape, and depth.
[0321] In another example, the encoding apparatus and the decoding apparatus can be set to not perform reconfiguration of the merge candidate list when the size of the current block is equal to or greater than or equal to or smaller than a predefined size and the current block is split by a binary tree or a quad tree.
[0322] In another example, the encoding apparatus or the decoding apparatus can be set to not perform reconfiguration of the merge candidate list when the size of the current block is equal to or greater than or smaller than or equal to a predefined size and the current block is split by a binary tree or a quad tree.
[0323] The decoding apparatus can determine whether to apply the merge candidate list reconfiguration method to the current target block according to flag information entropy-decoded in at least one of a picture / slice unit, a CTU unit, and a CTU lower level unit. Here, the lower level unit can include at least one of a CTU lower level unit, a quad tree unit, and a binary tree unit. In another example, the decoding apparatus can determine whether to perform the merge candidate list reconfiguration method according to a temporal layer of a current picture or slice to which the current block belongs.
[0324] Figure 15 FIG. 1 is a diagram illustrating a picture coding method according to an embodiment of the present application.
[0325] Referring to Figure 15 At S1500, the encoding apparatus can derive merge candidates of the current block. Subsequently, at S1510, the encoding apparatus can generate an initial merge candidate list of the merge candidates by using the derived merge candidate list. Subsequently, at S1520, the encoding apparatus can generate a reconfigured merge candidate list by using the initial merge candidate list.
[0326] The reconfiguring of the merge candidate list by the encoding apparatus using motion information of neighboring blocks to calculate a distortion value of a merge candidate and based on the distortion value corresponds to utilizing Figure 6 The operation of the described decoding apparatus will thus be omitted.
[0327] The above embodiments can be performed in the same method in the encoder and the decoder.
[0328] The order applied to the above embodiments can be different between the encoder and the decoder, or the order applied to the above embodiments can be the same in the encoder and the decoder.
[0329] The above embodiments can be performed for each of the luma signal and the chroma signal, or the above embodiments can be identically performed for the luma signal and the chroma signal.
[0330] The block shape to which the above embodiments of the present application are applied can have a square shape or a non-square shape.
[0331] The above embodiments of the present application can be applied in accordance with the size of at least one of the following blocks / units: an encoding block, a prediction block, a transform block, a block, a current block, an encoding unit, a prediction unit, a transform unit, a unit, and a current unit. Here, the size can be defined as a minimum size or a maximum size or both the minimum size and the maximum size, so that the above embodiments are applied, or the size can be defined as a fixed size to which the above embodiments are applied. Further, in the above embodiments, a first embodiment can be applied to a first size, and a second embodiment can be applied to a second size. In other words, the above embodiments can be applied in combination depending on the size. Further, when the size is equal to or greater than a minimum size and equal to or smaller than a maximum size, the above embodiments can be applied. In other words, when the block size is included in a certain range, the above embodiments can be applied.
[0332] For example, the above embodiments can be applied when the size of the current block is 8x8 or greater. For example, the above embodiments can be applied when the size of the current block is 4x4 or greater. For example, the above embodiments can be applied when the size of the current block is 16x16 or greater. For example, the above embodiments can be applied when the size of the current block is equal to or greater than 16x16 and equal to or smaller than 64x64.
[0333] The above embodiments of the present application can be applied in accordance with a temporal layer. In order to identify a temporal layer to which the above embodiments can be applied, a corresponding identifier can be signaled, and the above embodiments can be applied to a designated temporal layer identified by the corresponding identifier. Here, the identifier can be defined as a lowest layer or a highest layer or both the lowest layer and the highest layer to which the above embodiments can be applied, or can be defined as a specific layer indicating the application of the embodiments. Further, a fixed temporal layer to which the embodiments are applied can be defined.
[0334] For example, the above embodiment can be applied when the temporal layer of the current picture is the lowest layer. For example, the above embodiment can be applied when the temporal layer identifier of the current picture is 1. For example, the above embodiment can be applied when the temporal layer of the current picture is the highest layer.
[0335] A slice type to which the above embodiment of the present application is applied can be defined, and the above embodiment can be applied according to the corresponding slice type.
[0336] The above embodiment of the present application can also be applied when a motion vector has at least one of 16 pel units, 8 pel units, 4 pel units, integer pel units, 1 / 8 pel units, 1 / 16 pel units, 1 / 32 pel units, and 1 / 64 pel units. The motion vector can be selectively used for each pixel unit.
[0337] In the above embodiments, the method is described based on a flowchart having a series of steps or units, but the present application is not limited to the order of the steps, and some steps can be performed simultaneously with other steps or in a different order. In addition, it should be understood by one of ordinary skill in the art that the steps in the flowchart are not mutually exclusive, and other steps can be added to the flowchart or some of the steps can be deleted from the flowchart without affecting the scope of the present application.
[0338] The embodiments include various aspects of the examples. Not all possible combinations can be described, but one of ordinary skill in the art will be able to recognize different combinations. Accordingly, the present application can include all replacements, modifications, and changes within the scope of claims.
[0339] Embodiments of the present application can be implemented in the form of program instructions, which can be executed by various computer components and recorded in computer-readable recording media. The computer-readable recording media can include program instructions, data files, data structures, etc. alone or in combination. The program instructions recorded in the computer-readable recording media can be specifically designed and constructed for the present application, or well known to those skilled in the computer software technology field. Examples of the computer-readable recording media include magnetic recording media such as a hard disk, a floppy disk, and a magnetic tape, optical data storage media such as a CD-ROM or a DVD-ROM, a magneto-optical medium such as a floptical disk, and hardware devices specifically constructed to store and implement program instructions such as a read-only memory (ROM), a random access memory (RAM), a flash memory, etc. Examples of the program instructions include not only machine language codes formatted by a compiler, but also high-level language codes that can be implemented by a computer using an interpreter. The hardware devices can be configured to be operated by one or more software modules or vice versa to perform the processes according to the present application.
[0340] Although the present application has been described in terms of particular embodiments (such as detailed elements) and illustrative examples, they are merely provided to help a more complete understanding and are not intended to limit the present application to the embodiments described. Various modifications and changes can be made by persons of ordinary skill in the art to which the present application pertains without departing from the scope and spirit of the present application.
[0341] Accordingly, the spirit of the present application should not be limited to the above-described embodiments, and the entire scope of the claims and their equivalents will fall within the scope and spirit of the present application.
[0342] Industrial Applicability
[0343] The present application can be used to encode / decode an image.
Claims
1. An image decoding method performed by a decoding device, the method comprising: deriving a first merge candidate and a second merge candidate of a current block based on neighboring blocks of the current block; deriving a combined merge candidate of the current block based on a combination of the first merge candidate and the second merge candidate; generating a merge candidate list based on the combined merge candidate of the current block; and deriving a motion vector for the current block based on the merge candidate list, wherein based on the first merge candidate including an L0 motion vector and the second merge candidate including an L0 motion vector, an L0 motion vector of the combined merge candidate is derived based on an average of the L0 motion vector of the first merge candidate and the L0 motion vector of the second merge candidate, wherein based on the first merge candidate including an L0 motion vector and the second merge candidate including an L1 motion vector, an L0 motion vector of the combined merge candidate is derived based on the L0 motion vector of the first merge candidate, and wherein based on the first merge candidate including an L1 motion vector and the second merge candidate including an L0 motion vector, an L0 motion vector of the combined merge candidate is derived based on the L0 motion vector of the second merge candidate.
2. An image encoding method performed by an encoding device, the method comprising: deriving a first merge candidate and a second merge candidate of a current block based on neighboring blocks of the current block; deriving a combined merge candidate of the current block based on a combination of the first merge candidate and the second merge candidate; generating a merge candidate list based on the combined merge candidate of the current block; deriving motion vector information for the current block based on the merge candidate list; and encoding image information including prediction related information for the current block, wherein based on the first merge candidate including an L0 motion vector and the second merge candidate including an L0 motion vector, an L0 motion vector of the combined merge candidate is derived based on an average of the L0 motion vector of the first merge candidate and the L0 motion vector of the second merge candidate, wherein based on the first merge candidate including an L0 motion vector and the second merge candidate including an L1 motion vector, an L0 motion vector of the combined merge candidate is derived based on the L0 motion vector of the first merge candidate, and wherein based on the first merge candidate including an L1 motion vector and the second merge candidate including an L0 motion vector, an L0 motion vector of the combined merge candidate is derived based on the L0 motion vector of the second merge candidate.
3. A transmission method of image data, the method comprising: obtaining a bitstream of encoded picture information, wherein the encoded picture information is generated based on the following steps: deriving a first merge candidate and a second merge candidate of a current block based on neighboring blocks of the current block; deriving a combined merge candidate of the current block based on a combination of the first merge candidate and the second merge candidate; generating a merge candidate list based on the combined merge candidate of the current block; deriving a motion vector for the current block based on the merge candidate list; and encoding picture information including prediction-related information for the current block; and transmitting picture data including the bitstream, wherein, based on the first merge candidate including an L0 motion vector and the second merge candidate including an L0 motion vector, an L0 motion vector of the combined merge candidate is derived based on an average of the L0 motion vector of the first merge candidate and the L0 motion vector of the second merge candidate, wherein, based on the first merge candidate including an L0 motion vector and the second merge candidate including an L1 motion vector, an L0 motion vector of the combined merge candidate is derived based on the L0 motion vector of the first merge candidate, and wherein, based on the first merge candidate including an L1 motion vector and the second merge candidate including an L0 motion vector, an L0 motion vector of the combined merge candidate is derived based on the L0 motion vector of the second merge candidate.
Citation Information
Patent Citations
Method and device for sharing a candidate list
CN104185988A