Image coding / decoding method and device based on sub-block division
By adopting asymmetric sub-block partition structure and independent prediction methods, the problem of inefficiency in existing video encoding/decoding is solved, and a more efficient encoding/decoding effect is achieved.
Patent Information
- Application Number
- CN202510544119.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-25
- Filing Date
- 2020-06-16
- Publication Date
- 2025-07-15
AI Technical Summary
In the existing video encoding/decoding methods, the encoding/decoding blocks are usually square or rectangular, and the local characteristics in the video are not fully considered, resulting in low encoding/decoding efficiency.
Asymmetric sub-block partition structure is adopted, the current block is divided into the first sub-block and the second sub-block through a straight line, and its motion information is derived respectively. The prediction sample points of the current block are generated using weighted sum, and the binary tree after the quad-tree, combined quad-tree and binary tree block partition structure is supported to independently perform prediction.
Improves the efficiency of video encoding/decoding, supports independent prediction of asymmetric sub-blocks, and enhances the flexibility and efficiency of encoding/decoding.
Smart Images

Figure CN120321392A_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application for an invention titled "An Image Encoding / Decoding Method and Apparatus Based on Sub-Block Partitioning" with an application date of June 16, 2020, an application number of "202080036941.4". Technical Field
[0002] The present invention relates to a video encoding / decoding method, apparatus, and recording medium storing a bitstream. More specifically, the present invention relates to a video encoding / decoding method and apparatus based on at least one asymmetric sub-block. Background Art
[0003] Recently, in various applications, the demand for high-resolution and high-quality images (such as high-definition (HD) or ultra-high-definition (UHD) images) has increased. As the resolution and quality of images increase, the amount of data correspondingly increases. This is one of the reasons for the increase in transmission cost and storage cost when transmitting image data through existing transmission media (such as wired or wireless broadband channels) or when storing image data. To solve these problems of high-resolution and high-quality image data, efficient image encoding / decoding techniques are required.
[0004] There are various video compression techniques, such as an inter-frame prediction technique that predicts the value of a pixel in the current frame from the values of pixels in a previous frame or a subsequent frame, an intra-frame prediction technique that predicts the value of a pixel in another area of the current frame from the values of pixels in an area of the current frame, a transform and quantization technique that compresses the energy of a residual signal, and an entropy encoding technique that assigns short codes to frequently occurring pixel values and long codes to less frequently occurring pixel values.
[0005] In conventional video encoding / decoding methods and apparatuses, the encoding / decoding blocks always have a square shape or a rectangular shape or both a square shape and a rectangular shape, and are partitioned into a quadtree shape. Accordingly, encoding / decoding is performed while limitedly considering local characteristics within the video. Summary of the Invention
[0006] Technical Problem
[0007] An object of the present invention is to provide a method and apparatus for video encoding / decoding using various asymmetric sub-block partitioning structures.
[0008] In addition, another object of the present invention is to provide a method and apparatus for video encoding / decoding in which prediction is independently performed on asymmetric sub-blocks in the video encoding / decoding.
[0009] In addition, another object of the present invention is to provide a video encoding / decoding method and apparatus based on asymmetric sub-blocks in at least one of a block partitioning structure in which a quadtree is followed by a binary tree, a combined quadtree and binary tree block partitioning structure, and a separate PU / TU tree block partitioning structure, thereby improving encoding / decoding efficiency.
[0010] In addition, another object of the present invention is to provide a video encoding / decoding method and apparatus for partitioning a current block into at least one or more asymmetric sub-blocks or performing different predictions on each asymmetric sub-block.
[0011] In addition, another object of the present invention is to provide a method and apparatus for storing motion information of a current block that is asymmetrically subdivided.
[0012] In addition, another object of the present invention is to provide a recording medium for storing a bitstream generated by the video encoding / decoding method or apparatus of the present invention.
[0013] Technical Solution
[0014] A method for decoding an image according to the present invention may include: obtaining block partitioning information of a current block; partitioning the current block into a first sub-block and a second sub-block based on the block partitioning information; respectively deriving motion information of the first sub-block and motion information of the second sub-block; respectively generating prediction samples of the first sub-block and prediction samples of the second sub-block based on the motion information of the first sub-block and the motion information of the second sub-block; and generating prediction samples of the current block through a weighted sum of the prediction samples of the first sub-block and the prediction samples of the second sub-block, wherein the block partitioning information is index information indicating an index of a table, and wherein the table includes information indicating a plurality of predefined asymmetric partitioning shapes.
[0015] In the method for decoding an image according to the present invention, wherein the step of respectively deriving motion information of the first sub-block and motion information of the second sub-block includes: respectively obtaining a merge index of the first sub-block and a merge index of the second sub-block; generating a merge candidate list; deriving motion information of the first sub-block by using the merge candidate list and the merge index of the first sub-block; and deriving motion information of the second sub-block by using the merge candidate list and the merge index of the second sub-block.
[0016] In the method for decoding an image according to the present invention, wherein the merge candidate list is generated based on the current block.
[0017] In the method for decoding an image according to the present invention, wherein the step of partitioning the current block into a first sub-block and a second sub-block includes: partitioning the current block into the first sub-block and the second sub-block by a straight line.
[0018] In the method for decoding an image according to the present invention, information indicating the plurality of predefined asymmetric partition shapes includes at least one of angle information and distance information of the straight line.
[0019] In the method for decoding an image according to the present invention, the first sub-block and the second sub-block have any one of the shapes of a triangle, a rectangle, a trapezoid, and a pentagon.
[0020] In the method for decoding an image according to the present invention, the method further includes storing at least one of motion information of the first sub-block, motion information of the second sub-block, and third motion information, where when the motion information of the first sub-block and the motion information of the second sub-block refer to reference pictures in the same direction, the third motion information is derived as any one of the motion information of the first sub-block and the motion information of the second sub-block.
[0021] In the method for decoding an image according to the present invention, when the horizontal length and the vertical length of the current block are less than respective predetermined thresholds, the step of partitioning the current block into a first sub-block and a second sub-block based on the block partition information is not performed.
[0022] A method for encoding an image according to the present invention may include: determining a block partition structure of a current block; partitioning the current block into a first sub-block and a second sub-block based on the block partition structure; respectively deriving motion information of the first sub-block and motion information of the second sub-block; encoding block partition information based on the block partition structure; and encoding a merge index of the first sub-block and a merge index of the second sub-block respectively based on the motion information of the first sub-block and the motion information of the second sub-block, where the block partition information is index information indicating an index of a table, and the table includes information indicating a plurality of predefined asymmetric partition shapes.
[0023] In the method for encoding an image according to the present invention, the step of encoding the merge index of the first sub-block and the merge index of the second sub-block respectively includes: generating a merge candidate list; encoding the merge index of the first sub-block by using the merge candidate list and the motion information of the first sub-block; and encoding the merge index of the second sub-block by using the merge candidate list and the motion information of the second sub-block.
[0024] In the method for encoding an image according to the present invention, the merge candidate list is generated based on the current block.
[0025] In the method for encoding an image according to the present invention, the step of partitioning the current block into a first sub-block and a second sub-block includes partitioning the current block into a first sub-block and a second sub-block by a straight line.
[0026] In the method for encoding an image according to the present invention, information indicating the plurality of predefined asymmetric partition shapes includes at least one of angle information and distance information of the straight line.
[0027] In the method for encoding an image according to the present invention, the first sub-block and the second sub-block each have any one of a triangular shape, a rectangular shape, a trapezoidal shape, and a pentagonal shape.
[0028] The method for encoding an image according to the present invention further includes storing at least one of motion information of the first sub-block, motion information of the second sub-block, and third motion information. When the motion information of the first sub-block and the motion information of the second sub-block refer to reference pictures in the same direction, the third motion information is derived as any one of the motion information of the first sub-block and the motion information of the second sub-block.
[0029] In the method for encoding an image according to the present invention, when the horizontal length and the vertical length of the current block are less than respective predetermined thresholds, the step of partitioning the current block into a first sub-block and a second sub-block based on the block partition information is not performed.
[0030] A computer-readable recording medium according to the present invention can store a bitstream generated by the image encoding method according to the present invention.
[0031] Advantageous Effects
[0032] According to the present invention, a video encoding / decoding method and apparatus using various asymmetric sub-block partition structures can be provided.
[0033] In addition, according to the present invention, a video encoding / decoding method and apparatus for independently performing prediction on asymmetric sub-blocks of each partition can be provided.
[0034] In addition, according to the present invention, a recording medium for storing a bitstream generated by the video encoding / decoding method or apparatus of the present invention can be provided.
[0035] In addition, according to the present invention, video encoding and decoding efficiency can be improved. Description of the Drawings
[0036] Figure 1 is a block diagram showing the configuration of an encoding device according to an embodiment to which the present invention is applied.
[0037] Figure 2 is a block diagram showing the configuration of a decoding device according to an embodiment to which the present invention is applied.
[0038] Figure 3 is a diagram schematically showing a partition structure of an image when the image is encoded and decoded.
[0039] Figure 4 It is a diagram showing intra prediction processing.
[0040] Figure 5 It is a diagram showing an example of inter picture prediction processing.
[0041] Figure 6 It is a diagram showing transform and quantization processing.
[0042] Figure 7 It is a diagram showing reference sample points that can be used for intra prediction.
[0043] Figure 8 It is a diagram showing a method for deriving spatial merge candidates according to an embodiment of the present invention.
[0044] Figure 9 It is a diagram showing a method for deriving temporal merge candidates according to an embodiment of the present invention.
[0045] Figure 10 It is a diagram showing a method for deriving sub-block based spatio-temporal combination merge candidates according to an embodiment of the present invention.
[0046] Figure 11 It is a diagram showing a method for deriving inter prediction information by using a bidirectional matching method according to an embodiment of the present invention.
[0047] Figure 12 It is a diagram showing a method for deriving inter prediction information by using a template matching method according to an embodiment of the present invention.
[0048] Figure 13 It is a diagram showing a method for deriving inter prediction information based on overlapping block motion compensation (OMBC) according to an embodiment of the present invention.
[0049] Figure 14 It is a diagram showing quadtree partitioning, symmetric binary tree partitioning, and asymmetric binary tree partitioning according to an embodiment of the present invention.
[0050] Figure 15 It is a diagram showing symmetric binary tree partitioning after quadtree partitioning according to an embodiment of the present invention.
[0051] Figure 16 It is a diagram showing asymmetric partitioning according to an embodiment of the present invention.
[0052] Figure 17 It is a diagram showing a method for deriving motion prediction information of a sub-block by using the lowest level sub-block according to an embodiment of the present invention.
[0053] Figures 18 to 21It is a diagram showing the type of motion information for storing each block according to an embodiment of the present invention.
[0054] Figure 22 It is a flowchart showing a video decoding method according to an embodiment of the present invention.
[0055] Figure 23 It is a flowchart showing a video encoding method according to an embodiment of the present invention. Detailed implementation
[0056] Various modifications can be made to the present invention, and there are various embodiments of the present invention. Herein, examples of various embodiments of the present invention will now be provided and described in detail with reference to the accompanying drawings. However, the present invention is not limited thereto. Although the exemplary embodiments may be interpreted as including all modifications, equivalents, or alternatives within the technical concept and technical scope of the present invention. In all aspects, like reference numerals refer to the same or similar functions. In the drawings, for clarity, the shapes and sizes of elements may be exaggerated. In the following detailed description of the present invention, reference is made to the accompanying drawings, in which specific embodiments in which the present invention may be practiced are shown in a schematic manner. These embodiments are described in sufficient detail to enable those skilled in the art to practice the present disclosure. It should be understood that the various embodiments of the present disclosure, although different, are not necessarily mutually exclusive. For example, specific features, structures, and characteristics described herein in connection with one embodiment may be implemented in other embodiments without departing from the spirit and scope of the present disclosure. Additionally, it should be understood that the positions or arrangements of the individual elements within each disclosed embodiment may be modified without departing from the spirit and scope of the present disclosure. Accordingly, the following detailed description should not be considered limiting, and the scope of the present disclosure is defined only by the appended claims (interpreted appropriately, together with the full scope of equivalents claimed by the claims).
[0057] The terms "first", "second", etc. used in the specification may be used to describe various components, but the components should not be construed as limited to these terms. These terms are only used to distinguish one component from other components. For example, without departing from the scope of the present invention, the "first" component may be named the "second" component, and the "second" component may be similarly named the "first" component. The term "and / or" includes combinations of multiple items or any one of the multiple items.
[0058] It will be understood that in this specification, when an element is simply referred to as "connected to" or "coupled to" another element rather than "directly connected to" or "directly coupled to" another element, the element may be "directly connected to" or "directly coupled to" another element, or may be connected to or coupled to another element in the case where other elements are interposed between the element and the other element. Conversely, it should be understood that when an element is referred to as being "directly coupled" or "directly connected" to another element, there is no intermediate element.
[0059] In addition, the components shown in the embodiments of the present invention are independently shown to represent different characteristic functions from each other. Therefore, this does not mean that each component is constituted by a separate hardware or software component unit. In other words, for convenience, each component includes each of the listed components. Therefore, at least two components of each component may be combined to form one component, or one component may be partitioned into multiple components to perform each function. Embodiments in which each component is combined and embodiments in which one component is partitioned are also included in the scope of the present invention without departing from the essence of the present invention.
[0060] The terms used in this specification are only used to describe specific embodiments and are not intended to limit the present invention. Unless having a significantly different meaning in the context, expressions used in the singular form include expressions in the plural form. In this specification, it will be understood that terms such as "including", "having", etc. are intended to indicate the presence of the features, numbers, steps, actions, elements, components or combinations thereof disclosed in the specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, actions, elements, components or combinations thereof may exist or may be added. In other words, when a specific element is referred to as being "included", elements other than the corresponding element are not excluded, but additional elements may be included in the embodiments of the present invention or within the scope of the present invention.
[0061] In addition, some components may not be essential components for performing the basic functions of the present invention, but are only optional components for improving its performance. The present invention can be implemented by only including the essential components for realizing the essence of the present invention without including the components for improving performance. Structures that only include the essential components and do not include the optional components for only improving performance are also included in the scope of the present invention.
[0062] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. When describing the exemplary embodiments of the present invention, well-known functions or configurations will not be described in detail because they may unnecessarily obscure the understanding of the present invention. The same components in the drawings are denoted by the same reference numerals, and repeated descriptions of the same components will be omitted.
[0063] Hereinafter, an image may refer to a frame constituting a video, or may refer to the video itself. For example, "encoding or decoding an image, or both encoding and decoding" may refer to "encoding or decoding a moving picture, or both encoding and decoding", and may refer to "encoding or decoding, or both encoding and decoding, one image among the images of a moving picture".
[0064] Hereinafter, the terms "moving picture" and "video" may be used with the same meaning and may be interchangeable with each other.
[0065] Hereinafter, a target image may be an encoding target image as an encoding target and / or a decoding target image as a decoding target. Additionally, the target image may be an input image input to an encoding device and an input image input to a decoding device. Here, the target image may have the same meaning as the current frame.
[0066] Hereinafter, the terms "image", "frame", "picture", and "screen" may be used with the same meaning and may be interchangeable with each other.
[0067] Hereinafter, a target block may be an encoding target block as an encoding target and / or a decoding target block as a decoding target. Additionally, the target block may be a current block as a target of current encoding and / or decoding. For example, the terms "target block" and "current block" may be used with the same meaning and may be interchangeable with each other.
[0068] Hereinafter, the terms "block" and "unit" may be used with the same meaning and may be interchangeable with each other. Or a "block" may represent a specific unit.
[0069] Hereinafter, the terms "region" and "segment" may be interchangeable with each other.
[0070] Hereinafter, a specific signal may be a signal representing a specific block. For example, an original signal may be a signal representing a target block. A prediction signal may be a signal representing a prediction block. A residual signal may be a signal representing a residual block.
[0071] In an embodiment, each of specific information, data, flag, index, element, and attribute, etc. may have a value. A value of information, data, flag, index, element, and attribute equal to "0" may represent logical false or a first predefined value. In other words, the values "0", false, logical false, and the first predefined value may be interchangeable with each other. A value of information, data, flag, index, element, and attribute equal to "1" may represent logical true or a second predefined value. In other words, the values "1", true, logical true, and the second predefined value may be interchangeable with each other.
[0072] When the variable i or j is used to represent a column, row, or index, the value of i can be an integer equal to or greater than 0, or an integer equal to or greater than 1. That is, columns, rows, indices, etc. can be counted starting from 0, or can be counted starting from 1.
[0073] Term description
[0074] Encoder: Represents a device that performs encoding. That is, it represents an encoding device.
[0075] Decoder: Represents a device that performs decoding. That is, it represents a decoding device.
[0076] Block: Is an array of samples of M×N. Here, M and N can represent positive integers, and the block can represent an array of samples in a two-dimensional shape. A block can refer to a unit. The current block can represent an encoding target block that becomes the target during encoding, or a decoding target block that becomes the target during decoding. Additionally, the current block can be at least one of an encoding block, a prediction block, a residual block, and a transform block.
[0077] Sample: Is the basic unit that constitutes a block. According to the bit depth (Bd), the sample can be represented as a value from 0 to 2 Bd -1. In the present invention, the sample can be used in the meaning of a pixel. That is, a sample, a pel, and a pixel can have the same meaning as each other.
[0078] Unit: Can refer to an encoding and decoding unit. When encoding and decoding an image, the unit can be a region generated by partitioning a single image. Additionally, when a single image is partitioned into sub-partition units during encoding or decoding, the unit can represent a sub-partition unit. That is, an image can be partitioned into multiple units. When encoding and decoding an image, predetermined processing can be performed for each unit. A single unit can be partitioned into sub-units with a size smaller than the size of the unit. Depending on the function, the unit can represent a block, a macroblock, a coding tree unit, a coding tree block, an encoding unit, an encoding block, a prediction unit, a prediction block, a residual unit, a residual block, a transform unit, a transform block, etc. Additionally, in order to distinguish the unit from the block, the unit can include a luminance component block, a chrominance component block associated with the luminance component block, and syntax elements of each color component block. The unit can have various sizes and shapes. Specifically, the shape of the unit can be a two-dimensional geometric figure, such as a square, a rectangle, a trapezoid, a triangle, a pentagon, etc. Additionally, the unit information can include at least one of the unit type indicating an encoding unit, a prediction unit, a transform unit, etc., and the unit size, the unit depth, the order of encoding and decoding of the unit, etc.
[0079] Coding tree unit: A single coding tree block configured with the luminance component Y and two coding tree blocks related to the chrominance components Cb and Cr. Additionally, a coding tree unit may represent a block and the syntax elements of each block. Each coding tree unit can be partitioned by using at least one of a quadtree partitioning method, a binary tree partitioning method, and a ternary tree partitioning method to configure lower-level units such as coding units, prediction units, transform units, etc. A coding tree unit can be used as a term for specifying a sample block that becomes a processing unit when encoding / decoding an image as an input image. Here, a quadtree may represent a quadtree.
[0080] When the size of a coding block is within a predetermined range, partitioning can be performed using only quadtree partitioning. Here, the predetermined range can be defined as at least one of the maximum size and the minimum size of a coding block that can be partitioned using only quadtree partitioning. Information indicating the maximum / minimum size of a coding block allowing quadtree partitioning can be signaled by a bitstream, and the information can be signaled in at least one of a sequence, a picture parameter, a parallel block group, or a slice. Optionally, the maximum / minimum size of a coding block can be a fixed size predetermined in an encoder / decoder. For example, when the size of a coding block corresponds to 256×256 to 64×64, partitioning can be performed using only quadtree partitioning. Optionally, when the size of a coding block is larger than the size of the maximum transform block, partitioning can be performed using only quadtree partitioning. Here, the block to be partitioned can be at least one of a coding block and a transform block. In this case, information indicating the partitioning of a coding block (e.g., split_flag) can be a flag indicating whether quadtree partitioning is performed. When the size of a coding block falls within the predetermined range, partitioning can be performed using only binary tree or ternary tree partitioning. In this case, the above description of quadtree partitioning can be applied to binary tree partitioning or ternary tree partitioning in the same manner.
[0081] Coding tree block: A term that can be used to specify any one of a Y coding tree block, a Cb coding tree block, and a Cr coding tree block.
[0082] Neighboring block: A block that can represent a block adjacent to the current block. A block adjacent to the current block can represent a block that touches the boundary of the current block or a block located within a predetermined distance from the current block. A neighboring block can represent a block adjacent to the vertex of the current block. Here, a block adjacent to the vertex of the current block can represent a block that is vertically adjacent to a horizontally adjacent neighboring block of the current block or a block that is horizontally adjacent to a vertically adjacent neighboring block of the current block.
[0083] Reconstructed neighboring blocks: Can represent neighboring blocks that are adjacent to the current block and have been encoded or decoded in space / time. Here, the reconstructed neighboring blocks can represent reconstructed neighboring units. The reconstructed spatial neighboring blocks can be blocks within the current picture that have been reconstructed by encoding or decoding or both. The reconstructed temporal neighboring blocks are blocks at positions corresponding to the current block of the current picture in a reference picture or neighboring blocks of said blocks.
[0084] Unit depth: Can represent the degree of partitioning of a unit. In a tree structure, the highest node (root node) can correspond to a first unit that is not partitioned. Additionally, the highest node can have the smallest depth value. In this case, the depth of the highest node can be level 0. A node with a depth of level 1 can represent a unit generated by a first partitioning of the first unit. A node with a depth of level 2 can represent a unit generated by a second partitioning of the first unit. A node with a depth of level n can represent a unit generated by an nth partitioning of the first unit. A leaf node can be the lowest node and is a node that cannot be further partitioned. The depth of a leaf node can be the maximum level. For example, a predefined value for the maximum level can be 3. The depth of the root node can be the lowest, and the depth of the leaf node can be the deepest. Additionally, when a unit is represented as a tree structure, the level in which the unit exists can represent the unit depth.
[0085] Bitstream: Can represent a bitstream including encoded image information.
[0086] Parameter set: Corresponds to the header information among the configurations within the bitstream. At least one of a video parameter set, a sequence parameter set, a picture parameter set, and an adaptive parameter set can be included in the parameter set. In addition, the parameter set can include slice headers, tile group headers, and tile header information. The term "tile group" represents a group of tiles and has the same meaning as a slice.
[0087] The adaptive parameter set can represent a parameter set that can be shared by referring to different pictures, sub - pictures, slices, tile groups, tiles, or blocks. Additionally, the information in the adaptive parameter set can be used by referring to different adaptive parameter sets for sub - pictures, slices, tile groups, tiles, or blocks within a picture.
[0088] In addition, regarding the adaptive parameter set, different adaptive parameter sets can be referred to by using the identifiers of different adaptive parameter sets for sub - pictures, slices, tile groups, tiles, or blocks within a picture.
[0089] In addition, regarding the adaptive parameter set, different adaptive parameter sets can be referred to by using the identifiers of different adaptive parameter sets for slices, tile groups, tiles, or blocks within a sub - picture.
[0090] In addition, regarding the adaptive parameter set, different adaptive parameter sets can be referred to by using identifiers of different adaptive parameter sets for parallel blocks or chunks within a slice.
[0091] In addition, regarding the adaptive parameter set, different adaptive parameter sets can be referred to by using identifiers of different adaptive parameter sets for chunks within a parallel block.
[0092] Information on the adaptive parameter set identifier can be included in the parameter set or header of a sub-picture, and the adaptive parameter set corresponding to the adaptive parameter set identifier can be used for the sub-picture.
[0093] Information on the adaptive parameter set identifier can be included in the parameter set or header of a parallel block, and the adaptive parameter set corresponding to the adaptive parameter set identifier can be used for the parallel block.
[0094] Information on the adaptive parameter set identifier can be included in the header of a chunk, and the adaptive parameter set corresponding to the adaptive parameter set identifier can be used for the chunk.
[0095] A picture can be partitioned into one or more parallel block rows and one or more parallel block columns.
[0096] A sub-picture can be partitioned into one or more parallel block rows and one or more parallel block columns within the picture. The sub-picture can be a region with a rectangular / square shape within the picture and can include one or more CTUs. Additionally, at least one or more parallel blocks / chunks / slices can be included within one sub-picture.
[0097] A parallel block can be a region with a rectangular / square shape within the picture and can include one or more CTUs. Additionally, a parallel block can be partitioned into one or more chunks.
[0098] A chunk can represent one or more CTU rows within a parallel block. A parallel block can be partitioned into one or more blocks, and each block can have at least one or more CTU rows. A parallel block that is not partitioned into two or more can represent a chunk.
[0099] A slice can include one or more parallel blocks within the picture and can include one or more chunks within the parallel block.
[0100] Parsing: can represent determining the value of a syntax element by performing entropy decoding, or can represent entropy decoding itself.
[0101] Symbol: can represent at least one of a syntax element, coding parameter, and transform coefficient value of an encoding / decoding target unit. Additionally, a symbol can represent an entropy coding target or an entropy decoding result.
[0102] Prediction mode: It can be information indicating a mode encoded / decoded using intra prediction or a mode encoded / decoded using inter prediction.
[0103] Prediction unit: It can represent a basic unit when performing prediction (such as inter prediction, intra prediction, inter compensation, intra compensation, and motion compensation). A single prediction unit can be partitioned into multiple partitions with smaller sizes, or can be partitioned into multiple lower-level prediction units. Multiple partitions can be basic units when performing prediction or compensation. The partitions generated by partitioning the prediction unit can also be prediction units.
[0104] Prediction unit partition: It can represent a shape obtained by partitioning a prediction unit.
[0105] The reference picture list can refer to a list including one or more reference pictures for inter prediction or motion compensation. There are several types of available reference picture lists, including LC (List Combination), L0 (List 0), L1 (List 1), L2 (List 2), L3 (List 3).
[0106] The inter prediction indicator can refer to the direction of inter prediction of the current block (unidirectional prediction, bidirectional prediction, etc.). Optionally, the inter prediction indicator can refer to the number of reference pictures used to generate the prediction block of the current block. Optionally, the inter prediction indicator can refer to the number of prediction blocks used when performing inter prediction or motion compensation on the current block.
[0107] The prediction list utilization flag indicates whether at least one reference picture in a specific reference picture list is used to generate a prediction block. The prediction list utilization flag can be used to derive the inter prediction indicator, and conversely, the inter prediction indicator can be used to derive the prediction list utilization flag. For example, when the prediction list utilization flag has a first value of zero (0), it indicates that the reference pictures in the reference picture list are not used to generate a prediction block. On the other hand, when the prediction list utilization flag has a second value of one (1), it indicates that the reference picture list is used to generate a prediction block.
[0108] The reference picture index can refer to an index indicating a specific reference picture in the reference picture list.
[0109] The reference picture can represent a reference picture referred to by a specific block for the purpose of inter prediction or motion compensation of the specific block. Optionally, the reference picture can be a picture including reference blocks referred to by the current block for inter prediction or motion compensation. Hereinafter, the terms "reference picture" and "reference picture" have the same meaning and can be interchanged.
[0110] A motion vector can be a two-dimensional vector for inter-frame prediction or motion compensation. The motion vector can represent the offset between an encoded / decoded target block and a reference block. For example, (mvX, mvY) can represent the motion vector. Here, mvX can represent the horizontal component, and mvY can represent the vertical component.
[0111] The search range can be a two-dimensional region that is searched during inter-frame prediction to retrieve the motion vector. For example, the size of the search range can be M×N. Here, both M and N are integers.
[0112] A motion vector candidate can refer to a predicted candidate block or the motion vector of the predicted candidate block when predicting the motion vector. Additionally, the motion vector candidate can be included in a motion vector candidate list.
[0113] The motion vector candidate list can represent a list composed of one or more motion vector candidates.
[0114] The motion vector candidate index can represent an indicator that indicates the motion vector candidate in the motion vector candidate list. Optionally, it can be an index of a motion vector predictor.
[0115] Motion information can represent information including at least one of a motion vector, a reference picture index, an inter-frame prediction indicator, a prediction list utilization flag, reference picture list information, a reference picture, a motion vector candidate, a motion vector candidate index, a merge candidate, and a merge index.
[0116] The merge candidate list can represent a list composed of one or more merge candidates.
[0117] A merge candidate can represent a spatial merge candidate, a temporal merge candidate, a combined merge candidate, a combined bi-prediction merge candidate, or a zero merge candidate. The merge candidate can include motion information such as an inter-frame prediction indicator, a reference picture index for each list, a motion vector, a prediction list utilization flag, and an inter-frame prediction indicator.
[0118] The merge index can represent an indicator that indicates the merge candidate in the merge candidate list. Optionally, the merge index can indicate a block in a reconstructed block that is spatially / temporally adjacent to the current block, from which the merge candidate has been derived. Optionally, the merge index can indicate at least one motion information of the merge candidate.
[0119] The transform unit: can represent the basic unit when performing encoding / decoding (such as transformation, inverse transformation, quantization, dequantization, transform coefficient encoding / decoding) on the residual signal. A single transform unit can be partitioned into multiple lower-level transform units with smaller sizes. Here, the transform / inverse transform can include at least one of a first transform / first inverse transform and a second-level transform / second inverse transform.
[0120] Scaling: It can represent a process of multiplying the quantization levels by a factor. Transformation coefficients can be generated by scaling the quantization levels. Scaling can also be referred to as inverse quantization.
[0121] Quantization parameter: It can represent a value used when transformation coefficients are used to generate quantization levels during quantization. The quantization parameter can also represent a value used when transformation coefficients are generated by scaling the quantization levels during inverse quantization. The quantization parameter can be a value mapped on the quantization step.
[0122] Delta quantization parameter: It can represent the difference between the predicted quantization parameter and the quantization parameter of the coding / decoding target unit.
[0123] Scanning: It can represent a method of sorting coefficients within a unit, block, or matrix. For example, converting a two-dimensional matrix of coefficients into a one-dimensional matrix can be called scanning, and converting a one-dimensional matrix of coefficients into a two-dimensional matrix can be called scanning or inverse scanning.
[0124] Transformation coefficient: It can represent the coefficient value generated after performing a transformation in an encoder. The transformation coefficient can represent the coefficient value generated after performing at least one of entropy decoding and inverse quantization in a decoder. The quantization levels or quantization transformation coefficient levels obtained by quantizing the transformation coefficients or residual signals can also fall within the meaning of transformation coefficients.
[0125] Quantization level: It can represent the value generated by quantizing the transformation coefficients or residual signals in an encoder. Optionally, the quantization level can represent the value of the inverse quantization target that undergoes inverse quantization in a decoder. Similarly, the quantization transformation coefficient levels as a result of transformation and quantization can also fall within the meaning of quantization levels.
[0126] Non-zero transformation coefficient: It can represent a transformation coefficient with a value other than zero, or a transformation coefficient level or quantization level with a value other than zero.
[0127] Quantization matrix: It can represent a matrix used in the quantization process or inverse quantization process performed to improve subjective image quality or objective image quality. The quantization matrix can also be referred to as a scaling list.
[0128] Quantization matrix coefficient: It can represent each element within the quantization matrix. The quantization matrix coefficient can also be referred to as a matrix coefficient.
[0129] Default matrix: It can represent a predefined quantization matrix in an encoder or decoder.
[0130] Non-default matrix: It can represent a quantization matrix that is not predefined in an encoder or decoder but signaled by a user.
[0131] Statistical value: The statistical value for at least one of variables, coded parameters, constant values, etc. having computable specific values may be one or more of the average value, sum value, weighted average value, weighted sum value, minimum value, maximum value, most frequently occurring value, median value, interpolation of the corresponding specific values.
[0132] Figure 1 is a block diagram showing the configuration of an encoding device according to an embodiment to which the present invention is applied.
[0133] The encoding device 100 may be an encoder, a video encoding device, or an image encoding device. The video may include at least one image. The encoding device 100 may sequentially encode at least one image.
[0134] Refer to Figure 1 , the encoding device 100 may include a motion prediction unit 111, a motion compensation unit 112, an intra prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.
[0135] The encoding device 100 may perform encoding of an input image by using an intra mode or an inter mode or both the intra mode and the inter mode. In addition, the encoding device 100 may generate a bitstream including encoding information by encoding the input image and output the generated bitstream. The generated bitstream may be stored in a computer-readable recording medium or may be streamed through a wired / wireless transmission medium. When the intra mode is used as a prediction mode, the switch 115 may switch to intra. Optionally, when the inter mode is used as a prediction mode, the switch 115 may switch to the inter mode. Here, the intra mode may represent an intra prediction mode, and the inter mode may represent an inter prediction mode. The encoding device 100 may generate a prediction block for an input block of the input image. In addition, the encoding device 100 may encode a residual block by using the residual of the input block and the prediction block after generating the prediction block. The input image may be referred to as the current picture that is the current encoding target. The input block may be referred to as the current block that is the current encoding target or as the encoding target block.
[0136] When the prediction mode is the intra mode, the intra prediction unit 120 may use the samples of the blocks that have been encoded / decoded and are adjacent to the current block as reference samples. The intra prediction unit 120 may perform spatial prediction on the current block by using the reference samples or may generate prediction samples of the input block by performing spatial prediction. Here, the intra prediction may represent the prediction inside the frame.
[0137] When the prediction mode is an inter-frame mode, the motion prediction unit 111 may retrieve, during motion prediction, a region that best matches the input block from a reference picture, and derive a motion vector by using the retrieved region. In this case, the search region may be used as the region. The reference picture may be stored in the reference picture buffer 190. Here, when encoding / decoding of the reference picture is performed, the reference picture may be stored in the reference picture buffer 190.
[0138] The motion compensation unit 112 may perform motion compensation on the current block by using the motion vector to generate a prediction block. Here, inter-frame prediction may represent prediction or motion compensation between frames.
[0139] When the value of the motion vector is not an integer, the motion prediction unit 111 and the motion compensation unit 112 may generate a prediction block by applying an interpolation filter to a partial region of the reference picture. To perform inter-picture prediction or motion compensation on an encoding unit, it may be determined which one of a skip mode, a merge mode, an advanced motion vector prediction (AMVP) mode, and a current picture reference mode is used for motion prediction and motion compensation of a prediction unit included in the corresponding encoding unit. Then, inter-picture prediction or motion compensation may be performed differently according to the determined mode.
[0140] The subtractor 125 may generate a residual block by using the difference between the input block and the prediction block. The residual block may be referred to as a residual signal. The residual signal may represent the difference between the original signal and the prediction signal. In addition, the residual signal may be a signal generated by transforming or quantizing or both transforming and quantizing the difference between the original signal and the prediction signal. The residual block may be the residual signal of a block unit.
[0141] The transform unit 130 may generate transform coefficients by performing a transform on the residual block, and output the generated transform coefficients. Here, the transform coefficients may be coefficient values generated by performing a transform on the residual block. When the transform skip mode is applied, the transform unit 130 may skip the transform of the residual block.
[0142] Quantized levels may be generated by applying quantization to the transform coefficients or to the residual signal. Hereinafter, the quantized levels may also be referred to as transform coefficients in the embodiments.
[0143] The quantization unit 140 may generate quantized levels by quantizing the transform coefficients or the residual signal according to parameters, and output the generated quantized levels. Here, the quantization unit 140 may quantize the transform coefficients by using a quantization matrix.
[0144] The entropy encoding unit 150 may generate a bitstream by performing entropy encoding on the values calculated by the quantization unit 140 or on the encoding parameter values calculated during encoding according to a probability distribution, and output the generated bitstream. The entropy encoding unit 150 may perform entropy encoding on the sample information of the image and the information for decoding the image. For example, the information for decoding the image may include syntax elements.
[0145] When entropy encoding is applied, symbols are represented such that a smaller number of bits are assigned to symbols with a high generation probability, and a larger number of bits are assigned to symbols with a low generation probability. Thus, the size of the bitstream of the symbols to be encoded can be reduced. The entropy encoding unit 150 may use encoding methods for entropy encoding such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. For example, the entropy encoding unit 150 may perform entropy encoding by using a variable length coding / code (VLC) table. In addition, the entropy encoding unit 150 may derive a binarization method for a target symbol and a probability model of the target symbol / binary bits, and perform arithmetic encoding by using the derived binarization method and context model.
[0146] To encode the transform coefficient levels (quantized levels), the entropy encoding unit 150 may change the coefficients in a two-dimensional block shape into a one-dimensional vector shape by using a transform coefficient scanning method.
[0147] Coding parameters may include information such as syntax elements (flags, indices, etc.) that are encoded in an encoder and signaled to a decoder, as well as information derived during encoding or decoding. The coding parameters may represent the information required for encoding or decoding an image. For example, at least one value or combination of the following items may be included in the coding parameters: unit / block size, unit / block depth, unit / block partitioning information, unit / block shape, unit / block partitioning structure, whether to perform quadtree-shaped partitioning, whether to perform binary-tree-shaped partitioning, binary-tree-shaped partitioning direction (horizontal or vertical), binary-tree-shaped partitioning shape (symmetric partitioning or asymmetric partitioning), whether the current coding unit is partitioned by ternary-tree partitioning, ternary-tree partitioning direction (horizontal or vertical), ternary-tree partitioning type (symmetric type or asymmetric type), whether the current coding unit is partitioned by multi-type tree partitioning, multi-type tree partitioning direction (horizontal or vertical), multi-type tree partitioning type (symmetric type or asymmetric type), multi-type tree partitioning tree (binary tree or ternary tree) structure, prediction mode (intra prediction or inter prediction), luminance intra prediction mode / direction, chrominance intra prediction mode / direction, intra-partitioning information, inter-partitioning information, coding block partitioning flag, prediction block partitioning flag, transform block partitioning flag, reference sample filtering method, reference sample filter taps, reference sample filter coefficients, prediction block filtering method, prediction block filter taps, prediction block filter coefficients, prediction block boundary filtering method, prediction block boundary filter taps, prediction block boundary filter coefficients, intra prediction mode, inter prediction mode, motion information, motion vector, motion vector difference, reference picture index, inter prediction angle, inter prediction indicator, prediction list utilization flag, reference picture list, reference picture, motion vector predictor index, motion vector predictor candidate, motion vector candidate list, whether to use the merge mode, merge index, merge candidate, merge candidate list, whether to use the skip mode, interpolation filter type, interpolation filter taps, interpolation filter coefficients, motion vector size, representation precision of the motion vector, transform type, transform size, information on whether the first (primary) transform is used, information on whether the secondary transform is used, primary transform index, secondary transform index, information on whether a residual signal exists, coding block style, coding block flag (CBF), quantization parameter, quantization parameter of the residual, quantization matrix, whether to apply an intra-loop filter, intra-loop filter coefficients, intra-loop filter taps, intra-loop filter shape / shape, whether to apply a deblocking filter, deblocking filter coefficients, deblocking filter taps, deblocking filter strength, deblocking filter shape / shape, whether to apply an adaptive sample offset, adaptive sample offset value, adaptive sample offset category, adaptive sample offset type, whether to apply an adaptive loop filter, adaptive loop filter coefficients, adaptive loop filter taps, adaptive loop filter shape / shape,Binarization / inverse binarization method, context model determination method, context model update method, whether to execute the normal mode, whether to execute the bypass mode, context binary bit, bypass binary bit, valid coefficient flag, last valid coefficient flag, coding flag for the unit of the coefficient group, position of the last valid coefficient, flag indicating whether the value of the coefficient is greater than 1, flag indicating whether the value of the coefficient is greater than 2, flag indicating whether the value of the coefficient is greater than 3, information about the values of the remaining coefficients, sign information, reconstructed luminance sample, reconstructed chrominance sample, residual luminance sample, residual chrominance sample, luminance transform coefficient, chrominance transform coefficient, quantized luminance level, quantized chrominance level, transform coefficient level scanning method, size of the motion vector search area on the decoder side, shape of the motion vector search area on the decoder side, number of times of motion vector search on the decoder side, information about the CTU size, information about the minimum block size, information about the maximum block size, information about the maximum block depth, information about the minimum block depth, image display / output order, slice identification information, slice type, slice partition information, parallel block identification information, parallel block type, parallel block partition information, parallel block group identification information, parallel block group type, parallel block group partition information, picture type, bit depth of the input sample, bit depth of the reconstructed sample, bit depth of the residual sample, bit depth of the transform coefficient, bit depth of the quantized level, and information about the luminance signal or information about the chrominance signal.
[0148] Here, it can be indicated by a signaling flag or index that the encoder performs entropy coding on the corresponding flag or index and includes it in the bitstream, and it can be indicated that the decoder performs entropy decoding on the corresponding flag or index from the bitstream.
[0149] When the encoding device 100 performs encoding by inter prediction, the encoded current picture can be used as a reference picture for another image to be processed subsequently. Therefore, the encoding device 100 can reconstruct or decode the encoded current picture, or store the reconstructed or decoded image as a reference picture in the reference picture buffer 190.
[0150] The quantized level can be dequantized in the dequantization unit 160, or can be inverse-transformed in the inverse transform unit 170. The coefficient that has been dequantized or inverse-transformed or both can be added to the prediction block by the adder 175. By adding the coefficient that has been dequantized or inverse-transformed or both to the prediction block, a reconstructed block can be generated. Here, the coefficient that has been dequantized or inverse-transformed or both can represent a coefficient on which at least one of dequantization and inverse transformation has been performed, and can represent a reconstructed residual block.
[0151] The reconstructed block can pass through the filter unit 180. The filter unit 180 can apply at least one of a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF) to the reconstructed samples, reconstructed block, or reconstructed image. The filter unit 180 can be referred to as a loop filter.
[0152] The deblocking filter can remove block distortion generated at the boundary between blocks. To determine whether to apply the deblocking filter, it can be determined whether to apply the deblocking filter to the current block based on the samples included in several rows or columns included in the block. When applying the deblocking filter to a block, another filter can be applied according to the required deblocking filtering strength.
[0153] To compensate for coding errors, an appropriate offset value can be added to the sample value by using sample adaptive offset. The sample adaptive offset can correct the offset between the deblocked image and the original image on a sample-by-sample basis. A method that applies an offset considering edge information about each sample can be used, or a method can be used in which the samples of the image are partitioned into a predetermined number of regions, the regions to which the offset is applied are determined, and the offset is applied to the determined regions.
[0154] The adaptive loop filter can perform filtering based on the comparison result between the filtered reconstructed image and the original image. The samples included in the image can be partitioned into predetermined groups, the filter to be applied to each group can be determined, and differential filtering can be performed on each group. Information on whether to apply the ALF can be signaled through the coding unit (CU), and the shape and coefficients of the ALF to be applied to each block can vary.
[0155] The reconstructed block or reconstructed image that has passed through the filter unit 180 can be stored in the reference picture buffer 190. The reconstructed block processed by the filter unit 180 can be part of a reference picture. That is, the reference picture is a reconstructed image composed of the reconstructed blocks processed by the filter unit 180. The stored reference picture can be used later in inter prediction or motion compensation.
[0156] Figure 2 is a block diagram showing the configuration of a decoding device according to an embodiment and applying the present invention.
[0157] The decoding device 200 can be a decoder, a video decoding device, or an image decoding device.
[0158] Referring to Figure 2 , the decoding device 200 can include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, a motion compensation unit 250, an adder 255, a filter unit 260, and a reference picture buffer 270.
[0159] The decoding device 200 may receive the bitstream output from the encoding device 100. The decoding device 200 may receive the bitstream stored in a computer-readable recording medium, or may receive the bitstream streamed through a wired / wireless transmission medium. The decoding device 200 may decode the bitstream by using an intra mode or an inter mode. In addition, the decoding device 200 may generate a reconstructed image or a decoded image produced by decoding, and output the reconstructed image or the decoded image.
[0160] When the prediction mode used during decoding is the intra mode, the switcher may be switched to intra. Optionally, when the prediction mode used during decoding is the inter mode, the switcher may be switched to the inter mode.
[0161] The decoding device 200 may obtain a reconstructed residual block by decoding the input bitstream, and generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 may generate a reconstructed block to be decoded by adding the reconstructed residual block to the prediction block. The block to be decoded may be referred to as the current block.
[0162] The entropy decoding unit 210 may generate symbols by performing entropy decoding on the bitstream according to a probability distribution. The generated symbols may include symbols in a quantized level shape. Here, the entropy decoding method may be the inverse process of the above entropy encoding method.
[0163] In order to decode the transform coefficient levels (quantized levels), the entropy decoding unit 210 may change the coefficients in a one-way vector shape into a two-dimensional block shape by using a transform coefficient scanning method.
[0164] The quantized levels may be dequantized in the dequantization unit 220, or may be inverse-transformed in the inverse transform unit 230. The quantized levels may be the result of performing dequantization or inverse transformation or both dequantization and inverse transformation, and may be generated as a reconstructed residual block. Here, the dequantization unit 220 may apply a quantization matrix to the quantized levels.
[0165] When using the intra mode, the intra prediction unit 240 may generate a prediction block by performing spatial prediction on the current block, where the spatial prediction uses the sample values of the blocks adjacent to the block to be decoded and already decoded.
[0166] When using the inter mode, the motion compensation unit 250 may generate a prediction block by performing motion compensation on the current block, where the motion compensation uses a motion vector and a reference picture stored in the reference picture buffer 270.
[0167] The adder 225 may generate a reconstructed block by adding a reconstructed residual block and a prediction block. The filter unit 260 may apply at least one of a deblocking filter, sample adaptive offset, and adaptive loop filter to the reconstructed block or the reconstructed image. The filter unit 260 may output the reconstructed image. The reconstructed block or the reconstructed image may be stored in the reference picture buffer 270 and used during inter prediction. The reconstructed block processed by the filter unit 260 may be part of a reference picture. That is, the reference picture is a reconstructed image composed of the reconstructed blocks processed by the filter unit 260. The stored reference picture may be used later in inter prediction or motion compensation.
[0168] Figure 3 is a diagram schematically showing a partitioning structure of an image when encoding and decoding the image. Figure 3 schematically shows an example of partitioning a single unit into multiple lower-level units.
[0169] To partition an image effectively, a coding unit (CU) may be used when encoding and decoding. The coding unit may be used as a basic unit when encoding / decoding an image. In addition, the coding unit may be used as a unit for distinguishing an intra prediction mode from an inter prediction mode when encoding / decoding an image. The coding unit may be a basic unit for prediction, transformation, quantization, inverse transformation, dequantization, or encoding / decoding processing of transform coefficients.
[0170] Referring to Figure 3 , the image 300 is sequentially partitioned according to the largest coding unit (LCU), and the LCU unit is determined as the partitioning structure. Here, the LCU may be used with the same meaning as the coding tree unit (CTU). Unit partitioning may represent partitioning a block associated with the unit. In the block partitioning information, information on the unit depth may be included. The depth information may represent the number of times or the degree or both the number of times and the degree to which the unit is partitioned. A single unit may be partitioned into multiple lower-level units hierarchically associated with the depth information based on a tree structure. In other words, the unit and the lower-level units generated by partitioning the unit may correspond to a node and the children of the node, respectively. Each of the partitioned lower-level units may have depth information. The depth information may be information representing the size of the CU and may be stored in each CU. The unit depth represents the number and / or degree related to partitioning the unit. Therefore, the partitioning information of the lower-level units may include information on the size of the lower-level units.
[0171] The partition structure can represent the distribution of coding units (CUs) within the LCU 310. Such a distribution can be determined based on whether a single CU is partitioned into multiple (positive integers equal to or greater than 2, including 2, 4, 8, 16, etc.) CUs. The horizontal size and vertical size of the CUs generated by the partition can be half of the horizontal size and vertical size of the CU before the partition, respectively, or can have sizes smaller than the horizontal size and vertical size before the partition according to the number of partitions. A CU can be recursively partitioned into multiple CUs. Through recursive partitioning, at least one of the height and width of the CU after partitioning can be reduced compared to at least one of the height and width of the CU before partitioning. The partitioning of the CU can be recursively executed until a predetermined depth or a predetermined size is reached. For example, the depth of the LCU can be 0, and the depth of the smallest coding unit (SCU) can be a predetermined maximum depth. Here, as described above, the LCU can be a coding unit with the maximum coding unit size, and the SCU can be a coding unit with the smallest coding unit size. The partitioning starts from the LCU 310, and when the horizontal size or vertical size or both the horizontal size and vertical size of the CU are reduced by the partition, the CU depth increases by 1. For example, for each depth, the size of the unpartitioned CU can be 2N×2N. In addition, in the case of a partitioned CU, a CU with a size of 2N×2N can be partitioned into four CUs with a size of N×N. As the depth increases by 1, the size of N can be halved.
[0172] In addition, information indicating whether a CU is partitioned can be represented by using the partition information of the CU. The partition information can be 1-bit information. All CUs except the SCU can include the partition information. For example, when the value of the partition information is a first value, the CU can not be partitioned, and when the value of the partition information is a second value, the CU can be partitioned.
[0173] Referring to Figure 3 , the LCU with a depth of 0 can be a 64×64 block. 0 can be the minimum depth. The SCU with a depth of 3 can be an 8×8 block. 3 can be the maximum depth. The CUs of 32×32 blocks and 16×16 blocks can be represented as depth 1 and depth 2, respectively.
[0174] For example, when a single coding unit is partitioned into four coding units, the horizontal size and vertical size of the four partitioned coding units can be half of the horizontal size and vertical size of the CU before being partitioned. In one embodiment, when a coding unit with a size of 32×32 is partitioned into four coding units, each of the four partitioned coding units can have a size of 16×16. When a single coding unit is partitioned into four coding units, it can be said that the coding unit can be partitioned into a quadtree shape.
[0175] For example, when a coding unit is partitioned into two sub-coding units, the horizontal size or vertical size (width or height) of each of the two sub-coding units can be half of the horizontal size or vertical size of the original coding unit. For example, when a coding unit with a size of 32×32 is vertically partitioned into two sub-coding units, each of the two sub-coding units can have a size of 16×32. For example, when a coding unit with a size of 8×32 is horizontally partitioned into two sub-coding units, each of the two sub-coding units can have a size of 8×16. When a coding unit is partitioned into two sub-coding units, it can be said that the coding unit is bipartitioned or partitioned through a binary tree partitioning structure.
[0176] For example, when a coding unit is partitioned into three sub-coding units, the horizontal size or vertical size of the coding unit can be partitioned in a ratio of 1:2:1, thereby generating three sub-coding units with a horizontal size or vertical size ratio of 1:2:1. For example, when a coding unit with a size of 16×32 is horizontally partitioned into three sub-coding units, the three sub-coding units can have sizes of 16×8, 16×16, and 16×8 in order from the uppermost sub-coding unit to the lowermost sub-coding unit. For example, when a coding unit with a size of 32×32 is vertically partitioned into three sub-coding units, the three sub-coding units can have sizes of 8×32, 16×32, and 8×32 in order from the leftmost sub-coding unit to the rightmost sub-coding unit. When a coding unit is partitioned into three sub-coding units, it can be said that the coding unit is tripartitioned or partitioned according to a ternary tree partitioning structure.
[0177] In Figure 3 the coding tree unit (CTU) 320 is an example of a CTU to which a quadtree partitioning structure, a binary tree partitioning structure, and a ternary tree partitioning structure are all applied.
[0178] As described above, in order to partition a CTU, at least one of a quadtree partitioning structure, a binary tree partitioning structure, and a ternary tree partitioning structure can be applied. Various tree partitioning structures can be sequentially applied to the CTU according to a predetermined priority order. For example, the quadtree partitioning structure can be preferentially applied to the CTU. A coding unit for which the quadtree partitioning structure can no longer be used for partitioning can correspond to a leaf node of the quadtree. A coding unit corresponding to a leaf node of the quadtree can be used as the root node of a binary tree and / or a ternary tree partitioning structure. That is, a coding unit corresponding to a leaf node of the quadtree can be further partitioned according to a binary tree partitioning structure or a ternary tree partitioning structure, or may not be further partitioned. Therefore, by preventing coding units obtained from binary tree partitioning or ternary tree partitioning of a coding unit corresponding to a leaf node of the quadtree from undergoing further quadtree partitioning, the block partitioning operation and / or the operation of signaling partitioning information can be effectively performed.
[0179] The fact that a coding unit corresponding to a node of a quadtree is partitioned can be signaled using quadtree partition information. Quadtree partition information having a first value (e.g., "1") can indicate that the current coding unit is partitioned according to the quadtree partition structure. Quadtree partition information having a second value (e.g., "0") can indicate that the current coding unit is not partitioned according to the quadtree partition structure. The quadtree partition information can be a flag having a predetermined length (e.g., one bit).
[0180] There may be no priority between binary tree partitioning and ternary tree partitioning. That is, a coding unit corresponding to a leaf node of a quadtree can further undergo either binary tree partitioning or ternary tree partitioning. Additionally, a coding unit generated by binary tree partitioning or ternary tree partitioning can undergo further binary tree partitioning or further ternary tree partitioning, or may not be further partitioned.
[0181] A tree structure in which there is no priority between binary tree partitioning and ternary tree partitioning is referred to as a multi-type tree structure. A coding unit corresponding to a leaf node of a quadtree can be used as the root node of a multi-type tree. At least one of multi-type tree partition indication information, partition direction information, and partition tree information can be used to signal whether to partition a coding unit corresponding to a node of a multi-type tree. In order to partition a coding unit corresponding to a node of a multi-type tree, the multi-type tree partition indication information, partition direction information, and partition tree information can be signaled sequentially.
[0182] Multi-type tree partition indication information having a first value (e.g., "1") can indicate that the current coding unit will undergo multi-type tree partitioning. Multi-type tree partition indication information having a second value (e.g., "0") can indicate that the current coding unit will not undergo multi-type tree partitioning.
[0183] When a coding unit corresponding to a node of a multi-type tree is further partitioned according to the multi-type tree partition structure, the coding unit can include partition direction information. The partition direction information can indicate in which direction the current coding unit will be partitioned for multi-type tree partitioning. Partition direction information having a first value (e.g., "1") can indicate that the current coding unit will be vertically partitioned. Partition direction information having a second value (e.g., "0") can indicate that the current coding unit will be horizontally partitioned.
[0184] When a coding unit corresponding to a node of a multi-type tree is further partitioned according to the multi-type tree partition structure, the current coding unit can include partition tree information. The partition tree information can indicate the tree partition structure that will be used to partition the node of the multi-type tree. Partition tree information having a first value (e.g., "1") can indicate that the current coding unit will be partitioned according to the binary tree partition structure. Partition tree information having a second value (e.g., "0") can indicate that the current coding unit will be partitioned according to the ternary tree partition structure.
[0185] The partition indication information, the partition tree information, and the partition direction information can all be flags having a predetermined length (e.g., one bit).
[0186] At least any one of the quadtree partition indication information, the multi-type tree partition indication information, the partition direction information, and the partition tree information can be entropy-coded / entropy-decoded. To entropy-code / entropy-decode those types of information, information about neighboring coding units adjacent to the current coding unit can be used. For example, the probability that the partition type (partitioned or not partitioned, partition tree, and / or partition direction) of the left neighboring coding unit and / or the upper neighboring coding unit of the current coding unit is similar to the partition type of the current coding unit is very high. Therefore, context information for entropy-coding / entropy-decoding the information about the current coding unit can be derived from the information about the neighboring coding units. The information about the neighboring coding units can include at least any one of the quad-partition information, the multi-type tree partition indication information, the partition direction information, and the partition tree information.
[0187] As another example, in binary tree partitioning and ternary tree partitioning, binary tree partitioning can be preferentially performed. That is, the current coding unit can first undergo binary tree partitioning, and subsequently, the coding unit corresponding to the leaf node of the binary tree can be set as the root node for ternary tree partitioning. In this case, for the coding unit corresponding to the node of the ternary tree, neither quadtree partitioning nor binary tree partitioning can be performed.
[0188] The coding unit that cannot be partitioned according to the quadtree partition structure, the binary tree partition structure, and / or the ternary tree partition structure becomes the basic unit for coding, prediction, and / or transformation. That is, the coding unit cannot be further partitioned for prediction and / or transformation. Therefore, in the bitstream, there may be no partition structure information and partition information for partitioning the coding unit into a prediction unit and / or a transformation unit.
[0189] However, when the size of a coding unit (i.e., the basic unit for partitioning) is larger than the size of the maximum transform block, the coding unit can be recursively partitioned until the size of the coding unit is reduced to be equal to or smaller than the size of the maximum transform block. For example, when the size of the coding unit is 64×64 and when the size of the maximum transform block is 32×32, the coding unit can be partitioned into four 32×32 blocks for transformation. For example, when the size of the coding unit is 32×64 and the size of the maximum transform block is 32×32, the coding unit can be partitioned into two 32×32 blocks for transformation. In this case, the partitioning of the coding unit for transformation is not signaled separately, and the partitioning of the coding unit for transformation can be determined by comparing the horizontal size or vertical size of the coding unit with the horizontal size or vertical size of the maximum transform block. For example, when the horizontal size (width) of the coding unit is larger than the horizontal size (width) of the maximum transform block, the coding unit can be bisected vertically. For example, when the vertical size (height) of the coding unit is larger than the vertical size (height) of the maximum transform block, the coding unit can be bisected horizontally.
[0190] The information on the maximum and / or minimum size of the coding unit and the information on the maximum and / or minimum size of the transform block can be signaled or determined at a higher level of the coding unit. The higher level can be, for example, sequence level, picture level, slice level, parallel block group level, parallel block level, etc. For example, the minimum size of the coding unit can be determined to be 4×4. For example, the maximum size of the transform block can be determined to be 64×64. For example, the minimum size of the transform block can be determined to be 4×4.
[0191] The information on the minimum size (quadtree minimum size) of the coding unit corresponding to the leaf node of the quadtree and / or the information on the maximum depth (maximum tree depth of the multi-type tree) from the root node to the leaf node of the multi-type tree can be signaled or determined at a higher level of the coding unit. For example, the higher level can be sequence level, picture level, slice level, parallel block group level, parallel block level, etc. The information on the minimum size of the quadtree and / or the information on the maximum depth of the multi-type tree can be signaled or determined for each of the in-picture slices and inter-picture slices.
[0192] The difference information between the size of the CTU and the maximum size of the transform block can be signaled or determined at a higher level of the coding unit. For example, the higher level can be the sequence level, picture level, slice level, parallel block group level, parallel block level, etc. The information on the maximum size of the coding unit (hereinafter referred to as the maximum size of the binary tree) corresponding to each node of the binary tree can be determined based on the size of the coding tree unit and the difference information. The maximum size of the coding unit (hereinafter referred to as the maximum size of the ternary tree) corresponding to each node of the ternary tree can vary according to the type of slice. For example, for an intra-slice, the maximum size of the ternary tree can be 32×32. For example, for an inter-slice, the maximum size of the ternary tree can be 128×128. For example, the minimum size of the coding unit (hereinafter referred to as the minimum size of the binary tree) corresponding to each node of the binary tree and / or the minimum size of the coding unit (hereinafter referred to as the minimum size of the ternary tree) corresponding to each node of the ternary tree can be set to the minimum size of the coding block.
[0193] As another example, the maximum size of the binary tree and / or the maximum size of the ternary tree can be signaled or determined at the slice level. Optionally, the minimum size of the binary tree and / or the minimum size of the ternary tree can be signaled or determined at the slice level.
[0194] According to the size and depth information of the above various blocks, the quad-partition information, multi-type tree partition indication information, partition tree information, and / or partition direction information may or may not be included in the bitstream.
[0195] For example, when the size of the coding unit is not greater than the minimum size of the quadtree, the coding unit does not include the quad-partition information. Therefore, the quad-partition information can be inferred as a second value.
[0196] For example, when the size (horizontal size and vertical size) of the coding unit corresponding to a node of the multi-type tree is greater than the maximum size (horizontal size and vertical size) of the binary tree and / or the maximum size (horizontal size and vertical size) of the ternary tree, the coding unit may not be partitioned by the binary tree or the ternary tree. Therefore, the multi-type tree partition indication information may not be signaled, but the multi-type tree partition indication information can be inferred as a second value.
[0197] Optionally, when the size (horizontal size and vertical size) of a coding unit corresponding to a node of a multi-type tree is the same as the maximum size (horizontal size and vertical size) of a binary tree and / or twice as large as the maximum size (horizontal size and vertical size) of a ternary tree, the coding unit may not be further bipartitioned or tripartitioned. Therefore, instead of signaling multi-type tree partitioning indication information, the multi-type tree partitioning indication information may be derived from a second value. This is because when partitioning a coding unit through a binary tree partitioning structure and / or a ternary tree partitioning structure, coding units smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree are generated.
[0198] Optionally, the binary tree partitioning or ternary tree partitioning may be restricted based on the size of a virtual pipeline data unit (hereinafter, pipeline buffer size). For example, when partitioning a coding unit into sub-coding units that do not fit the pipeline buffer size through binary tree partitioning or ternary tree partitioning, the corresponding binary tree partitioning or ternary tree partitioning may be restricted. The pipeline buffer size may be the size of the largest transform block (e.g., 64×64). For example, when the pipeline buffer size is 64×64, the following partitioning may be restricted.
[0199] - Ternary tree partitioning for an N×M (N and / or M is 128) coding unit
[0200] - Binary tree partitioning in the horizontal direction for a 128×N (N <= 64) coding unit
[0201] - Binary tree partitioning in the vertical direction for an N×128 (N <= 64) coding unit
[0202] Optionally, when the depth of a coding unit corresponding to a node of a multi-type tree is equal to the maximum depth of the multi-type tree, the coding unit may not be further bipartitioned and / or tripartitioned. Therefore, instead of signaling multi-type tree partitioning indication information, the multi-type tree partitioning indication information may be inferred as a second value.
[0203] Optionally, the multi-type tree partitioning indication information may be signaled only when at least one of the binary tree partitioning in the vertical direction, the binary tree partitioning in the horizontal direction, the ternary tree partitioning in the vertical direction, and the ternary tree partitioning in the horizontal direction is possible for a coding unit corresponding to a node of a multi-type tree. Otherwise, the coding unit may not be bipartitioned and / or tripartitioned. Therefore, instead of signaling multi-type tree partitioning indication information, the multi-type tree partitioning indication information may be inferred as a second value.
[0204] Optionally, the partitioning direction information may be signaled only when both the vertical binary tree partitioning and the horizontal binary tree partitioning or both the vertical ternary tree partitioning and the horizontal ternary tree partitioning are possible for a coding unit corresponding to a node of the multi-type tree. Otherwise, the partitioning direction information may not be signaled, but the partitioning direction information may be derived from values indicating possible partitioning directions.
[0205] Optionally, the partitioning tree information may be signaled only when both the vertical binary tree partitioning and the vertical ternary tree partitioning or both the horizontal binary tree partitioning and the horizontal ternary tree partitioning are possible for a coding tree corresponding to a node of the multi-type tree. Otherwise, the partitioning tree information may not be signaled, but may be inferred as a value indicating a possible partitioning tree structure.
[0206] Figure 4 is a diagram showing the intra prediction process.
[0207] Figure 4 The arrows from the center to the outside in may represent the prediction direction of the intra prediction mode.
[0208] Intra coding and / or decoding may be performed by using reference samples of neighboring blocks of the current block. The neighboring blocks may be reconstructed neighboring blocks. For example, intra coding and / or decoding may be performed by using coding parameters or values of reference samples included in the reconstructed neighboring blocks.
[0209] A prediction block may represent a block generated by performing intra prediction. The prediction block may correspond to at least one of a CU, a PU, and a TU. The unit of the prediction block may have the size of one of a CU, a PU, and a TU. The prediction block may be a square block with a size of 2×2, 4×4, 16×16, 32×32, or 64×64, etc., or may be a rectangular block with a size of 2×8, 4×8, 2×16, 4×16, and 8×16, etc.
[0210] Intra prediction may be performed according to the intra prediction mode for the current block. The number of intra prediction modes that the current block may have may be a fixed value, and may be a value determined differently according to the attributes of the prediction block. For example, the attributes of the prediction block may include the size of the prediction block and the shape of the prediction block, etc.
[0211] Regardless of the block size, the number of intra prediction modes can be fixed at N. Alternatively, the number of intra prediction modes can be, for example, 3, 5, 9, 17, 34, 35, 36, 65, or 67. Optionally, the number of intra prediction modes can vary according to the block size or color component type or both the block size and color component type. For example, the number of intra prediction modes can vary according to whether the color component is a luminance signal or a chrominance signal. For example, as the block size increases, the number of intra prediction modes can increase. Optionally, the number of intra prediction modes for a luminance component block can be greater than the number of intra prediction modes for a chrominance component block.
[0212] The intra prediction mode can be a non - angular mode or an angular mode. The non - angular mode can be a DC mode or a planar mode, and the angular mode can be a prediction mode having a specific direction or angle. The intra prediction mode can be represented by at least one of a mode number, a mode value, a mode numbering, a mode angle, and a mode direction. The number of intra prediction modes can be M greater than 1, including the non - angular mode and the angular mode. To perform intra prediction on a current block, a step of determining whether samples included in reconstructed neighboring blocks can be used as reference samples for the current block can be performed. When there are samples that cannot be used as reference samples for the current block, values obtained by copying at least one sample value included in the reconstructed neighboring blocks or performing interpolation or performing both copying and interpolation can be used to replace the unavailable sample values of the samples, and thus the replaced sample values are used as reference samples for the current block.
[0213] Figure 7 is a diagram showing reference samples that can be used for intra prediction.
[0214] As Figure 7 shown, at least one of reference sample lines 0 to 3 can be used for intra prediction of the current block. In Figure 7 , the samples of segment A and segment F can be filled with the samples closest to segment B and segment E, respectively, instead of being retrieved from the reconstructed neighboring blocks. Index information indicating the reference sample line to be used for intra prediction of the current block can be signaled. For example, in Figure 7 , reference sample line indicators 0, 1, and 2 can be signaled as index information indicating reference sample lines 0, 1, and 2. When the upper boundary of the current block is the boundary of a CTU, only reference sample line 0 can be available. Therefore, in this case, the index information can be not signaled. When using a reference sample line other than reference sample line 0, filtering for the prediction block described later can be not performed.
[0215] When performing intra prediction, a filter can be applied to at least one of the reference samples and the prediction samples based on the intra prediction mode and the current block size / shape.
[0216] In the case of the planar mode, when generating a prediction block of a current block, according to the position of a prediction target sample within the prediction block, the sample value of the prediction target sample can be generated by using a weighted sum of the upper reference sample and the left reference sample of the current block and the upper-right reference sample and the lower-left reference sample of the current block. Additionally, in the case of the DC mode, when generating a prediction block of a current block, the average value of the upper reference sample and the left reference sample of the current block can be used. Additionally, in the case of the angular mode, the prediction block can be generated by using the upper reference sample, the left reference sample, the upper-right reference sample, and / or the lower-left reference sample of the current block. To generate the prediction sample value, interpolation of real number units can be performed.
[0217] In the case of intra prediction between color components, a prediction block of a current block of a second color component can be generated based on a corresponding reconstructed block of a first color component. For example, the first color component can be a luminance component, and the second color component can be a chrominance component. For intra prediction between color components, parameters of a linear model between the first color component and the second color component can be derived based on a template. The template can include upper and / or left neighboring samples of the current block and upper and / or left neighboring samples of the corresponding reconstructed block of the first color component. For example, the sample value of the first color component with the maximum value among the samples in the template and the corresponding sample value of the second color component, and the sample value of the first color component with the minimum value among the samples in the template and the corresponding sample value of the second color component can be used to derive the parameters of the linear model. When deriving the parameters of the linear model, the corresponding reconstructed block can be applied to the linear model to generate a prediction block of the current block. According to the video format, quadratic sampling can be performed on the reconstructed block of the first color component and neighboring samples of the corresponding reconstructed block. For example, when one sample of the second color component corresponds to four samples of the first color component, quadratic sampling can be performed on the four samples of the first color component to calculate one corresponding sample. In this case, derivation of the parameters of the linear model and intra prediction between color components can be performed based on the samples of the corresponding quadratic sampling. Whether to perform intra prediction between color components and / or the range of the template can be signaled as an intra prediction mode.
[0218] The current block can be partitioned into two sub - blocks or four sub - blocks either horizontally or vertically. The partitioned sub - blocks can be reconstructed sequentially. That is, intra - prediction can be performed on the sub - blocks to generate sub - prediction blocks. Additionally, inverse quantization and / or inverse transformation can be performed on the sub - blocks to generate sub - residual blocks. The reconstructed sub - blocks can be generated by adding the sub - prediction blocks to the sub - residual blocks. The reconstructed sub - blocks can be used as reference sample points for intra - prediction of sub - sub - blocks. A sub - block can be a block including a predetermined number (e.g., 16) or more sample points. Thus, for example, when the current block is an 8×4 block or a 4×8 block, the current block can be partitioned into two sub - blocks. Further, when the current block is a 4×4 block, the current block may not be partitioned into sub - blocks. When the current block has other dimensions, the current block can be partitioned into four sub - blocks. Information regarding whether to perform intra - prediction based on sub - blocks and / or the partition direction (horizontal or vertical) can be signaled. The intra - prediction based on sub - blocks can be limited to be performed only when using reference sample line 0. When performing intra - prediction based on sub - blocks, the filtering for the prediction block described later may not be performed.
[0219] The final prediction block can be generated by performing filtering on the prediction block that has been intra - predicted. The filtering can be performed by applying predetermined weights to the filtering target sample points, the left reference sample points, the upper reference sample points, and / or the upper - left reference sample points. The weights and / or reference sample points (range, position, etc.) for filtering can be determined based on at least one of the block size, the intra - prediction mode, and the position of the filtering target sample points in the prediction block. The filtering can be performed only in the case of predetermined intra - prediction modes (e.g., DC, planar, vertical, horizontal, diagonal, and / or adjacent - diagonal modes). The adjacent - diagonal mode can be a mode obtained by adding k to or subtracting k from the diagonal mode. For example, k can be a positive integer of 8 or less.
[0220] The intra - prediction mode of the current block can be entropy - encoded / entropy - decoded by predicting the intra - prediction modes of the blocks adjacent to the current block. Additionally, the information that the intra - prediction modes of the current block and the neighboring blocks are the same can be signaled by using predetermined flag information. Further, the indicator information of the intra - prediction mode among the intra - prediction modes of multiple neighboring blocks that is the same as the intra - prediction mode of the current block can be signaled. When the intra - prediction modes of the current block and the neighboring blocks are not the same, the intra - prediction mode information of the current block can be entropy - encoded / entropy - decoded by performing entropy - encoding / entropy - decoding based on the intra - prediction modes of the neighboring blocks.
[0221] Figure 5 It is a diagram showing an embodiment of the inter - prediction process.
[0222] In Figure 5 a rectangle can represent a picture. In Figure 5In this case, the arrow indicates the prediction direction. According to the coding type of the picture, the picture can be classified into an intra picture (I picture), a predictive picture (P picture), and a bi-predictive picture (B picture).
[0223] The I picture can be encoded by intra prediction without the need for inter-picture prediction. The P picture can be encoded by inter-picture prediction by using the reference picture existing in one direction (i.e., forward or backward) for the current block. The B picture can be encoded by inter-picture prediction by using the reference pictures existing in two directions (i.e., forward and backward) for the current block. When using inter-picture prediction, the encoder can perform inter-picture prediction or motion compensation, and the decoder can perform corresponding motion compensation.
[0224] Hereinafter, embodiments of inter-frame prediction will be described in detail.
[0225] The reference picture and motion information can be used to perform inter-picture prediction or motion compensation.
[0226] The motion information of the current block can be derived by each of the encoding device 100 and the decoding device 200 during inter-picture prediction. The motion information of the current block can be derived by using the motion information of the reconstructed neighboring block, the motion information of the co-located block (also referred to as the col block or co-located block), and / or the motion information of the block adjacent to the co-located block. The co-located block can represent a block that is spatially in the same position as the current block within a previously reconstructed co-located picture (also referred to as the col picture or co-located picture). The co-located picture can be one of one or more reference pictures included in the reference picture list.
[0227] The method for deriving the motion information may vary depending on the prediction mode of the current block. For example, the prediction modes applied to inter-frame prediction include the AMVP mode, the merge mode, the skip mode, the merge mode with motion vector difference, the sub-block merge mode, the geometric partitioning mode, the combined inter-intra prediction mode, the affine mode, etc. Here, the merge mode can be referred to as the motion merge mode.
[0228] For example, when AMVP is used as the prediction mode, at least one of the motion vectors of the reconstructed neighboring block, the motion vector of the co-located block, the motion vector of the block adjacent to the co-located block, and the (0,0) motion vector can be determined as the motion vector candidate for the current block, and a motion vector candidate list is generated by using the motion vector candidate. The motion vector candidate of the current block can be derived by using the generated motion vector candidate list. The motion information of the current block can be determined based on the derived motion vector candidate. The motion vector of the co-located block or the motion vector of the block adjacent to the co-located block can be referred to as the temporal motion vector candidate, and the motion vector of the reconstructed neighboring block can be referred to as the spatial motion vector candidate.
[0229] The encoding device 100 may calculate a motion vector difference (MVD) between a motion vector of a current block and a motion vector candidate, and may perform entropy encoding on the motion vector difference (MVD). Additionally, the encoding device 100 may perform entropy encoding on a motion vector candidate index and generate a bitstream. The motion vector candidate index may indicate the best motion vector candidate among the motion vector candidates included in a motion vector candidate list. The decoding device may perform entropy decoding on the motion vector candidate index included in the bitstream, and may select a motion vector candidate for a decoding target block from the motion vector candidates included in the motion vector candidate list by using the entropy-decoded motion vector candidate index. Additionally, the decoding device 200 may add the entropy-decoded MVD to the motion vector candidate extracted by entropy decoding, thereby deriving the motion vector of the decoding target block.
[0230] Additionally, the encoding device 100 may perform entropy encoding on resolution information of the calculated MVD. The decoding device 200 may use the MVD resolution information to adjust the resolution of the entropy-decoded MVD.
[0231] Additionally, the encoding device 100 calculates a motion vector difference (MVD) between the motion vector in the current block and the motion vector candidate based on an affine model, and performs entropy encoding on the MVD. The decoding device 200 derives an affine control motion vector of the decoded target block based on the sum of the entropy-decoded MVD and the affine control motion vector candidate to derive a motion vector for each sub-block.
[0232] The bitstream may include a reference picture index indicating a reference picture. The reference picture index may be entropy-encoded by the encoding device 100 and subsequently signaled as the bitstream to the decoding device 200. The decoding device 200 may generate a prediction block for the decoding target block based on the derived motion vector and the reference picture index information.
[0233] Another example of a method for deriving motion information of a current block may be a merge mode. The merge mode may represent a method of merging the motion of multiple blocks. The merge mode may represent a mode of deriving the motion information of the current block from the motion information of neighboring blocks. When the merge mode is applied, the motion information of the reconstructed neighboring blocks and / or the motion information of co-located blocks may be used to generate a merge candidate list. The motion information may include at least one of a motion vector, a reference picture index, and an inter-picture prediction indicator. The prediction indicator may indicate uni-directional prediction (L0 prediction or L1 prediction) or bi-directional prediction (L0 prediction and L1 prediction).
[0234] The merge candidate list may be a list of stored motion information. The motion information included in the merge candidate list may be at least one of the following: motion information of neighboring blocks adjacent to the current block (spatial merge candidates), motion information of collocated blocks of the current block in a reference picture (temporal merge candidates), new motion information generated by combining the motion information existing in the merge candidate list, motion information of blocks encoded / decoded before the current block (history-based merge candidates), and zero merge candidates.
[0235] The encoding device 100 may generate a bitstream by performing entropy encoding on at least one of the merge flag and the merge index, and may signal the bitstream to the decoding device 200. The merge flag may be information indicating whether the merge mode is performed for each block, and the merge index may be information indicating which neighboring block among the neighboring blocks of the current block is the merge target block. For example, the neighboring blocks of the current block may include a left neighboring block located on the left side of the current block, an upper neighboring block arranged above the current block, and a temporal neighboring block temporally adjacent to the current block.
[0236] In addition, the encoding device 100 performs entropy encoding on correction information for correcting a motion vector in the motion information of the merge candidate and signals it to the decoding device 200. The decoding device 200 may correct the motion vector of the merge candidate selected by the merge index based on the correction information. Here, the correction information may include at least one of information on whether correction is performed, correction direction information, and correction size information. As described above, the prediction mode of correcting the motion vector of the merge candidate based on the signaled correction information may be referred to as a merge mode with a motion vector difference.
[0237] The skip mode may be a mode of applying the motion information of a neighboring block to the current block as it is. When the skip mode is applied, the encoding device 100 may perform entropy encoding on information about which block's motion information will be used as the motion information of the current block to generate a bitstream, and may signal the bitstream to the decoding device 200. The encoding device 100 may not signal syntax elements regarding at least any one of motion vector difference information, coded block flag, and transform coefficient level to the decoding device 200.
[0238] The sub-block merge mode may represent a mode of deriving motion information in units of sub-blocks of a coding unit (CU). When the sub-block merge mode is applied, motion information of sub-blocks collocated with the current sub-block in a reference picture (sub-block-based temporal merge candidates) and / or affine control point motion vector merge candidates may be used to generate a sub-block merge candidate list.
[0239] The geometric partitioning mode may represent a mode of deriving motion information by partitioning a current block in a predefined direction, using each of the derived motion information to derive each predicted sample point, and deriving the predicted sample point of the current block by weighting each of the derived predicted sample points.
[0240] The inter-intra combined prediction mode may represent a mode of deriving the predicted sample point of a current block by weighting the predicted sample points generated by inter-frame prediction and the predicted sample points generated by intra-frame prediction.
[0241] The decoding device 200 may correct the derived motion information by itself. The decoding device 200 may search a predetermined area based on the reference block indicated by the derived motion information, and derive the motion information with the minimum SAD as the corrected motion information.
[0242] The decoding device 200 may compensate the predicted sample points derived via inter-frame prediction using optical flow.
[0243] Figure 6 is a diagram showing transform and quantization processing.
[0244] As Figure 6 shown, transform processing and / or quantization processing is performed on the residual signal to generate a quantized level signal. The residual signal is the difference between the original block and the predicted block (i.e., the intra-frame predicted block or the inter-frame predicted block). The predicted block is a block generated by intra-frame prediction or inter-frame prediction. The transform may be a primary transform, a secondary transform, or both a primary transform and a secondary transform. The primary transform of the residual signal generates transform coefficients, and the secondary transform of the transform coefficients generates secondary transform coefficients.
[0245] At least one scheme selected from various predefined transform schemes is used to perform the primary transform. For example, examples of the predetermined transform schemes include discrete cosine transform (DCT), discrete sine transform (DST), and Karhunen-Loève transform (KLT). The transform coefficients generated by the primary transform may undergo a secondary transform. The transform scheme used for the primary transform and / or the secondary transform may be determined according to the coding parameters of the current block and / or the neighboring blocks of the current block. Optionally, transform information indicating the transform scheme may be signaled. The DCT-based transform may include, for example, DCT-2, DCT-8, etc. The DST-based transform may include, for example, DST-7.
[0246] A quantized level signal (quantization coefficient) can be generated by performing quantization on a residual signal or a result of performing a primary transform and / or a secondary transform. Depending on the intra prediction mode or block size / shape of a block, the quantized level signal can be scanned according to at least one of a diagonal right-up scan, a vertical scan, and a horizontal scan. For example, when scanning coefficients in a diagonal right-up scan, the coefficients of the block shape are changed to a one-dimensional vector shape. In addition to the diagonal right-up scan, depending on the intra prediction mode and / or the size of the transform block, a horizontal scan that horizontally scans the coefficients of a two-dimensional block shape or a vertical scan that vertically scans the coefficients of a two-dimensional block shape can be used. The scanned quantized level coefficients can be entropy coded to be inserted into a bitstream.
[0247] The decoder performs entropy decoding on the bitstream to obtain the quantized level coefficients. The quantized level coefficients can be arranged in a two-dimensional block shape by inverse scanning. For the inverse scanning, at least one of a diagonal right-up scan, a vertical scan, and a horizontal scan can be used.
[0248] Then, the quantized level coefficients can be dequantized, then a secondary inverse transform can be performed as needed, and finally a primary inverse transform can be performed as needed to generate a reconstructed residual signal.
[0249] An inverse mapping in the dynamic range can be performed on the luminance component reconstructed by intra prediction or inter prediction before loop filtering. The dynamic range can be partitioned into 16 equal segments, and a mapping function for each segment can be signaled. The mapping function can be signaled at the slice level or the parallel block group level. An inverse mapping function for performing the inverse mapping can be derived based on the mapping function. Loop filtering, reference picture storage, and motion compensation are performed in the inverse mapping region, and a prediction block generated by inter prediction is transformed to the mapping region via mapping using the mapping function and then used to generate a reconstructed block. However, since intra prediction is performed in the mapping region, a prediction block generated by intra prediction can be used to generate a reconstructed block without mapping / inverse mapping.
[0250] When the current block is a residual block of a chrominance component, the residual block can be transformed to the inverse mapped region by performing scaling on the chrominance component of the mapped region. The availability of the scaling can be signaled at the slice level or at the parallel block group level. The scaling can be applied only when the mapping of the luminance component is available and the partitioning of the luminance component and the partitioning of the chrominance component follow the same tree structure. The scaling can be performed based on the average of the sample values of the luminance prediction block corresponding to the chrominance difference block. In this case, when inter prediction is used for the current block, the luminance prediction block can represent the mapped luminance prediction block. The value required for scaling can be derived by using the index of the segment to which the average of the sample values of the luminance prediction block belongs to look up a table. Finally, by scaling the residual block using the derived value, the residual block can be transformed to the inverse mapped region. Then, chrominance component block recovery, intra prediction, inter prediction, loop filtering, and reference picture storage can be performed in the inverse mapped region.
[0251] The information indicating whether the mapping / inverse mapping of the luminance component and the chrominance component is available can be signaled via the sequence parameter set.
[0252] The prediction block of the current block can be generated based on the block vector indicating the displacement between the current block in the current picture and the reference block. In this way, the prediction mode for generating the prediction block by referring to the current picture is called the Intra Block Copy (IBC) mode. The IBC mode can be applied to M×N (M <= 64, N <= 64) coding units. The IBC mode can include a skip mode, a merge mode, an AMVP mode, etc. In the case of the skip mode or the merge mode, a merge candidate list is constructed, and a merge index is signaled so that a merge candidate can be specified. The block vector of the specified merge candidate can be used as the block vector of the current block. The merge candidate list can include at least one of a spatial candidate, a history-based candidate, a candidate based on the average of two candidates, and a zero merge candidate. In the case of the AMVP mode, a differential block vector can be signaled. Additionally, the prediction block vector can be derived from the left neighboring block and the upper neighboring block of the current block. The index of the neighboring block to be used can be signaled. The prediction block in the IBC mode is included in the current CTU or the left CTU and is limited to the blocks in the already reconstructed region. For example, the value of the block vector can be restricted so that the prediction block of the current block is located in the region of three 64×64 blocks before the 64×64 block to which the current block belongs in the coding / decoding order. By restricting the value of the block vector in this way, the memory consumption and device complexity according to the IBC mode implementation can be reduced.
[0253] In the following, reference will be made to Figures 8 to 23 describe a sub-block partitioning method and / or a method for deriving prediction information between sub-blocks according to an embodiment of the present invention.
[0254] A method for deriving inter prediction information will be described.
[0255] When performing inter prediction of a current block according to a merge mode, merge candidates may include spatial merge candidates, temporal merge candidates, sub-block-based temporal merge candidates, sub-block-based spatio-temporal combined merge candidates, combined merge candidates, zero merge candidates, and the like. The merge candidates may include inter prediction information of at least one of an inter prediction indicator, a reference picture index of a reference picture list, a motion vector, and a picture order count (POC).
[0256] A method for deriving spatial merge candidates will be described.
[0257] Spatial merge candidates for a current block can be derived from reconstructed blocks that are spatially adjacent to the current block to be encoded / decoded.
[0258] Figure 8 FIG. is a diagram illustrating a method for deriving spatial merge candidates according to an embodiment of the present invention.
[0259] Referring to Figure 8 , motion information can be derived from a block corresponding to at least one of a block A1 located to the left of a current block X to be encoded / decoded, a block B1 located above the current block X, a block B0 located in the upper right corner of the current block X, a block A0 located in the lower left corner of the current block X, and a block B2 located in the upper left corner of the current block X. The spatial merge candidates of the current block can be determined by using the derived motion information. In the example, the derived motion information can be used as the spatial merge candidates of the current block.
[0260] Spatial merge candidates may represent blocks reconstructed adjacent to the block to be encoded / decoded in space (or motion information of reconstructed blocks adjacent in space). The blocks may have a square shape or a non-square shape. In addition, blocks reconstructed adjacent to the block to be encoded / decoded in space may be divided into units of low-level blocks (sub-blocks). At least one spatial merge candidate can be derived for each low-level block.
[0261] Deriving spatial merge candidates may represent deriving spatial merge candidates and adding them to a merge candidate list. Here, each of the merge candidates added to the merge candidate list may have different motion information.
[0262] Up to maxNumSpatialMergeCand spatial merge candidates can be derived. Here, maxNumSpatialMergeCand can be a positive integer including 0. In an example, maxNumSpatialMVPCand can be 5. MaxNumMergeCand can be the maximum number of merge candidates that can be included in a merge candidate list, and can be a positive integer including 0. Additionally, numMergeCand can represent the number of merge candidates included in an actual merge candidate list within a predefined MaxNumMergeCand. Additionally, the use of maxNumSpatialMergeCand, numMergeCand, MaxNumMergeCand does not limit the scope of the present invention. The encoding / decoding device can use the above information by using parameter values having the same meaning as numMergeCand and MaxNumMergeCand.
[0263] A method for deriving temporal merge candidates will be described.
[0264] Temporal merge candidates can be derived from blocks reconstructed in temporally adjacent pictures or reference pictures to be encoded / decoded. A reference picture temporally adjacent to the current block can represent a collocated picture (picture). Information about the collocated picture (e.g., at least one of an inter prediction indicator, a reference picture index, and motion vector information indicating a collocated block of the current block) can be sent from the encoder to the decoder for at least one coding block unit within a sequence / picture / strip / parallel block / CTU / CU. Additionally, information about the collocated picture can be signaled in multiple units. For example, information about the collocated picture can be signaled separately for a picture unit and a strip unit. Optionally, information about the collocated picture can be implicitly derived in the encoder / decoder by using at least one of a hierarchy according to the encoding / decoding order in the encoder / decoder, motion information of the currently encoded / decoded block or temporally and spatially neighboring blocks or both the current block and temporally and spatially neighboring blocks (e.g., an inter prediction indicator or a reference picture index or both an inter prediction indicator and a reference picture index), an inter prediction indicator of the collocated picture at the sequence / picture / strip / parallel block level, and reference picture index information. For example, when there is no information about the collocated picture at the strip level, information about the collocated picture at the picture level can be regarded as information about the collocated picture at the strip level.
[0265] Here, when deriving the temporal merge candidates for the current block, the position of the co-located picture and / or the co-located block within the co-located picture can be selected by using at least one piece of motion information of the temporally and spatially neighboring blocks that have been encoded / decoded. Accordingly, a block at the same position within the co-located picture can be selected based on the position of the current block. Optionally, the co-located block of the current block can be defined as the block that is shifted by a corresponding vector from the spatially same position of the current block within the selected co-located picture by using at least one piece of motion vector information of the temporally and spatially neighboring blocks that have been encoded / decoded.
[0266] Here, the motion information of the temporally and spatially neighboring blocks that have been encoded / decoded can be at least one of a motion vector, a reference picture index, an inter prediction indicator, a picture order count (POC), and information about the co-located picture at the level of the current encoded picture (or slice).
[0267] Deriving the temporal merge candidates can mean adding the derived temporal merge candidates to the merge candidate list when the motion information of the derived temporal merge candidates is different from the motion information of the merge candidate list.
[0268] The number of the thus-derived temporal merge candidates can be as many as maxNumTemporalMergeCand. Here, maxNumTemporalMergeCand can be a positive integer including 0. For example, maxNumTemporalMergeCand can be 1. However, the use of maxNumTemporalMergeCand does not limit the scope of the present invention. The encoder / decoder can use the above information by means of a parameter value having the same meaning as maxNumTemporalMergeCand.
[0269] In addition, the prediction using the temporal merge candidates can be referred to as TMVP (Temporal Motion Vector Prediction).
[0270] Figure 9 is a diagram illustrating a method of deriving temporal merge candidates according to an embodiment of the present invention.
[0271] Referring to Figure 9 , the temporal merge candidates can be derived in the block at position H or the block at position C3, where the block at position H exists outside the co-located block C at the spatially same position as the current block X to be encoded / decoded within the reference picture of the current picture to be encoded / decoded.
[0272] Here, when the temporal merge candidates are likely to be derived from the block at position H, the temporal merge candidates can be derived from the block at position H. Otherwise, when the temporal merge candidates are not derived from the block at position H, the temporal merge candidates can be derived from the block at position C3. The order of deriving the temporal merge candidates can vary.
[0273] In addition, when the predetermined position or position C3 is intra-coded, temporal merge candidates can be derived in the blocks at position H or position C3. The co-located block of the current block can have a square shape or a non-square shape.
[0274] When the distance between the picture including the current block and the reference picture of the current block is different from the distance between the picture including the co-located block and the reference picture of the co-located block, temporal merge candidates can be derived by scaling the motion vector of the co-located block. The scaling of the motion vector can be performed according to the ratio of tb to td (in the example, the ratio = (tb / td)). Here, td can represent the difference between the POC of the co-located picture and the POC of the reference picture of the co-located block. In addition, tb can represent the difference between the POC of the picture to be encoded / decoded and the POC of the reference picture of the current block.
[0275] Derivation of sub-block based temporal merge candidates will be described.
[0276] Temporal merge candidates can be derived from the co-located sub-blocks to be encoded / decoded in units of sub-blocks having at least one of a size, shape, and depth smaller than that of the current block. For example, the sub-block can be a block having a horizontal or vertical length smaller than that of the current block, or a block having a depth deeper than that of the current block or a minimized shape, or can be a block included in the current block.
[0277] The co-located sub-blocks of the sub-blocks to be encoded / decoded can have a square shape or a non-square shape. In addition, the co-located block of the current block can be divided in units of sub-blocks having at least one of a size, shape, and depth smaller or deeper than that of the current block. At least one temporal merge candidate can be derived for each sub-block.
[0278] When deriving at least one temporal merge candidate by performing division in units of sub-blocks, at least one of the size, shape, and depth of the sub-block can be used to derive temporal merge candidates in the co-located sub-blocks at position H or C3 or both H and C3 described in Figure 9 Optionally, at least one temporal merge candidate can be derived by using the motion information (at least one of a motion vector, a reference picture index, an inter prediction indicator, and a POC) stored in each sub-block unit of the co-located block associated with the position moved according to any motion information derived from the neighboring blocks of the current block.
[0279] In addition, it is possible to determine whether to derive a time merge candidate based on a sub-block by checking for the existence of motion information of a sub-block at a predefined position of a co-located block corresponding to a position moved according to random motion information derived from neighboring blocks of the current block. For example, a time merge candidate based on a sub-block can be derived only when there is motion information in the sub-blocks at the predefined positions of the co-located blocks. Here, the sub-blocks at the predefined positions of the co-located blocks can be the sub-blocks at the central positions.
[0280] In addition, when there is no available motion information at each sub-block unit of a co-located block corresponding to a position moved according to random motion information derived from neighboring blocks of the current block, at least one time merge candidate can be derived by using the motion information of the sub-blocks at the predefined positions of the co-located blocks for determining whether to derive a time merge candidate based on a sub-block.
[0281] When deriving a time merge candidate for the current block or a sub-block of the current block, the motion vectors of each reference picture list (e.g., L0 or L1 or both) brought from co-located sub-blocks within the co-located block can be scaled to motion vectors corresponding to any reference picture of the current block. Optionally, after obtaining multiple motion vectors by scaling the motion vectors generated from the co-located sub-blocks to motion vectors corresponding to at least one reference picture among all the reference pictures that can be referenced by the sub-blocks of the current block, at least one predicted block using the scaled motion vectors corresponding to each reference picture can be obtained. In addition, a predicted block for the current block or sub-block can be obtained by using a weighted sum of the obtained predicted blocks.
[0282] In addition, prediction based on a time merge candidate of a sub-block can be referred to as sub-block based time motion vector prediction (SbTMVP).
[0283] A method for deriving a time-spatial combined merge candidate based on a sub-block will be described.
[0284] It is possible to derive a merge candidate for the current block by dividing the current block into sub-blocks and by using at least one motion information of spatially adjacent sub-blocks and co-located sub-blocks within a co-located picture for each obtained sub-block unit.
[0285] Figure 10 is a diagram showing a method for deriving a time-spatial combined merge candidate based on a sub-block according to an embodiment of the present invention.
[0286] Figure 10 is a diagram showing a block structure in which the shaded area represented by an 8×8 current block is divided into four 4×4 sub-blocks (i.e., blocks A, B, C, and D). A time-spatial combined merge candidate based on a sub-block can be derived by using motion vector information of sub-blocks that are temporally and spatially adjacent to each sub-block. Here, the motion vector information can represent a motion vector, an inter-frame prediction indicator, a reference picture index, a POC, etc.
[0287] In Figure 10 , when deriving a residual signal based on motion compensation after dividing a current block into sub - blocks, motion information can be obtained by performing a scan starting from a sub - block located above the first sub - block A in a left - to - right direction. In an example, when encoding the first upper sub - block by using an intra - prediction method, the second upper sub - block can be scanned sequentially. In other words, the scan of the upper sub - blocks can be performed until an upper sub - block including available motion vector information is found.
[0288] In addition, after obtaining available motion information of the upper sub - block, available motion information can be obtained by performing a scan in a top - to - bottom direction at a sub - block c to the left of the first sub - block A.
[0289] In addition, after obtaining spatial adjacent motion information of a left - hand sub - block or an upper sub - block or both a left - hand sub - block and an upper sub - block, temporal motion information can be derived by obtaining motion information of a collocated sub - block of the current sub - block, or a collocated block, or both a collocated sub - block and a collocated block of the current sub - block.
[0290] Here, the position of a collocated block or a sub - block of a collocated block can be motion information at position C3 or position H described by Figure 9 , or can represent a collocated block at a position compensated by a motion vector derived adjacent to a current block or a sub - block of a collocated block. Motion information of at least one of blocks that are spatially adjacent and temporally adjacent to L0 or L1 or both L0 and L1 can be obtained by using the above method. In addition, based on at least one piece of obtained motion information, a sub - block - based spatio - temporal combination merge candidate of the current sub - block to be encoded / decoded can be derived.
[0291] In an example, for L0 or L1 or both L0 and L1, for at least one motion vector information derived in the described temporal / spatial sub - blocks of the sub - blocks of the current block, a scan of the motion vectors can be performed so as to be associated with the first reference picture of the current block. Subsequently, a motion vector or a spatio - temporal combination merge candidate of the first current sub - block A can be derived by using at least one of an average value, a maximum value, a minimum value, a median value, a weighted value, and a mode of up to three scaled motion vectors. In addition, spatio - temporal combination merge candidates of sub - blocks B, C, and D can be derived by using the above method.
[0292] In addition, prediction using a sub - block - based spatio - temporal combination merge candidate can be referred to as STMVP (Spatial - Temporal Motion Vector Prediction).
[0293] Derivation of additional merge candidates will be described.
[0294] As an additional merge candidate that can be used in the present invention, at least one of a derivable modified spatial merge candidate, a modified temporal merge candidate, a combined merge candidate, and a merge candidate having a predetermined motion information value can be derived.
[0295] Here, deriving an additional merge candidate can mean that when there is a merge candidate having motion information different from that of the merge candidates existing in the existing merge candidate list, the corresponding merge candidate is added to the merge candidate list.
[0296] A modified spatial merge candidate can mean a merge candidate obtained by modifying the motion information of at least one of the spatial merge candidates derived by using the above method.
[0297] A modified temporal merge candidate can mean a merge candidate obtained by modifying the motion information of at least one of the temporal merge candidates derived by using the above method.
[0298] A combined merge candidate can mean a merge candidate that uses at least one motion information of merge candidates, where the merge candidates are spatial merge candidates, temporal merge candidates, modified spatial merge candidates, modified temporal merge candidates, combined merge candidates, and merge candidates having a predetermined motion information value existing in the merge candidate list. Here, the combined merge candidate can mean a combined bidirectional prediction merge candidate. Additionally, the prediction using the combined merge candidate can be referred to as CMP (Combined Motion Prediction).
[0299] A merge candidate having a predetermined motion information value can mean a zero merge candidate having a motion vector (0, 0). Additionally, the prediction using the merge candidate having a predetermined motion information value can be referred to as ZMP (Zero Motion Prediction).
[0300] At least one of a modified spatial merge candidate, a spatial merge candidate, a modified temporal merge candidate, a temporal merge candidate, a combined merge candidate, and a merge candidate having a predetermined motion information value can be derived for each sub-block of the current block, and the merge candidates derived for each sub-block can be added to the merge candidate list.
[0301] Inter-frame prediction information can be derived in units of sub-blocks having at least one of a size, shape, and depth smaller or deeper than that of the current block to be encoded / decoded. In an example, the size can represent a horizontal size or a vertical size or both a horizontal size and a vertical size.
[0302] When inter-frame prediction information is derived in units of sub-blocks of the current block, the encoder / decoder can derive the inter-frame prediction information by using at least one of a bidirectional matching method and a template matching method.
[0303] When using the bidirectional matching method, an initial motion vector list can be configured. When configuring the initial motion vector list, the motion vectors adjacent to the current block can be used.
[0304] In an example, the initial motion vector list may be configured by using the predicted motion vector candidates of the AMVP mode of the current block.
[0305] In another example, the initial motion vector list may be configured by using the merge candidates of the merge mode of the current block.
[0306] In another example, the initial motion vector list may be configured with the unidirectional motion vectors of L0 or L1 or both L0 and L1 of the merge mode of the current block.
[0307] In another example, the initial motion vector list may be configured with the motion vectors of the remaining blocks other than the merge mode of the current block.
[0308] In another example, the initial motion vector list may be configured by combining at least N motion vectors of the above examples. Here, N may represent a positive integer greater than 0.
[0309] In another example, the initial motion vector list may be configured with the motion vectors in one direction of list 0 or list 1.
[0310] Figure 11 is a diagram illustrating a method for deriving inter-frame prediction information by using a bidirectional matching method according to an embodiment of the present invention.
[0311] Referring to Figure 11 , when the motion vector existing in the initial motion vector list is MV0 existing in the L0 list, in the reference picture in the opposite direction, MV1 existing on the same trajectory as MV0 and indicating the block that best matches the block indicated by MV0 may be derived. Here, the MV having the minimum SAD (Sum of Absolute Differences) between the blocks indicated by MV0 and MV1 may be derived as the inter-frame prediction information of the current sub-block.
[0312] Figure 12 is a diagram illustrating a method for deriving inter-frame prediction information by using a template matching method according to an embodiment of the present invention.
[0313] By using Figure 12 the template defined in, the neighboring blocks of the current block may be used as templates. Here, the horizontal (width) and vertical (height) dimensions of the template may be the same as or different from the horizontal (width) and vertical (height) dimensions of the current block.
[0314] In an example, above the current block (Cur block) may be used as a template.
[0315] In another example, the left part of the current block may be used as a template.
[0316] In another example, the left and upper parts of the current block can be used as templates.
[0317] In another example, in the reference picture (Ref0) of the current picture (Cur pic), the upper part, the left part, or both the upper and left parts of the co-located block of the current block can be used as templates.
[0318] In another example, the motion vector (MV) having the minimum sum of absolute differences (SAD) between the template of the current block and the template of the reference block can be derived as the inter-frame prediction information of the current sub-block.
[0319] When the inter-frame prediction information is derived in units of sub-blocks of the current block, luminance compensation can be performed. For example, the luminance changes of the spatially adjacent samples of the current block sampled at at least N samples and the luminance changes of the spatially adjacent samples of the reference block can be approximated by using a linear model, where N is any positive integer. Additionally, the linear model can be applied to the block to which the motion compensation of the current sub-block is applied to perform luminance compensation.
[0320] When the inter-frame prediction information is derived in units of sub-blocks of the current block, affine-based spatial motion prediction and compensation can be performed. For example, for the motion vector of the upper-left coordinate of the current block and the motion vector of the upper-right of the current block, motion vectors can be generated in units of sub-blocks of the current block by using an affine transformation formula. Additionally, motion compensation can be performed by using the generated motion vectors.
[0321] Figure 13 is a diagram showing a method for deriving inter-frame prediction information based on overlapped block motion compensation (OMBC) according to an embodiment of the present invention.
[0322] When the inter-frame prediction information is derived in units of sub-blocks of the current block, a prediction block based on overlapped block motion compensation (OBMC) of the sub-blocks of the current block can be generated by combining the block compensated by using the inter-frame prediction information of the current block with at least one sub-block compensated by using the inter-frame prediction information of at least one of the sub-blocks at the left, right, upper, and lower positions included in the current block.
[0323] In an example, the execution can be applied only to the sub-blocks at the boundary positions existing inside the current block.
[0324] In another example, the execution can be applied to all the sub-blocks inside the current block.
[0325] In another example, the execution can be applied to the sub-blocks at the left boundary position existing inside the current block.
[0326] In another example, the execution can be applied to the sub-blocks at the right boundary position existing inside the current block.
[0327] According to an embodiment of the present invention, a picture may be encoded / decoded by dividing the picture according to the number of sub-block units. Units and blocks having the same meaning may be used.
[0328] Figure 14 is a diagram showing quadtree partitioning, symmetric binary tree partitioning, and asymmetric binary tree partitioning according to an embodiment of the present invention. In Figure 14 where w may represent the horizontal size of a block and h may represent the vertical size of the block.
[0329] Referring to Figure 14 , quadtree partitioning is a partitioning shape in which one block is divided into four sub-blocks, and the horizontal and vertical sizes of the four sub-blocks may be half of the horizontal and vertical sizes of the block before division.
[0330] Binary tree partitioning is a partitioning shape in which one block is divided into two sub-blocks, and may include symmetric binary tree partitioning (symmetric partitioning) or asymmetric binary tree partitioning (asymmetric partitioning). Here, symmetric binary tree partitioning may include horizontal direction symmetric partitioning and vertical direction symmetric partitioning. In addition, asymmetric binary tree partitioning may include horizontal direction asymmetric partitioning or vertical direction asymmetric partitioning, or both horizontal direction asymmetric partitioning and vertical direction asymmetric partitioning. In addition, the leaf nodes of the binary tree may represent CUs.
[0331] The nodes obtained by symmetric binary tree partitioning may have the same size. In addition, the nodes obtained by asymmetric binary tree partitioning may have different sizes.
[0332] According to an embodiment of the present invention, as a partitioning structure, there may be quadtree (QT) partitioning.
[0333] Referring to Figure 14 , one CTU may be recursively divided into multiple CUs by using a quadtree structure. Whether to use intra prediction or inter prediction may be determined based on the CU unit.
[0334] In an example, one CU may be divided into at least M PUs. Here, M may be a positive integer equal to or greater than 2.
[0335] In another example, one CU may be divided into at least N TUs by using a quadtree structure. Here, N may be a positive integer equal to or greater than 2.
[0336] According to an embodiment of the present invention, as a partitioning structure, there may be binary tree partitioning after quadtree. Binary tree partitioning after quadtree may represent a partitioning structure in which quadtree partitioning is preferentially applied and then binary tree partitioning is applied. Here, the leaf nodes of the quadtree or the leaf nodes of the binary tree may represent CUs.
[0337] In the example, after quadtree partitioning, a CTU can be recursively partitioned into two or four CUs by using a binary tree. Here, when a CU is split into two, the partitioning can be performed by using a binary tree (BT) structure, and when a CU is split into four, the partitioning can be performed by using a quadtree structure. Since the CTU is partitioned by a quadtree and then a binary tree, the CUs can have a square shape or a non-square (rectangular) shape.
[0338] When partitioning a CU by using a binary tree after quadtree partitioning, at least one of a first flag (information indicating whether quadtree partitioning is performed or whether further partitioning is performed, or both whether quadtree partitioning is performed and whether further partitioning is performed) and a first index (information indicating whether horizontal symmetric partitioning or vertical symmetric partitioning is performed and whether further partitioning is performed, or both whether horizontal symmetric partitioning or vertical symmetric partitioning is performed and whether further partitioning is performed) can be signaled. Here, when the first flag indicates a first value, it can indicate partitioning by using a quadtree structure, and when the first flag indicates a second value, it can indicate not performing further partitioning. Additionally, when the first index indicates a first value, it can indicate not performing further partitioning, when the first index indicates a second value, it can indicate horizontal symmetric partitioning, and when the first index indicates a third value, it can indicate vertical symmetric partitioning. When the first flag indicates the second value, the first index can be signaled. Additionally, when it is determined that further partitioning of a CU is not possible based on the size or depth of the CU or both the size and depth of the CU, the first flag or the first index or both the first flag and the first index may not be signaled.
[0339] Figure 15 is a diagram showing symmetric binary tree partitioning after quadtree partitioning according to an embodiment of the present invention. In Figure 15 it, the QT split flag can indicate whether quadtree partitioning is performed, the BT split flag can indicate whether binary tree partitioning is performed, and the BT split type can indicate whether horizontal partitioning (or partitioning in the horizontal direction) or vertical partitioning (or partitioning in the vertical direction) is performed.
[0340] Referring to Figure 15 it, a CTU can be partitioned by using a quadtree structure. Additionally, the leaf nodes of the quadtree can be further partitioned by using a binary tree structure. Here, the leaf nodes of the quadtree or the leaf nodes of the binary tree can represent CUs.
[0341] In the binary tree partitioning structure after the quadtree, a CU can be used as a unit for performing prediction and transformation without further partitioning it. In other words, in the binary tree partitioning structure after the quadtree, the CUs, PUs, and TUs can have the same size. Additionally, it can be determined on a per-CU basis whether to use intra prediction or inter prediction. In other words, in the binary tree partitioning structure after the quadtree, at least one of intra prediction, inter prediction, transformation, inverse transformation, quantization, dequantization, entropy encoding / entropy decoding, and in-loop filtering can be performed in units of square blocks or non-square (rectangular) blocks.
[0342] A CU can include one luma (Y) component block and two chroma (Cb / Cr) component blocks. Additionally, a CU can include one luma component block or two chroma component blocks. Additionally, a CU can include one luma component, a Cr chroma component block, or a Cb chroma component block.
[0343] According to an embodiment of the present invention, the quadtree partitioning after the binary tree can exist as a partitioning structure.
[0344] According to an embodiment of the present invention, the combined quadtree and binary tree partitioning can exist as a partitioning structure. The combined quadtree and binary tree partitioning can represent a partitioning structure that applies the quadtree partitioning and the binary tree partitioning without priority. In the above-mentioned binary tree partitioning after the quadtree, the quadtree partitioning is preferentially applied. However, in the combined quadtree and binary tree partitioning, the quadtree partitioning is not prior and the binary tree partitioning can be applied first.
[0345] A CTU can be recursively partitioned into two or four CUs by using the combined quadtree and binary tree partitioning structure. In the combined quadtree and binary tree partitioning structure, either the quadtree partitioning or the binary tree partitioning can be applied to a CU. Here, when a CU is split into two, the partitioning can be performed by using a binary tree, and when a CU is split into four, the partitioning can be performed by using a quadtree. Additionally, since the CUs are obtained by partitioning the CTU by using the combined quadtree and binary tree structure, the CUs can have a square or non-square (rectangular) shape.
[0346] By using the block partitioning structure of combined quadtree and binary tree shapes, the picture can be encoded / decoded in all non-square block shapes having a predetermined horizontal size and vertical size or a larger size.
[0347] The luminance signal and the chrominance signal within a CTU can be partitioned by block partitioning structures different from each other. For example, in the case of a specific slice (I slice), the luminance signal and the chrominance signal within the CTU can be partitioned by block partitioning structures different from each other. In the case of other slices (P slice or B slice), the luminance signal and the chrominance signal within the CTU can be partitioned by the same block partitioning structure. Here, the Cb signal and the Cr signal can use different intra prediction modes, and entropy coding / decoding can be performed on the intra prediction mode of each of the Cb signal and the Cr signal. The intra prediction mode of the Cb signal can be entropy coded / decoded by using the intra prediction mode of the Cr signal. Conversely, the intra prediction mode of the Cr signal can be entropy coded / decoded by using the intra prediction mode of the Cb signal.
[0348] A method for deriving intra prediction or inter prediction information or both intra prediction and inter prediction information based on sub-blocks will be described.
[0349] Hereinafter, a sub-block partitioning method will be described.
[0350] A current block (CU) can have a square or rectangular shape or both a square and a rectangular shape, and can represent a leaf node of at least one of a quadtree, a binary tree, and a ternary tree. In addition, at least one of intra prediction or inter prediction or both intra prediction and inter prediction, primary / secondary transform and inverse transform, quantization, dequantization, entropy coding / decoding, and in-loop filtering coding / decoding can be performed in units of at least one of the size, shape, and depth of the current block (CU).
[0351] The current block can be partitioned into at least one of symmetric sub-blocks or asymmetric sub-blocks or both symmetric sub-blocks and asymmetric sub-blocks. Intra prediction or inter prediction information or both intra prediction and inter prediction information can be derived for each sub-block. Here, symmetric sub-blocks can represent sub-blocks obtained by using at least one of the quadtree, binary tree, and ternary tree partitioning structures described. In addition, asymmetric sub-blocks can represent sub-blocks obtained by using the partitioning structures to be described in conjunction with Figure 14 but are not limited thereto. Asymmetric sub-blocks can represent at least one sub-block having a shape other than a square or rectangular shape or both. Figure 16 In
[0352] When the current block is partitioned into two sub-blocks, the two sub-blocks can be respectively defined as a first sub-block and a second sub-block. In addition, the first sub-block can be referred to as sub-block A, and the second sub-block can be referred to as sub-block B. Figure 16
[0353] When the current block is partitioned into at least one of symmetric sub - blocks or asymmetric sub - blocks or both symmetric and asymmetric sub - blocks, the minimum size of the sub - block can be defined as M×N. Here, M and N can respectively represent positive integers greater than 0. In addition, M and N can have the same or different values from each other. In an example, a 4×4 block can be defined as the minimum - size sub - block.
[0354] When the current block is partitioned into at least one of symmetric sub - blocks or asymmetric sub - blocks or both symmetric and asymmetric sub - blocks, for a specific block size or a specific block depth or a smaller size / deeper depth, further block partitioning may not be performed. Information about the specific block size or the specific block depth can be entropy - encoded / entropy - decoded in at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a parallel block header, a slice header, a coding tree unit (CTU), and a coding unit (CU).
[0355] Information about the specific block size or the specific block depth can be entropy - encoded / entropy - decoded for each of the luminance and chrominance signals and can have different parameter values from each other.
[0356] Information about the specific block size or the specific block depth can be entropy - encoded / entropy - decoded for each of the Cb and Cr signals and can have different parameter values.
[0357] Information about the specific block size of the specific block depth can be entropy - encoded / entropy - decoded for each upper - level rank and can have different parameter values.
[0358] Information about the specific block size or the specific block depth can be determined based on a comparison between the size or the depth of the current block and a predetermined threshold. The predetermined threshold can represent a reference size or depth for determining the block structure. Additionally, the predetermined threshold can be represented in the form of at least one of the minimum and maximum values of the reference size or depth. Additionally, the predetermined threshold can be a value predefined in the encoder / decoder, can be variably derived based on the coding parameters of the current block, or can be signaled through the bitstream.
[0359] In an example, when the size or the depth of the current block is equal to or less than / equal to or greater than the predetermined threshold, partitioning the current block into at least one sub - block may not be performed. For example, when the sum of the horizontal length and the vertical length of the current block is equal to or less than the predetermined threshold, partitioning the current block into at least one sub - block may not be performed.
[0360] For another example, when the size or depth of the current block is less than or greater than a predetermined threshold, partitioning the current block into at least one sub-block may not be performed. For example, when the sum of the horizontal length and the vertical length of the current block is less than a predetermined threshold, the current block may not be partitioned into at least one sub-block. Additionally, when the horizontal and vertical lengths of the current block are each less than a predetermined threshold, the current block may not be partitioned into at least one sub-block. Here, the predetermined threshold may be 8.
[0361] For another example, when the current block is a quadtree leaf node with a predetermined threshold depth, partitioning the current block into at least one sub-block may not be performed.
[0362] For another example, when the current block is a binary tree leaf node with a predetermined threshold depth, partitioning the current block into at least one sub-block may not be performed.
[0363] For another example, when the current block is a quadtree, binary tree, and / or ternary tree leaf node that performs motion prediction / compensation through an affine transformation formula, partitioning the current block into at least one sub-block or an asymmetric sub-block may not be performed.
[0364] For another example, when the current block is a quadtree, binary tree, and / or ternary tree leaf node that derives inter-frame prediction information by using at least one of a bidirectional matching method and a template matching method, partitioning the current block into at least one sub-block or an asymmetric sub-block may not be performed.
[0365] When partitioning the current block into at least one asymmetric sub-block, at least one of the sub-blocks obtained thereby may have an arbitrary shape other than a square and / or a rectangle.
[0366] For example, the current block may be partitioned into two sub-blocks by a straight line. In this case, the sub-blocks obtained thereby may be triangular, square (rectangular, trapezoidal), and hexagonal block shapes.
[0367] According to the present invention, when the current block is partitioned into at least one asymmetric sub-block, at least one of the sub-blocks obtained thereby may have a triangular shape.
[0368] Figure 16 is a diagram showing asymmetric partitioning according to an embodiment of the present invention. In Figure 16 , w may represent the horizontal size of the block, and h may represent the vertical size of the block.
[0369] In Figure 16 (a) of, when partitioning the current block into two sub-blocks, the current block may be partitioned into two triangular sub-blocks by a diagonal boundary extending from the upper left corner to the lower right corner of the current block. Here, the remaining area except for the upper right area (the second sub-block or sub-block B) of the current block may be defined as the first sub-block or sub-block A.
[0370] In Figure 16 (b) of, when partitioning the current block into two sub-blocks, the current block can be partitioned into two triangular sub-blocks by a diagonal boundary extending from the upper right to the lower left of the current block. Here, the remaining area except for the lower right area (the second sub-block or sub-block B) of the current block can be defined as the first sub-block or sub-block A.
[0371] Referring to Figure 16 (a) of and Figure 16 (b) of, when partitioning the current block into four sub-blocks, the current block can be partitioned into four triangular sub-blocks obtained by drawing a diagonal boundary from the upper left to the lower right and then drawing a diagonal boundary from the upper right to the lower left. Optionally, the current block can be partitioned into four triangular sub-blocks by drawing a diagonal boundary from the upper right to the lower left and then drawing a diagonal boundary from the upper left to the lower right.
[0372] In addition, when the motion prediction / compensation method of the current block (CU) is at least one of the skip mode and the merge mode, the partitioning of the triangular sub-block can be applied.
[0373] In Figure 16 (c) of, when the current block is partitioned into two sub-blocks, the remaining area except for the lower right area (the second sub-block or sub-block B) of the current block can be defined as the first sub-block or sub-block A.
[0374] In Figure 16 (d) of, when the current block is partitioned into two sub-blocks, the remaining area except for the lower left area (the second sub-block or sub-block B) of the current block can be defined as the first sub-block or sub-block A.
[0375] In Figure 16 (e) of, when the current block is partitioned into two sub-blocks, the remaining area except for the upper right area (the second sub-block or sub-block B) of the current block can be defined as the first sub-block or sub-block A.
[0376] In Figure 16 (f) of, when the current block is partitioned into two sub-blocks, the remaining area except for the upper left area (the second sub-block or sub-block B) of the current block can be defined as the first sub-block or sub-block A.
[0377] In Figure 16 (g) of, when the current block is partitioned into two sub-blocks, the shaped area composed of the upper part area, the lower part area, and the left part area of the current block can be defined as the first sub-block or sub-block A. In addition, the remaining area except for the first sub-block or sub-block A can be defined as the second sub-block or sub-block B.
[0378] In Figure 16In (h), when the current block is partitioned into two sub - blocks, the shaped region composed of the upper partial region, lower partial region, and right - hand partial region of the current block can be defined as the first sub - block or sub - block A. Additionally, the remaining region except the first sub - block or sub - block A can be defined as the second sub - block or sub - block B.
[0379] In Figure 16 In (i), when the current block is partitioned into two sub - blocks, the shaped region composed of the lower partial region, right - hand partial region, and left - hand partial region of the current block can be defined as the first sub - block or sub - block A. Additionally, the remaining region except the first sub - block or sub - block A can be defined as the second sub - block or sub - block B.
[0380] In Figure 16 In (j), when the current block is partitioned into two sub - blocks, the "П" - shaped region composed of the upper partial region, right - hand partial region, and left - hand partial region of the current block can be defined as the first sub - block or sub - block A. Additionally, the remaining region except the first sub - block or sub - block A can be defined as the second sub - block or sub - block B.
[0381] In Figure 16 In (k), when the current block is partitioned into two sub - blocks, the remaining region of the current block except the central region (the second sub - block or sub - block B) can be defined as the first sub - block or sub - block A.
[0382] Additionally, in Figure 16 from (a) to Figure 16 in (k), the first sub - block (or sub - block A) and the second sub - block (or sub - block B) defined can be interchanged with each other.
[0383] The encoder / decoder can store a table or list including multiple asymmetric partitioning shapes. The asymmetric partitioning shape of the current block determined in the encoder can be sent to the decoder in the form of an index or a flag. In other words, based on the information indicating multiple partitioning shapes (e.g., angle and distance information) and the table including the corresponding index, the encoder can send the index of the table to the decoder.
[0384] Additionally, the encoder / decoder can determine the asymmetric partitioning shape of the current block based on the coding parameters of the current block. Additionally, the encoder / decoder can determine the asymmetric partitioning shape of the current block based on the neighboring blocks of the current block.
[0385] When the current block is partitioned into at least one asymmetric sub - block, the sub - blocks obtained therefrom can have a horizontal size and / or a vertical size equal to or less than the horizontal size (w) and / or the vertical size (h) of the current block.
[0386] In Figure 16In , when the current block is partitioned into two sub - blocks, the horizontal dimension and / or vertical dimension of the sub - blocks may be smaller than the horizontal dimension and / or vertical dimension of the current block.
[0387] In Figure 16 from (c) to Figure 16 in (f), when the current block is partitioned into two sub - blocks, compared with the current block, the second sub - block may have a horizontal dimension of (3 / 4)×w and a vertical dimension of (3 / 4)×h, respectively.
[0388] In Figure 16 from (g) to Figure 16 in (h), when the current block is partitioned into two sub - blocks, compared with the current block, the second sub - block may have a horizontal dimension of (3 / 4)×w and a vertical dimension of (2 / 4)×h, respectively.
[0389] In Figure 16 from (i) to (j) of 16, when the current block is partitioned into two sub - blocks, compared with the current block, the second sub - block may have a horizontal dimension of (2 / 4)×w and a vertical dimension of (3 / 4)×h, respectively.
[0390] In Figure 16 in (k), compared with the current block, the second sub - block may have a horizontal dimension of (2 / 4)×w and a vertical dimension of (2 / 4)×h, respectively.
[0391] In addition, the above - mentioned ratios of the horizontal dimension and / or vertical dimension of the second sub - block may be predefined in the encoder or obtained based on the information signaled from the encoder to the decoder.
[0392] The current block (CU) may have a square shape, a rectangular shape, or both a square shape and a rectangular shape. Additionally, the current block can be partitioned into at least one asymmetric sub - block by using the above - mentioned method so that intra - prediction and / or inter - prediction information can be derived. Here, each sub - block can derive intra - prediction and / or inter - prediction information in units of the lowest - level sub - block unit, and the lowest - level sub - block can represent the smallest block unit with a predetermined size. For example, a 4×4 block size can be defined as the lowest - level sub - block.
[0393] In addition, information about the size of the lowest - level sub - block can be entropy - encoded / entropy - decoded in units of at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a parallel block header, a slice header, a CTU, and a CU.
[0394] The current block (CU) can represent a quadtree, a binary tree, and / or a ternary tree leaf node, and can perform intra - prediction and / or inter - prediction, primary / secondary transform and inverse transform, quantization, de - quantization, entropy - encoding / entropy - decoding, and / or in - loop filter encoding / decoding in units of sub - block size, shape, and / or depth.
[0395] The current block (CU) may represent a quadtree, binary tree, and / or ternary tree leaf node. At least one of the encoding / decoding processes such as intra / inter prediction, primary / secondary transform and inverse transform, quantization, dequantization, entropy encoding / decoding, and in-loop filter encoding / decoding of the current block may be performed in units of sub-block size, shape, and / or depth.
[0396] For example, when encoding the current block (CU), intra / inter prediction may be performed in units of sub-block size, shape, and depth, and the remaining processes (such as primary / secondary transform and inverse transform, quantization, dequantization, entropy encoding / decoding, and in-loop filter) other than intra / inter prediction may be performed in units of the size, shape, and / or depth of the current block.
[0397] For another example, when the current block is divided into two sub-blocks (e.g., a first sub-block and a second sub-block), the first sub-block and the second sub-block may derive different sets of intra prediction information.
[0398] For yet another example, when the current block is divided into two sub-blocks (e.g., a first sub-block and a second sub-block), the first sub-block and the second sub-block may derive different sets of inter prediction information.
[0399] For yet another example, when the current block is divided into two sub-blocks (e.g., a first sub-block and a second sub-block), the first sub-block and the second sub-block may derive combined intra and / or inter prediction information. Here, the first sub-block may derive inter prediction information, and the second sub-block may derive intra prediction information. Optionally, the first sub-block may derive intra prediction information, and the second sub-block may derive inter prediction information.
[0400] For yet another example, when encoding the current block (CU), primary / secondary transform and inverse transform may be performed in units of sub-block size, shape, and / or depth. Except for primary / secondary transform and inverse transform, the remaining processes (such as intra and / or inter prediction, quantization, dequantization, entropy encoding / decoding, and in-loop filter) may be performed in units of the size, shape, and / or depth of the current block.
[0401] For yet another example, when the current block is divided into two sub-blocks (e.g., a first sub-block and a second sub-block), the primary / secondary transform process and the inverse transform process of the first sub-block and / or the second sub-block may be skipped. Optionally, different primary / secondary transform processes and inverse transform processes may be performed.
[0402] For yet another example, when the current block is divided into two sub-blocks (e.g., a first sub-block and a second sub-block), the secondary transform process and the inverse transform process of the first sub-block and / or the second sub-block may be skipped. Optionally, different primary / secondary transform processes and different inverse transform processes may be performed.
[0403] For another example, when encoding the current block (CU), quantization and dequantization can be performed in units of sub-block size, shape, and / or depth. In addition to quantization and dequantization, the remaining processing (such as intra and / or inter prediction, primary / secondary transform and inverse transform, entropy encoding / decoding, and in-loop filtering) can be performed in units of the size, shape, and / or depth of the current block.
[0404] For another example, when the current block is divided into two sub-blocks (e.g., a first sub-block and a second sub-block), the quantization and dequantization processing of the first sub-block and / or the second sub-block can be skipped. Optionally, different quantization and dequantization processing can be performed.
[0405] For another example, when the current block is divided into two sub-blocks (e.g., a first sub-block and a second sub-block), the first sub-block can be quantized according to the quantization parameter set in the first encoding, and the second sub-block can be encoded / decoded by using a quantization parameter different from the initially set quantization parameter. Here, according to the method set in the encoder / decoder, the quantization parameter and / or offset of the second sub-block different from the quantization parameter set in the first encoding can be explicitly sent or implicitly derived.
[0406] As another example, when the current block is divided into two sub-blocks (e.g., a first sub-block and a second sub-block), the second sub-block can be quantized according to the quantization parameter set in the first encoding, and the first sub-block can be encoded / decoded by using a quantization parameter different from the initially set quantization parameter. Here, according to the method set in the encoder / decoder, the quantization parameter and / or offset of the first sub-block different from the quantization parameter set in the first encoding can be explicitly sent or implicitly derived.
[0407] As another example, when encoding the current block (CU), entropy encoding / decoding can be performed in units of sub-block size, shape, and / or depth. In addition to entropy encoding / decoding, the remaining processing (such as intra and / or inter prediction, primary / secondary transform and inverse transform, quantization, dequantization, and in-loop filtering) can be performed in units of the size, shape, and / or depth of the current block.
[0408] For another example, when encoding the current block (CU), in-loop filtering can be performed in units of sub-block size, shape, and / or depth. In addition to in-loop filtering, the remaining processing (such as intra and / or inter prediction, primary / secondary transform and inverse transform, quantization, dequantization, and entropy encoding / decoding) can be performed in units of the size, shape, and / or depth of the current block.
[0409] For yet another example, when the current block is divided into two sub-blocks (eg, a first sub-block and a second sub-block), in-loop filtering of the first sub-block and / or the second sub-block may be skipped. Alternatively, different in-loop filtering processes may be performed.
[0410] When the current block is divided into at least one symmetric sub-block and / or an asymmetric sub-block, a flag indicating whether partitioning into sub-blocks is performed and / or index information about the sub-block partition type may be signaled by block (CU) units through the bitstream. For example, the index information about the asymmetric sub-block partition type may be signaled at at least one level of sequence, picture, sub-picture, slice, parallel block, CTU and CU. In addition, a flag indicating whether asymmetric sub-block partitioning is possible may be signaled at the sequence level. Optionally, the flag indicating whether asymmetric sub-block partitioning is possible may be variably derived based on the coding parameters of the current block. For example, the flag indicating whether asymmetric sub-block partitioning is possible may not be signaled through the bitstream, but may be implicitly derived based on the coding parameters of the current block (e.g., the size of the current block, the prediction mode of the current block, the slice type, and the flag indicating whether asymmetric sub-block partitioning is possible, etc.). Here, the flag may be signaled by using Figure 16 The subblock partition type may be defined by at least one of the asymmetric subblock partition types described in (a) to (k) of 16, and then encoding / decoding of the current block may be performed. Optionally, the subblock partition type may be predefined in the encoder / decoder and different from Figure 16 16 (a) to (k) of 16. At least one of the sub-blocks obtained by partitioning may have any block shape other than a square and / or a rectangle. In addition, the sub-block partition type may include information about a direction for partitioning the current block into sub-blocks, a sub-block shape, a shape relationship between the current block and the sub-block, and / or a shape relationship between the sub-blocks.
[0411] For example, when at least one of the types of 16 (a) to 16 (k) is used, a flag indicating whether to perform sub-block-based encoding and decoding and / or a sub-block partition index (or sub-block partition type) can be sent by a bitstream signal, or can be variably derived based on the encoding parameters of the current block. Here, when the index information is explicitly sent, at least one of a truncated Rice binarization method, a k-order exponential Columbus (exp_golomb) binarization method, a restricted k-order exp_golomb binarization method, a fixed length binarization method, a unary binarization method, and a truncated unary binarization method can be used. In addition, after binarization, CABAC (ae (v)) can be finally used to encode / decode the current block.
[0412] In addition, for example, when Figure 16When at least one of two types from (a) to 16 (b) is used, a flag indicating whether to perform triangular sub-block partitioning on the current block (CU) can be signaled.
[0413] The flag can be signaled in units of at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a parallel block header, a strip header, a CTU, and a CU. Also, in the case of a specific strip (e.g., a B strip), the flag can be signaled.
[0414] In addition, when triangular sub-block partitioning is performed on the current block (CU), an index indicating at least one of the direction of partitioning the CU into triangular sub-blocks and the motion information of the triangular sub-blocks can be signaled. The index can be variably derived based on the coding parameters of the current block.
[0415] In addition, when the flag indicates a first value, it can indicate that motion prediction / compensation based on triangular sub-blocks is used to generate prediction samples of the current block (CU). Also, simultaneously, the index can be signaled only when the flag indicates the first value.
[0416] The index range can be from 0 to M. M can be a positive integer greater than 0. For example, M can be 39.
[0417] In addition, the encoder / decoder can store a table or list for deriving the direction of partitioning the current block into arbitrary sub-blocks and / or the motion information of the sub-blocks from the index.
[0418] Table 1 is an example of a look-up table showing the direction of partitioning the current block into triangular sub-blocks. Based on the above index, the direction of partitioning into triangular sub-blocks can be derived.
[0419] Table 1
[0420] merge_triangle_idx[xCb][yCb] 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 TriangleDir 0 1 1 0 0 1 1 1 0 0 0 0 1 0 0 0 0 1 1 1 merge_triangle_idx[xCb][yCb] 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 TriangleDir 1 0 0 1 1 1 1 1 1 1 0 0 1 0 1 0 0 1 0 0
[0421] Referring to Table 1, when TriangleDir has a first value of 0, it can indicate that the current block is partitioned into two triangular sub-blocks by a diagonal boundary from the upper left to the lower right. For example, it can represent Figure 16 the partitioning shape of (a). In addition, when TriangleDir has a second value of 1, it can indicate that the current block is partitioned into two triangular sub-blocks by a diagonal boundary from the upper right to the lower left. For example, it can represent Figure 16 the partitioning shape of (b). Also, the first value and the second value can be interchanged with each other.
[0422] In addition, the index (merge_triangle_idx) indicating the partitioning direction of the sub-blocks can be in the range of 0 to 39. The index can be signaled for the current block. The index information can be the same as the index information indicating the merge candidates to be used in each sub-block described in Table 2 below.
[0423] In addition, for example, the current block can be divided into two sub-blocks by a straight line, and prediction samples of the current block can be generated by performing motion prediction / compensation on the sub-blocks. A flag indicating whether motion prediction / compensation based on the sub-blocks obtained by straight-line partitioning is possible can be signaled.
[0424] The flag can be signaled in units of at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a parallel block header, a slice header, a CTU, and a CU. Also, in the case of a specific slice (e.g., a B slice), the flag can be signaled.
[0425] In addition, when the flag indicates a first value, information indicating whether to perform motion prediction / compensation based on the sub-blocks obtained by straight-line partitioning can be derived. Also, information indicating whether to perform motion prediction / compensation based on the sub-blocks obtained by straight-line partitioning can be variably derived based on the coding parameters of the current block. For example, information indicating whether to perform motion prediction / compensation based on the sub-blocks obtained by straight-line partitioning can be implicitly derived based on the coding parameters of the current block (e.g., the size of the current block, the prediction mode of the current block, the slice type, etc.).
[0426] In addition, when performing sub-block partitioning of the current block (CU) using a straight line, an index indicating at least one of the angle information and the distance information of the straight line can be signaled. The index can be variably derived based on the coding parameters of the current block.
[0427] In addition, simultaneously, the index can be signaled only when the flag indicates the first value.
[0428] The index range can be from 0 to M. M can be a positive integer greater than 0. For example, M can be 63.
[0429] In addition, the encoder / decoder can store a table or a list for deriving the direction of partitioning the current block into arbitrary sub-blocks and / or the motion information of the sub-blocks from the index.
[0430] In addition, at least one of the flag and the index that are entropy-encoded in the encoder and entropy-decoded in the decoder can use at least one of the following binaryization methods.
[0431] Truncated Rice binaryization method
[0432] k-th order Exp_Golomb binaryization method
[0433] Finite k-th order Exp_Golomb binarization method
[0434] Fixed-length binarization method
[0435] Unary binarization method
[0436] Truncated unary binarization method
[0437] Hereinafter, a method for deriving intra-block and / or inter-block prediction information will be described.
[0438] When the current block is divided into at least one symmetric and / or asymmetric sub-block, each of the obtained sub-blocks can derive the prediction information of the current block by using at least one of the following methods: deriving different strip intra-prediction information between sub-blocks, deriving different strip inter-prediction information between sub-blocks, and deriving combined intra / inter-prediction information between sub-blocks.
[0439] Inter-prediction information may represent motion information for motion prediction / compensation (e.g., at least one of a motion vector, an inter-prediction indicator, a reference picture index, a picture order count, a skip flag, a merge flag, a merge index, an affine flag, an OBMC flag, a bidirectional matching and / or template matching flag, and a bidirectional optical flow (BIO) flag). Additionally, inter-prediction information and motion information may be defined to have the same meaning.
[0440] Intra-prediction information may represent intra-prediction mode information for generating an intra-prediction block (e.g., at least one of an MPM flag, an MPM index, a selected mode set flag, a selected mode index, and a residual mode index).
[0441] Inter- and / or intra-prediction information can be explicitly sent from the encoder to the decoder through a bitstream, or inter- and / or intra-prediction information can be variably derived based on the shape, size, and / or depth of the current block and / or sub-block. Additionally, inter- and / or intra-prediction information can be variably derived based on the coding parameters of the current block and / or sub-block, or inter- and / or intra-prediction information can be signaled through a bitstream.
[0442] Hereinafter, a method for deriving inter-prediction information between sub-blocks will be described.
[0443] When the current block is divided into at least one symmetric and / or asymmetric sub-block, each of the obtained sub-blocks can derive different strip inter-prediction information. Here, each sub-block can derive inter-prediction information by using at least one of the inter-prediction methods such as a skip mode, a merge mode, an AMVP mode, motion prediction / compensation using an affine transformation equation, motion prediction / compensation based on bidirectional matching, motion prediction / compensation based on template matching, and motion prediction / compensation based on OBMC.
[0444] When the current block is divided into two sub - blocks, the first sub - block (or sub - block A) and / or the second sub - block (or sub - block B) can derive inter - frame prediction information of different strips. When deriving the inter - frame prediction information of the first sub - block, motion information can be derived by using at least one inter - frame prediction method among skip mode, merge mode, AMVP mode, motion prediction / compensation using affine transformation equations, motion prediction / compensation based on bidirectional matching, motion prediction / compensation based on template matching, and motion prediction / compensation based on OBMC. In addition, when deriving the inter - frame prediction information of the second sub - block, motion information can be derived by using at least one inter - frame prediction method among skip mode, merge mode, AMVP mode, motion prediction / compensation using affine transformation equations, motion prediction / compensation based on bidirectional matching, motion prediction / compensation based on template matching, and motion prediction / compensation based on OBMC.
[0445] When the current block is divided into two sub - blocks, each sub - block can derive motion information in units of the lowest - level sub - block. Here, the lowest - level sub - block can represent the smallest block unit with a predetermined value. For example, a 4×4 block size can be defined as the lowest - level sub - block.
[0446] When the current block is divided into two sub - blocks, all sub - blocks can perform different motion prediction / compensation based on the skip mode according to the shape of each sub - block. Here, the current block can explicitly send two different pieces of motion information (for example, at least one of a skip flag and / or merge index information and picture order count).
[0447] When the current block is divided into two sub - blocks, all sub - blocks can perform different motion prediction / compensation based on the merge mode according to the shape of each sub - block. Here, the current block can explicitly send two different pieces of motion information (for example, at least one of a merge flag and / or merge index information and picture order count).
[0448] When the current block is divided into two sub - blocks, the first sub - block and / or the second sub - block can perform motion prediction / compensation based on different strip motion information according to the merge mode. Here, the current block can derive two different motion information based on different merge modes in each sub - block by configuring a single merge candidate list. For example, the current block can configure a merge candidate list including N merge candidates by using spatial merge candidates, temporal merge candidates, combined merge candidates, and zero merge candidates, and then can derive motion information by using different merge candidates in each sub - block obtained by partitioning. N can represent a natural number greater than 0. When the merge candidate list is configured, if the corresponding merge candidate has bidirectional motion information, the merge candidate list can consist of only unidirectional prediction candidates to reduce memory bandwidth. For example, in bidirectional motion information, only the L0 or L1 motion information can be added to the list. Optionally, the average value or weighted sum of the L0 and L1 motion information can be added to the list. Additionally, when different merge candidates are used in each sub - block, a predefined value can be used.
[0449] For example, when different merge candidates are used for each sub - block, the Nth candidate in the merge candidate list can be used for the first sub - block (or sub - block A), and the Mth candidate in the merge candidate list can be used for the second sub - block (or sub - block B). N and M can be natural numbers including 0 and can be different from each other. Additionally, N and M can be predefined values in the encoder / decoder.
[0450] For another example, in order to use different merge candidates for each sub - block, the first merge index and the second merge index can be signaled for the first sub - block and the second sub - block respectively.
[0451] For yet another example, when different merge candidates are used for each sub - block, for the merge candidates (or index information of the merge candidates) corresponding to each sub - block, a merge candidate group (or merge candidate group list or table) can be defined and used. The merge candidate group can include a pair of merge candidates for each sub - block as elements. Additionally, when the number of merge candidates configured with spatial merge candidates, temporal merge candidates, combined merge candidates, zero merge candidates, etc. is N, each merge candidate corresponding to each sub - block can have a value from 0 to N - 1. Here, N is a natural number including 0.
[0452] Table 2 shows an example of a look - up table representing the merge candidates used in each sub - block.
[0453] Table 2
[0454] merge_triangle_idx[xCb][yCb] 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 A 1 0 0 0 2 0 0 1 3 4 0 1 1 0 0 1 1 1 1 2 B 0 1 2 1 0 3 4 0 0 0 2 2 2 4 3 3 4 4 3 1 merge_triangle_idx[xCb][yCb] 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 A 2 2 4 3 3 3 4 3 2 4 4 2 4 3 4 3 2 2 4 3 B 0 1 3 0 2 4 0 1 3 1 1 3 2 2 3 1 4 4 2 4
[0455] Referring to Table 2, as Figure 16As shown in (a) of or (b) of 16, when the current block is divided into two triangular sub - blocks, A can represent the first sub - block (or sub - block A), and B can represent the second sub - block (or sub - block B). Additionally, when the number of merge candidates such as spatial merge candidates, temporal merge candidates, combined merge candidates, zero merge candidates, etc. is 5, A and B can each have values from 0 to 4.
[0456] Additionally, the range of the index (merge_triangle_idx) indicating the index information of the merge candidates mapped to each sub - block can be from 0 to M. M can be a positive integer greater than 0. For example, in Table 2, M can be 39. The index can be signaled for the current block. Accordingly, the motion information of the sub - block can be derived based on the index. Additionally, the index information can be the same as the index information indicating the partition direction information of the sub - block described in Table 1.
[0457] Additionally, the encoder / decoder can store a table or list for deriving the direction of partitioning the current block into arbitrary sub - blocks and / or the motion information of the sub - blocks from the index.
[0458] When the current block is divided into two sub - blocks, the first sub - block can perform motion prediction / compensation based on bidirectional matching, and the second sub - block can perform motion prediction / compensation based on template matching. Here, the motion information of each sub - block (e.g., at least one of motion vector, inter - frame prediction indicator, reference picture index, and POC) can be explicitly sent from the encoder or can be implicitly derived in the encoder / decoder.
[0459] When the current block is divided into two sub - blocks, the first sub - block can perform motion prediction / compensation based on template matching, and the second sub - block can perform motion prediction / compensation based on bidirectional matching. Here, the motion information of each sub - block (e.g., at least one of motion vector, inter - frame prediction indicator, reference picture index, and POC) can be explicitly sent from the encoder or can be implicitly derived in the encoder / decoder.
[0460] When the current block is divided into two sub - blocks, the first sub - block can perform motion prediction / compensation by using the motion information of spatially neighboring blocks, and the second sub - block can derive inter - frame prediction information by using at least one of the inter - frame prediction methods such as skip mode, merge mode, AMVP mode, motion prediction / compensation using an affine transformation equation, motion prediction / compensation based on bidirectional matching, motion prediction / compensation based on template matching, and motion prediction / compensation based on OBMC.
[0461] When the current block is divided into two sub-blocks, the second sub-block may perform motion prediction / compensation by using the motion information of spatially neighboring blocks, and the first sub-block may derive inter-frame prediction information by using at least one inter-frame prediction method among a skip mode, a merge mode, an AMVP mode, motion prediction / compensation using an affine transform equation, motion prediction / compensation based on bidirectional matching, motion prediction / compensation based on template matching, and motion prediction / compensation based on OBMC.
[0462] In addition, based on the inter-frame prediction information (i.e., motion information) of the first sub-block and the second sub-block derived by the above method, inter-frame prediction may be performed for each of the first sub-block and the second sub-block, and thus prediction samples may be generated for each sub-block. In addition, the final prediction samples of the current block may be derived by a weighted sum of the prediction samples generated for the two sub-blocks.
[0463] In addition, the first sub-block and / or the second sub-block that perform motion prediction / compensation by using the motion information of spatially neighboring blocks may derive motion information in units of the lowest-rank sub-blocks having a predetermined size.
[0464] When the current block is divided into two sub-blocks, the motion information of the current block may be stored by using at least one of the motion information of the first sub-block, the motion information of the second sub-block, and the third motion information generated by the first sub-block and the second sub-block. The motion information may be stored in a unit of N×N size. In addition, the motion information may be stored in a motion information buffer of temporally neighboring pictures for temporal motion information prediction, and may be used to predict the motion information of spatially neighboring blocks. N may be a positive integer greater than 0 and may have at least one value among 2, 4, 8, 16, 32, 64, 128, and 256. For example, the motion information may be stored in a 4×4 unit.
[0465] For example, when the current block is divided into two sub-blocks, each sub-block may have unidirectional motion information. The area for storing the unidirectional motion information of the first sub-block may be defined as type 0 (sType = 0), and the area for storing the unidirectional motion information of the second sub-block may be defined as type 1 (sType = 1). Type 0 and type 1 may be defined conversely to each other. In addition, the area for storing the third motion information generated by the motion information of the first sub-block and the motion information of the second sub-block may be defined as type 2 (sType = 2). The motion information corresponding to at least one of type 0, type 1, and type 2 may be stored in each N×N motion information storage unit. For example, as Figures 18 to 21 shown, when the motion information storage unit is 4×4, the type may be derived for each 4×4 size.
[0466] In addition, the motion information storage unit may vary according to the motion information of the first sub-block, the motion information of the second sub-block, or the third motion information. In addition, the motion information storage unit may vary according to type 0, type 1, or type 2.
[0467] As described below, the motion information type in each motion information storage unit can be derived based on at least one of the number of horizontal / vertical blocks, the aspect ratio, and the partition direction information.
[0468] -minSb = min(numSbX, numSbY) - 1
[0469] Here, numSbX is the number of N×N horizontal blocks, and numSbY is the number of N×N vertical blocks. For example, when the motion information storage unit is 4×4, numSbX and numSbY respectively represent the number of 4×4 blocks in the horizontal and vertical directions. Min() represents a function for calculating the minimum value.
[0470] -CbRatio = (cbWidth > cbHeight)? (cbWidth / cbHeight) : (cbHeight / cbWidth)
[0471] Here, cbWidth represents the width of the current block, and cbHeight represents the height of the current block.
[0472] -(xSbIdx, ySbIdx) xSbIdx = 0 to numSbx - 1, ySbIdx = 0 to numSby - 1
[0473] Here, (xSbIdx, ySbIdx) represents the index of the N×N sub-block of the current block. For example, when the motion information storage unit is 4×4, (xSbIdx, ySbIdx) represents the index of each 4×4 sub-block.
[0474] For each N×N sub-block, the motion information type can be determined as follows.
[0475] xIdx = (cbWidth > cbHeight)? (xSbIdx / cbRatio) : xSbIdx
[0476] yIdx = (cbWidth > cbHeight)? ySbIdx : (ySbIdx / cbRatio)
[0477] When the sub-block partition direction is from the upper left to the lower right (i.e., when the partition direction information has the first value (0)),
[0478] sType = (xIdx ==? yIdx)? 2 : ((xIdx > yIdx)? 0 : 1)
[0479] When the sub-block partition is from the upper right to the lower left (i.e., when the partition direction information has a second value (1)),
[0480] sType = (xIdx == yIdx)? 2 : ((xIdx > yIdx)? 0 : 1)
[0481] When minSb is defined as min(numSbX, numSbY), the equation can be defined as follows.
[0482] sType = (xIdx + yIdx == minSb)? 2 : ((xIdx + yIdx < minSb)? 0 : 1)
[0483] As Figure 16 of (a) and Figure 16 of (b) shown, when the current block is diagonally divided into two sub-blocks, the area for storing the unidirectional motion information of the upper sub-block can be defined as type 0 (sType = 0), and the area for storing the unidirectional motion information of the lower sub-block can be defined as type 1 (sType = 1). In each motion information storage unit, the motion information corresponding to at least one of type 1 and type 2 can be stored. For example, as Figures 19 to 20 shown, when the motion information storage unit is 4×4, the type can be deduced for each 4×4 size.
[0484] When the sub-block partition direction is from the upper left to the lower right (i.e., when the partition direction information has a first value (0)),
[0485] sType = ((xIdx >= yIdx)? 0 : 1)
[0486] When the sub-block partition is from the upper right to the lower left (i.e., when the partition direction information has a second value (1)), sType = ((xIdx + yIdx < minSb)? 0 : 1)
[0487] Here, minSb can be defined as min(numSbX, numSbY).
[0488] In addition, as Figure 16 of (a) and Figure 16 of (b) shown, when the current block is diagonally divided into two sub-blocks, the motion information can be stored by only using type 2 (sType = 2) (i.e., the third motion information generated by the unidirectional motion information of the upper sub-block and the unidirectional motion information of the lower sub-block). For example, as Figure 21As shown, when the motion information storage unit is 4×4, only type 2 can be used for each 4×4 size.
[0489] In addition, as Figure 16 (a) of Figure 16 and (b) of
[0490] show, when the current block is diagonally divided into two sub-blocks, the motion information of the current block can be stored in the N×N motion information storage unit by using only the unidirectional motion information (sType = 0) of the upper sub-block. Figure 16 Figure 16 In addition, as (a) of
[0491] and (b) of Figure 16 show, when the current block is diagonally divided into two sub-blocks, the motion information of the current block can be stored in the N×N motion information storage unit by using only the unidirectional motion information (sType = 1) of the lower sub-block. Figure 16
[0492] As Figure 16 (a) of Figure 16 and (b) of show, when the current block is diagonally divided into two sub-blocks, only the unidirectional motion information of the upper sub-block or the unidirectional motion information of the lower sub-block can be stored in the N×N motion information storage unit based on the sub-block partition direction information.
[0492] Figure 16 For example, as Figure 16 (a) shows, when the current block is divided from the upper left to the lower right, the unidirectional motion information of the upper sub-block can be stored in the N×N motion information storage unit. Conversely, the unidirectional motion information of the lower sub-block can be stored in the N×N motion information storage unit.
[0493] For example, as Figure 16 (b) shows, when the current block is divided from the upper right to the lower left, the unidirectional motion information of the lower sub-block can be stored in the N×N motion information storage unit. Conversely, the unidirectional motion information of the upper sub-block can be stored in the N×N motion information storage unit.
[0494] As described below, when the motion information type derived according to the N×N motion information storage unit is type 2 (sType = 2), the third motion information can be derived based on the direction of the reference picture (or the type of the reference picture list) referred to by the unidirectional motion information of each sub-block.
[0495] When the unidirectional motion information of the first sub-block and the unidirectional motion information of the second sub-block refer to reference pictures in different directions (for example, when the motion information of the first sub-block refers to the L0 reference picture and the motion information of the second sub-block refers to the L1 reference picture), the third motion information can be generated in the form of bidirectional motion information by combining the unidirectional motion information of the first sub-block and the unidirectional motion information of the second sub-block.
[0496] In contrast, when the unidirectional motion information of the first sub-block and the unidirectional motion information of the second sub-block refer to reference pictures in the same direction (e.g., when the motion information of the first sub-block refers to the L0 reference picture and the motion information of the second sub-block refers to the L0 reference picture), the third motion information may be set to any one of the motion information of the first sub-block and the motion information of the second sub-block.
[0497] In addition, as described below, when the motion information type is type 2 (sType = 2), the third motion information can be derived according to the direction of the reference picture (or the type of the reference picture list) referred to by the unidirectional motion information of each sub-block.
[0498] When the unidirectional motion information of the first sub-block and the unidirectional motion information of the second sub-block refer to reference pictures in different directions, the motion information of the first sub-block can be used as the L0 motion information of the third motion information, and the motion information of the second sub-block can be used as the L1 motion information of the third motion information. When the unidirectional motion information of the first sub-block and the unidirectional motion information of the second sub-block commonly refer to the L0 reference picture, the third motion information can be set as follows.
[0499] - When the motion information of the first sub-block is used as the L0 motion information and there is a reference picture in the L1 reference picture that has the same POC value as the L0 reference picture indicated by the motion information of the second sub-block, the motion vector of the second sub-block can be used as the L1 motion vector, and the reference picture index of the L1 reference picture with the same POC value can be used as the L1 reference picture index.
[0500] - As described above, when there is no reference picture in the L1 reference picture that has the same POC value as the L0 reference picture indicated by the motion information of the second sub-block, the motion information of the second sub-block can be used as the L0 motion information. In addition, when there is a reference picture in the L1 reference picture that has the same POC value as the L0 reference picture indicated by the motion information of the first sub-block, the motion vector of the first sub-block can be used as the L1 motion vector, and the reference picture index of the L1 reference picture with the same POC value can be used as the L1 reference picture index.
[0501] - When no condition is satisfied, the motion information of the first sub-block can be used as the L0 motion information, and the L1 motion information can be set to unavailable. In other words, for the horizontal direction and the vertical direction, the L1 motion vector can be set to (0, 0), and the L1 reference picture index can be set to a value of -1. For another example, the motion information of the second sub-block can be used as the L0 motion information, and the L1 motion information can be set to unavailable. In other words, for the horizontal direction and the vertical direction, the L1 motion vector can be set to (0, 0), and the L1 reference picture index can be set to a value of -1.
[0502] - For another example, in the case where there is no process of searching for a reference picture having the same POC value in the L1 direction, the L0 motion vector in the L0 motion information can be set to a value derived by using at least one of the average value, the minimum value, and the maximum value of the motion vectors of the first sub-block and the second sub-block, and the L0 reference picture index can be set to the reference picture index of the first sub-block or the reference picture index of the second sub-block. The L1 motion vector can be set to (0, 0) for both the horizontal direction and the vertical direction, and the L1 reference picture index can be set to a value of -1.
[0503] - For another example, in the case where there is no process of searching for a reference picture having the same POC value in the L1 direction, the motion information of a defined sub-block that is either the first sub-block or the second sub-block can be set to the L0 motion information, the L1 motion vector can be set to (0, 0) for both the horizontal direction and the vertical direction, and the L1 reference picture index can be set to the value -1. For example, the motion information of the first sub-block can be used as the L0 motion information. For example, the motion information of the second sub-block can be used as the L0 motion information.
[0504] - For another example, in the case where there is no process of searching for a reference picture having the same POC value in the L1 direction, the motion information of the sub-block can be set to the L0 motion information based on the sub-block partition direction information that is either the motion information of the first sub-block or the motion information of the second sub-block, the L1 motion vector can be set to (0, 0) for both the horizontal direction and the vertical direction, and the L1 reference picture index can be set to the value -1. For example, when the current block is partitioned from the upper left to the lower right as shown in (a) of Figure 16 , the motion information of the upper sub-block can be set to the L0 motion information. When the current block is partitioned from the upper right to the lower left as shown in (b) of Figure 16 , the motion information of the lower sub-block can be set to the L0 motion information. On the contrary, when the current block is partitioned from the upper left to the lower right as shown in (a) of Figure 16 , the motion information of the lower sub-block can be set to the L0 motion information. When the current block is partitioned from the upper right to the lower left as shown in (b) of Figure 16 , the motion information of the upper sub-block can be set to the L0 motion information.
[0505] When the unidirectional motion information of the first sub-block and the unidirectional motion information of the second sub-block commonly refer to the L1 reference picture, the third motion information can be set as follows.
[0506] - When the motion information of the first sub-block is used as the L1 motion information and there is a reference picture in the L0 reference picture that has the same POC value as the L1 reference picture indicated by the motion information of the second sub-block, the motion vector of the second sub-block can be used as the L0 motion vector, and the reference picture index of the L0 reference picture having the same POC value can be used as the L0 reference picture index.
[0507] - As described above, when there is no reference picture in the L0 reference picture having the same POC value as the L1 reference picture indicated by the motion information of the second sub-block, the motion information of the second sub-block can be used as the L1 motion information. Additionally, when there is a reference picture in the L0 reference picture having the same POC value as the L1 reference picture indicated by the motion information of the first sub-block, the motion vector of the first sub-block can be used as the L0 motion vector, and the reference picture index of the L0 reference picture having the same POC value can be used as the L0 reference picture index.
[0508] - When no condition is satisfied, the motion information of the first sub-block can be used as the L1 motion information, and the L0 motion information can be set to unavailable. In other words, the L0 motion vector can be set to (0, 0) for both the horizontal and vertical directions, and the L0 reference picture index can be set to a value of -1. For another example, the motion information of the second sub-block can be used as the L1 motion information, and the L0 motion information can be set to unavailable. In other words, the L0 motion vector can be set to (0, 0) for both the horizontal and vertical directions, and the L0 reference picture index can be set to a value of -1.
[0509] - For yet another example, in the case where there is no process for searching for a reference picture having the same POC value in the L0 direction, the L1 motion vector in the L1 motion information can be set to a value derived by using at least one of the average value, minimum value, and maximum value of the motion vectors of the first sub-block and the second sub-block, and the L1 reference picture index can be set to the reference picture index of the first sub-block or the reference picture index of the second sub-block. The L0 motion vector can be set to (0, 0) for both the horizontal and vertical directions, and the L0 reference picture index can be set to a value of -1.
[0510] - For yet another example, in the case where there is no process for searching for a reference picture having the same POC value in the L0 direction, the motion information of the defined sub-block that is the first sub-block or the second sub-block can be set as the L1 motion information, the L0 motion vector can be set to (0, 0) for both the horizontal and vertical directions, and the L0 reference picture index can be set to the value -1. For example, the motion information of the first sub-block can be used as the L1 motion information. For example, the motion information of the second sub-block can be used as the L1 motion information.
[0511] - For yet another example, in the case where there is no process for searching for a reference picture having the same POC value in the L0 direction, the motion information of the sub-block can be set as the L1 motion information based on the sub-block partition direction information that is the motion information of the first sub-block or the motion information of the second sub-block, the L0 motion vector can be set to (0, 0) for both the horizontal and vertical directions, and the L0 reference picture index can be set to the value -1. For example, when asFigure 16 When dividing the current block from the upper left to the lower right as shown in (a) of Figure 16 When dividing the current block from the upper right to the lower left as shown in (b) of Figure 16 When dividing the current block from the upper left to the lower right as shown in (a) of Figure 16 When dividing the current block from the upper right to the lower left as shown in (b) of
[0512] Figure 17 is a diagram showing a method of deriving motion prediction information of a sub - block by using the lowest - level sub - block according to an embodiment of the present invention.
[0513] In Figure 16 In (c) of Figure 17 , when the current block is divided into two asymmetric sub - blocks and the first sub - block performs motion prediction / compensation by using the motion information of spatially adjacent blocks, referring to
[0514] In Figure 17 , the motion prediction / compensation of the first sub - block can be performed in units of the lowest - level sub - blocks of the first sub - block. Therefore, the motion information of the spatially adjacent lowest - level sub - blocks located to the left and / or above the lowest - level sub - block can be implicitly derived as the motion information of the first sub - block. Here, the size of the lowest - level sub - block can be 4×4. Here, the second sub - block can explicitly derive motion information by using the AMVP mode.
[0515] In Figure 17 , the motion information of the upper - left lowest - level sub - block in the lowest - level sub - blocks of the first sub - block can be derived by using at least one of the motion information of the spatially adjacent upper - left lowest - level sub - block, the motion information of the spatially adjacent upper lowest - level sub - block, and the motion information of the spatially adjacent upper - left lowest - level sub - block. Here, the motion information of the upper - left lowest - level sub - block can use the motion information of one of the spatially adjacent left, upper, and upper - left lowest - level sub - blocks. Optionally, the motion information can be derived based on at least one of the average value, mode, and weighted sum of up to three adjacent lowest - level sub - blocks.
[0515] In Figure 17 , the motion information of the lowest - level sub - blocks of the first sub - block can be derived by using at least one of the spatially adjacent left - hand lowest - level sub - block and / or the spatially adjacent upper - hand lowest - level sub - block.
[0516] In Figure 17In [the case where] there is no motion information in the spatially adjacent left lowest-level sub-block or the spatially adjacent upper lowest-level sub-block, or both the spatially adjacent left lowest-level sub-block and the spatially adjacent upper lowest-level sub-block, the motion information of the lowest-level sub-block of the first sub-block can be derived in the lowest-level sub-blocks to the left and / or above the spatially adjacent left lowest-level sub-block and / or the spatially adjacent upper lowest-level sub-block.
[0517] In Figure 17 In [the case where] there is no motion information in the spatially adjacent left lowest-level sub-block or the spatially adjacent upper lowest-level sub-block, or both the spatially adjacent left lowest-level sub-block and the spatially adjacent upper lowest-level sub-block, the motion information of the lowest-level sub-block of the first sub-block can be replaced by the motion information derived in the second sub-block through the AMVP mode.
[0518] In Figure 16 In [the case where] the current block is divided into two sub-blocks, the first sub-block can perform motion prediction / compensation by using at least one piece of motion information in the merge candidate list, and the second sub-block can derive the inter-frame prediction information by using at least one of the inter-frame prediction methods including the skip mode, the merge mode, the AMVP mode, motion prediction / compensation using an affine transformation equation, motion prediction / compensation based on bidirectional matching, motion prediction / compensation based on template matching, and motion prediction / compensation based on OBMC. Optionally, the second sub-block can perform motion prediction / compensation by using at least one piece of motion information in the merge candidate list, and the first sub-block can derive the inter-frame prediction information by using at least one of the inter-frame prediction methods including the skip mode, the merge mode, the AMVP mode, motion prediction / compensation using an affine transformation equation, motion prediction / compensation based on bidirectional matching, motion prediction / compensation based on template matching, and motion prediction / compensation based on OBMC.
[0519] According to the above example, the first sub-block and / or the second sub-block that perform motion prediction / compensation by using at least one piece of motion information in the merge candidate list can derive the motion information in units of the lowest-level sub-blocks with a predetermined size.
[0520] In Figure 16 In (c) of [the case where], when the current block is divided into two asymmetric sub-blocks, the motion information of the first sub-block can be implicitly derived by using at least one piece of motion information in the merge candidate list.
[0521] For example, the first piece of motion information in the merge candidate list of the current block can be derived as the motion information of the first sub-block.
[0522] For another example, it can be done by using in Figure 9Derive the motion information of the first sub-block from at least one piece of motion information derived from A0, A1, B0, B1, B2, C3, and H.
[0523] For another example, the lowest-level sub-block located to the left of the first sub-block can derive its motion information by using at least one piece of motion information derived from Figure 9 A0, A1, and B2. Additionally, the highest-level sub-block located above the first sub-block can derive its motion information by using at least one piece of motion information derived from Figure 9 B0, B1, and B2.
[0524] In Figure 16 , when the current block is divided into two sub-blocks, the first sub-block can perform motion prediction / compensation by using at least one piece of motion information from the motion vector candidate list used in the AMVP mode, and the second sub-block can derive inter-frame prediction information by using at least one of the inter-frame prediction methods including the skip mode, merge mode, AMVP mode, motion prediction / compensation using an affine transformation equation, motion prediction / compensation based on bidirectional matching, motion prediction / compensation based on template matching, and motion prediction / compensation based on OBMC. Optionally, the second sub-block can perform motion prediction / compensation by using at least one piece of motion information from the motion vector candidate list used in the AMVP mode, and the first sub-block can derive inter-frame prediction information by using at least one of the inter-frame prediction methods including the skip mode, merge mode, AMVP mode, motion prediction / compensation using an affine transformation equation, motion prediction / compensation based on bidirectional matching, motion prediction / compensation based on template matching, and motion prediction / compensation based on OBMC.
[0525] According to the above examples, the first sub-block and / or the second sub-block that perform motion prediction / compensation by using at least one piece of motion information from the motion vector candidate list used in the AMVP mode can derive motion information in units of the lowest-level sub-blocks with a predetermined size.
[0526] In Figure 16 (a), when the current block is divided into two asymmetric sub-blocks, the motion information of the first sub-block can be implicitly derived by using at least one piece of motion information from the motion vector candidate list used in the AMVP mode.
[0527] For example, the first piece of motion information from the motion vector candidate list used in the AMVP mode for the current block can be derived as the motion information of the first sub-block.
[0528] For another example, the motion information of the first sub-block can be derived by using a zero motion vector.
[0529] Hereinafter, a method for deriving intra-frame prediction information between sub-blocks will be described.
[0530] When the current block is divided into at least one or more symmetric / asymmetric sub-blocks, different intra prediction information between the sub-blocks can be derived for each sub-block thus obtained. Here, the sub-blocks of the current block can derive different intra prediction information in units of the lowest-level sub-blocks. The lowest-level sub-block can represent the smallest block unit having a predetermined size. A 4×4 block size can be defined as the size of the lowest-level sub-block.
[0531] In Figure 16 when the current block is divided into two sub-blocks, the first sub-block (or sub-block A) and / or the second sub-block (or sub-block B) can derive intra prediction information in units of the lowest-level sub-blocks having a predetermined size.
[0532] In Figure 16 in (c) of Figure 17 when the first sub-block performs intra prediction by using the intra prediction information of spatially adjacent blocks, referring to Figure 17 the motion prediction / compensation of the first sub-block can be performed in units of the lowest-level sub-blocks of the first sub-block, so that the intra prediction information of the spatially adjacent lowest-level sub-blocks located on the left side and / or above the lowest-level sub-blocks can be implicitly derived as the intra prediction information of the first sub-block. Here, the size of the lowest-level sub-block can be 4×4. Here, the intra prediction information of the second sub-block can be explicitly derived by using the reference sample points adjacent to the current block to minimize the distortion value of the second sub-block for the intra prediction mode information.
[0533] For example, after generating a prediction block having the size of the current block by using the reference sample points adjacent to the current block, the distortion value can derive the intra prediction mode that minimizes the sum of absolute differences (SAD) and / or the sum of absolute transform differences (SATD) only in the actual second sub-block area as the intra prediction mode of the second sub-block.
[0534] In Figure 17 the intra prediction mode of the upper-left lowest-level sub-block in the lowest-level sub-blocks of the first sub-block can be derived by using the intra prediction mode information of at least one of the spatially adjacent left, upper, and upper-left lowest-level sub-blocks. Here, the intra prediction mode information of one of the spatially adjacent left, upper, and upper-left lowest-level sub-blocks can be used as the intra prediction mode information of the upper-left lowest-level sub-block. Optionally, the intra prediction mode information of the current lowest-level sub-block can be derived by using at least one of the average value, mode value, and weighted sum of the intra prediction modes of up to three lowest-level adjacent sub-blocks.
[0535] In Figure 17 the intra prediction mode information of the lowest-level sub-blocks of the first sub-block can be derived by using at least one of the spatially adjacent left and / or upper lowest-level sub-blocks.
[0536] In Figure 17 when there is no intra prediction mode information in the spatially adjacent left and / or upper lowest-level sub-blocks, the intra prediction mode information of the lowest-level sub-block of the first sub-block can be derived in the lowest-level sub-block adjacent to the spatially adjacent left and / or upper lowest-level sub-block. "Adjacent" can mean "left" and / or "above".
[0537] In Figure 17 when there is no intra prediction mode information in the spatially adjacent left and / or upper lowest-level sub-blocks, the intra prediction mode information of the lowest-level sub-block of the first sub-block can be replaced with the intra prediction mode information derived in the second sub-block.
[0538] When generating an intra prediction block for the first sub-block, after generating at least one intra prediction block in units of the lowest-level sub-blocks, the final prediction block can be generated by using the weighted sum of the prediction blocks.
[0539] For example, according to the above method, after generating a prediction block (pred_1) by using the intra prediction mode implicitly derived in the first sub-block in units of the lowest-level sub-blocks, a prediction block (pred_2) can be generated by applying the intra prediction mode derived in the second sub-block to the lowest-level sub-block of the first sub-block. Therefore, the prediction block of the lowest-level sub-block of the first sub-block can be generated by using the weighted sum and / or average of pred_1 and / or pred_2.
[0540] Hereinafter, a method of deriving combined intra / inter prediction information between sub-blocks will be described.
[0541] When the current block is divided into at least one or more symmetric and / or asymmetric sub-blocks, different sets of intra and / or inter prediction information between the sub-blocks can be derived for each of the sub-blocks thus obtained. Here, the sub-blocks of the current block can derive different sets of intra and / or inter prediction information in units of the lowest-level sub-blocks. The lowest-level sub-block can represent the smallest block unit having a predetermined size. A 4×4 block size can be defined as the size of the lowest-level sub-block.
[0542] In Figure 16 when the current block is divided into two sub-blocks, the first sub-block can derive intra prediction information and the second sub-block can derive inter prediction information.
[0543] In Figure 16 when the current block is divided into two sub-blocks, the first sub-block can derive intra prediction information in units of the lowest-level sub-blocks and the second sub-block can derive inter prediction information.
[0544] In Figure 16In [reference], when the current block is divided into two sub - blocks, the first sub - block can derive intra - prediction information, and the second sub - block can derive inter - prediction information at the lowest - level sub - block unit.
[0545] In Figure 16 In [reference], when the current block is divided into two sub - blocks, the first sub - block can derive intra - prediction information at the lowest - level sub - block unit, and the second sub - block can also derive inter - prediction information at the lowest - level sub - block unit.
[0546] In Figure 16 In [reference], when the current block is divided into two sub - blocks, the first sub - block can derive inter - prediction information, and the second sub - block can derive intra - prediction information.
[0547] In Figure 16 In [reference], when the current block is divided into two sub - blocks, the first sub - block can derive inter - prediction information at the lowest - level sub - block unit, and the second sub - block can derive intra - prediction information.
[0548] In Figure 16 In [reference], when the current block is divided into two sub - blocks, the first sub - block can derive inter - prediction information, and the second sub - block can derive intra - prediction information at the lowest - level sub - block unit.
[0549] In Figure 16 In [reference], when the current block is divided into two sub - blocks, the first sub - block can derive inter - prediction information at the lowest - level sub - block unit, and the second sub - block can also derive intra - prediction information at the lowest - level sub - block unit.
[0550] Figure 22 is a flowchart showing a video decoding method according to an embodiment of the present invention.
[0551] Referring to Figure 22 , the decoder can obtain block - partitioning information of the current block (S2201). Here, the block - partitioning information may be index information indicating an index of a table including information indicating a plurality of predefined asymmetric partitioning shapes.
[0552] In addition, the decoder can partition the current block into a first sub - block and a second sub - block based on the block - partitioning information (S2202).
[0553] More specifically, the decoder can partition the current block into a first sub - block and a second sub - block by a straight line.
[0554] In addition, at least one of the angle information and the distance information of the straight line may be included in the information indicating a plurality of predefined asymmetric partitioning shapes.
[0555] In addition, the first sub - block and the second sub - block can have any one of the shapes of a triangle, a rectangle, a trapezoid, and a pentagon.
[0556] In addition, when the horizontal length and the vertical length of the current block are respectively less than a predetermined threshold, step S2202 may not be performed.
[0557] In addition, the decoder may respectively derive the motion information of the first sub-block and the motion information of the second sub-block (S2203).
[0558] More specifically, the decoder may respectively obtain the merge index of the first sub-block and the merge index of the second sub-block, and generate a merge candidate list. In addition, the decoder may derive the motion information of the first sub-block by using the merge candidate list and the merge index of the first sub-block, and derive the motion information of the second sub-block by using the merge candidate list and the merge index of the second sub-block.
[0559] In addition, a merge candidate list may be generated based on the current block.
[0560] In addition, the decoder may respectively generate a predicted sample point of the first sub-block and a predicted sample point of the second sub-block based on the motion information of the first sub-block and the motion information of the second sub-block (S2204).
[0561] In addition, the decoder may generate a predicted sample point of the current block by a weighted sum of the predicted sample point of the first sub-block and the predicted sample point of the second sub-block (S2205).
[0562] In addition, the decoder may store at least one of the motion information of the first sub-block, the motion information of the second sub-block, and the third motion information. Here, when the motion information of the first sub-block and the motion information of the second sub-block refer to reference pictures in the same direction, the third motion information may be derived as any one of the motion information of the first sub-block and the motion information of the second sub-block.
[0563] Figure 23 It is a flowchart showing a video encoding method according to an embodiment of the present invention.
[0564] Referring to Figure 23 , the encoder may determine the block partition structure of the current block (S2301).
[0565] In addition, the encoder may partition the current block into a first sub-block and a second sub-block based on the block partition structure (S2302).
[0566] More specifically, the encoder may partition the current block into a first sub-block and a second sub-block by a straight line.
[0567] In addition, at least one of the angle information and the distance information of the straight line may be included in the information indicating a plurality of predefined asymmetric partition shapes.
[0568] In addition, the first sub-block and the second sub-block may have any one of a triangular shape, a rectangular shape, a trapezoidal shape, and a pentagonal shape.
[0569] In addition, when the horizontal length and the vertical length of the current block are respectively less than a predetermined threshold, step S2302 may not be performed.
[0570] In addition, the encoder may respectively derive the motion information of the first sub-block and the motion information of the second sub-block (S2303).
[0571] In addition, the encoder may encode the block partition information based on the block partition structure (S2304).
[0572] Here, the block partition information may be index information indicating an index of a table including information indicating a plurality of predefined asymmetric partition shapes.
[0573] In addition, the encoder may respectively encode the merge index of the first sub-block and the merge index of the second sub-block based on the motion information of the first sub-block and the motion information of the second sub-block (S2305).
[0574] More specifically, the encoder may generate a merge candidate list. In addition, the encoder may encode the merge index of the first sub-block by using the merge candidate list and the motion information of the first sub-block, and encode the merge index of the second sub-block by using the merge candidate list and the motion information of the second sub-block.
[0575] Here, the merge candidate list may be generated based on the current block.
[0576] In addition, the encoder may store at least one of the motion information of the first sub-block, the motion information of the second sub-block, and the third motion information. Here, when the motion information of the first sub-block and the motion information of the second sub-block refer to reference pictures in the same direction, the third motion information may be derived as any one of the motion information of the first sub-block and the motion information of the second sub-block.
[0577] The computer-readable non-transitory recording medium according to the present invention may store a bitstream generated by the Figure 23 video encoding method described above.
[0578] The above embodiments may be executed in the same manner in the encoder and the decoder.
[0579] At least one or a combination of the above embodiments may be used to encode / decode a video.
[0580] The order applied to the above embodiments may be different between the encoder and the decoder, or the order applied to the above embodiments may be the same in the encoder and the decoder.
[0581] The above embodiments may be executed for each luminance signal and chrominance signal, or the above embodiments may be executed identically for the luminance and chrominance signals.
[0582] The block shape according to the above embodiments of the present invention may have a square shape or a non-square shape.
[0583] At least one of the syntax elements (flags, indexes, etc.) entropy-encoded in the encoder and entropy-decoded in the decoder may use at least one of the following binarization, de-binarization, entropy-encoding / entropy-decoding methods.
[0584] Truncated Rice binarization
[0585] k-th order Exp_Golomb binarization
[0586] Limited k-th order Exp_Golomb binarization
[0587] Fixed-length binarization
[0588] Unary binarization
[0589] Truncated unary binarization
[0590] Truncated binary binarization
[0591] The above embodiments of the present invention may be applied according to the size of at least one of the coding block, prediction block, transform block, block, current block, coding unit, prediction unit, transform unit, unit, and current unit. Here, the size may be defined as the minimum size or the maximum size or both the minimum size and the maximum size such that the above embodiments are applied, or may be defined as a fixed size for applying the above embodiments. In addition, in the above embodiments, the first embodiment may be applied to the first size, and the second embodiment may be applied to the second size. In other words, the above embodiments may be applied according to the size combination. In addition, when the size is equal to or greater than the minimum size and equal to or less than the maximum size, the above embodiments may be applied. In other words, when the block size is included in a specific range, the above embodiments may be applied.
[0592] For example, when the size of the current block is 8×8 or larger, the above embodiments may be applied. For example, when the size of the current block is only 4×4, the above embodiments may be applied. For example, when the size of the current block is 16×16 or smaller, the above embodiments may be applied. For example, when the size of the current block is equal to or greater than 16×16 and equal to or less than 64×64, the above embodiments may be applied.
[0593] The above embodiments of the present invention can be applied according to time layers. To identify the time layer to which the above embodiments can be applied, a corresponding identifier can be signaled, and the above embodiments can be applied to the specified time layer identified by the corresponding identifier. Here, the identifier can be defined as the lowest layer or the highest layer or both the lowest layer and the highest layer to which the above embodiments can be applied, or can be defined as indicating a specific layer to which the embodiments are applied. Additionally, a fixed time layer for applying the embodiments can be defined.
[0594] For example, when the time layer of the current image is the lowest layer, the above embodiments can be applied. For example, when the time layer identifier of the current image is 1, the above embodiments can be applied. For example, when the time layer of the current image is the highest layer, the above embodiments can be applied.
[0595] The stripe type or parallel block group type for applying the above embodiments of the present invention can be defined, and the above embodiments can be applied according to the corresponding stripe type or parallel block group type.
[0596] In the above embodiments, the method is described based on a flowchart having a series of steps or units, but the present invention is not limited to the order of the steps, and rather some steps can be executed simultaneously with other steps or in a different order. Additionally, those of ordinary skill in the art should understand that the steps in the flowchart are not mutually exclusive, and other steps can be added to the flowchart or some steps can be deleted from the flowchart without affecting the scope of the present invention.
[0597] The embodiments include various aspects of the examples. All possible combinations for each aspect may not be described, but those skilled in the art will be able to recognize different combinations. Therefore, the present invention can include all substitutions, modifications, and changes within the scope of the claims.
[0598] Embodiments of the present invention can be implemented in the form of program instructions executable by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include independent program instructions, data files, data structures, etc., or a combination of program instructions, data files, data structures, etc. The program instructions recorded on the computer-readable recording medium may be specially designed and constructed for the present invention, or well-known to those of ordinary skill in the computer software art. Examples of the computer-readable recording medium include: magnetic recording media (such as hard disks, floppy disks, and magnetic tapes); optical data storage media (such as CD-ROMs or DVD-ROMs); magneto-optical media (such as optical floppy disks); and hardware devices specially configured to store and execute program instructions (such as read-only memories (ROMs), random access memories (RAMs), flash memories, etc.). Examples of program instructions include not only machine language codes formatted by compilers, but also high-level language codes that can be implemented by computers using interpreters. The hardware devices may be configured to be operated by one or more software modules to perform the processing according to the present invention, or vice versa.
[0599] Although the present invention has been described based on specific items such as detailed elements, and limited embodiments and drawings, they are only provided to help a more comprehensive understanding of the present invention, and the present invention is not limited to the above embodiments. Those skilled in the art to which the present invention pertains should understand that various modifications and changes can be made based on the above description.
[0600] Therefore, the spirit of the present invention should not be limited to the above embodiments, and the entire scope of the appended claims and their equivalents will fall within the scope and spirit of the present invention.
[0601] Industrial Applicability
[0602] The present invention can be used for encoding or decoding images.
Claims
1. A video decoding method, the method comprising: Obtaining block partition information of a current block; Deriving first motion information of a first sub-block of the current block and second motion information of a second sub-block of the current block respectively; Generating first predicted samples of the first sub-block and second predicted samples of the second sub-block respectively based on the first motion information and the second motion information; Generating predicted samples of the current block through a weighted sum of the first predicted samples and the second predicted samples; And Based on the value of sType, storing third motion information obtained by using the first motion information of the first sub-block and the second motion information of the second sub-block, Wherein the block partition information is index information indicating an index of a table, and the table includes information indicating a plurality of predefined asymmetric partition shapes of the first sub-block and the second sub-block, Wherein, in response to the value of sType being 2 and the first motion information and the second motion information commonly referring to an L1 reference picture in the same direction, without referring to any L0 reference picture, the L1 motion information of the third motion information is set to be the same as the first motion information of the first sub-block or the second motion information of the second sub-block, the L0 motion vectors of the L0 motion information in the horizontal direction and the vertical direction are set to 0, and the L0 reference picture index of the L0 motion information is set to the value -1, and Wherein the width and height of the current block are greater than or equal to a threshold.
2. The video decoding method according to claim 1, Among them, The steps of deriving the first motion information of the first sub-block and the second motion information of the second sub-block respectively include: Obtaining a first merge index of the first sub-block and a second merge index of the second sub-block respectively; Generating a merge candidate list; Deriving the first motion information of the first sub-block by using the merge candidate list and the first merge index; and Deriving the second motion information of the second sub-block by using the merge candidate list and the second merge index.
3. The video decoding method according to claim 2, Among them, The merge candidate list is generated based on the current block.
4. The video decoding method according to claim 1, Among them, The block partition information indicates a partition line of the first sub-block and the second sub-block.
5. The video decoding method according to claim 4, Among them, The information indicating the plurality of predefined asymmetric partition shapes includes at least one of angle information and distance information of the partition line.
6. The video decoding method according to claim 1, Among them, The asymmetric partition shape is any one of a triangle, a rectangle, a trapezoid, and a pentagon.
7. A video encoding method, the method comprising: Determining a block partition structure of a current block; Deriving first motion information of a first sub-block of the current block and second motion information of a second sub-block of the current block respectively; Encoding block partition information based on the block partition structure; Encoding a first merge index of the first sub-block and a second merge index of the second sub-block respectively based on the first motion information and the second motion information; And Based on the value of sType, storing third motion information obtained by using the first motion information of the first sub-block and the second motion information of the second sub-block, Among them, the block partition information is index information indicating an index of a table, where the table includes information on a plurality of predefined asymmetric partition shapes indicating a first sub-block and a second sub-block. Among them, in response to the value of sType being 2 and the first motion information and the second motion information commonly referring to an L1 reference picture, without referring to any L0 reference picture, the L1 motion information of the third motion information is set to be the same as the first motion information of the first sub-block or the second motion information of the second sub-block, the L0 motion vectors of the L0 motion information of the third motion information are set to 0 for both the horizontal direction and the vertical direction, and the L0 reference picture index of the L0 motion information is set to the value -1, and where the width and height of the current block are greater than or equal to a threshold.
8. The video coding method according to claim 7, Among them, The steps of separately encoding a first merge index and a second merge index include: Generating a merge candidate list; Encoding the first merge index by using the merge candidate list and the first motion information; and Encoding the second merge index by using the merge candidate list and the second motion information.
9. The video coding method according to claim 8, Among them, The merge candidate list is generated based on the current block.
10. The video coding method according to claim 7, Among them, The block partition information indicates a partition line between the first sub-block and the second sub-block.
11. The video coding method according to claim 10, Among them, The information indicating the plurality of predefined asymmetric partition shapes includes at least one of angle information and distance information of the partition line.
12. The video coding method according to claim 7, Among them, The asymmetric partition shape is any one of a triangle, a rectangle, a trapezoid, and a pentagon.
13. A non-transitory computer-readable recording medium for storing a bitstream generated by a video coding method, Among them, The video coding method includes: Determining a block partition structure of a current block; Separately deriving first motion information of a first sub-block of the current block and second motion information of a second sub-block of the current block; Encoding block partition information based on the block partition structure; Encoding a first merge index of the first sub-block and a second merge index of the second sub-block respectively based on the first motion information and the second motion information; and Storing third motion information obtained by using the first motion information of the first sub-block and the second motion information of the second sub-block based on the value of sType, where the block partition information is index information indicating an index of a table, where the table includes information on a plurality of predefined asymmetric partition shapes indicating a first sub-block and a second sub-block. Wherein, in response to the value of sType being 2 and the first motion information and the second motion information both referring to the L1 reference picture, without referring to any L0 reference pictures, the L1 motion information of the third motion information is set to be the same as the first motion information of the first sub-block or the second motion information of the second sub-block, the L0 motion vectors of the L0 motion information of the third motion information are set to 0 for both the horizontal direction and the vertical direction, and the L0 reference picture index of the L0 motion information is set to the value -1, and Wherein, the width and height of the current block are greater than or equal to a threshold value.