Method and apparatus for image encoding / decoding and recording medium for storing bitstream
Patent Information
- Application Number
- KR1020200028646
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-28
- Filing Date
- 2020-03-06
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2040-03-06
Smart Images

Figure 112020024346421-PAT00092_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a method and apparatus for encoding / decoding images, and more specifically, to a method and apparatus for encoding / decoding video images based on adaptive transformation type selection. Background Technology
[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing across various application fields. As video data becomes higher in resolution and quality, the relative volume of data increases compared to conventional video data; consequently, transmission and storage costs increase when video data is transmitted using existing wired or wireless broadband lines or stored using existing storage media. To address these issues arising from the increase in video data resolution and quality, high-efficiency video encoding and decoding technologies for video with higher resolution and quality are required.
[0003] Various video compression technologies exist, such as inter-frame prediction technology that predicts pixel values in the current picture from previous or subsequent pictures, intra-frame prediction technology that predicts pixel values in the current picture using pixel information within the current picture, transformation and quantization technology for compressing the energy of residual signals, and entropy coding technology that assigns short codes to values with high frequency and long codes to values with low frequency; by utilizing these video compression technologies, video data can be effectively compressed for transmission or storage. The problem to be solved
[0004] The present invention aims to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] In addition, the present invention aims to provide a video encoding / decoding method and apparatus based on conversion, shuffling, rearrangement and / or flipping to improve encoding / decoding efficiency.
[0006] In addition, the present invention aims to provide an image encoding / decoding method and apparatus based on adaptive conversion type selection to improve encoding / decoding efficiency.
[0007] In addition, the present invention aims to provide an image encoding / decoding method and apparatus for improving image conversion efficiency.
[0008] In addition, the present invention aims to provide a recording medium storing a bitstream generated by the image encoding / decoding method or device of the present invention. means of solving the problem
[0009] A video decoding method according to one embodiment of the present invention comprises: a step of determining a horizontal transformation type and a vertical transformation type of a current block; a step of performing an inverse transformation on the current block based on the determined horizontal transformation type and vertical transformation type to derive a residual block of the current block; and a step of restoring the current block based on the residual block, wherein the step of determining the horizontal transformation type and the vertical transformation type may be performed based on at least one of the horizontal size and the vertical size of the current block regardless of the intra-block prediction mode of the current block when the current block is in an ISP (Intra Sub-block Partitions) mode.
[0010] In the above-described image decoding method, the step of determining the horizontal conversion type and the vertical conversion type may further include the step of setting implicit multiple conversion selection information.
[0011] In the above-described image decoding method, the step of setting the implicit multiple conversion selection information may be such that, when the current block is in ISP (Intra Sub-block Partitions) mode, the implicit multiple conversion selection information is set to a value indicating an implicit multiple conversion selection.
[0012] In the above-described image decoding method, the step of setting the implicit multiple transformation selection information may be such that, if the intra-frame prediction explicit multiple transformation selection allowance information indicates that explicit multiple transformation selection is not allowed, and the prediction mode of the current block is an intra-frame prediction mode, and a second inverse transformation is not performed on the current block, and the current block is not a matrix-based intra-frame prediction mode, the implicit multiple transformation selection information may be set to a value indicating implicit multiple transformation selection.
[0013] In the above image decoding method, when the implicit multiple transformation selection information indicates an implicit multiple transformation selection, the horizontal transformation type and the vertical transformation type may be determined based on whether the current block is in SBT (Sub-Block Transform) mode.
[0014] In the above video decoding method, the implicit multiple transformation selection information indicates an implicit multiple transformation selection, and if the current block is not in SBT (Sub-Block Transform) mode, the horizontal transformation type and the vertical transformation type can be determined regardless of the intra-frame prediction mode of the current block.
[0015] In the above video decoding method, when the current block is in ISP (Intra Sub-block Partitions) mode and a second inverse transformation is performed, regardless of the implicit multiple transformation selection information, the horizontal transformation type and the vertical transformation type can be determined as a first transformation indicating a DCT-2 based integer transformation.
[0016] In the above image decoding method, the step of determining the horizontal conversion type and the vertical conversion type may be such that, when the current block is a chrominance component, the horizontal conversion type and the vertical conversion type are determined as a first conversion indicating a DCT-2 based integer conversion, regardless of the implicit multiple conversion selection information.
[0017] A video encoding method according to another embodiment of the present invention comprises: a step of determining a horizontal conversion type and a vertical conversion type of a current block; a step of performing a conversion on a remaining block of the current block based on the determined horizontal conversion type and vertical conversion type; and a step of encoding the current block based on the converted remaining block, wherein the step of determining the horizontal conversion type and the vertical conversion type may be performed based on at least one of the horizontal size and the vertical size of the current block regardless of the intra-block prediction mode of the current block when the current block is in an ISP (Intra Sub-block Partitions) mode.
[0018] In the above video encoding method, the step of determining the horizontal conversion type and the vertical conversion type may further include the step of setting implicit multiple conversion selection information.
[0019] In the above video encoding method, the step of setting the implicit multiple conversion selection information may be such that, when the current block is in ISP (Intra Sub-block Partitions) mode, the implicit multiple conversion selection information is set to a value indicating an implicit multiple conversion selection.
[0020] In the above video encoding method, the step of setting the implicit multiple transformation selection information may be set to a value indicating implicit multiple transformation selection when the intra-frame prediction explicit multiple transformation selection allowance information indicates that explicit multiple transformation selection is not allowed, the prediction mode of the current block is an intra-frame prediction mode, a second inverse transformation is not performed on the current block, and the current block is not a matrix-based intra-frame prediction mode.
[0021] In the above video encoding method, when the implicit multiple transform selection information indicates an implicit multiple transform selection, the horizontal transform type and the vertical transform type may be determined based on whether the current block is in SBT (Sub-Block Transform) mode.
[0022] In the above video encoding method, the implicit multiple transform selection information indicates an implicit multiple transform selection, and if the current block is not in SBT (Sub-Block Transform) mode, the horizontal transform type and the vertical transform type can be determined regardless of the intra-frame prediction mode of the current block.
[0023] In the above video encoding method, when the current block is in ISP (Intra Sub-block Partitions) mode and a second inverse transformation is performed, the horizontal transformation type and the vertical transformation type can be determined as a first transformation indicating a DCT-2 based integer transformation, regardless of the implicit multiple transformation selection information.
[0024] In the above video encoding method, the step of determining the horizontal conversion type and the vertical conversion type may be such that, when the current block is a chrominance component, the horizontal conversion type and the vertical conversion type are determined as a first conversion indicating a DCT-2 based integer conversion, regardless of the implicit multiple conversion selection information.
[0025] A computer-readable recording medium according to another embodiment of the present invention is a non-transient computer-readable recording medium that stores a bitstream generated by an image encoding method, wherein the image encoding method comprises: a step of determining a horizontal conversion type and a vertical conversion type of a current block; a step of performing a conversion on a remaining block of the current block based on the determined horizontal conversion type and vertical conversion type; and a step of encoding the current block based on the converted remaining block, wherein the step of determining the horizontal conversion type and the vertical conversion type may be performed based on at least one of the horizontal size and the vertical size of the current block regardless of the intra-block prediction mode of the current block when the current block is in an ISP (Intra Sub-block Partitions) mode. Effects of the invention
[0026] According to the present invention, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency may be provided.
[0027] In addition, according to the present invention, a video encoding / decoding method and apparatus based on conversion, shuffling, rearrangement and / or flipping for improving encoding / decoding efficiency may be provided.
[0028] In addition, according to the present invention, an image encoding / decoding method and apparatus based on adaptive conversion type selection for improving encoding / decoding efficiency may be provided.
[0029] In addition, according to the present invention, an image encoding / decoding method and apparatus for improving image conversion efficiency may be provided.
[0030] In addition, according to the present invention, a recording medium storing a bitstream generated by the image encoding / decoding method or device of the present invention may be provided. Brief explanation of the drawing
[0031] FIG. 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied. FIG. 2 is a block diagram showing the configuration according to one embodiment of a decoding device to which the present invention is applied. Figure 3 is a diagram schematically showing the segmentation structure of an image when encoding and decoding an image. Figure 4 is a diagram illustrating an example of an in-screen prediction process. Figure 5 is a diagram illustrating an example of an inter-frame prediction process. Figure 6 is a diagram illustrating the process of transformation and quantization. FIG. 7 is a diagram illustrating the basis vector in the frequency domain of DCT-2 according to the present invention. FIG. 8 is a diagram illustrating the basis vectors in each frequency domain of DST-7 according to the present invention. Figure 9 is a diagram showing the distribution of average residual values according to the position within the 2Nx2N prediction unit (PU) of the 8x8 coding unit (CU) predicted in inter mode, obtained by experimenting with the “Cactus” sequence in a Low Delay-P profile environment. Figure 10 is a three-dimensional graph showing the residual signal distribution characteristics within the 2Nx2N prediction unit (PU) of the 8x8 coding unit (CU) predicted in inter-mode prediction. FIG. 11 is a diagram illustrating the distribution characteristics of residual signals in the 2Nx2N prediction unit (PU) mode of a coding unit (CU) according to the present invention. FIG. 12 is a diagram illustrating the residual signal distribution characteristics before and after shuffling of a 2Nx2N prediction unit (PU) according to the present invention. FIG. 13 is a diagram illustrating an example of 4x4 residual data rearrangement of a subblock according to the present invention. FIGS. 14(a) and FIGS. 14(b) are drawings for illustrating an example of a conversion unit (TU) splitting structure of a coding unit (CU) according to a prediction unit (PU) mode and a shuffling method of a conversion unit (TU). Figure 15 is a diagram illustrating the results of performing DCT-2 transformation and SDST transformation according to the residual signal distribution of the 2Nx2N prediction unit (PU). FIG. 16 is a diagram illustrating the SDST process according to the present invention. FIG. 17 is a diagram illustrating the distribution characteristics of the size of the division and residual absolute value of the conversion unit (TU) according to the partition mode of the prediction unit (PU) of the inter-frame predicted coding unit (CU) according to the present invention. FIG. 18 is a diagram illustrating the residual signal scanning sequence and relocation sequence of a conversion unit (TU) with a depth of 0 within a prediction unit (PU) according to one embodiment of the present invention. FIG. 19 is a flowchart illustrating the DCT-2 or SDST selective encoding process through rate-distortion optimization (RDO) according to the present invention. FIG. 20 is a flowchart illustrating the process of selecting and decoding DCT-2 or SDST according to the present invention. FIG. 21 is a flowchart illustrating a decoding process using SDST according to the present invention. FIGS. 22 and FIGS. 23 each show the locations where residual rearrangement is performed in the encoder and decoder according to the present invention. FIG. 24 is a diagram illustrating an embodiment of a decoding method using the SDST method according to the present invention. FIG. 25 is a diagram illustrating an embodiment of an encoding method using the SDST method according to the present invention. FIG. 26 is a diagram illustrating an example of an encoding process of a method for performing a transformation after flipping. FIG. 27 is a diagram illustrating an example of a decoding process of a method for performing flipping after inverse transformation. FIG. 28 is a diagram illustrating an example of an encoding process for performing flipping after conversion. FIG. 29 is a diagram illustrating an example of a decoding process of a method for performing inverse transformation after flipping. FIG. 30 is a diagram illustrating an example of an encoding process of a method for performing flipping after quantization. FIG. 31 is a diagram illustrating an example of a decoding process of a method for performing inverse quantization after flipping. FIG. 32 is a diagram illustrating the performance of flipping on the remaining blocks. FIG. 33 is a diagram illustrating an embodiment for implementing a flipping operation on an 8x8 size residual block in hardware. FIG. 34 is a diagram illustrating the flipping and transformation of the remaining blocks. FIGS. 35 to 37 are drawings for explaining embodiments of the first sub-block division mode according to the present invention. FIG. 38 is a drawing for explaining an embodiment of the second sub-block division mode according to the present invention. FIG. 39 is a diagram illustrating an example of a diagonal scan. FIG. 40 is a drawing for illustrating an example of a horizontal scan. FIG. 41 is a drawing for illustrating an example of a vertical scan. FIG. 42 is a diagram illustrating an example of a block-based diagonal scan. FIG. 43 is a drawing for illustrating an embodiment of a block-based horizontal scan. FIG. 44 is a drawing for illustrating an embodiment of a block-based vertical scan. FIG. 45 is a drawing for illustrating an embodiment of a block-based horizontal scan. FIG. 46 is a diagram illustrating an embodiment of a block-based vertical scan. FIG. 47 is a drawing for illustrating various embodiments of scanning based on the shape of a block. Figure 48 is a diagram illustrating an in-screen prediction mode. FIG. 49 is a diagram illustrating reference samples available for in-frame prediction. FIGS. 50 to 54 are examples of an encoding process or a decoding process using a conversion according to an embodiment of the present invention. FIG. 55 is a diagram illustrating an image decoding method according to an embodiment of the present invention. FIG. 56 is a diagram illustrating an image encoding method according to an embodiment of the present invention. Specific details for implementing the invention
[0032] The present invention is susceptible to various modifications and may have various embodiments; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the invention to specific embodiments, and it should be understood that the invention includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention. Similar reference numerals in the drawings refer to the same or similar functions across various aspects. The shapes and sizes of elements in the drawings may be exaggerated for clearer explanation. The detailed description of exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments as examples. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that various embodiments are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the invention in relation to one embodiment. Furthermore, it should be understood that the location or arrangement of individual components within each disclosed embodiment may be changed without departing from the spirit and scope of the embodiment. Accordingly, the following detailed description is not intended to be taken in a limiting sense, and the scope of exemplary embodiments is limited only by the appended claims, together with all equivalents to those claimed therein, provided they are properly described.
[0033] In the present invention, terms such as "first," "second," etc. may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.
[0034] When it is stated that a component of the present invention is “connected” or “connected” to another component, it should be understood that it may be directly connected to or connected to the other component, or that other components may exist in between. On the other hand, when it is stated that a component is “directly connected” or “directly connected” to another component, it should be understood that no other components exist in between.
[0035] The components shown in the embodiments of the present invention are illustrated independently to represent different characteristic functions and do not imply that each component consists of separate hardware or a single software unit. That is, each component is listed and included as a separate component for the convenience of explanation; however, at least two of the components may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separated embodiments of each component are included within the scope of the present invention as long as they do not deviate from the essence of the invention.
[0036] The terms used in this invention are used merely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this invention, terms such as "comprising" or "having" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. That is, the description in this invention that a specific configuration "comprising" does not exclude configurations other than that configuration, but means that additional configurations may be included within the scope of the practice or technical concept of this invention.
[0037] Some components of the present invention may not be essential components performing an essential function in the present invention, but may be optional components merely for enhancing performance. The present invention may be implemented by including only the components essential for realizing the essence of the present invention, excluding components used merely for performance enhancement, and a structure including only the essential components, excluding optional components used merely for performance enhancement, is also included within the scope of the rights of the present invention.
[0038] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In describing the embodiments of this specification, if it is determined that a detailed description of related known configurations or functions may obscure the gist of this specification, such detailed description is omitted; similar reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted.
[0039] In the following, "image" may refer to a single picture constituting a video, or it may refer to the video itself. For example, "encoding and / or decoding of an image" may mean "encoding and / or decoding of an image," and may also mean "encoding and / or decoding of one of the images constituting the video."
[0040] In the following, the terms "video" and "video" may be used interchangeably with the same meaning.
[0041] In the following, the target image may be an image to be encoded and / or an image to be decoded. Additionally, the target image may be an input image input to an encoding device and an input image input to a decoding device. Here, the target image may have the same meaning as the current image.
[0042] In the following, the terms "image," "picture," "frame," and "screen" may be used interchangeably with the same meaning.
[0043] In the following, the target block may be an encoding target block that is the target of encoding and / or a decoding target block that is the target of decoding. Additionally, the target block may be a current block that is the target of current encoding and / or decoding. For example, the terms "target block" and "current block" may be used interchangeably with the same meaning.
[0044] In the following, the terms "block" and "unit" may be used interchangeably with the same meaning. Alternatively, "block" may refer to a specific unit.
[0045] In the following, the terms "region" and "segment" may be used interchangeably.
[0046] In the following, a specific signal may be a signal representing a specific block. For example, the original signal may be a signal representing the target block. The prediction signal may be a signal representing the prediction block. The residual signal may be a signal representing the residual block.
[0047] In the embodiments, each of the specified information, data, flag, index and element, attribute, etc., may have a value. A value "0" of the information, data, flag, index and element, attribute, etc., may represent logical false or a first predefined value. That is to say, the value "0", false, logical false, and the first predefined value may be used interchangeably. A value "1" of the information, data, flag, index and element, attribute, etc., may represent logical true or a second predefined value. That is to say, the value "1", true, logical true, and the second predefined value may be used interchangeably.
[0048] When a variable such as i or j is used to represent a row, column, or index, the value of i may be an integer greater than or equal to 0, or an integer greater than or equal to 1. That is to say, in the embodiments, the row, column, and index, etc. may be counted from 0, or from 1.
[0050] Glossary of Terms
[0051] Encoder: Refers to a device that performs encoding. In other words, it can mean an encoding device.
[0052] Decoder: Refers to a device that performs decoding. In other words, it can mean a decoding device.
[0053] Block: An MxN array of samples. Here, M and N may represent positive integer values, and a block may commonly represent a two-dimensional array of samples. A block may represent a unit. The current block may represent a block to be encoded during encoding, or a block to be decoded during decoding. Additionally, the current block may be at least one of an encoding block, a prediction block, a residual block, or a transformation block.
[0054] Sample: The basic unit that makes up a block. Bit depth (B d From 0 to 2 depending on ) Bd - It can be expressed as a value up to 1. In the present invention, the term "sample" can be used interchangeably with "pixel" or "pixel." That is, "sample," "pixel," and "pixel" can have the same meaning.
[0055] Unit: This may refer to a unit of image encoding and decoding. In image encoding and decoding, a unit may be a region into which a single image is divided. Additionally, when an image is divided into subdivided units for encoding or decoding, a unit may refer to the divided unit. In other words, a single image can be divided into multiple units. In image encoding and decoding, predefined processing may be performed for each unit. A single unit may be further subdivided into sub-units that have a smaller size than the unit. Depending on the function, a unit may refer to a Block, Macroblock, Coding Tree Unit, Coding Tree Block, Coding Unit, Coding Block, Prediction Unit, Prediction Block, Residual Unit, Residual Block, Transform Unit, Transform Block, etc. Additionally, to distinguish it from a block, a unit may refer to a block of luminance (Luma) components, a corresponding block of chroma (Chroma) components, and syntactic elements for each block. A unit may have various sizes and shapes, and in particular, the shape of a unit may include not only squares but also geometric shapes that can be represented in two dimensions, such as rectangles, trapezoids, triangles, and pentagons. Additionally, unit information may include at least one of the following: the type of unit indicating an encoding unit, a prediction unit, a residual unit, a transformation unit, etc., the size of the unit, the depth of the unit, and the encoding and decoding order of the unit.
[0056] Coding Tree Unit: Consists of a single luminance component (Y) coding tree block and two chrominance component (Cb, Cr) coding tree blocks associated with it. It may also refer to the blocks and the syntactic elements for each block. Each coding tree unit may be partitioned using one or more partitioning methods, such as a quad tree, binary tree, or ternary tree, to form sub-units such as a coding unit, a prediction unit, and a transform unit. It may be used as a term to refer to a sample block that serves as a processing unit in the image decoding process, such as the partitioning of an input image. Here, a quad tree may refer to a quaternary tree.
[0057] If the size of the encoding block falls within a predetermined range, it may be possible to split it into a quadtree only. Here, the predetermined range may be defined as at least one of the maximum size and minimum size of the encoding block that can be split into a quadtree only. Information indicating the maximum / minimum size of the encoding block for which quadtree-type splitting is allowed may be signaled via a bitstream, and such information may be signaled in at least one unit among a sequence, picture parameter, tile group, or slice (segment). Alternatively, the maximum / minimum size of the encoding block may be a fixed size pre-set in the encoder / decoder. For example, if the size of the encoding block corresponds to 256x256 to 64x64, it may be possible to split it into a quadtree only. Or, if the size of the encoding block is larger than the maximum size of the conversion block, it may be possible to split it into a quadtree only. In this case, the block being split may be at least one of the encoding block or the conversion block. In such cases, information indicating the splitting of the encoding block (e.g., split_flag) may be a flag indicating whether to split it into a quadtree. If the size of the encoding block falls within a predetermined range, it may be divided only into a binary tree or a triad tree. In this case, the above description regarding the quad tree may be applied equally to the binary tree or the triad tree.
[0058] Coding Tree Block: This term may be used to refer to any one of the Y coding tree block, Cb coding tree block, or Cr coding tree block.
[0059] Neighbor block: This may refer to a block adjacent to the current block. A block adjacent to the current block may refer to a block whose boundary meets the current block or a block located within a certain distance from the current block. A neighbor block may refer to a block adjacent to a vertex of the current block. Here, a block adjacent to a vertex of the current block may be a block vertically adjacent to a neighbor block horizontally adjacent to the current block, or a block horizontally adjacent to a neighbor block vertically adjacent to the current block. A neighbor block may also refer to a restored neighbor block.
[0060] Reconstructed Neighbor Block: This may refer to a neighbor block that has already been encoded or decoded spatially or temporally around the current block. In this case, a reconstructed neighbor block may refer to a reconstructed neighbor unit. A reconstructed spatial neighbor block may be a block within the current picture that has already been reconstructed through encoding and / or decoding. A reconstructed temporal neighbor block may be a reconstructed block or its neighbor block located at a position corresponding to the current block of the current picture within the reference image.
[0061] Unit Depth: This refers to the degree to which a unit is divided. In a tree structure, the topmost node (Root Node) corresponds to the initial, undivided unit. This topmost node can be referred to as the root node. Additionally, the topmost node can have a minimum depth value. In this case, the topmost node can have a depth of Level 0. A node with a depth of Level 1 can represent a unit created as the initial unit is divided once. A node with a depth of Level 2 can represent a unit created as the initial unit is divided twice. A node with a depth of Level n can represent a unit created as the initial unit is divided n times. A Leaf Node can be the lowest node and can be a node that cannot be further divided. The depth of a Leaf Node can be the maximum level. For example, the predefined value for the maximum level can be 3. It can be said that the Root Node has the shallowest depth, and the Leaf Node has the deepest depth. Additionally, when units are represented as a tree structure, the level at which a unit exists can represent the unit depth.
[0062] Bitstream: Can refer to a sequence of bits containing encoded image information.
[0063] Parameter Set: This corresponds to header information within the structure of the bitstream. At least one of the video parameter set, sequence parameter set, picture parameter set, and adaptation parameter set may be included in the parameter set. Additionally, the parameter set may include tile group, slice header, and tile header information. Furthermore, the tile group may refer to a group containing multiple tiles and may have the same meaning as a slice.
[0064] An adaptive parameter set may refer to a set of parameters that can be referenced and shared across different pictures, subpictures, slices, tile groups, tiles, or bricks. Additionally, subpictures, slices, tile groups, tiles, or bricks within a picture may reference different adaptive parameter sets to utilize information within those sets.
[0065] Additionally, within a picture, different adaptation parameter sets can be referenced using identifiers of different adaptation parameter sets in subpictures, slices, tile groups, tiles, or bricks.
[0066] Additionally, within a slice, tile group, tile, or brick in a subpicture, different adaptation parameter sets can be referenced using the identifiers of different adaptation parameter sets.
[0067] Additionally, an adaptive parameter set can refer to different adaptive parameter sets within a tile or brick using the identifier of a different adaptive parameter set.
[0068] Additionally, within a brick in a tile, different adaptation parameter sets can be referenced using the identifiers of different adaptation parameter sets.
[0069] Information regarding an adaptive parameter set identifier is included in the parameter set or header of the above subpicture, so that an adaptive parameter set corresponding to the adaptive parameter set identifier can be used in the subpicture.
[0070] Information regarding an adaptive parameter set identifier is included in the parameter set or header of the above tile, so that an adaptive parameter set corresponding to the adaptive parameter set identifier can be used in the tile.
[0071] The header of the above brick includes information regarding an adaptive parameter set identifier, so that an adaptive parameter set corresponding to the adaptive parameter set identifier can be used in the brick.
[0072] The above picture can be divided into one or more rows of tiles and one or more columns of tiles.
[0073] The above subpicture may be divided into one or more tile rows and one or more tile columns within the picture. The above subpicture is an area having a rectangular / square shape within the picture and may include one or more CTUs. Additionally, at least one tile / brick / slice may be included within a single subpicture.
[0074] The above tile is an area within the picture that has a rectangular or square shape and may include one or more CTUs. Additionally, the tile may be divided into one or more bricks.
[0075] The above brick may refer to one or more CTU rows within a tile. A tile may be divided into one or more bricks, and each brick may have at least one CTU row. A tile that is not divided into two or more may also refer to a brick.
[0076] The above slice may include one or more tiles within the picture and one or more bricks within the tile.
[0077] Parsing: This refers to determining the value of a syntax element by entropy decoding a bitstream, or it may refer to entropy decoding itself.
[0078] Symbol: May represent at least one of the following: a syntactic element of the unit to be encoded / decoded, a coding parameter, or a value of a transform coefficient. Additionally, the symbol may represent the target of entropy encoding or the result of entropy decoding.
[0079] Prediction Mode: This may be information indicating a mode of encoding / decoding by intra-frame prediction or a mode of encoding / decoding by inter-frame prediction.
[0080] Prediction Unit: This refers to the basic unit used when performing predictions, such as inter-frame prediction, intra-frame prediction, inter-frame reward, intra-frame reward, and motion reward. A single prediction unit may be divided into multiple partitions or multiple sub-prediction units of smaller sizes. Multiple partitions may also serve as basic units for performing prediction or reward. Partitions created by the division of a prediction unit may also be prediction units.
[0081] Prediction Unit Partition: This can refer to a form in which prediction units are divided.
[0082] Reference Picture List: This may refer to a list containing one or more reference pictures used for inter-frame prediction or motion compensation. The types of reference picture lists may include LC (List Combined), L0 (List 0), L1 (List 1), L2 (List 2), L3 (List 3), etc., and one or more reference picture lists may be used for inter-frame prediction.
[0083] Inter Prediction Indicator: May indicate the inter-frame prediction direction (unidirectional prediction, bidirectional prediction, etc.) of the current block. Alternatively, it may indicate the number of reference images used when generating the prediction blocks for the current block. Alternatively, it may indicate the number of prediction blocks used when performing inter-frame prediction or motion compensation for the current block.
[0084] Prediction list utilization flag: Indicates whether a prediction block is generated using at least one reference image within a specific reference image list. A prediction list utilization flag can be used to derive a prediction indicator between frames, and conversely, a prediction list utilization flag can be used to derive a prediction list utilization flag. For example, if the prediction list utilization flag indicates a first value of 0, it may indicate that a prediction block is not generated using a reference image within the reference image list, and if it indicates a second value of 1, it may indicate that a prediction block can be generated using the reference image list.
[0085] Reference Picture Index: This can refer to an index in a reference picture list that points to a specific reference picture.
[0086] Reference Picture: This may refer to an image referenced by a specific block for inter-frame prediction or motion compensation. Alternatively, the reference picture may be an image containing a reference block referenced by the current block for inter-frame prediction or motion compensation. Hereinafter, the terms "reference picture" and "reference image" may be used interchangeably with the same meaning.
[0087] Motion Vector: This can be a 2D vector used for inter-frame prediction or motion compensation. A motion vector can represent the offset between the block to be encoded / decoded and the reference block. For example, (mvX, mvY) can represent a motion vector. mvX can represent the horizontal component, and mvY can represent the vertical component.
[0088] Search Range: The search range may be a 2-dimensional area where a search for motion vectors takes place during cross-frame prediction. For example, the size of the search range may be MxN. M and N may each be positive integers.
[0089] Motion Vector Candidate: This may refer to a block that serves as a prediction candidate when predicting a motion vector, or the motion vector of that block. Additionally, a motion vector candidate may be included in the motion vector candidate list.
[0090] Motion Vector Candidate List: This can refer to a list composed of one or more motion vector candidates.
[0091] Motion Vector Candidate Index: May refer to an indicator pointing to a motion vector candidate within the motion vector candidate list. May be the index of a Motion Vector Predictor.
[0092] Motion Information: This may refer to information including at least one of motion vectors, reference image indices, and cross-frame prediction indicators, as well as prediction list utilization flags, reference image list information, reference images, motion vector candidates, motion vector candidate indices, merge candidates, merge indices, etc.
[0093] Merge Candidate List: Can refer to a list composed of one or more merge candidates.
[0094] Merge Candidate: This may refer to spatial merge candidates, temporal merge candidates, combined merge candidates, combined positive prediction merge candidates, zero merge candidates, etc. A merge candidate may include motion information such as cross-frame prediction indicators, reference image indices for each list, motion vectors, prediction list utilization flags, and cross-frame prediction indicators.
[0095] Merge Index: This may refer to an indicator pointing to a merge candidate within a merge candidate list. Additionally, the merge index may indicate the block that induced the merge candidate among the blocks restored spatially or temporally adjacent to the current block. Furthermore, the merge index may indicate at least one of the movement information possessed by the merge candidate.
[0096] Transform Unit: This may refer to a basic unit for performing residual signal encoding / decoding, such as transform, inverse transform, quantization, inverse quantization, and transform coefficient encoding / decoding. A single transform unit may be divided into multiple sub-transform units of smaller sizes. Here, the transform / inverse transform may include at least one of a first-order transform / inverse transform and a second-order transform / inverse transform.
[0097] Scaling: This can refer to the process of multiplying a factor by a quantized level. Transformation coefficients can be generated as a result of scaling the quantized level. Scaling can also be called dequantization.
[0098] Quantization Parameter: This may refer to a value used to generate a quantized level using a transform factor in quantization. Alternatively, it may refer to a value used to generate a transform factor by scaling the quantized level in inverse quantization. The quantization parameter may be a value mapped to the quantization step size.
[0099] Delta Quantization Parameter: This may refer to the difference between the predicted quantization parameter and the quantization parameter of the unit to be encoded / decoded.
[0100] Scan: This can refer to a method of sorting the order of units, blocks, or coefficients within a matrix. For example, sorting a 2D array into a 1D array is called a scan. Alternatively, sorting a 1D array into a 2D array can also be called a scan or inverse scan.
[0101] Transform Coefficient: This may refer to the coefficient value generated after performing a transformation in the encoder. Alternatively, it may refer to the coefficient value generated after performing at least one of entropy decoding and inverse quantization in the decoder. Quantized levels or quantized transform coefficient levels obtained by applying quantization to the transform coefficient or residual signal may also be included in the meaning of transform coefficient.
[0102] Quantized Level: This may refer to a value generated by performing quantization on transform coefficients or residual signals in an encoder. Alternatively, it may refer to the value subject to inverse quantization before it is performed in a decoder. Similarly, the quantized transform coefficient level resulting from transform and quantization may also be included within the meaning of quantized level.
[0103] Non-zero Transform Coefficient: This may refer to a transform coefficient whose magnitude is not zero, a transform coefficient level whose magnitude is not zero, or a quantized level.
[0104] Quantization Matrix: This refers to a matrix used in the quantization or inverse quantization process to improve the subjective or objective quality of an image. A quantization matrix can also be called a scaling list.
[0105] Quantization Matrix Coefficient: This can refer to each element within the quantization matrix. Quantization matrix coefficients can also be called matrix coefficients.
[0106] Default Matrix: This may refer to a predetermined quantization matrix predefined in the encoder and decoder.
[0107] Non-default Matrix: This can refer to a quantization matrix that is not predefined in the encoder and decoder and is signaled by the user.
[0108] Statistic value: A statistical value for at least one variable, encoding parameter, constant, etc., having specific values that can be computed, may be at least one of the average value, weighted average value, weighted sum value, minimum value, maximum value, mode, median value, and interpolation value of said specific values.
[0109] FIG. 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied.
[0110] The encoding device (100) may be an encoder, a video encoding device, or an image encoding device. The video may include one or more images. The encoding device (100) may sequentially encode one or more images.
[0111] Referring to FIG. 1, the encoding device (100) may include a motion prediction unit (111), a motion compensation unit (112), an intra prediction unit (120), a switch (115), a subtractor (125), a converter (130), a quantization unit (140), an entropy encoding unit (150), an inverse quantization unit (160), an inverse converter (170), an adder (175), a filter unit (180), and a reference picture buffer (190).
[0112] The encoding device (100) can perform encoding on an input image in intra mode and / or inter mode. Additionally, the encoding device (100) can generate a bitstream containing encoded information through encoding of the input image and can output the generated bitstream. The generated bitstream can be stored on a computer-readable recording medium or streamed via a wired / wireless transmission medium. When intra mode is used as the prediction mode, the switch (115) can be switched to intra, and when inter mode is used as the prediction mode, the switch (115) can be switched to inter. Here, intra mode may refer to an intra-frame prediction mode, and inter mode may refer to an inter-frame prediction mode. The encoding device (100) can generate a prediction block for an input block of the input image. Additionally, after the prediction block is generated, the encoding device (100) can encode a residual block using the residual of the input block and the prediction block. The input image may be referred to as the current image that is the target of the current encoding. The input block may be referred to as the current block or the block to be encoded, which is the target of the current encoding.
[0113] When the prediction mode is an intra mode, the intra prediction unit (120) may use a sample of a block that has already been encoded / decoded around the current block as a reference sample. The intra prediction unit (120) may perform spatial prediction for the current block using the reference sample and generate prediction samples for the input block through spatial prediction. Here, intra prediction may mean intra-frame prediction.
[0114] When the prediction mode is an inter mode, the motion prediction unit (111) can search for the region that best matches the input block from the reference image during the motion prediction process and derive a motion vector using the searched region. At this time, the search region can be used as the region. The reference image can be stored in the reference picture buffer (190). Here, the reference image can be stored in the reference picture buffer (190) when encoding / decoding of the reference image is processed.
[0115] The motion compensation unit (112) can generate a prediction block for the current block by performing motion compensation using a motion vector. Here, inter-prediction may mean inter-frame prediction or motion compensation.
[0116] The motion prediction unit (111) and motion compensation unit (112) can generate a prediction block by applying an interpolation filter to a portion of the reference image when the value of the motion vector does not have an integer value. To perform inter-frame prediction or motion compensation, based on the encoding unit, it can determine whether the motion prediction and motion compensation method of the prediction unit included in the corresponding encoding unit is a Skip Mode, Merge Mode, Advanced Motion Vector Prediction (AMVP) Mode, or Current Picture Reference Mode, and can perform inter-frame prediction or motion compensation according to each mode.
[0117] The subtractor (125) can generate a residual block using the difference between the input block and the prediction block. The residual block may also be referred to as a residual signal. The residual signal may represent the difference between the original signal and the prediction signal. Alternatively, the residual signal may be a signal generated by transforming, quantizing, or transforming and quantizing the difference between the original signal and the prediction signal. The residual block may be a residual signal in block units.
[0118] The transformation unit (130) can generate a transform coefficient by performing a transform on the remaining block and output the generated transform coefficient. Here, the transform coefficient may be a coefficient value generated by performing a transform on the remaining block. When a transform skip mode is applied, the transformation unit (130) may skip the transform on the remaining block.
[0119] A quantized level can be generated by applying quantization to a conversion coefficient or a residual signal. In the following embodiments, the quantized level may also be referred to as a conversion coefficient.
[0120] The quantization unit (140) can generate a quantized level by quantizing a transformation coefficient or residual signal according to a quantization parameter, and can output the generated quantized level. At this time, the quantization unit (140) can quantize the transformation coefficient using a quantization matrix.
[0121] The entropy encoding unit (150) can generate a bitstream and output a bitstream by performing entropy encoding according to a probability distribution on values calculated by the quantization unit (140) or coding parameter values calculated during the encoding process. The entropy encoding unit (150) can perform entropy encoding on information regarding a sample of an image and information for decoding an image. For example, information for decoding an image may include syntax elements, etc.
[0122] When entropy coding is applied, a small number of bits are allocated to symbols with a high probability of occurrence and a large number of bits are allocated to symbols with a low probability of occurrence, thereby representing the symbols and reducing the size of the bit sequence for the symbols to be encoded. The entropy coding unit (150) may use encoding methods such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding) for entropy coding. For example, the entropy coding unit (150) may perform entropy coding using a Variable Length Coding (VLC) table. In addition, the entropy encoding unit (150) may perform arithmetic encoding using the derived binarization method, probability model, and context model after deriving a binarization method of the target symbol and a probability model of the target symbol / bin.
[0123] The entropy encoding unit (150) can convert a 2-dimensional block form coefficient into a 1-dimensional vector form through a transform coefficient scanning method to encode a transform coefficient level (quantized level).
[0124] Coding parameters may include not only information (flags, indices, etc.) that is encoded in the encoder and signaled to the decoder, such as syntax elements, but also information derived during the encoding or decoding process, and may refer to information required when encoding or decoding images. For example, unit / block size, unit / block depth, unit / block partitioning information, unit / block shape, unit / block partitioning structure, whether to partition in quadtree form, whether to partition in binary tree form, binary tree partitioning direction (horizontal or vertical), binary tree partitioning type (symmetrical or asymmetrical), whether to partition in triad tree form, triad tree partitioning direction (horizontal or vertical), triad tree partitioning type (symmetrical or asymmetrical), whether to partition in complex tree form, complex tree partitioning direction (horizontal or vertical), complex tree partitioning type (symmetrical or asymmetrical), complex tree partitioning tree (binary tree or triad tree), prediction mode (intra-frame prediction or inter-frame prediction), intra-frame luminance prediction mode / direction, intra-frame chrominance prediction mode / direction, intra-frame partitioning information, inter-frame partitioning information, encoded block partitioning flag, predicted block partitioning flag, transform block partitioning flag, reference sample filtering method, reference sample filter tab, reference sample filter coefficients, predicted block filtering method, predicted block filter tab, predicted block Filter coefficients, Prediction block boundary filtering method, Prediction block boundary filter tab, Prediction block boundary filter coefficients, Intra-frame prediction mode, Inter-frame prediction mode, Motion information, Motion vector, Motion vector difference, Reference image index, Inter-frame prediction direction, Inter-frame prediction indicator, Prediction list utilization flag, Reference image list, Reference image, Motion vector prediction index, Motion vector prediction candidate, Motion vector candidate list, Whether to use merge mode, Merge index, Merge candidate, Merge candidate list, Whether to use skip mode, Interpolation filter type,Interpolation Filter Tab, Interpolation Filter Coefficients, Motion Vector Magnitude, Motion Vector Representation Accuracy, Transform Type, Transform Magnitude, 1st Transform Usage Information, 2nd Transform Usage Information, 1st Transform Index, 2nd Transform Index, Residual Signal Presence Information, Coded Block Pattern, Coded Block Flag, Quantization Parameters, Residual Quantization Parameters, Quantization Matrix, Intra-Loop Filter Application Status, Intra-Loop Filter Coefficients, Intra-Loop Filter Tab, Intra-Loop Filter Shape / Form, Deblocking Filter Application Status, Deblocking Filter Coefficients, Deblocking Filter Tab, Deblocking Filter Strength, Deblocking Filter Shape / Form, Adaptive Sample Offset Application Status, Adaptive Sample Offset Value, Adaptive Sample Offset Category, Adaptive Sample Offset Type, Adaptive Loop Filter Application Status, Adaptive Loop Filter Coefficients, Adaptive Loop Filter Tab, Adaptive Loop Filter Shape / Form, Binarization / Debinarization Method, Context Model Determination Method, Context Model Update Method, Regular Mode Execution Status, Bypass Mode Execution Status, Context Bin, Bypass Bin, Important Factor Flag, Last Important Factor Flag, Factor Group Unit Encoding Flag, Last Important Factor Position, Flag for whether the factor value is greater than 1, Flag for whether the factor value is greater than 2, Flag for whether the factor value is greater than 3, Remaining Factor Value Information, Sign Information, Reconstructed Luminance Sample, Reconstructed Chromaticity Sample, Residual Luminance Sample, Residual Chromaticity Sample, Luminance Conversion Factor, Chromaticity Conversion Factor, Luminance Quantized Level, Chromaticity Quantized Level, Conversion Factor Level Scanning Method, Size of Decoder Side Motion Vector Search Area, Shape of Decoder Side Motion Vector Search Area, Number of Decoder Side Motion Vector Searches, CTU Size Information, Minimum Block Size Information, Maximum Block Size Information, Maximum Block Depth Information, Minimum Block Depth Information, Image Display / Output Order, Slice Identification Information, Slice Type, Slice Splitting Information,At least one value or a combination of tile group identification information, tile group type, tile group splitting information, tile identification information, tile type, tile splitting information, picture type, input sample bit depth, restored sample bit depth, residual sample bit depth, transform factor bit depth, quantized level bit depth, information about the luminance signal, and information about the chrominance signal may be included in the encoding parameter.
[0125] Here, signaling a flag or index may mean that in an encoder, the corresponding flag or index is entropy encoded and included in a bitstream, and in a decoder, the corresponding flag or index is entropy decoded from the bitstream.
[0126] When the encoding device (100) performs encoding through inter-prediction, the encoded current image can be used as a reference image for another image to be processed later. Accordingly, the encoding device (100) can restore or decode the encoded current image again, and can store the restored or decoded image as a reference image in the reference picture buffer (190).
[0127] The quantized level can be dequantized in the dequantization unit (160) and inverse transformed in the inverse transform unit (170). The dequantized and / or inverse transformed coefficients can be added to the prediction block through the adder (175). A reconstructed block can be generated by adding the dequantized and / or inverse transformed coefficients and the prediction block. Here, the dequantized and / or inverse transformed coefficients refer to coefficients for which at least one of dequantization and inverse transformation has been performed, and may refer to the reconstructed residual block.
[0128] The restoration block may pass through a filter section (180). The filter section (180) may apply at least one of a deblocking filter, a Sample Adaptive Offset (SAO), an Adaptive Loop Filter (ALF), etc., to the restoration sample, restoration block, or restoration image. The filter section (180) may also be referred to as an in-loop filter.
[0129] Deblocking filters can remove block distortion occurring at the boundaries between blocks. To determine whether to perform deblocking, the decision to apply the filter to the current block can be made based on samples contained in a few columns or rows within the block. When applying a deblocking filter to a block, different filters can be applied depending on the required deblocking filtering intensity.
[0130] To compensate for encoding errors using a sample adaptive offset, an appropriate offset value can be added to the sample value. The sample adaptive offset can correct the offset from the original image on a sample-by-sample basis for the deblocked image. One method may be to divide the samples included in the image into a certain number of regions, determine the region to be offset, and apply the offset to that region, or to apply the offset by considering the edge information of each sample.
[0131] An adaptive loop filter can perform filtering based on a comparison of the reconstructed image and the original image. After dividing the samples included in the image into predetermined groups, a filter to be applied to each group can be determined, thereby performing filtering differently for each group. Information regarding whether to apply an adaptive loop filter can be signaled per coding unit (CU), and the shape and filter coefficients of the adaptive loop filter to be applied may vary depending on each block.
[0132] The restored block or restored image that has passed through the filter unit (180) can be stored in the reference picture buffer (190). The restored block that has passed through the filter unit (180) may be part of the reference image. That is to say, the reference image may be a restored image composed of the restored blocks that have passed through the filter unit (180). The stored reference image may subsequently be used for inter-frame prediction or motion compensation.
[0133] FIG. 2 is a block diagram showing the configuration according to one embodiment of a decoding device to which the present invention is applied.
[0134] The decoding device (200) may be a decoder, a video decoding device, or an image decoding device.
[0135] Referring to FIG. 2, the decoding device (200) may include an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (230), an intra prediction unit (240), a motion compensation unit (250), an adder (255), a filter unit (260), and a reference picture buffer (270).
[0136] The decoding device (200) can receive a bitstream output from the encoding device (100). The decoding device (200) can receive a bitstream stored in a computer-readable recording medium or a bitstream stream streamed through a wired / wireless transmission medium. The decoding device (200) can perform decoding on the bitstream in intra mode or inter mode. Additionally, the decoding device (200) can generate a restored image or a decoded image through decoding and can output the restored image or the decoded image.
[0137] If the prediction mode used for decoding is intra mode, the switch can be switched to intra. If the prediction mode used for decoding is inter mode, the switch can be switched to inter.
[0138] The decoding device (200) can decode the input bitstream to obtain a reconstructed residual block and generate a prediction block. Once the reconstructed residual block and the prediction block are obtained, the decoding device (200) can generate a reconstructed block to be decoded by adding the reconstructed residual block and the prediction block. The block to be decoded may be referred to as the current block.
[0139] The entropy decoding unit (210) can generate symbols by performing entropy decoding according to the probability distribution of the bitstream. The generated symbols may include symbols in the form of quantized levels. Here, the entropy decoding method may be the inverse process of the entropy encoding method described above.
[0140] The entropy decoding unit (210) can convert a one-dimensional vector-shaped coefficient into a two-dimensional block-shaped coefficient through a conversion coefficient scanning method to decode a conversion coefficient level (quantized level).
[0141] The quantized level can be inversely quantized in the inverse quantization unit (220) and inversely transformed in the inverse transformation unit (230). The quantized level can be generated as a restored residual block as a result of inverse quantization and / or inverse transformation being performed. At this time, the inverse quantization unit (220) can apply a quantization matrix to the quantized level.
[0142] When intra mode is used, the intra prediction unit (240) can generate a prediction block by performing a spatial prediction on the current block using sample values of already decoded blocks around the block to be decoded.
[0143] When an inter mode is used, the motion compensation unit (250) can generate a prediction block by performing motion compensation on the current block using a motion vector and a reference image stored in the reference picture buffer (270). The motion compensation unit (250) can generate a prediction block by applying an interpolation filter to a portion of the reference image when the value of the motion vector does not have an integer value. To perform motion compensation, it can determine whether the motion compensation method of the prediction unit included in the corresponding encoding unit is a skip mode, merge mode, AMVP mode, or current picture reference mode based on the encoding unit, and can perform motion compensation according to each mode.
[0144] The adder (255) can generate a restored block by adding the restored residual block and the prediction block. The filter unit (260) can apply at least one of a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the restored block or the restored image. The filter unit (260) can output the restored image. The restored block or the restored image can be stored in a reference picture buffer (270) and used for inter-prediction. The restored block that has passed through the filter unit (260) may be part of the reference image. That is to say, the reference image may be a restored image composed of the restored blocks that have passed through the filter unit (260). The stored reference image may subsequently be used for inter-frame prediction or motion compensation.
[0145] FIG. 3 is a diagram schematically illustrating the segmentation structure of an image when encoding and decoding an image. FIG. 3 schematically illustrates an embodiment in which a single unit is divided into a plurality of sub-units.
[0146] To efficiently segment the image, a coding unit (CU) may be used in encoding and decoding. The coding unit may be used as the basic unit of image encoding / decoding. Additionally, the coding unit may be used as a unit to distinguish between intra-frame prediction mode and inter-frame prediction mode during image encoding / decoding. The coding unit may be the basic unit used for the processes of prediction, transform, quantization, inverse transform, inverse quantization, or encoding / decoding of transform coefficients.
[0147] Referring to FIG. 3, the image (300) is sequentially divided into Largest Coding Units (LCUs), and the division structure is determined in LCU units. Here, LCU can be used with the same meaning as Coding Tree Unit (CTU). The division of a unit may refer to the division of a block corresponding to the unit. The block division information may include information regarding the depth of the unit. The depth information may indicate the number and / or degree of division of the unit. A unit may be hierarchically divided into multiple sub-units based on a tree structure and having depth information. That is to say, the unit and the sub-units generated by the division of the unit may correspond to a node and a child node of the node, respectively. Each divided sub-unit may have depth information. The depth information may be information indicating the size of the CU and may be stored for each CU. Since the unit depth indicates the number and / or degree of division of the unit, the division information of the sub-unit may include information regarding the size of the sub-unit.
[0148] The partitioning structure may refer to the distribution of coding units (CUs) within the CTU (310). This distribution may be determined by whether to partition a single CU into multiple CUs (positive integers of 2 or more, including 2, 4, 8, 16, etc.). The width and height of the CUs generated by partitioning may be half the width and height of the CUs before partitioning, respectively, or may have a size smaller than the width and height of the CUs before partitioning, depending on the number of partitions. The CUs may be recursively partitioned into multiple CUs. Through recursive partitioning, at least one of the width and height of the partitioned CUs may be reduced compared to at least one of the width and height of the CUs before partitioning. The partitioning of the CUs may be performed recursively up to a predefined depth or a predefined size. For example, the depth of the CTU may be 0, and the depth of the Smallest Coding Unit (SCU) may be a predefined maximum depth. Here, the CTU may be a coding unit having the maximum coding unit size as described above, and the SCU may be a coding unit having the minimum coding unit size. Division begins from the CTU (310), and the depth of the CU increases by 1 each time the horizontal and / or vertical size of the CU is reduced by division. For example, for each depth, the CU that is not divided may have a size of 2Nx2N. Also, for the CU that is divided, the CU of size 2Nx2N may be divided into 4 CUs of size NxN. The size of N may be reduced by half each time the depth increases by 1.
[0149] Additionally, information regarding whether a CU is divided can be expressed through the division information of the CU. The division information may be 1 bit of information. All CUs except the SCU may include division information. For example, if the value of the division information is a first value, the CU may not be divided, and if the value of the division information is a second value, the CU may be divided.
[0150] Referring to FIG. 3, a CTU with a depth of 0 can be a 64x64 block. 0 can be the minimum depth. A SCU with a depth of 3 can be an 8x8 block. 3 can be the maximum depth. CUs of 32x32 blocks and 16x16 blocks can be represented as depth 1 and depth 2, respectively.
[0151] For example, if a single encoding unit is divided into four encoding units, the width and height of the four divided encoding units may each have half the size of the encoding unit before division. For example, if a 32x32 encoding unit is divided into four encoding units, the four divided encoding units may each have a size of 16x16. When a single encoding unit is divided into four encoding units, the encoding unit can be said to have been divided into a quad-tree form (quad-tree partition).
[0152] For example, if a single encoding unit is divided into two encoding units, the width or height of the two divided encoding units may be half the size of the encoding unit before division. For example, if a 32x32 encoding unit is divided vertically into two encoding units, the two divided encoding units may each have a size of 16x32. For example, if an 8x32 encoding unit is divided horizontally into two encoding units, the two divided encoding units may each have a size of 8x16. When a single encoding unit is divided into two encoding units, the encoding unit can be said to have been partitioned into a binary tree form (binary-tree partition).
[0153] For example, when a single encoding unit is divided into three encoding units, the encoding unit can be divided into three encoding units by dividing the horizontal or vertical dimensions of the encoding unit before division in a ratio of 1:2:1. For example, if a 16x32 encoding unit is divided horizontally into three encoding units, the three divided encoding units may have dimensions of 16x8, 16x16, and 16x8, respectively, starting from the top. For example, if a 32x32 encoding unit is divided vertically into three encoding units, the three divided encoding units may have dimensions of 8x32, 16x32, and 8x32, respectively, starting from the left. When a single encoding unit is divided into three encoding units, the encoding unit can be said to have been partitioned in the form of a ternary-tree (ternary-tree partition).
[0154] The CTU (320) of Fig. 3 is an example of a CTU to which quadtree splitting, binary tree splitting and 3-part tree splitting are all applied.
[0155] As described above, to partition a CTU, at least one of quadtree partitioning, binary tree partitioning, and triad tree partitioning may be applied. Each partitioning may be applied based on a predetermined priority. For example, quadtree partitioning may be applied preferentially to the CTU. A coding unit that can no longer be quadtree partitioned may correspond to a leaf node of a quadtree. A coding unit corresponding to a leaf node of a quadtree may become a root node of a binary tree and / or a triad tree. That is, a coding unit corresponding to a leaf node of a quadtree may be binary tree partitioned, triad tree partitioned, or not partitioned further. At this time, by ensuring that quadtree partitioning is not performed again on the coding unit created by binary tree partitioning or triad tree partitioning of the coding unit corresponding to a leaf node of a quadtree, the partitioning of the block and / or signaling of partitioning information can be effectively performed.
[0156] The division of a coding unit corresponding to each node of a quadtree can be signaled using quad division information. Quad division information having a first value (e.g., '1') can indicate that the corresponding coding unit is quadtree divided. Quad division information having a second value (e.g., '0') can indicate that the corresponding coding unit is not quadtree divided. Quad division information may be a flag having a predetermined length (e.g., 1 bit).
[0157] There may be no priority between binary tree splitting and triad tree splitting. That is, encoding units corresponding to the leaf nodes of a quadtree can be binary tree split or triad tree split. Additionally, encoding units generated by binary tree splitting or triad tree splitting may be binary tree split or triad tree split again, or may not be split any further.
[0158] A partition in which there is no priority between binary tree partitioning and triad tree partitioning can be referred to as a multi-type tree partition. That is, the encoding unit corresponding to the leaf node of a quadtree can become the root node of a multi-type tree. The partition of the encoding unit corresponding to each node of the multi-type tree can be signaled using at least one of the partition status information, partition direction information, and partition tree information of the multi-type tree. For the partition of the encoding unit corresponding to each node of the multi-type tree, the partition status information, partition direction information, and partition tree information may be signaled sequentially.
[0159] Information on whether a composite tree is split with a first value (e.g., '1') may indicate that the corresponding encoding unit is split into a composite tree. Information on whether a composite tree is split with a second value (e.g., '0') may indicate that the corresponding encoding unit is not split into a composite tree.
[0160] When a encoding unit corresponding to each node of a composite tree is split, the corresponding encoding unit may further include splitting direction information. The splitting direction information may indicate the splitting direction of the composite tree split. Splitting direction information having a first value (e.g., '1') may indicate that the corresponding encoding unit is split in the vertical direction. Splitting direction information having a second value (e.g., '0') may indicate that the corresponding encoding unit is split in the horizontal direction.
[0161] When a encoding unit corresponding to each node of a composite tree is partitioned, the encoding unit may further include partition tree information. The partition tree information may indicate the tree used for the composite tree partition. Partition tree information having a first value (e.g., '1') may indicate that the encoding unit is partitioned into a binary tree. Partition tree information having a second value (e.g., '0') may indicate that the encoding unit is partitioned into a triad tree.
[0162] The splitting information, the splitting tree information, and the splitting direction information may each be flags having a predetermined length (e.g., 1 bit).
[0163] At least one of quad splitting information, information on whether a composite tree is split, splitting direction information, and splitting tree information can be entropy encoded / decoded. For the entropy encoding / decoding of the above information, information from a neighboring encoding unit adjacent to the current encoding unit may be used. For example, the splitting form (segmentation status, splitting tree, and / or splitting direction) of the left encoding unit and / or the upper encoding unit is highly likely to be similar to the splitting form of the current encoding unit. Therefore, context information for the entropy encoding / decoding of the information of the current encoding unit can be derived based on the information of the neighboring encoding unit. At this time, the information of the neighboring encoding unit may include at least one of the quad splitting information, information on whether a composite tree is split, splitting direction information, and splitting tree information of the corresponding encoding unit.
[0164] In another embodiment, among binary tree partitioning and three-part tree partitioning, binary tree partitioning may be performed first. That is, binary tree partitioning is applied first, and a encoding unit corresponding to a leaf node of the binary tree may be set as the root node of the three-part tree. In this case, quadtree partitioning and binary tree partitioning may not be performed for the encoding unit corresponding to a node of the three-part tree.
[0165] A encoding unit that is no longer divided by quadtree splitting, binary tree splitting, and / or ternary tree splitting can be a unit of encoding, prediction, and / or conversion. That is, the encoding unit may no longer be divided for prediction and / or conversion. Therefore, a splitting structure, splitting information, etc., for splitting the encoding unit into a prediction unit and / or conversion unit may not exist in the bitstream.
[0166] However, if the size of the encoding unit serving as the unit of division is larger than the size of the maximum conversion block, the encoding unit may be recursively divided until it becomes equal to or smaller than the size of the maximum conversion block. For example, if the size of the encoding unit is 64x64 and the size of the maximum conversion block is 32x32, the encoding unit may be divided into four 32x32 blocks for conversion. For example, if the size of the encoding unit is 32x64 and the size of the maximum conversion block is 32x32, the encoding unit may be divided into two 32x32 blocks for conversion. In this case, whether the encoding unit is divided for conversion is not separately signaled, but may be determined by comparing the width or height of the encoding unit with the width or height of the maximum conversion block. For example, if the width of the encoding unit is larger than the width of the maximum conversion block, the encoding unit may be divided vertically into two. In addition, if the vertical dimension of the encoding unit is greater than the vertical dimension of the maximum conversion block, the encoding unit can be divided horizontally into two halves.
[0167] Information regarding the maximum and / or minimum size of the encoding unit and information regarding the maximum and / or minimum size of the conversion block may be signaled or determined at an upper level of the encoding unit. The upper level may be, for example, a sequence level, a picture level, a tile level, a tile group level, a slice level, etc. For example, the minimum size of the encoding unit may be determined to be 4x4. For example, the maximum size of the conversion block may be determined to be 64x64. For example, the minimum size of the conversion block may be determined to be 4x4.
[0168] Information regarding the minimum size of an encoding unit corresponding to a leaf node of a quadtree (quadtree minimum size) and / or information regarding the maximum depth from the root node to a leaf node of a composite tree (composite tree maximum depth) may be signaled or determined at an upper level of the encoding unit. The upper level may be, for example, a sequence level, a picture level, a slice level, a tile group level, a tile level, etc. Information regarding the quadtree minimum size and / or information regarding the composite tree maximum depth may be signaled or determined for each of the in-frame slice and the inter-frame slice.
[0169] Difference information regarding the size of the CTU and the maximum size of the transform block may be signaled or determined at an upper level of the encoding unit. The upper level may be, for example, a sequence level, a picture level, a slice level, a tile group level, a tile level, etc. Information regarding the maximum size of the encoding unit corresponding to each node of the binary tree (binary tree maximum size) may be determined based on the size of the encoding tree unit and the difference information. The maximum size of the encoding unit corresponding to each node of the triad tree (triad tree maximum size) may have different values depending on the type of slice. For example, in the case of an in-frame slice, the triad tree maximum size may be 32x32. Also, for example, in the case of an inter-frame slice, the triad tree maximum size may be 128x128. For example, the minimum size of the encoding unit corresponding to each node of the binary tree (binary tree minimum size) and / or the minimum size of the encoding unit corresponding to each node of the triad tree (triad tree minimum size) can be set as the minimum size of the encoding block.
[0170] As another example, the maximum size of a binary tree and / or the maximum size of a triad tree can be signaled or determined at the slice level. Also, the minimum size of a binary tree and / or the minimum size of a triad tree can be signaled or determined at the slice level.
[0171] Based on the size and depth information of the various blocks mentioned above, quad splitting information, information on whether a composite tree is split, splitting tree information and / or splitting direction information, etc., may or may not exist in the bitstream.
[0172] For example, if the size of the encoding unit is not larger than the minimum size of the quadtree, the encoding unit does not include quad splitting information, and the said quad splitting information can be inferred as a second value.
[0173] For example, if the size (width and height) of a encoding unit corresponding to a node of a composite tree is larger than the maximum size (width and height) of a binary tree and / or the maximum size (width and height) of a three-part tree, the encoding unit may not be divided into a binary tree and / or a three-part tree. Accordingly, information on whether the composite tree is divided is not signaled and can be inferred as a second value.
[0174] Alternatively, if the size (width and height) of the encoding unit corresponding to the node of the composite tree is equal to the minimum size (width and height) of the binary tree, or if the size (width and height) of the encoding unit is equal to twice the minimum size (width and height) of the 3-partition tree, the encoding unit may not be divided into a binary tree and / or a 3-partition tree. Accordingly, information regarding whether the composite tree is divided is not signaled and can be inferred as a second value. This is because if the encoding unit is divided into a binary tree and / or a 3-partition tree, an encoding unit smaller than the minimum size of the binary tree and / or the minimum size of the 3-partition tree is generated.
[0175] Alternatively, binary tree splitting or three-part tree splitting may be limited based on the size of a virtual pipeline data unit (hereinafter referred to as the pipeline buffer size). For example, if a encoding unit is divided into sub-coding units that are not suitable for the pipeline buffer size by binary tree splitting or three-part tree splitting, said binary tree splitting or three-part tree splitting may be limited. The pipeline buffer size may be the size of a maximum transform block (e.g., 64X64). For example, when the pipeline buffer size is 64X64, the following splitting may be limited.
[0176] - 3-partition tree partitioning for NxM (N and / or M are 128) encoding units
[0177] - Horizontal binary tree partitioning for 128xN (N <= 64) encoding units
[0178] - Vertical binary tree partitioning for Nx128 (N <= 64) encoding units
[0179] Alternatively, if the depth of the encoding unit within the composite tree corresponding to the node of the composite tree is equal to the maximum depth of the composite tree, the encoding unit may not be divided into a binary tree and / or a three-way tree. Accordingly, information regarding whether the composite tree is divided is not signaled and can be inferred as a second value.
[0180] Alternatively, information on whether the composite tree is divided can be signaled only when at least one of vertical binary tree splitting, horizontal binary tree splitting, vertical three-way splitting, and horizontal three-way splitting is possible for the encoding unit corresponding to the node of the composite tree. Otherwise, the encoding unit may not be divided into a binary tree and / or a three-way splitting. Accordingly, information on whether the composite tree is divided is not signaled and can be inferred as a second value.
[0181] Alternatively, the division direction information may be signaled only when both vertical binary tree division and horizontal binary tree division are possible for the encoding unit corresponding to the node of the composite tree, or when both vertical 3-division tree division and horizontal 3-division tree division are possible. Otherwise, the division direction information is not signaled and may be inferred as a value indicating a direction in which division is possible.
[0182] Alternatively, the split tree information may be signaled only when both vertical binary tree splitting and vertical triad tree splitting are possible for the encoding unit corresponding to the node of the composite tree, or when both horizontal binary tree splitting and horizontal triad tree splitting are possible. Otherwise, the split tree information is not signaled and may be inferred as a value indicating a splittable tree.
[0183] Figure 4 is a diagram illustrating an example of an in-screen prediction process.
[0184] The arrows from the center to the outer edge of Fig. 4 may indicate the prediction directions of the in-screen prediction modes.
[0185] Intra-frame encoding and / or decoding may be performed using reference samples from neighboring blocks of the current block. Neighboring blocks may be restored neighboring blocks. For example, intra-frame encoding and / or decoding may be performed using the values of reference samples or encoding parameters contained in the restored neighboring blocks.
[0186] A prediction block may refer to a block generated as a result of performing in-frame prediction. A prediction block may correspond to at least one of CU, PU, and TU. The unit of a prediction block may be at least one of the sizes of CU, PU, and TU. A prediction block may be a square-shaped block with sizes such as 2x2, 4x4, 16x16, 32x32, or 64x64, or a rectangular-shaped block with sizes such as 2x8, 4x8, 2x16, 4x16, and 8x16.
[0187] In-frame prediction can be performed according to the in-frame prediction mode for the current block. The number of in-frame prediction modes that the current block may have can be a predefined fixed value or a value determined differently based on the attributes of the prediction block. For example, the attributes of the prediction block may include the size and shape of the prediction block.
[0188] The number of in-frame prediction modes may be fixed at N regardless of the block size. Or, for example, the number of in-frame prediction modes may be 3, 5, 9, 17, 34, 35, 36, 65, or 67. Or, the number of in-frame prediction modes may vary depending on the block size and / or the type of color component. For example, the number of in-frame prediction modes may differ depending on whether the color component is a luminance signal or a chroma signal. For example, as the block size increases, the number of in-frame prediction modes may increase. Or, the number of in-frame prediction modes for a luminance component block may be greater than the number of in-frame prediction modes for a chroma component block.
[0189] The in-frame prediction mode may be a non-directional mode or a directional mode. The non-directional mode may be a DC mode or a Planar mode, and the angular mode may be a prediction mode having a specific direction or angle. The in-frame prediction mode may be represented by at least one of a mode number, a mode value, a mode number, a mode angle, or a mode direction. The number of in-frame prediction modes may be one or more M, including the non-directional and directional modes. A step of checking whether samples included in the restored surrounding blocks can be used as reference samples for the current block to predict the current block in-frame may be performed. If there are samples that cannot be used as reference samples for the current block, the sample value of the sample that cannot be used as a reference sample may be replaced with a value obtained by copying and / or interpolating at least one sample value among the samples included in the restored surrounding blocks, and then used as a reference sample for the current block.
[0190] FIG. 49 is a diagram illustrating reference samples available for in-frame prediction.
[0191] As illustrated in FIG. 49, at least one of reference sample lines 0 to 3 may be used for the in-frame prediction of the current block. In FIG. 49, samples of segment A and segment F may be padded with the nearest samples of segment B and segment E, respectively, instead of being taken from the restored neighboring block. Index information indicating the reference sample line to be used for the in-frame prediction of the current block may be signaled. If the top boundary of the current block is the boundary of the CTU, only reference sample line 0 may be available. Therefore, in this case, the index information may not be signaled. If a reference sample line other than reference sample line 0 is used, filtering for the prediction block described below may not be performed.
[0192] When predicting within a frame, a filter may be applied to at least one of the reference sample or the prediction sample based on at least one of the within-frame prediction mode and the size of the current block.
[0193] In Planner mode, when generating a prediction block for the current block, the sample value of the target sample can be generated using the weighted sum of the top and left reference samples of the current sample and the top-right and bottom-left reference samples of the current block, depending on the position of the target sample within the prediction block. Additionally, in DC mode, when generating a prediction block for the current block, the average value of the top and left reference samples of the current block can be used. Furthermore, in Directional mode, a prediction block can be generated using the top, left, top-right, and / or bottom-left reference samples of the current block. Real-valued interpolation may also be performed to generate the prediction sample value.
[0194] In the case of intra-frame prediction between color components, a prediction block for the current block of the second color component can be generated based on the corresponding restoration block of the first color component. For example, the first color component may be a luminance component, and the second color component may be a chrominance component. For intra-frame prediction between color components, parameters of a linear model between the first color component and the second color component may be derived based on a template. The template may include upper and / or left peripheral samples of the current block and corresponding upper and / or left peripheral samples of the restoration block of the first color component. For example, the parameters of the linear model may be derived using the sample value of the first color component having the maximum value among the samples in the template and the corresponding sample value of the second color component, and the sample value of the first color component having the minimum value among the samples in the template and the corresponding sample value of the second color component. Once the parameters of the linear model are derived, the corresponding restoration block can be applied to the linear model to generate a prediction block for the current block. Depending on the image format, subsampling may be performed on the peripheral samples of the restoration block of the first color component and the corresponding restoration block. For example, if one sample of the second color component corresponds to four samples of the first color component, one corresponding sample can be calculated by subsampling the four samples of the first color component. In this case, parameter derivation of the linear model and intra-frame prediction between color components can be performed based on the subsampled corresponding sample. Whether to perform intra-frame prediction between color components and / or the range of the template can be signaled as an intra-frame prediction mode.
[0195] The current block can be divided into two or four sub-blocks in the horizontal or vertical direction. The divided sub-blocks can be restored sequentially. That is, an intra-frame prediction can be performed on the sub-blocks to generate sub-predicted blocks. Additionally, inverse quantization and / or inverse transformation can be performed on the sub-blocks to generate sub-residual blocks. A restored sub-block can be generated by adding the sub-predicted block to the sub-residual block. The restored sub-block can be used as a reference sample for intra-frame prediction of a lower-priority sub-block. A sub-block may be a block containing a predetermined number (e.g., 16) or more samples. Thus, for example, if the current block is an 8x4 block or a 4x8 block, the current block can be divided into two sub-blocks. Also, if the current block is a 4x4 block, the current block cannot be divided into sub-blocks. If the current block has other sizes, the current block can be divided into four sub-blocks. Information regarding whether the sub-block-based intra-frame prediction is performed and / or the division direction (horizontal or vertical) may be signaled. The sub-block-based intra-frame prediction may be restricted to be performed only when using reference sample line 0. When the sub-block-based intra-frame prediction is performed, filtering for the prediction block described below may not be performed.
[0196] A final prediction block can be generated by performing filtering on the predicted prediction block within the screen. The filtering can be performed by applying a predetermined weight to the filtering target sample, the left reference sample, the top reference sample, and / or the top-left reference sample. The weight and / or reference samples (range, position, etc.) used for the filtering can be determined based on at least one of the block size, the in-screen prediction mode, and the position of the filtering target sample within the prediction block. The filtering can be performed only in the case of a predetermined in-screen prediction mode (e.g., DC, planar, vertical, horizontal, diagonal, and / or adjacent diagonal mode). The adjacent diagonal mode may be a mode obtained by adding or subtracting k from the diagonal mode. For example, k may be a positive integer less than or equal to 8.
[0197] The in-frame prediction mode of the current block can be entropy encoded / decoded by predicting it from the in-frame prediction mode of blocks existing in the vicinity of the current block. If the in-frame prediction modes of the current block and the surrounding blocks are identical, information indicating that the in-frame prediction modes of the current block and the surrounding blocks are identical can be signaled using predetermined flag information. Additionally, indicator information regarding the in-frame prediction mode that is identical to the in-frame prediction mode of the current block among multiple in-frame prediction modes of surrounding blocks can be signaled. If the in-frame prediction modes of the current block and the surrounding blocks are different, the in-frame prediction mode information of the current block can be entropy encoded / decoded by performing entropy encoding / decoding based on the in-frame prediction modes of the surrounding blocks.
[0198] Figure 5 is a diagram illustrating an example of an inter-frame prediction process.
[0199] The rectangle shown in Fig. 5 can represent an image. Additionally, the arrow in Fig. 5 can indicate the prediction direction. Each image can be classified into I-picture (Intra Picture), P-picture (Predictive Picture), B-picture (Bi-predictive Picture), etc., depending on the encoding type.
[0200] Picture I can be encoded / decoded through intra-frame prediction without inter-frame prediction. Picture P can be encoded / decoded through inter-frame prediction using only reference images existing in a unidirectional direction (e.g., forward or reverse). Picture B can be encoded / decoded through inter-frame prediction using reference images existing in both directions (e.g., forward and reverse). Additionally, in the case of Picture B, it can be encoded / decoded through inter-frame prediction using reference images existing in both directions, or through inter-frame prediction using reference images existing in either the forward or reverse direction. Here, the bidirectional direction may be the forward and reverse directions. Here, when inter-frame prediction is used, the encoder may perform inter-frame prediction or motion compensation, and the decoder may perform corresponding motion compensation.
[0201] Below, the inter-screen prediction according to the embodiment is described in detail.
[0202] Inter-frame prediction or motion compensation can be performed using reference images and motion information.
[0203] Motion information for the current block can be derived during inter-frame prediction by each of the encoding device (100) and the decoding device (200). Motion information can be derived using motion information of restored surrounding blocks, motion information of a collocated block, and / or a block adjacent to the collocated block. A collocated block may be a block corresponding to the spatial position of the current block within an already restored collocated picture. Here, the collocated picture may be one picture among at least one reference picture included in a reference picture list.
[0204] The method of deriving motion information may vary depending on the prediction mode of the current block. For example, prediction modes applied for inter-frame prediction may include AMVP mode, merge mode, skip mode, merge mode with motion vector difference, sub-block merge mode, triangulation mode, inter-intra combined prediction mode, and affine inter mode. Here, the merge mode can be referred to as the motion merge mode.
[0205] For example, when AMVP is applied as a prediction mode, a motion vector candidate list can be generated by determining at least one of the motion vector of a restored surrounding block, the motion vector of a call block, the motion vector of a block adjacent to the call block, and the (0, 0) motion vector as a motion vector candidate. Motion vector candidates can be derived using the generated motion vector candidate list. Motion information of the current block can be determined based on the derived motion vector candidates. Here, the motion vector of the call block or the motion vector of a block adjacent to the call block can be referred to as a temporal motion vector candidate, and the motion vector of a restored surrounding block can be referred to as a spatial motion vector candidate.
[0206] The encoding device (100) can calculate the Motion Vector Difference (MVD) between the motion vector of the current block and a motion vector candidate, and can entropy-encode the MVD. Additionally, the encoding device (100) can generate a bitstream by entropy-encoding a motion vector candidate index. The motion vector candidate index can indicate the optimal motion vector candidate selected from among the motion vector candidates included in the motion vector candidate list. The decoding device (200) entropy-decodes the motion vector candidate index from the bitstream and can select a motion vector candidate for the block to be decoded from among the motion vector candidates included in the motion vector candidate list using the entropy-decoded motion vector candidate index. Additionally, the decoding device (200) can derive the motion vector of the block to be decoded through the sum of the entropy-decoded MVD and the motion vector candidate.
[0207] Meanwhile, the encoding device (100) can entropy-encode the resolution information of the calculated MVD. The decoding device (200) can adjust the resolution of the entropy-decoded MVD using the MVD resolution information.
[0208] Meanwhile, the encoding device (100) can calculate the Motion Vector Difference (MVD) between the motion vector of the current block and motion vector candidates based on an affine model, and can entropy encode the MVD. The decoding device (200) can derive the affine control motion vector of the block to be decoded by deriving the affine control motion vector of the block to be decoded through the sum of the entropy decoded MVD and the affine control motion vector candidates, thereby deriving motion vectors in sub-block units.
[0209] The bitstream may include a reference image index indicating a reference image. The reference image index may be entropy encoded and signaled from the encoding device (100) to the decoding device (200) via the bitstream. The decoding device (200) may generate a prediction block for a block to be decoded based on the induced motion vector and the reference image index information.
[0210] Another example of a method for deriving motion information is merge mode. Merge mode may refer to the merging of motions for multiple blocks. Merge mode may refer to a mode in which motion information of the current block is derived from the motion information of surrounding blocks. When merge mode is applied, a merge candidate list can be generated using the restored motion information of surrounding blocks and / or the motion information of the call block. Motion information may include at least one of 1) a motion vector, 2) a reference image index, and 3) an inter-frame prediction indicator. The prediction indicator may be unidirectional (L0 prediction, L1 prediction) or bidirectional.
[0211] The merge candidate list may represent a list in which motion information is stored. The motion information stored in the merge candidate list may be at least one of motion information of neighboring blocks adjacent to the current block (spatial merge candidate), motion information of a block collocated with the current block in a reference image (temporal merge candidate), new motion information generated by a combination of motion information already existing in the merge candidate list, motion information of a block encoded / decoded prior to the current block (history-based merge candidate), and zero merge candidate.
[0212] The encoding device (100) can generate a bitstream by entropy encoding at least one of a merge flag and a merge index and then signal it to the decoding device (200). The merge flag may be information indicating whether to perform a merge mode on a block-by-block basis, and the merge index may be information regarding which block among the surrounding blocks adjacent to the current block will be merged with. For example, the surrounding blocks of the current block may include at least one of the left adjacent block, the top adjacent block, and the temporally adjacent block of the current block.
[0213] Meanwhile, the encoding device (100) can entropy-encode correction information for correcting the motion vector among the motion information of the merge candidate and signal it to the decoding device (200). The decoding device (200) can correct the motion vector of the merge candidate selected by the merge index based on the correction information. Here, the correction information may include at least one of correction status information, correction direction information, and correction magnitude information. As described above, the prediction mode for correcting the motion vector of the merge candidate based on the signaled correction information can be referred to as a merge mode having a motion vector difference.
[0214] Skip mode may be a mode that applies the motion information of surrounding blocks directly to the current block. When skip mode is used, the encoding device (100) may entropy-encode information regarding which block's motion information to use as the motion information of the current block and signal it to the decoding device (200) via a bitstream. At this time, the encoding device (100) may not signal to the decoding device (200) any syntax elements regarding at least one of motion vector difference information, encoding block flags, and transform coefficient levels (quantized levels).
[0215] The sub-block merge mode may refer to a mode that derives motion information at the sub-block level of a coding block (CU). When the sub-block merge mode is applied, a sub-block merge candidate list may be generated using motion information of the sub-block collocated with the current sub-block in the reference image (sub-block based temporal merge candidate) and / or affine control point motion vector merge candidate.
[0216] The triangle partition mode may refer to a mode in which the current block is divided diagonally to derive movement information for each, each prediction sample is derived using the derived movement information, and each prediction sample is derived by weighted summing the derived prediction samples to derive the prediction sample of the current block.
[0217] The inter-intra combined prediction mode may refer to a mode that derives the prediction sample of the current block by weighting the prediction sample generated by inter-frame prediction and the prediction sample generated by intra-frame prediction.
[0218] The decoding device (200) can self-correct the derived motion information. The decoding device (200) can derive the motion information having the minimum SAD into corrected motion information by searching a predefined area based on the reference block indicated by the derived motion information.
[0219] The decoding device (200) can compensate for prediction samples derived through inter-frame prediction using optical flow.
[0220] Figure 6 is a diagram illustrating the process of transformation and quantization.
[0221] As illustrated in FIG. 6, a quantized level may be generated by performing a transformation and / or quantization process on the residual signal. The residual signal may be generated as the difference between the original block and the prediction block (intra-frame prediction block or inter-frame prediction block). Here, the prediction block may be a block generated by intra-frame prediction or inter-frame prediction. Here, the transformation may include at least one of a first transformation and a second transformation. Transformation coefficients may be generated by performing a first transformation on the residual signal, and second transformation coefficients may be generated by performing a second transformation on the transformation coefficients.
[0222] The primary transform may be performed using at least one of a plurality of predefined transform methods. For example, the plurality of predefined transform methods may include a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or a Karhunen-Loeve Transform (KLT)-based transform. A secondary transform may be performed on the transform coefficients generated after the primary transform is performed. The transform method applied during the primary transform and / or secondary transform may be determined based on at least one of the encoding parameters of the current block and / or surrounding blocks. Alternatively, transform information indicating the transform method may be signaled. The DCT-based transform may include, for example, DCT2, DCT-8, etc. The DST-based transform may include, for example, DST-7.
[0224] Quantized levels can be generated by performing quantization on the result of a first transformation and / or a second transformation, or on the residual signal. The quantized levels can be scanned according to at least one of an up-right diagonal scan, a vertical scan, and a horizontal scan based on at least one of an in-frame prediction mode or block size / shape. For example, the coefficients of a block can be converted into a one-dimensional vector form by scanning them using an up-right diagonal scan. Depending on the size of the transformed block and / or the in-frame prediction mode, a vertical scan that scans the two-dimensional block shape coefficients in the column direction, or a horizontal scan that scans the two-dimensional block shape coefficients in the row direction, may be used instead of an up-right diagonal scan. The scanned quantized levels can be entropy-encoded and included in a bitstream.
[0225] In the decoder, the bitstream can be entropy decoded to generate quantized levels. The quantized levels can be inverse scanned and aligned into a two-dimensional block shape. At this time, at least one of an upper-right diagonal scan, a vertical scan, and a horizontal scan can be performed as a method of inverse scanning.
[0226] Inverse quantization can be performed on the quantized level, and depending on whether a second inverse transform is performed, a second inverse transform can be performed, and depending on whether a first inverse transform is performed on the result of the second inverse transform, a first inverse transform can be performed to generate a restored residual signal.
[0227] Inverse mapping of the dynamic range can be performed on the luminance component restored through intra-frame prediction or inter-frame prediction before in-loop filtering. The dynamic range can be divided into 16 equal pieces, and a mapping function for each piece can be signaled. The mapping function can be signaled at the slice level or the tile group level. An inverse mapping function for performing the inverse mapping can be derived based on the mapping function. In-loop filtering, saving of the reference picture, and motion compensation are performed in the inversely mapped area, and the prediction block generated through inter-frame prediction can be used to generate the restoration block after being converted to the mapped area by mapping using the mapping function. However, since intra-frame prediction is performed in the mapped area, the prediction block generated by intra-frame prediction can be used to generate the restoration block without mapping / inverse mapping.
[0228] If the current block is a residual block of a chrominance component, the residual block can be converted into an inversely mapped area by performing scaling on the chrominance component of the mapped area. The availability of the scaling can be signaled at the slice level or the tile group level. The scaling may be applied only when the mapping for the luminance component is available and the division of the luminance component and the division of the chrominance component follow the same tree structure. The scaling may be performed based on the average of the sample values of the luminance prediction block corresponding to the chrominance block. In this case, if the current block uses inter-frame prediction, the luminance prediction block may refer to the mapped luminance prediction block. By referencing a lookup table using the index of the piece to which the average of the sample values of the luminance prediction block belongs, the value required for the scaling can be derived. Finally, by scaling the residual block using the derived value, the residual block can be converted into an inversely mapped area. Subsequent restoration of color difference component blocks, intra-frame prediction, inter-frame prediction, in-loop filtering, and saving of the reference picture can be performed in the inversely mapped area.
[0229] Information indicating whether mapping / inverse mapping of the above luminance component and color difference component is available can be signaled through a sequence parameter set.
[0230] The predicted block of the current block can be generated based on a block vector representing the displacement between the current block and the reference block within the current picture. In this way, the prediction mode that generates the predicted block by referencing the current picture can be named the Intra Block Copy (IBC) mode. The IBC mode can be applied to an MxN (M <= 64, N <= 64) encoding unit. The IBC mode may include skip mode, merge mode, AMVP mode, etc. In the case of skip mode or merge mode, a merge candidate list is constructed, and a merge index is signaled to identify a single merge candidate. The block vector of the identified merge candidate can be used as the block vector of the current block. The merge candidate list may include at least one of the following: a spatial candidate, a history-based candidate, a candidate based on the average of two candidates, or a zero-merge candidate. In the case of AMVP mode, a difference block vector may be signaled. Additionally, the predicted block vector can be derived from the left neighbor block and the top neighbor block of the current block. An index regarding which neighbor block to use can be signaled. The predicted block in IBC mode may be limited to a block within a previously restored area that is included in the current CTU or the left CTU. For example, the value of the block vector may be restricted so that the predicted block of the current block is located within the three 64x64 block areas that precede the 64x64 block to which the current block belongs in terms of encoding / decoding order. By restricting the value of the block vector in this way, memory consumption and device complexity associated with the implementation of IBC mode can be reduced.
[0232] In the following, we will explain a method for improving video compression efficiency by improving the transform method, which is one of the video coding processes. More specifically, the encoding of conventional video coding largely involves an intra / inter-frame prediction step that predicts the original block, which is a part of the current original image; a transformation and quantization step for the residual block, which is the difference between the predicted prediction block and the original block; and an entropy coding step, which is a probability-based lossless compression method for the coefficients of the transformed and quantized block and the compression information obtained from the preceding stage, to form a bitstream, which is the compressed form of the original image, and transmit it to a decoder or store it on a recording medium. The Shuffling and Discrete Sine Transform (hereinafter “SDST”) described below in this specification is intended to improve compression efficiency by increasing the efficiency of the transformation.
[0233] The SDST method according to the present invention can better reflect the common frequency characteristics of images by using Discrete Sine Transform type-7 (Discrete Sine Transform type-7, hereinafter “DST-VII” or “DST-7”) instead of Discrete Cosine Transform type-2 (hereinafter “DCT-II” or “DCT-2”), which is a transform kernel widely used in video coding.
[0234] According to the conversion method of the present invention, high objective video quality can be obtained even with a relatively low bit rate compared to conventional video coding methods.
[0235] DST-7 can be applied to the data of a residual block. The application of DST-7 to the residual block can be performed based on a prediction mode corresponding to the residual block. For example, it can be applied to a residual block encoded in an inter-mode (inter-frame mode). According to one embodiment of the present invention, DST-7 can be applied after rearranging or shuffling the data of the residual block. Here, shuffling refers to the rearrangement of image data and can be referred to as residual signal rearrangement or flipping in an equivalent sense. Here, residual block may have the same meaning as residual, residual block, residual signal, residual signal, residual data, or residual data. In addition, the residual block may have the same meaning as the reconstructed residual, reconstructed residual block, reconstructed residual signal, reconstructed residual signal, reconstructed residual data, or reconstructed residual data, which are the forms in which the residual block is reconstructed from the encoder and decoder.
[0236] An SDST according to one embodiment of the present invention may use DST-7 as a transformation kernel. At this time, the transformation kernel of the SDST is not limited to DST-7, but may include Discrete Sine Transform type-1 (DST-1), Discrete Sine Transform type-2 (DST-2), Discrete Sine Transform type-3 (DST-3), … , Discrete Sine Transform type-n (DST-n), Discrete Cosine Transform type-1 (DT-1), Discrete Cosine Transform type-2 (DT-2), Discrete Cosine Transform type-3 (DT-3), … At least one of various types of DST and DCT, such as Discrete Cosine Transform type-n (DCT-n), may be used. (Here, n is a positive integer greater than or equal to 1)
[0237] Mathematical Equation 1 below can represent a method for performing a one-dimensional DCT-2 according to an embodiment of the present invention. Here, N is the block size, k is the position of the frequency component, and x n It can represent the value of the nth coefficient in the spatial domain.
[0238]
[0239] DCT-2 in a two-dimensional domain can be achieved by performing a horizontal transform and a vertical transform on the remaining block using the above mathematical formula 1.
[0240] The DCT-2 transformation kernel can be defined by Equation 2 as follows. Here, Xk is a basis vector based on position in the frequency domain, and N can represent the magnitude of the frequency domain.
[0241]
[0242] Meanwhile, FIG. 7 is a diagram illustrating the basis vectors of DCT-2 in the frequency domain according to the present invention. FIG. 7 shows the frequency characteristics of DCT-2 in the frequency domain. Here, the value calculated through the X0 basis vector of DCT-2 may represent the DC component.
[0243] DCT-2 can be used in the conversion process for residual blocks of sizes such as 4x4, 8x8, 16x16, and 32x32.
[0244] Meanwhile, DCT-2 may be selectively used based on at least one of the size of the residual block, the color component of the residual block (e.g., luminance component, chrominance component), or the prediction mode corresponding to the residual block. For example, if the residual block is 4x4 in size and encoded in intra mode (in-frame mode), and the component of the residual block is a luminance component, DCT-2 may not be used. For example, if the horizontal length of the residual block encoded in intra mode falls within a predetermined range (e.g., 4 pixels or more and 16 pixels or less) and the horizontal length is not greater than the vertical length, a first transformation kernel may be used for horizontal transformation. Otherwise, a second transformation kernel may be used for horizontal transformation. For example, if the vertical length of the residual block encoded in intra mode falls within 4 pixels or more and 16 pixels or less and the vertical length is not greater than the horizontal length, a first transformation kernel may be used for vertical transformation. Otherwise, a second transformation kernel may be used for vertical transformation. The first and second transformation kernels mentioned above may differ. That is, the horizontal and vertical transformation methods of a block encoded in intra mode may be implicitly determined based on the block shape under certain conditions. For example, the first transformation kernel may be DST-7 and the second transformation kernel may be DCT-2. In this case, since the residual block is the target of transformation, it may have the same meaning as the transformation block. Here, the prediction mode may refer to inter-frame prediction or intra-frame prediction. Furthermore, if the prediction mode is intra-frame prediction, it may refer to the intra-frame prediction mode or the intra-frame prediction direction.
[0245] Transformation using the DCT-2 transform kernel can demonstrate high compression efficiency for blocks with characteristics of minimal variation between neighboring pixels, such as image backgrounds. However, it may not be suitable as a transform kernel for regions with complex patterns, such as texture images. This is because transforming blocks with low correlation between neighboring pixels using DCT-2 can result in a large number of transform factors occurring in the high-frequency components of the frequency domain. Frequent occurrence of transform factors in the high-frequency region can reduce image compression efficiency. To improve compression efficiency, large values should occur near low-frequency components, while values should be as close to zero as possible in high-frequency components.
[0246] Mathematical Equation 3 below may represent a method for performing a one-dimensional DST-7 according to an embodiment of the present invention. Here, N is the block size, k is the position of the frequency component, and x n can mean the value of the nth coefficient in the spatial domain.
[0247]
[0248] DST-7 in a two-dimensional domain can be achieved by performing a horizontal transform and a vertical transform on the remaining block using the above mathematical formula 3.
[0249] The DST-7 transformation kernel can be defined by the following Equation 4. Here, X k can represent the Kth basis vector of DST-7, i can represent the position in the frequency domain, and N can represent the magnitude of the frequency domain.
[0250]
[0251] DST-7 can be used in a conversion process for residual blocks of at least one size among 2x2, 4x4, 8x8, 16x16, 32x32, 64x64, 128x128, etc.
[0252] Meanwhile, DST-7 can be applied to rectangular blocks rather than square blocks. For example, DST-7 can be applied to at least one of the vertical and horizontal transformations of rectangular blocks with different widths and heights, such as 8x4, 16x8, 32x4, and 64x16. If the selective application of multiple transformation methods is available, DCT-2 can be applied to the horizontal and vertical transformations of square blocks. If the selective application of multiple transformation methods is not available, DST-7 can be applied to the horizontal and vertical transformations of square blocks.
[0253] Additionally, DST-7 may be selectively used based on at least one of the size of the residual block, the color components of the residual block (e.g., luminance component, chrominance component), the prediction mode corresponding to the residual block, the intra-frame prediction mode (direction), and the shape of the residual block. For example, DST-7 may be used when the residual block is 4x4 in size and the component of the residual block is a luminance component. In this case, the prediction mode may mean inter-frame prediction or intra-frame prediction. Additionally, if the prediction mode is intra-frame prediction, it may mean the intra-frame prediction mode or the intra-frame prediction direction. For example, for the chrominance component, the selection of a conversion method based on the block shape may not be available. For example, if the intra-frame prediction mode is between color components, the selection of a conversion method based on the block shape may not be available. For example, the conversion method for the chrominance component may be specified by information signaled through the bitstream. When a current block is divided into multiple sub-blocks and intra-frame prediction is performed for each sub-block, the transformation method for the current block may be determined based on the intra-frame prediction mode and / or the size of the block (horizontal and / or vertical size). For example, if the intra-frame prediction mode is non-directional (DC or Planar), and the horizontal length (vertical length) falls within a predetermined range, a first transformation kernel may be used for horizontal transformation (vertical transformation), and if not, a second transformation kernel may be used. The first transformation kernel and the second transformation kernel may be different. For example, the first transformation kernel may be DST-7 and the second transformation kernel may be DCT-2. The predetermined range may be, for example, 4 pixels to 16 pixels. If the size of the block does not fall within the predetermined range, the same kernel (e.g., the second transformation kernel) may be used for horizontal and vertical transformation.When the size of the block falls within the above-mentioned predetermined range, different transformation kernels may be used for adjacent in-frame prediction modes. For example, if the second transformation kernel and the first transformation kernel are used for the horizontal transformation and vertical transformation of mode 27, respectively, the first transformation kernel and the second transformation kernel may be used for the horizontal transformation and vertical transformation of mode 26 and mode 28 adjacent to mode 27, respectively.
[0254] Meanwhile, FIG. 8 is a diagram illustrating the basis vectors of DST-7 in each frequency domain according to the present invention. Referring to FIG. 8, the first basis vector (x0) of DST-7 has a curved shape. Through this, it can be predicted that DST-7 will exhibit higher transformation performance for blocks with large spatial changes in the image compared to DCT-2.
[0255] DST-7 can be used for transformations on 4x4 transformation units (TU) within intra-predicted coding units (CU). This allows for the use of DST-7, which exhibits higher transformation efficiency, by reflecting the fact that the amount of error increases as the distance from the reference sample increases due to the nature of intra prediction. In other words, for blocks in the spatial domain where the amount of residual signal increases as the distance from the (0, 0) position within the block increases, DST-7 can be used to efficiently compress the block.
[0256] As described above, it may be important to use a conversion kernel that matches the frequency characteristics of the image to increase conversion efficiency. In particular, since the conversion is performed on the residual block relative to the original block, the conversion efficiency of DST-7 and DCT-2 can be determined by checking the distribution characteristics of the residual signal within the CU, PU, or TU block.
[0257] Figure 9 is a diagram showing the distribution of average residual values according to the position within the 2Nx2N prediction unit (PU) of the 8x8 coding unit (CU) predicted in inter mode, obtained by experimenting with the “Cactus” sequence in a Low Delay-P profile environment.
[0258] Referring to Fig. 9, the left diagram of Fig. 9 separately displays the top 30% of the relatively large values among the average residual signal values within the block, and the right diagram separately displays the top 70% of the relatively large values among the average residual signal values within the block, which are the same as the left diagram.
[0259] Through Figure 9, it can be seen that the residual signal distribution within the 2Nx2N PU of the 8x8 CU predicted in inter mode has the characteristic that values with small residual signal magnitudes are mainly concentrated near the center of the block, and the residual signal value increases as it moves away from the middle point of the block. In other words, it can be seen that the residual signal value increases at the block boundary. Such residual signal distribution characteristics may be a common feature of the residual signal within the PU regardless of the CU size and the PU partitioning mode (2Nx2N, 2NxN, Nx2N, NxN, nRx2N, nLx2N, 2NxnU, 2NxnD) that the predicted CU between screens can have.
[0260] Figure 10 is a three-dimensional graph showing the residual signal distribution characteristics within the 2Nx2N prediction unit (PU) of the 8x8 coding unit (CU) predicted in inter-mode prediction.
[0261] Referring to Figure 10, it can be seen that residual signals of relatively small values are concentrated near the center of the block, and residual signals closer to the block boundary have relatively larger values.
[0262] Based on the residual signal distribution characteristics according to Figures 9 and 10, the conversion of residual signals within the PU of the predicted inter-frame CU can be more efficient by using DST-7 instead of DCT-2.
[0263] Below, we will explain SDST, one of the conversion methods that uses DST-7 as the conversion kernel.
[0264] In the following, Block may mean any one of CU, PU, and TU.
[0265] SDST according to the present invention can be performed in two steps. The first step is to shuffle residual signals within the PU of the CU predicted in inter mode (inter-frame mode) or intra mode (intra-frame mode). The second step is to apply DST-7 to the residual signals within the shuffled block.
[0266] Residual signals arranged within the current block (e.g., CU, PU, or TU) can be scanned according to a first direction and rearranged according to a second direction. That is, shuffling can be performed by scanning the residual signals arranged within the current block according to a first direction and rearranging them according to a second direction. In this case, the residual signal may refer to a signal representing the difference signal between the original signal and the predicted signal. That is, the residual signal may refer to a signal prior to performing at least one of transformation and quantization. Alternatively, the residual signal may refer to a signal form prior to performing at least one of transformation and quantization. Furthermore, the residual signal may refer to a restored residual signal. That is, the residual signal may refer to a signal prior to performing at least one of inverse transformation and inverse quantization. Furthermore, the residual signal may refer to a signal prior to performing at least one of inverse transformation and inverse quantization.
[0267] Meanwhile, the first direction (or scan direction) may be any one of a raster scan order, an up-right diagonal scan order, a horizontal scan order, or a vertical scan order. Additionally, the first direction may be defined as at least one of (1) to (10) below.
[0268] (1) Scan from the top row to the bottom row, but scan from left to right within each row.
[0269] (2) Scan from the top row to the bottom row, but scan from right to left within each row.
[0270] (3) Scan from the bottom row to the top row, but scan from left to right within each row.
[0271] (4) Scan from the bottom row to the top row, but scan from right to left within each row.
[0272] (5) Scan from the left column to the right column, but scan from top to bottom in each column.
[0273] (6) Scan from the left column to the right column, but scan from the bottom to the top in each column.
[0274] (7) Scan from the right column to the left column, but scan from top to bottom in each column.
[0275] (8) Scan from the right column to the left column, but scan from the bottom to the top in each column.
[0276] (9) Scan in a spiral shape: Scan from inside (or outside) the block to outside (or inside) the block, clockwise / counterclockwise.
[0277] (10) Diagonal scan: Start from one vertex within the block and scan diagonally in the direction of top-left, top-right, bottom-left, or bottom-right.
[0278] Meanwhile, at least one of the scan directions of (1) to (10) above may also be selectively used as the second direction (or rearrangement direction). The first direction and the second direction may be the same or different from each other.
[0280] The scanning and rearrangement process for residual signals can currently be performed in block units.
[0281] Here, rearrangement may mean arranging residual signals scanned according to a first direction within a block according to a second direction into blocks of the same size. In this case, the size of the block scanned according to the first direction and the size of the block rearranged according to the second direction may be different from each other.
[0282] Additionally, although scanning and rearrangement are described here as being performed separately according to the first and second directions, respectively, scanning and rearrangement can be performed as a single process for the first direction. For example, residual signals within a block can be scanned from the top row to the bottom row, while scanning from right to left within a single row to store (rearrange) them in the block.
[0283] Meanwhile, the scanning and rearrangement process for residual signals can be performed in units of a predetermined sub-block within the current block. Here, the sub-block may be a block equal to or smaller than the current block. The sub-block may be a block divided from the current block into a quadtree, binary tree, etc.
[0284] Subblock units may have a fixed size and / or shape (e.g., 4x4, 4x8, 8x8, … NxM, where N and M are positive integers). Additionally, the size and / or shape of subblock units may be derived variably. For example, the size and / or shape of subblock units may be determined dependently on the size, shape, and / or prediction mode (inter, intra) of the current block.
[0285] The scan direction and / or relocation direction may be adaptively determined based on the location of the subblock. In this case, different scan directions and / or relocation directions may be used for each subblock, or all or some of the subblocks belonging to the current block may use the same scan direction and / or relocation direction.
[0286] FIG. 11 is a diagram illustrating the distribution characteristics of residual signals in the 2Nx2N prediction unit (PU) mode of a coding unit (CU) according to the present invention.
[0287] Referring to Fig. 11, the PU is divided into four sub-blocks in a quadtree structure, and the direction of the arrow in each sub-block indicates the characteristics of the residual signal distribution. Specifically, the direction of the arrow in each sub-block indicates the direction in which the residual signal increases. This is due to the distribution characteristics that the residual signals within the PU have in common, regardless of the PU partitioning mode. Therefore, a shuffling operation can be performed to rearrange the residual signals in each sub-block to have distribution characteristics suitable for DST-7 transformation.
[0288] FIG. 12 is a diagram illustrating the residual signal distribution characteristics before and after shuffling of a 2Nx2N prediction unit (PU) according to the present invention.
[0289] Referring to Fig. 12, the upper block shows the distribution of residual signals within the 2Nx2N PU of the 8x8 CU predicted in inter-mode prior to shuffling. Equation 5 below represents the value according to the position of each residual signal within the upper block of Fig. 12.
[0290]
[0291] Due to the distribution characteristics of residual signals within the PU of the CU predicted in inter-mode, there are many residual signals with relatively small values distributed in the central region within the upper block of Fig. 12, and as one proceeds toward the boundary of the upper block, there are many residual signals with large values distributed.
[0292] The lower block of Fig. 12 shows the residual signal distribution characteristics within the 2Nx2N PU after shuffling. This shows that the distribution of residual signals per sub-block of the shuffled PU is a distribution of residual signals suitable for the first basis vector of DST-7. That is, since the residual signal within each sub-block has a larger value as it moves further away from the (0, 0) position, when performing the transformation, the transformation coefficient values frequency-transformed through DST-7 may appear concentrated in the low-frequency region.
[0293] Mathematical Equation 6 below represents a method for performing shuffling according to the position of each sub-block within the PU in four sub-blocks divided into a quadtree structure in the PU.
[0294]
[0295] Here, Wk and Hk are the k-th sub-blocks (k) in PU, respectively. {blk0, blk1, blk2, blk3}) represents the width or height, and blk0~blk3 represent each sub-block divided into a quadtree structure in the PU. Also, x and y represent the horizontal and vertical positions within each sub-block. a(x,y), b(x,y), c(x,y), and d(x,y) represent the positions of each residual signal before shuffling, as shown in the upper block of Fig. 12. a'(x,y), b'(x,y), c'(x,y), and d'(x,y) represent the positions of the residual signals changed through shuffling, as shown in the lower block of Fig. 12.
[0296] FIG. 13 is a diagram illustrating an example of 4x4 residual data rearrangement of a subblock according to the present invention.
[0297] Referring to FIG. 13, a subblock refers to any one of a plurality of subblocks belonging to an 8x8 prediction block. FIG. 13 (a) indicates the location of the original residual data before rearrangement, and FIG. 13 (b) indicates the rearranged location of the residual data.
[0298] Referring to Fig. 13 (c), the value of the residual data can gradually increase from the position (0,0) to the position (3,3). Here, the horizontal and / or vertical one-dimensional residual data within each sub-block may have a data distribution in the form of the basis vectors shown in Fig. 8.
[0299] That is, the shuffling according to the present invention can rearrange the residual data of each sub-block so that the residual data distribution is suitable for the form of the aforementioned DST-7 basis vector. After shuffling for each sub-block, a DST-7 transformation can be applied to the data rearranged on a sub-block basis.
[0300] Meanwhile, sub-blocks may be additionally partitioned into a quadtree structure based on the depth of the TU, or a rearrangement process may be optionally performed. For example, if the depth of the TU is 2, an NxN sub-block belonging to a 2Nx2N PU may be partitioned into N / 2xN / 2 blocks, and a rearrangement process may be applied to each N / 2xN / 2 block. Here, quadtree-based TU partitioning may be performed repeatedly until the minimum TU size is reached.
[0301] In addition, when the depth of the TU is 0, the DCT-2 transformation may be applied to 2Nx2N blocks. In this case, rearrangement of the remaining data may not be performed.
[0302] Meanwhile, since the SDST method according to the present invention utilizes the distribution characteristics of residual signals within the PU block, the partitioning structure of the TU performing SDST can be defined as being partitioned into a quadtree structure based on the PU.
[0303] FIGS. 14(a) and FIGS. 14(b) are diagrams illustrating an example of a conversion unit (TU) partition structure of a coding unit (CU) according to a prediction unit (PU) mode and a shuffling method of a conversion unit (TU). FIGS. 14(a) and FIGS. 14(b) show the quadtree partition structure of a TU according to TU depth for each asymmetric partition mode (2NxnU, 2NxnD, nRx2N, nLx2N) of an inter-predicted PU.
[0304] Referring to FIGS. 14(a) and FIGS. 14(b), the thick solid line in each block represents a PU within a CU, and the thin solid line represents a TU. Also, S0, S1, S2, and S3 within each TU represent the shuffling method of the residual signal within the TU defined in Equation 6 described above.
[0305] Meanwhile, in FIG. 14(a) and FIG. 14(b), the depth 0 TU of each PU has the same block size as the corresponding PU (e.g., in a 2Nx2N PU, the size of the depth 0 TU is the same as the size of the PU). Here, shuffling of the residual signal within the depth 0 TU is described later with reference to FIG. 18.
[0306] In addition, if at least one of CU, PU, and TU has a rectangular shape (e.g., 2NxnU, 2NxnD, nRx2N, nLx2N), at least one of CU, PU, and TU may be divided into N sub-blocks, such as 2, 4, 6, 8, or 16, before residual signal rearrangement, and residual signal rearrangement may be applied to the divided sub-blocks.
[0307] In addition, if at least one of CU, PU, and TU has a square shape (e.g., 2Nx2N, NxN), at least one of CU, PU, and TU may be divided into N sub-blocks, such as 4, 8, or 16, before residual signal rearrangement, and residual signal rearrangement may be applied to the divided sub-blocks.
[0308] In addition, when the above TU is divided from the CU or PU, if the TU is at the highest depth (if not divided), the TU can be divided into N sub-blocks, such as 2, 4, 6, 8, or 16, and then residual signal rearrangement can be performed on the divided sub-blocks.
[0309] In the above example, an example of performing residual signal rearrangement was shown when CU, PU, and TU each have different shapes or different sizes, but the residual signal rearrangement can also be applied when at least two of CU, PU, and TU have the same shape or the same size.
[0310] Meanwhile, FIGS. 14(a) and FIGS. 14(b) describe an asymmetric partitioning mode of an inter-predicted PU, but are not limited thereto, and the partitioning and shuffling of the TU can also be applied to a symmetric partitioning mode of the PU (2NxN, Nx2N).
[0311] DST-7 transformation can be performed on each TU within the shuffling PU. In this case, if the CU, PU, and TU all have the same size and shape, DST-7 transformation can be performed on a single block.
[0312] Considering the residual signal distribution characteristics of the inter-predicted PU block, performing the DST-7 transformation after shuffling may be more efficient than performing the DCT-2 transformation, regardless of the size of the CU and the PU splitting mode.
[0313] After conversion, if the conversion coefficients are more distributed near the low-frequency component (especially the DC component), the residual signal distribution shows higher compression efficiency compared to cases where they are not, in terms of i) minimizing energy loss after quantization and ii) reducing bit usage in the entropy coding process.
[0314] Figure 15 is a diagram illustrating the results of performing DCT-2 transformation and SDST transformation according to the residual signal distribution of the 2Nx2N prediction unit (PU).
[0315] The diagram on the left side of Fig. 15 shows a distribution in which the residual signal increases from the center toward the boundary when the PU partitioning mode of the CU is 2Nx2N. Additionally, the diagram in the middle of Fig. 15 shows the distribution of the residual signal after performing a DCT-2 transformation on a TU at depth 1 within the PU, and the diagram on the right side of Fig. 15 shows the distribution of the residual signal after performing a DST-7 transformation (SDST) after shuffling on a TU at depth 1 within the PU.
[0316] Referring to Figure 15, it can be seen that when SDST is performed on the TU of a PU having the residual signal distribution characteristics mentioned above, more coefficients are concentrated near the low-frequency components and coefficients on the high-frequency components have smaller values compared to when DCT-2 is performed. According to the above transformation characteristics, it can be seen that when transforming the residual signal of an inter-predicted PU, performing SDST instead of DCT-2 is advantageous in terms of compression efficiency.
[0317] The unit of the block on which the DST-7 transformation is performed is the TU unit defined in PU, and as explained with reference to FIG. 14, the TU can be divided into quadtrees or binary trees from the PU unit to the maximum depth. This means that the DST-7 transformation can be performed after shuffling not only on square blocks but also on rectangular blocks.
[0318] For example, for a block predicted between frames, a residual block of the same size as the block may be decoded, or a sub-residual block corresponding to a part of the block may be decoded. Information for this purpose may be signaled to the block, and said information may be, for example, a flag. When a residual block of the same size as the block is decoded, information regarding the transformation kernel may be determined by decoding information included in the bitstream. When a sub-residual block corresponding to a part of the block is decoded, a transformation kernel for the sub-residual block may be determined based on information specifying the type and / or location within the block of the sub-residual block. For example, information regarding the type and / or location within the block of said sub-residual block may be included in the bitstream and signaled. In this case, if the block is larger than 32x32, the determination of the transformation kernel based on the type and / or location within the block of the sub-residual block may not be performed. For example, for blocks larger than 32x32, a predetermined conversion kernel (e.g., DCT-2) may be applied, or information about the conversion kernel may be explicitly signaled. Alternatively, if the width or height of the block is greater than 32, the determination of the conversion kernel based on the type and / or location within the block of the sub-residual block may not be performed. For example, for 64x8 blocks, a predetermined conversion kernel (e.g., DCT-2) may be applied, or information about the conversion kernel may be explicitly signaled.
[0319] Information regarding the type of the sub-remaining block may be partition information of the block. The partition information of the block may be partition direction information indicating, for example, one of horizontal partition or vertical partition. Alternatively, the partition information of the block may include partition ratio information. For example, the partition ratio may include 1:1, 1:3, and / or 3:1. The partition direction information and the partition ratio information may be signaled as separate syntactic elements or as a single syntactic element.
[0320] Information regarding the location of the above-mentioned sub-remaining block may indicate a location within the block. For example, if the block is divided vertically, the information regarding the location may indicate either the left or the right. Additionally, if the block is divided horizontally, the information regarding the location may indicate either the top or the bottom.
[0321] The transformation kernel of the above sub-remaining block may be determined based on the type information and / or position information. The transformation kernel may be determined independently for horizontal transformation and vertical transformation. For example, the transformation kernel may be determined based on the division direction. For example, in the case of vertical division, a first transformation kernel may be applied to the vertical transformation, and in the case of horizontal division, a first transformation kernel may be applied to the horizontal transformation. For example, a first transformation kernel or a second transformation kernel may be applied to the horizontal transformation in the case of vertical division and the vertical transformation in the case of horizontal division. For example, in the case of vertical division, a second transformation kernel may be applied to the horizontal transformation of the left position, and a first transformation kernel may be applied to the horizontal transformation of the right position. In addition, in the case of horizontal division, a second transformation kernel may be applied to the vertical transformation of the upper position, and a first transformation kernel may be applied to the vertical transformation of the lower position. For example, the first transformation kernel and the second transformation kernel may be DST-7 and DCT-8, respectively. For example, the first transformation kernel and the second transformation kernel may be DST-7 and DCT-2, respectively. However, this is not limited thereto, and any two different transformation kernels among the various transformation kernels mentioned in this specification may be used as the first and second transformation kernels, respectively. Here, the block may mean a CU or a TU. Additionally, the sub-residual block may mean a Sub-TU.
[0322] FIG. 16 is a diagram illustrating the SDST process according to the present invention.
[0323] The residual signal of the TU to be transformed is input (S1610). At this time, the TU may be a divided TU within a PU in which the prediction mode is inter-mode. Shuffling may be performed on the TU to be transformed (S1620). Then, the SDST process may proceed in the order of performing DST-7 transformation on the shuffled TU (S1630), quantization (S1640), and a series of subsequent processes.
[0324] Meanwhile, shuffling and DST-7 conversion can also be performed on blocks where the prediction mode is intra mode.
[0325] Hereinafter, as an example of implementing SDST transformation in an encoder, i) a method of performing SDST for all TUs within an inter-predicted PU and ii) a method of selectively performing SDST or DCT-2 through Rate-Distortion Optimization will be described. Although the following method describes an inter-predicted block, it is not limited thereto and the following method may also be applied to an intra-predicted block.
[0326] FIG. 17 is a diagram illustrating the distribution characteristics of the size of the division and residual absolute value of the conversion unit (TU) according to the partition mode of the prediction unit (PU) of the inter-frame predicted coding unit (CU) according to the present invention.
[0327] Referring to FIG. 17, in the inter-prediction mode, TU can be partitioned into a quadtree or a binary tree from CU up to a maximum depth, and there can be a total of K cases for the partition mode of PU. Here, K is a positive integer, for example, K is 8 in FIG. 17.
[0328] The SDST according to the present invention utilizes the residual signal distribution characteristics in the PU within the inter-predicted CU as described with reference to FIG. 10. Additionally, the TU can be divided from the PU into a quadtree structure or a binary tree structure. That is, a TU with a depth of 0 can correspond to a PU, and a TU with a depth of 1 can correspond to each of the subblocks obtained by dividing the PU into a quadtree structure or a binary tree structure once.
[0329] Each block of FIG. 17 shows the form in which a TU is divided to a depth of 2 for each of the PU division modes of the inter-predicted CU. Here, a thick solid line represents a PU, a thin solid line represents a TU, and the direction of the arrow of each TU may indicate the direction in which the residual signal value within the corresponding TU increases. Each TU may perform the shuffling mentioned in the description of the shuffling step according to its position within the PU.
[0330] In particular, for TUs with a depth of 0, shuffling can be performed in various ways in addition to the method presented for the shuffling step above.
[0331] One method involves starting to scan the residual signal at the center of the PU block, scanning the surrounding residual signals in a circular motion toward the block's boundary, and then repositioning the scanned residual signals in a zig-zag scanning sequence starting from the PU's (0,0) position.
[0332] FIG. 18 is a diagram illustrating the residual signal scanning sequence and relocation sequence of a conversion unit (TU) with a depth of 0 within a prediction unit (PU) according to one embodiment of the present invention.
[0333] FIG. 18 (a) and (b) illustrate the scanning sequence for shuffling, and FIG. 18 (c) illustrates the rearrangement sequence for SDST.
[0334] For the residual signals within each shuffled TU, a DST-7 transformation is performed, and quantization and entropy coding can be performed. This shuffling method utilizes the residual signal distribution characteristics within the TU according to the PU partitioning mode, and the residual signal distribution can be optimized to increase the efficiency of the next step, the DST-7 transformation.
[0335] SDST can be performed for all TUs within the inter-predicted PU in the encoder according to the SDST process of FIG. 16 described above. Depending on the PU partitioning mode of the inter-predicted CU, TU partitioning can be performed from the PU up to a maximum depth of 2 in the form shown in FIG. 17. Shuffling can be performed on the residual signals within each TU using the residual signal distribution characteristics within the TU in FIG. 17. Afterward, a transformation using the DST-7 transformation kernel is performed, followed by quantization and entropy coding, etc.
[0336] When the decoder performs residual signal recovery of a TU within an inter-predicted PU, it can obtain the recovered residual signal by performing an inverse DST-7 transformation for each TU within the inter-predicted PU and inverse shuffling the recovered residual signal. This SDST method has the advantage that there are no additional flags or information that need to be signaled to the decoder because SDST is applied to the transformation method of all TUs within the inter-predicted PU. In other words, the SDST method can be performed without separate signaling for the SDST method.
[0337] Meanwhile, even if SDST is performed for all TUs within the inter-predicted PU, the encoder may determine some of the rearrangement methods of the residual signals described above with respect to the shuffling step as the optimal rearrangement method, and signal information about the determined rearrangement method to the decoder.
[0338] As another embodiment for performing SDST, at least one of two or more conversion methods (e.g., DCT-2 and SDST) may be selected and applied for the conversion of the corresponding PU. According to this method, the amount of computation in the encoder may increase compared to an embodiment in which SDST is performed for all TUs within the inter-predicted PU. However, compression efficiency may be improved because the more efficient conversion method between DCT-2 and SDST is selected.
[0339] FIG. 19 is a flowchart illustrating the DCT-2 or SDST selective encoding process through rate-distortion optimization (RDO) according to the present invention.
[0340] Referring to FIG. 19, the residual signal of the TU to be converted can be input (S1910). By comparing the cost of the TU obtained by performing DCT-2 on each TU within the PU predicted in inter-mode (S1920) with the cost of the TU obtained by performing shuffling (S1930) and DST-7 (S1940), the optimal conversion mode of the TU (e.g., DST-2 or SDST) in terms of rate distortion can be determined (S1950). Then, quantization (S1960) and entropy encoding, etc., can be performed on the converted TU according to the determined conversion mode.
[0341] Meanwhile, TU can select the optimal conversion mode between SDST and DCT-2 only if any one of the following conditions is satisfied.
[0342] i) Regardless of the PU partition mode, TU must be a CU or a quadtree partition or binary tree partition of CU.
[0343] ii) TU must be a PU depending on the PU partitioning mode, or a PU partitioned into a quadtree or binary tree.
[0344] iii) TU is not split from CU regardless of PU split mode.
[0345] Condition i) is a method of selecting DCT-2 or SDST as the transformation mode from the perspective of rate-distortion optimization for TUs obtained by splitting the CU into quadtrees or binary trees or splitting it into CU sizes, regardless of the PU splitting mode.
[0346] Condition ii) relates to an example in which SDST is performed on all TUs within an inter-predicted PU. That is, DCT-2 and SDST are performed on TUs obtained by splitting the PU into quadtrees or binary trees or splitting them into PU sizes according to the PU splitting mode, and the transformation mode of the TU is determined by considering the respective costs.
[0347] Condition iii) determines the transformation mode of the TU by performing DCT-2 and SDST without dividing the CU or TU in a CU unit having the same size as the TU, regardless of the PU division mode.
[0348] When comparing the rate-distortion cost (RD cost) for a depth 0 TU block of a specific PU split mode, the conversion mode of depth 0 TU can be selected by comparing the cost of the result of performing SDST on depth 0 TU and the cost of the result of performing DCT-2 on depth 0 TU.
[0349] FIG. 20 is a flowchart illustrating the process of selecting and decoding DCT-2 or SDST according to the present invention.
[0350] Referring to FIG. 20, the signaled SDST flag can be referenced for each TU (S2010). Here, the SDST flag may be a flag indicating whether to use SDST in the conversion mode.
[0351] When the SDST flag is true (S2020 - Yes), the transformation mode of the TU is determined to be SDST mode, and the DST-7 inverse transformation is performed on the residual signal within the TU (S2030). Then, the DST-7 inverse transformation is performed on the residual signal within the TU, and inverse shuffling is performed using the above-described mathematical formula 6 according to the position of the TU within the PU (S2040), thereby obtaining the finally restored residual signal (S2060).
[0352] Meanwhile, if the SDST flag is not true (S2520 - No), the conversion mode of the TU is determined to be DCT-2 mode, and the DCT-2 inverse conversion is performed on the residual signal within the TU (S2550) to obtain the restored residual signal (S2560).
[0353] When the SDST method is used, residual data can be rearranged. Here, residual data may refer to residual data corresponding to the inter-predicted PU. An integer transformation derived from DST-7 using the separable property can be used as the SDST method.
[0354] Meanwhile, sdst_flag may be signaled for the optional use of DCT-2 or DST-7. sdst_flag may be signaled on a TU basis. sdst_flag may indicate that it is required to identify whether SDST is performed.
[0355] FIG. 21 is a flowchart illustrating a decoding process using SDST according to the present invention.
[0356] Referring to FIG. 21, sdst_flag can be entropy decoded in TU units (S2110).
[0357] First, if the depth of the TU is 0 (S2120 - Yes), SDST is not used, and DCT-2 is used to restore the TU (S2170 and S2180). This is because SDST can be performed between the depth of the TU of 1 and the maximum TU depth value.
[0358] In addition, even if the depth of the TU is not 0 (S2120 - No), if the transformation mode of the TU is a transformation skip mode and / or if the value of the coded block flag (cbf) of the TU is 0 (S2130 - Yes), the TU can be restored without performing an inverse transformation (S2180).
[0359] Meanwhile, if the depth of the TU is not 0 (S2120 - No), the conversion mode of the TU is not conversion skip mode, and the cbf value of the TU is not 0 (S2130 - No), the sdst_flag value can be checked (S2140).
[0360] Here, if the value of sdst_flag is 1 (S2140 - Yes), a DST-7-based inverse transformation is performed (S2150), and the TU can be restored by performing inverse shuffling on the remaining data of the TU (S2160) (S2180). On the other hand, if the value of sdst_flag is 0 (S2140 - No), a DCT-2-based inverse transformation is performed (S2170) and the TU can be restored (S2180).
[0361] Here, the signal subject to shuffling or rearrangement may be at least one of the residual signal before inverse transformation, the residual signal before inverse quantization, the residual signal after inverse transformation, the residual signal after inverse quantization, the restored residual signal, and the restored block signal.
[0362] Meanwhile, although FIG. 21 describes that sdst_flag is signaled on a TU basis, sdst_flag may be selectively signaled based on at least one of the TU's transformation mode or the TU's cbf value. For example, if the TU's transformation mode is a transformation skip mode and / or the TU's cbf value is 0, sdst_flag may not be signaled. Additionally, sdst_flag may not be signaled even if the TU's depth is 0.
[0363] Meanwhile, although it was explained that sdst_flag is signaled in TU units, it may be signaled in a predetermined unit. For example, sdst_flag may be signaled in at least one of a video, sequence, picture, slice, tile, encoding tree unit, encoding unit, prediction unit, or transformation unit.
[0364] As in the example of the SDST flag of FIG. 20 and the sdst_flag of FIG. 21, selected conversion mode information can be entropy encoded / decoded at the TU unit through an n-bit flag or index (n is a positive integer greater than or equal to 1). The conversion mode information can indicate at least one of the following: whether the TU performed the conversion through DCT-2, the conversion through SDST, or the conversion through DST-7.
[0365] Entropy encoding / decoding of the corresponding transform mode information in bypass mode may be performed only when the TU within the inter-predicted PU is present. Additionally, entropy encoding / decoding of the transform mode information may be omitted and not signaled even if it is in transform skip mode, or at least one of RDPCM (Residual Differential PCM) mode and lossless mode.
[0366] In addition, even if the coded block flag of a block is 0, the entropy encoding / decoding of the conversion mode information may be omitted and not signaled. Since the inverse conversion process in the decoder is omitted when the coded block flag is 0, the block can be restored even if the conversion mode information does not exist in the decoder.
[0367] However, the conversion mode information is not limited to indicating the conversion mode through a flag, but may also be implemented in the form of predefined tables and indexes. Here, the predefined table may be one in which available conversion modes are defined for each index.
[0368] Meanwhile, although it has been described in FIGS. 19 to 21 that SDST or DCT-2 is selectively used, it is not limited thereto, and DCT-n or DST-n (where n is a positive integer) may be applied instead of DCT-2.
[0369] In addition, DCT-2 or SDST conversion can be performed separately for the horizontal and vertical directions. The same conversion mode may be used for the horizontal and vertical directions, or different conversion modes may be used for each.
[0370] Additionally, for the horizontal and vertical directions, respectively, conversion mode information regarding whether DCT-2, SDST, or DST-7 was used can be entropy encoded / decoded. The said conversion mode information can be signaled, for example, as an index, and the conversion kernel indicated by the same index can be the same for intra-frame predicted blocks and inter-frame predicted blocks.
[0371] In addition, conversion mode information can be entropy encoded / decoded in at least one of the CU, PU, TU, and block units.
[0372] Additionally, the transformation mode information can be signaled according to the luminance component or the chrominance component. In other words, the transformation mode information can be signaled according to the Y component, the Cb component, or the Cr component. For example, if transformation mode information regarding whether DCT-2 or SDST was performed is signaled for the Y component, the transformation mode information signaled in the Y component can be used as the transformation mode of the corresponding block without separate transformation mode information signaling in at least one of the Cb component and the Cr component.
[0373] Here, the transformation mode information can be entropy encoded / decoded using an arithmetic encoding method that utilizes a context model. If the transformation mode information is implemented in the form of predefined tables and indexes, entropy encoded / decoded can be performed using an arithmetic encoding method that utilizes a context model for all or some of the multiple bins.
[0374] Additionally, the conversion mode information may be optionally entropy encoded / decoded depending on the block size. For example, if the current block size is 64x64 or larger, the conversion mode information is not entropy encoded / decoded, and if it is 32x32 or smaller, the conversion mode information may be entropy encoded / decoded.
[0375] Additionally, if there are L non-zero transformation coefficients or quantized levels within the current block, one of the DCT-2, DST-7, or SDST methods may be performed without entropy encoding / decoding the transformation mode information. In this case, the transformation mode information may not be entropy encoded / decoded regardless of the position of the non-zero transformation coefficients or quantized levels within the block. Additionally, the transformation mode information may not be entropy encoded / decoded only if the non-zero transformation coefficients or quantized levels are located at the top-left position within the block. Here, L may be a positive integer including 0, for example, 1.
[0376] Additionally, if there are J or more non-zero transformation coefficients or quantized levels within the current block, the transformation mode information can be entropy encoded / decoded. In this case, J is a positive integer.
[0377] In addition, the binarization method of the conversion mode information may vary depending on the conversion mode of the collocated block, such as by restricting the use of some conversion modes or by representing the conversion mode of the collocated block with fewer bits.
[0378] The above-described SDST may be used restrictively based on at least one of the prediction mode, intra-frame prediction mode, inter-frame prediction mode, TU depth (depth), size, and shape of the current block.
[0379] For example, SDST can be used when the current block is encoded in inter mode.
[0380] A minimum / maximum depth at which SDST is allowed may be defined. In this case, SDST may be used when the depth of the current block is equal to or greater than the minimum depth, or SDST may be used when the depth of the current block is equal to or less than the maximum depth. Here, the minimum / maximum depth may be a fixed value, or it may be determined variably based on information indicating the minimum / maximum depth. Information indicating the minimum / maximum depth may be signaled from an encoder, or it may be derived from a decoder based on the attributes of the current / neighboring block (e.g., size, depth, and / or shape).
[0381] A minimum / maximum size for which SDST is allowed may be defined. Similarly, SDST may be used when the current block size is equal to or greater than the minimum size, or when the current block size is equal to or smaller than the maximum size. Here, the minimum / maximum size may be a fixed value or may be determined variably based on information indicating the minimum / maximum size. Information indicating the minimum / maximum size may be signaled from the encoder or may be derived from the decoder based on the attributes of the current / neighboring block (e.g., size, depth, and / or shape). For example, if the current block is 4x4, DCT-2 may be used as the conversion method, and the conversion mode information regarding whether DCT-2 or SDST was used may not be entropy encoded / decoded.
[0382] The block types for which SDST is allowed may be defined. In this case, SDST may be used if the current block type is the defined block type. Alternatively, block types for which SDST is not allowed may be defined. In this case, SDST may not be used if the current block type is the defined block type. The block types for which SDST is allowed or not allowed may be fixed, or information regarding this may be signaled from the encoder. Alternatively, they may be derived from the decoder based on the attributes of the current / neighboring blocks (e.g., size, depth, and / or shape). The block types for which SDST is allowed or not allowed may mean, for example, MxN blocks, M, N, and / or the ratio of M to N.
[0383] In addition, when the depth of the TU is 0, DCT-2 or DST-7 is used as the conversion method, and the conversion mode information regarding which conversion method was used can be entropy encoded / decoded. If DST-7 is used as the conversion method, the rearrangement process of the residual signal can be performed. In addition, when the depth of the TU is 1 or greater, DCT-2 or SDST is used as the conversion method, and the conversion mode information regarding which conversion method was used can be entropy encoded / decoded.
[0384] In addition, the conversion method can be selectively used depending on the partitioning form of CU and PU or the form of the current block.
[0385] According to one embodiment, DCT-2 is used when the split form of CU and PU or the current block form is 2Nx2N, and DCT-2 or SDST can be optionally used for the remaining split and block forms.
[0386] In addition, DCT-2 is used when the partition form of CU and PU or the current block form is 2NxN or Nx2N, and DCT-2 or SDST can be optionally used for the remaining partition forms and block forms.
[0387] In addition, DCT-2 is used when the partition form of CU and PU or the current block form is nRx2N or nLx2N or 2NxnU or 2NxnD, and DCT-2 or SDST can be optionally used for the remaining partition forms and block forms.
[0388] Meanwhile, when SDST or DST-7 is performed on a block unit divided from the current block, scanning and inverse scanning of conversion coefficients (quantized levels) may be performed on a block unit divided from the current block. Additionally, when SDST or DST-7 is performed on a block unit divided from the current block, scanning and inverse scanning of conversion coefficients (quantized levels) may be performed on an undivided current block unit.
[0389] In addition, the conversion / inverse conversion using the SDST or DST-7 can be performed according to at least one of the in-frame prediction mode (direction) of the current block, the size of the current block, and the component of the current block (whether it is a luminance component or a chrominance component).
[0390] In addition, DST-1 may be used instead of DST-7 when performing conversion / inverse conversion using the above SDST or DST-7. Furthermore, DCT-4 may be used instead of DST-7 when performing conversion / inverse conversion using the above SDST or DST-7.
[0391] In addition, when performing conversion / inverse conversion using the above DCT-2, the rearrangement method used for rearranging the residual signal of SDST or DST-7 can be applied. That is, even when using DCT-2, rearrangement of the residual signal or rotation of the residual signal using a predetermined angle can be performed.
[0392] Hereinafter, various variations and embodiments regarding the shuffling method and the signaling method are described.
[0393] The SDST of the present invention aims to improve image compression efficiency by changing the conversion, shuffling, rearrangement, and / or flipping methods. Performing DST-7 through shuffling of residual signals can exhibit high compression efficiency because it effectively reflects the residual signal distribution characteristics within the PU.
[0394] In the above description regarding the shuffling step, we examined the residual signal rearrangement method. Below, we will examine other implementation methods in addition to the shuffling method for residual signal rearrangement described above.
[0395] The rearrangement method described below can be applied to at least one of the embodiments of the SDST method described above.
[0396] In order to minimize hardware complexity in implementing residual signal rearrangement, the residual signal rearrangement process can be implemented using horizontal flipping and vertical flipping methods. The residual signal rearrangement method can be implemented through flipping as shown in (1) to (4) below. The rearrangement described below may refer to flipping.
[0397] (1) r'(x,y)=r(x,y) ; no flipping
[0398] (2) r'(x,y)=r(w-1-x,y) ; horizontal flipping
[0399] (3) r'(x,y)=r(x,h-1-y) ; vertical flipping
[0400] (4) r'(x,y)=r(w-1-x,h-1-y) ; horizontal and vertical flipping
[0401] r'(x,y) is the residual signal after rearrangement, and r(x,y) is the residual signal before rearrangement. w and h represent the width and height of the block, respectively, and x and y represent the positions of the residual signals within the block. The inverse rearrangement method of a rearrangement method using flipping can be performed through the same process as the rearrangement method. That is, the residual signals rearranged using horizontal flipping can be restored to their original residual signal array by performing horizontal flipping once again. The rearrangement method performed in the encoder and the inverse rearrangement method performed in the decoder may be the same flipping method.
[0402] For example, if horizontal flipping is performed on the remaining block after horizontal flipping, the remaining block before flipping can be obtained. This can be expressed mathematically as follows.
[0403] r'(w-1-x, y) = r(w-1-(w-1-x), y) = r(x,y).
[0404] For example, if vertical flipping is performed on the remaining block after vertical flipping, the remaining block before flipping can be obtained, which can be expressed as a formula as follows.
[0405] r'(x, h-1-y)=r(x,h-1-(h-1-y)) = r(x,y).
[0406] For example, if horizontal and vertical flipping is performed on the remaining blocks after horizontal and vertical flipping, the remaining blocks before flipping can be obtained, which can be expressed by the following formula.
[0407] r'(w-1-x, h-1-y) = r(w-1-(w-1-x), h-1-(h-1-y)) = r(x,y).
[0408] The above-described flipping-based residual signal shuffling / rearrangement method can be used without dividing the current block. That is, while the above-described SDST method describes dividing the current block (TU, etc.) into sub-blocks and using DST-7 for each sub-block, when using the flipping-based residual signal shuffling / rearrangement method, the current block can be flipped on the entire current block or a part thereof without dividing it into sub-blocks, and then the DST-7 transformation can be performed. Additionally, when using the flipping-based residual signal shuffling / rearrangement method, the current block can be flipped on the entire current block or a part thereof after performing the inverse DST-7 transformation without dividing it into sub-blocks.
[0409] A maximum size (MxN) and / or minimum size (OxP) of a block capable of performing flipping-based residual signal shuffling / rearrangement may be defined. Here, the size may include at least one of a width (M or O) which is the horizontal size and a height (N or P) which is the vertical size. M, N, O, and P may be positive integers. The maximum size and / or minimum size of the block may be a value predefined in the encoder / decoder, or it may be information signaled from the encoder to the decoder.
[0410] For example, if the current block size is smaller than the minimum size for performing the flipping method, only the DCT-2 transformation can be performed without performing the flipping and DST-7 transformations. In this case, the SDST flag, which is transformation mode information indicating whether to use flipping and DST-7 as the transformation mode, may not be signaled.
[0411] For example, if only the width of the block is smaller than the minimum width at which the flipping method can be performed and the height of the block is greater than the minimum height at which the flipping method can be performed, the horizontal 1D transformation can be performed using only DCT-2, and the vertical 1D transformation can be performed by using DST-7 after the vertical flipping, or by using DST-7 without flipping. In this case, the SDST flag, which is transformation mode information indicating whether to use flipping as the transformation mode, can be signaled only for the vertical 1D transformation.
[0412] For example, if only the height of the block is smaller than the minimum width at which the flipping method can be performed, and the width of the block is larger than the minimum width at which the flipping method can be performed, the horizontal 1D transformation can be performed using DST-7 after the horizontal flipping, or the horizontal transformation can be performed using DST-7 without flipping, and the vertical 1D transformation can be performed using only DCT-2. In this case, the SDST flag, which is transformation mode information indicating whether to use flipping as the transformation mode, can be signaled only for the horizontal 1D transformation.
[0413] For example, if the current block size is larger than the maximum size for which the flipping method can be performed, only the DCT-2 transformation can be used without using flipping and the DST-7 transformation. In this case, the SDST flag, which is transformation mode information indicating whether to use flipping and the DST-7 transformation as the transformation mode, may not be signaled.
[0414] For example, if the current block size is larger than the maximum size for which the flipping method can be performed, only the DCT-2 transformation or the DST-7 transformation can be used.
[0415] For example, if the maximum size for which the flipping method can be performed is 32x32 and the minimum size is 4x4, then for a 64x64 block, flipping and DST-7 conversion are not used, and only DCT-2 conversion may be used. In this case, for a 64x64 block, the SDST flag, which is conversion mode information indicating whether to use flipping and DST-7 in the conversion mode, may not be signaled. Additionally, for blocks ranging from 4x4 to 32x32, the SDST flag, which is conversion mode information indicating whether to use flipping and DST-7 in the conversion mode, may be signaled. In this case, since DST-7 conversion is not used for a 64x64 block, memory space for storing DST-7 conversion used for a 64x64 block can be saved.
[0416] For example, if the maximum size for which the flipping method can be performed is 32x32 and the minimum size is 4x4, then for a block of size 64x64, the flipping method is not used, and DCT-2 or DST-7 transformations may be used.
[0417] For example, a square block of size MxN can be partitioned into four blocks using a quadtree, and a shuffling / rearrangement method using flipping can be performed on each sub-block, followed by a DST-7 transformation. In this case, the flipping method can be explicitly signaled for each sub-block. The flipping method can be signaled by a 2-bit fixed-length code or by a truncated unary code. Additionally, a binarization method based on the probability of the flipping method occurring for each partitioned block may be used. Here, M and N can be positive integers, for example, 64x64.
[0418] For example, an MxN square block can be partitioned into four quadtrees, and a shuffling / rearrangement method using flipping can be performed on each subblock, followed by a DST-7 transformation. The flipping method for each subblock may be implicitly determined. For instance, horizontal and vertical flipping may be determined for the first (top-left) subblock, vertical flipping for the second (top-right) subblock, horizontal flipping for the third (bottom-left) subblock, and no flipping for the fourth (bottom-right) subblock. When the flipping method is implicitly determined in this way, signaling for the flipping method is not required. Here, M and N can be positive integers, for example, 64x64.
[0419] For example, a rectangular block of size 2MxN can be binary tree partitioned into two MxN square blocks, and a shuffling / rearrangement method using flipping can be performed on each partitioned block, followed by a DST-7 transformation. In this case, the flipping method can be explicitly signaled for each sub-block. The flipping method can be signaled by a 2-bit fixed-length code or by a truncated unary code. Additionally, a binarization method based on the probability of the flipping method occurring for each sub-block may be used. Here, M and N can be positive integers, for example, 8x8.
[0420] For example, a 2MxN rectangular block can be partitioned into two MxN square blocks using a binary tree, and a shuffling / rearrangement method using flipping can be performed on each subblock, followed by a DST-7 transformation. The flipping method for each subblock may be implicitly determined. For the first (left) subblock, horizontal flipping may be determined, and for the second (right) subblock, no flipping may be performed. When the flipping method is implicitly determined in this way, signaling for the flipping method is not required. Here, M and N can be positive integers, for example, 4x4.
[0421] For example, a rectangular block of size Mx2N can be partitioned into two MxN square blocks using a binary tree, and a shuffling / rearrangement method using flipping can be performed on each subblock, followed by a DST-7 transformation. The flipping method for each subblock may be implicitly determined. For the first (top) subblock, vertical flipping may be determined, and for the second (bottom) subblock, no flipping may be performed. When the flipping method is implicitly determined in this way, signaling for the flipping method is not required. Here, M and N can be positive integers, for example, 4x4.
[0422] At least one of two methods may be applied: performing DCT-2 transformation / inverse transformation on an MxN block, or dividing the block into a quadtree or binary tree to create subblocks, and then performing DST-7 transformation / inverse transformation on each subblock after flipping. In this case, the flipping method may be performed differently depending on the relative position of each subblock in its parent block, and this may be implicitly determined. Here, M and N are positive integers, for example, M and N may be 64. That is, the MxN block may be a block with a relatively large size.
[0423] - In the case of the top-left subblock, the flipping for that subblock can be determined as horizontal and vertical flipping.
[0424] - In the case of the upper right subblock, the flipping for that subblock can be determined as vertical flipping.
[0425] - In the case of a bottom-left subblock, the flipping for that subblock can be determined as a horizontal flipping.
[0426] - In the case of a bottom-right subblock, it may be decided not to perform flipping for that subblock.
[0427] Using the transformation mode information, the information on using the flipping-based residual signal shuffling / rearrangement method (sdst_flag or sdst flag) can be entropy encoded / decoded. That is, through signaling of the transformation mode information, the same method performed in the encoder can be performed in the decoder. For example, if the flag bit indicating the transformation mode information has a first value, the flipping-based residual signal shuffling / rearrangement method and DST-7 can be used as the transformation / inverse transformation method, and if the flag bit has a second value, a different transformation / inverse transformation method can be used. In this case, the transformation mode information can be entropy encoded / decoded block by block. Here, the different transformation / inverse transformation method may be the DCT-2 transformation / inverse transformation method. In addition, if it is a transform skip mode, or any one of RDPCM (Residual Differential PCM) mode and lossless mode, the entropy encoding / decoding of the transform mode information may be omitted and not signaled.
[0428] The above transformation mode information can be entropy encoded / decoded using at least one of the depth of the current block, the size of the current block, the shape of the current block, the transformation mode information of surrounding blocks, the encoding block flag of the current block, and whether the transformation omission mode of the current block is used. For example, if the encoding block flag of the current block is 0, the entropy encoding / decoding of the transformation mode information may be omitted and not signaled. Additionally, the above transformation mode information can be predictively encoded / decoded from the transformation mode information of the blocks restored around the current block during entropy encoding / decoding. Additionally, the above transformation mode information can be signaled based on at least one of the encoding parameters of the current block and surrounding blocks.
[0429] Additionally, using the flipping method information, at least one of the four flipping methods (no flipping, horizontal flipping, vertical flipping, horizontal and vertical flipping) can be entropy encoded / decoded in the form of a flag or an index (flipping_idx). That is, by signaling the flipping method information, the same flipping method performed in the encoder can be performed in the decoder. The conversion mode information may include the flipping method information.
[0430] In addition, if the transform skip mode, RDPCM (Residual Differential PCM) mode, or lossless mode is used, the entropy encoding / decoding of the flipping method information may be omitted and not signaled. The flipping method information may be entropy encoded / decoded using at least one of the following: the depth of the current block, the size of the current block, the shape of the current block, the flipping method information of surrounding blocks, the encoding block flag of the current block, and whether the transform skip mode of the current block is used. For example, if the encoding block flag of the current block is 0, the entropy encoding / decoding of the flipping method information may be omitted and not signaled. Additionally, the flipping method information may be predictively encoded / decoded from the flipping method information of the blocks restored around the current block during entropy encoding / decoding. Additionally, the flipping method information may be signaled based on at least one of the encoding parameters of the current block and surrounding blocks.
[0431] In addition, the residual signal rearrangement method is not limited to the residual signal rearrangement previously described above, and shuffling may be implemented by rotating the residual signal within the block by a predetermined angle. Here, the predetermined angle may mean 0 degrees, 90 degrees, 180 degrees, -90 degrees, -180 degrees, 270 degrees, -270 degrees, 45 degrees, -45 degrees, 135 degrees, -135 degrees, etc. In this case, information regarding the angle may be entropy encoded / decoded in the form of a flag or an index, and may be performed similarly to the signaling method for the conversion mode information.
[0432] Additionally, the above angle information can be predicted during entropy encoding / decoding from the angle information of the restored blocks surrounding the current block. When rearranging using angle information, SDST or DST-7 may be performed on the current block after partitioning, but SDST or DST-7 may also be performed on the current block unit without partitioning the current block.
[0433] A predetermined angle may be determined differently depending on the position of the sub-block. A method of rearranging via rotation may also be used restrictively only for the sub-block at a specific position (e.g., the first sub-block) among the sub-blocks. Additionally, rearrangement using a predetermined angle may be applied to the entire current block. In this case, the current block subject to rearrangement may be at least one of the residual block before inverse transformation, the residual block before inverse quantization, the residual block after inverse transformation, the residual block after inverse quantization, the restored residual block, and the restored block.
[0434] Meanwhile, the coefficients of the transformation matrix for the transformation may be rearranged or rotated to produce the same effect as residual signal rearrangement or rotation, and the transformation may be performed by applying this to the pre-arranged residual signal. That is, by performing the transformation using rearrangement of the transformation matrix instead of residual signal rearrangement, the same effect as the method of performing residual signal rearrangement and transformation can be obtained. In this case, the rearrangement of the coefficients of the transformation matrix can be performed in the same manner as the residual signal rearrangement methods described above, and the signaling method for the information required for this can also be performed in the same manner as the signaling method for the information required for the residual signal rearrangement methods described above.
[0435] Meanwhile, the encoder may determine some of the residual signal rearrangement methods mentioned in the above description regarding the shuffling step as the optimal rearrangement method, and signal information about the determined rearrangement method (flipping method information) to the decoder. For example, if four rearrangement methods are used, the encoder may signal 2 bits of information about the residual signal rearrangement method to the decoder.
[0436] In addition, when the probability of occurrence of each of the rearrangement methods used is different, the rearrangement method with a high probability of occurrence can be encoded using fewer bits, and the rearrangement method with a low probability of occurrence can be encoded using relatively more bits. For example, four rearrangement methods can be signaled as truncated unary codes (e.g., (0, 10, 110, 111) or (1, 01, 001, 000)) in order of highest probability of occurrence.
[0437] In addition, since the probability of occurrence of a rearrangement method may vary depending on encoding parameters such as the current CU's prediction mode, the PU's intra-frame prediction mode (direction), and the motion vector of surrounding blocks, the encoding method for information about the rearrangement method (flipping method information) may be used differently depending on the encoding parameters. For example, since the probability of occurrence of a rearrangement method may differ depending on the intra-frame prediction mode, for each intra-mode, a small number of bits may be allocated to rearrangement methods with a high probability of occurrence and a large number of bits may be allocated to rearrangement methods with a low probability of occurrence, or in some cases, rearrangement methods with a very low probability of occurrence may not be used and no bits may be allocated.
[0438] A rearrangement set including at least one of residual signal rearrangement methods can be configured according to at least one of the prediction mode of the current block (inter mode or intra mode), intra-frame prediction mode (including directional mode and non-directional mode), inter-frame prediction mode, block size, block shape (square or non-square), whether there is a luminance / chrominance signal, and transformation mode information. The rearrangement may refer to flipping. Additionally, a rearrangement set including at least one of residual signal rearrangement methods can be configured based on at least one of the encoding parameters of the current block and surrounding blocks.
[0439] Additionally, at least one of the following rearrangement sets may be selected based on at least one of the prediction mode of the current block, intra-frame prediction mode, inter-frame prediction mode, block size, block shape, presence of luminance / chrominance signal, transformation mode information, etc. Additionally, at least one of the rearrangement sets may be selected based on at least one of the encoding parameters of the current block and surrounding blocks.
[0440] A rearrangement set may include at least one of 'no flipping', 'horizontal flipping', 'vertical flipping', and 'horizontal and vertical flipping'. Examples of rearrangement sets are shown below.
[0441] 1. Do not perform flipping
[0442] 2. Horizontal flipping
[0443] 3. Vertical flipping
[0444] 4. Horizontal and vertical flipping
[0445] 5. Do not perform flipping, horizontal flipping
[0446] 6. Do not perform flipping, vertical flipping
[0447] 7. Do not perform flipping, horizontal and vertical flipping
[0448] 8. Horizontal flipping, vertical flipping
[0449] 9. Horizontal flipping, horizontal and vertical flipping
[0450] 10. Vertical flipping, horizontal and vertical flipping
[0451] 11. Do not perform flipping, horizontal flipping, vertical flipping
[0452] 12. Do not perform flipping, horizontal flipping, horizontal and vertical flipping
[0453] 13. Do not perform flipping, portrait flipping, landscape and portrait flipping
[0454] 14. Horizontal flipping, vertical flipping, horizontal and vertical flipping
[0455] 15. Do not perform flipping, horizontal flipping, vertical flipping, horizontal and vertical flipping
[0456] Based on the above rearrangement set, at least one of the residual signal rearrangement methods can be used for the rearrangement of the current block.
[0457] In addition, at least one of the residual signal rearrangement methods in the rearrangement set may be selected according to at least one of the prediction mode of the current block, intra-frame prediction mode, inter-frame prediction mode, block size, block shape, whether there is a luminance / chrominance signal, transformation mode information, flipping method information, etc. In addition, at least one of the residual signal rearrangement methods in the rearrangement set may be selected based on at least one of the encoding parameters of the current block and surrounding blocks.
[0458] Depending on the prediction mode of the current block, at least one rearrangement set may be configured. For example, if the prediction mode of the current block is intra-frame prediction, multiple rearrangement sets may be configured, and if the prediction mode of the current block is inter-frame prediction, one rearrangement set may be configured.
[0459] At least one rearrangement set may be configured depending on the in-frame prediction mode of the current block. For example, if the in-frame prediction mode of the current block is a non-directional mode, one rearrangement set may be configured, and if the in-frame prediction mode of the current block is a directional mode, multiple rearrangement sets may be configured.
[0460] Depending on the size of the current block, at least one rearrangement set may be configured. For example, if the size of the current block is greater than 16x16, one rearrangement set may be configured, and if the size of the current block is less than or equal to 16x16, multiple rearrangement sets may be configured.
[0461] Depending on the shape of the current block, at least one rearrangement set may be formed. For example, if the current block is square, one rearrangement set may be formed, and if the current block is non-square, multiple rearrangement sets may be formed.
[0462] At least one rearrangement set may be configured depending on whether the current block is a luminance / chrominance signal. For example, if the current block is a chrominance signal, one rearrangement set may be configured, and if the current block is a luminance signal, multiple rearrangement sets may be configured.
[0463] Additionally, an index for a residual signal rearrangement method based on the above rearrangement set can be entropy encoded / decoded. In this case, the index can be entropy encoded / decoded using a variable-length code or a fixed-length code.
[0464] In addition, binarization and debinarization of the index for the residual signal rearrangement method can be performed based on the above rearrangement set. At this time, the index can be binarized and debinarized into a variable-length code or a fixed-length code.
[0465] In addition, the above rearrangement set can be in the form of a table in the encoder and decoder and can be calculated through an expression.
[0466] Additionally, the above rearrangement set may be configured to have symmetry. For example, a table for the above rearrangement set may be configured to have symmetry. In this case, the table may be configured to have symmetry with respect to the in-frame prediction mode.
[0467] Additionally, the above rearrangement set may be configured according to at least one of whether the in-frame prediction mode is included in a specific range and whether the in-frame prediction mode is even or odd.
[0468] The tables below show examples of methods for encoding / decoding residual signal rearrangement methods based on the prediction mode of the current block and the intra-frame prediction mode (direction).
[0469] In addition, the use of at least one of the residual signal rearrangement methods in the tables below can be indicated using flipping method information.
[0470] Prediction mode In-screen prediction direction (In-screen prediction mode) Residual signal rearrangement method (1) (2) (3) (4) In-screen Horizontal or near-horizontal mode 0 - 1 - In-screen Vertical or near-vertical mode 0 1 - - In-screen A mode at a 45-degree diagonal or close to a 45-degree diagonal. * - - - In-screen even number 0 10 11 - In-screen odd number 0 - 10 11 In-screen The otherwise 0 110 10 111 Between screens N / A 00 01 10 11
[0471] (1) to (4) in the column of residual signal rearrangement methods in Table 1 may specify residual signal rearrangement methods, such as an index for a scanning / rearrangement order for residual signal rearrangement previously described, an index for a predetermined angle value, or an index for a predetermined flipping method. The * symbol in the column of residual signal rearrangement methods in Table 1 indicates that the corresponding rearrangement method is used implicitly without signaling, and the - symbol indicates that the corresponding rearrangement method is not used in that case. The meaning of implicitly using the corresponding rearrangement method may mean that the corresponding rearrangement method is used by using the conversion mode information (sdst_flag or sdst flag) without entropy encoding / decoding of the index for residual signal rearrangement methods.
[0472] The above residual signal rearrangement methods (1) to (4) may each mean (1) no flipping, (2) horizontal flipping, (3) vertical flipping, and (4) horizontal and vertical flipping. Additionally, the above 0, 1, 10, 11, 110, 111, etc. may be the result of binarization / debinarization used to entropy-encode / decode the above residual signal rearrangement method. A fixed-length code, a truncated unary code, or a unary code may be used as the above binarization / debinarization method.
[0473] As shown in Table 1 above, when the current block corresponds to at least one of each prediction mode and each in-frame prediction mode (direction), at least one rearrangement method may be used in the encoder and decoder. Here, the 45-degree diagonal direction may mean a direction toward the top-left position from the current block or a direction toward the current block from the top-left position of the current block.
[0474] Prediction mode In-screen prediction direction (In-screen prediction mode) Residual signal rearrangement method (1) (2) (3) (4) In-screen Horizontal or near-horizontal mode 0 - 1 - In-screen Vertical or near-vertical mode 0 1 - - In-screen A mode at a 45-degree diagonal or close to a 45-degree diagonal. * - - - In-screen The otherwise 00 01 10 11 Between screens N / A 00 01 10 11
[0475] As another example, as shown in Table 2 above, when the current block corresponds to at least one of each prediction mode and each in-frame prediction mode (direction), at least one rearrangement method can be used in the encoder and decoder.
[0476] Prediction mode In-screen prediction direction (In-screen prediction mode) Residual signal rearrangement method (1) (2) (3) (4) In-screen Horizontal or near-horizontal mode 0 - 1 - In-screen Vertical or near-vertical mode 0 1 - - In-screen A mode at a 45-degree diagonal or close to a 45-degree diagonal. * - - - In-screen The otherwise 0 110 10 111 Between screens N / A 0 110 10 111
[0477] As another example, as shown in Table 3 above, when the current block corresponds to at least one of each prediction mode and each in-frame prediction mode (direction), at least one rearrangement method can be used in the encoder and decoder.
[0478] Prediction mode In-screen prediction direction (In-screen prediction mode) Residual signal rearrangement method (1) (2) (3) (4) In-screen even number 0 10 11 - In-screen odd number 0 - 10 11 Between screens N / A 00 01 10 11
[0479] As another example, as shown in Table 4 above, when the current block corresponds to at least one of each prediction mode and each intra-frame prediction mode (direction), at least one rearrangement method may be used in the encoder and decoder. For example, if the current block is in intra mode and the intra-frame prediction direction is even, at least one of the methods of not performing flipping, horizontal direction flipping, and vertical direction flipping may be used as a residual signal rearrangement method. Additionally, if the current block is in intra mode and the intra-frame prediction direction is odd, at least one of the methods of not performing flipping, vertical direction flipping, and horizontal direction and vertical direction flipping may be used as a residual signal rearrangement method.
[0480] Prediction mode In-screen prediction direction (In-screen prediction mode) Residual signal rearrangement method (1) (2) (3) (4) In-screen even number 0 10 11 - In-screen odd number 0 - 10 11 In-screen Non-directional mode (DC mode or Planar mode) 0 110 10 111 Between screens N / A 00 01 10 11
[0481] As another example, as shown in Table 5 above, when the current block corresponds to at least one of each prediction mode and each in-frame prediction mode (direction), at least one rearrangement method can be used in the encoder and decoder.
[0482] Prediction mode In-screen prediction direction (In-screen prediction mode) Residual signal rearrangement method (1) (2) (3) (4) In-screen Even mode that is not non-directional mode 0 10 11 - In-screen Odd mode, not non-directional mode 0 - 10 11 In-screen Non-directional mode (DC mode or Planar mode) 00 01 10 11 Between screens N / A 00 01 10 11
[0483] As another example, as shown in Table 6 above, when the current block corresponds to at least one of each prediction mode and each in-frame prediction mode (direction), at least one rearrangement method can be used in the encoder and decoder.
[0484] Prediction mode In-screen prediction direction (In-screen prediction mode) Residual signal rearrangement method (1) (2) (3) (4) In-screen Horizontal or near-horizontal mode 0 11 10 - In-screen Vertical or near-vertical mode 0 10 11 - In-screen A mode at a 45-degree diagonal or close to a 45-degree diagonal. * - - - In-screen The otherwise 0 110 10 111 Between screens N / A 0 110 10 111
[0485] As another example, as shown in Table 7 above, when the current block corresponds to at least one of each prediction mode and each in-frame prediction mode (direction), at least one rearrangement method can be used in the encoder and decoder.
[0486] Prediction mode In-screen prediction direction (In-screen prediction mode) Residual signal rearrangement method (1) (2) (3) (4) In-screen Horizontal or near-horizontal mode 0 - 10 11 In-screen Vertical or near-vertical mode 0 10 - 11 In-screen A mode at a 45-degree diagonal or close to a 45-degree diagonal. * - - - In-screen The otherwise 0 110 10 111 Between screens N / A 0 110 10 111
[0487] As another example, as shown in Table 8 above, when the current block corresponds to at least one of each prediction mode and each in-frame prediction mode (direction), at least one rearrangement method can be used in the encoder and decoder.
[0488] Prediction mode Residual signal rearrangement method (1) (2) (3) (4) In-screen 0 110 10 111 Between screens 00 01 10 11
[0489] As another example, as shown in Table 9 above, when the current block corresponds to at least one of each prediction mode and each in-frame prediction mode (direction), at least one rearrangement method can be used in the encoder and decoder.
[0490] Prediction mode In-screen prediction direction (In-screen prediction mode) Residual signal rearrangement method (1) (2) (3) (4) In-screen Horizontal or near-horizontal mode 0 - 10 11 In-screen Vertical or near-vertical mode 0 10 - 11 In-screen A mode at a 45-degree diagonal or close to a 45-degree diagonal. * - - - In-screen A mode at a 135-degree diagonal or close to a 135-degree diagonal. 0 10 11 - In-screen -45 degree diagonal direction or a mode close to the -45 degree diagonal direction 0 10 11 - In-screen The otherwise 00 01 10 11 Between screens N / A 0 110 10 111
[0491] As another example, as shown in Table 10 above, if the current block corresponds to at least one of each prediction mode and each intra-frame prediction mode (direction), at least one rearrangement method may be used in the encoder and decoder. Here, the 135-degree diagonal direction may mean a direction toward the top-right position from the current block or a direction toward the current block from the top-right position of the current block. For example, the value for the 135-degree diagonal direction mode may be 6. Here, the -45-degree diagonal direction may mean a direction toward the bottom-right position from the current block or a direction toward the current block from the bottom-right position of the current block. For example, the value for the -45-degree diagonal direction mode may be 2.
[0492] Prediction mode In-screen prediction direction (In-screen prediction mode) Residual signal rearrangement method (1) (2) (3) (4) In-screen It is a horizontal or near-horizontal mode and an odd mode 0 - 1 - In-screen It is a horizontal or near-horizontal mode and an even mode 0 1 - - In-screen It is a vertical or near-vertical mode and an odd mode 0 1 - - In-screen It is a vertical or near-vertical mode and an even mode 0 - 1 - In-screen It is a mode that is diagonally 45 degrees or close to a diagonal 45 degrees, and is an odd mode * - - - In-screen It is a mode that is diagonally 45 degrees or close to a diagonal 45 degrees, and is an even mode - - - * In-screen The otherwise 0 110 10 111 Between screens N / A 0 110 10 111
[0493] As another example, as shown in Table 11 above, when the current block corresponds to at least one of each prediction mode and each in-frame prediction mode (direction), at least one rearrangement method can be used in the encoder and decoder.
[0494] Prediction mode In-screen prediction direction (In-screen prediction mode) Residual signal rearrangement method (1) (2) (3) (4) In-screen It is a horizontal or near-horizontal mode and an odd mode 0 - 1 - In-screen It is a horizontal or near-horizontal mode and an even mode 0 - - 1 In-screen It is a vertical or near-vertical mode and an odd mode 0 1 - - In-screen It is a vertical or near-vertical mode and an even mode 0 - - 1 In-screen A mode at a 45-degree diagonal or close to a 45-degree diagonal. * - - - In-screen The otherwise 0 10 11 - Between screens N / A 0 110 10 111
[0495] As another example, as shown in Table 12 above, when the current block corresponds to at least one of each prediction mode and each in-frame prediction mode (direction), at least one rearrangement method can be used in the encoder and decoder.
[0496] Prediction mode Residual signal rearrangement method (1) (2) (3) (4) In-screen and inter-screen 0 10 110 1110
[0497] As another example, as shown in Table 13 above, when the current block corresponds to at least one of each prediction mode and each intra-frame prediction mode (direction), at least one reordering method may be used in the encoder and decoder. Here, the residual signal reordering method may represent a type of transformation. For example, if the residual signal reordering method is (1), both horizontal transformation and vertical transformation may represent the first transformation kernel. As another example, if the residual signal reordering method is (2), horizontal transformation and vertical transformation may represent the second transformation kernel and the first transformation kernel, respectively. As another example, if the residual signal reordering method is (3), horizontal transformation and vertical transformation may represent the first transformation kernel and the second transformation kernel, respectively. As another example, if the residual signal reordering method is (4), horizontal transformation and vertical transformation may represent the second transformation kernel and the second transformation kernel, respectively. For example, the first transformation kernel may be DST-7 and the second transformation kernel may be DCT-8.
[0498] When the in-frame prediction mode is Planar mode or DC mode, information on four rearrangement methods (flipping method information) can be entropy encoded / decoded using a truncated unary code based on the frequency of occurrence.
[0499] If the predicted direction within the screen is a horizontal direction or a mode close to the horizontal direction, the probability of rearrangement method (1) and / or rearrangement method (3) may be high. In such cases, information about the rearrangement method can be entropy encoded / decoded using 1 bit for each of the two rearrangement methods. Here, the meaning of a mode close to the horizontal direction may be that the value of a specific mode falls between the value for the horizontal direction mode - K and the value for the horizontal direction mode + K. Here, K may be an integer. For example, if the value for the horizontal direction mode is 18, K is 4, and the specific mode is 20, the specific mode can be said to be a mode close to the horizontal direction. For example, if the value for the horizontal direction mode is 18, K is 4, and the specific mode is 26, the specific mode cannot be said to be a mode close to the horizontal direction.
[0500] If the predicted direction within the screen is a vertical direction or a mode close to the vertical direction, the probability of rearrangement method (1) and / or rearrangement method (2) may be high. In such cases, information about the rearrangement method can be entropy encoded / decoded using 1 bit for each of the two methods. Here, the meaning of a mode close to the vertical direction may be that the value of a specific mode falls between the value for the vertical direction mode - K and the value for the vertical direction mode + K. Here, K may be an integer. For example, if the value for the vertical direction mode is 50, K is 2, and the specific mode is 51, the specific mode can be said to be a mode close to the vertical direction. For example, if the value for the vertical direction mode is 50, K is 8, and the specific mode is 20, the specific mode cannot be said to be a mode close to the vertical direction.
[0501] When the predicted direction within the screen is a 45-degree diagonal direction or a mode close to a 45-degree diagonal direction, the probabilities of the remaining rearrangement methods (2), (3), and (4) may be very low compared to the probability of rearrangement method (1). In such cases, only the above-mentioned method is applied, and the above-mentioned method may be used implicitly without signaling information about the rearrangement method. Here, the meaning of a mode close to a 45-degree diagonal direction may mean that the value of a specific mode is included between the value for a 45-degree diagonal direction mode - K and the value for a 45-degree diagonal direction mode + K. Here, K may be an integer. For example, if the value for a 45-degree diagonal direction mode is 34, K is 2, and the specific mode is 36, the specific mode may be said to be a mode close to a 45-degree diagonal direction. For example, if the value for the 45-degree diagonal mode is 34, K is 8, and the specific mode is 10, the specific mode cannot be said to be a mode close to the 45-degree diagonal direction.
[0502] When the prediction direction within the screen is even, information about the rearrangement method can be entropy encoded / decoded with a truncated unary code or unary code only for rearrangement methods (1), (2) and (3).
[0503] When the number of predicted directions within the screen is odd, information about the rearrangement method can be entropy encoded / decoded with a truncated unary code or unary code only for rearrangement methods (1), (3) and (4).
[0504] For other predicted directions within the screen, the probability of occurrence of the rearrangement method (4) may be low, so information about the rearrangement method can be entropy encoded / decoded with a single code or a single code that is truncated only for the rearrangement methods (1), (2) and (3).
[0505] In the case of inter-frame prediction, the probability of occurrence of rearrangement methods (1) to (4) can be viewed equally, and information about the rearrangement method can be entropy encoded / decoded with a 2-bit fixed-length code.
[0506] Arithmetic coding / decoding can be used for the above code. In addition, entropy coding / decoding can be performed in bypass mode without using arithmetic coding that uses a context model for the above code.
[0507] For a region within a picture, a CTU, the entire picture, or the current block within a picture group, one of two methods can be selected to perform conversion / inversion to DST-7 or conversion / inversion to DCT-2 without flipping. In this case, 1-bit flag information (conversion mode information) indicating whether to use DST-7 or DCT-2 on a current block basis can be entropy encoded / decoded. This method may be used when the energy of the residual signal increases with distance from the reference sample, or to reduce computational complexity during encoding and decoding. Information regarding the region where this method is used may be signaled in CTU units, slice units, PPS units, SPS units, or other units indicating specific regions, and a 1-bit flag may be signaled in an on / off format.
[0508] For a region within a picture, a CTU, the entire picture, or the current block within a picture group, a transformation / inverse transformation can be performed by selecting one of three methods: DCT-2 transformation / inverse transformation, DST-7 transformation / inverse transformation without flipping, or DST-7 transformation / inverse transformation after performing vertical flipping. Information regarding which of the three methods to select may be implicitly selected using the surrounding information of the current block, or it may be explicitly selected through index (transformation mode information or flipping method information) signaling. Index signaling may be signaled as a truncated unary code, such as 0 for DCT-2, 10 for DST-7 without flipping, and 11 for DST-7 after vertical flipping. Additionally, the binaryization of DCT-2 and DST-7 may be switched and signaled depending on the size of the current block and surrounding information. Furthermore, among the above binary numbers, the first binary number may be signaled in CU units, and the remaining binary numbers may be signaled in TU or PU units. Information about the area where this method is used can be signaled in CTU units, slice units, PPS units, SPS units, or other units representing specific areas, and a 1-bit flag can be signaled in an on / off format.
[0509] Transformation / inverse transformation can be performed by selecting one of four methods for a region within a picture, a CTU, the entire picture, or the current block within a picture group: DCT-2 transform / inverse transformation, DST-7 transform / inverse transformation without flipping, DST-7 transform / inverse transformation after horizontal flipping, or DST-7 transform / inverse transformation after vertical flipping. Information regarding which of the four methods to select may be implicitly selected using the surrounding information of the current block, or it may be explicitly selected through index (transformation mode information or flipping method information) signaling. Index signaling may be signaled as a truncated unary code, such as 0 for DCT-2, 10 for DST-7 without flipping, 110 for DST-7 after horizontal flipping, and 111 for DST-7 after vertical flipping. Additionally, the binarization of DCT-2 and DST-7 may be swapped and signaled depending on the size and surrounding information of the current block. Additionally, the first binary number among the above binary numbers may be signaled in CU units, and the remaining binary numbers may be signaled in TU or PU units. Depending on the intra-frame prediction mode, only some of the four methods may be used. For example, if the intra-frame prediction mode is smaller than the diagonal prediction mode, or is DC mode or Planar mode, only three methods may be used: DCT-2, DST-7 without flipping, and DST-7 after vertical flipping. In this case, transformation mode information or flipping method information may be signaled as DCT-2 0, DST-7 without flipping 10, and DST-7 after vertical flipping 11. For example, if the intra-frame prediction mode is larger than the diagonal prediction mode, only three methods may be used: DCT-2, DST-7 without flipping, and DST-7 after horizontal flipping. In this case, conversion mode information or flipping method information may be signaled, such as DCT-2 being 0, DST-7 being 10 without flipping, and DST-7 being 11 after horizontal flipping.Information about the area where this method is used can be signaled in CTU units, slice units, PPS units, SPS units, or other units representing specific areas, and a 1-bit flag can be signaled in an on / off format.
[0510] Transformation / inverse transformation can be performed by selecting one of five methods for a region or CTU within a picture, the entire picture, or the current block within a picture group: DCT-2 transform / inverse transformation, DST-7 transform / inverse transformation without flipping, DST-7 transform / inverse transformation after horizontal flipping, DST-7 transform / inverse transformation after vertical flipping, or DST-7 transform / inverse transformation after horizontal and vertical flipping. Information on which of the five methods to select for transformation may be implicitly selected using surrounding information of the current block, or may be explicitly selected through index (transformation mode information or flipping method information) signaling. Index signaling can be signaled as a truncated unary code, such as DCT-2 being 0, DST-7 being 10 without flipping, DST-7 being 110 after horizontal flipping, DST-7 being 1110 after vertical flipping, and DST-7 being 1111 after both horizontal and vertical flipping. Additionally, the binarization of DCT-2 and DST-7 may be switched and signaled depending on the size of the current block and surrounding information. Furthermore, the first binary number among the above binary numbers may be signaled in CU units, and the remaining binary numbers may be signaled in TU or PU units. Additionally, the first binary number and the second and third binary numbers among the above binary numbers may be distinguished and signaled as information using a fixed-length code. For example, transformation mode information or flipping method information may be signaled as follows: DCT-2 is 0, DST-7 without flipping is 000, DST-7 after horizontal flipping is 001, DST-7 after vertical flipping is 010, and DST-7 after both horizontal and vertical flipping is 011. Additionally, only some of the five methods may be used depending on the intra-frame prediction mode. For example, if the intra-frame prediction mode is a prediction mode close to the horizontal prediction mode, only three transformation methods may be used: DCT-2, DST-7 without flipping, and DST-7 after vertical flipping.In this case, conversion mode information or flipping method information may be signaled, such as DCT-2 being 0, DST-7 being 10 without flipping, and DST-7 being 11 after vertical flipping.
[0511] For example, if the in-screen prediction mode is a prediction mode close to the vertical prediction mode, only three conversion methods can be used: DCT-2, DST-7 without flipping, and DST-7 after horizontal flipping. In this case, conversion mode information or flipping method information can be signaled as DCT-2 is 0, DST-7 without flipping is 10, and DST-7 after horizontal flipping is 11.
[0512] For example, if the in-frame prediction mode is a prediction mode close to the diagonal prediction mode, only two transformation methods can be used: DCT-2 and DST-7 without flipping. In this case, transformation mode information or flipping method information can be signaled, such as DCT-2 being 0 and DST-7 without flipping being 1.
[0513] For example, if none of the above three cases are present, all five transformation methods may be used: DCT-2, DST-7 without flipping, DST-7 after horizontal flipping, DST-7 after vertical flipping, and DST-7 after both horizontal and vertical flipping. The index for the above transformation method may be signaled by the truncated unary code or fixed-length code method or other methods.
[0514] For example, if the in-frame prediction mode is a non-directional mode, all five transformation methods can be used: DCT-2, DST-7 without flipping, DST-7 after horizontal flipping, DST-7 after vertical flipping, and DST-7 after horizontal and vertical flipping. The index for the transformation method can be signaled by the truncated unary code or fixed-length code method or other methods.
[0515] For example, if the in-screen prediction mode is an odd mode, four transformation methods can be used: DCT-2, DST-7 without flipping, DST-7 after vertical flipping, and DST-7 after horizontal and vertical flipping. In this case, transformation mode information or flipping method information can be signaled as DCT-2 is 0, DST-7 without flipping is 10, DST-7 after vertical flipping is 110, and DST-7 after horizontal and vertical flipping is 111.
[0516] For example, if the in-frame prediction mode is an even mode, four transformation methods can be used: DCT-2, DST-7 without flipping, DST-7 after horizontal flipping, and DST-7 after vertical flipping. In this case, transformation mode information or flipping method information may be signaled, such as 0 for DCT-2, 10 for DST-7 without flipping, 110 for DST-7 after horizontal flipping, and 111 for DST-7 after vertical flipping. Information regarding the area where this method is used may be signaled in CTU units, slice units, PPS units, SPS units, or other units representing specific areas, and a 1-bit flag may be signaled in an on / off format.
[0517] FIGS. 22 and FIGS. 23 each show the locations where residual rearrangement is performed in the encoder and decoder according to the present invention.
[0518] Referring to FIG. 22, residual signal rearrangement may be performed in the encoder before the DST-7 conversion process. Although not shown in FIG. 22, residual signal rearrangement may be performed in the encoder between the conversion process and the quantization process, or after the quantization process.
[0519] Referring to FIG. 23, residual signal rearrangement may be performed in the decoder after the DST-7 inverse transform process. Although not shown in FIG. 23, residual signal rearrangement may be performed between the inverse quantization process and the inverse transform process in the decoder, or residual signal rearrangement may be performed before the inverse quantization process.
[0521] The SDST method according to the present invention has been described above with reference to FIGS. 7 to 23. Below, the decoding method, encoding method, decoder, encoder, and bitstream to which the SDST method according to the present invention is applied will be described in detail with reference to FIGS. 24 and 25.
[0522] FIG. 24 is a diagram illustrating an embodiment of a decoding method using the SDST method according to the present invention.
[0523] Referring to FIG. 24, first, the conversion mode of the current block is determined (S2401), and the remaining data of the current block can be inversely converted according to the conversion mode of the current block (S2402).
[0524] And, depending on the transformation mode of the current block, the remaining data of the inversely transformed current block can be rearranged (S2403).
[0525] Here, the transformation mode may include at least one of SDST (Shuffling Discrete Sine Transform), SDCT (Shuffling Discrete Cosine Transform), DST (Discrete Sine Transform), and DCT (Discrete Cosine Transform).
[0526] SDST mode can indicate a mode that performs inverse transformation in DST-7 transformation mode and performs rearrangement of the inverse transformed residual data.
[0527] SDCT mode can indicate a mode that performs inverse transformation in DCT-2 transformation mode and performs rearrangement of the inverse transformed residual data.
[0528] DST mode can indicate a mode that performs inverse transformation in DST-7 transformation mode and does not perform rearrangement of the remaining inverse transformed data.
[0529] The DCT mode may indicate a mode that performs inverse transformation in DCT-2 transformation mode and does not perform rearrangement of the remaining inverse transformed data.
[0530] Therefore, the step of rearranging the remaining data can be performed only when the transformation mode of the current block is either SDST or SDCT.
[0531] Although it was explained that the inverse conversion to the DST-7 conversion mode is performed for the aforementioned SDST and DST modes, other DST-based conversion modes such as DST-1 and DST-2 may also be used.
[0532] Meanwhile, the step of determining the conversion mode of the current block (S2401) may include the step of obtaining conversion mode information of the current block from a bitstream and the step of determining the conversion mode of the current block based on the conversion mode information.
[0533] Additionally, the step of determining the transformation mode of the current block (S2401) can be determined based on at least one of the prediction mode of the current block, the depth information of the current block, the size of the current block, and the shape of the current block.
[0534] Specifically, if the prediction mode of the current block is an inter-prediction mode, the transformation mode of the current block can be determined as either SDST or SDCT.
[0535] Meanwhile, the step of rearranging the remaining data of the inversely transformed current block (S2403) may include the step of scanning the remaining data arranged within the inversely transformed current block in a first direction order and the step of rearranging the remaining data scanned in the first direction in a second direction order within the inversely transformed current block. Here, the first direction order may be any one of a raster scan order, an up-right diagonal scan order, a horizontal scan order, and a vertical scan order. Additionally, the first direction order may be defined as follows.
[0536] (1) Scan from the top row to the bottom row, but scan from left to right within each row.
[0537] (2) Scan from the top row to the bottom row, but scan from right to left within each row.
[0538] (3) Scan from the bottom row to the top row, but scan from left to right within each row.
[0539] (4) Scan from the bottom row to the top row, but scan from right to left within each row.
[0540] (5) Scan from the left column to the right column, but scan from top to bottom in each column.
[0541] (6) Scan from the left column to the right column, but scan from the bottom to the top in each column.
[0542] (7) Scan from the right column to the left column, but scan from top to bottom in each column.
[0543] (8) Scan from the right column to the left column, but scan from the bottom to the top in each column.
[0544] (9) Scan in a spiral shape: Scan from inside (or outside) the block to outside (or inside) the block, clockwise / counterclockwise.
[0545] Meanwhile, for the second direction sequence, any one of the aforementioned directions may be selectively used. The first direction and the second direction may be the same or different from each other.
[0546] In addition, the step (S2403) of rearranging the remaining data of the inversely transformed current block can be rearranged in units of sub-blocks within the current block. In this case, the remaining data can be rearranged based on the location of the sub-blocks within the current block. Since rearranging the remaining data based on the location of the sub-blocks has been explained in detail in Equation 6 above, a redundant explanation will be avoided.
[0547] Additionally, the step of rearranging the remaining data of the inversely transformed current block (S2403) can rearrange the remaining data arranged within the inversely transformed current block by rotating them at a predetermined angle.
[0548] Additionally, the step of rearranging the remaining data of the inversely transformed current block (S2403) may rearrange the remaining data arranged within the inversely transformed current block by flipping it according to a flipping method. In this case, the step of determining the transformation mode of the current block (S2401) may include the step of obtaining flipping method information from a bitstream and the step of determining the flipping method of the current block based on the flipping method information.
[0549] FIG. 25 is a diagram illustrating an embodiment of an encoding method using the SDST method according to the present invention.
[0550] Referring to FIG. 25, the conversion mode of the current block can be determined (S2501).
[0551] And, the remaining data of the current block can be rearranged according to the conversion mode of the current block (S2502).
[0552] And, the remaining data of the current block rearranged according to the conversion mode of the current block can be converted (S2503).
[0553] Here, the transformation mode may include at least one of SDST (Shuffling Discrete Sine Transform), SDCT (Shuffling Discrete Cosine Transform), DST (Discrete Sine Transform), and DCT (Discrete Cosine Transform). As the SDST, SDCT, DST, and DCT modes have been described in FIG. 24, a redundant description will be avoided.
[0554] Meanwhile, steps to rearrange the remaining data can be performed only when the transformation mode of the current block is either SDST or SDCT.
[0555] Additionally, the step of determining the conversion mode of the current block (S2501) can be determined based on at least one of the prediction mode of the current block, depth information of the current block, size of the current block, and shape of the current block.
[0556] Here, if the prediction mode of the current block is an inter-prediction mode, the transformation mode of the current block can be determined as either SDST or SDCT.
[0557] Meanwhile, the step of rearranging the remaining data of the current block (S2502) may include the step of scanning the remaining data arranged within the current block according to a first direction order and the step of rearranging the remaining data scanned in the first direction within the current block according to a second direction order.
[0558] Additionally, the step of rearranging the remaining data of the current block (S2502) can be rearranged in units of sub-blocks within the current block.
[0559] In this case, the step of rearranging the remaining data of the current block (S2502) can rearrange the remaining data based on the location of the sub-block within the current block.
[0560] Meanwhile, the step of rearranging the remaining data of the current block (S2502) can rearrange the remaining data arranged within the current block by rotating them at a predetermined angle.
[0561] Meanwhile, the step of rearranging the remaining data of the current block (S2502) can rearrange the remaining data arranged within the current block by flipping it according to a flipping method.
[0562] An image decoder using the SDST method according to the present invention may include an inverse transformation unit that determines a transformation mode of a current block, inversely transforms the residual data of the current block according to the transformation mode of the current block, and rearranges the inversely transformed residual data of the current block according to the transformation mode of the current block. Here, the transformation mode may include at least one of SDST (Shuffling Discrete Sine Transform), SDCT (Shuffling Discrete Cosine Transform), DST (Discrete Sine Transform), and DCT (Discrete Cosine Transform).
[0563] An image decoder using the SDST method according to the present invention may include an inverse transformation unit that determines a transformation mode of a current block, rearranges the remaining data of the current block according to the transformation mode of the current block, and inversely transforms the rearranged remaining data of the current block according to the transformation mode of the current block. Here, the transformation mode may include at least one of SDST (Shuffling Discrete Sine Transform), SDCT (Shuffling Discrete Cosine Transform), DST (Discrete Sine Transform), and DCT (Discrete Cosine Transform).
[0564] An image encoder using the SDST method according to the present invention may include a transformation unit that determines a transformation mode of a current block, rearranges the residual data of the current block according to the transformation mode of the current block, and transforms the residual data of the current block rearranged according to the transformation mode of the current block. Here, the transformation mode may include at least one of SDST (Shuffling Discrete Sine Transform), SDCT (Shuffling Discrete Cosine Transform), DST (Discrete Sine Transform), and DCT (Discrete Cosine Transform).
[0565] An image encoder using the SDST method according to the present invention may include a transformation unit that determines a transformation mode of a current block, transforms residual data of a current block according to the transformation mode of a current block, and rearranges the transformed residual data of a current block according to the transformation mode of a current block. Here, the transformation mode may include at least one of SDST (Shuffling Discrete Sine Transform), SDCT (Shuffling Discrete Cosine Transform), DST (Discrete Sine Transform), and DCT (Discrete Cosine Transform).
[0566] A bitstream generated by an encoding method using an SDST method according to the present invention comprises the steps of determining a transformation mode of a current block, rearranging residual data of a current block according to the transformation mode of a current block, and transforming the rearranged residual data of a current block according to the transformation mode of a current block, wherein the transformation mode may include at least one of SDST (Shuffling Discrete Sine Transform), SDCT (Shuffling Discrete Cosine Transform), DST (Discrete Sine Transform), and DCT (Discrete Cosine Transform).
[0567] FIGS. 26 to 31 show examples of locations where a flipping method is performed in an encoder or decoder according to the present invention.
[0568] FIG. 26 is a diagram illustrating an example of an encoding process of a method for performing a transformation after flipping.
[0569] FIG. 27 is a diagram illustrating an example of a decoding process of a method for performing flipping after inverse transformation.
[0570] Referring to FIG. 26, a residual signal is generated by subtracting an inter-frame or intra-frame prediction signal from the original signal for the current block, and then one of DCT-2 transformation or flipping and DST-7 transformation can be selected as the transformation method. If the transformation method is DCT-2 transformation, transformation coefficients can be generated by performing a transformation on the residual signal using DCT-2 transformation. If the transformation method is flipping and DST-7 transformation, one of four flipping methods (no flipping, horizontal flipping, vertical flipping, horizontal and vertical flipping) can be selected to perform flipping on the residual signal, and then transformation coefficients can be generated by performing a transformation on the flipped residual signal using DST-7 transformation. Quantization can be performed on the transformation coefficients to generate quantized levels.
[0571] Referring to FIG. 27, a quantized level can be input and inverse quantization can be performed to generate transform coefficients. A method corresponding to the method selected during the encoding process can be selected from among DCT-2 inverse transform or DST-7 inverse transform and flipping. That is, if DCT-2 transform is performed during the encoding process, DCT-2 inverse transform can be performed during the decoding process. In addition, if flipping and DST-7 transform methods are performed during the encoding process, DST-7 inverse transform and flipping can be performed during the decoding process. If the inverse transform method is DCT-2 inverse transform, an inverse transform can be performed on the transform coefficients using DCT-2 inverse transform to generate a restored residual signal. If the inverse transform method is the DST-7 inverse transform and flipping method, the inverse transform is performed on the residual coefficients using the DST-7 inverse transform to generate a restored residual signal, and then one of four flipping methods (no flipping, horizontal flipping, vertical flipping, horizontal and vertical flipping) is selected to perform flipping on the restored residual signal to generate a flipped and restored residual signal. An inter-frame or intra-frame prediction signal can be added to the restored residual signal or the flipped and restored residual signal to generate a restored signal.
[0572] FIG. 28 is a diagram illustrating an example of an encoding process for performing flipping after conversion.
[0573] FIG. 29 is a diagram illustrating an example of a decoding process of a method for performing inverse transformation after flipping.
[0574] Referring to FIG. 28, a residual signal is generated by subtracting an inter-frame or intra-frame prediction signal from the original signal for the current block, and then one of DCT-2 transformation or DST-7 transformation and flipping can be selected as the transformation method. If the transformation method is DCT-2 transformation, transformation coefficients can be generated by performing a transformation on the residual signal using DCT-2 transformation. If the transformation method is DST-7 transformation and flipping, transformation can be performed on the residual signal using DST-7 transformation, and then one of four flipping methods (no flipping, horizontal flipping, vertical flipping, horizontal and vertical flipping) can be selected to perform flipping on the transformation coefficients to generate flipped transformation coefficients. Quantized levels can be generated by performing quantization on the transformation coefficients or the flipped transformation coefficients. Additionally, when performing flipping on the transformation coefficients, rearrangement of the transformation coefficients can be performed. The method of performing realignment may be the same as flipping, may be a method in which a second transformation is performed to rotate the axis at the zero point of the transformation basis, or may be a method of changing the positive and negative signs of the transformation coefficients.
[0575] Referring to FIG. 29, a quantized level can be input and inverse quantization can be performed to generate transform coefficients. A method corresponding to the method selected during the encoding process is selected among DCT-2 inverse transform, flipping, and DST-7 inverse transform. That is, if DCT-2 transform is performed during the encoding process, DCT-2 inverse transform can be performed during the decoding process. In addition, if DST-7 transform and flipping methods are performed during the encoding process, flipping and DST-7 inverse transform can be performed during the decoding process. If the inverse transform method is DCT-2 inverse transform, the DCT-2 inverse transform can be used to perform an inverse transform on the transform coefficients to generate a restored residual signal. When the inverse transform method is a flipping and DST-7 inverse transform method, a restored residual signal can be generated by selecting one of four flipping methods (no flipping, horizontal flipping, vertical flipping, horizontal and vertical flipping) to perform flipping on the transform coefficients, and then performing an inverse transform on the flipped transform coefficients using the DST-7 inverse transform. A restored signal can be generated by adding an inter-frame or intra-frame prediction signal to the restored residual signal.
[0576] FIG. 30 is a diagram illustrating an example of an encoding process of a method for performing flipping after quantization.
[0577] FIG. 31 is a diagram illustrating an example of a decoding process of a method for performing inverse quantization after flipping.
[0578] Referring to FIG. 30, a residual signal is generated by subtracting an inter-frame or intra-frame prediction signal from the original signal for the current block, and then one of DCT-2 transformation or DST-7 transformation can be selected as the transformation method. If the transformation method is DCT-2 transformation, transformation coefficients can be generated by performing a transformation on the residual signal using DCT-2 transformation. If the transformation method is DST-7 transformation, transformation coefficients can be generated by performing a transformation on the residual signal using DST-7 transformation. Quantized levels can be generated by performing quantization on the transformation coefficients. Additionally, if the transformation method is DST-7 transformation, flipped quantized levels can be generated by performing flipping on the quantized levels by selecting one of four flipping methods (no flipping, horizontal flipping, vertical flipping, horizontal and vertical flipping). Additionally, when performing flipping on the quantized levels, rearrangement of the quantized levels can be performed. The method of performing the reordering may be the same as the flipping method, may be a method in which a second transformation is performed to rotate the axis at the zero of the transformation basis, and may be a method of changing the positive and negative signs of the quantized levels.
[0579] Referring to FIG. 31, a quantized level is input, and an inverse transform method corresponding to the method selected during the encoding process, such as DCT-2 inverse transform or DST-7 inverse transform, is selected. That is, if DCT-2 transform is performed during the encoding process, DCT-2 inverse transform can be performed during the decoding process. In addition, if DST-7 transform is performed during the encoding process, DST-7 inverse transform can be performed during the decoding process. If the inverse transform method is DCT-2 inverse transform, a restored residual signal can be generated by performing inverse quantization on the quantized level to generate transform coefficients, and then performing an inverse transform on the transform coefficients using DCT-2 inverse transform. When the inverse transform method is the DST-7 inverse transform method, one of four flipping methods (no flipping, horizontal flipping, vertical flipping, horizontal and vertical flipping) can be selected to perform flipping on the quantized level, and then inverse quantization can be performed on the flipped quantized level to generate transform coefficients. A restored residual signal can be generated by performing an inverse transform on the transform coefficients using the DST-7 inverse transform. An inter-frame or intra-frame prediction signal can be added to the restored residual signal to generate a restored signal.
[0580] Meanwhile, the location where the flipping method is performed in the decoder can be determined based on information regarding the flipping location signaled from the encoder.
[0581] FIG. 32 is a diagram illustrating the performance of flipping on the remaining blocks.
[0582] Referring to FIG. 32, at least one of ‘no flipping’, ‘horizontal flipping’, ‘vertical flipping’, and ‘horizontal and vertical flipping’ can be performed on the remaining block. As shown in FIG. 32, the position of the sample within the remaining block may be changed depending on the type of flipping.
[0583] FIG. 33 is a diagram illustrating an embodiment for implementing a flipping operation on an 8x8 size residual block in hardware.
[0584] Referring to FIG. 33, when implementing vertical flipping for an MxN residual block in hardware, vertical flipping for the MxN block can be performed by changing the address value (addr) used when reading data from the residual block memory to M-1-addr. That is, vertical flipping can be implemented by reading data within the residual block by changing the memory row address for the MxN block instead of performing a vertical flipping operation.
[0585] When implementing a hardware version of horizontal flipping for MxN remaining blocks, horizontal flipping for MxN blocks can be performed by changing the order of data values in the remaining block memory and reading them. That is, horizontal flipping can be implemented by changing the order in which data values for MxN blocks are read instead of the horizontal flipping operation. For example, if the order of data stored in memory is a, b, c, d, e, f, g, h, horizontal flipping can be performed by reading the data values in the order h, g, f, e, d, c, b, a.
[0586] FIG. 34 is a diagram illustrating the flipping and transformation of the remaining blocks.
[0587] Referring to FIG. 34, at least one of ‘No-Flip’, ‘Horizontal Flipping (H-Flip)’, ‘Vertical Flipping (V-Flip)’, and ‘Horizontal and Vertical Flipping (HV-Flip)’ can be performed on the remaining block, and a DST-7 transformation can be performed. As shown in FIG. 34, depending on the type of flipping, the position of the sample within the remaining block can be changed so that a DST-7 transformation can be performed.
[0588] The following illustrates an example of using an Adaptive Multiple Transform (AMT) method with at least one of the transformations used in this specification.
[0589] An AMT set may be constructed using at least one of the transformations used in this specification. For example, at least one transformation may be added to the AMT transformation set for each intra-frame and inter-frame encoded / decoded block, as well as transformations such as DCT-2, DCT-5, DCT-8, DST-1, and DST-7. Specifically, DST-4 and the identity transform may be added to the AMT transformation set for the inter-frame encoded / decoded block, and KLT-1 and KLT-2 may be added to the AMT transformation set for the intra-frame encoded / decoded block.
[0590] Transformations corresponding to blocks of sizes such as 4x24 and 8x48, rather than powers of 2, may be added. For example, in the in-frame encoding / decoding process, seven sets of transformations, each having four transformation pairs, may be defined as shown in Table 14 below.
[0591] Prediction mode Transformation pair set T 0, 화면내 { (DST-4, DST-4), (DST-7, DST-7), (DST-4, DCT-8), (DCT-8, DST-4)} T 1, 화면내 { (DST-7, DST-7), (DST-7, DCT-5), (DCT-5, DST-7), (DST-1, DCT-5)} T 2, 화면내 { (DST-7, DST-7), (DST-7, DCT-8), (DCT-8, DST-7), (DCT-5, DCT-5)} T 3, 화면내 { (DST-4, DST-4), (DST-4, DCT-5), (DCT-8, DST-4), (DST-1, DST-7)} T 4, 화면내 { (DST-4, DST-7), (DST-7, DCT-5), (DCT-8, DST-7), (DST-1, DST-7)} T 5, 화면내 { (DST-7, DST-7), (DST-7, DCT-5), (DCT-8, DST-7), (DST-1, DST-7)} T 6, 화면내 { (DST-7, DST-7), (DST-7, DCT-5), (DCT-5, DST-7), (DST-1, DST-7)}.
[0592] In Table 14 above, the first item of a transformation pair may represent a vertical transformation, and the second item may represent a horizontal transformation. The sets of transformation pairs in Table 14 above may be defined so that each of the seven transformation sets is assigned based on different in-frame prediction modes and different block sizes. In Table 14 above, T0 through T6 may represent sets of transformation pairs available corresponding to each block size. For example, T0 may be used for a 2x2 block size, T1 for a 4x4 block size, T2 for an 8x8 block size, T3 for a 16x16 block size, T4 for a 32x32 block size, T5 for a 64x64 block size, and T6 for a 128x128 block size.
[0593] The identity transformation may be applied to blocks that do not exceed 16x16. Additionally, the identity transformation may be applied to blocks having a mode close to the horizontal and vertical in-frame prediction directions, and said mode close to the horizontal and / or vertical in-frame prediction directions may be defined by a threshold based on the size of the block. For example, if the transformation index is 3 and the block satisfies the above conditions, the horizontal and / or vertical identity transformation may be applied.
[0594] Meanwhile, in the inter-frame encoding / decoding process, two sets of transformations can be defined, each having four transformation pairs, as shown in Table 15 below.
[0595] T 0, 화면간 { (DCT-8, DCT-8), (DCT-8, DST-7), (DST-7, DCT-8), (DST-7, DST-7)} T 1, 화면간 { (KLT-1, KLT-1), (KLT-1, KLT-2), (KLT-2, KLT-1), (KLT-2, KLT-2)}
[0596] In Table 15 above, T0 and T1 may represent sets of available transformation pairs corresponding to the block size. For example, in Table 15 above, for a block having a size smaller than or equal to 16x16, a set of transformations including KLTs (i.e., T 1, 화면간 ) may be applied, and for blocks larger than 16x16, T 0, 화면간 This can be applied.
[0597] In addition, a method can be used to approximate the AMT transform using only DCT-2 series transform and adjustment steps. The adjustment steps can be defined using block-band orthogonal matrices to transform the DCT-2 series transform into a form similar to the AMT transform.
[0598] The set of primary transforms for AMT used in this specification may consist of DCT-2, DCT-8, DST-4, DST-7 transforms, etc., and the set of primary transforms may also consist of DCT-8, DST-4, and DST-7 transforms. In addition, the DST-7 transform matrix may be implemented by performing flipping, sign changing, etc. based on the DCT-8 transform matrix.
[0599] For example, a set of two-dimensional transformations (i.e., horizontal and vertical transformations) can be constructed using the above transformations and used in the inter-frame encoding / decoding process. In the intra-frame encoding / decoding process, a set of two-dimensional transformations as shown in Table 16 below can be used.
[0600] TrIdxpredModIdx 0 1 2 3 0 DST4,DST4 DST7,DST7 DST4,DCT8 DCT8,DST4 1 DST7,DST7 DST7,DCT2 DCT2,DST7 DCT2, DCT8 2 DST7,DST7 DST7,DCT8 DCT8,DST7 DCT2,DST7 3 DST4,DST4 DST4,DCT2 DCT8,DST4 DCT2,DST7 4 DST4, DST7 DST7,DCT2 DCT8,DST7 DCT2,DST7 5 DST7,DST7 DST7,DCT2 DCT8,DST7 DCT2,DST7 6 DST7,DST7 DST7,DCT2 DCT2,DST7 DCT2,DST7
[0601] Table 16 above shows the sets of transformations for vertical and horizontal transformations for each prediction mode (predModIdx) and transformation index (TrIdx).
[0602] In addition, the above AMT conversion set can be replaced with a conversion set using DCT-8 and DST-7.
[0603] In addition, if the width or height of a block exceeds 32 pixels, the AMT conversion is not applied to the block, and at least one of the AMT conversion usage information (AMT flag) and conversion index information (AMT index) may not be signaled to the decoder.
[0604] The transformation matrices of DCT-8, DST-1, and DCT-5 included in the AMT transformation set used in this specification may be replaced with other transformation matrices. Flipped DST-7 may be used instead of DCT-8. DST-6 may be used instead of DST-1. DCT-2 may be used instead of DCT-5.
[0605] The transformation matrices of the flipped DST-7 and DST-6 above can be derived from DST-7, respectively, as shown in Equation 7 below.
[0606]
[0607] Here, is in the NxN transformation matrix of DST-7 k of the i-th basis vector l It represents the nth component.
[0608] In addition, an AMT transformation including the transformation matrices of the DCT-8, DST-1, and DCT-5 can be applied to both the luminance component and the chrominance component.
[0609] The transformation for the luminance component can be determined based on an explicitly signaled AMT index representing a set of mode-dependent transformations and horizontal and vertical transformations.
[0610] For color difference components / in-frame modes, the conversion can be determined in the same way as the conversion determination method for the luminance component, but the number of conversion candidates may be smaller than the number of conversion candidates for the luminance component.
[0611] For the color difference component / inter-frame mode, the conversion can be determined by a 1-bit flag indicating whether the AMT index is the same as the call block of the luminance component or the base conversion (DCT-2xDCT-2).
[0612] Additionally, AMT can select horizontal and vertical transformations in DCT-2, DST-7, and flipped DST-7 (FDST-7). Additionally, an AMT flag may be defined. An AMT flag of 0 may indicate that DCT-2 is used for both horizontal and vertical transformations, while an AMT flag of 1 may indicate that a different transformation based on the AMT index is used. AMT usage may be permitted only when both the width and height of the block are less than or equal to 64. The AMT flag may be determined by the intra-frame prediction mode. An even intra-frame prediction mode may implicitly assign the AMT flag to 1, and an odd intra-frame prediction mode may implicitly assign the AMT flag to 0. Additionally, an odd intra-frame prediction mode may implicitly assign the AMT flag to 1, and an even intra-frame prediction mode may implicitly assign the AMT flag to 0.
[0613] A set of transformations including two transformations, DST-7 and DCT-8, may be used, and the maximum block size to which AMT is applied may be limited to 32x32. A forward N×N DST-7 with a Discrete Fourier Transform (DFT) of length 2N+1 may be implemented to obtain an N×N DST-7. The 2N+1 FFT may be reconstructed into a 2D FFT. DCT-8 may be derived from DST-7 through sign changes and reordering immediately before and after the DST-7 calculation. Thus, DST-7 may be reused to implement DCT-8.
[0614] The transformation or inverse transformation for the current block may be performed only on the sub-blocks within the current block. For example, the sub-block may be the sub-block located at the top-left position of the current block. The width and height of the sub-block may be determined independently. For example, the width (or height) of the sub-block may be determined according to the type of transformation kernel applied to the horizontal transformation or inverse transformation (or vertical transformation or inverse transformation). For example, if the transformation kernel applied to the horizontal transformation or inverse transformation is DCT-2, the width may be 32 samples. For example, if the transformation kernel applied to the horizontal transformation or inverse transformation is not DCT-2, but, for example, DST-7 or DCT-8, the width may be 16 samples. Likewise, for example, if the transformation kernel applied to the vertical transformation or inverse transformation is DCT-2, the height may be 32 samples. For example, if the transformation kernel applied to the vertical transformation or inverse transformation is not DCT-2, but, for example, DST-7 or DCT-8, the vertical length may be 16 samples. Also, since a sub-block cannot be larger than the current block, if the length of the current block is smaller than the length of the derived sub-block (e.g., 32 samples or 16 samples), the length of the block in which the transformation or inverse transformation is performed may be determined by the length of the current block. For samples within the region of the current block that are not included in the sub-block, the transformation or inverse transformation is not performed, and the sample values of those samples may all be set to '0'. Here, the sub-block may include a residual signal, which is the difference between the input signal and the predicted signal, or a transformation coefficient, which is the transformed form of the residual signal.
[0615] AMT conversion can be implicitly determined during intra-frame and inter-frame encoding / decoding processes.
[0616] The intra prediction mode dependent transforms of the luminance component and the chrominance component can be represented as shown in Tables 17 and 18 below, respectively.
[0617] In-screen prediction mode Horizontal transformation Vertical transformation Block size limit PlanarAng. 31,32,34,36,37 DST-7 DST-7 Width <= 64 && Height <= 64 DCAng. 33, 35 DCT-2 DCT-2 Width <= 64 && Height <= 64 Ang. 2, 4, 6...28,30Ang. 39,41,43 ...63,65 DST-7 DCT-2 Width <= 64 && Height <= 64 Ang. 3,5,7...27,29Ang. 38,40,42...64,66 DCT-2 DST-7 Width <= 64 && Height <= 64
[0618] In-screen prediction mode Horizontal transformation Vertical transformation Block size limit LM modes DST-7 DST-7 Width <= 8 && Height <= 8 Planar DST-7 DST-7 Width <= 16 && Height <= 16 Hor DST-7 DCT-2 Width <= 16 && Height <= 32 Ver, VDIA DCT-2 DST-7 Width <= 32 && Height <= 16
[0619] Here, Table 17 shows the conversion mapping table for the luminance component, and Table 18 shows the conversion mapping table for the chrominance component.
[0620] Additionally, a position-dependent transform can be used for the residual signal of the merge mode. The transform for the residual signal of the merge mode may vary depending on the spatial motion vector predictor (MVP) candidate used for motion compensation of the current block.
[0621] Table 19 below shows the mapping table between MVP positions and transformations.
[0622] MVP position Horizontal transformation Vertical transformation Block size limit L (left) DST-7 DCT-2 Width <= 32 && Height <= 32 A (above) DCT-2 DST-7 Width <= 32 && Height <= 32
[0623] In Table 19 above, for the left MVP candidate, DST-7 and DCT-2 can be used as horizontal and vertical transformations, respectively. Also, for the above MVP candidate, DCT-2 and DST-7 can be used as horizontal and vertical transformations, respectively. In other cases, DCT-2 can be used as the default transformation.
[0624] Transformation usage information combining primary transformation AMT transformation usage information and secondary transformation NSST (non-separable secondary transform) transformation usage information can be entropy encoded / decoded, and the use of AMT and NSST can be represented by a single transform index. Instead of signaling the indices of the primary and secondary transforms independently, the primary and secondary transforms can be combined and signaled by a single transform index. Additionally, the combined transform index can be used for both the luminance component and the chrominance component.
[0625] Additionally, the transformation used in this specification may be selected from a set of N predefined transformation candidates for each block. Here, N may be a positive integer. Each of the transformation candidates may specify a first horizontal transformation, a first vertical transformation, and a second transformation (which may be identical to the identity transformation). The list of transformation candidates may vary depending on the block size and the prediction mode. The selected transformation may be signaled as follows: If the encoding block flag is 1, a flag specifying whether the first transformation of the candidate list is used may be transmitted. If the flag specifying whether the first transformation of the candidate list is used is 0, the following may apply: if the number of non-zero transformation coefficient levels is greater than a threshold, a transformation index indicating the transformed candidate used may be transmitted; otherwise, the second transformation of the list may be used.
[0626] In addition, the second-order transformation NSST can be used only when DCT-2 is used as the primary transformation. Also, for the horizontal or vertical transformation, DST-7 can be selected without signaling when the horizontal or vertical is independently less than or equal to 4.
[0627] Additionally, the AMT flag may be signaled when the number of non-zero conversion factors is greater than the threshold. For inter-frame blocks, the threshold may be set to 2. For intra-frame blocks, the threshold may be set to 0. If the number of non-zero conversion factors is greater than 2, the AMT index may be signaled. Otherwise, it may be estimated to be 0. For NSST, for intra-frame luminance component blocks, the NSST index may be signaled if the sum of the number of non-zero conversion factors for the top-left 8x8 or 4x4 luminance and the number of non-zero AC factors for the top-left 8x8 or 4x4 chrominance components is greater than 2.
[0628] For the residual block, if the width of the block is less than or equal to K, DST-7 may be used instead of DCT-2 for the one-dimensional horizontal transformation. If the height of the block is less than or equal to L, DST-7 may be used instead of DCT-2 for the one-dimensional vertical transformation. Additionally, even if the width or height of the block is less than or equal to K, DCT-2 may be used if the intra-frame prediction mode is the LM (linear model) chroma mode. Here, K and L are positive integers, for example, 4. Also, K and L may have the same or different values. Additionally, the residual block may be a block encoded in intra mode. Also, the residual block may be a chroma block.
[0629] Instead of performing the above flipping method on the residual signal, the transformation / inverse transformation can be performed using the flipped transformation kernel or transformation matrix. Here, the flipped transformation / inverse transformation kernel or transformation / inverse transformation matrix may be a kernel or matrix that has been predefined in the encoder / decoder after the flipping is performed. In this case, since the transformation / inverse transformation is performed using the flipped transformation / inverse transformation matrix, the same effect as performing the flipping on the residual signal can be obtained. Here, the flipping may be at least one of no flipping, horizontal flipping, vertical flipping, or horizontal and vertical flipping. In this case, information regarding whether the flipped transformation / inverse transformation is used can be signaled. Additionally, information regarding whether the flipped transformation / inverse transformation is used can be signaled for the horizontal transformation / inverse transformation and the vertical transformation / inverse transformation, respectively.
[0630] In addition, instead of performing the above-mentioned flipping method on the residual signal, the transformation / inverse transformation can be performed by performing flipping on the transformation kernel or transformation matrix during the encoding / decoding process. In this case, since the transformation / inverse transformation is performed by performing flipping on the transformation / inverse transformation matrix, the same effect as performing flipping on the residual signal can be obtained. Here, the flipping may be at least one of no flipping, horizontal flipping, vertical flipping, or horizontal and vertical flipping. In this case, information on whether or not to perform flipping on the transformation / inverse transformation matrix can be signaled. Additionally, information on whether or not to perform flipping on the transformation / inverse transformation matrix can be signaled for the horizontal transformation / inverse transformation and the vertical transformation / inverse transformation, respectively.
[0631] The above flipping method is determined based on the in-frame prediction mode, and if two or more are used as the in-frame prediction mode of the current block, flipping can be performed before and after the transformation / inverse transformation of the current block as a flipping method for the non-directional mode.
[0632] In addition, the above flipping method is determined based on an in-frame prediction mode, and if two or more are used as in-frame prediction modes for the current block, flipping can be performed before and after the transformation / inverse transformation for the current block as a flipping method for the main directional mode. Here, the main directional mode may be at least one of a vertical mode, a horizontal mode, and a diagonal mode.
[0633] If the size of the above transformation is greater than or equal to MxN, the transformation coefficients existing in the range of M / 2 to M and N / 2 to N during or after the transformation can all be set to a value of 0. Here, M and N are positive integers, and for example, can be 64x64.
[0634] To reduce memory requirements, a right shift operation of K can be performed on the conversion coefficient generated after the above conversion. Additionally, a right shift operation of K can be performed on the temporary conversion coefficient generated after the above horizontal conversion. Additionally, a right shift operation of K can be performed on the temporary conversion coefficient generated after the above vertical conversion. Here, K is a positive integer.
[0635] To reduce memory requirements, a right shift operation of K can be performed on the restored residual signal generated after the above inverse transform. Additionally, a right shift operation of K can be performed on the temporary transformation coefficient generated after the above horizontal inverse transform. Additionally, a right shift operation of K can be performed on the temporary transformation coefficient generated after the above vertical inverse transform. Here, K is a positive integer.
[0636] At least one of the above flipping methods can be performed on at least one of the signals generated before performing horizontal direction conversion / inverse conversion, after performing horizontal direction conversion / inverse conversion, before performing vertical direction conversion / inverse conversion, and after performing vertical direction conversion / inverse conversion. In this case, information on the flipping method used in the horizontal direction conversion / inverse conversion or vertical direction conversion / inverse conversion can be signaled.
[0637] In addition, DCT-4 may be used instead of the above DST-7. 2 N In a DCT-2 transform / inverse transform matrix of size 2 N-1 Since DCT-4 transform / inverse transform matrices of size 1 can be extracted and used, only DCT-2 transform / inverse transform matrices can be stored in the encoder / decoder instead of DCT-4, thereby reducing the memory requirements of the encoder / decoder. In addition, 2 N-1 DCT-4 convert / inverse conversion logic of size 2 N Since it can be utilized from the DCT-2 transform / inverse transform logic of the specified size, the chip area required to implement the encoder / decoder can be reduced. Here, the above example is not applicable only to the DCT-2 and DCT-4, but may also be applied if there is a transform matrix or transform logic shared among at least one of the types of DST transform / inverse transforms and at least one of the types of DCT transform / inverse transforms. That is, another transform / inverse transform matrix or logic can be extracted and used from one transform / inverse transform matrix or logic. Furthermore, for a specific transform / inverse transform size, another transform / inverse transform matrix or logic can be extracted and used from one transform / inverse transform matrix or logic. Additionally, another transform / inverse transform matrix can be extracted from one transform / inverse transform matrix in at least one of the matrix unit, basis vector unit, or matrix coefficient unit.
[0638] Additionally, if the current block is smaller than MxN, a different transform / inverse transform may be used for the transform / inverse transform of the current block instead of a specific transform / inverse transform. Additionally, if the current block is larger than MxN, a different transform / inverse transform may be used for the transform / inverse transform of the current block instead of a specific transform / inverse transform. Here, M and N are positive integers. The specific transform / inverse transform and the other transform / inverse transform may be transforms / inverse transforms predefined in the encoder / decoder.
[0639] Additionally, at least one of the transformations such as DCT-4, DCT-8, DCT-2, DST-4, DST-1, and DST-7 used in this specification may be replaced with at least one of the transformations calculated based on the transformations such as DCT-4, DCT-8, DCT-2, DST-4, DST-1, and DST-7. Here, the calculated transformation may be a transformation calculated by changing the coefficient values within the transformation matrix of DCT-4, DCT-8, DCT-2, DST-4, DST-1, and DST-7. Furthermore, the coefficient values within the transformation matrix of DCT-4, DCT-8, DCT-2, DST-4, DST-1, and DST-7 may have integer values. That is, the transformations such as DCT-4, DCT-8, DCT-2, DST-4, DST-1, and DST-7 may be integer transforms. Additionally, the coefficient values within the transformation matrix calculated above may have integer values. That is, the calculated transformation may be an integer transformation. Furthermore, the calculated transformation may be the result of performing a left shift operation by N on the coefficient values within the transformation matrices of DCT-4, DCT-8, DCT-2, DST-4, DST-1, DST-7, etc., where N may be a positive integer.
[0640] The above DCT-Q and DST-W transformations may include the above DCT-Q and DST-W transformations and the above DCT-Q and DST-W inverse transformations. Here, Q and W may have positive values greater than or equal to 1, and for example, 1 to 9 may be used with the same meaning as I to IX.
[0641] Additionally, the transformations DCT-4, DCT-8, DCT-2, DST-4, DST-1, DST-7, etc. used in this specification are not limited to those transformations only, and at least one of the DCT-Q and DST-W transformations may be used in place of the DCT-4, DCT-8, DCT-2, DST-4, DST-1, and DST-7 transformations. Here, Q and W may have positive values of 1 or more, and, for example, 1 to 9 may be used with the same meaning as I to IX.
[0642] Additionally, the transformation used in this specification may be performed in the form of a square transformation in the case of a square block, and in the form of a non-square transformation in the case of a non-square block; in the case of a square region comprising at least one square block and a non-square block, the transformation may be performed in the form of a square transformation in the region; and in the case of a non-square region comprising at least one square block and a non-square block, the transformation may be performed in the form of a non-square transformation in the region.
[0643] Additionally, the information regarding the rearrangement method in this specification may be flipping method information.
[0644] Additionally, the term "conversion" as used in this specification may mean at least one of a conversion and an inverse conversion.
[0646] In the encoder, to improve the subjective / objective quality of the image, a transformation is performed on the residual block to generate transformation coefficients, the transformation coefficients are quantized to generate quantized coefficient levels, and the quantized coefficient levels can be entropy encoded.
[0647] In the decoder, the quantized coefficient levels can be entropy decoded, the quantized coefficient levels can be inversely quantized to generate transform coefficients, and the transformed coefficients can be inversely transformed to generate the restored residual block.
[0648] Transformation type information regarding which transformation was used for the above transformation and inverse transformation can be explicitly entropy encoded / decoded. Additionally, transformation type information regarding which transformation was used for the above transformation and inverse transformation can be implicitly determined based on at least one of the encoding parameters without entropy encoding / decoding.
[0649] The following describes an embodiment of an image encoding / decoding method, apparatus, and a recording medium storing a bitstream for performing at least one of conversion and inverse conversion in the present invention.
[0651] By using at least one of the following embodiments, a block can be divided into N sub-blocks to perform at least one of prediction, transform / inverse transform, quantization / inverse quantization, and entropy encoding / decoding. Such a mode may be referred to as a first sub-block partitioning mode (e.g., ISP mode, Intra Sub-Partitions mode).
[0652] The above block may mean at least one of an encoding block, a prediction block, and a transformation block. For example, the above block may be a transformation block.
[0653] Additionally, the divided subblock may mean at least one of an encoding block, a prediction block, and a transformation block. For example, the divided subblock may be a transformation block.
[0654] Additionally, the block or the divided sub-block may be at least one of an intra-screen block, an inter-screen block, or an intra-screen block copy. For example, the block and the sub-block may be intra-screen blocks.
[0655] Additionally, the block or the divided sub-block may be at least one of an intra-frame prediction block, an inter-frame prediction block, or an intra-frame block copy prediction block. For example, the sub-block may be an intra-frame prediction block.
[0656] Additionally, the block or the divided sub-block may be at least one of a luminance signal block and a chrominance signal block. For example, the block and the sub-block may be luminance signal blocks.
[0658] When the above block is divided into N sub-blocks, the block before division may be an encoding block, and the divided sub-block may be at least one of a prediction block and a transformation block. That is, prediction, transformation / inverse transformation, quantization / inverse quantization, and entropy encoding / decoding of transformation coefficients may be performed on the size of the divided sub-block.
[0659] Additionally, when the above block is divided into N sub-blocks, the block before division may be at least one of an encoding block and a prediction block, and the divided sub-block may be a transformation block. That is, the prediction is performed with the size of the block before division, while the transformation / inverse transformation, quantization / inverse quantization, and entropy encoding / decoding of the transformation coefficients may be performed with the size of the divided sub-block.
[0661] It is possible to determine whether to divide the block into multiple sub-blocks based on at least one of the area (the product of the width and height, etc.), size (width, height, or a combination of width and height), and shape (rectangle, square, etc.) of the block.
[0662] For example, if the current block is a 64x64 block, the current block can be divided into multiple sub-blocks.
[0663] As another example, if the current block is a 32x32 block, the current block can be divided into multiple sub-blocks.
[0664] As another example, if the current block is a 32x16 block, the current block can be divided into multiple sub-blocks.
[0665] As another example, if the current block is a 16x32 block, the current block can be divided into multiple sub-blocks.
[0666] As another example, if the current block is a 4x4 block, the current block may not be divided into multiple sub-blocks.
[0667] As another example, if the current block is a 2x4 block, the current block may not be divided into multiple sub-blocks.
[0668] As another example, if the area of the current block is 32 or more, the current block can be divided into multiple sub-blocks.
[0669] As another example, if the area of the current block is less than 32, the current block may not be divided into multiple sub-blocks.
[0670] As another example, if the area of the current block is 256 and the shape of the current block is rectangular, the current block can be divided into multiple sub-blocks.
[0671] As another example, if the area of the current block is 16 and the shape of the current block is a square, the current block may not be divided into multiple sub-blocks.
[0673] When dividing the above block, the block may be divided into a plurality of sub-blocks in at least one of the vertical or horizontal dividing directions.
[0674] For example, the current block can be divided into two sub-blocks in the vertical direction.
[0675] As another example, the current block can be divided into two sub-blocks in the horizontal direction.
[0676] As another example, the current block can be divided into four sub-blocks in the horizontal direction.
[0677] As another example, the current block can be divided into four sub-blocks in the vertical direction.
[0679] When dividing a block into N sub-blocks, N can be a positive integer, for example, 2 or 4. Additionally, N can be determined using at least one of the area, size, shape, and division direction of the block.
[0680] For example, if the current block is a 4x8 or 8x4 block, it can be divided into two sub-blocks in the horizontal direction or two sub-blocks in the vertical direction.
[0681] As another example, if the current block is a 16x8 or 16x16 block, it can be divided into 4 sub-blocks in the vertical direction or 4 sub-blocks in the horizontal direction.
[0682] As another example, if the current block is an 8x32 or 32x32 block, it can be divided into 4 sub-blocks in the horizontal direction or 4 sub-blocks in the vertical direction.
[0683] As another example, if the current block is a 16x4, 32x4, or 64x4 block, it can be divided into 4 sub-blocks in the vertical direction. Also, if the current block is a 16x4, 32x4, or 64x4 block, it can be divided into 2 sub-blocks in the horizontal direction.
[0684] As another example, if the current block is 4x16, 4x32, or 4x64, it can be divided into 4 sub-blocks in the horizontal direction. Also, if the current block is 4x16, 4x32, or 4x64, it can be divided into 2 sub-blocks in the vertical direction.
[0685] As another example, if the current block is Jx4, it can be divided into 2 subblocks horizontally. Here, J can be a positive integer.
[0686] As another example, if the current block is 4xK, it can be divided into 2 subblocks in the vertical direction. Here, J can be a positive integer.
[0687] As another example, if the current block is JxK (K>4), it can be divided into 4 subblocks in the horizontal direction. Here, J can be a positive integer.
[0688] As another example, if the current block is JxK (J>4), it can be divided into 4 subblocks in the vertical direction. Here, J can be a positive integer.
[0689] As another example, if the current block is JxK (K>4), it can be divided into 4 subblocks in the vertical direction. Here, J can be a positive integer.
[0691] As another example, if the current block area is 64, it can be divided into 4 sub-blocks in the horizontal or vertical direction.
[0692] As another example, if the current block is 16x4 and the shape of the current block is rectangular, it can be divided into 4 sub-blocks in the vertical direction.
[0693] As another example, if the current block is 1024 and the shape of the current block is a square, it can be divided into 4 sub-blocks in the horizontal or vertical direction.
[0695] In addition, the above sub-block may have at least one of a minimum area, a minimum width, and a minimum height.
[0696] For example, the above sub-block may have a minimum area S. Here, S can be a positive integer, and for example, 16.
[0697] As another example, the above subblock may have a minimum width of J. Here, J can be a positive integer, for example, 4.
[0698] As another example, the above subblock may have a minimum vertical size of K. Here, K can be a positive integer, for example, 4.
[0700] Each divided subblock can generate a restored block by adding a residual block (or restored residual block) and a prediction block. Here, at least one of the restored samples within each restored subblock can be used as a reference sample during in-frame prediction of the subblock that is subsequently encoded / decoded.
[0702] The encoding / decoding order of each sub-block divided from the block can be determined according to at least one of the division directions.
[0703] For example, the encoding / decoding order of each horizontally divided sub-block can be determined from the top direction to the bottom direction.
[0704] As another example, the encoding / decoding order of each vertically divided sub-block can be determined from left to right.
[0706] Each of the above-mentioned divided sub-blocks can share and use the in-screen prediction mode.
[0707] At this time, the in-frame prediction mode information for each sub-block can be entropy encoded / decoded only once in the block prior to splitting.
[0708] Each of the above-mentioned divided sub-blocks can share and use the in-screen block copy mode.
[0709] At this time, the in-screen block copy mode information for each sub-block can be entropy encoded / decoded only once in the block prior to splitting.
[0711] To indicate a subblock partitioning mode that divides the above block into N subblocks and performs at least one of prediction, transformation / inverse transformation, quantization / inverse quantization, and entropy encoding / decoding, at least one of subblock partitioning mode information and partitioning direction information may be entropy encoded / decoded.
[0712] Here, the subblock partitioning mode information may be used to indicate the subblock partitioning mode. When indicating that the subblock partitioning mode is used (second value), the block may be partitioned into subblocks to perform at least one of prediction, transform / inverse transform, quantization / inverse quantization, and entropy encoding / decoding. When indicating that the subblock partitioning mode is not used (first value), the block may not be partitioned into subblocks to perform at least one of prediction, transform / inverse transform, quantization / inverse quantization, and entropy encoding / decoding. Here, the first value may be 0, and the second value may be 1.
[0713] Additionally, the division direction information may be used to indicate whether the sub-block division mode divides in a vertical direction or a horizontal direction. If the division direction information is a first value, the block may be divided in a horizontal direction, and the first value may be 0. Additionally, if the division direction information is a second value, the block may be divided in a vertical direction, and the second value may be 1.
[0715] If the current block does not use the nearest reference sample line (first reference sample line) as a reference sample line, at least one of the sub-block partitioning mode information and partitioning direction information may not be entropy encoded / decoded. In this case, the sub-block partitioning mode information can be inferred that the current block is not partitioned into sub-blocks.
[0716] Here, if the current block does not use the nearest reference sample line (the first reference sample line) as the reference sample line, it may mean that a second or more reference sample lines are used as the reference sample lines restored around the current block.
[0717] That is, only when the current block uses the nearest reference sample line as the reference sample line, at least one of the sub-block partitioning mode information and partitioning direction information can be entropy encoded / decoded.
[0719] At least one of the area, size, shape, and partitioning direction of the coefficient group used for entropy encoding / decoding of the transformation coefficients can be determined based on at least one of the area, size, shape, and partitioning direction of the subblock.
[0720] For example, if the area of the sub-block is 16, the area of the coefficient group can be determined as 16.
[0721] As another example, if the area of the sub-block is 32, the area of the coefficient group can be determined to be 16.
[0722] As another example, if the size of the subblock is 1x16 or 16x1, the size of the coefficient group can be determined to be 1x16 or 16x1.
[0723] As another example, if the size of the subblock is 2x8 or 8x2, the size of the coefficient group can be determined to be 2x8 or 8x2.
[0724] As another example, if the size of the subblock is 4x4, the size of the coefficient group can be determined to be 4x4.
[0725] As another example, if the width of the sub-block is 2, the width of the coefficient group can be determined as 2.
[0726] As another example, if the width of the sub-block is 4, the width of the coefficient group can be determined to be 4.
[0727] As another example, if the vertical size of the sub-block is 2, the vertical size of the coefficient group can be determined as 2.
[0728] As another example, if the vertical size of the sub-block is 4, the vertical size of the coefficient group can be determined to be 4.
[0729] As another example, if the shape of the subblock is rectangular, the shape of the coefficient group can be determined to be rectangular.
[0730] As another example, if the shape of the sub-block is a square, the shape of the coefficient group can be determined to be a square.
[0731] As another example, if the size of the sub-block is 16x4 and the shape is rectangular, the size of the coefficient group can be determined as at least one of 4x4 or 8x2.
[0732] As another example, if the size of the sub-block is 4x8 and the shape is rectangular, the size of the coefficient group can be determined as at least one of 4x4 or 2x8.
[0733] As another example, if the size of the sub-block is 32x4 and the shape is rectangular, the size of the coefficient group can be determined as at least one of 4x4 or 8x2.
[0734] As another example, if the size of the sub-block is 8x64 and the shape is rectangular, the size of the coefficient group can be determined as at least one of 4x4 or 2x8.
[0735] As another example, if the size of the sub-block is 16x4 and the division direction is vertical, the size of the coefficient group can be determined to be 4x4.
[0736] As another example, if the size of the sub-block is 4x8 and the division direction is horizontal, the size of the coefficient group can be determined to be 4x4.
[0737] As another example, if the size of the sub-block is 32x4 and the division direction is horizontal, the size of the coefficient group can be determined to be 8x2.
[0738] As another example, if the size of the subblock is 8x64 and the partitioning direction is vertical, the size of the coefficient group can be determined to be 2x8.
[0740] Each divided sub-block can entropy encode / decode a coded block flag that indicates whether there is at least one conversion factor having a non-zero value on a sub-block basis.
[0741] For example, at least one subblock may be indicated as having at least one conversion factor having a non-zero value for the encoding block flag.
[0742] As another example, if m represents the total number of subblocks and indicates that there are no transformation factors with non-zero values in the encoding block flags of the first m-1 subblocks, it can be inferred that there is at least one transformation factor with a non-zero value in the encoding block flag of the m-th subblock.
[0743] As another example, if the encoding block flag is entropy encoded / decoded at the sub-block level, the encoding block flag may not be entropy encoded / decoded at the block level before splitting.
[0744] As another example, if the encoding block flag is entropy encoded / decoded at the block level before splitting, the encoding block flag may not be entropy encoded / decoded at the sub-block level.
[0746] Meanwhile, if the current block is in the first subblock partitioning mode and the size of the current block is a predefined size, the size of the subblock for in-frame prediction and the size of the subblock for transformation may be different from each other. That is, the subblock partitioning for in-frame prediction and the subblock partitioning for transformation may be different from each other. Here, the predefined size may be 4xN or 8xN (N>4). Here, the subblock partitioning may mean vertical partitioning.
[0747] For example, if the current block is in a first subblock partitioning mode and the size of the current block is 4xN (N>4), the current block may be vertically partitioned into 4xN subblocks for in-frame prediction and may be vertically partitioned into 1xN subblocks for transformation. In this case, a one-dimensional transformation / inverse transformation may be performed to perform a 1xN transformation. That is, a one-dimensional transformation / inverse transformation may be performed based on at least one of the partitioning mode of the current block and the size of the current block.
[0748] As another example, if the current block is in the first subblock partitioning mode and the size of the current block is 8xN (N>4), the current block may be vertically partitioned into 4xN subblocks for in-frame prediction and may be vertically partitioned into 2xN subblocks for transformation. In this case, a 2D transformation / inverse transformation may be performed to perform the 2xN transformation. That is, a 2D transformation / inverse transformation may be performed based on at least one of the partitioning mode of the current block and the size of the current block.
[0749] Here, N can mean a positive integer and can be a positive integer smaller than 64 or 128.
[0750] In addition, the size of the current block may mean at least one of the size of the encoding block, the size of the prediction block, and the size of the transformation block of the current block.
[0752] As shown in the example of FIG. 35, according to an embodiment of the first sub-block division mode, the current block can be divided into two sub-blocks in the horizontal direction and into two sub-blocks in the vertical direction.
[0754] As shown in the example of FIG. 36, according to an embodiment of the first sub-block division mode, the current block can be divided into two sub-blocks in the horizontal direction and into two sub-blocks in the vertical direction.
[0756] As shown in the example of FIG. 37, according to an embodiment of the first sub-block division mode, the current block can be divided into four sub-blocks in the horizontal direction and into four sub-blocks in the vertical direction.
[0758] By using at least one of the following embodiments, a block can be divided into N sub-blocks to perform at least one of transform / inverse transform, quantization / inverse quantization, and entropy encoding / decoding. Such a mode may be referred to as a second sub-block division mode (e.g., SBT mode, Sub-Block Transform mode).
[0759] The above block may mean at least one of an encoding block, a prediction block, and a transformation block. For example, the above block may be a transformation block.
[0760] Additionally, the divided subblock may mean at least one of an encoding block, a prediction block, and a transformation block. For example, the divided subblock may be a transformation block.
[0761] Additionally, the block or the divided sub-block may be at least one of an intra-screen block, an inter-screen block, or an intra-screen block copy block. For example, the block and the sub-block may be inter-screen blocks.
[0762] Additionally, the block or the divided sub-block may be at least one of an intra-frame prediction block, an inter-frame prediction block, or an intra-frame block copy prediction block. For example, the block may be an inter-frame prediction block.
[0763] Additionally, the block or the divided sub-block may be at least one of a luminance signal block and a chrominance signal block. For example, the block and the sub-block may be luminance signal blocks.
[0765] When the above block is divided into N sub-blocks, the block before division may be an encoding block, and the divided sub-block may be at least one of a prediction block and a transformation block. That is, prediction, transformation / inverse transformation, quantization / inverse quantization, and entropy encoding / decoding of transformation coefficients may be performed on the size of the divided sub-block.
[0766] Additionally, when the above block is divided into N sub-blocks, the block before division may be at least one of an encoding block and a prediction block, and the divided sub-block may be a transformation block. That is, the prediction is performed with the size of the block before division, while the transformation / inverse transformation, quantization / inverse quantization, and entropy encoding / decoding of the transformation coefficients may be performed with the size of the divided sub-block.
[0768] It is possible to determine whether to divide the block into multiple sub-blocks based on at least one of the area (the product of the width and height, etc.), size (width, height, or a combination of width and height), and shape (rectangle, square, etc.) of the block.
[0769] For example, if the current block is a 64x64 block, the current block can be divided into multiple sub-blocks.
[0770] As another example, if the current block is a 32x32 block, the current block can be divided into multiple sub-blocks.
[0771] As another example, if the current block is a 32x16 block, the current block can be divided into multiple sub-blocks.
[0772] As another example, if the current block is a 16x32 block, the current block can be divided into multiple sub-blocks.
[0773] As another example, if the current block is a 4x4 block, the current block may not be divided into multiple sub-blocks.
[0774] As another example, if the current block is a 2x4 block, the current block may not be divided into multiple sub-blocks.
[0775] As another example, if at least one of the width and height of the current block is greater than the maximum size of the transformation block, the current block may not be divided into multiple sub-blocks. That is, if at least one of the width and height of the current block is less than or equal to the maximum size of the transformation block, the current block may use a second sub-block division mode.
[0776] Addit...
Claims
Claim 1 A video decoding method comprising: a step of determining whether to perform a second inverse transformation on a current block; a step of determining a horizontal transformation type and a vertical transformation type for performing a first inverse transformation on the current block; a step of performing the first inverse transformation on the current block based on the horizontal transformation type and the vertical transformation type to derive a residual block of the current block; and a step of restoring the current block based on the residual block, wherein when the second inverse transformation is performed on the current block, the horizontal transformation type and the vertical transformation type are determined as DCT-2, and when the current block is generated by dividing coding blocks according to an ISP (Intra Sub-block Partitions) mode, the horizontal transformation type is determined based on the horizontal size of the current block and the vertical transformation type is determined based on the vertical size of the current block, regardless of the intra-frame prediction mode of the current block. Claim 2 A video decoding method according to claim 1, characterized in that the determination of the horizontal conversion type and the vertical conversion type is based on implicit multiple conversion selection information. Claim 3 A video decoding method according to claim 2, wherein, when the current block is generated by dividing the coding block according to an ISP (Intra Sub-block Partitions) mode, the implicit multiple transformation selection information is set to a value indicating an implicit multiple transformation selection. Claim 4 delete Claim 5 A video decoding method according to claim 2, characterized in that when the implicit multiple transformation selection information indicates an implicit multiple transformation selection, the horizontal transformation type and the vertical transformation type are determined based on whether the current block is in SBT (Sub-Block Transform) mode. Claim 6 A video decoding method according to claim 2, characterized in that the implicit multiple transformation selection information indicates an implicit multiple transformation selection, and when the current block is not in SBT (Sub-Block Transform) mode, the horizontal transformation type and the vertical transformation type are determined regardless of the in-frame prediction mode of the current block. Claim 7 delete Claim 8 A video decoding method according to claim 2, characterized in that when the current block is a color difference component, the horizontal conversion type and the vertical conversion type are determined as DCT-2 regardless of the implicit multiple conversion selection information. Claim 9 A video encoding method comprising: a step of determining a horizontal conversion type and a vertical conversion type of a current block; a step of performing a conversion on a remaining block of the current block based on the horizontal conversion type and the vertical conversion type; a step of determining whether to perform a secondary conversion on the current block; and a step of encoding the converted remaining block, wherein the secondary conversion is performed only when the horizontal conversion type and the vertical conversion type are DCT-2, and when the current block is generated by dividing coding blocks according to an ISP (Intra Sub-block Partitions) mode, the horizontal conversion type is determined based on the horizontal size of the current block regardless of the intra-frame prediction mode of the current block, and the vertical conversion type is determined based on the vertical size of the current block. Claim 10 A video encoding method according to claim 9, characterized in that the determination of the horizontal conversion type and the vertical conversion type is based on implicit multiple conversion selection information. Claim 11 A video encoding method according to claim 10, wherein, when the current block is generated by dividing the coding block according to an ISP (Intra Sub-block Partitions) mode, the implicit multiple conversion selection information is set to a value indicating an implicit multiple conversion selection. Claim 12 delete Claim 13 A video encoding method according to claim 10, characterized in that when the implicit multiple transformation selection information indicates an implicit multiple transformation selection, the horizontal transformation type and the vertical transformation type are determined based on whether the current block is in SBT (Sub-Block Transform) mode. Claim 14 A video encoding method according to claim 10, characterized in that the implicit multiple transformation selection information indicates an implicit multiple transformation selection, and when the current block is not in SBT (Sub-Block Transform) mode, the horizontal transformation type and the vertical transformation type are determined regardless of the intra-frame prediction mode of the current block. Claim 15 delete Claim 16 A video encoding method according to claim 10, wherein when the current block is a color difference component, the horizontal conversion type and the vertical conversion type are determined as DCT-2 regardless of the implicit multiple conversion selection information. Claim 17 A non-transient computer-readable recording medium storing a bitstream generated by a video encoding method, wherein the video encoding method comprises: a step of determining a horizontal conversion type and a vertical conversion type of a current block; a step of performing a conversion on a remaining block of the current block based on the horizontal conversion type and the vertical conversion type; a step of determining whether to perform a secondary conversion on the current block; and a step of encoding the converted remaining block, wherein the secondary conversion is performed only when the horizontal conversion type and the vertical conversion type are DCT-2, and wherein, when the current block is generated by dividing coding blocks according to an ISP (Intra Sub-block Partitions) mode, the horizontal conversion type is determined based on the horizontal size of the current block regardless of the intra-frame prediction mode of the current block, and the vertical conversion type is determined based on the vertical size of the current block.
Citation Information
Patent Citations
Implicit Transform Selection in Video Coding
KR1020210135245A
Method and apparatus of transform type assignment for intra sub-partition in video coding
WO2020156454A1
An encoder, a decoder, and corresponding methods that are used for transform process
WO2020177509A1