Method and apparatus for transform-based image encoding / decoding
Through the transform-based image encoding/decoding method, the image is encoded and decoded using different transformation modes, the problem of increasing the amount of high-resolution image data is solved, the encoding/decoding efficiency is improved and the cost is reduced.
Patent Information
- Application Number
- CN202210230960.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-03-14
- Filing Date
- 2017-06-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2037-06-23
AI Technical Summary
The prior art is difficult to effectively solve the problem of increased transmission and storage costs caused by the increase in the amount of data of high-resolution and high-quality images.
By using a transform-based image encoding/decoding method, the transformation mode of the current block is determined, inverse transformation and rearranged, and transformation modes such as SDST, SDCT, DST and DCT are used to improve the encoding/decoding efficiency.
The encoding/decoding efficiency and transformation efficiency of images are improved, and the transmission and storage costs of image data are reduced.
Smart Images

Figure CN114401407B_ABST
Abstract
Description
[0001] This application is a divisional application of an invention patent application with an application date of June 23, 2017, application number "201780040348.5", and title "Method and device for transform-based image encoding / decoding". Technical Field
[0002] The present invention relates to a method and apparatus for encoding and decoding an image. More particularly, the present invention relates to a method and apparatus for encoding and decoding a video image based on a transform. Background Art
[0003] Recently, the demand for high-resolution and high-quality images such as high-definition (HD) images and ultra-high-definition (UHD) images has been increasing in various application fields. However, the data volume of image data with higher resolution and quality has increased compared to traditional image data. Therefore, when image data is transmitted by using a medium such as a traditional wired broadband network and a wireless broadband network, or when image data is stored by using a traditional storage medium, the cost of transmission and storage increases. In order to solve these problems arising as the resolution and quality of image data are improved, efficient image encoding / decoding technology is required for images with higher resolution and higher quality.
[0004] The image compression technology includes various technologies, including: an inter-frame prediction technology for predicting a pixel value included in a current picture from a previous picture or a subsequent picture of the current picture; an intra-frame prediction technology for predicting a pixel value included in the current picture by using pixel information in the current picture; a transform and quantization technology for compressing the energy of a residual signal; an entropy coding technology for assigning a short code to a value with a high frequency of occurrence and a long code to a value with a low frequency of occurrence; etc. By using such an image compression technology, image data can be effectively compressed and can be transmitted or stored. Summary of the invention
[0005] Technical issues
[0006] An object of the present invention is to provide a method and apparatus for encoding and decoding a video image based on transformation to improve encoding / decoding efficiency of the image.
[0007] Another object of the present invention is to provide a method and apparatus for encoding and decoding video images based on transformation to improve image transformation efficiency.
[0008] Solution
[0009] A method for decoding a video according to the present invention includes: determining a transform mode of a current block; inversely transforming residual data of the current block according to the transform mode of the current block; and rearranging the inversely transformed residual data of the current block according to the transform mode of the current block, wherein the transform mode includes at least one of SDST (reordered discrete sine transform), SDCT (reordered discrete cosine transform), DST (discrete sine transform) and DCT (discrete cosine transform).
[0010] In the method for decoding a video according to the present invention, the rearrangement may be performed only when a transform mode of a current block is one of SDST and SDCT.
[0011] In the method for decoding a video according to the present invention, the step of determining the transform mode of the current block may include: obtaining transform mode information of the current block from a bitstream; and determining the transform mode of the current block based on the transform mode information.
[0012] In the method for decoding a video according to the present invention, the step of determining the transformation mode of the current block may be performed based on at least one of the following items: a prediction mode of the current block, depth information of the current block, a size of the current block, and a shape of the current block.
[0013] In the method for decoding a video according to the present invention, when a prediction mode of a current block is an inter prediction mode, a transform mode of the current block may be determined as one of SDST and SDCT.
[0014] In a method for decoding a video according to the present invention, the rearrangement may include: scanning the inverse transformed residual data arranged in the current block along a first direction; and rearranging the inverse transformed residual data scanned along the first direction in the current block along a second direction.
[0015] In the method for decoding a video according to the present invention, the rearrangement may be performed on each subblock in the current block.
[0016] In the method for decoding a video according to the present invention, the rearrangement may be to rearrange the residual data based on the position of a sub-block in the current block.
[0017] In the method for decoding a video according to the present invention, the rearrangement is performed by rotating the inverse-transformed residual data arranged in the current block by a predefined angle.
[0018] A method for encoding a video according to the present invention includes: determining a transformation mode of a current block; rearranging residual data of the current block according to the transformation mode of the current block; and transforming the rearranged residual data of the current block according to the transformation mode of the current block, wherein the transformation mode includes at least one of SDST (reordered discrete sine transform), SDCT (reordered discrete cosine transform), DST (discrete sine transform) and DCT (discrete cosine transform).
[0019] In the method for encoding a video according to the present invention, the rearrangement may be performed only when a transform mode of a current block is one of SDST and SDCT.
[0020] In the method for encoding a video according to the present invention, the step of determining the transformation mode of the current block may be performed based on at least one of the following items: a prediction mode of the current block, depth information of the current block, a size of the current block, and a shape of the current block.
[0021] In the method for encoding a video according to the present invention, when a prediction mode of a current block is an inter prediction mode, a transform mode of the current block may be determined as one of SDST and SDCT.
[0022] In the method for encoding a video according to the present invention, the rearranging may include: scanning the residual data arranged in the current block along a first direction; and rearranging the residual data scanned along the first direction in the current block along a second direction.
[0023] In the method for encoding a video according to the present invention, the rearrangement may be performed on each subblock in the current block.
[0024] In the method for encoding a video according to the present invention, the rearrangement may be to rearrange the residual data based on the position of a sub-block in the current block.
[0025] In the method for encoding a video according to the present invention, the rearrangement may be performed by rotating the residual data arranged in the current block at a predefined angle.
[0026] A device for decoding a video according to the present invention includes an inverse transform unit, which determines a transform mode of a current block, inversely transforms residual data of the current block according to the transform mode of the current block, and rearranges the inversely transformed residual data of the current block according to the transform mode of the current block, wherein the transform mode includes at least one of SDST (reordered discrete sine transform), SDCT (reordered discrete cosine transform), DST (discrete sine transform) and DCT (discrete cosine transform).
[0027] A device for encoding a video according to the present invention includes a transform unit, which determines a transform mode of a current block, rearranges residual data of the current block according to the transform mode of the current block, and transforms the rearranged residual data of the current block according to the transform mode of the current block, wherein the transform mode includes at least one of SDST (reordered discrete sine transform), SDCT (reordered discrete cosine transform), DST (discrete sine transform) and DCT (discrete cosine transform).
[0028] A recording medium according to the present invention stores a bit stream, which is formed by a method for encoding a video, the method comprising: determining a transform mode of a current block; rearranging residual data of the current block according to the transform mode of the current block; and transforming the rearranged residual data of the current block according to the transform mode of the current block, wherein the transform mode includes at least one of SDST (shuffled discrete sine transform), SDCT (shuffled discrete cosine transform), DST (discrete sine transform) and DCT (discrete cosine transform).
[0029] Beneficial Effects
[0030] According to the present invention, the encoding / decoding efficiency of an image can be improved.
[0031] According to the present invention, the image conversion efficiency can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a block diagram showing the configuration of an encoding device according to an embodiment of the present invention.
[0033] Figure 2 is a block diagram showing the configuration of a decoding device according to an embodiment of the present invention.
[0034] Figure 3 is a diagram schematically showing a partition structure of an image when the image is encoded and decoded.
[0035] Figure 4 is a diagram illustrating a form of a prediction unit (PU) that may be included in a coding unit (CU).
[0036] Figure 5 is a diagram illustrating a form of a transform unit (TU) that may be included in a coding unit (CU).
[0037] Figure 6 is a diagram for explaining an embodiment of a process of intra prediction.
[0038] Figure 7 is a diagram for explaining an embodiment of a process of inter-frame prediction.
[0039] Figure 8 is a diagram for explaining a transform set according to an intra prediction mode.
[0040] Fig. 9 is a diagram for explaining the process of conversion.
[0041] Fig.10 is a diagram for explaining scanning of quantized transform coefficients.
[0042] Fig.11 is a diagram for explaining block partitioning.
[0043] Fig.12 is a diagram showing basis vectors in the DCT-2 frequency domain according to the present invention.
[0044] Fig.13 is a diagram showing basis vectors in the DST-7 frequency domain according to the present invention.
[0045] Fig.14 is a diagram illustrating distribution of average residual values of positions in a 2N×2N prediction unit (PU) of an 8×8 coding unit (CU) predicted in an inter mode according to a Cactus sequence according to the present invention.
[0046] Fig.15 is a three-dimensional graph showing distribution characteristics of residual values in a 2N×2N prediction unit (PU) of an 8×8 coding unit (CU) predicted in an inter prediction mode (inter mode) according to the present invention.
[0047] Fig.16 is a diagram illustrating distribution characteristics of a residual signal in a 2N×2N prediction unit (PU) mode of a coding unit (CU) according to the present invention.
[0048] Fig.17 is a diagram illustrating distribution characteristics of residual signals before and after rearrangement of a 2N×2N prediction unit (PU) according to the present invention.
[0049] Fig.18 is a diagram showing an example of rearrangement of 4×4 residual data of a subblock according to the present invention.
[0050] Fig.19a and Fig.19b is a diagram illustrating a partition structure of a transform unit (TU) according to a prediction unit (PU) mode of a coding unit (CU) and a rearrangement method of a transform unit (TU) according to the present invention.
[0051] Fig. 20 is a diagram illustrating a result of performing DCT-2 and SDST based on a residual signal distribution of a 2N×2N prediction unit (PU) according to the present invention.
[0052] Fig.21 is a diagram illustrating an SDST process according to the present invention.
[0053] Fig. 22 is a diagram illustrating distribution characteristics of a residual absolute value and a transform unit (TU) partition mode of a coding unit (CU) based on inter-prediction according to the present invention.
[0054] Fig.23a , Figure 23b and Fig.23c is a diagram illustrating a scanning order and a rearrangement order of a residual signal for a transform unit (TU) having a depth zero in a prediction unit (PU) according to the present invention.
[0055] Fig.24 is a flow chart illustrating an encoding process of selecting DCT-2 or SDST through rate-distortion optimization (RDO) according to the present invention.
[0056] Fig.25 is a flowchart illustrating a decoding process of selecting DCT-2 or SDST according to the present invention.
[0057] Fig.26 is a flow chart illustrating a decoding process using SDST according to the present invention.
[0058] Fig. 27 and Fig.28 is a diagram illustrating a location where residual signal rearrangement (residual rearrangement) is performed in an encoder and a decoder according to the present invention.
[0059] Fig.29 is a flowchart illustrating a decoding method using the SDST method according to the present invention.
[0060] Fig.30 is a flowchart illustrating an encoding method using the SDST method according to the present invention. DETAILED DESCRIPTION
[0061] Various modifications may be made to the present invention, and there are various embodiments of the present invention, wherein examples of the various embodiments will now be provided with reference to the accompanying drawings and examples of the various embodiments will be described in detail. However, the present invention is not limited thereto, although the exemplary embodiments may be interpreted as including all modifications, equivalents or replacement forms within the technical concept and technical scope of the present invention. Similar reference numerals refer to functions that are identical or similar in all respects. In the accompanying drawings, the shapes and sizes of the elements may be exaggerated for clarity. In the following detailed description of the present invention, reference is made to the accompanying drawings that illustrate specific embodiments in which the present invention may be implemented by way of illustration. These embodiments are described in sufficient detail to enable those skilled in the art to implement the present disclosure. It should be understood that the various embodiments of the present disclosure, although different, are not necessarily mutually exclusive. For example, without departing from the spirit and scope of the present disclosure, the specific features, structures and characteristics associated with one embodiment described herein may be implemented in other embodiments. In addition, it should be understood that without departing from the spirit and scope of the present disclosure, the position or arrangement of each element within each disclosed embodiment may be modified. Therefore, the following detailed description is not intended to be limiting, and the scope of the present disclosure is defined by the appended claims (with the full range of equivalents claimed by the claims, in the case of appropriate interpretation).
[0062] The terms "first", "second", etc. used in the specification may be used to describe various components, but these components are not to be construed as limiting the terms. The terms are only used to distinguish one component from another component. For example, without departing from the scope of the present invention, a "first" component may be referred to as a "second" component, and a "second" component may also be similarly referred to as a "first" component. The term "and / or" includes a combination of a plurality of items or any one of a plurality of items.
[0063] It will be understood that in this specification, when an element is simply referred to as being “connected to” or “coupled to” another element rather than being “directly connected to” or “directly coupled to” another element, it may be “directly connected to” or “directly coupled to” another element, or connected to or coupled to another element with other elements interposed therebetween. Conversely, it will be understood that when an element is referred to as being “directly coupled to” or “directly coupled to” another element, there are no intervening elements.
[0064] In addition, the components shown in the embodiments of the present invention are shown independently to present characteristic functions different from each other. Therefore, this does not mean that each component is composed of a component unit of separate hardware or software. In other words, for convenience, each component includes each of the enumerated components. Therefore, at least two components in each component can be combined to form a component, or a component can be divided into multiple components to perform each function. Without departing from the essence of the present invention, the embodiment in which each component is combined and the embodiment in which a component is divided are also included in the scope of the present invention.
[0065] The terms used in this specification are only used to describe specific embodiments and are not intended to limit the present invention. Expressions used in the singular include plural expressions unless it has a significantly different meaning in the context. In this specification, it will be understood that terms such as "including ...", "having ..." etc. are intended to indicate the existence of features, quantities, steps, behaviors, elements, parts, or combinations thereof disclosed in the specification, and are not intended to exclude the possibility that one or more other features, quantities, steps, behaviors, elements, parts, or combinations thereof may exist or may be added. In other words, when a particular element is referred to as "included", elements other than the corresponding element are not excluded, but other elements may be included in embodiments of the present invention or in the scope of the present invention.
[0066] In addition, some components may not be indispensable components for performing the necessary functions of the present invention, but optional components that only improve their performance. The present invention can be implemented by only including the indispensable components for implementing the essence of the present invention and excluding the components used when improving performance. The structure that only includes the indispensable components and excludes the optional components used when only improving performance is also included in the scope of the present invention.
[0067] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In describing the exemplary embodiments of the present invention, well-known functions or structures will not be described in detail because they will unnecessarily obscure the understanding of the present invention. The same components in the accompanying drawings are represented by the same reference numerals, and repeated descriptions of the same components will be omitted.
[0068] In addition, hereinafter, an image may mean a picture constituting a video, or may mean a video itself. For example, "encode or decode an image, or both" may mean "encode or decode a video, or both", and may mean "encode or decode one image among a plurality of images of a video, or both". Here, a picture and an image may have the same meaning.
[0069] Terminology Description
[0070] Encoder: Can mean a device that performs encoding.
[0071] Decoder: May mean a device that performs decoding.
[0072] Parsing: may mean determining the value of a syntax element by performing entropy decoding, or may mean the entropy decoding itself.
[0073] Block: can be meant as a sample point of an M×N matrix. Here, M and N are positive integers, and a block can be meant as a sample point matrix in a two-dimensional form.
[0074] Sample: is the basic unit of a block and can indicate a range of 0 to 2 depending on the bit depth (Bd). Bd The value of -1. A sample point in the present invention may mean a pixel.
[0075] Unit: It may mean a unit for encoding and decoding an image. When encoding and decoding an image, a unit may be a region generated by partitioning one image. In addition, a unit may mean a sub-division unit when one image is partitioned into a plurality of sub-division units during encoding or decoding. When encoding and decoding an image, a predetermined process for each unit may be performed. One unit may be partitioned into sub-units having a size smaller than that of the unit. Depending on the function, a unit may mean a block, a macroblock, a coding tree unit, a coding tree block, a coding unit, a coding block, a prediction unit, a prediction block, a transform unit, a transform block, etc. In addition, in order to distinguish a unit from a block, a unit may include a luminance component block, a chrominance component block of a luminance component block, and a syntax element of each color component block. A unit may have various sizes and shapes, and specifically, the shape of a unit may be a two-dimensional geometric figure such as a rectangle, a square, a trapezoid, a triangle, a pentagon, etc. In addition, the unit information may include at least one of a unit type (indicating a coding unit, a prediction unit, a transform unit, etc.), a unit size, a unit depth, an order in which a unit is encoded and decoded, and the like.
[0076] Reconstruction neighboring unit: may mean a reconstruction unit that has been previously encoded or decoded in space / time, and the reconstruction unit is adjacent to the encoding / decoding target unit. Here, the reconstruction neighboring unit may mean a reconstruction neighboring block.
[0077] Neighboring block: may mean a block adjacent to the encoding / decoding target block. A block adjacent to the encoding / decoding target block may mean a block having a boundary in contact with the encoding / decoding target block. A neighboring block may mean a block located at an adjacent vertex of the encoding / decoding target block. A neighboring block may mean a reconstructed neighboring block.
[0078] Unit depth: It can mean the degree to which a unit is partitioned. In a tree structure, the root node can be the highest node and the leaf node can be the lowest node.
[0079] Symbol: It can mean the syntax element, coding parameter, value of transform coefficient, etc. of the encoding / decoding target unit.
[0080] Parameter set: may refer to header information in the structure of a bitstream. A parameter set may include at least one of a video parameter set, a sequence parameter set, a picture parameter set, or an adaptation parameter set. In addition, a parameter set may refer to slice header information and tile header information, etc.
[0081] Bitstream: can mean a string of bits that includes encoded image information.
[0082] Prediction unit: may mean a basic unit when performing inter-frame prediction or intra-frame prediction and compensation for prediction. One prediction unit may be partitioned into a plurality of partitions. In this case, each of the plurality of partitions may be a basic unit when performing prediction and compensation, and each partition obtained from the prediction unit partition may be a prediction unit. In addition, one prediction unit may be partitioned into a plurality of small prediction units. The prediction unit may have various sizes and shapes, and specifically, the shape of the prediction unit may be a two-dimensional geometric figure such as a rectangle, a square, a trapezoid, a triangle, a pentagon, and the like.
[0083] Prediction unit partition: may refer to the shape of the partitioned prediction unit.
[0084] Reference picture list: may mean a list including at least one reference picture, wherein the at least one reference picture is used for inter prediction or motion compensation. The type of the reference picture list may be List Combined (LC), List 0 (L0), List 1 (L1), List 2 (L2), List 3 (L3), etc. At least one reference picture list may be used for inter prediction.
[0085] Inter-frame prediction indicator: may mean one of the following: the inter-frame prediction direction (unidirectional prediction, bidirectional prediction, etc.) of the encoding / decoding target block in the case of inter-frame prediction, the number of reference pictures used to generate the prediction block through the encoding / decoding target block, and the number of reference blocks used to perform inter-frame prediction or motion compensation through the encoding / decoding target block.
[0086] Reference picture index: may mean the index of a specific reference picture in a reference picture list.
[0087] Reference picture: may mean a picture that a specific unit refers to for inter-frame prediction or motion compensation. A reference image may be referred to as a reference picture.
[0088] Motion vector: is a two-dimensional vector used for inter-frame prediction or motion compensation, and may mean the offset between the encoding / decoding target picture and the reference picture. For example, (mvX, mvY) may indicate a motion vector, mvX may indicate a horizontal component, and mvY may indicate a vertical component.
[0089] Motion vector candidate: may mean a unit that becomes a prediction candidate when predicting a motion vector, or may mean a motion vector of the unit.
[0090] Motion vector candidate list: may mean a list configured by using motion vector candidates.
[0091] Motion vector candidate index: may mean an indicator indicating a motion vector candidate in a motion vector candidate list. The motion vector candidate index may be referred to as an index of a motion vector predictor.
[0092] Motion information: may mean a motion vector, a reference picture index, and an inter prediction indicator, as well as information including at least one of reference picture list information, a reference picture, a motion vector candidate, a motion vector candidate index, and the like.
[0093] Merge candidate list: may mean a list configured by using merge candidates.
[0094] Merge candidates: may include spatial merge candidates, temporal merge candidates, combined merge candidates, combined bi-predictive merge candidates, zero merge candidates, etc. The merge candidates may include motion information such as prediction type information, reference picture index for each list, motion vector, etc.
[0095] Merge index: may mean information indicating a merge candidate in a merge candidate list. In addition, the merge index may indicate a block of a derived merge candidate among reconstructed blocks that are spatially / temporally adjacent to the current block. In addition, the merge index may indicate at least one of a plurality of motion information of the merge candidate.
[0096] Transform unit: may mean a basic unit when encoding / decoding similar to transform, inverse transform, quantization, inverse quantization, and transform coefficient encoding / decoding is performed on a residual signal. One transform unit may be partitioned into a plurality of small transform units. The transform unit may have various sizes and shapes. Specifically, the shape of the transform unit may be a two-dimensional geometric figure such as a rectangle, a square, a trapezoid, a triangle, a pentagon, etc.
[0097] Scaling: may refer to the process of multiplying a factor by the transform coefficient levels, as a result of which transform coefficients may be generated. Scaling may also be referred to as inverse quantization.
[0098] Quantization parameter: may mean a value used when scaling transform coefficient levels during quantization and inverse quantization. Here, the quantization parameter may be a value mapped to a step size of quantization.
[0099] Delta quantization parameter: may refer to the difference between the quantization parameter of the encoding / decoding target unit and the predicted quantization parameter.
[0100] Scan: may mean a method of sorting the order of coefficients within a block or matrix. For example, the operation of sorting a two-dimensional matrix into a one-dimensional matrix may be called scanning, and the operation of sorting a one-dimensional matrix into a two-dimensional matrix may be called scanning or inverse scanning.
[0101] Transform coefficient: may mean a coefficient value generated after performing a transform. In the present invention, a quantized transform coefficient level (ie, a transform coefficient to which quantization is applied) may be referred to as a transform coefficient.
[0102] Non-zero transform coefficient: may mean a transform coefficient whose value is not zero, or may mean a transform coefficient level whose value is not zero.
[0103] Quantization matrix: may refer to a matrix used in quantization and inverse quantization in order to improve the subject quality or object quality of an image. A quantization matrix may be referred to as a scaling list.
[0104] Quantization matrix coefficient: may refer to each element of the quantization matrix. Quantization matrix coefficient may be referred to as matrix coefficient.
[0105] Default matrix: may mean a predetermined quantization matrix that is predefined in an encoder and a decoder.
[0106] Non-default matrix: may mean a quantization matrix sent / received by the user without being pre-defined in the encoder and decoder.
[0107] Coding tree unit: may be composed of one luminance component (Y) coding tree unit and two associated chrominance component (Cb, Cr) coding tree units. Each coding tree unit may be partitioned by using at least one partitioning method (such as a quadtree, a binary tree, etc.) to constitute sub-units such as coding units, prediction units, transform units, etc. Coding tree unit may be used as a term for indicating a pixel block (i.e., a processing unit in the decoding / encoding process of an image, such as a partition of an input image).
[0108] Coding tree block: may be used as a term used to indicate one of a Y coding tree unit, a Cb coding tree unit, and a Cr coding tree unit.
[0109] Figure 1 is a block diagram showing the configuration of an encoding device according to an embodiment of the present invention.
[0110] The encoding device 100 may be a video encoding device or an image encoding device. A video may include one or more images. The encoding device 100 may encode one or more images of a video in a time sequence.
[0111] Reference Figure 1 , the encoding device 100 may include a motion prediction unit 111, a motion compensation unit 112, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180 and a reference picture buffer 190.
[0112] The encoding device 100 may encode the input picture in an intra mode or an inter mode or in both the intra mode and the inter mode. In addition, the encoding device 100 may generate a bit stream by encoding the input picture, and may output the generated bit stream. When the intra mode is used as a prediction mode, the switch 115 may switch to the intra mode. When the inter mode is used as a prediction mode, the switch 115 may switch to the inter mode. Here, the intra mode may be referred to as an intra prediction mode, and the inter mode may be referred to as an inter prediction mode. The encoding device 100 may generate a prediction block of an input block of the input picture. In addition, after generating the prediction block, the encoding device 100 may encode the residual between the input block and the prediction block. The input picture may be referred to as a current image as a target of current encoding. The input block may be referred to as a current block or may be referred to as a coding target block as a target of current encoding.
[0113] When the prediction mode is intra mode, the intra prediction unit 120 may use the pixel value of the previous encoding block adjacent to the current block as a reference pixel. The intra prediction unit 120 may perform spatial prediction by using the reference pixel, and may generate a prediction sample of the input block by using the spatial prediction. Here, intra prediction may mean intra-frame prediction.
[0114] When the prediction mode is the inter mode, the motion prediction unit 111 may search for an area that best matches the input block from the reference picture in the motion prediction process and may derive a motion vector by using the searched area. The reference picture may be stored in the reference picture buffer 190.
[0115] The motion compensation unit 112 may generate a prediction block by performing motion compensation using a motion vector. Here, the motion vector may be a two-dimensional vector for inter-frame prediction. In addition, the motion vector may indicate an offset between a current picture and a reference picture. Here, inter-frame prediction may mean inter-frame prediction.
[0116] When the value of the motion vector is not an integer, the motion prediction unit 111 and the motion compensation unit 112 may generate a prediction block by applying an interpolation filter to a partial area in the reference picture. In order to perform inter-frame prediction or motion compensation based on the coding unit, it may be determined which methods are used by the motion prediction and compensation methods of the prediction unit in the coding unit among the skip mode, merge mode, AMVP mode, and current picture reference mode. Inter-frame prediction or motion compensation may be performed according to each mode. Here, the current picture reference mode may mean a prediction mode using a pre-reconstructed area of the current picture with the encoding target block. In order to specify the pre-reconstructed area, a motion vector for the current picture reference mode may be defined. Whether the encoding target block is encoded according to the current picture reference mode may be encoded by using the reference picture index of the encoding target block.
[0117] The subtractor 125 may generate a residual block by using a residual between the input block and the prediction block. The residual block may be referred to as a residual signal.
[0118] The transform unit 130 may generate a transform coefficient by transforming the residual block, and may output the transform coefficient. Here, the transform coefficient may be a coefficient value generated by transforming the residual block. In the transform skip mode, the transform unit 130 may skip transforming the residual block.
[0119] The quantized transform coefficient levels may be generated by applying quantization to the transform coefficients. Hereinafter, in an embodiment of the present invention, the quantized transform coefficient levels may be referred to as transform coefficients.
[0120] The quantization unit 140 may generate quantized transform coefficient levels by quantizing the transform coefficients according to the quantization parameters, and may output the quantized transform coefficient levels. Here, the quantization unit 140 may quantize the transform coefficients by using a quantization matrix.
[0121] The entropy encoding unit 150 may generate a bitstream by performing entropy encoding on the value calculated by the quantization unit 140 or the encoding parameter value calculated in the encoding process according to the probability distribution, and may output the generated bitstream. The entropy encoding unit 150 may perform entropy encoding on information for decoding an image, and may perform entropy encoding on information of pixels of the image. For example, the information for decoding an image may include a syntax element, etc.
[0122] When entropy coding is applied, a symbol is represented by allocating a small number of bits to a symbol with a high probability of occurrence and a large number of bits to a symbol with a low probability of occurrence, thereby reducing the size of the bit stream for encoding the target symbol. Therefore, by entropy coding, the compression performance of image coding can be improved. For entropy coding, the entropy coding unit 150 can use a coding method such as exponential Golomb, context adaptive variable length coding (CAVLC) and context adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 can perform entropy coding by using a variable length coding / code (VLC) table. In addition, the entropy coding unit 150 can derive a binarization method of the target symbol and a probability model of the target symbol / binary bit, and can then perform arithmetic coding by using the derived binarization method or the derived probability model.
[0123] In order to encode the transformation coefficient level, the entropy coding unit 150 may change the coefficients in the two-dimensional block form into a one-dimensional vector form by using a transformation coefficient scanning method. For example, by scanning the coefficients of the block with an upper right scan, the coefficients in the two-dimensional form may be changed into a one-dimensional vector form. Depending on the size of the transformation unit and the intra-frame prediction mode, a vertical direction scan for scanning the coefficients in the two-dimensional block form in the column direction and a horizontal direction scan for scanning the coefficients in the two-dimensional block form in the row direction may be used instead of using an upper right scan. That is, depending on the size of the transformation unit and the intra-frame prediction mode, it may be determined which scanning method among the upper right scan, the vertical direction scan, and the horizontal direction scan will be used.
[0124] The coding parameters may include information such as syntax elements that are encoded by the encoder and sent to the decoder, and may include information that can be derived in the encoding or decoding process. The coding parameters may mean information necessary to encode or decode an image. For example, the coding parameters may include at least one value or combination of the following items: block size, block depth, block partition information, unit size, unit depth, unit partition information, partition flag in quadtree form, partition flag in binary tree form, partition direction in binary tree form, intra-frame prediction mode, intra-frame prediction direction, reference sample filtering method, prediction block boundary filtering method, filter taps, filter coefficients, inter-frame prediction mode, motion information, motion vector, reference picture index, inter-frame prediction direction, inter-frame prediction indicator, reference picture list, motion vector predictor, motion vector candidate list, information on whether motion merge mode is used, motion merge candidate, motion merge candidate list, information on whether skip mode is used, interpolation filter type type, motion vector size, accuracy of motion vector representation, transform type, transform size, information on whether an additional (secondary) transform is used, information on whether a residual signal exists, coding block pattern, coding block flag, quantization parameter, quantization matrix, filter information within the loop, information on whether the filter is applied within the loop, filter coefficients within the loop, binarization / debinarization method, context model, context binary bit, bypass binary bit, transform coefficient, transform coefficient level, transform coefficient level scanning method, image display / output order, slice identification information, slice type, slice partition information, parallel block identification information, parallel block type, parallel block partition information, picture type, bit depth, and information of luminance signal or chrominance signal.
[0125] The residual signal may mean the difference between the original signal and the prediction signal. Alternatively, the residual signal may be a signal generated by transforming the difference between the original signal and the prediction signal. Alternatively, the residual signal may be a signal generated by transforming and quantizing the difference between the original signal and the prediction signal. The residual block may be a residual signal of a block unit.
[0126] When the encoding device 100 performs encoding by using inter-frame prediction. The encoded current picture may be used as a reference picture for another image to be processed subsequently. Therefore, the encoding device 100 may decode the encoded current picture and may store the decoded image as a reference picture. In order to perform decoding, inverse quantization and inverse transformation may be performed on the encoded current picture.
[0127] The quantized coefficients may be dequantized by the dequantization unit 160, and may be inversely transformed by the inverse transform unit 170. The dequantized and inversely transformed coefficients may be added to the prediction block by the adder 175, and thus a reconstructed block may be generated.
[0128] The reconstructed block may pass through the filter unit 180. The filter unit 180 may apply at least one of a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF) to the reconstructed block or the reconstructed picture. The filter unit 180 may be referred to as a loop filter.
[0129] The deblocking filter can remove block distortion that occurs at the boundary between blocks. In order to determine whether the deblocking filter is run, it can be determined whether the deblocking filter is applied to the current block based on the pixels in a number of rows or columns included in the block. When the deblocking filter is applied to the block, a strong filter or a weak filter can be applied according to the required deblocking filter strength. In addition, when applying the deblocking filter, horizontal filtering and vertical filtering can be processed in parallel.
[0130] Sample adaptive offset can add an optimal offset value to a pixel value to compensate for a coding error. Sample adaptive offset can correct the offset between a deblocked filtered image and an original picture for each pixel. In order to perform offset correction on a specific picture, a method of applying an offset considering edge information of each pixel can be used, or a method of partitioning pixels of an image into a predetermined number of regions, determining the region where offset correction is to be performed, and applying offset correction to the determined region can be used.
[0131] The adaptive loop filter may perform filtering based on a value obtained by comparing the reconstructed picture with the original picture. The pixels of the image may be partitioned into predetermined groups, a filter applied to each group is determined, and different filtering may be performed in each group. Information about whether the adaptive loop filter is applied to the luminance signal may be sent for each coding unit (CU). The shape and filter coefficients of the adaptive loop filter applied to each block may vary. In addition, an adaptive loop filter having the same form (fixed form) may be applied without considering the characteristics of the target block.
[0132] The reconstructed block passed through the filter unit 180 may be stored in the reference picture buffer 190 .
[0133] Figure 2 is a block diagram showing the configuration of a decoding device according to an embodiment of the present invention.
[0134] The decoding device 200 may be a video decoding device or an image decoding device.
[0135] Reference Figure 2 , the decoding device 200 may include an entropy decoding unit 210, a dequantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 255, a filter unit 260 and a reference picture buffer 270.
[0136] The decoding apparatus 200 may receive a bitstream output from the encoding apparatus 100. The decoding apparatus 200 may decode the bitstream in an intra mode or an inter mode. In addition, the decoding apparatus 100 may generate a reconstructed picture by performing decoding, and may output the reconstructed picture.
[0137] When the prediction mode used in decoding is the intra mode, the switch may switch to the intra mode. When the prediction mode used in decoding is the inter mode, the switch may switch to the inter mode.
[0138] The decoding device 200 may obtain a reconstructed residual block from the input bit stream and may generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 may generate a reconstructed block as a decoding target block by adding the reconstructed residual block to the prediction block. The decoding target block may be referred to as a current block.
[0139] The entropy decoding unit 210 may generate symbols by performing entropy decoding on the bit stream according to the probability distribution. The generated symbols may include symbols with quantized transform coefficient levels. Here, the entropy decoding method may be similar to the above-mentioned entropy encoding method. For example, the entropy decoding method may be an inverse process of the above-mentioned entropy encoding method.
[0140] In order to decode the transform coefficient level, the entropy decoding unit 210 may perform transform coefficient scanning, whereby the coefficients in the one-dimensional vector form may be changed into the two-dimensional block form. For example, by scanning the coefficients of the block with an upper right scan, the coefficients in the one-dimensional vector form may be changed into the two-dimensional block form. Depending on the size of the transform unit and the intra-frame prediction mode, vertical direction scanning and horizontal direction scanning may be used instead of using the upper right scan. That is, depending on the size of the transform unit and the intra-frame prediction mode, it may be determined which scanning method is used among the upper right scan, the vertical direction scan, and the horizontal direction scan.
[0141] The quantized transform coefficient levels may be dequantized by the dequantization unit 220 and may be inversely transformed by the inverse transform unit 230. The quantized transform coefficient levels are dequantized and inversely transformed to generate a reconstructed residual block. Here, the dequantization unit 220 may apply a quantization matrix to the quantized transform coefficient levels.
[0142] When the intra mode is used, the intra prediction unit 240 may generate a prediction block by performing spatial prediction using pixel values of a previously decoded block adjacent to a decoding target block.
[0143] When the inter-frame mode is used, the motion compensation unit 250 may generate a prediction block by performing motion compensation, wherein the motion compensation uses both the reference picture stored in the reference picture buffer 270 and the motion vector. When the value of the motion vector is not an integer, the motion compensation unit 250 may generate a prediction block by applying an interpolation filter to a partial area in the reference picture. In order to perform motion compensation, based on the coding unit, which method is used by the motion compensation method of the prediction unit in the coding unit may be determined among the skip mode, the merge mode, the AMVP mode, and the current picture reference mode. In addition, motion compensation may be performed according to the mode. Here, the current picture reference mode may mean a prediction mode using a previously reconstructed area within the current picture with the decoding target block. The previously reconstructed area may not be adjacent to the decoding target block. In order to indicate the previously reconstructed area, a fixed vector may be used for the current picture reference mode. In addition, a flag or index indicating whether the decoding target block is a block decoded according to the current picture reference mode may be sent with a signal and may be derived by using the reference picture index of the decoding target block. The current picture for the current picture reference mode may exist at a fixed position (e.g., a position where the reference picture index is 0 or the last position) within the reference picture list for the decoding target block. In addition, the current picture may be variably located within the reference picture list, for which a reference picture index indicating the position of the current picture may be signaled. Here, signaling a flag or index may mean that an encoder entropy encodes a corresponding flag or index and includes the corresponding flag or index into a bitstream and a decoder entropy decodes the corresponding flag or index from the bitstream.
[0144] The reconstructed residual block may be added to the prediction block by the adder 255. A block generated by adding the reconstructed residual block and the prediction block may pass through the filter unit 260. The filter unit 260 may apply at least one of a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the reconstructed block or the reconstructed picture. The filter unit 260 may output the reconstructed picture. The reconstructed picture may be stored in the reference picture buffer 270 and may be used for inter-frame prediction.
[0145] Figure 3 is a diagram schematically showing a partition structure of an image when the image is encoded and decoded. Figure 3 An embodiment of partitioning a unit into a plurality of sub-units is schematically shown.
[0146] In order to effectively partition an image, a coding unit (CU) may be used in encoding and decoding. Here, a coding unit may mean a unit for encoding. A unit may be a combination of 1) a syntax element and 2) a block including image samples. For example, "partition of a unit" may mean "partition of a block associated with a unit". Block partition information may include information about the depth of the unit. The depth information may indicate the number of times a unit is partitioned or the degree to which a unit is partitioned, or both.
[0147] Reference Figure 3 , the image 300 is sequentially partitioned for each maximum coding unit (LCU), and the partition structure is determined for each LCU. Here, LCU and coding tree unit (CTU) have the same meaning. A unit may have depth information based on a tree structure and may be partitioned hierarchically. Each partitioned sub-unit may have depth information. The depth information indicates the number of times the unit is partitioned or the degree to which the unit is partitioned, or both, and therefore, the depth information may include information about the size of the sub-unit.
[0148] The partition structure may mean the distribution of coding units (CUs) in the LCU 310. A CU may be a unit for efficiently encoding / decoding an image. The distribution may be determined based on whether a CU will be partitioned multiple times (i.e., a positive integer equal to or greater than 2, including 2, 4, 8, 16, etc.). The width size and height size of the partitioned CU may be half the width size and half the height size of the original CU, respectively. Optionally, depending on the number of partitions, the width size and height size of the partitioned CU may be smaller than the width size and height size of the original CU, respectively. The partitioned CU may be recursively partitioned into multiple further partitioned CUs, wherein, according to the same partitioning method, the further partitioned CU has a width size and height size that is smaller than the width size and height size of the partitioned CU.
[0149] Here, partitioning of the CU may be recursively performed up to a predetermined depth. The depth information may be information indicating the size of the CU and may be stored in each CU. For example, the depth of the LCU may be 0, and the depth of the minimum coding unit (SCU) may be a predetermined maximum depth. Here, the LCU may be a coding unit having the above-mentioned maximum size, and the SCU may be a coding unit having the minimum size.
[0150] Whenever the LCU 310 starts to be partitioned and the width and height sizes of the CU are reduced by the partition operation, the depth of the CU increases by 1. In the case where the CU cannot be partitioned, the CU may have a 2N×2N size for each depth. In the case where the CU can be partitioned, a CU having a 2N×2N size may be partitioned into a plurality of N×N sized CUs. Whenever the depth increases by 1, the size of N is halved.
[0151] For example, when one coding unit is partitioned into four sub-coding units, the width size and height size of one of the four sub-coding units may be half the width size and half the height size of the original coding unit, respectively. For example, when a coding unit of 32×32 size is partitioned into four sub-coding units, each of the four sub-coding units may have a size of 16×16. When one coding unit is partitioned into four sub-coding units, the coding unit may be partitioned in a quadtree form.
[0152] For example, when one coding unit is partitioned into two sub-coding units, the width size or height size of one of the two sub-coding units may be half the width size or half the height size of the original coding unit, respectively. For example, when a coding unit of size 32×32 is partitioned vertically into two sub-coding units, each of the two sub-coding units may have a size of 16×32. For example, when a coding unit of size 32×32 is partitioned horizontally into two sub-coding units, each of the two sub-coding units may have a size of 32×16. When one coding unit is partitioned into two sub-coding units, the coding unit may be partitioned in a binary tree form.
[0153] Reference Figure 3 , the size of an LCU having a minimum depth of 0 may be 64×64 pixels, and the size of an SCU having a maximum depth of 3 may be 8×8 pixels. Here, a CU having 64×64 pixels (ie, LCU) may be represented by a depth of 0, a CU having 32×32 pixels may be represented by a depth of 1, a CU having 16×16 pixels may be represented by a depth of 2, and a CU having 8×8 pixels (ie, SCU) may be represented by a depth of 3.
[0154] In addition, information about whether a CU will be partitioned may be indicated by partition information of the CU. The partition information may be 1-bit information. The partition information may be included in all CUs except the SCU. For example, when the value of the partition information is 0, the CU may not be partitioned, and when the value of the partition information is 1, the CU may be partitioned.
[0155] Figure 4 is a diagram illustrating a form of a prediction unit (PU) that may be included in a coding unit (CU).
[0156] A CU that is no longer partitioned among a plurality of CUs partitioned from the LCU may be partitioned into at least one prediction unit (PU). This process may also be referred to as partitioning.
[0157] PU may be a basic unit for prediction. PU may be encoded and decoded in any one of skip mode, inter mode, and intra mode. PU may be partitioned in various forms according to the mode.
[0158] Also, a coding unit may not be partitioned into a plurality of prediction units, and the coding unit and the prediction unit may have the same size.
[0159] like Figure 4 As shown, in skip mode, the CU may not be partitioned. In skip mode, a 2N×2N mode 410 having the same size as a non-partitioned CU may be supported.
[0160] In inter mode, eight partition modes may be supported in a CU. For example, in inter mode, 2N×2N mode 410, 2N×N mode 415, N×2N mode 420, N×N mode 425, 2N×nU mode 430, 2N×nD mode 435, nL×2N mode 440, and nR×2N mode 445 may be supported. In intra mode, 2N×2N mode 410 and N×N mode 425 may be supported.
[0161] One coding unit may be partitioned into one or more prediction units. One prediction unit may be partitioned into one or more sub-prediction units.
[0162] For example, when one prediction unit is partitioned into four sub-prediction units, the width size and height size of one of the four sub-prediction units may be half the width size and half the height size of the original prediction unit. For example, when a 32×32-sized prediction unit is partitioned into four sub-prediction units, each of the four sub-prediction units may have a 16×16 size. When one prediction unit is partitioned into four sub-prediction units, the prediction unit may be partitioned in a quadtree form.
[0163] For example, when one prediction unit is partitioned into two sub-prediction units, the width size or height size of one of the two sub-prediction units may be half the width size or half the height size of the original prediction unit. For example, when a 32×32-sized prediction unit is partitioned vertically into two sub-prediction units, each of the two sub-prediction units may have a size of 16×32. For example, when a 32×32-sized prediction unit is partitioned horizontally into two sub-prediction units, each of the two sub-prediction units may have a size of 32×16. When one prediction unit is partitioned into two sub-prediction units, the prediction unit may be partitioned in a binary tree form.
[0164] Figure 5 is a diagram illustrating a form of a transform unit (TU) that may be included in a coding unit (CU).
[0165] A transform unit (TU) may be a basic unit for transform, quantization, inverse transform, and inverse quantization within a CU. A TU may have a square shape or a rectangular shape, etc. A TU may be independently determined according to the size of a CU or the form of a CU or both.
[0166] The CU that is no longer partitioned among the CUs partitioned from the LCU may be partitioned into at least one TU. Here, the partition structure of the TU may be a quadtree structure. For example, Figure 5 As shown, a CU 510 may be partitioned one or more times according to a quadtree structure. The case where a CU is partitioned at least once may be referred to as recursive partitioning. By partitioning, a CU 510 may be formed by TUs of different sizes. Alternatively, a CU may be partitioned into at least one TU according to the number of vertical lines for partitioning the CU or the number of horizontal lines for partitioning the CU, or both. The CU may be partitioned into TUs that are symmetrical to each other, or may be partitioned into TUs that are asymmetrical to each other. In order to partition the CU into TUs that are symmetrical to each other, information on the size / shape of the TU may be sent with a signal and may be derived from the information on the size / shape of the CU.
[0167] Also, a coding unit may not be partitioned into transformation units, and the coding unit and the transformation unit may have the same size.
[0168] One coding unit may be partitioned into at least one transformation unit, and one transformation unit may be partitioned into at least one sub-transformation unit.
[0169] For example, when one transformation unit is partitioned into four sub-transformation units, the width size and height size of one of the four sub-transformation units may be half the width size and half the height size of the original transformation unit, respectively. For example, when a 32×32-sized transformation unit is partitioned into four sub-transformation units, each of the four sub-transformation units may have a size of 16×16. When one transformation unit is partitioned into four sub-transformation units, the transformation unit may be partitioned in a quadtree form.
[0170] For example, when one transformation unit is partitioned into two sub-transformation units, the width size or height size of one of the two sub-transformation units may be half the width size or half the height size of the original transformation unit, respectively. For example, when a 32×32-sized transformation unit is vertically partitioned into two sub-transformation units, each of the two sub-transformation units may have a size of 16×32. For example, when a 32×32-sized transformation unit is horizontally partitioned into two sub-transformation units, each of the two sub-transformation units may have a size of 32×16. When one transformation unit is partitioned into two sub-transformation units, the transformation unit may be partitioned in a binary tree form.
[0171] When performing the transformation, the residual block may be transformed by using at least one of the predetermined transformation methods. For example, the predetermined transformation method may include discrete cosine transform (DCT), discrete sine transform (DST), KLT, etc. Which transformation method is applied to transform the residual block may be determined by using at least one of the following items: inter-frame prediction mode information of the prediction unit, intra-frame prediction mode information of the prediction unit, and the size / shape of the transform block. Information indicating the transformation method may be sent by signaling.
[0172] Figure 6 is a diagram for explaining an embodiment of a process of intra prediction.
[0173] The intra prediction mode may be a non-directional mode or a directional mode. The non-directional mode may be a DC mode or a planar mode. The directional mode may be a prediction mode having a specific direction or angle, and the number of directional modes may be M which is equal to or greater than 1. The directional mode may be indicated as at least one of a mode number, a mode value, and a mode angle.
[0174] The number of intra prediction modes may be N which is equal to or greater than 1, including non-directional modes and directional modes.
[0175] The number of intra prediction modes may vary depending on the size of the block. For example, when the block size is 4×4 or 8×8, the number of intra prediction modes may be 67, when the block size is 16×16, the number of intra prediction modes may be 35, when the block size is 32×32, the number of intra prediction modes may be 19, and when the block size is 64×64, the number of intra prediction modes may be 7.
[0176] The number of intra prediction modes may be fixed to N regardless of the size of the block. For example, the number of intra prediction modes may be fixed to at least one of 35 or 67 regardless of the size of the block.
[0177] The number of intra prediction modes may vary depending on the type of color component. For example, the number of prediction modes may vary depending on whether the color component is a luminance signal or a chrominance signal.
[0178] Intra-coding and / or decoding may be performed by using sample values or encoding parameters included in the reconstructed neighboring blocks.
[0179] In order to encode / decode the current block according to intra prediction, it is possible to identify whether a sample included in a reconstructed neighboring block can be used as a reference sample of a coding / decoding target block. When there are samples that cannot be used as reference samples of a coding / decoding target block, by using at least one sample among the samples included in the reconstructed neighboring block, sample values are copied and / or interpolated to the samples that cannot be used as reference samples, whereby the samples that cannot be used as reference samples can be used as reference samples of a coding / decoding target block.
[0180] In intra prediction, based on at least one of the intra prediction mode and the size of the encoding / decoding target block, a filter may be applied to at least one of the reference sample or the prediction sample. Here, the encoding / decoding target block may mean the current block, and may mean at least one of the encoding block, the prediction block, and the transform block. The type of filter applied to the reference sample or the prediction sample may vary depending on at least one of the intra prediction mode or the size / shape of the current block. The type of filter may vary depending on at least one of the number of filter taps, the filter coefficient value, or the filter strength.
[0181] In a non-directional planar mode among intra prediction modes, when a prediction block of an encoding / decoding target block is generated, a sample value in the prediction block may be generated according to the sample position by using a weighted sum of an upper reference sample of a current sample, a left reference sample of the current sample, an upper right reference sample of the current block, and a lower left reference sample of the current block.
[0182] In the non-directional DC mode among the intra prediction modes, when generating a prediction block of a coding / decoding target block, the prediction block may be generated by the average of the upper reference sample of the current block and the left reference sample of the current block. In addition, filtering may be performed on one or more upper rows and one or more left columns adjacent to the reference sample in the coding / decoding block by using the reference sample value.
[0183] In the case of multiple directional modes (angle modes) among intra prediction modes, a prediction block may be generated by using an upper right reference sample and / or a lower left reference sample, and the multiple directional modes may have different directions. In order to generate prediction sample values, interpolation of real number units may be performed.
[0184] In order to perform the intra prediction method, the intra prediction mode of the current prediction block may be predicted from the intra prediction mode of the neighboring prediction block adjacent to the current prediction block. In the case of predicting the intra prediction mode of the current prediction block by using the mode information predicted from the neighboring intra prediction mode, when the current prediction block and the neighboring prediction block have the same intra prediction mode, information that the current prediction block and the neighboring prediction block have the same intra prediction mode may be transmitted by using predetermined flag information. When the intra prediction mode of the current prediction block is different from the intra prediction mode of the neighboring prediction block, the intra prediction mode information of the encoding / decoding target block may be encoded by performing entropy encoding.
[0185] Figure 7 is a diagram for explaining an embodiment of a process of inter-frame prediction.
[0186] Figure 7 The quadrilateral shown in may indicate an image (or picture). Figure 7 The arrow may indicate the prediction direction. That is, the image may be encoded or decoded or encoded and decoded according to the prediction direction. According to the encoding type, each image may be classified into an I picture (intra picture), a P picture (unidirectional prediction picture), a B picture (bidirectional prediction picture), etc. Each picture may be encoded and decoded according to the encoding type of each picture.
[0187] When the image targeted for encoding is an I picture, the picture itself can be intra-coded without inter-frame prediction. When the image targeted for encoding is a P picture, the image can be encoded by performing inter-frame prediction or motion compensation using a reference picture only in the forward direction. When the image targeted for encoding is a B picture, the image can be encoded by performing inter-frame prediction or motion compensation using reference pictures in both the forward and reverse directions. Alternatively, the image can be encoded by performing inter-frame prediction or motion compensation using a reference picture in one of the forward and reverse directions. Here, when the inter-frame prediction mode is used, the encoder can perform inter-frame prediction or motion compensation, and the decoder can perform motion compensation in response to the encoder. Images of P pictures and B pictures that are encoded or decoded or encoded and decoded by using reference pictures can be regarded as images for inter-frame prediction.
[0188] Hereinafter, inter prediction according to an embodiment will be described in detail.
[0189] Inter prediction or motion compensation may be performed by using both reference pictures and motion information. In addition, inter prediction may use the above-described skip mode.
[0190] The reference picture may be at least one of a previous picture and a subsequent picture of the current picture. Here, inter-frame prediction may predict a block of the current picture based on the reference picture. Here, the reference picture may mean an image used when predicting a block. Here, an area within the reference picture may be indicated by using a reference picture index (refIdx) indicating the reference picture, a motion vector, etc.
[0191] Inter prediction can select a reference picture and a reference block related to the current block in the reference picture. A prediction block of the current block can be generated by using the selected reference block. The current block can be a block of the current picture that is a current encoding target or a current decoding target.
[0192] The motion information may be derived from the process of inter-frame prediction by the encoding device 100 and the decoding device 200. In addition, the derived motion information may be used when performing inter-frame prediction. Here, the encoding device 100 and the decoding device 200 may improve the encoding efficiency or the decoding efficiency or both by using the motion information of the reconstructed neighboring blocks or the motion information of the co-located blocks (col blocks) or the motion information of both. The col block may be a block related to the spatial position of the encoding / decoding target block within the previously reconstructed co-located picture (col picture). The reconstructed neighboring block may be a block within the current picture, and a block previously reconstructed by encoding or decoding or both encoding or decoding. In addition, the reconstructed block may be a block adjacent to the encoding / decoding target block, or a block located at the outer corner of the encoding / decoding target block, or both. Here, the block located at the outer corner of the encoding / decoding target block may be a block vertically adjacent to the neighboring block horizontally adjacent to the encoding / decoding target block. Alternatively, the block located at the outer corner of the encoding / decoding target block may be a block horizontally adjacent to the neighboring block vertically adjacent to the encoding / decoding target block.
[0193] The encoding device 100 and the decoding device 200 may respectively determine a block existing at a position spatially related to the encoding / decoding target block within the col picture, and may determine a predefined relative position based on the determined block. The predefined relative position may be an internal position or an external position of the block existing at a position spatially related to the encoding / decoding target block, or both. In addition, the encoding device 100 and the decoding device 200 may respectively derive the col block based on the determined predefined relative position. Here, the col picture may be one of at least one reference picture included in the reference picture list.
[0194] The method of deriving motion information may vary according to the prediction mode of the encoding / decoding target block. For example, the prediction mode applied to inter prediction may include advanced motion vector prediction (AMVP), merge mode, etc. Here, merge mode may be referred to as motion merge mode.
[0195] For example, when AMVP is applied as a prediction mode, the encoding device 100 and the decoding device 200 may generate a motion vector candidate list by using a motion vector of a reconstructed neighboring block or a motion vector of a col block, or both. The motion vector of the reconstructed neighboring block or the motion vector of the col block, or both may be used as a motion vector candidate. Here, the motion vector of the col block may be referred to as a temporal motion vector candidate, and the motion vector of the reconstructed neighboring block may be referred to as a spatial motion vector candidate.
[0196] The encoding device 100 may generate a bitstream, and the bitstream may include a motion vector candidate index. That is, the encoding device 100 may generate a bitstream by entropy encoding the motion vector candidate index. The motion vector candidate index may indicate an optimal motion vector candidate selected from the motion vector candidates included in the motion vector candidate list. The motion vector candidate index may be sent from the encoding device 100 to the decoding device 200 through the bitstream.
[0197] The decoding apparatus 200 may entropy-decode a motion vector candidate index from a bitstream, and may select a motion vector candidate of a decoding target block among motion vector candidates included in a motion vector candidate list by using the entropy-decoded motion vector candidate index.
[0198] The encoding device 100 may calculate a motion vector difference (MVD) between a motion vector of a decoding target block and a motion vector candidate, and may entropy encode the MVD. The bitstream may include the entropy-encoded MVD. The MVD may be sent from the encoding device 100 to the decoding device 200 through the bitstream. Here, the decoding device 200 may entropy decode the MVD received from the bitstream. The decoding device 200 may derive the motion vector of the decoding target block by summing the decoded MVD and the motion vector candidate.
[0199] The bitstream may include a reference picture index indicating a reference picture, etc., and the reference picture index may be entropy encoded and transmitted from the encoding device 100 to the decoding device 200 through the bitstream. The decoding device 200 may predict a motion vector of a decoding target block by using motion information of a neighboring block, and may derive a motion vector of the decoding target block by using the predicted motion vector and a motion vector difference. The decoding device 200 may generate a prediction block of the decoding target block based on the derived motion vector and the reference picture index information.
[0200] As another method of deriving motion information, merge mode is used. Merge mode may mean the merging of motions of multiple blocks. Merge mode may mean that the motion information of one block is applied to another block. When merge mode is applied, the encoding device 100 and the decoding device 200 may generate a merge candidate list by using the motion information of the reconstructed adjacent block or the motion information of the col block, or both. The motion information may include at least one of the following items: 1) motion vector, 2) reference picture index, and 3) inter-frame prediction indicator. The prediction indicator may indicate unidirectional (L0 prediction, L1 prediction) or bidirectional.
[0201] Here, the merge mode may be applied to each CU or each PU. When the merge mode is performed on each CU or each PU, the encoding device 100 may generate a bitstream by entropy decoding the predefined information, and may send the bitstream to the decoding device 200. The bitstream may include the predefined information. The predefined information may include: 1) a merge flag as information indicating whether the merge mode is performed for each block partition, 2) a merge index as information indicating which block among the neighboring blocks adjacent to the encoding target block is merged. For example, the neighboring blocks adjacent to the encoding target block may include the left neighboring block of the encoding target block, the upper neighboring block of the encoding target block, the temporal neighboring block of the encoding target block, etc.
[0202] The merge candidate list may indicate a list storing motion information. In addition, the merge candidate list may be generated before executing the merge mode. The motion information stored in the merge candidate list may be at least one of the following motion information: motion information of a neighboring block adjacent to the encoding / decoding target block, motion information of a co-located block related to the encoding / decoding target block in a reference picture, motion information newly generated by pre-combining motion information existing in the motion candidate list, and a zero merge candidate. Here, the motion information of a neighboring block adjacent to the encoding / decoding target block may be referred to as a spatial merge candidate. The motion information of a co-located block related to the encoding / decoding target block in a reference picture may be referred to as a temporal merge candidate.
[0203] The skip mode may be a mode in which the mode information of the neighboring block itself is applied to the encoding / decoding target block. The skip mode may be one of the modes for inter-frame prediction. When the skip mode is used, the encoding device 100 may entropy encode the information about which block's motion information is used as the motion information of the encoding target block, and may send the information to the decoding device 200 through the bitstream. The encoding device 100 may not send other information (e.g., syntax element information) to the decoding device 200. The syntax element information may include at least one of motion vector difference information, a coding block flag, and a transform coefficient level.
[0204] The residual signal generated after intra prediction or inter prediction can be transformed into the frequency domain by a transform process as a part of the quantization process. Here, the first transform can use DCT type 2 (DCT-II) and various DCT, DST cores. These transform cores can perform a separable transform that performs a 1D transform in the horizontal and / or vertical direction on the residual signal, or can perform a 2D non-separable transform on the residual signal.
[0205] For example, in the case of 1D transformation, the DCT and DST types used in the transformation can use DCT-II, DCT-V, DCT-VIII, DST-I and DST-VII as shown in the following table. For example, as shown in Tables 1 and 2, the DCT or DST type used in the transformation by synthesizing the transformation set can be derived.
[0206] [Table 1]
[0207] Transformation Sets Transform 0 DST_VII, DCT-VIII 1 DST-VII, DST-I 2 DST-VII, DCT-V
[0208] [Table 2]
[0209] Transformation Sets Transform 0 DST_VII, DCT-VIII, DST-I 1 DST-VII, DST-I, DCT-VIII 2 DST-VII, DCT-V, DST-I
[0210] For example, Figure 8 As shown, different transform sets are defined for the horizontal direction and the vertical direction according to the intra prediction mode. Next, the encoder / decoder can perform transform and / or inverse transform by using the intra prediction mode of the current encoding / decoding target block and the transform of the relevant transform set. In this case, entropy coding / decoding is not performed on the transform set, and the encoder / decoder can define the transform set according to the same rule. In this case, entropy coding / decoding indicating which transform among the transforms of the transform set is used can be performed. For example, when the size of the block is equal to or less than 64×64, according to the intra prediction mode, three transform sets are synthesized as shown in Table 2, and three transforms are used for each horizontal direction transform and vertical direction transform to combine and perform a total of nine multi-transformation methods. Next, the residual signal is encoded / decoded by using the optimal transform method, thereby, the coding efficiency can be improved. Here, in order to entropy encode / decode information about which transform method is used among the three transforms in a transform set, truncated unary binarization can be used. Here, in order to perform at least one of the vertical transform and the horizontal transform, entropy coding / decoding can be performed on information indicating which transform among the transforms of the transform set is used.
[0211] After completing the above first transformation, Fig. 9As shown in , the encoder may perform a secondary transform on the transform coefficients to improve energy concentration. The secondary transform may perform a separable transform for performing a 1D transform in the horizontal and / or vertical direction, or may perform a 2D inseparable transform. The transform information used may be sent, or may be derived by the encoder / decoder based on current encoding information and neighboring encoding information. For example, as with a 1D transform, a transform set for a secondary transform may be defined. Entropy encoding / decoding is not performed on the transform set, and the encoder / decoder may define the transform set according to the same rules. In this case, information indicating which transform among the transforms of the transform set is used may be sent, and the information may be applied to at least one residual signal through intra-frame prediction or inter-frame prediction.
[0212] At least one of the number or type of transform candidates is different for each transform set. At least one of the number or type of transform candidates may be determined differently based on at least one of the following items: the position, size, partition form, and prediction mode (intra / inter mode) of a block (CU, PU, TU, etc.) or the direction / non-direction of an intra prediction mode.
[0213] The decoder may perform the secondary inverse transform depending on whether the secondary inverse transform is performed, and may perform the first inverse transform depending on whether the first inverse transform is performed from a result of the secondary inverse transform.
[0214] The above-mentioned first transform and second transform may be applied to at least one signal component in the luminance / chrominance component, or may be applied according to the size / shape of any coding block. Entropy encoding / decoding may be performed on an index indicating whether the first transform / secondary transform is used and the first transform / secondary transform used in any coding block. Alternatively, the index may be derived by default by the encoder / decoder according to at least one current / neighboring coding information.
[0215] The residual signal generated after intra-frame prediction or inter-frame prediction is quantized after the first transform and / or the second transform, and the quantized transform coefficients are entropy encoded. Fig.10 As shown in , the quantized transform coefficients may be scanned in a diagonal direction, a vertical direction, and a horizontal direction based on at least one of an intra prediction mode or a size / shape of a minimum block.
[0216] In addition, the quantized transform coefficients on which entropy decoding has been performed may be arranged in a block form by being inversely scanned, and at least one of inverse quantization or inverse transformation may be performed on the relevant block. Here, as a method of inverse scanning, at least one of diagonal scanning, horizontal scanning, and vertical scanning may be performed.
[0217] For example, when the size of the current coding block is 8×8, the residual signal for the 8×8 block may be subjected to a first transform, a second transform, and quantization. Next, Fig.10 At least one of the three scanning order methods shown in the above is used to perform scanning and entropy encoding on the quantized transform coefficients for each of the four 4×4 sub-blocks. In addition, the quantized transform coefficients may be inversely scanned by performing entropy decoding. The quantized transform coefficients on which the inverse scanning is performed become transform coefficients after inverse quantization, and at least one of a secondary inverse transform or a primary inverse transform is performed, whereby a reconstructed residual signal may be generated.
[0218] In the video encoding process, a block can be Fig.11 The partition information may be partitioned as shown, and an indicator corresponding to the partition information may be sent with a signal. Here, the partition information may be at least one of the following items: a partition flag (split_flag), a quad / binary tree flag (QB_flag), a quadtree partition flag (quadtree_flag), a binary tree partition flag (binarytree_flag), and a binary tree partition type flag (Btype_flag). Here, split_flag is a flag indicating whether the block is partitioned, QB_flag is a flag indicating whether the block is partitioned in a quadtree form or a binary tree form, quadtree_flag is a flag indicating whether the block is partitioned in a quadtree form, binarytree_flag is a flag indicating whether the block is partitioned in a binary tree form, and Btype_flag is a flag indicating whether the block is partitioned vertically or horizontally in the case of partitioning in a binary tree form.
[0219] When the partition flag is 1, it may indicate that partitioning is performed, and when the partition flag is 0, it may indicate that partitioning is not performed. In the case of the quad / binary tree flag, 0 may indicate quadtree partitioning, and 1 may indicate binary tree partitioning. Alternatively, 0 may indicate binary tree partitioning, and 1 may indicate quadtree partitioning. In the case of the binary tree partition type flag, 0 may indicate horizontal direction partitioning, and 1 may indicate vertical direction partitioning. Alternatively, 0 may indicate vertical direction partitioning, and 1 may indicate horizontal direction partitioning.
[0220] For example, it may be derived by signaling at least one of quadtree_flag, binarytree_flag, and Btype_flag as shown in Table 3. Fig.11 Partition information.
[0221] [Table 3]
[0222]
[0223] For example, it may be derived by signaling at least one of split_flag, QB_flag, and Btype_flag as shown in Table 4. Fig.11 Partition information.
[0224] [Table 4]
[0225]
[0226] The partitioning method may be performed only in a quadtree form or only in a binary tree form according to the size / shape of the block. In this case, split_flag may mean a flag indicating whether partitioning is performed in a quadtree form or a binary tree form. The size / shape of the block may be derived according to the depth information of the block, and the depth information may be transmitted with a signal.
[0227] When the size of the block is in a predetermined range, partitioning can be performed only in the form of a quadtree. Here, the predetermined range can be defined as at least one of the size of the largest block or the size of the smallest block that can only be partitioned in the form of a quadtree. Information indicating the size of the largest block / smallest block that allows partitioning in the form of a quadtree can be signaled through a bitstream, and the information can be signaled in units of at least one of a sequence, a picture parameter, or a slice (segment). Alternatively, the size of the largest block / smallest block can be a fixed size preset in the encoder / decoder. For example, when the size of the block ranges from 256x256 to 64x64, partitioning can be performed only in the form of a quadtree. In this case, split_flag can be a flag indicating whether partitioning is performed in the form of a quadtree.
[0228] When the size of the block is in a predetermined range, partitioning can be performed only in the form of a binary tree. Here, the predetermined range can be defined as at least one of the size of the largest block or the size of the smallest block that can only be partitioned in the form of a binary tree. Information indicating the size of the largest block / smallest block that allows partitioning in the form of a binary tree can be signaled by a bitstream, and the information can be signaled in units of at least one of a sequence, a picture parameter, or a slice (segment). Alternatively, the size of the largest block / smallest block can be a fixed size preset in an encoder / decoder. For example, when the size of the block ranges from 16x16 to 8x8, partitioning can be performed only in the form of a binary tree. In this case, split_flag can be a flag indicating whether partitioning is performed in the form of a binary tree.
[0229] After partitioning a block in the binary tree form, when the partitioned block is further partitioned, the partitioning may be performed only in the binary tree form.
[0230] When the width size or the length size of the partitioned block cannot be further partitioned, at least one indicator may not be signaled.
[0231] In addition to the quadtree-based binary tree partitioning, quadtree-based partitioning may be performed after the binary tree partitioning.
[0232] Hereinafter, a method for improving video compression efficiency by a transformation method as part of a video encoding process will be described. More specifically, the encoding process of conventional video encoding includes the following steps: an intra / inter prediction step for predicting an original block as part of a current original image; a transformation and quantization step for a residual block as the difference between the predicted prediction block and the original block; entropy coding, i.e., a probability-based lossless compression method for the coefficients of the compressed information obtained from the previous stage and the transformed and quantized blocks. By encoding, a bit stream as a compressed form of the original image is formed, and the bit stream is sent to a decoder or stored in a recording medium. The rearrangement and discrete sine transform (hereinafter referred to as SDST) described here are used to improve the efficiency of the transformation method, thereby improving the compression efficiency.
[0233] The SDST method according to the present invention uses discrete sine transform type 7 or DST-VII (hereinafter referred to as DST-7) instead of discrete cosine transform type 2 or DCT-II (hereinafter referred to as DCT-2), which is a transform kernel widely used in video encoding, so that the common frequency characteristics of the image can be better utilized.
[0234] By using the transformation method according to the present invention, an objectively high-definition video can be obtained at a lower bit rate than that of a traditional video encoding method.
[0235] DST-7 may be applied to the data of the residual block. The operation of applying DST-7 to the residual block may be performed based on the prediction mode corresponding to the residual block. For example, DST-7 may be applied to the residual block encoded in the inter-frame mode. According to an embodiment of the present invention, DST-7 may be applied to the data of the residual block after rearrangement or reordering. Here, reordering means the rearrangement of the image data, and may be referred to as the residual signal rearrangement. Here, the residual block may have the same meaning as the residual, the residual signal, and the residual data. In addition, the residual block may have the same meaning as the reconstructed residual, the reconstructed residual block, the reconstructed residual signal, and the reconstructed residual data, which are the forms after the residual block is reconstructed by the encoder and the decoder.
[0236] According to an embodiment of the present invention, SDST may use DST-7 as a transformation kernel. Here, the transformation kernel of SDST is not limited to DST-7, and may be at least one of various types of DST (such as discrete sine transform type 1 (DST-1), discrete sine transform type 2 (DST-2), discrete sine transform type 3 (DST-3), ..., discrete sine transform type n (DST-n), etc. (here, n is a positive integer equal to or greater than 1)).
[0237] The method of performing one-dimensional DCT-2 according to an embodiment of the present invention can be expressed as the following Formula 1. Here, the block size is designated as N, the frequency component position is designated as k, and the value of the nth coefficient in the spatial domain is designated as x n .
[0238] [Formula 1]
[0239]
[0240] By performing horizontal transform and vertical transform on the residual block via Equation 1, DCT-2 in a two-dimensional domain is possible.
[0241] The DCT-2 transform kernel may be defined as follows: Formula 2. Here, the basis vector according to the position in the frequency domain may be designated as X k , the size of the frequency domain can be specified as N.
[0242] [Formula 2]
[0243]
[0244] at the same time, Fig.12 is a diagram showing basis vectors in the DCT-2 frequency domain according to the present invention. Fig.12 The frequency characteristics of DCT-2 in the frequency domain are shown. Here, the value calculated by the X0 basis vector of DCT-2 may be referred to as a DC component.
[0245] DCT-2 may be used in the transform process for residual blocks of sizes 4×4, 8×8, 16×16, 32×32, and the like.
[0246] Meanwhile, DCT-2 may be selectively used based on at least one of the residual block size, the color component (e.g., luminance component and chrominance component) of the residual block, or the prediction mode corresponding to the residual block. For example, when the component of the residual block of size 4×4 encoded in intra mode is the luminance component, DCT-2 may not be used. Here, the prediction mode may mean inter prediction or intra prediction. In addition, in the case of intra prediction, the prediction mode may mean an intra prediction mode or an intra prediction direction.
[0247] Transformation by the DCT-2 transform kernel can have high compression efficiency in blocks with characteristics of small changes between adjacent pixels (such as the background in an image). However, the DCT-2 transform kernel may not be suitable as a transform kernel for areas with complex patterns (such as textures in an image). When blocks with low correlation between adjacent pixels are transformed by DCT-2, a large number of transform coefficients may appear in the high-frequency components of the frequency domain. In video compression, the frequent occurrence of transform coefficients in high-frequency areas may reduce compression efficiency. In order to improve compression efficiency, coefficients are expected to be large values near low-frequency components, and coefficients are expected to be close to zero in high-frequency components.
[0248] The method of performing one-dimensional DST-7 according to an embodiment of the present invention may be expressed as the following Formula 3. Here, the block size is designated as N, the frequency component position is designated as k, and the value of the nth coefficient in the spatial domain is designated as x n .
[0249] [Formula 3]
[0250]
[0251] By performing horizontal transform and vertical transform on the residual block via Equation 3, DST-7 in the two-dimensional domain is possible.
[0252] The DST-7 transform kernel can be defined as follows in Formula 4. Here, the Kth basis vector of DST-7 is designated as X k , the position in the frequency domain is designated as i, and the size of the frequency domain is designated as N.
[0253] [Formula 4]
[0254]
[0255] DST-7 may be used in a transform process for a residual block having at least one size of 2x2, 4x4, 8x8, 16x16, 32x32, 64x64, 128x128, etc.
[0256] Meanwhile, DST-7 may be applied to rectangular blocks instead of square blocks. For example, DST-7 may be applied to at least one of vertical transformation and horizontal transformation of rectangular blocks having different width and height sizes (e.g., 8×4, 16×8, 32×4, 64×16, etc.).
[0257] In addition, DST-7 may be selectively used based on at least one of the residual block size, the color component (e.g., luminance component and chrominance component) of the residual block, or the prediction mode corresponding to the residual block. For example, when the component of the 4×4 size residual block encoded in the intra mode is the luminance component, DST-7 may be used. Here, the prediction mode may mean inter prediction or intra prediction. In addition, in the case of intra prediction, the prediction mode may mean an intra prediction mode or an intra prediction direction.
[0258] at the same time, Fig.13 : is a diagram showing basis vectors in the DST-7 frequency domain according to the present invention. Fig.13 , the first basis vector (x0) of DST-7 has a curved shape. Therefore, compared with DCT-2, DST-7 can provide higher transformation performance for blocks with large spatial variations in an image.
[0259] DST-7 can be used when performing a transform on a 4×4 transform unit (TU) in an intra-predicted coding unit (CU). Due to the characteristics of intra-prediction, the error rate increases as it moves away from the reference sample, but this is suitable for DST-7, so DST-7 can provide higher transform efficiency. That is, in the spatial domain, when a block has a residual signal that increases as it moves away from the (0, 0) position in the block, the block can be effectively compressed by using DST-7.
[0260] As described above, in order to improve the conversion efficiency, it is important to use a conversion kernel suitable for the frequency characteristics of the image. Specifically, the transformation is performed on the residual block for the original block, so the conversion efficiency of DST-7 and DCT-2 can be identified by identifying the distribution characteristics of the residual signal in the CU or PU or TU.
[0261] Fig.14 is a diagram showing the distribution of average residual values according to positions in a 2Nx2N prediction unit (PU) of an 8x8 coding unit (CU), which is predicted in inter-frame mode by experimenting with a "Cactus" sequence in a low-latency P-profile environment.
[0262] Reference Fig.14 , Fig.14 The left side shows the top 30% of the relatively large values among the average residual signal values in the block, Fig.14 The right side of shows the relatively large top 70% values among the average residual signal values in the same block as the left side.
[0263] exist Fig.14, the residual signal distribution in the 2Nx2N PU of the 8×8 CU predicted in inter mode shows the following characteristics: small residual signal values are mainly concentrated near the center of the block and the residual signal values become larger as they move away from the center of the block. That is, the residual signal value is larger at the block boundary. This distribution characteristic of the residual signal is a common feature of the residual signal in the PU, regardless of the CU size and the PU partition mode (2Nx2N, 2NxN, Nx2N, NxN, nRx2N, nLx2N, 2NxnU, 2NxnD) that enables the CU to be predicted in inter mode.
[0264] Fig.15 is a three-dimensional graph showing distribution characteristics of a residual signal in a 2N×2N prediction unit (PU) of an 8×8 coding unit (CU) predicted in an inter prediction mode (inter mode) according to the present invention.
[0265] Reference Fig.15 , small residual signal values are concentrated near the center of the block, and the residual signal values are relatively larger as they approach the block boundary.
[0266] based on Fig.14 and Fig.15 Taking into account the distribution characteristics of the residual signal in the inter-frame mode, the transformation of the residual signal in the PU of the CU predicted in the inter-frame mode can be more efficient by using DST-7 instead of DCT-2.
[0267] Hereinafter, SDST, which is one of the transformation methods using DST-7 as the transformation kernel, will be described.
[0268] The SDST according to the present invention can be performed in two steps. The first step is to rearrange the residual signal in the PU of the CU predicted in the inter-frame mode or the intra-frame mode. The second step is to apply DST-7 to the rearranged residual signal in the block.
[0269] The residual signal arranged in the current block (e.g., CU, PU, or TU) may be scanned along a first direction and may be rearranged along a second direction. That is, the residual signal arranged in the current block may be scanned along a first direction and may be rearranged along a second direction, so that a rearrangement operation may be performed. Here, the residual signal may mean a signal indicating a residual signal between an original signal and a predicted signal. That is, the residual signal may mean a signal before at least one of transformation and quantization is performed. Alternatively, the residual signal may mean a signal on which at least one of transformation and quantization is performed.
[0270] In addition, the residual signal may mean a reconstructed residual signal. That is, the residual signal may mean a signal that has been subjected to at least one of inverse transformation and inverse quantization. In addition, the residual signal may mean a signal before at least one of inverse transformation and inverse quantization is performed.
[0271] Meanwhile, the first direction (or scanning direction) may be one of a raster scanning order, an upper right diagonal scanning order, a horizontal scanning order, and a vertical scanning order. In addition, the first direction may be defined as follows.
[0272] (1) Scan from top to bottom, and scan from left to right within a row
[0273] (2) Scan from the top row to the bottom row, and scan from the right to the left in a row
[0274] (3) Scan from the bottom row to the top row, and scan from the left to the right in a row
[0275] (4) Scan from the bottom row to the top row, and scan from the right to the left in a row
[0276] (5) Scan from the left column to the right column, and scan from top to bottom within a column
[0277] (6) Scan from the left column to the right column, and scan from the bottom to the top within a column
[0278] (7) Scan from the right column to the left column, and scan from top to bottom within a column
[0279] (8) Scan from the right column to the left column, and scan from the bottom to the top within a column
[0280] (9) Scanning in a spiral shape: Scanning from the inside (or outside) of the block to the outside (or inside) of the block, and scanning in a clockwise / counterclockwise direction
[0281] (10) Diagonal scan: Scan diagonally from one corner of the block to the upper left, upper right, lower left, or lower right.
[0282] Meanwhile, one of the above-mentioned scanning directions may be selectively used as the second direction (or rearrangement direction). The first direction and the second direction may be the same, or may be different from each other.
[0283] The scanning and rearrangement process for the residual signal may be performed on each current block.
[0284] Here, rearranging may mean arranging the residual signal scanned in the block along the first direction in the block of the same size along the second direction. In addition, the size of the block scanned along the first direction may be different from the size of the block rearranged along the second direction.
[0285] In addition, here, scanning and rearranging are performed separately according to the first direction and the second direction, respectively, but scanning and rearranging may be performed as one process with respect to the first direction. For example, the residual signal in the block is scanned from the top row to the bottom row, but in one row, the residual signal is scanned from the right to the left, and the residual signal may be stored (rearranged) in the block.
[0286] At the same time, the scanning and rearrangement process for the residual signal may be performed on each predetermined subblock in the current block. Here, the subblock may be a block having a size equal to or smaller than the current block.
[0287] The sub-block may have a fixed size / shape (e.g., 4x4, 4x8, 8x8, ..., NxM, where N and M are positive integers). In addition, the size and / or shape of the sub-block may be obtained differently. For example, the size and / or shape of the sub-block may be determined according to the size, shape, and / or prediction mode (inter, intra) of the current block.
[0288] The scanning direction and / or rearrangement direction may be adaptively determined according to the position of the sub-block. In this case, each sub-block may use a different scanning direction and / or a different rearrangement direction. Alternatively, all or part of the sub-blocks in the current block may use the same scanning direction and / or the same rearrangement direction.
[0289] Fig.16 is a diagram illustrating distribution characteristics of a residual signal in a 2N×2N prediction unit (PU) mode of a coding unit (CU) according to the present invention.
[0290] Reference Fig.16 , the PU is partitioned into four sub-blocks in a quadtree structure, and the direction of the arrow of each sub-block indicates the distribution characteristics of the residual signal. Specifically, the direction of the arrow of each sub-block indicates the direction in which the residual signal increases. Regardless of the PU partition mode, the residual signal in the PU has a common distribution characteristic. Therefore, in order to have a distribution characteristic suitable for DST-7 transformation, rearrangement may be performed to rearrange the residual signal of each sub-block.
[0291] Fig.17 is a diagram illustrating distribution characteristics of residual signals before and after reordering 2N×2N prediction units (PUs) according to the present invention.
[0292] Reference Fig.17 , the upper block shows the distribution of the residual signal in the 2N×2NPU of the 8×8CU predicted in the inter-frame mode before rearrangement. The following formula 5 indicates that according to Fig.17 The value of each residual signal position in the upper block.
[0293] [Formula 5]
[0294]
[0295] Due to the distribution characteristics of the residual signal in the PU of the CU predicted in the inter-frame mode, the residual signal with a relatively small value is basically distributed in Fig.17 The residual signal having a large value is basically distributed at the boundary of the upper block.
[0296] exist Fig.17 , the lower block shows the distribution of the residual signal in the 2N×2N PU after rearrangement. This shows that the residual signal distribution for each subblock of the PU on which rearrangement is performed is a residual signal distribution suitable for the first basis vector of DST-7. That is, the residual signal in each subblock has a large value when it is far away from the (0, 0) position. Therefore, when the transformation is performed, the transformation coefficient value frequency-transformed by DST-7 may appear concentrically in the low-frequency region.
[0297] The following Formula 6 indicates a rearrangement method according to the position of each subblock in a PU, where the PU is partitioned into four subblocks in a quadtree structure.
[0298] [Formula 6]
[0299]
[0300]
[0301]
[0302]
[0303] 0≤x≤W k , 0≤y≤H k , k∈{blk0, blk1, blk2, blk3}
[0304] Here, the width and height of the k-th subblock (k∈{blk0, blk1, blk2, blk3}) in the PU are designated as Wk and Hk, and the subblocks partitioned from the PU in a quadtree structure are designated as blk0 to blk3. In addition, the horizontal position and vertical position in each subblock are designated as x and y. The position of the residual signal before rearrangement is designated as Fig.17 The changed positions of the residual signal after rearrangement are designated as a(x, y), b(x, y), c(x, y) and d(x, y) shown in the upper block of FIG. Fig.17 a′(x, y), b′(x, y), c′(x, y) and d′(x, y) shown in the lower block of .
[0305] Fig.18 is a diagram showing an example of rearrangement of 4×4 residual data of a subblock according to the present invention.
[0306] Reference Fig.18 , a sub-block may mean one of several sub-blocks in an 8×8 prediction block. Fig.18 (a) shows the location of the original residual data before rearrangement, and Fig.18 (b) shows the rearranged position of the residual data.
[0307] Reference Fig.18 (c), the value of the residual data gradually increases from the position (0, 0) to the position (3, 3). Here, the horizontal and / or vertical one-dimensional residual data in each sub-block may have Fig.13 The data distribution in the type of basis vectors shown in .
[0308] That is, the rearrangement according to the present invention can rearrange the residual data of each sub-block so that the residual data distribution is suitable for the type of DST-7 basis vectors.
[0309] After rearrangement for each sub-block, a DST-7 transform may be applied to the data rearranged for each sub-block.
[0310] At the same time, based on the depth of the TU, the sub-blocks can be partitioned in a quadtree structure, or rearrangement processing can be selectively performed. For example, when the depth of the TU is 2, the N×N sub-blocks in the 2Nx2N PU can be partitioned into N / 2×N / 2 blocks, and the rearrangement processing can be applied to each N / 2×N / 2 block. Here, the quadtree-based TU partitioning can be performed continuously until the minimum TU size is reached.
[0311] Furthermore, when the depth of the TU is zero, the DCT-2 transform may be applied to the 2N×2N block. Here, rearrangement of residual data may not be performed.
[0312] Meanwhile, the SDST method according to the present invention uses the distribution characteristics of the residual signal in the PU block, so that the partition structure of the TU performing SDST can be defined as partitioned in a quadtree structure based on the PU.
[0313] Fig.19a and Fig.19b is a diagram illustrating a partition structure of a transform unit (TU) according to a prediction unit (PU) mode of a coding unit (CU) and a rearrangement method of a transform unit (TU) according to the present invention. Fig.19a and Fig.19b The quadtree partition structure of a TU according to the depth of the TU for each of the asymmetric partition modes (2NxnU, 2NxnD, nRx2N, nLx2N) of an inter-predicted PU is shown.
[0314] Reference Fig.19a and Fig.19b , the thick line of each block indicates the PU in the CU, and the thin line indicates the TU. In addition, S0, S1, S2, and S3 in the TU indicate the rearrangement method of the residual signal in the TU defined in the formula.
[0315] At the same time, Fig.19a and Fig.19b In the PU, the TU with depth zero in each PU has the same block size as the PU (for example, in a 2Nx2N PU, the size of the TU with depth zero is the same as the size of the PU). Fig.23a , Figure 23b and Fig.23c A reordering operation for a residual signal in a TU with depth zero is disclosed.
[0316] In addition, when at least one of the CU, PU and TU has a rectangular shape (e.g., 2NxnU, 2NxnD, nRx2N and nLx2N), at least one of the CU, PU and TU is partitioned into N sub-blocks (such as 2, 4, 6, 8, 16 sub-blocks, etc.) before rearranging the residual signal, and the rearrangement of the residual signal can be applied to the partitioned sub-blocks.
[0317] In addition, when at least one of the CU, PU and TU has a square shape (e.g., 2Nx2N and NxN), at least one of the CU, PU and TU is partitioned into N sub-blocks (such as 4, 8, 16 sub-blocks, etc.) before rearranging the residual signal, and the rearrangement of the residual signal can be applied to the partitioned sub-blocks.
[0318] In addition, when a TU is partitioned from a CU or PU and the TU has the highest depth (cannot be partitioned), the TU may be partitioned into N subblocks, such as 2, 4, 6, 8, 16 subblocks, etc., and rearrangement of the residual signal may be applied to the partitioned subblocks.
[0319] The above example shows that the rearrangement of the residual signal is performed when the CU, PU, and TU have different shapes or sizes. However, even when at least two of the CU, PU, and TU have the same shape or size, the rearrangement of the residual signal may be applied.
[0320] at the same time, Fig.19a and Fig.19b An asymmetric partition mode of an inter-predicted PU is shown, but the asymmetric partition mode is not limited thereto. The partitioning and rearrangement operations of a TU may be applied to a symmetric partition mode (2NxN and Nx2N) of a PU.
[0321] The DST-7 transform may be performed on each TU in the PU on which the rearrangement is performed. Here, when the CU, the PU, and the TU have the same size and shape, the DST-7 transform may be performed on one block.
[0322] When considering the distribution characteristics of the residual signal of the inter-predicted PU block, performing DST-7 transform after rearrangement is a more efficient transform method rather than performing DCT-2 transform regardless of the size of the CU and the PU partition mode.
[0323] The fact that after the transformation, the transform coefficients are basically distributed near the low-frequency components (especially the DC component) means: compared with the opposite case of the residual signal distribution, i) the energy loss after quantization is minimized, and ii) there is a higher compression efficiency in the entropy coding process in terms of reduced bit usage.
[0324] Fig. 20 is a diagram illustrating a result of performing DCT-2 transform and SDST transform based on a residual signal distribution of a 2N×2N prediction unit (PU) according to the present invention.
[0325] Fig. 20 The left side shows the distribution of the residual signal increasing from the center to the boundary when the PU partition mode of the CU is 2N×2N. In addition, Fig. 20 The middle shows the distribution of the residual signal after DCT-2 transform is performed on the TU with depth 1 in the PU. Fig. 20 The right side of shows the distribution of the residual signal after DST-7 transform (SDST) is performed on the TU with depth 1 in the PU after rearrangement.
[0326] Reference Fig. 20 When SDST is performed on a TU of a PU having a distribution characteristic of a residual signal, a larger number of coefficients are concentrated near low-frequency components compared to performing DCT-2. Smaller coefficient values appear on the high-frequency component side. When the residual signal of an inter-predicted PU is transformed based on such a transformation characteristic, higher compression efficiency can be obtained by performing SDST instead of DCT-2.
[0327] SDST is performed on a block, which is a TU defined in a PU and on which a DST-7 transform is performed. The partition structure of a TU is as follows Fig.19a and Fig.19b The quadtree structure or binary tree structure from the PU size to the maximum depth is shown. This means that after reordering, the DST-7 transform can be performed on square blocks, and the DST-7 transform can also be performed on rectangular blocks.
[0328] Fig.2121 is a diagram showing SDST processing according to the present invention. First, when the residual signal in the TU is transformed in step S2110, the partitioned TU in the PU whose prediction mode is the inter-frame mode may be rearranged in step S2120. Next, DST-7 transformation is performed on the rearranged TU in step S2130, and quantization may be performed in step S2140, and a series of subsequent steps may be performed.
[0329] Meanwhile, rearrangement and DST-7 transform may be performed on a block whose prediction mode is in an intra mode.
[0330] Hereinafter, the following methods will be disclosed as embodiments for implementing SDST in an encoder: i) a method of performing SDST on all TUs in an inter-predicted PU, and ii) a method of selectively performing SDST or DCT-2 through rate-distortion optimization. The following method is described for an inter-prediction block, but is not limited thereto, and the following method may be applied to an intra-prediction block.
[0331] Fig. 22 is a diagram illustrating distribution characteristics of a residual absolute value and a transform unit (TU) partition mode of a coding unit (CU) based on inter-prediction according to the present invention.
[0332] Reference Fig. 22 , in the inter prediction mode, the CU may be partitioned into TUs in a quadtree structure or a binary tree structure up to a maximum depth, and the number of partition modes of the PU may be K. Here, K is a positive integer, and Fig. 22 The K in is 8.
[0333] The SDST according to the present invention uses Fig.15 The distribution characteristics of the residual signal of the PU in the inter-predicted CU described in . In addition, the PU may be partitioned into TUs in a quadtree structure or a binary tree structure. A TU with a depth of zero may correspond to the PU, and a TU with a depth of 1 may correspond to each sub-block partitioned from the PU in a quadtree structure or a binary tree structure.
[0334] Fig. 22 Each block of the CU is partitioned into TUs with a depth of 2 according to each PU partition mode of the inter-frame predicted CU. Here, the thick line indicates the PU and the thin line indicates the TU. The direction of the arrow of each TU indicates the direction in which the value of the residual signal in the TU increases. Each TU can be reordered according to the position of each TU in the PU.
[0335] Specifically, when a TU has a depth of zero, reordering may be performed in various methods in addition to the above-described method for reordering.
[0336] One of the methods is to scan from a residual signal at the center of the PU, scan adjacent residual signals in a spiral direction, and rearrange the scanned residual signals starting from a (0, 0) position of the PU in a zigzag scanning order.
[0337] Fig.23a , Figure 23b and Fig.23c is a diagram illustrating a scanning order and a rearrangement order of a residual signal for a transform unit (TU) having a depth of zero in a prediction unit (PU).
[0338] Fig.23a and Figure 23b shows the scan order for rearrangement, and Fig.23c The rearrangement sequence for SDST is shown.
[0339] The DST-7 transform is performed on the rearranged residual signal in each TU, and quantization and entropy encoding, etc. may be performed thereon. This rearrangement method uses the distribution characteristics of the residual signal in the TU according to the PU partition mode. The rearrangement method can optimize the distribution of the residual signal so as to improve the efficiency of the DST-7 transform as a subsequent step.
[0340] In the encoder, according to Fig.21 The SDST process in performs SDST on all TUs in the inter-frame predicted PU. According to the PU partition mode of the inter-frame predicted CU, such as Fig. 22 As shown, a PU can be partitioned into TUs up to a maximum depth of 2. Fig. 22 The residual signal in each TU is rearranged according to the distribution characteristics of the residual signal in the TU in the DST-7 transform kernel. Next, quantization and entropy encoding, etc. may be performed after transformation using the DST-7 transform kernel.
[0341] In the decoder, when reconstructing the residual signal of the TU in the inter-frame predicted PU, the DST-7 inverse transform is performed on each TU of the inter-frame predicted PU. The reconstructed residual signal can be obtained by inversely rearranging the reconstructed residual signal. According to the SDST method, SDST is applied to the transform method of all TUs in the inter-frame predicted PU, so that there is no flag or information to be sent to the decoder. That is, the SDST method can be performed without using a signal for the SDST method.
[0342] Meanwhile, even when SDST is performed on all TUs in an inter-predicted PU, the encoder determines a part of the above-mentioned residual signal rearrangement method regarding the rearrangement operation as an optimal rearrangement method. Information regarding the determined rearrangement method may be transmitted to the decoder.
[0343] As another embodiment for implementing SDST, a transform method for applying a PU by using one of DCT-2 and SDST via RDO will be disclosed. Compared with the previous embodiment in which SDST is performed on all TUs in the inter-predicted PU, the amount of calculation of the encoder is increased according to this method. However, a more efficient transform method is selected from DCT-2 and SDST, so that higher compression efficiency can be obtained than in the previous embodiment.
[0344] Fig.24 is a flow chart illustrating an encoding process of selecting DCT-2 or SDST through rate-distortion optimization (RDO) according to the present invention.
[0345] Reference Fig.24 , in step S2410, the residual signal of the TU is transformed, the cost of the TU obtained by performing DCT-2 on each TU in the PU performing prediction in the inter-frame mode in step S2420 is compared with the cost of the TU obtained by performing SDST in steps S2430 and S2440, and the optimal transformation mode (DCT-7 or SDST) of the TU can be determined according to the rate distortion in step S2450. Next, in step S2460, quantization and entropy encoding, etc. can be performed on the transformed TU according to the determined transformation mode.
[0346] Meanwhile, the optimal transform mode may be selected by RDO only when a TU to which SDST or DCT-2 is applied satisfies one of the following conditions.
[0347] i) Regardless of the PU partition mode, a TU performing DCT-2 and SDST is partitioned in a quadtree structure or a binary tree structure based on a CU or a CU size.
[0348] ii) A TU performing DCT-2 and SDST is partitioned from a PU in a quadtree structure or a binary tree structure according to a PU partition mode or a PU size.
[0349] iii) Regardless of the PU partition mode, a TU performing DCT-2 and SDST is not partitioned on a CU basis.
[0350] Condition i) is a method of selecting DCT-2 or SDST as a transform mode for a TU partitioned from a CU in a quadtree structure or a binary tree structure or a CU size according to rate-distortion optimization regardless of a PU partition mode.
[0351] Condition ii) is: according to the PU partition mode described in the embodiment for performing SDST on all TUs in the inter-frame predicted PU, DCT-2 and SDST are performed on the TU partitioned in a quadtree structure or a binary tree structure or by PU size, and the transformation mode of the TU is determined by using the cost.
[0352] Condition iii) is that regardless of the PU partition mode, in a case where the CU is not partitioned or the TU has the same size as the CU, DCT-2 and SDST are performed, and a transform mode of the TU is determined.
[0353] When comparing RD costs for TUs with depth zero in a specific PU partition mode, the cost of a result of performing SDST on the TU with depth zero is compared with the cost of a result of performing DCT-2 on the TU with depth zero. A transform mode for TUs with depth zero may be selected.
[0354] Fig.25 is a flowchart illustrating a decoding process of selecting DCT-2 or SDST according to the present invention.
[0355] Reference Fig.25 In step S2510, the transmitted SDST flag may be referred to for each TU. Here, the SDST flag may be a flag indicating whether SDST is used as a transform mode.
[0356] When the SDST flag is true (i.e., "yes" in step S2520), in step S2530, the SDST mode is determined as the transform mode of the TU, and the DST-7 inverse transform is performed on the residual signal in the TU. In step S2540, the residual signal in the TU on which the DST-7 inverse transform is performed is inversely rearranged using the above formula 6 according to the position of the TU in the PU, and thus, in step S2560, a reconstructed residual signal can be obtained.
[0357] At the same time, when the SDST flag is false (i.e., "No" in step S2520), in step S2550, the DCT-2 mode is determined as the transformation mode of the TU, and the DCT-2 inverse transform is performed on the residual signal in the TU, and therefore, in step S2560, a reconstructed residual signal can be obtained.
[0358] When the SDST method is used, the residual data may be rearranged. Here, the residual data may mean residual data corresponding to the inter-predicted PU. An integer transform derived from DST-7 by using a separable property may be used as the SDST method.
[0359] Meanwhile, sdst_flag may be signaled to selectively use DCT-2 or DST-7. sdst_flag may be signaled for each TU. sdst_flag is used to identify whether SDST is performed.
[0360] Fig.26 is a flow chart illustrating a decoding process using SDST according to the present invention.
[0361] Reference Fig.26 , in step S2610, for each TU, sdst_flag may be signaled and may be entropy decoded.
[0362] First, when the depth of a TU is zero (Yes at step S2620), the TU may be reconstructed by using DCT-2 instead of SDST at steps S2670 and S2680. SDST may be performed on TUs having depths from 1 to a maximum.
[0363] In addition, even if the depth of the TU is not zero ("No" in step S2620), when the transform mode of the TU is the transform skip mode and / or when the value of the coded block flag (cbf) of the TU is zero ("Yes" in step S2630), the TU can be reconstructed without performing an inverse transform in step S2680.
[0364] Meanwhile, when the depth of the TU is not zero (No in step S2620 ) and the transform mode of the TU is not the transform skip mode and the value of the cbf of the TU is not zero (No in step S2630 ), in step S2640 , the value of sdst_flag may be identified.
[0365] Here, when the value of sdst_flag is 1 ("Yes" in step S2640), an inverse transform based on DST-7 may be performed in step S2650, and inverse rearrangement is performed on the residual data of the TU in step S2660, and the TU may be reconstructed in step S2680. Conversely, when the value of sdst_flag is not 1 ("No" in step S2640), an inverse transform based on DCT-2 may be performed in step S2670, and the TU may be reconstructed in step S2680.
[0366] Here, the rearranged or re-arranged target signal may be at least one of: a residual signal before inverse transformation, a residual signal before inverse quantization, a residual signal after inverse transformation, a residual signal after inverse quantization, a reconstructed residual signal, and a reconstructed block signal.
[0367] At the same time, Fig.26 In the embodiment, sdst_flag is signaled for each TU, but sdst_flag may be selectively signaled based on at least one of a transform mode of the TU or a value of a cbf of the TU. For example, when the transform mode of the TU is a transform skip mode and / or when the value of the cbf of the TU is zero, sdst_flag may not be signaled. In addition, when the depth of the TU is zero, sdst_flag may not be signaled.
[0368] Meanwhile, sdst_flag is signaled for each TU, but sdst_flag may also be signaled for a predetermined unit. For example, sdst_flag may be signaled for at least one of the following items: video, sequence, picture, slice, tile, coding tree unit, coding unit, prediction unit, and transform unit.
[0369] As in Fig.25 The SDST logo and Fig.25 As in the embodiment of sdst_flag of , the selected transform mode information can be entropy encoded / decoded for each TU through an n-bit flag (n is a positive integer equal to or greater than 1). The transform mode information may indicate at least one of the following items: whether to perform transform on the TU through DCT-2, SDST, DST-7, etc.
[0370] Only in the case of TUs in inter-predicted PUs, the transform mode information may be entropy encoded / decoded in bypass mode. In addition, when the transform mode is at least one of a transform skip mode, an RDPCM (residual differential PCM) mode, or a lossless mode, the transform mode information may not be entropy encoded / decoded and may not be signaled.
[0371] In addition, when the value of the coded block flag of the block is zero, the transform mode information may not be entropy encoded / decoded and may not be signaled. When the value of the coded block flag is zero, the inverse transform process is omitted in the decoder. Therefore, even if the transform mode information does not exist in the decoder, the block can be reconstructed.
[0372] However, the transformation mode information is not limited to indicating the transformation mode through a flag, and may be implemented as a predefined table and an index. Here, the transformation mode available for each index may be defined as the predefined table.
[0373] Furthermore, DCT or SDST may be performed in the horizontal direction and the vertical direction, respectively. The same transform mode may be used in the horizontal direction and the vertical direction, or different transform modes may be used.
[0374] Also, transform mode information regarding whether DCT-2, SDST, and DST-7 are used in horizontal and vertical directions may be entropy encoded / decoded, respectively.
[0375] Also, transform mode information may be entropy encoded / decoded in at least one of a CU, a PU, and a TU.
[0376] In addition, the transform mode information may be transmitted according to the luminance component or the chrominance component. That is, the transform mode information may be transmitted according to the Y component or the Cb component or the Cr component. For example, when the transform mode information as to whether DCT-2 is performed on the Y component or SDST is performed on the Y component is signaled, the transform mode information signaled in the Y component without signaling the transform mode information in at least one of the Cb component and the Cr component may be used as the transform mode of the block.
[0377] Here, the transformation mode information may be entropy encoded / decoded by arithmetic coding method using a context model. When the transformation mode information is implemented as a predefined table and index, all or part of several binary bits may be entropy encoded / decoded by arithmetic coding method using a context model.
[0378] In addition, the transform mode information may be selectively entropy encoded / decoded according to the block size. For example, when the size of the current block is equal to or greater than 64×64, the transform mode information may not be entropy encoded / decoded. When the size of the current block is equal to or less than 32×32, the transform mode information may be entropy encoded / decoded.
[0379] In addition, when there is a non-zero transform coefficient or a quantization level in the current block, the transform mode information may not be entropy encoded / decoded, and at least one method of DCT-2, DST-7, or SDST may be performed on the transform mode information. Here, regardless of the position of the non-zero transform coefficient or quantization level in the block, the transform mode information may not be entropy encoded / decoded. In addition, the transform mode information may not be entropy encoded / decoded only when there is a non-zero transform coefficient or quantization level in the upper left of the block.
[0380] In addition, when there are J or more non-zero transform coefficients or quantization levels in the current block, the transform mode information may be entropy encoded / decoded. Here, J is a positive integer.
[0381] Furthermore, the transform mode information may vary due to some transform modes being restrictively used according to the transform mode of the co-located block, or may vary due to a binarization method for indicating the transform mode of the co-located block as transform information of fewer bits.
[0382] SDST may be restrictively used based on at least one of a prediction mode of a current block, a depth, a size, and a shape of a TU.
[0383] For example, when the current block is encoded in the inter mode, SDST may be used.
[0384] A minimum / maximum depth for allowing SDST may be defined. In this case, SDST may be used when the depth of the current block is equal to or greater than the minimum depth. Alternatively, SDST may be used when the depth of the current block is equal to or greater than the maximum depth. Here, the minimum / maximum depth may be a fixed value or may be determined differently based on information indicating the minimum / maximum depth. The information indicating the minimum / maximum depth may be signaled from the encoder and may be derived by the decoder based on properties of the current block / neighboring block (e.g., size, depth, and / or shape).
[0385] A minimum / maximum size for allowing SDST may be defined. Similarly, SDST may be used when the size of the current block is equal to or greater than the minimum size. Alternatively, SDST may be used when the size of the current block is equal to or greater than the maximum size. Here, the minimum / maximum size may be a fixed value, or may be determined differently based on information indicating the minimum / maximum size. Information indicating the minimum / maximum size may be signaled from the encoder and may be derived by the decoder based on properties of the current block / neighboring block (e.g., size, depth, and / or shape). For example, when the size of the current block is 4×4, DCT-2 may be used as a transform method, and transform mode information regarding whether DCT-2 or SDST is used may not be entropy encoded / decoded.
[0386] The shape of a block for allowing SDST may be defined. In this case, SDST may be used when the shape of the current block is the defined shape of the block. Alternatively, the shape of a block for not allowing SDST may be defined. In this case, SDST may not be used when the shape of the current block is the defined shape of the block. The shape of a block for allowing or not allowing SDST may be fixed, and information on the shape of a block for allowing or not allowing SDST may be signaled from an encoder. Alternatively, information on the shape of the current block may be derived by a decoder based on properties of the current block / neighboring block (e.g., size, depth, and / or shape). The shape of a block for allowing or not allowing SDST may mean, for example, M, N, and / or a ratio of M to N in an MxN block.
[0387] In addition, when the depth of the TU is zero, DCT-2 or DST-7 may be used as a transform method, and transform mode information about which transform method is used may be entropy encoded / decoded. When DST-7 is used as a transform method, rearrangement of the residual signal may be performed. In addition, when the depth of the TU is equal to or greater than 1, DCT-2 or SDST may be used as a transform method, and transform mode information about which transform method is used may be entropy encoded / decoded.
[0388] Furthermore, a transform method may be selectively used according to partition shapes of CU and PU or the shape of a current block.
[0389] According to an embodiment, when the partition shape of a CU and a PU or the shape of a current block is 2N×2N, DCT-2 may be used and DCT-2 or SDST may be selectively used for the remaining partitions and block shapes.
[0390] Also, when the partition shape of the CU and PU or the shape of the current block is 2NxN or Nx2N, DCT-2 may be used and DCT-2 or SDST may be selectively used for the remaining partitions and block shapes.
[0391] Also, when the partition shape of the CU and PU or the shape of the current block is nRx2N or nLx2N or 2NxnU or 2NxnD, DCT-2 may be used and DCT-2 or SDST may be selectively used for the remaining partitions and block shapes.
[0392] At the same time, SDST or DST-7 is performed on each block partitioned from the current block, and scanning and inverse scanning for transform coefficients (quantization levels) may be performed on each partitioned block. In addition, SDST or DST-7 is performed on each block partitioned from the current block, and scanning and inverse scanning for transform coefficients (quantization levels) may be performed on each unpartitioned current block.
[0393] Also, transformation / inverse transformation using SDST or DST-7 may be performed according to at least one of an intra prediction mode (direction) of a current block, a size of a current block, and a component (luminance component or chrominance component) of the current block.
[0394] Furthermore, when transform / inverse transform is performed using SDST or DST-7, DST-1 may be used instead of DST-7. Furthermore, when transform / inverse transform is performed using SDST or DST-7, DCT-4 may be used instead of DST-7.
[0395] In addition, when transform / inverse transform is performed using DCT-2, the arrangement method used when arranging the residual signal of SDST or DST-7 can be applied. That is, even when DCT-2 is used, the residual signal can be rearranged or rotated by a predetermined angle.
[0396] Hereinafter, various modifications and embodiments for the rearrangement method and the signal transmission method will be disclosed.
[0397] The SDST of the present invention is used to improve image compression efficiency by changing the transformation method. By rearranging the residual signal, the distribution characteristics of the residual signal in the PU are effectively applied when performing DST-7, so high compression efficiency can be obtained.
[0398] In the above description of rearrangement, a method for rearranging a residual signal has been disclosed. In the following, in addition to the rearrangement operation, other embodiments of the method for rearranging a residual signal will be disclosed.
[0399] In order to minimize the hardware complexity for rearranging the residual signal, the rearranging method of the residual signal may be implemented by horizontal flipping and vertical flipping methods. The rearranging methods (1) to (4) of the residual signal may be implemented by the following flipping operation.
[0400] (1): r'(x, y) = r(x, y); no flipping is performed (no flipping)
[0401] (2): r'(x, y) = r(w-1-x, y); horizontal flip
[0402] (3): r'(x, y) = r(x, h-1-y); vertical flip
[0403] (4): r'(x, y) = r(w-1-x, h-1-y); flip horizontally and vertically
[0404] r'(x, y) is the residual signal after rearrangement, and r(x, y) is the residual signal before rearrangement. The width and height of the block are designated as w and h, respectively. The position of the residual signal in the block is designated as x, y. The inverse rearrangement method of the rearrangement method using flipping can be performed in the same process of the rearrangement method. That is, the rearranged residual signal using horizontal flipping is further flipped in the horizontal direction, so that the original arrangement of the residual signal can be reconstructed. The rearrangement method performed in the encoder and the inverse rearrangement method performed in the decoder can use the same flipping method.
[0405] The residual signal rearrangement / rearrangement method using the flip operation can use the current block without partitioning. That is, in the SDST method, the current block (TU, etc.) is partitioned into sub-blocks and DST-7 is used for each sub-block. However, in the residual signal rearrangement / rearrangement method using the flip operation, flipping can be performed on all or part of the current block without partitioning the current block into sub-blocks, and then DST-7 can be used.
[0406] Information about whether to use a residual signal rearrangement / rearrangement method using a flip operation may be entropy encoded / decoded by using transform mode information. For example, when a flag bit indicating transform mode information has a first value, a residual signal rearrangement / rearrangement method using a flip operation and DST-7 may be used as a transform / inverse transform method. When the flag bit has a second value, another transform / inverse transform method may be used instead of using a residual signal rearrangement / rearrangement method using a flip operation. Here, the transform mode information may be entropy encoded / decoded for each block.
[0407] In addition, at least one of the four flipping methods (no flipping, horizontal flipping, vertical flipping, horizontal and vertical flipping) can be entropy encoded / decoded as a flag or an index by using the flipping method information. That is, the flipping method performed in the encoder by signaling the flipping method information can be performed in the decoder. The transform mode information may include the flipping method information.
[0408] In addition, the rearrangement method of the residual signal is not limited to the rearrangement of the residual signal described previously, and the rearrangement operation can be achieved by rotating the residual signal in the block at a predetermined angle. Here, the predetermined angle may mean zero degree, 90 degrees, 180 degrees, -90 degrees, -180 degrees, 270 degrees, -270 degrees, 45 degrees, -45 degrees, 135 degrees, -135 degrees, etc. Here, the information about the angle may be entropy encoded / decoded as a flag or an index, and may be performed similarly to the signal transmission method for transform mode information (transform mode).
[0409] In addition, when entropy encoding / decoding is performed, angle information may be predictively encoded / decoded from angle information of a reconstructed block adjacent to the current block. When rearrangement is performed by using angle information, SDST or DST-7 may be performed after partitioning the current block. Alternatively, SDST or DST-7 may be performed on each current block without partitioning the current block.
[0410] The predetermined angle may be determined differently according to the position of the sub-block. A rearrangement method of rotating only one sub-block (e.g., the first sub-block) at a specific position among the plurality of sub-blocks may be used restrictively. In addition, the rearrangement using the predetermined angle may be applied to the entire current block. Here, the rearranged target current block may be at least one of the following: a residual block before inverse transformation, a residual block before dequantization, a residual block after inverse transformation, a residual block after dequantization, a reconstructed residual block, and a reconstructed block.
[0411] At the same time, in order to obtain the same effect as the rearrangement or rotation of the residual signal, the coefficients of the transformation matrix used for the transformation may be rearranged or rotated, and the coefficients may be applied to the pre-arranged residual signal, thereby performing the transformation. That is, the transformation is performed by using the rearrangement of the transformation matrix instead of the rearrangement of the residual signal, so that the same effect as the method of performing the residual signal rearrangement and transformation can be obtained. Here, the rearrangement of the coefficients of the transformation matrix may be performed in the same manner as the residual signal rearrangement method. Therefore, the signal transmission method for the information of the rearrangement of the coefficients of the transformation matrix may be performed in the same manner as the signal transmission method for the information of the residual signal rearrangement method.
[0412] At the same time, the encoder may determine a part of the residual signal rearrangement method described in the above rearrangement operation description as the optimal rearrangement method, and may send information about the determined rearrangement method to the decoder. For example, when four rearrangement methods are used, the encoder may send information about the residual signal rearrangement method to the decoder in 2 bits.
[0413] In addition, when the rearrangement methods have different occurrence probabilities, the rearrangement methods with high occurrence probabilities can be encoded by using a small number of bits, and the rearrangement methods with low occurrence probabilities can be encoded by using a relatively large number of bits. For example, four rearrangement methods can be encoded as truncated unary codes (0, 10, 110, 111) in order of higher occurrence probabilities.
[0414] In addition, the probability of occurrence of the rearrangement method can be changed according to the coding parameters (such as the prediction mode of the current CU, the intra-frame prediction mode (direction) of the PU, the motion vector of the neighboring block, etc.). Therefore, the coding method of the information about the rearrangement method can be used differently according to the coding parameters. For example, the probability of occurrence of the rearrangement method may vary according to the prediction mode of the intra-frame prediction. Therefore, for each intra-frame mode, a small number of bits are allocated to the rearrangement method with a high probability of occurrence, and a large number of bits are allocated to the rearrangement method with a low probability of occurrence. In some cases, a rearrangement method with an extremely low probability of occurrence may not be used, and bits may not be allocated to the rearrangement method with an extremely low probability of occurrence.
[0415] The following Table 5 shows an example of encoding a residual signal rearrangement method according to a prediction mode of a CU and an intra prediction mode (direction) of a PU.
[0416] [Table 5]
[0417]
[0418] The residual signal rearrangement methods (1) to (4) in Table 5 may specify a residual signal rearrangement method, such as an index for a scanning / rearrangement order for rearranging a residual signal, an index for a predetermined angle value, an index for a predetermined flipping method, etc.
[0419] As shown in Table 5, when the current block is associated with at least one of a prediction mode and an intra prediction mode (direction), at least one rearrangement method may be used in the encoder and the decoder.
[0420] For example, when the current block is in intra mode and the intra prediction direction is an even number, at least one of a non-flipping method, a horizontal flipping method, and a vertical flipping method may be used as a residual signal rearrangement method. In addition, when the current block is in intra mode and the intra prediction direction is an odd number, at least one of a non-flipping method, a vertical flipping method, and a horizontal and vertical flipping method may be used as a residual signal rearrangement method.
[0421] In case of planar / DC prediction of intra prediction, information about four rearrangement methods may be entropy encoded / decoded as truncated unary codes based on the occurrence frequencies of the four rearrangement methods.
[0422] When the intra prediction direction is horizontal or close to horizontal mode, the probability of rearrangement methods (1) and (3) may be high. In this case, 1 bit is used for each of the two rearrangement methods, and information about the rearrangement method can be entropy encoded / decoded.
[0423] When the intra prediction direction is vertical or close to the vertical mode, the probability of rearrangement methods (1) and (2) may be high. In this case, 1 bit is used for each of the two rearrangement methods, and information about the rearrangement method can be entropy encoded / decoded.
[0424] When the intra prediction direction is an even number, information about the rearrangement methods (1), (2), and (3) may be entropy encoded / decoded as a truncated unary code.
[0425] When the intra prediction direction is an odd number, information about the rearrangement methods (1), (3), and (4) may be entropy encoded / decoded as a truncated unary code.
[0426] In other intra prediction directions, the occurrence probability of the rearrangement method (4) may be low. Therefore, information about the rearrangement methods (1), (2), and (3) may be entropy encoded / decoded as truncated unary codes.
[0427] In the case of inter-frame prediction, rearrangement methods (1) to (4) have the same probability of occurrence, and information about the rearrangement method can be entropy encoded / decoded as a 2-bit fixed-length code.
[0428] Here, each coded bit value may be arithmetically coded / decoded. In addition, each coded bit value may be entropy coded / decoded in a bypass without using arithmetic coding.
[0429] Fig. 27 and Fig.28 is a diagram showing where residual signal rearrangement (residual rearrangement) is performed in an encoder and a decoder according to the present invention.
[0430] Reference Fig. 27 , in the encoder, the residual signal rearrangement may be performed before the DST-7 transform. Fig. 27 Not shown in FIG. 1 , in the encoder, residual signal rearrangement may be performed between transform and quantization, and residual signal rearrangement may be performed after quantization.
[0431] Reference Fig.28 , in the decoder, residual signal rearrangement may be performed after the inverse DST-7 transform. Fig.28 Not shown in FIG. 1 , in the decoder, residual signal rearrangement may be performed between dequantization and inverse transform, and residual signal rearrangement may be performed before dequantization.
[0432] Referenced above Figures 12 to 28 The SDST method according to the present invention is described. Fig.29 and Fig.30 A decoding method, an encoding method, a decoder, an encoder, and a bit stream to which the SDST method is applied according to the present invention are described in detail.
[0433] Fig.29 is a diagram illustrating a decoding method using the SDST method according to the present invention.
[0434] Reference Fig.29 First, in step S2910, a transformation mode of the current block may be determined. In step S2920, residual data of the current block may be inversely transformed according to the transformation mode of the current block.
[0435] Furthermore, in step S2930, the inverse-transformed residual data of the current block may be rearranged according to the transformation mode of the current block.
[0436] Here, the transform mode may include at least one of SDST (shuffled discrete sine transform), SDCT (shuffled discrete cosine transform), DST (discrete sine transform), or DCT (discrete cosine transform).
[0437] In the SDST mode, inverse transformation may be performed in the DST-7 transformation mode, and a mode for performing rearrangement on inversely transformed residual data may be commanded.
[0438] In the SDCT mode, inverse transform may be performed in the DCT-2 transform mode, and a mode for performing rearrangement on inverse-transformed residual data may be commanded.
[0439] In the DST mode, inverse transformation may be performed in the DST-7 transform mode, and a mode for not performing rearrangement on inverse-transformed residual data may be commanded.
[0440] In the DCT mode, inverse transform may be performed in the DCT-2 transform mode, and a mode for not performing rearrangement on inverse-transformed residual data may be commanded.
[0441] Therefore, the rearrangement of the residual data may be performed only when the transform mode of the current block is one of SDST and SDCT.
[0442] For the SDST mode and the DST mode, inverse transformation may be performed in the DST-7 transformation mode, but a transformation mode based on another DST such as DST-1, DST-2, etc. may be used.
[0443] Meanwhile, the step of determining the transformation mode of the current block in step S2910 may include: obtaining transformation mode information of the current block from a bitstream; and determining the transformation mode of the current block based on the transformation mode information.
[0444] Also, the step of determining the transform mode of the current block at step S2910 may be performed based on at least one of the prediction mode of the current block, depth information of the current block, the size of the current block, and the shape of the current block.
[0445] Specifically, when the prediction mode of the current block is the inter prediction mode, the transform mode of the current block may be determined as one of SDST and SDCT.
[0446] Meanwhile, the step of rearranging the inversely transformed residual data of the current block at step S2930 may include: scanning the inversely transformed residual data arranged in the current block in a first direction order; and rearranging the residual data scanned along the first direction in the current block in a second direction order. Here, the first direction order may be one of a raster scanning order, an upper right diagonal scanning order, a horizontal scanning order, and a vertical scanning order. In addition, the first direction order may be defined as follows.
[0447] (1) Scan from top to bottom, and scan from left to right within a row
[0448] (2) Scan from the top row to the bottom row, and scan from the right to the left in a row
[0449] (3) Scan from the bottom row to the top row, and scan from the left to the right in a row
[0450] (4) Scan from the bottom row to the top row, and scan from the right to the left in a row
[0451] (5) Scan from the left column to the right column, and scan from top to bottom within a column
[0452] (6) Scan from the left column to the right column, and scan from the bottom to the top within a column
[0453] (7) Scan from the right column to the left column, and scan from top to bottom within a column
[0454] (8) Scan from the right column to the left column, and scan from the bottom to the top within a column
[0455] (9) Scanning in a spiral shape: Scanning from the inside (or outside) of the block to the outside (or inside) of the block in a clockwise / counterclockwise direction
[0456] Meanwhile, one of the above directions may be selectively used as the second direction. The first direction and the second direction may be the same, or may be different from each other.
[0457] In addition, the operation of rearranging the inverse transformed residual data of the current block in step S2930 may be performed on each subblock in the current block. In this case, the residual data may be rearranged based on the position of the subblock in the current block. The operation of rearranging the residual data based on the position of the subblock has been described in Formula 6, and therefore, its repeated description will be omitted.
[0458] Also, the operation of rearranging the inverse transformed residual data of the current block at step S2930 may be performed by rotating the inverse transformed residual data arranged in the current block at a predetermined angle.
[0459] In addition, the operation of rearranging the inverse transformed residual data of the current block at step S2930 may be performed by flipping the inverse transformed residual data arranged in the current block according to the flipping method. In this case, the step of determining the transform mode of the current block at step S2910 may include: obtaining flipping method information from the bitstream; and determining the flipping method of the current block based on the flipping method information.
[0460] Fig.30 is a diagram illustrating an encoding method using the SDST method according to the present invention.
[0461] Reference Fig.30, in step S3010, the transformation mode of the current block can be determined.
[0462] In addition, in step S3020, the residual data of the current block may be rearranged according to the transform mode of the current block.
[0463] Furthermore, in step S3030, the rearranged residual data of the current block may be transformed according to a transform mode of the current block.
[0464] Here, the transform mode may include at least one of SDST (reordered discrete sine transform), SDCT (reordered discrete cosine transform), DST (discrete sine transform) or DCT (discrete cosine transform). Fig.29 The SDST, SDCT, DST, and DCT modes are described, and thus, a repeated description thereof will be omitted.
[0465] Meanwhile, rearrangement of residual data may be performed only when a transform mode of a current block is one of SDST and SDCT.
[0466] Also, the operation of determining the transform mode of the current block in step S3010 may be performed based on at least one of the prediction mode of the current block, depth information of the current block, the size of the current block, and the shape of the current block.
[0467] Here, when the prediction mode of the current block is the inter prediction mode, the transform mode of the current block may be determined as one of SDST and SDCT.
[0468] Meanwhile, the operation of rearranging the residual data of the current block in step S3020 may include: scanning the residual data arranged in the current block in a first direction sequence; and rearranging the residual data scanned along the first direction in the current block in a second direction sequence.
[0469] In addition, the operation of rearranging the residual data of the current block in step S3020 may be performed on each subblock in the current block.
[0470] In this case, the operation of rearranging the residual data of the current block in step S3020 may be performed based on the positions of the subblocks in the current block.
[0471] Meanwhile, the operation of rearranging the residual data of the current block at step S3020 may be performed by rotating the residual data arranged in the current block at a predetermined angle.
[0472] Meanwhile, the operation of rearranging the residual data of the current block at step S3020 may be performed by flipping the residual data arranged in the current block according to a flipping method.
[0473] According to the present invention, an apparatus for decoding a video by using an SDST method may include: an inverse transform unit, determining a transform mode of a current block, inversely transforming residual data of the current block according to the transform mode of the current block, and rearranging the inversely transformed residual data of the current block according to the transform mode of the current block. Here, the transform mode may include at least one of SDST (reordered discrete sine transform), SDCT (reordered discrete cosine transform), DST (discrete sine transform) or DCT (discrete cosine transform).
[0474] According to the present invention, an apparatus for decoding a video by using an SDST method may include: an inverse transform unit, determining a transform mode of a current block, rearranging residual data of the current block according to the transform mode of the current block, and inversely transforming the rearranged residual data of the current block according to the transform mode of the current block. Here, the transform mode may include at least one of SDST (rearranged discrete sine transform), SDCT (rearranged discrete cosine transform), DST (discrete sine transform) or DCT (discrete cosine transform).
[0475] According to the present invention, an apparatus for encoding a video by using an SDST method may include: a transform unit, determining a transform mode of a current block, rearranging residual data of the current block according to the transform mode of the current block, and transforming the rearranged residual data of the current block according to the transform mode of the current block. Here, the transform mode may include at least one of SDST (rearranged discrete sine transform), SDCT (rearranged discrete cosine transform), DST (discrete sine transform) or DCT (discrete cosine transform).
[0476] According to the present invention, an apparatus for encoding a video by using an SDST method may include: a transform unit, determining a transform mode of a current block, transforming residual data of the current block according to the transform mode of the current block, and rearranging the transformed residual data of the current block according to the transform mode of the current block. Here, the transform mode may include at least one of SDST (reordered discrete sine transform), SDCT (reordered discrete cosine transform), DST (discrete sine transform) or DCT (discrete cosine transform).
[0477] A bitstream is formed by a method for encoding a view using an SDST method according to the present invention. The method may include: determining a transform mode of a current block; rearranging residual data of the current block according to the transform mode of the current block; and transforming the rearranged residual data of the current block according to the transform mode of the current block. Here, the transform mode may include at least one of SDST (rearranged discrete sine transform), SDCT (rearranged discrete cosine transform), DST (discrete sine transform), or DCT (discrete cosine transform).
[0478] Inter-frame encoding / decoding processing may be performed for each of the luminance signal and the chrominance signal. For example, in the inter-frame encoding / decoding processing, at least one method of obtaining an inter-frame prediction indicator, generating a motion vector candidate list, deriving a motion vector, and performing motion compensation may be applied differently to the luminance signal and the chrominance signal.
[0479] Inter-frame encoding / decoding processing can be performed similarly for luma signals and chroma signals. For example, in the inter-frame encoding / decoding processing applied to the luma signal, at least one of the inter-frame prediction indicator, the motion vector candidate list, the motion vector candidate, the motion vector, and the reference picture can be applied to the chroma signal.
[0480] The method may be performed in the same manner in an encoder and a decoder. For example, in an inter-frame encoding / decoding process, at least one of deriving a motion vector candidate list, deriving a motion vector candidate, deriving a motion vector, and performing motion compensation may be equally applied in an encoder and a decoder. In addition, the order in which the method is applied may be different in an encoder and a decoder.
[0481] Embodiments of the present invention may be applied according to the size of at least one of a coding block, a prediction block, a block, and a unit. Here, the size may be defined as a minimum size and / or a maximum size in order to apply the embodiment, and may be defined as a fixed size to which the embodiment is applied. In addition, the first embodiment may be applied according to the first size, and the second embodiment may be applied according to the second size. That is, the embodiment may be applied multiple times according to the size. In addition, embodiments of the present invention may be applied only when the size is equal to or greater than the minimum size and equal to or less than the maximum size. That is, the embodiment may be applied only when the block size is within a predetermined range.
[0482] For example, the embodiment can be applied only when the size of the encoding / decoding target block is equal to or greater than 8×8. For example, the embodiment can be applied only when the size of the encoding / decoding target block is equal to or greater than 16×16. For example, the embodiment can be applied only when the size of the encoding / decoding target block is equal to or greater than 32×32. For example, the embodiment can be applied only when the size of the encoding / decoding target block is equal to or greater than 64×64. For example, the embodiment can be applied only when the size of the encoding / decoding target block is equal to or greater than 128×128. For example, the embodiment can be applied only when the size of the encoding / decoding target block is 4×4. For example, the embodiment can be applied only when the size of the encoding / decoding target block is equal to or less than 8×8. For example, the embodiment can be applied only when the size of the encoding / decoding target block is equal to or greater than 16×16. For example, the embodiment can be applied only when the size of the encoding / decoding target block is equal to or greater than 8×8 and equal to or less than 16×16. For example, the embodiment may be applied only when the size of the encoding / decoding target block is equal to or larger than 16×16 and equal to or smaller than 64×64.
[0483] Embodiments of the present invention may be applied according to time layers. An identifier for identifying a time layer to which an embodiment may be applied may be sent with a signal, and the embodiment may be applied to the time layer indicated by the indicator. Here, the identifier may be defined as indicating a minimum layer and / or a maximum layer to which an embodiment may be applied, and may be defined as indicating a specific layer to which an embodiment may be applied.
[0484] For example, the embodiment can be applied only when the temporal layer of the current picture is the lowest layer. For example, the embodiment can be applied only when the temporal layer identifier of the current picture is 0. For example, the embodiment can be applied only when the temporal layer identifier of the current picture is equal to or greater than 1. For example, the embodiment can be applied only when the temporal layer of the current picture is the highest layer.
[0485] As described in the embodiments of the present invention, a reference picture set used in the processes of reference picture list construction and reference picture list modification may use at least one of reference picture lists L0, L1, L2, and L3.
[0486] According to an embodiment of the present invention, when the deblocking filter calculates the boundary strength, at least one to a maximum of N motion vectors of the encoding / decoding target block may be used. Here, N indicates a positive integer equal to or greater than 1, such as 2, 3, 4, etc.
[0487] In motion vector prediction, when a motion vector has at least one of the following units, an embodiment of the present invention may be applied: a 16-pixel (16-pel) unit, an 8-pixel (8-pel) unit, a 4-pixel (4-pel) unit, an integer-pixel (integer-pel) unit, a 1 / 2-pixel (1 / 2-pel) unit, a 1 / 4-pixel (1 / 4-pel) unit, a 1 / 8-pixel (1 / 8-pel) unit, a 1 / 16-pixel (1 / 16-pel) unit, a 1 / 32-pixel (1 / 32-pel) unit, and a 1 / 64-pixel (1 / 64-pel) unit. In addition, when performing motion vector prediction, a motion vector may be optionally used for each pixel unit.
[0488] A stripe type to which the embodiment of the present invention is applied may be defined and the embodiment of the present invention may be applied according to the stripe type.
[0489] For example, when the slice type is T (three-way prediction)-slice, the prediction block may be generated by using at least three motion vectors, and may be used as a final prediction block of the encoding / decoding target block by calculating a weighted sum of the at least three prediction blocks. For example, when the slice type is Q (four-way prediction)-slice, the prediction block may be generated by using at least four motion vectors, and may be used as a final prediction block of the encoding / decoding target block by calculating a weighted sum of the at least four prediction blocks.
[0490] The embodiments of the present invention may be applied to an inter prediction and motion compensation method using motion vector prediction as well as an inter prediction and motion compensation method using a skip mode, a merge mode, and the like.
[0491] The shape of a block to which an embodiment of the present invention is applied may have a square shape or a non-square shape.
[0492] In the above embodiments, the method is described based on a flow chart with a series of steps or units, but the present invention is not limited to the order of the steps, but some steps may be performed simultaneously with other steps, or may be performed in a different order from other steps. In addition, it should be understood by those of ordinary skill in the art that the steps in the flow chart are not mutually exclusive, and other steps may be added to the flow chart, or some steps may be deleted from the flow chart without affecting the scope of the present invention.
[0493] Embodiments include various aspects of examples. All possible combinations for each aspect may not be described, but those skilled in the art will be able to recognize different combinations. Therefore, the present invention may include all alternatives, modifications and changes within the scope of the claims.
[0494] Embodiments of the present invention may be implemented in the form of program instructions, which may be executed by various computer components and recorded on a computer-readable recording medium. A computer-readable recording medium may include a separate program instruction, a data file, a data structure, etc., or a combination of program instructions, data files, data structures, etc. The program instructions recorded in the computer-readable recording medium may be specially designed and constructed for the present invention, or may be known to a person of ordinary skill in the field of computer software technology. Examples of computer-readable recording media include: magnetic recording media (such as hard disks, floppy disks, and magnetic tapes); optical data storage media (such as CD-ROMs or DVD-ROMs); magneto-optical media (such as floppy disks); and hardware devices (such as read-only memory (ROM), random access memory (RAM), flash memory, etc.) specially constructed to store and implement program instructions. Examples of program instructions include not only machine language codes formed by a compiler, but also high-level language codes that can be implemented by a computer using an interpreter. A hardware device may be configured to be operated by one or more software modules to perform processing according to the present invention, or vice versa.
[0495] Although the present invention has been described according to specific terms (such as detailed elements) and limited embodiments and drawings, they are only provided to help more popularly understand the present invention, and the present invention is not limited to the above embodiments. It will be appreciated by those skilled in the art that various modifications and changes can be made from the above description.
[0496] Therefore, the spirit of the present invention should not be limited to the above-described embodiments, and the full scope of the appended claims and their equivalents will fall within the scope and spirit of the present invention.
Claims
1. A method for decoding a video, the method comprising: Obtaining inverse quantized transform coefficients of the current block; Obtaining transform mode information of a current block from a bitstream; Determining a transformation kernel of a current block based on the transformation mode information; Obtain residual samples of the current block by performing an inverse transform on the inverse quantized transform coefficients based on a transform kernel of the current block; performing prediction to generate prediction samples for the current block; and Reconstructing a current block based on the residual samples and the prediction samples, The transformation mode information is index information for indicating a transformation core of a current block among a plurality of predefined transformation cores, and When there is only one non-zero transform coefficient in the current block and the non-zero transform coefficient is located at the upper left of the current block, the transform mode information is not obtained from the bitstream.
2. The method of claim 1, wherein: When there are non-zero transform coefficients in the current block, the transform mode information is obtained.
3. The method of claim 1, wherein: When transform skip is not performed in the current block, the transform mode information is obtained.
4. The method of claim 1, wherein: The transform mode information is information indicating a transform kernel for each of a horizontal direction and a vertical direction.
5. The method of claim 1, wherein: The transform mode information is obtained only when the current block is a luma component.
6. The method of claim 1, wherein: The transform mode information is obtained only when the size of the current block is less than or equal to a predetermined size.
7. A method for encoding a video, the method comprising: Obtaining inverse quantized transform coefficients of the current block; Determine the transformation kernel of the current block; Obtain residual samples of the current block by performing an inverse transform on the inverse quantized transform coefficients based on a transform kernel of the current block; performing prediction to generate prediction samples for the current block; Reconstructing a current block based on the residual samples and the prediction samples; and encoding transform mode information indicating a transform kernel of the current block, The transformation mode information is index information for indicating a transformation core of a current block among a plurality of predefined transformation cores, and When there is only one non-zero transform coefficient in the current block and the non-zero transform coefficient is located at the upper left of the current block, the transform mode information is not encoded.
8. A method for providing encoded video data to a video decoding device, the method comprising: Generate a bitstream by encoding the current block; as well as sending the bitstream to a video decoding device, Wherein, generating the bit stream comprises: Obtaining inverse quantized transform coefficients of the current block; Determine the transformation kernel of the current block; Obtain residual samples of the current block by performing an inverse transform on the inverse quantized transform coefficients based on a transform kernel of the current block; performing prediction to generate prediction samples for the current block; Reconstructing a current block based on the residual samples and the prediction samples; and encoding transform mode information indicating a transform kernel of the current block, The transformation mode information is index information for indicating a transformation core of a current block among a plurality of predefined transformation cores, and When there is only one non-zero transform coefficient in the current block and the non-zero transform coefficient is located at the upper left of the current block, the transform mode information is not encoded.