Image encoding / decoding method and apparatus, and recording medium storing bit stream

By optimizing the residual coding method and performing entropy coding and decoding based on the effective area of ​​the transform coefficients, the problem of low efficiency in high-resolution image coding in the existing technology is solved, and more efficient coding and decoding are achieved.

CN120660348APending Publication Date: 2025-09-16HYUNDAI MOTOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480009014.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-29
Filing Date
2024-03-28
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing residual coding technologies are inefficient in high-resolution and high-quality image coding and cannot correctly reflect the transform coefficient characteristics of the current block, resulting in unnecessary bit consumption and reduced decoding efficiency.

Method used

By determining the transform type and context information of the current block, entropy coding and decoding are performed based on the valid area of ​​the transform coefficients, and only the coefficient groups within the valid area are encoded to optimize the context model selection.

Benefits of technology

The efficiency of image encoding and decoding is improved, the number of binary bits for context decoding is reduced, and the encoding efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120660348A_ABST
    Figure CN120660348A_ABST
Patent Text Reader

Abstract

Provided are an image encoding / decoding method, an apparatus, a recording medium storing a bitstream, and a transmission method. The image decoding method comprises the steps of: determining a transform type of a current block; determining context information of transform coefficient information of the current block based on the transform type; and entropy decoding the transform coefficient information on the basis of context information, which may be determined on the basis of whether the transform type of the current block is an inseparable transform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image encoding / decoding method and apparatus, and a recording medium storing a bitstream. In particular, the present invention relates to an image encoding / decoding method and apparatus, and a recording medium storing a bitstream, which use a method of performing residual coding by utilizing information about a valid region where non-zero transform coefficients exist. Background Art

[0002] Recently, in various application fields, the demand for high-resolution and high-quality images (such as ultra-high-definition (UHD) images) has increased. As the resolution and quality of image data become higher, the amount of data increases relatively compared to existing image data. Therefore, when such image data is transmitted using existing media (such as wired or wireless broadband channels), or when the image data is stored using existing storage media, both transmission and storage costs increase. In order to solve these problems that occur as the resolution and quality of image data become higher, high-efficiency image encoding / decoding technology is required for images with higher resolution and image quality.

[0003] According to the existing residual coding technology used in image encoding / decoding technology, a transform block or an N×N square area of ​​a fixed size from a reference point is used to perform binarization of any syntax element. However, according to the existing residual coding technology, when the transform coefficient after separable transform or non-separable transform extends beyond the square area, unnecessary bits may be consumed.

[0004] Furthermore, existing residual coding techniques use a fixed-size N×N square region to select the context model for any syntax element. Because the residual features of the current block cannot be accurately reflected, the optimal context model cannot be selected, reducing decoding efficiency. Summary of the Invention

[0005] Technical issues

[0006] An object of the present invention is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0007] Another object of the present invention is to provide a recording medium for storing a bit stream generated by the image decoding method or apparatus according to the present invention.

[0008] Another object of the present invention is to provide a residual coding method performed based on a valid area reflecting the characteristics of the transformation coefficients of the current block, thereby solving the above-mentioned problems in residual coding.

[0009] Technical Solution

[0010] According to an embodiment of the present invention, an image decoding method may include: determining a transform type of a current block; determining context information of the transform coefficient information of the current block; and entropy decoding the transform coefficient information based on the context information, and the context information may be determined based on whether the transform type of the current block is a non-separable transform.

[0011] In the image decoding method, the transformation coefficient information can be determined based on the transformation coefficient information of the adjacent transformation coefficients adjacent to the current transformation coefficient, the adjacent transformation coefficients can be located within the valid area of ​​the current block, and the valid area of ​​the current block can be determined based on whether the transformation type of the current block is a non-separable transformation.

[0012] In the image decoding method, the current block may consist of a zeroed area and a valid area.

[0013] In the image decoding method, adjacent transform coefficients may be determined based on a scan direction of a transform of a current block.

[0014] In the image decoding method, transform coefficient information may be binarized based on a valid region.

[0015] In the image decoding method, the transformation coefficient information may include prefix information indicating a position value of a last non-zero coefficient in a current block and suffix information indicating a position value of the last non-zero coefficient in the current block.

[0016] In the image decoding method, a maximum value of an information length of a prefix indicating a position value of a last non-zero coefficient in a current block may be determined based on a valid area.

[0017] In the image decoding method, the transform coefficient information may indicate information about each coefficient group partitioned from the current block.

[0018] In the image decoding method, the transform coefficient information may indicate whether a coefficient group includes a non-zero transform coefficient and may be encoded only for the coefficient group included in the valid area.

[0019] According to an embodiment of the present invention, an image encoding method may include: determining a transform type of a current block; determining context information of the transform coefficient information of the current block; and entropy encoding the transform coefficient information based on the context information, and the context information may be determined based on whether the transform type of the current block is a non-separable transform.

[0020] According to an embodiment of the present invention, a non-transitory computer-readable recording medium can store a bit stream generated by an image encoding method, the image encoding method including: determining a transform type of a current block; determining context information of transform coefficient information of the current block; and entropy encoding the transform coefficient information based on the context information, and the context information can be determined based on whether the transform type of the current block is a non-separable transform.

[0021] According to an embodiment of the present invention, a method for transmitting a bit stream can transmit a bit stream generated by an image encoding method, the image encoding method including: determining a transform type of a current block; determining context information of the transform coefficient information of the current block; based on the context information, entropy encoding the transform coefficient information, and the context information can be determined based on whether the transform type of the current block is a non-separable transform.

[0022] The features briefly summarized above with respect to the present disclosure are provided merely as examples to explain specific embodiments and are not to be construed as limiting the scope of the present disclosure.

[0023] Beneficial effects

[0024] According to the present invention, it is possible to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0025] In addition, according to the present invention, it is possible to provide a residual encoding method performed based on a valid area reflecting characteristics of a transform coefficient of a current block.

[0026] Furthermore, according to the present invention, the number of binary bits of context coding used in residual encoding can be reduced.

[0027] Effects obtainable from the present disclosure are not limited to the above-mentioned effects, and other effects not mentioned will be clearly understood from the following description by those skilled in the art. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a block diagram showing the configuration of an encoding apparatus according to an embodiment of the present invention.

[0029] Figure 2 is a block diagram showing the configuration of a decoding device according to an embodiment of the present invention.

[0030] Figure 3 is a diagram schematically illustrating a video decoding system to which the present invention is applicable.

[0031] Figure 4 is a view for describing a annihilation method in a non-separable transform according to an embodiment of the present invention.

[0032] Figure 5is a view for describing a valid area derived as a transformation result of a current block according to an embodiment of the present invention.

[0033] Figure 6 is a view showing an implementation of a CG segmented from a transform block.

[0034] Figure 7 It is a view for describing a method of encoding a flag indicating whether there is a non-zero coefficient in a CG according to an embodiment of the present invention.

[0035] Figure 8 is a view illustrating an embodiment of a method for determining a context model based on a valid area according to an embodiment of the present invention.

[0036] Figure 9 is a view illustrating an embodiment of a method for determining a context model based on a valid area according to an embodiment of the present invention.

[0037] Figure 10 is a flowchart illustrating an image decoding method according to an embodiment of the present invention.

[0038] Figure 11 FIG. 1 is a diagram illustrating a content streaming system to which an embodiment of the present invention is applicable. DETAILED DESCRIPTION

[0039] The present invention may have various modifications and embodiments, and specific embodiments are shown in the drawings and described in detail in the specific embodiments. However, this is not intended to limit the present invention to specific embodiments, but should be understood to include all modifications, equivalents or replacements included in the spirit and scope of the present invention. In all aspects, similar reference numerals in the drawings indicate the same or similar functions. The shapes and sizes of the elements in the drawings can be provided by way of example for a clearer description. The detailed description of the exemplary embodiments described below refers to the drawings, which illustrate specific embodiments by way of example. These embodiments are described in sufficient detail to enable those skilled in the art to practice these embodiments. It should be understood that various embodiments are different from each other, but not necessarily mutually exclusive. For example, the specific shapes, structures and characteristics described herein can be implemented in other embodiments without departing from the spirit and scope of the present invention regarding an embodiment. It should also be understood that the position or arrangement of the various components in each disclosed embodiment can be changed without departing from the spirit and scope of the embodiment. Therefore, the specific embodiments set forth below are not intended to be restrictive, and the scope of the exemplary embodiments is limited only by the full range of equivalents to which the appended claims and these claims are entitled (if appropriately described).

[0040] In the present invention, the terms first, second, etc. may be used to describe various components, but these components should not be limited by these terms. These terms are used only to distinguish one component from another. For example, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component without departing from the scope of the present invention. The term is and / or includes a combination of multiple related descriptive items or any item among multiple related descriptive items.

[0041] The components shown in the embodiments of the present invention are depicted independently to indicate different characteristic functions, and do not mean that each component is formed as a separate hardware or software configuration unit. That is, for ease of explanation, each component is listed and included as a separate component, and at least two components may be combined to form a single component, or one component may be divided into multiple components to perform functions, and embodiments in which components are integrated and embodiments in which each component is divided are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.

[0042] The terms used in the present invention are only used to describe specific embodiments and are not intended to limit the present invention. Unless the context clearly indicates otherwise, singular expressions include plural expressions. In addition, some components of the present invention are not essential components for performing the basic functions of the present invention and may be optional components that are only used to improve performance. The present invention can be implemented by only including essential components for realizing the essence of the present invention and not including components that are only used to improve performance, and structures that only include essential components except optional components that are only used to improve performance are also included within the scope of the present invention.

[0043] In an embodiment, the term "at least one" may mean a number greater than or equal to 1, such as 1, 2, 3, and 4. In an embodiment, the term "plurality" may mean a number greater than or equal to 2, such as 2, 3, and 4.

[0044] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. When describing the embodiments of this specification, if it is determined that a detailed description of a related known configuration or function may obscure the subject matter of this specification, the detailed description will be omitted, and the same reference numerals will be used for the same components in the drawings, and repeated description of the same components will be omitted.

[0045] Terminology Description

[0046] Hereinafter, "image" may refer to a picture constituting a video, or may refer to the video itself. For example, "encoding and / or decoding of an image" may refer to "encoding and / or decoding of a video", or may refer to "encoding and / or decoding of a picture constituting a video".

[0047] Hereinafter, "moving image" and "video" may be used interchangeably with each other in the same meaning. Furthermore, a target image may be an encoding target image serving as an encoding target and / or a decoding target image serving as a decoding target. Furthermore, a target image may be an input image input to an encoding device or an input image input to a decoding device. Here, the target image may have the same meaning as the current image.

[0048] Hereinafter, an encoder and an image encoding device may be used with the same meaning and may be used interchangeably.

[0049] Hereinafter, a decoder and an image decoding device may be used with the same meaning and may be used interchangeably.

[0050] Hereinafter, “image,” “picture,” “frame,” and “picture” may be used with the same meaning and may be used interchangeably.

[0051] Hereinafter, a "target block" may be an encoding target block that is an encoding target and / or a decoding target block that is a decoding target. In addition, a target block may be a current block that is a target of current encoding and / or decoding. For example, "target block" and "current block" may be used with the same meaning and may be used interchangeably.

[0052] Hereinafter, "block" and "unit" may be used interchangeably with each other. Furthermore, a "unit" may include a luma component block and its corresponding chroma component blocks to distinguish a unit from a block. For example, a coding tree unit (CTU) may consist of one luma component (Y) coding tree block (CTB) and its associated two chroma component (Cb, Cr) coding tree blocks.

[0053] Hereinafter, "sample", "picture element" and "pixel" may be used with the same meaning and may be used interchangeably. Herein, a sample may represent a basic unit constituting a block.

[0054] Hereinafter, “inter-frame” and “inter-screen” may be used with the same meaning and may be used interchangeably.

[0055] Hereinafter, “intra-frame” and “intra-screen” may be used with the same meaning and may be used interchangeably.

[0056] Figure 1 is a block diagram showing the configuration of an encoding apparatus according to an embodiment of the present invention.

[0057] The encoding device 100 may be an encoder, a video encoding device, or an image encoding device. A video may include one or more images. The encoding device 100 may encode one or more images sequentially.

[0058] refer to Figure 1 , the encoding device 100 may include an image segmentation unit 110, an intra-frame prediction unit 120, a motion prediction unit 121, a motion compensation unit 122, a switch 115, a subtractor 113, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, a dequantization unit 160, an inverse transform unit 170, an adder 117, a filter unit 180 and a reference picture buffer 190.

[0059] In addition, the encoding device 100 can generate a bit stream including information encoded by encoding the input image and output the generated bit stream. The generated bit stream can be stored in a computer-readable recording medium or can be streamed through a wired / wireless transmission medium.

[0060] The image segmentation unit 110 can segment the input image into various forms to improve the efficiency of video encoding / decoding. That is, the input video is composed of multiple pictures, and for compression efficiency, parallel processing, etc., a picture can be segmented and processed hierarchically. For example, a picture can be segmented into one or more tiles or slices, and then segmented again into multiple coding tree units (CTUs). Alternatively, a picture can first be segmented into multiple sub-pictures defined as rectangular slice groups, and each sub-picture can be segmented into tiles / slices. Here, sub-pictures can be used to support the functions of partially independent encoding / decoding and transmitting pictures. Since multiple sub-pictures can be reconstructed separately, it has the advantage of easy editing in applications that configure multi-channel input into one picture. In addition, tiles can be divided horizontally to generate small blocks. Here, small blocks can be used as basic units for parallel processing within the picture. In addition, a CTU can be recursively segmented into a quadtree (QT), and the end node of the segmentation can be defined as a decoding unit (CU). The CU can be split into a prediction unit (PU) as a prediction unit and a transform unit (TU) as a transform unit to perform prediction and segmentation. At the same time, the CU can be used as a prediction unit and / or a transform unit itself. Here, for flexible segmentation, each CTU can be recursively split into a multi-type tree (MTT) and a quadtree (QT). The splitting of the CTU into a multi-type tree can start from the end node of the QT, and the MTT can be composed of a binary tree (BT) and a ternary tree (TT). For example, the MTT structure can be classified into a vertical binary split mode (SPLIT_BT_VER), a horizontal binary split mode (SPLIT_BT_HOR), a vertical ternary split mode (SPLIT_TT_VER), and a horizontal ternary split mode (SPLIT_TT_HOR). In addition, the minimum block size (MinQTSize) of the quadtree of the luma block during partitioning can be set to 16×16, the maximum block size (MaxBtSize) of the binary tree can be set to 128×128, and the maximum block size (MaxTtSize) of the ternary tree can be set to 64×64. In addition, the minimum block size (MinBtSize) of the binary tree and the minimum block size (MinTtSize) of the ternary tree can be specified as 4×4, and the maximum depth (MaxMttDepth) of the multi-type tree can be specified as 4. In addition, in order to improve the coding efficiency of the I slice, a dual tree of the CTU partition structure that uses the luma component and the chroma component differently can be applied. On the other hand, in P and B slices, the luma and chroma coding tree blocks (CTBs) within the CTU can be partitioned into a single tree that shares the coding tree structure.

[0061] The encoding device 100 can perform encoding on the input image in intra mode and / or inter mode. Alternatively, the encoding device 100 can perform encoding on the input image in a third mode other than intra mode and inter mode (for example, IBC mode, palette mode, etc.). However, if the third mode has functional characteristics similar to those of intra mode or inter mode, the third mode may be classified as intra mode or inter mode for ease of explanation. In the present invention, the third mode is separately classified and described only when a specific description of the third mode is required.

[0062] When the intra mode is used as the prediction mode, the switch 115 can be switched to intra, and when the inter mode is used as the prediction mode, the switch 115 can be switched to inter. Here, the intra mode can mean the intra prediction mode, and the inter mode can mean the inter prediction mode. The encoding device 100 can generate a prediction block for the input block of the input image. In addition, the encoding device 100 can encode the residual block using the residual of the input block and the prediction block after the prediction block is generated. The input image can be referred to as the current image of the current encoding target. The input block can be referred to as the current block that is the current encoding target or the encoding target block.

[0063] When the prediction mode is intra mode, the intra prediction unit 120 can use samples of blocks that have been encoded / decoded around the current block as reference samples. The intra prediction unit 120 can perform spatial prediction on the current block using the reference samples, or generate prediction samples of the input block through spatial prediction. Here, intra prediction can refer to intra-frame prediction.

[0064] As the intra prediction method, non-directional prediction modes such as DC mode and planar mode and directional prediction modes (for example, 65 directions) may be applied. Here, the intra prediction method may be expressed as an intra prediction mode or an intra-screen prediction mode.

[0065] When the prediction mode is inter mode, the motion prediction unit 121 can retrieve the area that best matches the input block from the reference image during the motion prediction process and derive a motion vector by using the retrieved area. In this case, the area can be used as a search area. The reference image can be stored in the reference picture buffer 190. Here, when encoding / decoding is performed on the reference image, it can be stored in the reference picture buffer 190.

[0066] The motion compensation unit 122 may generate a prediction block of the current block by performing motion compensation using a motion vector. Here, inter prediction may mean inter-picture prediction or motion compensation.

[0067] When the value of the motion vector is not an integer, the motion prediction unit 121 and the motion compensation unit 122 may generate a prediction block by applying an interpolation filter to a partial area of ​​the reference picture. In order to perform inter-frame prediction or motion compensation, it may be determined whether the motion prediction and motion compensation mode of the prediction unit included in the decoding unit is one of skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and decoding unit-based intra block copy (IBC) mode, and inter-frame prediction or motion compensation may be performed according to each mode.

[0068] In addition, based on the above inter-frame prediction methods, the AFFINE mode based on sub-PU prediction, the sub-block based temporal motion vector prediction (SbTMVP) mode, the PU-based prediction and MVD merge (MMVD) mode and the geometric partitioning mode (GPM) can be applied. In addition, in order to improve the performance of each mode, the history-based MVP (HMVP), pairwise average MVP (PAMVP), joint intra / inter prediction (CIIP), adaptive motion vector resolution (AMVR), bidirectional optical flow (BDOF), bidirectional prediction with CU weight (BCW), local illumination compensation (LIC), template matching (TM), overlapped block motion compensation (OBMC), etc. can be applied.

[0069] Among them, the AFFINE mode is a technology used in both AMVP and MERGE modes, and also has high coding efficiency. In the existing video coding standard, because motion compensation (MC) is performed by considering only the parallel movement of blocks, its disadvantage is that it cannot properly compensate for actual motions such as zooming in / out and rotation. To supplement this, a four-parameter affine motion model using two control point motion vectors (CPMV) and a six-parameter affine motion model using three control point motion vectors can be utilized and applied to inter-frame prediction. Here, CPMV is a vector representing an affine motion model of one of the upper left, upper right, and lower left sides of the current block.

[0070] The subtractor 113 may generate a residual block by using the difference between the input block and the prediction block. The residual block may be referred to as a residual signal. The residual signal may represent the difference between the original signal and the prediction signal. Alternatively, the residual signal may be a signal generated by transforming or quantizing, or transforming and quantizing, the difference between the original signal and the prediction signal. The residual block may be a residual signal of a block unit.

[0071] The transform unit 130 may generate a transform coefficient by performing a transform on the residual block and output the generated transform coefficient. Here, the transform coefficient may be a coefficient value generated by performing a transform on the residual block. When the transform skip mode is applied, the transform unit 130 may skip transforming the residual block.

[0072] The quantization level may be generated by applying quantization to the transform coefficient or the residual signal. In the following embodiments, the quantization level may also be referred to as the transform coefficient.

[0073] For example, a 4×4 luminance residual block generated by intra prediction is transformed using a basis vector based on discrete sine transform (DST), and the remaining residual block can be transformed using a basis vector based on discrete cosine transform (DCT). In addition, using the residual quadtree (RQT) technology, for one block, the transform block is divided into a quadtree shape, and after transforming and quantizing each transform block divided by RQT, a coded block flag (CBF) can be transmitted when all coefficients become 0 to improve coding efficiency.

[0074] As another alternative, a multiple transform selection (MTS) technique that selectively uses multiple transform bases to perform transforms can be applied. That is, without splitting the CU into TUs through RQT, a function similar to TU splitting can be performed through sub-block transform (SBT) technology. Specifically, SBT is only applied to inter-frame prediction blocks, and unlike RQT, the current block can be split into 1 / 2 or 1 / 4 sizes in the vertical or horizontal direction, and then the transform can be performed on only one block in the block. For example, if it is split vertically, the transform can be performed on the leftmost or rightmost block, and if it is split horizontally, the transform can be performed on the topmost or bottommost block.

[0075] In addition, a low-frequency non-separable transform (LFNST) can be applied. LFNST is a secondary transform technique that additionally transforms the residual signal transformed to the frequency domain by DCT or DST. LFNST additionally transforms the 4×4 or 8×8 low-frequency region on the upper left side, so that the residual coefficients can be concentrated on the upper left side.

[0076] The quantization unit 140 may generate a quantization level by quantizing the transform coefficient or the residual signal according to a quantization parameter (QP) and output the generated quantization level. Here, the quantization unit 140 may quantize the transform coefficient by using a quantization matrix.

[0077] For example, a quantizer using a QP value of 0 to 51 may be employed. Alternatively, if the image size is large and high coding efficiency is required, a QP of 0 to 63 may be used. In addition, a dependent quantization (DQ) method using two quantizers instead of one quantizer may be applied. DQ performs quantization using two quantizers (e.g., Q0 and Q1), but even without signaling information about the use of a specific quantizer, the quantizer to be used for the next transform coefficient can be selected based on the current state through a state transition model.

[0078] The entropy coding unit 150 may generate a bitstream by performing entropy coding on the values ​​calculated by the quantization unit 140 or the coding parameter values ​​calculated when performing coding according to a probability distribution, and output the bitstream. The entropy coding unit 150 may perform entropy coding on information about samples of an image and information for decoding the image. For example, the information for decoding the image may include syntax elements.

[0079] When entropy coding is applied, symbols are represented so that a smaller number of bits are allocated to symbols with a high probability of occurrence and a larger number of bits are allocated to symbols with a low probability of occurrence, and therefore, the size of the bit stream for the symbol to be encoded can be reduced. The entropy coding unit 150 can use a coding method such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. to perform entropy coding. For example, the entropy coding unit 150 can perform entropy coding by using a variable length coding / coding (VLC) table. In addition, the entropy coding unit 150 can derive a binarization method of the target symbol and a probability model of the target symbol / bin, and perform arithmetic decoding by using the derived binarization method and context model.

[0080] In this regard, when CABAC is applied, in order to reduce the size of the probability table stored in the decoding device, the table probability update method can be changed to a table update method using a simple equation and applied. In addition, two different probability models can be used to obtain more accurate symbol probability values.

[0081] In order to encode a transform coefficient level (quantized level), the entropy encoding unit 150 may change a two-dimensional block form coefficient into a one-dimensional vector form through a transform coefficient scanning method.

[0082] The decoding parameters may include information (flags, indexes, etc.) encoded in the encoding device 100 and signaled to the decoding device 200, such as syntax elements, and information derived in the encoding or decoding process, and may represent information required when encoding or decoding an image.

[0083] Here, signaling a flag or an index may mean that a corresponding flag or index is entropy-encoded in an encoder and included in a bitstream, and may mean that a corresponding flag or index is entropy-decoded from a bitstream in a decoder.

[0084] The encoded current image can be used as a reference image for other images to be processed subsequently. Therefore, the encoding device 100 can reconstruct or decode the encoded current image again and store the reconstructed or decoded image as a reference image in the reference picture buffer 190.

[0085] The quantized level may be dequantized in the dequantization unit 160 or inversely transformed in the inverse transform unit 170. The dequantized and / or inversely transformed coefficients may be added to the prediction block by the adder 117. Here, the dequantized and / or inversely transformed coefficients may refer to coefficients on which at least one of dequantization and inverse transformation is performed, and may refer to a reconstructed residual block. The dequantization unit 160 and the inverse transform unit 170 may be performed as inverse processes of the quantization unit 140 and the transform unit 130.

[0086] The reconstructed block may pass through the filter unit 180. The filter unit 180 may apply all or some filtering techniques such as a deblocking filter, sample adaptive offset (SAO), an adaptive loop filter (ALF), a bilateral filter (BIF), and luma mapping with chroma scaling (LMCS) to the reconstructed sample, reconstructed block, or reconstructed image. The filter unit 180 may be referred to as a loop filter. In this case, the term loop filter is also used to exclude LMCS.

[0087] A deblocking filter can remove block distortion generated at the boundaries between blocks. To determine whether to apply a deblocking filter, a determination can be made based on samples included in a number of rows or columns contained in the block. When applying a deblocking filter to a block, different filters can be applied depending on the desired deblocking filter strength.

[0088] To compensate for coding errors using sample adaptive offset, an appropriate offset value can be added to the sample value. Sample adaptive offset can correct the offset of the deblocked image from the original image on a sample-by-sample basis. A method can be used in which the samples contained in the image are divided into a predetermined number of regions, the regions to which an offset is applied are determined, and the offset is applied to the determined regions, or a method in which the offset is applied taking into account edge information about each sample.

[0089] The bilateral filter (BIF) can also correct the offset from the original image on a sample-level basis for the deblocked image.

[0090] The adaptive loop filter can perform filtering based on the comparison result of the reconstructed image and the original image. The samples contained in the image can be divided into predetermined groups, the filter to be applied to each group can be determined, and differential filtering can be performed for each group. Information on whether to apply the ALF can be transmitted by decoding unit (CU) signaling, and the form and coefficients of the adaptive loop filter to be applied to each block can be varied.

[0091] In luma mapping with chroma scaling (LMCS), luma mapping (LM) means remapping luma values ​​using a piecewise linear model, and chroma scaling (CS) means a technique for scaling the residual values ​​of chroma components according to the average luma value of the prediction signal. Specifically, LMCS can be used as an HDR correction technique that reflects the characteristics of high dynamic range (HDR) images.

[0092] The reconstructed block or reconstructed image that has passed through the filter unit 180 may be stored in the reference picture buffer 190. The reconstructed block that has passed through the filter unit 180 may be part of a reference image. That is, the reference image is a reconstructed image composed of the reconstructed block that has passed through the filter unit 180. The stored reference image may be used later in inter-frame prediction or motion compensation.

[0093] Figure 2 is a block diagram showing the configuration of a decoding device according to an embodiment of the present invention.

[0094] The decoding device 200 may be a decoder, a video decoding device, or an image decoding device.

[0095] refer to Figure 2 , the decoding device 200 may include an entropy decoding unit 210, a dequantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 201, a switch 203, a filter unit 260 and a reference picture buffer 270.

[0096] The decoding device 200 can receive the bitstream output from the encoding device 100. The decoding device 200 can receive the bitstream stored in a computer-readable recording medium, or can receive the bitstream streamed via a wired / wireless transmission medium. The decoding device 200 can decode the bitstream in intra-frame mode or inter-frame mode. In addition, the decoding device 200 can generate a reconstructed image or a decoded image generated by decoding, and output the reconstructed image or the decoded image.

[0097] When the prediction mode for decoding is intra mode, the switch 203 may be switched to intra. Alternatively, when the prediction mode for decoding is inter mode, the switch 203 may be switched to inter.

[0098] The decoding device 200 can obtain a reconstructed residual block by decoding the input bit stream and generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block that becomes the decoding target by adding the reconstructed residual block and the prediction block. The decoding target block can be referred to as the current block.

[0099] The entropy decoding unit 210 may generate symbols by entropy decoding the bit stream according to the probability distribution. The generated symbols may include symbols in the form of quantized levels. Here, the entropy decoding method may be the inverse process of the above-mentioned entropy encoding method.

[0100] The entropy decoding unit 210 may change the coefficients of the one-dimensional vector shape into coefficients of the two-dimensional block shape through a transform coefficient scanning method to decode transform coefficient levels (quantized levels).

[0101] The quantized levels may be dequantized in the dequantization unit 220 or inversely transformed in the inverse transform unit 230. The quantized levels may be the result of dequantization and / or inverse transformation and may be generated as a reconstructed residual block. Here, the dequantization unit 220 may apply a quantization matrix to the quantized levels. The dequantization unit 220 and the inverse transform unit 230 used in the decoding device may apply the same techniques as the dequantization unit 160 and the inverse transform unit 170 used in the encoding device described above.

[0102] When intra mode is used, the intra prediction unit 240 can generate a prediction block by performing spatial prediction on the current block using sample values ​​of blocks decoded around the decoding target block. The intra prediction unit 240 applied to the decoding device can apply the same technology as the intra prediction unit 120 applied to the above-mentioned encoding device.

[0103] When using inter-frame mode, the motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block using the motion vector and the reference image stored in the reference picture buffer 270. When the value of the motion vector is not an integer value, the motion compensation unit 250 can generate a prediction block by applying an interpolation filter to a partial area within the reference image. In order to perform motion compensation, it can be determined based on the decoding unit whether the motion compensation method of the prediction unit included in the corresponding decoding unit is skip mode, merge mode, AMVP mode, or current picture reference mode, and motion compensation can be performed according to each mode. The motion compensation unit 250 applied to the decoding device can apply the same technology as the motion compensation unit 122 applied to the above-mentioned encoding device.

[0104] The adder 201 can generate a reconstructed block by adding the reconstructed residual block and the prediction block. The filter unit 260 can apply at least one of an inverse LMCS, a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the reconstructed block or the reconstructed image. The filter unit 260 applied to the decoding device can apply the same filtering technology as the filtering technology applied to the filter unit 180 (applied to the above-mentioned encoding device).

[0105] The filter unit 260 may output a reconstructed image. The reconstructed block or image may be stored in the reference picture buffer 270 and used for inter-frame prediction. The reconstructed block that has passed through the filter unit 260 may be part of a reference image. In other words, the reference image may be a reconstructed image composed of the reconstructed block that has passed through the filter unit 260. The stored reference image may be subsequently used for inter-frame prediction or motion compensation.

[0106] Figure 3 is a diagram schematically illustrating a video decoding system to which the present invention is applicable.

[0107] The video decoding system according to the embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in the form of a file or stream via a digital storage medium or a network.

[0108] The encoding device 10 according to the embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. The decoding device 20 according to the embodiment may include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The reception unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, and the display unit may be configured as a separate device or an external component.

[0109] The video source generation unit 11 can obtain a video / image by capturing, synthesizing, or generating a video / image. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, a smartphone, etc., and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process for generating relevant data.

[0110] The encoding unit 12 can encode the input video / image. The encoding unit 12 can perform a series of processes for compression and encoding efficiency, such as prediction, transformation, and quantization. The encoding unit 12 can output encoded data (encoded video / image information) in the form of a bit stream. The detailed configuration of the encoding unit 12 can also be the same as above. Figure 1 The encoding device 100 is configured in the same manner.

[0111] The transmission unit 13 transmits the encoded video / image information or data output in the form of a bitstream to the receiving unit 21 of the decoding device 20 in the form of a file or stream via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit 13 may include components for generating a media file in a predetermined file format and may also include components for transmission via a broadcast / communication network. The receiving unit 21 may extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.

[0112] The decoding unit 22 can decode the video / image by performing a series of processes corresponding to the operations of the encoding unit 12, such as dequantization, inverse transformation, and prediction. The detailed configuration of the decoding unit 22 can also be the same as that described above. Figure 2 The decoding device 200 is configured in the same manner.

[0113] The rendering unit 23 may render the decoded video / image. The rendered video / image may be displayed by the display unit.

[0114] Transform is a technique for converting a signal in the spatial domain into a signal in the frequency domain. In order to improve the compression performance for high-resolution videos such as high-definition (HD) images or ultra-high-definition (UHD) images, the latest video compression standards support transforms for large-sized transform blocks. For example, the H.264 / AVC standard only supports transforms for transform blocks of 4×4 and 8×8 sizes, but the HEVC standard supports transforms for transform blocks of sizes ranging from 4×4 up to 32×32. In addition, the VVC standard supports transforms for transform blocks of up to 64×64 in size.

[0115] Generally, decoding efficiency improves as the size of the transform kernel increases. However, as the size of the transform kernel used in the transform increases, the computational complexity increases exponentially. Furthermore, as the size of the transform kernel increases, the memory required to store the kernel in the encoder and decoder also increases. Therefore, various methods are used in video compression standards to reduce the size of the transform kernel and reduce computational complexity.

[0116] As one method, when performing the transform, a method of zeroing high-frequency transform coefficients may be used.

[0117] In the case of separable transforms, by performing zeroing, for transform blocks with a length greater than X, up to Y transform coefficients starting from the DC position can be retained. Therefore, for X-length transforms (e.g., X-point DCT2, X-point DCT7, X-point DCT8), each transform requires only Y×X transform kernels, thereby reducing the memory required to store the transform kernels and reducing the computational load of the transform. As a result of performing zeroing on both the horizontal kernel and the vertical kernel during the transform process, non-zero transform coefficients can exist only within a predetermined region.

[0118] On the other hand, zeroing in a non-separable transform can be performed as follows.

[0119] Figure 4 is a view for describing a annihilation method in a non-separable transform according to an embodiment of the present invention.

[0120] refer to Figure 4 , the transform coefficients of the input block on which the inseparable transform is performed may be scanned in a predetermined direction and rearranged into a one-dimensional vector of a 1×M shape. Here, the predetermined direction may be one of a row-first direction, a column-first direction, and a diagonal direction. Here, M may be the product of the width and height of the transformed input block, i.e., (TbW×TbH). Alternatively, M may be the product of the width and height of a predetermined region of interest (ROI) within the transform block.

[0121] An M×N transform kernel may be applied to the rearranged 1×M vector. Here, N may be a positive integer equal to or less than M. If M is greater than N, the high-frequency transform coefficients may be zeroed. As a result of the transform, a one-dimensional vector having a 1×N shape may be output.

[0122] The one-dimensional vector having a 1×N shape generated as a result of the transformation can be scanned again in a predetermined direction in units of blocks or coefficient groups (CGs) and rearranged into a two-dimensional form. Here, the predetermined direction can be one of a row-first direction, a column-first direction, and a diagonal direction.

[0123] Figure 4 This figure illustrates the case where zeroing is performed in a non-separable transform. As a result of performing zeroing in a non-separable transform, zero or non-zero transform coefficients may exist in a specific region of the block with reference to the (0, 0) position of the block. On the other hand, all transform coefficients in the remaining region may be zero. Here, the region where non-zero transform coefficients exist may be referred to as a valid region.

[0124] For example, when a non-separable transform is performed on an 8×8 transform block, the transform coefficients of the 8×8 transform block can be scanned in a predetermined direction and rearranged into a 1×64 vector. In addition, a 64×32 transform kernel can be applied to the rearranged 1×64 vector. The 1×32 vector obtained as a result of the transform can be scanned in a predetermined direction and rearranged into a two-dimensional form. As a result, non-zero transform coefficients may only exist in a 4×8 block, which is a valid area within the 8×8 transform block, and each transform coefficient outside the valid area becomes 0.

[0125] By zeroing, the computational load used in the transform and / or inverse transform and the complexity of the transform and / or inverse transform can be reduced. In addition, instead of an M×M kernel, only an M×N kernel (where N<M) is required for the transform and / or inverse transform, and thus the memory used to store the transform kernel can also be reduced.

[0126] Figure 5 is a diagram for describing a valid area derived as a transformation result of a current block according to an embodiment of the present invention.

[0127] refer to Figure 5 (a), the size of the current block can be TbW×TbH. In addition, refer to Figure 5 (b), the current block may be divided into a region including N zero or non-zero transform coefficients and a region including only zero transform coefficients. Here, the region including N zero or non-zero transform coefficients may be referred to as a valid region. On the other hand, the region including only zero transform coefficients may be referred to as a zeroing region. Figure 5 As shown in (b), the size of the effective area can be ZoTbW×ZoTbH.

[0128] According to the present invention, encoding and / or decoding of transform coefficient information can be performed based on a valid region derived as a transform result of a current block. Here, the transform coefficient information may include a syntax element indicating information about residual encoding of the current block.

[0129] That is, based on the valid region, binarization of transform coefficient information and encoding and / or decoding of the transform coefficient information may be performed, and a context model of the transform coefficient information may be determined.

[0130] When transform coefficient information of a current block is encoded, the transform coefficient information may be binarized based on a valid region derived as a transform result of the current block.

[0131] According to conventional residual coding, transform coefficient information is binarized using a transform block or an N×N square region of a fixed size. For example, when indicating the position of the last non-zero transform coefficient within a transform block, information indicating the x-coordinate position of the last non-zero transform coefficient and information indicating the y-coordinate position of the last non-zero transform coefficient may be signaled separately.

[0132] [Table 1]

[0133]

[0134]

[0135] Referring to Table 1, the prefix information indicating the position of the last non-zero transform coefficient may be binarized using a truncated unary (TU), and the suffix information may be binarized using a fixed-length code (FLC). Here, if the x-coordinate or y-coordinate of the lower rightmost corner within the transform block or the fixed-size N×N square area is equal to the last number in the coordinate sequence defined in Table 1, the "0" contained in the brackets shown in Table 1 may not be encoded.

[0136] According to the present invention, the transform coefficient information can be binarized by using the width / height of a predetermined K×L size area determined based on various information, rather than the width / height of a transform block or a fixed-size N×N square area. Here, the values ​​of K and L can be the same or different from each other.

[0137] According to an embodiment, the width K and height L of the predetermined area may be determined using the width ZoTbW and / or height ZoTbH of the effective area derived as a zeroing result of separable transformation or non-separable transformation. Here, the values ​​of ZoTbW and ZoTbH may be the same as or different from each other.

[0138] The maximum length of the codeword of the prefix information indicating the x-coordinate position of the last non-zero transform coefficient may be determined according to Formula 1. In addition, the maximum length of the codeword of the prefix information indicating the y-coordinate position of the last non-zero transform coefficient may be determined according to Formula 2.

[0139] [Formula 1]

[0140] cMax x =min(TbW,ZoTbW)-1

[0141] [Formula 2]

[0142] cMax y =min(TbH,ZoTbH)-1

[0143] For example, when the current transform block is 8×8 and the width of the valid region of the current transform block to which zeroing of the inseparable transform is applied is 4×8, if the position of the last non-zero transform coefficient is (3, 7), the information indicating the position of the last non-zero transform coefficient can be binarized as (111, 111111). Since the proposed method uses a specific region instead of a transform block as the valid region, the number of binary bits used to binarize the x-coordinate value 3 can be saved compared to the conventional technology.

[0144] According to another embodiment, when a separable transform is applied to a current block of size 32×32, the valid region may have a size of 4×16. In the valid region, if the position of the last non-zero transform coefficient is (3, 15), information indicating the position of the last non-zero transform coefficient may be binarized as (111, 1111111011).

[0145] Unlike conventional techniques, the method proposed in the present invention can set a non-square area as a valid area, thereby saving the number of binary bits used to binarize the x-coordinate value 3 of the position of the last non-zero transform coefficient.

[0146] According to another embodiment of the present invention, the value of K (K is the width of the predetermined area) and the value of L (L is the height of the predetermined area) can be positive integers equal to or greater than 1 predefined in the encoder and decoder. Alternatively, any positive integers K and L can be determined based on one or more of the block size, the type of transform applied to the current block, the transform kernel size of the current block, the aspect ratio of the block, the quantization parameter (QP), and the prediction mode of the block. Here, the values ​​of K and L can be the same or different from each other.

[0147] According to another embodiment of the present invention, the value of K (K is the width of the predetermined region) and the value of L (L is the height of the predetermined region) may be determined based on a pre-encoding syntax element. The pre-encoding syntax element may be at least one of a syntax element containing information about a current block or a neighboring block, a syntax element indicating the type of transform to be applied, and a separable transform or non-separable transform-related syntax element containing a transform index and an index indicating transform core information / a transform core list.

[0148] When using K and L as predetermined arbitrary positive integers or determined according to precoding syntax elements, the maximum length of the codeword of the x-coordinate prefix can be determined according to Formula 3. In addition, the maximum length of the codeword of the y-coordinate prefix can be determined according to Formula 4.

[0149] [Formula 3]

[0150] cMax x=min(TbW,K)-1

[0151] [Formula 4]

[0152] cMax y =min(TbH,L)-1

[0153] The proposed method binarizes transform coefficient information using arbitrary regions where non-zero transform coefficients are likely to be located with high probability. Consequently, the number of bins generated thereby can be reduced, and decoding efficiency can be improved. Furthermore, if the reduced bins are context-coded bins, the proposed method can improve residual coding throughput.

[0154] By using a predetermined area instead of the current block, not only the information indicating the position of the last non-zero transform coefficient within the current block can be binarized, but also other information about the transform coefficients can be binarized. That is, other transform coefficient information can be binarized using arbitrary positive integers K and / or L instead of the width and / or height of the block or the width and / or height of a predetermined square area. Here, K and L can be determined based on the width and / or height of the valid area after zeroing, a predetermined arbitrary positive integer, and a previously encoded syntax element.

[0155] The transformation coefficient information of the current block can be encoded based on the valid area derived as the transformation result of the current block. The transformation coefficient information of the current block can be encoded in units of CGs segmented from the current block. The implementation of the CG segmented from the current block can be described as follows.

[0156] Figure 6 is a diagram illustrating an implementation of a CG segmented from a transform block.

[0157] refer to Figure 6 , the size of the transform block can be 8×8. In addition, the transform block can be divided into four CGs. The size of each CG can be 4×4.

[0158] In addition, a flag indicating whether there is a non-zero transform coefficient in each CG may be encoded. The flag indicating whether there is a non-zero transform coefficient may be sb_coded_flag.

[0159] The implementation of the flag indicating whether there is at least one non-zero transform coefficient in the CG partitioned from the transform block can be described as follows.

[0160] Figure 7 It is a view for describing a method of encoding a flag indicating whether there is a non-zero coefficient in a CG according to an embodiment of the present invention.

[0161] refer to Figure 7(a), as in the related art, when sb_coded_flag, which is information about the transform coefficient of the current block, is encoded using the width and / or height of the conventional transform block, sb_coded_flag is encoded to have values ​​of 0, 0, 1, and 1 in the reverse diagonal direction starting from the lower right corner of the transform block. In this case, since information about the right CG having no non-zero transform coefficient is also encoded, unnecessary bits are used.

[0162] refer to Figure 7 (b), as in the prior art, when sb_coded_flag (which is information about the transform coefficient of the current block) is encoded using the width and / or height of a square area (e.g., 4×4), for the upper left corner CG, sb_coded_flag is encoded to have a value of 1. In this case, sb_coded_flag may not be encoded for some CGs where non-zero transform coefficients are likely to exist with high probability. Therefore, all residual information is lost, which can lead to a significant reduction in decoding efficiency.

[0163] On the other hand, according to an embodiment of the present invention, Figure 7 As shown in (c) of FIG. 1 , the width and / or height of a K×L region may be used to encode information about the transform coefficients of the current block. Here, the values ​​of K and L may be determined using the width ZoTbW and / or height ZoTbH of the valid region after zeroing of a separable transform or a non-separable transform. Here, the values ​​of ZoTbW and ZoTbH may be the same as or different from each other.

[0164] According to the proposed method, when sb_coded_flag is encoded using ZoTbW=4 and ZbTbH=8, only "1" and "1" can be encoded for the top-left corner CG and the bottom-left corner CG. In this case, the transform coefficient information of the CG with non-zero transform coefficients can be retained. On the other hand, for the CG without non-zero transform coefficients, the transform coefficient information can be left unencoded. Therefore, decoding efficiency is improved.

[0165] According to another embodiment of the present invention, the value of K (K is the width of the predetermined area) and the value of L (L is the height of the predetermined area) can be positive integers equal to or greater than 1 predefined in the encoder and decoder. Alternatively, the arbitrary integers K and L can be determined based on one or more of the block size, the type of transform applied to the current block, the transform kernel size of the current block, the aspect ratio of the block, the quantization parameter (QP), and the prediction mode of the block. Here, the values ​​of K and L can be the same or different from each other.

[0166] According to another embodiment of the present invention, the value of K (K is the width of the predetermined region) and the value of L (L is the height of the predetermined region) may be determined based on pre-encoded syntax elements. The pre-encoded syntax elements may be at least one of syntax elements for the current block and neighboring blocks, syntax elements indicating the type of applied transform, and syntax elements related to separable transform or non-separable transform including a transform index and an index indicating transform core information / transform core list.

[0167] The proposed method encodes transform coefficient information using arbitrary regions where non-zero transform coefficients are likely to be located. Consequently, the number of bins generated thereby can be reduced, improving decoding efficiency. Furthermore, if the reduced bins are context-coded bins, the throughput of residual coding can be improved.

[0168] By using a predetermined area instead of the current block, not only the information indicating whether there is at least one non-zero transform coefficient in the CG can be encoded, but also other information about the transform coefficient can be encoded. That is, the transform coefficient related information can be encoded using arbitrary positive integers K and / or L instead of the width and / or height of the block or the width and / or height of a predetermined square area. Here, K and L can be determined based on the width and / or height of the valid area after zeroing, a predetermined arbitrary positive integer, and a previously encoded syntax element.

[0169] When transform coefficient information of a current block is encoded, a context model of the transform coefficient information may be determined based on a valid region derived from a transform result of the current block.

[0170] According to an embodiment of the present invention, information about a predetermined K×L size area determined based on various information (rather than the current block or a fixed-size N×N square area) can be used to determine the context model of the transform coefficient information. Here, the values ​​of K and L can be the same or different from each other.

[0171] According to an embodiment, a value of K (K is the width of the predetermined area) and a value of L (L is the height of the predetermined area) may be determined as the width of the current transform block and the height of the current transform block.

[0172] According to another embodiment of the present invention, the value of K (K is the width of the predetermined area) and the value of L (L is the height of the predetermined area) can be determined as the width of the valid area and the height of the valid area derived as a result of performing zeroing on the current transform block.

[0173] According to another embodiment of the present invention, the value of K (K is the width of the predetermined area) and the value of L (L is the height of the predetermined area) can be any positive integer that is a power of 2. The arbitrary positive integer can be determined as a positive integer that is a power of 2 predefined by the encoder and the decoder. Alternatively, the arbitrary positive integer can be determined based on one or more of the block size, the type of transform applied to the current block, the transform kernel size of the current block, the aspect ratio of the block, the quantization parameter (QP), and the prediction mode of the block. Here, the values ​​of K and L can be the same as or different from each other.

[0174] According to another embodiment of the present invention, the value of K (K is the width of the predetermined region) and the value of L (L is the height of the predetermined region) may be determined based on a pre-encoded syntax element. Here, the pre-encoded syntax element may be at least one of a syntax element containing information about the current block or a neighboring block, a syntax element indicating the type of transform to be applied, and a separable transform or non-separable transform-related syntax element containing a transform index and an index indicating transform core information and / or a transform core list.

[0175] According to an embodiment of the present invention, a context model of information about a transform coefficient within a current block may be determined based on a valid region.

[0176] If a non-separable transform is applied to the current block, the context model for information about the transform coefficient can be determined using adjacent transform coefficients within the valid area of ​​the current block. Specifically, during the encoding of the transform coefficient information, the adjacent transform coefficients included in the valid area of ​​the current block can be scanned. Furthermore, the context model for information about the current transform coefficient can be determined based on information about multiple adjacent transform coefficients scanned prior to the current transform coefficient. That is, when selecting multiple adjacent transform coefficients in consideration of the scan order, only the significant transform coefficients can be selected for determining the context model.

[0177] A method of determining a context model of transform coefficient information based on a valid region may be described as follows.

[0178] Figures 8 and 9 is a view illustrating an embodiment of a method for determining a context model based on a valid area according to an embodiment of the present invention.

[0179] refer to Figure 8 and Figure 9 , the size of the current block may be 8×8 and include 64 transform coefficients. Here, the valid area may be the 4×8 area on the left. The valid area may include non-zero transform coefficients. On the other hand, the transform coefficients outside the valid area have a value of 0.

[0180] According to relevant technologies, such as Figure 8As shown in (a), adjacent transform coefficients 29, 30, 36, 37, and 44 may be selected to determine the context model of the transform coefficient 28. However, here, the adjacent transform coefficients 29, 30, and 37 are transform coefficients outside the valid area, and the values ​​of the transform coefficients may be 0.

[0181] On the other hand, according to an embodiment of the present invention, Figure 8 As shown in (b), adjacent transform coefficients 33, 34, 35, 41, and 42 (which are adjacent transform coefficients scanned before transform coefficient 28) may be selected to determine the context model of transform coefficient 28. The selected adjacent transform coefficients may all be transform coefficients within the valid area.

[0182] In addition, according to related technologies, such as Figure 9 As shown in (a), adjacent transform coefficients 37, 38, 44, 45, and 52 may be selected to determine a context model for the transform coefficient 36. However, here, the transform coefficients 37, 38, and 45 are transform coefficients outside the valid region, and the values ​​of the transform coefficients may be 0.

[0183] On the other hand, according to an embodiment of the present invention, Figure 9 As shown in (b), transform coefficients 44, 51, 52, 58, and 59 (which are adjacent transform coefficients scanned before transform coefficient 36) may be selected to determine a context model for transform coefficient 36. Here, the selected transform coefficients may all be transform coefficients within the valid region.

[0184] According to another embodiment of the present invention, context information of the transform coefficient information may be determined based on the size of the valid region, as described below.

[0185] [Table 2]

[0186]

[0187] In Table 2, log2K and log2L may represent values ​​obtained by applying binary logarithms to K and L, which are width and height values ​​of the predetermined area, respectively. Here, K and L may be determined by the above method.

[0188] Referring to Table 2, context information of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix, which are syntax elements indicating position information of the last non-zero transform coefficient within a transform block, may be determined based on log2K and log2L, respectively.

[0189] By using a predetermined area instead of the current block, not only can the information indicating the position of the last non-zero transform coefficient within the transform block be selected, but also the context model of other information about the transform coefficient can be selected. That is, K or / and L can be used instead of the width or / and height of the predetermined square area to select the context model of arbitrary information. Here, K and L can be determined based on the width or / and height of the transform block, the width or / and height of the valid area after zeroing, a predetermined arbitrary positive integer (i.e., a power of 2), and a precoding syntax element.

[0190] According to an embodiment of the present invention, a context model of transform coefficient information of a current block can be selected by considering an area reflecting characteristics of the current block, thereby improving decoding efficiency.

[0191] Figure 10 is a flowchart illustrating an image decoding method according to an embodiment of the present invention. Figure 10 The image decoding method can be performed by an image decoding device.

[0192] A transform type of a current block may be determined ( S1010 ).

[0193] Context information of transform coefficient information of the current block may be determined based on the transform type ( S1020 ).

[0194] Based on the context information, the transform coefficient information may be entropy decoded ( S1030 ).

[0195] Here, the context information may be determined based on whether the transform type of the current block is a non-separable transform.

[0196] Here, the transformation coefficient information may be determined based on the transformation coefficient information of adjacent transformation coefficients adjacent to the current transformation coefficient, the adjacent transformation coefficients may be located within the valid area of ​​the current block, and the valid area of ​​the current block may be determined based on whether the transformation type of the current block is a non-separable transformation.

[0197] Here, the current block can be composed of a zeroed area and a valid area. The valid area where the non-zero transform coefficients are located can be as follows: Figure 4 and related content described.

[0198] Here, the adjacent transform coefficients can be determined based on the scanning direction of the transform of the current block. The method for determining the context information of the residual coding information can be as follows: Figures 8 and 9 and related content described.

[0199] Here, the transform coefficient information may be a syntax element binarized based on the valid region.

[0200] Here, the transformation coefficient information may include prefix information indicating a position value of the last non-zero coefficient in the current block, and suffix information indicating a position value of the last non-zero coefficient in the current block.

[0201] Here, the maximum value of the information length of the prefix indicating the position value of the last non-zero coefficient in the current block can be determined based on the valid region. The syntax element binarized based on the valid region can be as follows: Figure 5 and related content described.

[0202] Here, the transform coefficient information may indicate information about each coefficient group partitioned from the current block.

[0203] Here, the transform coefficient information may indicate whether a coefficient group includes a non-zero transform coefficient and may be encoded only for the coefficient group included in the valid area.

[0204] The syntax element indicating information about each coefficient group partitioned from the current block may be as follows Figures 6 and 7 and related content described.

[0205] at the same time, Figure 10 The steps described in the above can be performed in the same or corresponding manner in the image coding method. In addition, the bit stream can be processed by including Figure 10 The image encoding method of the steps described in the above is used to generate the bit stream. The bit stream can be stored in a non-transitory computer-readable recording medium and can also be transmitted (or streamed).

[0206] Figure 11 The following exemplifies a content streaming system to which the embodiments of the present invention can be applied.

[0207] like Figure 11 As shown, the content streaming media system applying the embodiment of the present invention may mainly include an encoding server, a streaming media server, a web server, a media storage device, a user device and a multimedia input device.

[0208] The encoding server compresses the content received from the multimedia input device (such as a smart phone, camera, CCTV, etc.) into digital data to generate a bit stream and transmits it to the streaming server. As another example, if the multimedia input device (e.g., a smart phone, camera, CCTV, etc.) directly generates a bit stream, the encoding server can be omitted.

[0209] A bit stream may be generated by applying the image encoding method and / or the image encoding device according to the embodiments of the present invention, and during transmission or reception of the bit stream, the streaming server may temporarily store the bit stream.

[0210] The streaming server transmits multimedia data to the user device via a web server based on user requests. The web server can also serve as a medium for notifying the user of any available services. When the user requests a desired service from the web server, the web server transmits it to the streaming server, which then sends the multimedia data to the user. In this case, the content streaming system may include a separate control server, which can control commands and responses between devices within the content streaming system.

[0211] The streaming media server can receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming media service, the streaming media server can store the bitstream for a certain period of time.

[0212] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, electronic tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.

[0213] Each server in the above content streaming system may operate as a distributed server, in which case the data received from each server may be distributed and processed.

[0214] The above embodiments may be performed in the same or corresponding manner in an encoding device and a decoding device. In addition, an image may be encoded / decoded using at least one of the above embodiments or a combination of at least one of the above embodiments.

[0215] The order of applying the above embodiments may be different in the encoding device and the decoding device. Alternatively, the order of applying the above embodiments may be the same in the encoding device and the decoding device.

[0216] The above embodiments may be performed for each of the luminance signal and the chrominance signal. Alternatively, the above embodiments for the luminance signal and the chrominance signal may be performed identically.

[0217] In the above embodiments, the method is described based on a flowchart having a series of steps or units, but the present invention is not limited to the order of the steps, but some steps can be performed simultaneously with other steps or in a different order. In addition, it should be understood by those skilled in the art that the steps in the flowchart are not mutually exclusive, and other steps can be added to the flowchart or some steps can be deleted from the flowchart without affecting the scope of the present invention.

[0218] The embodiments may be implemented in the form of program instructions executable by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium may include independent program instructions, data files, data structures, etc., or a combination of program instructions, data files, data structures, etc. The program instructions recorded in the computer-readable recording medium may be specially designed and constructed for the present invention, or may be well known to those skilled in the art in the field of computer software technology.

[0219] The bitstream generated by the encoding method according to the above embodiment can be stored in a non-transitory computer-readable recording medium. In addition, the bitstream stored in the non-transitory computer-readable recording medium can be decoded by the decoding method according to the above embodiment.

[0220] Examples of computer-readable recording media include: magnetic recording media such as hard disks, floppy disks, and magnetic tapes; optical data storage media such as CD-ROMs or DVD-ROMs; magneto-optical media such as magneto-optical disks; and hardware devices such as read-only memory (ROM), random access memory (RAM), flash memory, etc., which are specifically configured to store and implement program instructions. Examples of program instructions include not only machine language codes formatted by a compiler, but also high-level language codes that can be implemented by a computer using an interpreter. A hardware device can be configured to be operated by one or more software modules, or vice versa, to perform the process according to the present invention.

[0221] Although the present invention has been described with respect to specific items such as detailed elements and limited embodiments and drawings, these items are provided only to help a more comprehensive understanding of the present invention, and the present invention is not limited to the above embodiments. It will be understood by those skilled in the art that various modifications and changes can be made based on the above description.

[0222] Therefore, the spirit of the present invention should not be limited to the above-described embodiments, and the full scope of the appended claims and their equivalents should fall within the scope and spirit of the present invention.

[0223] Industrial Applicability

[0224] The present invention can be used in an image encoding / decoding device and a recording medium storing a bit stream.

Claims

1. A method for decoding an image, the method comprising: Determine the transform type of the current block; Determining context information of transform coefficient information of the current block based on the transform type; and Entropy decoding is performed on the transform coefficient information based on the context information, wherein the context information is determined based on whether the transform type of the current block is a non-separable transform.

2. The method according to claim 1, wherein The transform coefficient information is determined based on transform coefficient information of adjacent transform coefficients adjacent to the current transform coefficient, The adjacent transform coefficients are located within a valid area of ​​the current block, and the valid area of ​​the current block is determined based on whether the transform type of the current block is the inseparable transform.

3. The method according to claim 2, wherein: The current block consists of a zeroed area and the valid area.

4. The method according to claim 2, wherein: The neighboring transform coefficients are determined based on a scan direction of the transform of the current block.

5. The method according to claim 2, wherein: The transform coefficient information is binarized based on the valid area.

6. The method according to claim 5, wherein: The transform coefficient information includes prefix information indicating a position value of a last non-zero coefficient in the current block and suffix information indicating the position value of the last non-zero coefficient in the current block.

7. The method according to claim 6, wherein: A maximum value of the information length of the prefix indicating the position value of the last non-zero coefficient in the current block is determined based on the valid area.

8. The method according to claim 2, wherein: The transform coefficient information indicates information on each coefficient group (coefficient group) partitioned from the current block.

9. The method according to claim 8, wherein The transform coefficient information indicates whether the coefficient group includes a non-zero transform coefficient and is encoded only for the coefficient group included in the significant area.

10. A method for encoding an image, the method comprising: Determine the transform type of the current block; Determining context information of transform coefficient information of the current block based on the transform type; and Entropy encoding is performed on the transform coefficient information based on the context information, wherein the context information is determined based on whether the transform type of the current block is a non-separable transform.

11. A non-transitory computer-readable recording medium for storing a bit stream, the bit stream being generated by an image encoding method, wherein: The image encoding method comprises: Determine the transform type of the current block; Determining context information of transform coefficient information of the current block based on the transform type; and The transform coefficient information is entropy encoded based on the context information, and wherein the context information is determined based on whether the transform type of the current block is a non-separable transform.

12. A method for transmitting a bit stream, the bit stream being generated by an image encoding method, the method for transmitting the bit stream comprising transmitting the bit stream, in, The image encoding method comprises: Determine the transform type of the current block; Determining context information of transform coefficient information of the current block based on the transform type; and The transform coefficient information is entropy encoded based on the context information, and wherein the context information is determined based on whether the transform type of the current block is a non-separable transform.