Image encoding / decoding method, device, and recording medium storing bitstreams

WO2024205274A3PCT designated stage expired Publication Date: 2025-06-19HYUNDAI MOTOR CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/003974
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2024-03-28
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing video encoding/decoding technologies face inefficiencies due to the use of fixed-size N x N square areas in residual coding, leading to unnecessary bit consumption and suboptimal context modeling, which reduces coding efficiency as image resolution and quality increase.

Method used

The method determines a transform type for each block, calculates context information based on the transform type, and performs entropy encoding/decoding using an effective area that reflects the characteristics of the current block, reducing the number of context encoding bins and improving encoding/decoding efficiency.

Benefits of technology

This approach enhances encoding/decoding efficiency by optimizing residual coding based on the effective area, reducing unnecessary bit usage and improving context modeling, thereby improving the overall coding efficiency for high-resolution and high-quality video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024003974_19062025_PF_FP_ABST
    Figure KR2024003974_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an image encoding / decoding method, a device, a recording medium storing bitstreams, and a transmission method. The image decoding method comprises the steps of: determining the transform type of the current block; determining context information about transform coefficient information about the current block on the basis of the transform type; and entropy-decoding the transform coefficient information on the basis of the context information, wherein the context information can be determined on the basis of whether the transform type of the current block is a non-separable transform.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method, device, and recording medium storing bitstream

[0001] The present invention relates to a video encoding / decoding method, a device, and a recording medium storing a bitstream. Specifically, the present invention relates to a video encoding / decoding method, a device, and a recording medium storing a bitstream that utilizes a method of performing residual coding using information on a valid region where non-zero transform coefficients exist.

[0002] Recently, the demand for high-resolution, high-quality images, such as UHD (Ultra High Definition) images, is increasing across various application fields. As image data becomes higher in resolution and quality, the relative amount of data increases compared to conventional image data. Therefore, transmitting image data using existing media such as wired or wireless broadband lines or storing it using existing storage media leads to increased transmission and storage costs. To address these issues arising from the increasing resolution and quality of image data, high-efficiency image encoding / decoding technologies for higher-resolution and higher-quality images are required.

[0003] According to the existing residual coding technique used in image encoding / decoding technology, when performing binarization or coding of an arbitrary syntax element, an N x N square area of ​​a fixed size from a transform block or reference point is used. However, according to the existing residual coding technique, if the transform coefficient after a separable or non-separable transform exceeds the square area, unnecessary bits may be consumed.

[0004] Furthermore, conventional residual coding techniques use a fixed-size N x N square region when selecting a context model for an arbitrary syntactic element. Conventional residual coding techniques have the problem of failing to properly reflect the residual characteristics of the current block, preventing the optimal context model from being selected, and consequently reducing coding efficiency.

[0005] The purpose of the present invention is to provide a video encoding / decoding method and device with improved encoding / decoding efficiency.

[0006] In addition, the present invention aims to provide a recording medium storing a bitstream generated by an image decoding method or device according to the present invention.

[0007] In addition, the present invention aims to provide a residual coding method performed based on a valid region reflecting the characteristics of the transform coefficients of the current block in order to solve the problem of the residual coding.

[0008] An image decoding method according to one embodiment of the present invention includes a step of determining a transform type of a current block, a step of determining context information for transform coefficient information of the current block based on the transform type, and a step of entropy decoding the transform coefficient information based on the context information, wherein the context information can be determined based on whether the transform type of the current block is a non-separable transform.

[0009] In the above image decoding method, the transform coefficient information may be determined based on transform coefficient information of a neighboring transform coefficient adjacent to a current transform coefficient, the neighboring transform coefficient is located within a valid area of ​​the current block, and the valid area of ​​the current block may be determined based on whether the transform type of the current block is a non-separable transform.

[0010] In the above image decoding method, the current block may be characterized by being composed of a zero area and a valid area.

[0011] In the above image decoding method, the neighboring transform coefficient may be determined based on the scan direction for transforming the current block.

[0012] In the above image decoding method, the transform coefficient information may be characterized in that it is binarized based on the valid area.

[0013] In the above image decoding method, the transform coefficient information may include information indicating a prefix for a position value of the last non-zero coefficient in the current block, and information indicating a suffix for a position value of the last non-zero coefficient in the current block.

[0014] In the above image decoding method, the length of information indicating a prefix for the position value of the last non-zero coefficient in the current block can be set to a maximum value determined based on the valid area.

[0015] In the above image decoding method, the transform coefficient information may indicate information about each coefficient group divided from the current block.

[0016] In the above image decoding method, the transform coefficient information may indicate whether a coefficient group includes a non-zero transform coefficient, and may be encoded only for a coefficient group included within the valid area.

[0017] A video encoding method according to one embodiment of the present invention includes a step of determining a transform type of a current block, a step of determining context information for transform coefficient information of the current block based on the transform type, and a step of entropy encoding the transform coefficient information based on the context information, wherein the context information can be determined based on whether the transform type of the current block is a non-separable transform.

[0018] A non-transitory computer-readable recording medium according to one embodiment of the present invention can store a bitstream generated by a video encoding method, including the steps of determining a transform type of a current block, determining context information for transform coefficient information of the current block based on the transform type, and entropy encoding the transform coefficient information based on the context information, wherein the context information is determined based on whether the transform type of the current block is a non-separable transform.

[0019] A bitstream transmission method according to one embodiment of the present invention comprises a step of determining a transform type of a current block, a step of determining context information for transform coefficient information of the current block based on the transform type, and a step of entropy encoding the transform coefficient information based on the context information, wherein the context information is determined based on whether the transform type of the current block is a non-separable transform, and a bitstream generated by a video encoding method can be transmitted.

[0020] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.

[0021] According to the present invention, a video encoding / decoding method and device with improved encoding / decoding efficiency can be provided.

[0022] In addition, according to the present invention, a residual coding method performed based on a valid region reflecting the characteristics of the transform coefficients of the current block can be provided.

[0023] In addition, according to the present invention, the number of context encoding bins used for residual coding can be reduced.

[0024] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0025] Figure 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied.

[0026] Figure 2 is a block diagram showing the configuration according to one embodiment of a decryption device to which the present invention is applied.

[0027] FIG. 3 is a diagram schematically showing a video coding system to which the present invention can be applied.

[0028] FIG. 4 is a diagram for explaining a zeroing method in non-separable transformation according to one embodiment of the present invention.

[0029] FIG. 5 is a drawing for explaining a valid area derived from a transformation result for a current block according to an embodiment of the present invention.

[0030] FIG. 6 is a diagram illustrating one embodiment of a CG divided from a transformation block.

[0031] FIG. 7 is a diagram for explaining a method of encoding a flag indicating the presence or absence of a non-zero coefficient in a CG according to one embodiment of the present invention.

[0032] FIG. 8 is a diagram illustrating one embodiment of a method for determining a context model based on a valid area according to one embodiment of the present invention.

[0033] FIG. 9 is a diagram illustrating one embodiment of a method for determining a context model based on a valid area according to one embodiment of the present invention.

[0034] Figure 10 is a flowchart illustrating an image decoding method according to an embodiment of the present invention.

[0035] FIG. 11 is a drawing exemplifying a content streaming system to which an embodiment according to the present invention can be applied.

[0036] The present invention is susceptible to various modifications and embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and substitutes falling within the spirit and scope of the present invention. In the drawings, similar reference numerals designate the same or similar functions throughout. The shape and size of elements in the drawings may be provided by way of example only for clarity. The detailed description of the exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that the various embodiments, while different, are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the present invention. Furthermore, it should be understood that the location or arrangement of individual components within each disclosed embodiment may be modified without departing from the spirit and scope of the embodiment. Accordingly, the detailed description set forth below is not intended to be taken in a limiting sense, and the scope of the exemplary embodiments, if properly described, is defined only by the appended claims, along with the full scope equivalents to which such claims are entitled.

[0037] In the present invention, terms such as first, second, etc. may be used to describe various components, but the components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component may be referred to as the second component, and similarly, the second component may also be referred to as the first component. The term "and / or" includes a combination of multiple related described items or any of multiple related described items.

[0038] The components shown in the embodiments of the present invention are independently depicted to represent different characteristic functions, and do not mean that each component is composed of separate hardware or a single software component. That is, each component is listed and included as a separate component for convenience of explanation, and at least two components among each component may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separate embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.

[0039] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In addition, some components of the present invention are not essential components that perform essential functions in the present invention and may be optional components merely for performance enhancement. The present invention may be implemented by including only components essential to realizing the essence of the present invention, excluding components used only for performance enhancement, and a structure including only essential components, excluding optional components used only for performance enhancement, is also within the scope of the present invention.

[0040] In embodiments, the term "at least one" may mean one of a number greater than or equal to 1, such as 1, 2, 3, and 4. In embodiments, the term "a plurality of" may mean one of a number greater than or equal to 2, such as 2, 3, and 4.

[0041] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In describing the embodiments of this specification, if it is determined that a detailed description of a related known configuration or function may obscure the gist of this specification, the detailed description will be omitted. The same reference numerals will be used for identical components in the drawings, and duplicate descriptions of identical components will be omitted.

[0042] Glossary of Terms

[0043] Hereinafter, “video” may mean a single picture constituting a video, or may refer to the video itself. For example, “encoding and / or decoding of a video” may mean “encoding and / or decoding of a video,” or may mean “encoding and / or decoding of one of the videos constituting the video.”

[0044] Hereinafter, the terms "video" and "movie" may be used interchangeably and have the same meaning. Furthermore, the target image may be an encoding target image, which is the target of encoding, and / or a decoding target image, which is the target of decoding. Furthermore, the target image may be an input image input to an encoding device, or an input image input to a decoding device. Here, the target image may have the same meaning as the current image.

[0045] Hereinafter, the terms encoder and image encoding device may be used interchangeably and have the same meaning.

[0046] Hereinafter, the terms decoder and image decoding device may be used interchangeably and have the same meaning.

[0047] Hereinafter, “image”, “picture”, “frame” and “screen” may be used with the same meaning and may be used interchangeably.

[0048] Hereinafter, the term "target block" may refer to an encoding target block, which is the target of encoding, and / or a decoding target block, which is the target of decoding. Furthermore, the target block may refer to a current block, which is the target of current encoding and / or decoding. For example, the terms "target block" and "current block" may be used interchangeably and have the same meaning.

[0049] Hereinafter, "block" and "unit" may be used with the same meaning and may be used interchangeably. In addition, "unit" may mean including a luminance component block and a corresponding chroma component block to distinguish it from a block. For example, a coding tree unit (CTU) may be composed of one luma component (Y) coding tree block (CTB) and two chroma component (Cb, Cr) coding tree blocks associated with it.

[0050] Hereinafter, the terms “sample,” “pixel,” and “pixel” may be used interchangeably and have the same meaning. Here, a sample may represent a basic unit that constitutes a block.

[0051] Hereinafter, “inter” and “between screens” may be used interchangeably and have the same meaning.

[0052] Hereinafter, “intra” and “within screen” may be used interchangeably and have the same meaning.

[0053]

[0054] Figure 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied.

[0055] The encoding device (100) may be an encoder, a video encoding device, or an image encoding device. A video may include one or more images. The encoding device (100) may sequentially encode one or more images.

[0056] Referring to FIG. 1, the encoding device (100) may include an image segmentation unit (110), an intra prediction unit (120), a motion prediction unit (121), a motion compensation unit (122), a switch (115), a subtractor (113), a transformation unit (130), a quantization unit (140), an entropy encoding unit (150), an inverse quantization unit (160), an inverse transformation unit (170), an adder (117), a filter unit (180), and a reference picture buffer (190).

[0057] Additionally, the encoding device (100) can generate a bitstream including encoded information through encoding an input image and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium or can be streamed via a wired / wireless transmission medium.

[0058] The video segmentation unit (110) can segment the input video into various forms to increase the efficiency of video encoding / decoding. That is, the input video is composed of multiple pictures, and one picture can be hierarchically segmented and processed for compression efficiency, parallel processing, etc. For example, one picture can be segmented into one or more tiles or slices, which can then be segmented into multiple Coding Tree Units (CTUs). Alternatively, one picture can first be segmented into multiple sub-pictures defined as groups of rectangular slices, and each sub-picture can then be segmented into the tiles / slices. Here, the sub-pictures can be utilized to support the function of partially independently encoding / decoding and transmitting the picture. Since multiple sub-pictures can each be individually restored, there is an advantage of easy editing in applications that configure multi-channel input into a single picture. In addition, tiles can be segmented horizontally to generate bricks. Here, a brick can be utilized as the basic unit of intra-picture parallel processing. In addition, one CTU can be recursively split into a quadtree (QT), and the terminal node of the split can be defined as a coding unit (CU). The CU can be split into a prediction unit (PU) and a transformation unit (TU), and prediction and splitting can be performed. Meanwhile, the CU can be utilized as a prediction unit and / or a transformation unit itself. Here, for flexible splitting, each CTU can be recursively split into a multi-type tree (MTT) as well as a quadtree (QT). Splitting of a CTU into a multi-type tree can start from the terminal node of a QT, and the MTT can be composed of a binary tree (BT) and a triple tree (TT).For example, the MTT structure can be divided into vertical binary split mode (SPLIT_BT_VER), horizontal binary split mode (SPLIT_BT_HOR), vertical ternary split mode (SPLIT_TT_VER), and horizontal ternary split mode (SPLIT_TT_HOR). In addition, the minimum block size (MinQTSize) of the quad tree of the luminance block during splitting can be set to 16x16, the maximum block size (MaxBtSize) of the binary tree can be set to 128x128, and the maximum block size (MaxTtSize) of the triple tree can be set to 64x64. In addition, the minimum block size (MinBtSize) of the binary tree and the minimum block size (MinTtSize) of the triple tree can be set to 4x4, and the maximum depth (MaxMttDepth) of the multi-type tree can be set to 4. Additionally, to improve the encoding efficiency of the I slice, a dual tree can be applied that uses different CTU partition structures for luminance and chrominance components. On the other hand, in the P and B slices, the luminance and chrominance CTBs (Coding Tree Blocks) within the CTU can be partitioned into a single tree that shares the coding tree structure.

[0059] The encoding device (100) may perform encoding on the input image in intra mode and / or inter mode. Alternatively, the encoding device (100) may perform encoding on the input image in a third mode (e.g., IBC mode, Palette mode, etc.) other than the intra mode and inter mode. However, if the third mode has functional characteristics similar to the intra mode or inter mode, it may be classified as intra mode or inter mode for convenience of explanation. In the present invention, the third mode will be classified and described separately only when a specific description is required.

[0060] When the intra mode is used as the prediction mode, the switch (115) can be switched to intra, and when the inter mode is used as the prediction mode, the switch (115) can be switched to inter. Here, the intra mode can mean an intra-screen prediction mode, and the inter mode can mean an inter-screen prediction mode. The encoding device (100) can generate a prediction block for an input block of an input image. In addition, after the prediction block is generated, the encoding device (100) can encode a residual block using a residual of the input block and the prediction block. The input image can be referred to as a current image that is currently a target of encoding. The input block can be referred to as a current block that is currently a target of encoding or an encoding target block.

[0061] When the prediction mode is intra mode, the intra prediction unit (120) can use samples of blocks already encoded / decoded around the current block as reference samples. The intra prediction unit (120) can perform spatial prediction on the current block using the reference samples, and can generate prediction samples for the input block through spatial prediction. Here, intra prediction can mean prediction within the screen.

[0062] As an intra prediction method, non-directional prediction modes such as DC mode and Planar mode, as well as directional prediction modes (e.g., 65 directions) can be applied. Here, the intra prediction method can be expressed as an intra prediction mode or an intra-screen prediction mode.

[0063] When the prediction mode is inter mode, the motion prediction unit (121) can search for an area that best matches the input block from the reference image during the motion prediction process and derive a motion vector using the searched area. At this time, the area can be used as a search area. The reference image can be stored in the reference picture buffer (190). Here, when encoding / decoding for the reference image is processed, it can be stored in the reference picture buffer (190).

[0064] The motion compensation unit (122) can generate a prediction block for the current block by performing motion compensation using a motion vector. Here, inter prediction may mean inter-screen prediction or motion compensation.

[0065] The above motion prediction unit (121) and motion compensation unit (122) can generate a prediction block by applying an interpolation filter to a portion of the reference image when the value of the motion vector does not have an integer value. In order to perform inter-screen prediction or motion compensation, it is possible to determine whether the motion prediction and motion compensation method of the prediction unit included in the corresponding encoding unit is one of Skip Mode, Merge Mode, Advanced Motion Vector Prediction (AMVP) mode, and Intra Block Copy (IBC) mode based on the encoding unit, and perform inter-screen prediction or motion compensation according to each mode.

[0066] In addition, based on the above inter-screen prediction method, the AFFINE mode of sub-PU based prediction, the SbTMVP (Subblock-based Temporal Motion Vector Prediction) mode, and the MMVD (Merge with MVD) mode and the GPM (Geometric Partitioning Mode) mode of PU based prediction can be applied. In addition, in order to improve the performance of each mode, the HMVP (History based MVP), the PAMVP (Pairwise Average MVP), the CIIP (Combined Intra / Inter Prediction), the AMVR (Adaptive Motion Vector Resolution), the BDOF (Bi-Directional Optical-Flow), the BCW (Bi-predictive with CU Weights), the LIC (Local Illumination Compensation), the TM (Template Matching), and the OBMC (Overlapped Block Motion Compensation) can be applied.

[0067] Among these, AFFINE mode is a technology that is used in both AMVP and MERGE modes and also has high encoding efficiency. In the existing video coding standard, since MC (Motion Compensation) is performed by considering only the parallel translation of the block, there was a disadvantage in that it could not properly compensate for motions that occur in reality, such as zoom in / out and rotation. To supplement this, a 4-parameter affine motion model using two control point motion vectors (CPMV) and a 6-parameter affine motion model using three control point motion vectors can be applied to inter prediction. Here, CPMV is a vector representing the affine motion model of one of the upper left, upper right, and lower left of the current block.

[0068] The subtractor (113) can generate a residual block using the difference between the input block and the predicted block. The residual block may also be referred to as a residual signal. The residual signal may refer to the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming, quantizing, or transforming and quantizing the difference between the original signal and the predicted signal. The residual block may be a residual signal in block units.

[0069] The transform unit (130) can perform a transform on the residual block to generate a transform coefficient and output the generated transform coefficient. Here, the transform coefficient may be a coefficient value generated by performing a transform on the residual block. When the transform skip mode is applied, the transform unit (130) may also skip the transform on the residual block.

[0070] Quantized levels can be generated by applying quantization to transform coefficients or residual signals. In the following embodiments, quantized levels may also be referred to as transform coefficients.

[0071] For example, a 4x4 luminance residual block generated through within-screen prediction can be transformed using a basis vector based on DST (Discrete Sine Transform), and the remaining residual blocks can be transformed using a basis vector based on DCT (Discrete Cosine Transform). In addition, through RQT (Residual Quad Tree) technology, the transform block is divided into a quad tree shape for one block, and after performing transformation and quantization on each transform block divided through RQT, a coded block flag (cbf) can be transmitted to increase encoding efficiency when all coefficients become 0.

[0072] Another alternative is to apply Multiple Transform Selection (MTS) technology, which selectively performs transformation using multiple transformation bases. That is, instead of dividing CUs into TUs via RQT, a Sub-block Transform (SBT) technology can perform a function similar to TU division. Specifically, SBT is applied only to inter-screen prediction blocks, and unlike RQT, it can divide the current block into ½ or ¼ blocks vertically or horizontally, and then perform transformation on only one of the blocks. For example, in a vertically divided block, the transformation can be performed on the leftmost or rightmost block, and in a horizontally divided block, the transformation can be performed on the topmost or bottommost block.

[0073] Additionally, LFNST (Low Frequency Non-Separable Transform), a secondary transform technique that further transforms the residual signal converted to the frequency domain through DCT or DST, can be applied. LFNST additionally performs a transformation on the 4x4 or 8x8 low-frequency region in the upper left, which allows the residual coefficients to be concentrated in the upper left.

[0074] The quantization unit (140) can generate a quantized level by quantizing a transform coefficient or residual signal according to a quantization parameter (QP), and can output the generated quantized level. At this time, the quantization unit (140) can quantize the transform coefficient using a quantization matrix.

[0075] For example, a quantizer with QP values ​​of 0 to 51 can be used. Alternatively, if the image size is larger and high encoding efficiency is required, a QP of 0 to 63 can be used. In addition, a Dependent Quantization (DQ) method that uses two quantizers instead of a single quantizer can be applied. DQ performs quantization using two quantizers (e.g., Q0 and Q1), but even without signaling information about the use of a specific quantizer, the quantizer to be used for the next transform coefficient can be selected based on the current state through a state transition model.

[0076] The entropy encoding unit (150) can generate a bitstream by performing entropy encoding according to a probability distribution on values ​​produced by the quantization unit (140) or coding parameter values ​​produced during the encoding process, and can output the bitstream. The entropy encoding unit (150) can perform entropy encoding on information about image samples and information for decoding the image. For example, the information for decoding the image can include syntax elements, etc.

[0077] When entropy encoding is applied, a small number of bits are allocated to symbols with a high occurrence probability, and a large number of bits are allocated to symbols with a low occurrence probability, thereby representing the symbols, whereby the size of the bit string for the symbols to be encoded can be reduced. The entropy encoding unit (150) can use an encoding method such as exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), or Context-Adaptive Binary Arithmetic Coding (CABAC) for entropy encoding. For example, the entropy encoding unit (150) can perform entropy encoding using a Variable Length Coding / Code (VLC) table. In addition, the entropy encoding unit (150) may perform arithmetic encoding using the binarization method, probability model, and context model derived from the binarization method of the target symbol and the probability model of the target symbol / bin.

[0078] In this regard, when applying CABAC, the table probability update method can be changed to a simple formula-based table update method to reduce the size of the probability table stored in the decryption device. Furthermore, two different probability models can be used to obtain more accurate symbol probability values.

[0079] The entropy encoding unit (150) can change a two-dimensional block form coefficient into a one-dimensional vector form through a transform coefficient scanning method to encode a transform coefficient level (quantized level).

[0080] Coding parameters may include not only information (flags, indices, etc.) encoded in an encoding device (100) and signaled to a decoding device (200), such as syntax elements, but also information derived during an encoding or decoding process, and may mean information necessary when encoding or decoding an image.

[0081] Here, signaling a flag or index may mean that the encoder entropy encodes the flag or index and includes it in the bitstream, and that the decoder entropy decodes the flag or index from the bitstream.

[0082] The encoded current image can be used as a reference image for other images to be processed later. Accordingly, the encoding device (100) can reconstruct or decode the encoded current image again and store the reconstructed or decoded image as a reference image in the reference picture buffer (190).

[0083] The quantized level can be dequantized in the dequantization unit (160) and inversely transformed in the inverse transformation unit (170). The dequantized and / or inversely transformed coefficients can be combined with a prediction block through an adder (117), and a reconstructed block can be generated by combining the dequantized and / or inversely transformed coefficients and the prediction block. Here, the dequantized and / or inversely transformed coefficients refer to coefficients on which at least one of dequantization and inverse transformation has been performed, and may refer to a reconstructed residual block. The dequantization unit (160) and the inverse transformation unit (170) can be performed in the reverse process of the quantization unit (140) and the transformation unit (130).

[0084] The restoration block may pass through a filter unit (180). The filter unit (180) may apply a deblocking filter, a sample adaptive offset (SAO), an adaptive loop filter (ALF), a bilateral filter (BIF), a Luma Mapping with Chroma Scaling (LMCS), etc. as a filtering technique, in whole or in part, to the restoration sample, restoration block, or restoration image. The filter unit (180) may also be referred to as an in-loop filter. In this case, the in-loop filter is also used as a name excluding LMCS.

[0085] A deblocking filter can remove block distortion that occurs at the boundaries between blocks. Whether to apply a deblocking filter to the current block can be determined based on the samples contained in several columns or rows within the block. When applying a deblocking filter to a block, different filters can be applied depending on the required deblocking filtering strength.

[0086] Sample adaptive offset can be used to compensate for encoding errors by adding an appropriate offset value to sample values. Sample adaptive offset can compensate for the offset from the original image on a sample-by-sample basis for deblocked images. This can be done by dividing the samples contained in the image into a fixed number of regions, determining the regions to be offset, and applying the offset to those regions. Alternatively, the offset can be applied by considering the edge information of each sample.

[0087] Bilateral filter (BIF) can also compensate for the offset from the original image on a sample-by-sample basis for the deblocked image.

[0088] An adaptive loop filter can perform filtering based on a comparison between a reconstructed image and the original image. By dividing the samples contained in the image into predetermined groups and determining the filter to be applied to each group, filtering can be performed differentially for each group. Information regarding whether to apply an adaptive loop filter can be signaled for each coding unit (CU), and the shape and filter coefficients of the adaptive loop filter applied to each block can vary.

[0089] In LMCS (Luma Mapping with Chroma Scaling), luma mapping (LM) refers to remapping luminance values ​​through a piece-wise linear model, and chroma scaling (CS) refers to a technique that scales the residual values ​​of chrominance components according to the average luminance value of the prediction signal. In particular, LMCS can be utilized as an HDR correction technique that reflects the characteristics of HDR (High Dynamic Range) images.

[0090] The restored block or restored image that has passed through the filter unit (180) may be stored in the reference picture buffer (190). The restored block that has passed through the filter unit (180) may be a part of the reference image. In other words, the reference image may be a restored image composed of restored blocks that have passed through the filter unit (180). The stored reference image may be used for inter-screen prediction or motion compensation thereafter.

[0091] Figure 2 is a block diagram showing the configuration according to one embodiment of a decryption device to which the present invention is applied.

[0092] The decoding device (200) may be a decoder, a video decoding device, or an image decoding device.

[0093] Referring to FIG. 2, the decoding device (200) may include an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (230), an intra prediction unit (240), a motion compensation unit (250), an adder (201), a switch (203), a filter unit (260), and a reference picture buffer (270).

[0094] The decoding device (200) can receive a bitstream output from the encoding device (100). The decoding device (200) can receive a bitstream stored in a computer-readable recording medium, or a bitstream streamed through a wired / wireless transmission medium. The decoding device (200) can perform decoding on the bitstream in intra mode or inter mode. In addition, the decoding device (200) can generate a restored image or a decoded image through decoding, and can output the restored image or the decoded image.

[0095] If the prediction mode used for decryption is intra mode, the switch (203) can be switched to intra. If the prediction mode used for decryption is inter mode, the switch (203) can be switched to inter.

[0096] The decoding device (200) can decode the input bitstream to obtain a reconstructed residual block and generate a prediction block. Once the reconstructed residual block and the prediction block are obtained, the decoding device (200) can generate a reconstructed block to be decoded by adding the reconstructed residual block and the prediction block. The block to be decoded may be referred to as a current block.

[0097] The entropy decoding unit (210) can generate symbols by performing entropy decoding according to a probability distribution for the bitstream. The generated symbols may include symbols in the form of quantized levels. Here, the entropy decoding method may be the reverse process of the entropy encoding method described above.

[0098] The entropy decoding unit (210) can change a one-dimensional vector-shaped coefficient into a two-dimensional block-shaped coefficient through a transform coefficient scanning method to decode a transform coefficient level (quantized level).

[0099] The quantized level can be inversely quantized in the inverse quantization unit (220) and inversely transformed in the inverse transformation unit (230). The quantized level can be generated as a restored residual block as a result of performing inverse quantization and / or inverse transformation. At this time, the inverse quantization unit (220) can apply a quantization matrix to the quantized level. The inverse quantization unit (220) and inverse transformation unit (230) applied to the decoding device can apply the same technology as the inverse quantization unit (160) and inverse transformation unit (170) applied to the encoding device described above.

[0100] When the intra mode is used, the intra prediction unit (240) can generate a predicted block by performing spatial prediction on the current block using sample values ​​of already decoded blocks surrounding the block to be decoded. The intra prediction unit (240) applied to the decoding device can apply the same technology as the intra prediction unit (120) applied to the encoding device described above.

[0101] When the inter mode is used, the motion compensation unit (250) can generate a prediction block by performing motion compensation using a motion vector and a reference image stored in the reference picture buffer (270) on the current block. The motion compensation unit (250) can generate a prediction block by applying an interpolation filter to a portion of the reference image when the value of the motion vector does not have an integer value. In order to perform motion compensation, it is possible to determine whether the motion compensation method of the prediction unit included in the corresponding encoding unit is skip mode, merge mode, AMVP mode, or current picture reference mode based on the encoding unit, and motion compensation can be performed according to each mode. The motion compensation unit (250) applied to the decoding device can apply the same technology as the motion compensation unit (122) applied to the encoding device described above.

[0102] The adder (201) can add the restored residual block and the predicted block to generate a restored block. The filter unit (260) can apply at least one of an Inverse-LMCS, a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the restored block or restored image. The filter unit (260) applied to the decoding device can apply the same filtering technology as that applied to the filter unit (180) applied to the encoding device described above.

[0103] The filter unit (260) can output a restored image. The restored block or restored image can be stored in the reference picture buffer (270) and used for inter prediction. The restored block that has passed through the filter unit (260) can be a part of the reference image. In other words, the reference image can be a restored image composed of restored blocks that have passed through the filter unit (260). The stored reference image can be used for inter-screen prediction or motion compensation thereafter.

[0104] FIG. 3 is a diagram schematically showing a video coding system to which the present invention can be applied.

[0105] A video coding system according to one embodiment may include an encoding device (10) and a decoding device (20). The encoding device (10) may transmit encoded video and / or image information or data to the decoding device (20) in the form of a file or streaming via a digital storage medium or a network.

[0106] An encoding device (10) according to one embodiment may include a video source generation unit (11), an encoding unit (12), and a transmission unit (13). A decoding device (20) according to one embodiment may include a reception unit (21), a decoding unit (22), and a rendering unit (23). The encoding unit (12) may be referred to as a video / image encoding unit, and the decoding unit (22) may be referred to as a video / image decoding unit. The transmission unit (13) may be included in the encoding unit (12). The reception unit (21) may be included in the decoding unit (22). The rendering unit (23) may include a display unit, and the display unit may be configured as a separate device or an external component.

[0107] The video source generation unit (11) can obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source generation unit (11) can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced with a process of generating related data.

[0108] The encoding unit (12) can encode the input video / image. The encoding unit (12) can perform a series of procedures such as prediction, transformation, and quantization for compression and encoding efficiency. The encoding unit (12) can output encoded data (encoded video / image information) in the form of a bitstream. The detailed configuration of the encoding unit (12) can also be configured in the same manner as the encoding device (100) of FIG. 1 described above.

[0109] The transmission unit (13) can transmit encoded video / image information or data output in the form of a bitstream to the reception unit (21) of the decoding device (20) via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) can include an element for generating a media file through a predetermined file format and can include an element for transmission via a broadcasting / communication network. The reception unit (21) can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit (22).

[0110] The decoding unit (22) can decode video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding unit (12). The detailed configuration of the decoding unit (22) can also be configured in the same manner as the decoding device (200) of FIG. 2 described above.

[0111] The rendering unit (23) can render the decrypted video / image. The rendered video / image can be displayed through the display unit.

[0112]

[0113] Transform is a technology that converts a signal in the spatial domain into a signal in the frequency domain. To improve compression performance for high-resolution videos such as HD (high definition) or UHD (ultra-high definition) videos, the latest video compression standards support transforms for transform blocks with large sizes. For example, the H.264 / AVC standard only supported transforms for transform blocks of sizes 4x4 and 8x8, but the HEVC standard supports transforms for transform blocks of sizes from 4x4 to 32x32. In addition, the VVC standard supports transforms for transform blocks of up to 64x64 in size.

[0114] Generally, coding efficiency improves as the transform kernel size increases. However, as the transform kernel used for the transformation increases in size, computational complexity increases exponentially. Furthermore, as the transform kernel size increases, the memory required to store the kernel in the encoder and decoder also increases. Therefore, various methods are applied in video compression standards to reduce the transform kernel size and computational complexity.

[0115] As one method, a method of zeroing the high frequency transform coefficients when performing the transform can be used.

[0116] In the case of a separable transform, zeroing can be performed so that at most Y transform coefficients are left from the DC position for a transform block with a side length greater than X. Therefore, only Y x X transform kernels are required for each transform of length X (e.g., X-point DCT2, X-point DST7, X-point DCT8, etc.), which saves the memory required to store the transform kernels and reduces the amount of computation required for the transform. As a result of performing zeroing on both horizontal and vertical kernels during the transform process, non-zero transform coefficients can exist only for a certain region.

[0117] On the other hand, zeroing in non-separable transformations can be performed as follows.

[0118]

[0119] FIG. 4 is a diagram for explaining a zeroing method in non-separable transformation according to one embodiment of the present invention.

[0120] Referring to Fig. 4, the transform coefficients of the input block to be non-separably transformed can be scanned in a fixed direction and rearranged into a one-dimensional vector in the form of 1 x M. Here, the fixed direction can be one of the row-major direction, the column-major direction, and the diagonal direction. Here, M can be (TbW x TbH), which is the product of the width and height of the input block of the transformation. Alternatively, M can be the product of the width and height of a fixed region-of-interest (ROI) within the transformation block.

[0121] An M x N transform kernel can be applied to a rearranged 1 x M vector, where N can be a positive integer less than or equal to M. If M is greater than N, the high-frequency transform coefficients can be zeroed out. The result of the transform can be a one-dimensional vector of the form 1 x N.

[0122] The 1-dimensional vector in the form of 1 x N generated as a result of the transformation can be rearranged into a 2-dimensional form by scanning in a predetermined direction in block units or CG (coefficient group) units. Here, the predetermined direction can be one of the directions such as the row-major direction, the column-major direction, and the diagonal direction.

[0123] Figure 4 illustrates a case where zeroing is performed in a non-separable transform. As a result of performing zeroing in a non-separable transform, a specific region of a block may have zero or non-zero transform coefficients based on the block's (0, 0) position. On the other hand, the transform coefficients in the remaining regions may all be zero. Here, the region where non-zero transform coefficients exist can be referred to as a valid region.

[0124] For example, when performing a non-separable transform on a transform block of size 8 x 8, the transform coefficients of the 8 x 8 transform block can be scanned in a predetermined direction and rearranged into a 1 x 64 vector. Then, a 64 x 32 transform kernel can be applied to the rearranged 1 x 64 vector. The 1 x 32 vector obtained as a result of performing the transform can be rearranged into a two-dimensional form by scanning in a predetermined direction. As a result, non-zero transform coefficients can exist only in the 4 x 8 block, which is the valid area within the 8 x 8 transform block, and all transform coefficients located outside the valid area become 0.

[0125] Zeroing can reduce the computational complexity and the complexity of transformation and / or inverse transformation. Furthermore, since transformation and / or inverse transformation only require M x N (where N < M) kernels instead of M x M kernels, memory for storing transformation kernels can also be reduced.

[0126]

[0127] FIG. 5 is a drawing for explaining a valid area derived from a transformation result for a current block according to an embodiment of the present invention.

[0128] Referring to (a) of Fig. 5, the size of the current block may be TbW x TbH. And, referring to (b) of Fig. 5, the current block may be divided into a region including N transform coefficients that are 0 or not 0 and a region including only 0 transform coefficients. Here, the region including N transform coefficients that are 0 or not 0 may be referred to as a valid region. On the other hand, the region including only 0 transform coefficients may be referred to as a zeroing region. As illustrated in (b) of Fig. 5, the size of the valid region may be ZoTbW x ZoTbH.

[0129]

[0130] According to the present invention, encoding and / or decoding of transform coefficient information can be performed based on a valid area derived from a transform result for a current block. Here, the transform coefficient information can include a syntax element indicating information regarding residual coding of the current block.

[0131] That is, binarization of transform coefficient information, encoding and / or decoding of transform coefficient information can be performed based on the valid area, and a context model for transform coefficient information can be determined.

[0132]

[0133] When encoding transform coefficient information for the current block, the transform coefficient information can be binarized based on the effective area derived from the transform result of the current block.

[0134] According to conventional residual coding, transform coefficient information is binarized using the size of a transform block or an N x N square region of fixed size. For example, in indicating the position of the last non-zero transform coefficient within a transform block, information indicating the x-coordinate position of the last non-zero transform coefficient and information indicating the y-coordinate position can be independently signaled.

[0135]

[0136] Referring to Table 1, information indicating a prefix for the position of the last non-zero transform coefficient can be binarized into TU (truncated unary), and information indicating a suffix can be binarized into FLC (fixed length code). At this time, if the x-coordinate or y-coordinate of the lower right corner within a transform block or an N x N square region of fixed size is the same as the last number of the coordinate column defined in Table 1, the '0' in the parentheses shown in Table 1 may not be encoded.

[0137]

[0138] According to the present invention, instead of the width / height of a transform block or a fixed-size N x N square area, the transform coefficient information can be binarized using the width / height of a predetermined area of ​​size K x L determined based on various information. At this time, the values ​​of K and L may be the same or different.

[0139] According to one embodiment, the values ​​of the width K and the height L of a given region can be determined using the width ZoTbW and / or the height ZoTbH of the effective region derived as a result of zeroing a separable or non-separable transformation. Here, the values ​​of ZoTbW and ZoTbH can be the same or different.

[0140] The maximum length of a codeword of information indicating a prefix for the x-coordinate position of the last non-zero transform coefficient can be determined according to mathematical expression 1. And, the maximum length of a codeword of a prefix of information indicating a prefix for the y-coordinate position of the last non-zero transform coefficient can be determined according to mathematical expression 2.

[0141]

[0142]

[0143] For example, if the current transform block is 8 x 8 and the width of the valid region of the current transform block with the zeroing of the non-separable transform applied is 4 x 8, and the position of the last non-zero transform coefficient is (3, 7), the information indicating the position of the last non-zero transform coefficient can be binarized to (111, 111111). By using a specific region instead of the transform block as the valid region through the proposed method, the number of bins used to binarize the x-coordinate value 3 can be reduced compared to the prior art.

[0144] In another embodiment, when a separation transform is applied to a current block of size 32 x 32, the valid region may be a region of size 4 x 16. In the valid region, if the position of the last non-zero transform coefficient is (3, 15), the information indicating the position of the last non-zero transform coefficient may be binarized to (111, 1111111011).

[0145] Unlike the prior art, the method proposed in the present invention can set a non-square area as a valid area, thereby reducing the number of bins for binarizing the x-coordinate value 3 of the position of the last non-zero transformation coefficient.

[0146]

[0147] According to another embodiment of the present invention, the width K value of a predetermined region and the height L value of the predetermined region may be positive integers greater than or equal to 1 predefined in the encoder and the decoder. Alternatively, K and L, which are arbitrary positive integers, may be determined based on one or more pieces of information, such as the size of a block, the type of transformation applied to the current block, the size of the transformation kernel of the current block, the aspect ratio of the block, the quantization parameter (QP), and the prediction mode of the block. At this time, the values ​​of K and L may be the same or different from each other.

[0148]

[0149] According to another embodiment of the present invention, the width K value of a predetermined region and the height L value of the predetermined region can be determined based on pre-encoded syntax elements. The pre-encoded syntax elements can be at least one of a separable transformation or a non-separable transformation related syntax element including a syntax element for information of a current block or a neighboring block, a syntax element indicating a type of applied transformation, a transformation index, and an index indicating transformation kernel information / transformation kernel list.

[0150] When K, L are determined by predetermined random positive integers or pre-coded syntax elements, the maximum length of a codeword for an x-coordinate prefix can be determined according to Equation 3. In addition, the maximum length of a codeword for a y-coordinate prefix can be determined according to Equation 4.

[0151]

[0152]

[0153] The proposed method reduces the number of bins generated and improves coding efficiency by binarizing transform coefficient information using random regions where non-zero transform coefficients are likely to be located. Furthermore, if the reduced bins are context-coded, the proposed method can improve the throughput of residual coding.

[0154] By using a given region instead of the current block, not only the information indicating the position of the last non-zero transform coefficient in the current block, but also other information about the transform coefficients can be binarized. That is, other transform coefficient information can be binarized using arbitrary positive integers K or / and L instead of the width or / and height of the block or the width or / and height of a predetermined square region. Here, K and L can be determined according to the width or / and height of the effective region after zeroing, an arbitrary predetermined positive integer, or an already encoded syntax element.

[0155]

[0156] The transform coefficient information of the current block may be encoded based on the valid area derived from the transform result of the current block. The transform coefficient information of the current block may be encoded in CG units divided from the current block. One example of a CG divided from the current block may be as described below.

[0157]

[0158] FIG. 6 is a diagram illustrating one embodiment of a CG divided from a transformation block.

[0159] Referring to Fig. 6, the size of the transformation block may be 8 x 8. And, the transformation block may be divided into four CGs. The size of each CG may be 4 x 4.

[0160] Additionally, a flag indicating the presence of a non-zero transform coefficient in each CG may be encoded. The flag indicating the presence of a non-zero transform coefficient may be sb_coded_flag.

[0161] An embodiment in which a flag indicating the presence of at least one non-zero transform coefficient in CGs split from a transform block is encoded may be as described below.

[0162]

[0163] FIG. 7 is a diagram illustrating a method for encoding a flag indicating the presence or absence of a non-zero transform coefficient in a CG according to an embodiment of the present invention.

[0164] Referring to (a) of Fig. 7, when encoding sb_coded_flag, which is information on the transform coefficients of the current block, using the width and / or height of the existing transform block as in the prior art, sb_coded_flag having values ​​of '0', '0', '1', '1' in the anti-diagonal direction from the lower right of the transform block is encoded. In this case, information of the right CGs where non-zero transform coefficients are not located is also encoded, so unnecessary bits are used.

[0165] Referring to (b) of Fig. 7, when encoding sb_coded_flag, which is information on the transform coefficient of the current block, using the width and / or height of a square area (e.g., 4 x 4) as in the prior art, sb_coded_flag having a value of '1' is encoded for the CG at the upper left. In this case, sb_coded_flag may not be coded for some CGs where a non-zero transform coefficient may exist with a high probability. Therefore, all residual information is lost, which may significantly reduce the coding efficiency.

[0166] On the other hand, according to one embodiment of the present invention, information about the transform coefficients of the current block can be encoded using the width and / or height of the K x L region, as illustrated in (c) of FIG. 7. Here, the values ​​of K and L can be determined using the width ZoTbW and / or the height ZoTbH of the effective region after zeroing of the separable or non-separable transform. Here, the values ​​of ZoTbW and ZoTbH can be the same or different from each other.

[0167] When coding sb_coded_flag using ZoTbW = 4 and ZbTbH = 8 as in the proposed method, only '1' and '1' can be coded for the CGs at the upper left and lower left. In this case, the transform coefficient information of CGs with non-zero transform coefficients can be maintained. On the other hand, the transform coefficient information for CGs without non-zero transform coefficients may not be encoded. Therefore, the coding efficiency is improved.

[0168] According to another embodiment of the present invention, the width K value of a predetermined region and the height L value of the predetermined region may be positive integers greater than or equal to 1 predefined in the encoder and the decoder. Alternatively, K and L, which are arbitrary positive integers, may be determined based on one or more pieces of information, such as the size of a block, the type of transformation applied to the current block, the size of the transformation kernel of the current block, the aspect ratio of the block, the quantization parameter (QP), and the prediction mode of the block. At this time, the values ​​of K and L may be the same or different from each other.

[0169] According to another embodiment of the present invention, the width K value of a predetermined region and the height L value of the predetermined region can be determined based on pre-encoded syntax elements. The pre-encoded syntax elements can be at least one of syntax elements related to a separable transformation or a non-separable transformation, including syntax elements for a current block and surrounding blocks, syntax elements indicating a type of applied transformation, a transformation index, and an index indicating transformation kernel information / transformation kernel list.

[0170]

[0171] The proposed method reduces the number of bins generated and improves coding efficiency by encoding transform coefficient information using any region where non-zero transform coefficients can be located. Furthermore, if the reduced bins are context-coded, the throughput of residual coding can be improved.

[0172]

[0173] Instead of the current block, a given region can be used to encode not only information indicating whether there is at least one non-zero transform coefficient in the CG, but also other information about the transform coefficients. That is, instead of the width and / or height of the block or the width and / or height of a predetermined square region, the transform coefficient-related information can be encoded using arbitrary positive integers K or / and L. Here, K and L can be determined according to the width and / or height of the effective region after zeroing, an arbitrary predetermined positive integer, or a syntax element that has already been encoded.

[0174]

[0175] When encoding transform coefficient information for the current block, a context model of the transform coefficient information can be determined based on the effective area derived from the transform result of the current block.

[0176] According to one embodiment of the present invention, instead of a current block or a fixed-size N x N square region, a context model of transform coefficient information can be determined using information of a predetermined region of size K x L determined based on various information. Here, the values ​​of K and L may be the same or different from each other.

[0177] According to one embodiment, the width K value of a given region and the height L value of the given region can be determined by the width of the current transformation block and the height of the current transformation block.

[0178] According to another embodiment of the present invention, the values ​​of the width K of a given region and the height L of a given region can be determined as the width and height of a valid region derived as a result of performing zeroing on the current transformation block.

[0179] According to another embodiment of the present invention, the width K value of a predetermined region and the height L value of the predetermined region may be any positive number that is a power of 2. The arbitrary positive number may be determined as a positive integer that is a power of 2 predefined in the encoder and the decoder. Alternatively, the arbitrary positive number may be determined according to one or more pieces of information from among the size of the block, the type of transformation applied to the current block, the size of the transformation kernel of the current block, the aspect ratio of the block, the quantization parameter (QP), the prediction mode of the block, and the like. At this time, the values ​​of K and L may be the same or different from each other.

[0180] According to another embodiment of the present invention, the width K value of a predetermined region and the height L value of the predetermined region can be determined based on already encoded syntax elements. Here, the already encoded syntax elements can be at least one of separable or non-separable transformation-related syntax elements, including syntax elements for information on a current block or a neighboring block, syntax elements indicating a type of applied transformation, a transformation index, an index indicating transformation kernel information and / or a transformation kernel list, etc.

[0181]

[0182] According to one embodiment of the present invention, a context model for information of transform coefficients within a current block can be determined based on a valid area.

[0183] When a non-separable transform is applied to the current block, a context model for information about transform coefficients can be determined using neighboring transform coefficients within the valid region of the current block. Specifically, during the encoding process of transform coefficient information, neighboring transform coefficients included in the valid region of the current block can be scanned. Then, a context model for information about the current transform coefficient can be determined based on information about a plurality of neighboring transform coefficients scanned before the current transform coefficient. That is, since a plurality of neighboring transform coefficients are selected by considering the scanning order, only transform coefficients that are valid for determining the context model can be selected.

[0184] A method for determining a context model for transform coefficient information based on the valid area may be as described below.

[0185]

[0186] FIGS. 8 and 9 are diagrams illustrating an embodiment of a method for determining a context model based on a valid area according to one embodiment of the present invention.

[0187] Referring to FIGS. 8 and 9, the current block may have a size of 8 x 8 and may include 64 transform coefficients. Here, the valid region may be a 4x8 region on the left. The valid region may include non-zero transform coefficients. On the other hand, the values ​​of the transform coefficients located outside the valid region are 0.

[0188] According to the prior art, in order to determine a context model for the 28th transform coefficient as illustrated in (a) of FIG. 8, the neighboring transform coefficients 29, 30, 36, 37, and 44 may be selected. However, since the neighboring transform coefficients 29, 30, and 37 are transform coefficients outside the valid region, the values ​​of the transform coefficients may be 0.

[0189] On the other hand, according to one embodiment of the present invention, in order to determine a context model for transform coefficient 28 as illustrated in (b) of FIG. 8, neighboring transform coefficients 33, 34, 35, 41, and 42, which are neighboring transform coefficients scanned before transform coefficient 28, may be selected. Here, the selected neighboring transform coefficients may all be transform coefficients within the valid region.

[0190]

[0191] In addition, according to the prior art, neighboring transform coefficients 37, 38, 44, 45, and 52 may be selected to determine a context model for transform coefficient 36 as illustrated in (a) of FIG. 9. However, here, transform coefficients 37, 38, and 45 are transform coefficients outside the valid region, and thus the values ​​of the transform coefficients may be 0.

[0192] On the other hand, according to one embodiment of the present invention, in order to determine a context model for transform coefficient 36 as illustrated in (b) of FIG. 9, transform coefficients 44, 51, 52, 58, and 59, which are neighboring transform coefficients scanned before transform coefficient 36, may be selected. Here, the selected transform coefficients may all be transform coefficients within the valid region.

[0193]

[0194] According to another embodiment of the present invention, context information for conversion coefficient information can be determined based on the size of the valid area as follows.

[0195]

[0196] In Table 2, log2K and log2L can represent values ​​obtained by applying binary logarithm to K and L, which are the width and height values ​​of a given area. Here, K and L can be determined by the methods described above.

[0197] Referring to Table 2, the context information of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix, which are syntax elements indicating the location information of the last non-zero transform coefficient within a transform block, can be determined based on log2K and log2L, respectively.

[0198] Instead of the current block, a given region can be used to select a context model of not only information indicating the last non-zero transform coefficient position in the transform block, but also other information about the transform coefficients. That is, instead of the width or / and height of a predetermined square region, a context model of arbitrary information can be selected using K or / and L. Here, K and L can be determined according to the width or / and height of the transform block, the width or / and height of the effective region after zeroing, any predetermined positive integer that is a power of 2, or already encoded syntax elements.

[0199] According to embodiments of the present invention, a context model for transform coefficient information of a current block can be selected by considering a region reflecting the characteristics of the current block. Accordingly, coding efficiency can be improved.

[0200]

[0201] Fig. 10 is a flowchart illustrating an image decoding method according to an embodiment of the present invention. The image decoding method of Fig. 10 can be performed by an image decoding device.

[0202] The conversion type of the current block can be determined (S1010).

[0203] Context information about the transformation coefficient information of the current block can be determined based on the transformation type (S1020).

[0204] Transformation coefficient information can be entropy decoded based on context information (S1030).

[0205] Here, context information can be determined based on whether the transformation type of the current block is a non-separable transformation.

[0206] Here, the transform coefficient information is determined based on the transform coefficient information of a neighboring transform coefficient adjacent to the current transform coefficient, and the neighboring transform coefficient is located within the valid area of ​​the current block, and the valid area of ​​the current block can be determined based on whether the transform type of the current block is a non-separable transform.

[0207] Here, the current block may be composed of a zero region and a valid region. The valid region where non-zero transform coefficients are located is as described in FIG. 4 and related content.

[0208] Here, the neighboring transform coefficients can be determined based on the scan direction for transforming the current block. The method for determining context information for residual coding information is as described in FIGS. 8 and 9 and their related contents.

[0209] Here, the conversion coefficient information may be a binarized syntax element based on the valid area.

[0210] Here, the transform coefficient information may include information indicating a prefix for the position value of the last non-zero coefficient in the current block and information indicating a suffix for the position value of the last non-zero coefficient in the current block.

[0211] Here, the length of the information indicating the prefix for the position value of the last non-zero coefficient in the current block can be determined based on the valid area and set to a maximum value. The binarized syntax elements based on the valid area are as described in FIG. 5 and related content.

[0212] Here, the transformation coefficient information can indicate information about each coefficient group divided from the current block.

[0213] Here, the transform coefficient information indicates whether the coefficient group includes non-zero transform coefficients, and can be encoded only for coefficient groups included within the valid region.

[0214] The syntax elements indicating information about each of the coefficient groups divided from the current block are as described in FIGS. 6 to 7 and their related contents.

[0215] Meanwhile, the steps described in FIG. 10 can be performed identically or correspondingly in an image encoding method. Furthermore, a bitstream can be generated by an image encoding method including the steps described in FIG. 10. The bitstream can be stored on a non-transitory computer-readable recording medium and can also be transmitted (or streamed).

[0216]

[0217] FIG. 11 is a drawing exemplifying a content streaming system to which an embodiment according to the present invention can be applied.

[0218] As illustrated in FIG. 11, a content streaming system to which an embodiment of the present invention is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0219] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and CCTVs into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and CCTVs directly generate bitstreams, the encoding server may be omitted.

[0220] The above bitstream can be generated by an image encoding method and / or an image encoding device to which an embodiment of the present invention is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0221] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server may control commands / responses between each device within the content streaming system.

[0222] The streaming server can receive content from a media repository and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0223] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.

[0224] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.

[0225]

[0226] The above embodiments can be performed in the same or corresponding manner in an encoding device and a decoding device. In addition, an image can be encoded / decoded using at least one or a combination of at least one of the above embodiments.

[0227] The order in which the above embodiments are applied may be different in the encoding device and the decoding device. Alternatively, the order in which the above embodiments are applied may be the same in the encoding device and the decoding device.

[0228] The above embodiments can be performed for each of the luminance and chrominance signals. Alternatively, the above embodiments can be performed identically for the luminance and chrominance signals.

[0229] In the above embodiments, the methods are described based on a flowchart as a series of steps or units. However, the present invention is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps depicted in the flowchart are not exclusive, and that other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the present invention.

[0230] The above embodiments may be implemented in the form of program commands that can be executed by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, etc., either singly or in combination. The program commands recorded on the computer-readable recording medium may be those specifically designed and constructed for the present invention, or may be known and usable by those skilled in the art of computer software.

[0231] The bitstream generated by the encoding method according to the above embodiment can be stored in a non-transitory computer-readable recording medium. In addition, the bitstream stored in the non-transitory computer-readable recording medium can be decoded by the decoding method according to the above embodiment.

[0232] Here, examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions such as ROMs, RAMs, and flash memories. Examples of program instructions include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.

[0233] Although the present invention has been described above with specific details such as specific components and limited examples and drawings, these are provided only to help a more general understanding of the present invention, and the present invention is not limited to the above examples, and those with ordinary knowledge in the technical field to which the present invention pertains can make various modifications and variations from this description.

[0234] Therefore, the idea of ​​the present invention should not be limited to the embodiments described above, and all things that are modified equally or equivalently to the following claims as well as the claims are considered to fall within the scope of the idea of ​​the present invention.

[0235] The present invention can be used in a device for encoding / decoding an image and a recording medium storing a bitstream.

Claims

1. In the video decryption method, Step for determining the transformation type of the current block; A step of determining context information for the transformation coefficient information of the current block based on the transformation type; and A step of entropy decoding the transform coefficient information based on the context information is included, An image decoding method, characterized in that the context information is determined based on whether the transformation type of the current block is a non-separable transformation.

2. In paragraph 1, The above conversion coefficient information is determined based on the conversion coefficient information of the neighboring conversion coefficient adjacent to the current conversion coefficient, The above neighboring transformation coefficients are located within the valid area of ​​the current block, An image decoding method, characterized in that the valid area of ​​the current block is determined based on whether the transformation type of the current block is a non-separable transformation.

3. In paragraph 2, An image decoding method, characterized in that the current block is composed of a zeroing area and a valid area.

4. In paragraph 2, The above neighboring transformation coefficients are, A method characterized in that the scan direction for transformation of the current block is determined based on the scan direction.

5. In paragraph 2, The above conversion coefficient information is, An image decoding method characterized in that the image is binarized based on the above valid area.

6. In paragraph 5, The above conversion coefficient information is, An image decoding method, comprising information indicating a prefix for a position value of the last non-zero coefficient in the current block and information indicating a suffix for a position value of the last non-zero coefficient in the current block.

7. In paragraph 6, The length of the information indicating the prefix for the position value of the last non-zero coefficient in the current block is An image decoding method characterized in that a value determined based on the above valid area is set as the maximum value.

8. In paragraph 2, The above conversion coefficient information is, An image decoding method characterized by indicating information about each coefficient group divided from the current block.

9. In paragraph 8, The above conversion coefficient information is, Indicates whether the coefficient group contains non-zero transform coefficients, An image decoding method, characterized in that encoding is performed only for coefficient groups included within the above valid area.

10. In the video encoding method, Step for determining the transformation type of the current block; A step of determining context information for the transformation coefficient information of the current block based on the transformation type; and A step of entropy encoding the transform coefficient information based on the context information is included, An image encoding method, characterized in that the context information is determined based on whether the transformation type of the current block is a non-separable transformation.

11. A non-transitory computer-readable recording medium storing a bitstream generated by a video encoding method, The above image encoding method is, Step for determining the transformation type of the current block; A step of determining context information for the transformation coefficient information of the current block based on the transformation type; and A step of entropy encoding the transform coefficient information based on the context information is included, A non-transitory computer-readable recording medium, characterized in that the context information is determined based on whether the transformation type of the current block is a non-separable transformation.

12. A method for transmitting a bitstream generated by a video encoding method, The above transmission method comprises a step of transmitting the bitstream, The above image encoding method is, Step for determining the transformation type of the current block; A step of determining context information for the transformation coefficient information of the current block based on the transformation type; and A step of entropy encoding the transform coefficient information based on the context information is included, A transmission method, characterized in that the context information is determined based on whether the transformation type of the current block is a non-separable transformation.

Citation Information

Patent Citations

  • Coding code information of video data

    JP2018537908A

  • Conversion method and device in a video coding system

    KR102171362B1

  • Determining contexts for coding transform coefficient data in video coding

    KR102187013B1

  • A method and apparatus for encoding, decoding a video signal

    KR102509347B1

  • Encoding / decoding method for video signal and device therefor

    US20220394300A1