Transform coefficient level coding method and apparatus

By optimizing the binarization and decoding order of residual information and limiting flags in transform coefficient level coding, the method addresses inefficiencies in high-resolution image/video compression, enhancing coding efficiency and reducing data transmission/storage costs.

JP7775509B2Active Publication Date: 2025-11-25NOKIA TECHNOLOGIES OY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025014949
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-09-21
Filing Date
2025-01-31
Publication Date
2025-11-25
Estimated Expiration
2039-09-20

AI Technical Summary

Technical Problem

Existing image coding technologies face inefficiencies in compressing high-resolution, high-quality images and videos, particularly in handling residual coding and transform coefficient level coding, leading to increased transmission and storage costs.

Method used

The method and apparatus improve coding efficiency by performing a binarization process on residual information based on a Rice parameter, determining the decoding order of parity and transform coefficient level flags, and limiting the number of certain flags to a threshold, optimizing the encoding and decoding of transform coefficients.

Benefits of technology

This approach enhances overall image/video compression efficiency, reduces data coded based on context, and improves residual coding and transform coefficient level coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007775509000035
    Figure 0007775509000035
  • Figure 0007775509000036
    Figure 0007775509000036
  • Figure 0007775509000037
    Figure 0007775509000037
Patent Text Reader

Abstract

To provide a method and a device for enhancing the image coding efficiency.SOLUTION: A method for decoding an image performed by a decoding device according to the present disclosure comprises the steps of: receiving a bit stream including residual information; deriving a plurality of conversion factors for a conversion block on the basis of the residual information; deriving a plurality of residual samples for the conversion block by performing conversion on the basis of the plurality of conversion factors; and generating a reconstructed picture on the basis of the plurality of residual samples. The method determines a threshold on the basis of the size of the conversion block, and derives one conversion factor on the basis of a value of an effective factor flag for one conversion factor decoded on the basis of the threshold, a value of a parity level flag, a value of a first conversion factor level flag, a value of a second conversion factor level flag, a value of a residual absolute value syntax element and a value of a sign flag.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to image coding techniques, and more particularly to a transform coefficient level coding method and apparatus for an image coding system. [Background technology]

[0002] Recently, demand for high-resolution, high-quality images / videos, such as 4K or 8K or higher UHD (Ultra High Definition) images / videos, is increasing in various fields. As the resolution and quality of image / video data increases, the amount of information or bits to be transmitted increases relatively compared to existing image / video data. Therefore, when image data is transmitted using a medium such as an existing wired or wireless broadband line, or when image / video data is stored using an existing storage medium, transmission costs and storage costs increase.

[0003] In addition, interest in and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content and holograms has recently increased, and broadcasts of images / videos with image characteristics different from real images, such as game images, are on the rise.

[0004] Therefore, in order to effectively compress and transmit, store, or play back high-resolution, high-quality image / video information having the above-mentioned various characteristics, a highly efficient image / video compression technique is required. Summary of the Invention [Problem to be solved by the invention]

[0005] A technical problem of the present disclosure is to provide a method and apparatus for improving image coding efficiency.

[0006] Another technical problem of the present disclosure is to provide a method and apparatus for improving the efficiency of residual coding.

[0007] Yet another technical object of the present disclosure is to provide a method and apparatus for improving the efficiency of transform coefficient level coding.

[0008] It is still another technical object of the present disclosure to provide a method and apparatus for improving residual coding efficiency by performing a binarization process on residual information based on a Rice parameter.

[0009] Yet another technical object of the present disclosure is to provide a method and apparatus for improving coding efficiency by determining (or changing) the decoding procedure of a parity level flag regarding the parity of a transform coefficient level for a quantized transform coefficient and a first transform coefficient level flag regarding whether the transform coefficient level is greater than a first threshold (critical) value.

[0010] Another technical objective of the present disclosure is to provide a method and apparatus for improving coding efficiency by limiting the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags for quantized transform coefficients in a current block contained in residual information to a predetermined threshold or less.

[0011] Yet another technical object of the present disclosure is to provide a method and apparatus for reducing data coded based on context by limiting the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags for quantized transform coefficients in a current block included in residual information to a predetermined threshold or less. [Means for solving the problem]

[0012] According to an embodiment of the present disclosure, there is provided an image decoding method performed by a decoding device, the method including the steps of receiving a bitstream having residual information, deriving quantized transform coefficients for a current block based on the residual information included in the bitstream, deriving residual samples for the current block based on the quantized transform coefficients, and generating a reconstructed picture based on the residual samples for the current block, wherein the residual information includes a parity level flag indicating a parity of a transform coefficient level for the quantized transform coefficients and a first transform coefficient level flag indicating whether the transform coefficient level is greater than a first threshold, the deriving the quantized transform coefficients includes decoding the first transform coefficient level flag and decoding the parity level flag, and deriving the quantized transform coefficients based on the value of the decoded parity level flag and the value of the decoded first transform coefficient level flag, wherein the decoding of the first transform coefficient level flag is performed before the decoding of the parity level flag.

[0013] According to another embodiment of the present disclosure, there is provided a decoding device for performing image decoding, the decoding device including: an entropy decoding unit configured to receive a bitstream having residual information and derive quantized transform coefficients for a current block based on the residual information included in the bitstream; an inverse transform unit configured to derive residual samples for the current block based on the quantized transform coefficients; and an adder configured to generate a reconstructed picture based on the residual samples for the current block, wherein the residual information includes a parity level flag indicating a parity of a transform coefficient level for the quantized transform coefficient and a first transform coefficient level flag indicating whether the transform coefficient level is greater than a first threshold, wherein deriving the quantized transform coefficients by the entropy decoding unit includes decoding the first transform coefficient level flag and decoding the parity level flag, and deriving the quantized transform coefficients based on the value of the decoded parity level flag and the value of the decoded first transform coefficient level flag, wherein decoding the first transform coefficient level flag is performed before decoding the parity level flag.

[0014] According to yet another embodiment of the present disclosure, there is provided an image encoding method performed by an encoding device, the method including the steps of: deriving residual samples for a current block; deriving quantized transform coefficients based on the residual samples for the current block; and encoding residual information having information on the quantized transform coefficients, the residual information including a parity level flag related to a parity of a transform coefficient level for the quantized transform coefficients and a first transform coefficient level flag related to whether the transform coefficient level is greater than a first threshold, the encoding of the residual information including the steps of: deriving a value of the parity level flag and a value of the first transform coefficient level flag based on the quantized transform coefficients, encoding the first transform coefficient level flag, and encoding the parity level flag, the encoding of the first transform coefficient level flag being performed before the encoding of the parity level flag.

[0015] According to yet another embodiment of the present disclosure, there is provided an encoding device for performing image encoding, the encoding device including: a subtraction unit that derives residual samples for a current block; a quantization unit that derives quantized transform coefficients based on the residual samples for the current block; and an entropy encoding unit that encodes residual information having information about the quantized transform coefficients, the residual information including a parity level flag regarding a parity of a transform coefficient level for a quantized transform coefficient level and a first transform coefficient level flag regarding whether the transform coefficient level is greater than a first threshold, the encoding of the residual information performed by the entropy encoding unit including deriving a value of the parity level flag and a value of the first transform coefficient level flag based on the quantized transform coefficients, encoding the first transform coefficient level flag, and encoding the parity level flag, the encoding of the first transform coefficient level flag being performed prior to encoding the parity level flag.

[0016] According to yet another embodiment of the present disclosure, there is provided a decoder-readable storage medium storing information relating to instructions for causing a video decoding device to perform a decoding method according to some embodiments.

[0017] According to yet another embodiment of the present disclosure, there is provided a decoder-readable storage medium storing information relating to instructions for causing a video decoding device to perform a decoding method according to an embodiment. The decoding method according to an embodiment includes the steps of receiving a bitstream having residual information, deriving quantized transform coefficients for a current block based on the residual information included in the bitstream, deriving residual samples for the current block based on the quantized transform coefficients, and generating a reconstructed picture based on the residual samples for the current block, wherein the residual information includes a parity level flag indicating a parity of a transform coefficient level for the quantized transform coefficients and a first transform coefficient level flag indicating whether the transform coefficient level is greater than a first threshold, wherein the deriving the quantized transform coefficients includes decoding the first transform coefficient level flag and the parity level flag, and deriving the quantized transform coefficients based on the value of the decoded parity level flag and the value of the decoded first transform coefficient level flag, wherein the decoding of the first transform coefficient level flag is performed before the decoding of the parity level flag. [Effects of the Invention]

[0018] The present disclosure can improve overall image / video compression efficiency.

[0019] According to the present disclosure, the efficiency of residual coding can be improved.

[0020] According to the present disclosure, residual coding efficiency can be improved by performing a binarization process on residual information based on the Rice parameter.

[0021] The present disclosure allows for increased efficiency in transform coefficient level coding.

[0022] According to the present disclosure, residual coding efficiency can be improved by performing a binarization process on residual information based on the Rice parameter.

[0023] According to the present disclosure, coding efficiency can be improved by determining (or changing) the decoding order of a parity level flag regarding the parity of a transform coefficient level for a quantized transform coefficient and a first transform coefficient level flag regarding whether the transform coefficient level is greater than a first threshold.

[0024] According to the present disclosure, the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags for the quantized transform coefficients in the current block contained in the residual information can be limited to a predetermined threshold or less, thereby improving coding efficiency.

[0025] According to the present disclosure, the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags for the quantized transform coefficients in the current block included in the residual information can be limited to a predetermined threshold or less, thereby reducing the data coded based on the context. [Brief explanation of the drawings]

[0026] [Figure 1] FIG. 1 illustrates a schematic diagram of an example of a video / image coding system to which the present disclosure can be applied. [Figure 2] 1 is a diagram illustrating the configuration of a video / image encoding device to which the present disclosure can be applied. [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device to which the present disclosure can be applied. [Figure 4] FIG. 1 illustrates a block diagram of a CABAC encoding system according to one embodiment. [Figure 5] FIG. 10 is a diagram illustrating an example of transform coefficients in a 4×4 block. [Figure 6] FIG. 10 is a diagram illustrating an example of transform coefficients in a 2×2 block. [Figure 7]10 is a flowchart illustrating an operation of an encoding device according to an embodiment. [Figure 8] FIG. 1 is a block diagram showing a configuration of an encoding device according to an embodiment. [Figure 9] 10 is a flowchart illustrating an operation of a decoding device according to an embodiment. [Figure 10] 1 is a block diagram showing a configuration of a decoding device according to an embodiment; [Figure 11] 1 illustrates an example of a content streaming system to which the inventions disclosed in this document can be applied. DETAILED DESCRIPTION OF THE INVENTION

[0027] According to an embodiment of the present disclosure, there is provided an image decoding method performed by a decoding device, the method including the steps of receiving a bitstream including residual information, deriving quantized transform coefficients for a current block based on the residual information included in the bitstream, deriving transform coefficients for the current block from the quantized transform coefficients based on an inverse quantization process, and performing an inverse transform on the derived transform coefficients. transform) to derive residual samples for the current block, and generating a reconstructed picture based on the residual samples for the current block, wherein the residual information includes a parity level flag for parity of a transform coefficient level for the quantized transform coefficients and a first transform coefficient level flag for whether the transform coefficient level is greater than a first threshold, and the deriving the quantized transform coefficients includes decoding the first transform coefficient level flag, decoding the parity level flag, and deriving the quantized transform coefficients based on a value of the decoded parity level flag and a value of the decoded first transform coefficient level flag, and the decoding the first transform coefficient level flag is performed before the decoding the parity level flag. [Best Mode for Carrying Out the Invention]

[0028] Because the present disclosure can be modified in various ways and can have various embodiments, specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit the disclosure to the specific embodiments. Common terms used in this specification are used merely to describe specific embodiments and are not intended to limit the technical ideas of the present disclosure. A singular expression includes a plural expression unless the context clearly dictates otherwise. In this specification, terms such as "comprise" or "have" specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0029] Meanwhile, each component in the drawings described in this disclosure is shown independently for the convenience of describing different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also within the scope of the present invention as long as they do not deviate from the essence of the present invention.

[0030] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used to refer to the same components in the drawings, and duplicated descriptions of the same components will be omitted.

[0031] FIG. 1 illustrates schematically an example of a video / image coding system to which the present disclosure can be applied.

[0032] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the Versatile Video Coding (VVC) standard, the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the 2nd generation of Audio Video coding Standard (AVS2) or next generation video / image coding standards (e.g., H.267 or H.268).

[0033] In this document, various embodiments relating to video / image coding are presented, and unless otherwise stated, said embodiments may also be performed in combination with each other.

[0034] In this document, video may refer to a collection of a series of images over time. A picture generally refers to a unit that shows one image at a specific time, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more Coding Tree Units (CTUs). A picture consists of one or more slices / tiles. A picture consists of one or more tile groups. A tile group contains one or more tiles. A brick may represent a rectangular region of CTU rows within a tile in a picture. A tile may be partitioned into multiple bricks, each consisting of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick.A brick scan may indicate a specific sequential ordering of CTUs partitioning a picture, where the CTUs may be ordered consecutively in a CTU raster scan within a brick, the bricks within a tile may be ordered consecutively in a raster scan of the bricks of the tile, and the tiles within a picture may be ordered consecutively in a raster scan of the tiles of the picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set.The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the height of the picture. A tile scan indicates a specific sequential ordering of CTUs partitioning a picture, in which the CTUs are ordered consecutively in a CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of bricks of a picture that may be exclusively contained in a single NAL unit. A slice may consist of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile. In this document, the terms tile group and slice may be used interchangeably.For example, in this document, a tile group / tile group header may also be referred to as a slice / slice header.

[0035] A pixel or pel refers to the smallest unit that makes up a picture (or image). The term "sample" may also be used as a term corresponding to a pixel. A sample can generally refer to a pixel or a pixel value, or can refer to only a pixel / pixel value of a luma component, or can refer to only a pixel / pixel value of a chroma component.

[0036] A unit refers to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information about that region. One unit includes one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be mixed with terms such as block or area. In a general case, an M×N block includes a set (or array) of samples or transform coefficients consisting of M columns and N rows.

[0037] In this document, " / " and "," are interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Additionally, "A / B / C" means "at least one of A, B, and / or C." Also, "A, B, C" means "at least one of A, B, and / or C." JPEG0007775509000001.jpg22159

[0038] Additionally, in this document, "or" is interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." In other words, "or" in this document can mean "additionally or alternatively." JPEG0007775509000002.jpg28165

[0039] As shown in Figure 1, a video / image coding system includes a first device (source device) and a second device (receiving device). The source device transmits encoded video / image information or data to the receiving device in the form of a file or streaming data via a digital storage medium or a network.

[0040] The source device includes a video source, an encoding device, and a sending unit. The receiving device includes a receiving unit, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The sending unit may be included in the encoding device. The receiving unit may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.

[0041] A video source acquires video / images, such as through a video / image capture, synthesis, or generation process. A video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet, a smartphone, etc., which (electronically) generate video / images. For example, a virtual video / image may be generated by a computer, in which case a process for generating associated data may replace the video / image capture process.

[0042] An encoding device encodes input video / images. The encoding device performs a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) is output in the form of a bit stream.

[0043] The transmitter transmits the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital storage medium or a network in the form of a file or streaming. Digital storage media include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter may include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver receives / extracts the bitstream and transmits it to a decoding device.

[0044] The decoding device decodes the video / image by performing a series of steps such as inverse quantization, inverse transformation, and prediction, which correspond to the operations of the encoding device.

[0045] The renderer renders the decoded video / image, which is then displayed via the display unit.

[0046] 2 is a diagram for explaining the configuration of a video / image encoding device to which the present disclosure can be applied. Hereinafter, the term "video encoding device" may include an image encoding device.

[0047] As shown in FIG. 2, the encoding device 200 includes an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 includes an inter predictor 221 and an intra predictor 222. The residual processor 230 includes a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 further includes a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Furthermore, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The above hardware components may further include the memory 270 as an internal / external component.

[0048] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing units may be called coding units (CUs). In this case, the coding units are recursively divided from coding tree units (CTUs) or largest coding units (LCUs) according to a quad-tree, binary-tree, and ternary-tree (QTBTTT) structure. For example, one coding unit is divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or the ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. A coding procedure according to the present disclosure may be performed based on a final coding unit that is not further divided. In this case, the largest coding unit may be immediately used as the final coding unit based on coding efficiency according to image characteristics, or, if necessary, the coding unit may be recursively divided into coding units of lower depths, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the final coding unit.The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0049] The term "unit" may be mixed with terms such as "block" or "area" in some cases. In a general case, an MxN block refers to a set of samples or transform coefficients consisting of M columns and N rows. A sample generally refers to a pixel or pixel value, and may refer to only a pixel / pixel value of the luma component, or only a pixel / pixel value of the chroma component. A sample may also be used as a term corresponding to one pixel or pel of a picture (or image).

[0050] The encoding apparatus 200 subtracts a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, a unit in the encoding apparatus 200 that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) is called the subtraction unit 231. The prediction unit predicts a block to be processed (hereinafter referred to as a current block) and generates a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is to be applied for each current block or CU. The prediction unit generates various information related to prediction, such as prediction mode information, and transmits the information to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The prediction information is encoded in the entropy encoding unit 240 and output in the form of a bitstream.

[0051] The intra prediction unit 222 predicts the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away, depending on the prediction mode. In intra prediction, prediction modes include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes include, for example, DC mode and planar mode. The directional modes include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of precision of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0052] The inter prediction unit 221 derives a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. Here, to reduce the amount of motion information transmitted in inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information includes a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (such as L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The temporally surrounding blocks may be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and a reference picture including the temporally surrounding blocks may be referred to as a collocated picture (colPic). For example, the inter prediction unit 221 constructs a motion information candidate list based on the surrounding blocks and generates information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of a skip mode or a merge mode, the inter prediction unit 221 may use motion information of surrounding blocks as motion information for the current block. In the case of the skip mode, unlike in the merge mode, a residual signal may not be transmitted.In the case of the Motion Vector Prediction (MVP) mode, the motion vector of the current block can be indicated by using the motion vector of the surrounding block as a motion vector predictor and signaling the motion vector difference.

[0053] The predictor 220 can generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor can apply intra prediction or inter prediction for predicting a block, or can simultaneously apply intra prediction and inter prediction. This is called CIIP (Combined Inter and Intra Prediction). The predictor can also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode can be used for content image / video coding, such as games, such as Screen Content Coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. Palette mode can be seen as an example of intra coding or intra prediction. When palette mode is applied, sample values ​​within a picture can be signaled based on information about a palette table and a palette index.

[0054] The prediction signal generated by the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) is used to generate a reconstructed signal or a residual signal. The transform unit 232 applies a transform technique to the residual signal to generate transform coefficients. For example, the transform technique may include at least one of the following: a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. The transform process may be applied to square pixel blocks of the same size or non-square blocks of variable sizes.

[0055] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit 240 may encode information required for video / image restoration (e.g., values ​​of syntax elements) together with or separately from the quantized transform coefficients. Encoded information (e.g., encoded video / image information) is transmitted or stored in the form of a bitstream in Network Abstraction Layer (NAL) units. The video / image information may further include information on various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS). The video / image information may also include general constraint information. In this document, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium.Here, the network includes a broadcast network and / or a communication network, and the digital storage medium includes various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting the signal output from the entropy encoding unit 240 and / or a storage unit (not shown) for storing the signal may be configured as an internal / external element of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.

[0056] The quantized transform coefficients output from the quantization unit 233 are used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235. The adder 155 generates a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the current block, such as when skip mode is applied, a predicted block can be used as the reconstructed block. The adder 250 may also be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, or may be used for inter prediction of the next picture after filtering, as described below.

[0057] Meanwhile, Luma Mapping With Chroma Scaling (LMCS) can be applied during picture encoding and / or restoration.

[0058] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include, for example, deblock filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit the information to the entropy encoding unit 240, as will be described later in the description of each filtering method. The filtering information is encoded in the entropy encoding unit 240 and output in a bitstream format.

[0059] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter prediction unit 221. This allows the encoding device to avoid prediction mismatch between the encoding device 100 and the decoding device when inter prediction is applied, and also improves coding efficiency.

[0060] The DPB of the memory 270 stores the modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 stores motion information of blocks in the current picture from which motion information is derived (or encoded) and / or motion information of blocks in already reconstructed pictures. The stored motion information is transmitted to the inter predictor 221 to be used as motion information of spatially or temporally surrounding blocks. The memory 270 stores reconstructed samples of reconstructed blocks in the current picture and transmits them to the intra predictor 222.

[0061] FIG. 3 is a diagram illustrating the configuration of a video / image decoding device to which the present disclosure can be applied.

[0062] As shown in FIG. 3, the decoding device 300 includes an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 includes an inter-predictor 331 and an intra-predictor 332. The residual processor 320 includes a dequantizer 321 and an inverse transformer 322. Depending on the embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured as a single hardware component (e.g., a decoder chipset or processor). Furthermore, the memory 360 may include a decoded picture buffer (DPB) or may be configured as a digital storage medium. The above hardware components may further include a memory 360 as an internal / external component.

[0063] When a bitstream containing video / image information is input, the decoding device 300 reconstructs an image corresponding to the process by which the video / image information was processed in the encoding device of FIG. 2. For example, the decoding device 300 derives units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 performs decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding is, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximal coding unit using a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 is then reproduced by a playback device.

[0064] The decoding device 300 receives a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal is decoded by the entropy decoding unit 310. For example, the entropy decoding unit 310 parses the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The decoding device can decode pictures further based on the information on the parameter sets and / or the general constraint information. Signaling / received information and / or syntax elements, which will be described later in this document, can be obtained from the bitstream by being decoded by the decoding procedure. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients related to the residual.More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in a bitstream, determines a context model using decoding target syntax element information, decoding information of a decoding target block, or information of a symbol / bin decoded in a previous stage, predicts the occurrence probability of the bin according to the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to each syntax element. Here, the CABAC entropy decoding method determines a context model and then updates the context model using information on the decoded symbol / bin for the context model of the next symbol / bin. Information related to prediction among the information decoded by the entropy decoding unit 310 is provided to a prediction unit (an inter prediction unit 332 and an intra prediction unit 331), and residual values ​​entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, are input to a residual processing unit 320. The residual processing unit 320 may derive a residual signal (a residual block, a residual sample, a residual sample array). Information related to filtering among the information decoded by the entropy decoding unit 310 is provided to a filtering unit 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310.Meanwhile, the decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the addition unit 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.

[0065] The inverse quantization unit 321 inverse-quantizes the quantized transform coefficients and outputs the transform coefficients. The inverse quantization unit 321 rearranges the quantized transform coefficients in the form of a two-dimensional block. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit 321 inverse-quantizes the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0066] The inverse transform unit 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0067] The prediction unit performs prediction on the current block and generates a predicted block including prediction samples for the current block. The prediction unit determines whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoding unit 310, and determines a specific intra / inter prediction mode.

[0068] The prediction unit 320 generates a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The prediction unit may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but may be similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode can be seen as an example of intra coding or intra prediction. When the palette mode is applied, information regarding a palette table and a palette index may be included in the video / image information and signaled.

[0069] The intra prediction unit 331 predicts the current block by referring to samples in the current picture. The referenced samples are located either in the neighborhood of the current block or far away from it depending on the prediction mode. In intra prediction, prediction modes can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0070] The inter prediction unit 332 derives a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector in a reference picture. Here, to reduce the amount of motion information transmitted in inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 332 constructs a motion information candidate list based on the neighboring blocks and derives a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction can be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.

[0071] The adder 340 generates a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the current block, such as when skip mode is applied, the predicted block can be used as the reconstructed block.

[0072] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of the next block to be processed in the current picture, or may be output after filtering, as described below, or may be used for inter prediction of the next picture.

[0073] Meanwhile, Luma Mapping with Chroma Scaling (LMCS) may be applied during picture decoding.

[0074] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and transmit the modified reconstructed picture to the memory 360, specifically, to the DPB of the memory 360. Examples of the various filtering methods include deblock filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0075] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 stores motion information of blocks in the current picture from which motion information is derived (or decoded) and / or motion information of blocks in already reconstructed pictures. The stored motion information is transmitted to the inter prediction unit 260 to be used as motion information of spatially or temporally surrounding blocks. The memory 360 stores reconstructed samples of blocks reconstructed in the current picture and transmits them to the intra prediction unit 331.

[0076] In this specification, the embodiments described for the filtering unit 260, inter prediction unit 221 and intra prediction unit 222 of the encoding device 100 can also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332 and intra prediction unit 331 of the decoding device 300, respectively.

[0077] As described above, prediction is performed to improve compression efficiency during video coding. Accordingly, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same way in both an encoding device and a decoding device. The encoding device can improve image coding efficiency by signaling information (residual information) regarding the residual between the original block and the predicted block to the decoding device, rather than the original sample values ​​of the original block. The decoding device can derive a residual block including residual samples based on the residual information, combine the residual block with the predicted block to generate a reconstructed block including reconstructed samples, and generate a reconstructed picture including the reconstructed block.

[0078] The residual information is generated through a transform and quantization procedure. For example, an encoding device may derive a residual block between the original block and the predicted block, perform a transform procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, and signal the related residual information to a decoding device (via a bitstream). Here, the residual information includes information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding device performs an inverse quantization / inverse transform procedure based on the residual information to derive residual samples (or a residual block). The decoding device generates a reconstructed picture based on the predicted block and the residual block. The encoding device also inverse quantizes / inverse transforms the quantized transform coefficients to derive a residual block for reference for inter-prediction of a subsequent picture, and generates a reconstructed picture based on the residual block.

[0079] FIG. 4 illustrates a block diagram of a CABAC encoding system according to one embodiment.

[0080] FIG. 4 shows a block diagram of CABAC for encoding a single syntax element. During the CABAC encoding process, if the input signal is a syntax element that is not a binary value, the input signal can be converted to a binary value by binarization. If the input signal is already a binary value, it can be bypassed without binarization. Here, each binary digit 0 or 1 that constitutes the binary value may be called a bin. For example, if the binary string after binarization is 110, each of 1, 1, and 0 can be called one bin.

[0081] The binarized bins are input to a regular coding engine or a bypass coding engine. The regular coding engine assigns a context model reflecting a probability value to the bin and encodes the bin based on the assigned context model. In the regular coding engine, the probability model for each bin can be updated after encoding. Bins coded in this way are called context-coded bins. The bypass coding engine can omit the steps of estimating the probability for the input bin and updating the probability model applied to the bin after encoding. The coding speed can be improved by coding the input bins using a uniform probability distribution instead of assigning a context. Bins coded in this way are called "bypass bins."

[0082] Entropy encoding can determine whether to encode using a regular encoding engine or a bypass encoding engine, and can switch between encoding paths. Entropy decoding can perform the same process as encoding, but in reverse order.

[0083] In one embodiment, the (quantized) transform coefficients can be coded and / or decoded based on syntax elements such as transform_skip_flag, last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix, coded_sub_block_flag, sig_coeff_flag, par_level_flag, rem_abs_gt1_flag, rem_abs_gt2_flag, abs_remainder, coeff_sign_flag, mts_idx, etc. Table 1 below shows syntax elements related to residual data coding.

[0084] [Table 1] [Table 1-1]

[0085] [Table 1-2]

[0086] [Table 1-3]

[0087] [Table 1-4]

[0088] The transform_skip_flag indicates whether a transform is skipped in an associated block. The associated block may be a coding block (CB) or a transform block (TB). The terms CB and TB may be used interchangeably in the transform (and quantization) and residual coding procedures. For example, as described above, residual samples may be derived for the CB, and (quantized) transform coefficients may be derived by transforming and quantizing the residual samples. Information (e.g., synth elements) efficiently indicating the position, size, sign, etc. of the (quantized) transform coefficients may be generated and signaled in the residual coding procedure. Quantized transform coefficients may simply be referred to as transform coefficients. Generally, if the CB is not larger than the maximum TB, the size of the CB is the same as the size of the TB. In this case, the target block to be transformed (and quantized) and residual coded may be referred to as a CB or a TB. On the other hand, if the CB is larger than the maximum TB, the target block to be transformed (and quantized) and residual coded may be referred to as a TB. Hereinafter, it will be described that syntax elements related to residual coding are signaled in units of transform blocks (TBs), but this is merely an example, and as mentioned above, the TBs can be mixed with coding blocks (CBs).

[0089] In one embodiment, the (x, y) position information of the last non-zero transform coefficient in a transform block may be coded based on the syntax elements last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix and last_coeff_y_suffix. More specifically, last_sig_coeff_x_prefix indicates the prefix of the column position of the last significant coefficient in the scanning order within the transform block, last_sig_coeff_y_prefix indicates the prefix of the row position of the last significant coefficient in the scanning order within the transform block, last_sig_coeff_x_suffix indicates the suffix of the column position of the last significant coefficient in the scanning order within the transform block, and last_sig_coeff_y_suffix indicates the suffix of the row position of the last significant coefficient in the scanning order within the transform block. Here, the significant coefficient may indicate the non-zero coefficient. The scan order may be a diagonal scan order from top right to bottom right. Alternatively, the scan order may be a horizontal scan order or a vertical scan order. The scan order may be determined based on whether intra / inter prediction is applied to the current block (CB or CB including TB) and / or a specific intra / inter prediction mode.

[0090] Next, after dividing the transform block into 4x4 sub-blocks, a 1-bit syntax element coded_sub_block_flag can be used for each 4x4 sub-block to indicate whether there are any non-zero coefficients in the current sub-block.

[0091] If the value of coded_sub_block_flag is 0, there is no further information to transmit, so the encoding process for the current sub-block can be terminated. Conversely, if the value of coded_sub_block_flag is 1, the encoding process for sig_coeff_flag can be continued. Since the last sub-block containing a non-zero coefficient does not require coding for coded_sub_block_flag, and since sub-blocks containing DC information of transform blocks are likely to contain non-zero coefficients, coded_sub_block_flag can be assumed to be 1 without being coded.

[0092] If it is determined that a non-zero coefficient exists in the current sub-block because the value of coded_sub_block_flag is 1, sig_coeff_flag, which has a binary value, can be coded in reverse according to the scanned order. A 1-bit syntax element sig_coeff_flag can be coded for each coefficient in the scanned order. If the value of a transform coefficient at the current scanned position is non-zero, the value of sig_coeff_flag can be 1. For a sub-block containing the last non-zero coefficient, sig_coeff_flag does not need to be coded for the last non-zero coefficient, so the coding process for the sub-block may be omitted. Level information coding can be performed only when sig_coeff_flag is 1, and four syntax elements can be used in the level information coding process. More specifically, each sig_coeff_flag[xC][yC] can indicate whether the level (value) of a corresponding transform coefficient at each transform coefficient position (xC, yC) in the current TB is non-zero. In one embodiment, the sig_coeff_flag may be an example of a significant coefficient flag indicating whether the quantized transform coefficient is a significant coefficient other than zero.

[0093] The remaining level value after encoding for sig_coeff_flag is expressed as follows: That is, the syntax element remAbsLevel indicating the level value to be encoded is expressed as follows: Here, coeff means the actual transform coefficient value.

[0094] [Formula 1] remAbsLevel = |coeff| - 1

[0095] Using par_level_flag, the least significant coefficient (LSB) value of remAbsLevel described in Equation 1 can be coded as shown in Equation 2 below. Here, par_level_flag[n] can indicate the parity of the transform coefficient level (value) at scan position n. After coding par_level_flag, the transform coefficient level value remAbsLevel to be coded can be updated as shown in Equation 3 below.

[0096] [Formula 2] par_level_flag = remAbsLevel & 1

[0097] [Formula 3] remAbsLevel' = remAbsLevel >> 1

[0098] rem_abs_gt1_flag may indicate whether remAbsLevel' at the corresponding scan position (n) is greater than 1, and rem_abs_gt2_flag may indicate whether remAbsLevel' at the corresponding scan position (n) is greater than 2. Only when rem_abs_gt2_flag is 1, can coding be performed for abs_remainder. The relationship between coeff, which is an actual transform coefficient value, and each syntax element can be summarized as shown in Equation 4 below, and Table 2 below shows an example related to Equation 4. In addition, the sign of each coefficient can be coded using coeff_sign_flag, which is a 1-bit symbol. |coeff| indicates the transform coefficient level (value) and may be expressed as AbsLevel for the transform coefficient.

[0099] [Formula 4] | coeff | = sig_coeff_flag + par_level_flag + 2 * (rem_abs_gt1_flag + rem_abs_gt2_flag + abs_remainder)

[0100] [Table 2]

[0101] On the other hand, in one embodiment, the par_level_flag indicates an example of a parity level flag regarding the parity of the transform coefficient level for the quantized transform coefficient, the rem_abs_gt1_flag indicates an example of a first transform coefficient level flag regarding whether the transform coefficient level is greater than a first threshold, and the rem_abs_gt2_flag indicates an example of a second transform coefficient level flag regarding whether the transform coefficient level is greater than a second threshold.

[0102] In another embodiment, rem_abs_gt2_flag may be called rem_abs_gt3_flag, and in yet another embodiment, rem_abs_gt1_flag and rem_abs_gt2_flag may appear based on abs_level_gtx_flag[n][j]. abs_level_gtx_flag[n][j] may be a flag indicating whether the absolute value of the transform coefficient level (or the value obtained by shifting the transform coefficient level by 1 to the right) at scan position n is greater than (j<<1)+1. In one example, rem_abs_gt1_flag may perform the same and / or similar function as abs_level_gtx_flag[n][0], and rem_abs_gt2_flag may perform the same and / or similar function as abs_level_gtx_flag[n][1]. That is, the abs_level_gtx_flag[n][0] may be an example of the first transform coefficient level flag, and the abs_level_gtx_flag[n][1] may be an example of the second transform coefficient level flag. The (j<<1)+1 may be replaced with a predetermined threshold such as a first threshold or a second threshold, depending on the case.

[0103] FIG. 5 is a diagram showing an example of transform coefficients in a 4×4 block.

[0104] The 4x4 block in Figure 5 shows an example of quantized coefficients. The block shown in Figure 5 may be a 4x4 transform block or a 4x4 sub-block of an 8x8, 16x16, 32x32, or 64x64 transform block. The 4x4 block in Figure 5 may represent a luma block or a chroma block. The encoding results for the reverse diagonal scanned coefficients in Figure 5 are shown in Table 3, for example. In Table 3, scan_pos indicates the position of the coefficient according to the reverse diagonal scan. scan_pos15 indicates the coefficient scanned first in the 4x4 block, i.e., the bottom right corner, and scan_pos0 indicates the coefficient scanned last, i.e., the top left corner. Meanwhile, in one embodiment, scan_pos may also be referred to as a scan position. For example, scan_pos0 may be referred to as scan position 0.

[0105] [Table 3]

[0106] As described in Table 1, in one embodiment, main syntax elements for a 4x4 sub-block unit include sig_coeff_flag, par_level_flag, rem_abs_gt1_flag, rem_abs_gt2_flag, abs_remainder, coeff_sign_flag, etc. Among these, sig_coeff_flag, par_level_flag, rem_abs_gt1_flag, and rem_abs_gt2_flag indicate context coding bins that are coded using a regular coding engine, and abs_remainder and coeff_sign_flag indicate bypass bins that are coded using a bypass coding engine.

[0107] Context coding bins may exhibit high data dependency because they use updated probability states and ranges while processing previous bins. That is, because the context coding bins can encode / decode the next bin only after the current bin has been fully encoded / decoded, parallel processing may be difficult. In addition, reading the probability intervals and determining the current state may require a lot of time. Therefore, in one embodiment, a method is proposed to improve the CABAC processing amount by reducing the number of context coding bins and increasing the number of bypass bins.

[0108] In one embodiment, coefficient level information may be coded in reverse scan order. That is, the coefficients are coded starting from the coefficient at the bottom right of the unit block and scanned toward the top left. In one example, coefficient levels scanned earlier in the reverse scan order may indicate smaller values. Signaling the par_level_flag, rem_abs_gt1_flag, and rem_abs_gt2_flag for such earlier scanned coefficients may reduce the length of the binarized bins used to indicate coefficient levels, and each syntax element may be efficiently coded using arithmetic coding based on a previously coded context using a defined context.

[0109] However, for some coefficient levels with large values, i.e., coefficient levels located at the upper left corner of a unit block, signaling par_level_flag, rem_abs_gt1_flag, and rem_abs_gt2_flag may not help improve compression performance, and using par_level_flag, rem_abs_gt1_flag, and rem_abs_gt2_flag may reduce coding efficiency.

[0110] In one embodiment, the number of context coding bins can be reduced by quickly switching syntax elements (par_level_flag, rem_abs_gt1_flag, rem_abs_gt2_flag) that are coded in context coding bins to the abs_remainder syntax element, which is coded using the bypass coding engine, i.e., coded in the bypass bins.

[0111] In one embodiment, the number of coefficients for which rem_abs_gt2_flag is coded can be limited. The maximum number of explicitly coded rem_abs_gt2_flag in a 4x4 block can be 16. That is, rem_abs_gt2_flag may be coded for all coefficients whose absolute value is greater than 2. However, in one embodiment, rem_abs_gt2_flag may be coded only for the first N coefficients in the scan order that have an absolute value greater than 2 (i.e., coefficients whose rem_abs_gt1_flag is 1). N may be selected by the encoder or may be set to any value between 0 and 16. Table 4 shows an application example of the above embodiment when N is 1. According to one embodiment, in a 4x4 block, the number of coding operations for rem_abs_gt2_flag can be reduced by the number of Xs in Table 4 below, thereby reducing the number of context coding bins. When comparing the abs_remainder value of the coefficient for the scan position where coding for rem_abs_gt2_flag is not performed with Table 3, it is changed as shown in Table 4 below.

[0112] [Table 4]

[0113] In one embodiment, the number of coefficients for which coding is performed for rem_abs_gt1_flag is limited. The number of explicitly coded rem_abs_gt2_flag in a 4x4 block can be up to 16. That is, rem_abs_gt1_flag is coded for all coefficients whose absolute value is greater than 0, but in one embodiment, rem_abs_gt1_flag can be coded only for the first M coefficients (i.e., coefficients whose sig_coeff_flag is 1) that have an absolute value greater than 0 according to the scan order. M can be selected by the encoder or can be set to any value between 0 and 16. Table 5 shows an example of application of the above embodiment when M is 4. If rem_abs_gt1_flag is not coded, rem_abs_gt2_flag is also not coded. Therefore, according to the above embodiment, the number of coding operations for rem_abs_gt1_flag can be reduced by the number of Xs in a 4x4 block, and therefore the number of context coding bins can be reduced. For scan positions where coding for rem_abs_gt1_flag is not performed, the values ​​of rem_abs_gt2_flag and abs_remainder of the coefficients are changed as shown in Table 6 below, when compared with Table 3. Table 6 shows an application example of the above embodiment when M is 8.

[0114] [Table 5]

[0115] [Table 6]

[0116] In one embodiment, the above-described embodiment in which the number of rem_abs_gt1_flag and rem_abs_gt2_flag is limited may be combined. Both M, which indicates the limit on the number of rem_abs_gt1_flag, and N, which indicates the limit on the number of rem_abs_gt2_flag, may have values ​​from 0 to 16, but N may be equal to or smaller than M. Table 7 shows an application example of this embodiment when M is 8 and N is 1. Since the syntax element is not coded at the position marked with X, the number of context coding bins can be reduced.

[0117] [Table 7]

[0118] In one embodiment, the number of coefficients for which par_level_flag is coded may be limited. The maximum number of par_level_flag explicitly coded in a 4x4 block may be 16. That is, par_level_flag may be coded for all coefficients whose absolute value is greater than 0. However, in one embodiment, par_level_flag may be coded only for the first L coefficients (i.e., coefficients whose sig_coeff_flag is 1) having an absolute value greater than 0 in the scan order. L may be selected by the encoder or may be set to any value between 0 and 16. Table 8 shows an application example of the above embodiment when L is 8. According to the above embodiment, the number of coding operations for par_level_flag may be reduced by the number indicated as X in the 4x4 block, thereby reducing the number of context coding bins. For scan positions where coding for par_level_flag is not performed, the values ​​of rem_abs_gt1_flag, rem_abs_gt2_flag, and abs_remainder of the coefficients are changed as shown in Table 8 below, when compared with Table 3.

[0119] [Table 8]

[0120] In one embodiment, the above-described embodiment of limiting the number of par_level_flag and rem_abs_gt2_flag can be combined. Both L, which indicates the limit on the number of par_level_flag, and N, which indicates the limit on the number of rem_abs_gt2_flag, can have values ​​from 0 to 16. Table 9 shows an application example of this embodiment when L is 8 and N is 1. Since the syntax element is not coded at the position marked with X, the number of context coding bins can be reduced.

[0121] [Table 9]

[0122] In one embodiment, the above-described embodiment of limiting the number of par_level_flag and rem_abs_gt1_flag can be combined. Both L, which indicates the limit on the number of par_level_flag, and M, which indicates the limit on the number of rem_abs_gt1_flag, can have values ​​from 0 to 16. Table 10 shows an application example of this embodiment when L is 8 and M is 8. Since the syntax element is not coded at the position marked with X, the number of context coding bins can be reduced.

[0123] [Table 10]

[0124] In one embodiment, the above-described embodiments for limiting the number of par_level_flag, rem_abs_gt1_flag, and rem_abs_gt2_flag may be combined. L, which indicates the limit on the number of par_level_flag, M, which indicates the limit on the number of rem_abs_gt1_flag, and N, which indicates the limit on the number of rem_abs_gt2_flag, may all have values ​​from 0 to 16, but N may be equal to or less than M. Table 11 shows an application example of this embodiment when L is 8, M is 8, and N is 1. Since the syntax element is not coded at the position marked with X, the number of context coding bins can be reduced.

[0125] [Table 11]

[0126] In one embodiment, a method of limiting the sum of the numbers of sig_coeff_flag, par_level_flag, and rem_abs_gt1_flag is proposed. When the sum of the numbers of sig_coeff_flag, par_level_flag, and rem_abs_gt1_flag is limited to K, K can have a value from 0 to 48. In this embodiment, if the sum of the numbers of sig_coeff_flag, par_level_flag, and rem_abs_gt1_flag exceeds K and no coding is performed for sig_coeff_flag, par_level_flag, and rem_abs_gt1_flag, rem_abs_gt2_flag may also not be coded. Table 12 shows the case where K is limited to 30.

[0127] [Table 12]

[0128] In one embodiment, the method of limiting the sum of the numbers of sig_coeff_flag, par_level_flag, and rem_abs_gt1_flag can be combined with the method of limiting the number of rem_abs_gt2_flag described above. When the sum of the numbers of sig_coeff_flag, par_level_flag, and rem_abs_gt1_flag is limited to K and the number of rem_abs_gt2_flag is limited to N, K can have a value from 0 to 48, and N can have a value from 0 to 16. Table 13 shows the case where K is limited to 30 and N is limited to 2.

[0129] [Table 13]

[0130] In one embodiment, encoding may be performed in the order of par_level_flag and rem_abs_gt1_flag, but in an embodiment according to the present disclosure, the encoding order may be changed so that encoding is performed in the order of rem_abs_gt1_flag and par_level_flag. If the encoding order of par_level_flag and rem_abs_gt1_flag is changed in this way, rem_abs_gt1_flag is encoded after sig_coeff_flag, and par_level_flag is encoded only when rem_abs_gt1_flag is 1. Therefore, the relationship between coeff, which is an actual transform coefficient value, and each syntax element may be changed as shown in Equation 5 below. Table 14 below shows an example of a case where the encoding order of par_level_flag and rem_abs_gt1_flag is changed. Compared to Table 2 above, in Table 14 below, par_level_flag is not coded when |coeff| is 1, which may be advantageous in terms of processing load and coding. Of course, in Table 14, when |coeff| is 2, rem_abs_gt2_flag needs to be coded unlike Table 2, and when |coeff| is 4, abs_remainder needs to be coded unlike Table 2. However, since |coeff| is 1 more often than |coeff| is 2 or 4, Table 14 can exhibit higher processing load and coding performance than Table 2. The result of coding the 4x4 sub-block as shown in FIG. 5 is shown in Table 15 below.

[0131] [Formula 5] | coeff | = sig_coeff_flag + rem_abs_gt1_flag + par_level_flag + 2 * (rem_abs_gt2_flag + abs_remainder)

[0132] [Table 14]

[0133] [Table 15]

[0134] In one embodiment, when encoding is performed in the order of sig_coeff_flag, rem_abs_gt1_flag, par_level_flag, rem_abs_gt2_flag, abs_remainder, and coeff_sig_flag, a method is proposed to limit the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag. When the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag is limited to K, K can have a value from 0 to 48. In this embodiment, when sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag are not further encoded, rem_abs_gt2_flag may also not be encoded. Table 16 shows the case where K is limited to 25.

[0135] [Table 16]

[0136] In one embodiment, sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag can be coded within one for loop in the (residual) syntax. Since the sum of the numbers of the three syntax elements (sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag) does not exceed K, coding can be stopped at the same scan position even if the sum of the numbers of the three syntax elements does not exactly equal K. Table 17 below shows the case where K is limited to 27. When coding is performed up to scan position 3, the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag is 25. Although this value does not exceed K, the encoding device does not know the coefficient level value of the next scan position, scan_pos2, and therefore does not know whether the number of context coding bins generated at scan_pos2 will have a value of 1, 2, or 3. Therefore, the encoding device can encode only up to scan_pos3 and then end the encoding. Although the K value is different, the encoding results can be the same in Table 16 above and Table 17 below.

[0137] [Table 17]

[0138] In one embodiment, a method is proposed in which the coding order of par_level_flag and rem_abs_gt1_flag is changed, but the number of rem_abs_gt2_flag is limited. That is, coding is performed in the order of sig_coeff_flag, rem_abs_gt1_flag, par_level_flag, rem_abs_gt2_flag, abs_remainder, and coeff_sign_flag, but the number of coefficients coded by rem_abs_gt2_flag is limited.

[0139] The maximum number of rem_abs_gt2_flag explicitly coded in a 4x4 block is 16. That is, rem_abs_gt2_flag can be coded for all coefficients whose absolute value is greater than 2. In contrast, in this embodiment, rem_abs_gt2_flag can be coded only for the first N coefficients (i.e., coefficients whose rem_abs_gt1_flag is 1) having an absolute value greater than 2 according to the scan order. N can be selected by the encoder or set to any value between 0 and 16. Table 18 below shows an application example of this embodiment when N is 1. The number of times rem_abs_gt2_flag is coded can be reduced by the number indicated as X in a 4x4 block, thereby reducing the number of context coding bins. For scan positions where rem_abs_gt2_flag is not coded, the abs_remainder value of the coefficient is changed as shown in Table 18 below compared to Table 15.

[0140] [Table 18]

[0141] In one embodiment, the encoding order of par_level_flag and rem_abs_gt1_flag is changed, but a method is provided for limiting the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag and the number of rem_abs_gt2_flag. In one example, when encoding is performed in the order of sig_coeff_flag, rem_abs_gt1_flag, par_level_flag, rem_abs_gt2_flag, abs_remainder, and coeff_sig_flag, the method for limiting the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag can be combined with the method for limiting the number of rem_abs_gt2_flag. When the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag is limited to K, and the number of rem_abs_gt2_flag is limited to N, K can have values ​​from 0 to 48, and N can have values ​​from 0 to 16. Table 19 shows the case where K is limited to 25 and N is limited to 2.

[0142] [Table 19]

[0143] FIG. 6 is a diagram illustrating an example of transform coefficients in a 2×2 block.

[0144] The 2x2 block in Figure 6 shows an example of quantized coefficients. The block shown in Figure 6 may be a 2x2 transform block or a 2x2 sub-block of a 4x4, 8x8, 16x16, 32x32, or 64x64 transform block. The 2x2 block in Figure 6 shows a luma block or a chroma block. The coding result for the reverse diagonal scanned coefficients in Figure 6 is shown in Table 20, for example. In Table 20, scan_pos indicates the position of the coefficient according to the reverse diagonal scan. scan_pos3 indicates the coefficient that is scanned first in the 2x2 block, i.e., the bottom right corner, and scan_pos0 indicates the coefficient that is scanned last, i.e., the top left corner.

[0145] [Table 20]

[0146] As described in Table 1, in one embodiment, main syntax elements for a 2x2 sub-block unit include sig_coeff_flag, par_level_flag, rem_abs_gt1_flag, rem_abs_gt2_flag, abs_remainder, coeff_sign_flag, etc. Among them, sig_coeff_flag, par_level_flag, rem_abs_gt1_flag, and rem_abs_gt2_flag indicate context coding bins that are coded using a regular coding engine, and abs_remainder and coeff_sign_flag indicate bypass bins that are coded using a bypass coding engine.

[0147] Context coding bins exhibit high data dependency because they use updated probability states and ranges while processing previous bins. That is, because the context coding bins can only encode / decode the next bin after the current bin has been fully encoded / decoded, parallel processing may be difficult. In addition, reading the probability intervals and determining the current state may require a lot of time. Therefore, in one embodiment, a method is proposed to improve the CABAC processing amount by reducing the number of context coding bins and increasing the number of bypass bins.

[0148] In one embodiment, coefficient level information may be coded in reverse scan order. That is, coefficients are coded starting from the coefficient at the bottom right of the unit block and scanned toward the top left. In one example, coefficient levels scanned earlier in the reverse scan order may indicate smaller values. Signaling the par_level_flag, rem_abs_gt1_flag, and rem_abs_gt2_flag for coefficients scanned earlier in this manner may reduce the length of the binarization bins used to indicate coefficient levels, and each syntax element may be efficiently coded by arithmetic coding based on a previously coded context using a defined context.

[0149] However, for some coefficient levels with large values, i.e., coefficient levels located at the upper left corner of a unit block, signaling par_level_flag, rem_abs_gt1_flag, and rem_abs_gt2_flag may not help improve compression performance, and using par_level_flag, rem_abs_gt1_flag, and rem_abs_gt2_flag may also reduce coding efficiency.

[0150] In one embodiment, the number of context coding bins can be reduced by quickly switching syntax elements (par_level_flag, rem_abs_gt1_flag, rem_abs_gt2_flag) that are coded in context coding bins to abs_remainder syntax elements that are coded using a bypass coding engine, i.e., coded in bypass bins.

[0151] In one embodiment, the number of coefficients coded by rem_abs_gt2_flag may be limited. The maximum number of explicitly coded rem_abs_gt2_flag in a 2x2 block may be four. That is, rem_abs_gt2_flag may be coded for all coefficients whose absolute value is greater than 2. However, in one embodiment, rem_abs_gt2_flag may be coded only for the first N coefficients (i.e., coefficients whose rem_abs_gt1_flag is 1) having an absolute value greater than 2 in the scan order. N may be selected by the encoder or set to any value between 0 and 4. If the encoder limits the context coding bins for luma or chroma 4x4 sub-blocks in a manner similar to this embodiment, N may be calculated using the limit value used. As a method for calculating N, the context coding bin limit value (N4×4) for a luminance or chrominance 4×4 sub-block is used as is as shown in Equation 6, or since the number of pixels in a 2×2 sub-block is 4, N is calculated using Equation 7. Here, a and b are constants and are not limited to specific values ​​in the present disclosure.

[0152] Similarly, N can be calculated using the horizontal / vertical size values ​​of the sub-block. Since the sub-block is square, the horizontal size value and the vertical size value are the same. Since the horizontal or vertical size value of a 2x2 sub-block is 2, N can be calculated using Equation 8.

[0153] [Formula 6] N = N4x4

[0154] [Formula 7] N = {N 4x4 >> (4 - a)} + b

[0155] [Formula 8] N = {N 4x4 >> (a - 2)} + b

[0156] Table 21 shows an application example of this embodiment when N is 1. The number of coding operations for rem_abs_gt2_flag can be reduced by the number of Xs in a 2x2 block, thereby reducing the number of context coding bins. Comparing the abs_remainder values ​​of coefficients for scan positions where rem_abs_gt2_flag is not coded with Table 20, the values ​​are changed as shown in Table 21 below.

[0157] [Table 21]

[0158] In 2x2 sub-block coding of chrominance blocks according to one embodiment, the sum of the numbers of sig_coeff_flag, par_level_flag, and rem_abs_gt1_flag may be limited. When the sum of the numbers of sig_coeff_flag, par_level_flag, and rem_abs_gt1_flag is limited to K, K may have a value from 0 to 12. In this embodiment, when the sum of sig_coeff_flag, par_level_flag, and rem_abs_gt1_flag does not exceed K and is not coded, rem_abs_gt2_flag may not be coded either.

[0159] K can be selected by the encoder or set to any value between 0 and 12. When the context coding bins for the luma or chroma 4x4 sub-blocks are restricted in the encoder using a method similar to this embodiment, K can be calculated using a restriction value used. As a method for calculating K, the context coding bin restriction value (K 4×4 ) is used as it is, or since the number of 2×2 subblock pixels is 4, K is calculated by Equation 10. Here, a and b are constants and are not limited to specific values ​​in the present disclosure.

[0160] Similarly, K can be calculated using the horizontal / vertical size values ​​of the sub-block. Since the sub-block is square, the horizontal size value and the vertical size value are the same. Since the horizontal or vertical size value of a 2x2 sub-block is 2, K can be calculated using Equation 11.

[0161] [Formula 9] K = K 4x4

[0162] [Formula 10] K = {K 4x4 >> (4 - a)} + b

[0163] [Formula 11] K = {K 4x4 >> (a - 2)} + b

[0164] Table 22 shows the case where K is limited to 6.

[0165] [Table 22]

[0166] In one embodiment, the method of limiting the sum of the numbers of sig_coeff_flag, par_level_flag, and rem_abs_gt1_flag and the method of limiting the number of rem_abs_gt2_flag described above can be combined. When the sum of the numbers of sig_coeff_flag, par_level_flag, and rem_abs_gt1_flag is limited to K and the number of rem_abs_gt2_flag is limited to N, K can have a value from 0 to 12, and N can have a value from 0 to 4.

[0167] K and N can be determined in the encoder or calculated based on the method described above in connection with Equations 6 to 11.

[0168] Table 23 shows the case where K is limited to 6 and N is limited to 1.

[0169] [Table 23]

[0170] In one embodiment, a method for changing the encoding order of par_level_flag and rem_abs_gt1_flag is proposed. More specifically, instead of encoding par_level_flag and rem_abs_gt1_flag in this embodiment, when encoding a 2×2 size sub-block of a chrominance block, encoding is performed in the order of rem_abs_gt1_flag and par_level_flag. By changing the encoding order of par_level_flag and rem_abs_gt1_flag, rem_abs_gt1_flag is encoded after sig_coeff_flag, and par_level_flag can be encoded only when rem_abs_gt1_flag is 1. Therefore, the relationship between coeff, which is an actual transform coefficient value, and each syntax element is changed as shown in Equation 12 below.

[0171] [Formula 12] |coeff| = sig_coeff_flag + rem_abs_gt1_flag + par_level_flag + 2 * (rem_abs_gt2_flag + abs_remainder)

[0172] Table 24 below shows some examples related to Equation 12. Compared with Table 2, according to Table 24, when |coeff| is 1, par_level_flag is not coded, which can be advantageous in terms of throughput and coding. Of course, when |coeff| is 2, rem_abs_gt2_flag needs to be coded, unlike Table 2, and when |coeff| is 4, abs_remainder needs to be coded, unlike Table 2. However, since |coeff| is 1 occurs more frequently than when |coeff| is 2 or 4, the method according to Table 24 can exhibit higher throughput and coding performance than the method according to Table 2. Table 25 shows the result of coding the 4x4 sub-block as shown in FIG. 6.

[0173] [Table 24]

[0174] [Table 25]

[0175] In one embodiment, a method is proposed in which the encoding order of par_level_flag and rem_abs_gt1_flag is changed, but the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag is limited. For example, when encoding is performed in the order of sig_coeff_flag, rem_abs_gt1_flag, par_level_flag, rem_abs_gt2_flag, abs_remainder, and coeff_sign_flag, the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag is limited. When the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag is limited to K, K can have a value from 0 to 12. K can be selected by the encoder and can be set to any value from 0 to 12. Also, it can be calculated based on the above-described contents in relation to Equations 9 to 11.

[0176] In one embodiment, when sig_coeff_flag, rem_abs_gt1_flag and par_level_flag are not further coded, rem_abs_gt2_flag may also not be coded. Table 26 shows the case where K is limited to 6.

[0177] [Table 26]

[0178] In one embodiment, sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag can be coded within a single for loop. The sum of the three syntax elements (sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag) does not exceed K, so coding can be stopped at the same scan position even if the sum does not exactly match K. Table 27 shows the case where K is limited to 8. When coding is performed up to scan position 2 (scan_pos2), the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag is 6. Although this value does not exceed K, the encoder does not know the coefficient level value of the next scan position 1 (scan_pos1), and therefore does not know whether the number of context coding bins generated at scan_pos1 will have a value of 1, 2, or 3. Therefore, coding can be terminated by coding only up to scan_pos2. Therefore, although the values ​​of K are different, the encoding results may be the same in Table 26 and Table 27.

[0179] [Table 27]

[0180] In one embodiment, a method is proposed in which the coding order of par_level_flag and rem_abs_gt1_flag is changed, but the number of rem_abs_gt2_flag is limited. For example, when coding is performed in the order of sig_coeff_flag, rem_abs_gt1_flag, par_level_flag, rem_abs_gt2_flag, abs_remainder, and coeff_sign_flag, the number of coefficients coded by rem_abs_gt2_flag is limited.

[0181] In one embodiment, the maximum number of rem_abs_gt2_flag coded in a 2x2 block is 4. That is, rem_abs_gt2_flag is coded for all coefficients whose absolute value is greater than 2. In contrast, another embodiment of the present disclosure proposes a method of coding rem_abs_gt2_flag only for the first N coefficients having an absolute value greater than 2 according to the scan order (i.e., coefficients whose rem_abs_gt1_flag is 1).

[0182] N can be selected by the encoder or can be set to any value between 0 and 4. It can also be calculated in the manner described above in relation to Equations 6 to 8.

[0183] Table 28 shows an application example of this embodiment when N is 1. The number of times rem_abs_gt2_flag is coded can be reduced by the number of Xs in a 4x4 block, thereby reducing the number of context coding bins. Compared with Table 25, the abs_remainder value of the coefficient for the scan position where coding for rem_abs_gt2_flag is not performed is changed as shown in Table 28 below.

[0184] [Table 28]

[0185] In one embodiment, a method is provided in which the encoding order of par_level_flag and rem_abs_gt1_flag is changed, but the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag and the number of rem_abs_gt2_flag are limited. For example, when encoding is performed in the order of sig_coeff_flag, rem_abs_gt1_flag, par_level_flag, rem_abs_gt2_flag, abs_remainder, and coeff_sign_flag, the method of limiting the number of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag can be combined with the method of limiting the number of rem_abs_gt2_flag. When the sum of the numbers of sig_coeff_flag, rem_abs_gt1_flag, and par_level_flag is limited to K and the number of rem_abs_gt2_flag is limited to N, K can have a value from 0 to 12 and N can have a value from 0 to 4. K and N can be selected by the encoder, or K can be set to any value from 0 to 12 and N can be set to any value from 0 to 4. They can also be calculated in the manner described above in relation to Equations 6 to 11.

[0186] Table 29 below shows the case where K is limited to 6 and N is limited to 1.

[0187] [Table 29]

[0188] FIG. 7 is a flowchart showing the operation of the encoding device according to one embodiment, and FIG. 8 is a block diagram showing the configuration of the encoding device according to one embodiment.

[0189] The encoding apparatus according to Figures 7 and 8 can perform operations corresponding to those of the decoding apparatus according to Figures 9 and 10. Therefore, the operations of the decoding apparatus described below with reference to Figures 9 and 10 can be similarly applied to the encoding apparatus according to Figures 7 and 8.

[0190] Each step disclosed in Fig. 7 may be performed by the encoding device 200 disclosed in Fig. 2. More specifically, S700 is performed by the subtraction unit 231 disclosed in Fig. 2, S710 is performed by the quantization unit 233 disclosed in Fig. 2, and S720 is performed by the entropy encoding unit 240 disclosed in Fig. 2. In addition, the operations according to S700 to S720 are based on part of the content described above with reference to Figs. 4 to 6. Therefore, the description of specific content that overlaps with the content described above with reference to Figs. 2 and 4 to 6 will be omitted or simplified.

[0191] 8, the encoding device according to one embodiment includes a subtraction unit 231, a transformation unit 232, a quantization unit 233, and an entropy encoding unit 240. However, in some cases, not all of the components shown in FIG. 8 may be essential components of the encoding device, and the encoding device may be realized with more or fewer components than those shown in FIG. 8.

[0192] In an encoding device according to one embodiment, the subtraction unit 231, the transformation unit 232, the quantization unit 233, and the entropy encoding unit 240 are each implemented on a separate chip, or at least two or more components are implemented on one chip.

[0193] An encoding apparatus according to an embodiment derives residual samples for a current block (S700). More specifically, a subtraction unit 231 of the encoding apparatus derives residual samples for the current block.

[0194] According to an embodiment, an encoding apparatus derives quantized transform coefficients based on residual samples for a current block (S710). More specifically, a quantization unit 233 of the encoding apparatus derives quantized transform coefficients based on residual samples for the current block.

[0195] According to an embodiment, the encoding apparatus encodes residual information including information about the quantized transform coefficients (S720). More specifically, the entropy encoding unit 240 of the encoding apparatus encodes the residual information including information about the quantized transform coefficients.

[0196] In one embodiment, the residual information includes a parity level flag regarding the parity of a transform coefficient level for the quantized transform coefficient and a first transform coefficient level flag regarding whether the transform coefficient level is greater than a first threshold. In one example, the parity level flag indicates par_level_flag, the first transform coefficient level flag indicates rem_abs_gt1_flag or abs_level_gtx_flag[n][0], and the second transform coefficient level flag indicates rem_abs_gt2_flag or abs_gtx_flag[n][1].

[0197] In one embodiment, the step of encoding the residual information includes the steps of deriving a value of the parity level flag and a value of the first transform coefficient level flag based on the quantized transform coefficients, and encoding the first transform coefficient level flag and encoding the parity level flag.

[0198] In one embodiment, the encoding of the first transform coefficient level flag may be performed before the encoding of the parity level flag, for example, the encoding device may encode rem_abs_gt1_flag or abs_level_gtx_flag[n][0] before encoding par_level_flag.

[0199] In one embodiment, the residual information further includes a significance flag indicating whether the quantized transform coefficient is a significance coefficient other than 0 and a second transform coefficient level flag regarding whether the transform coefficient level of the quantized transform coefficient is greater than a second threshold. In one example, the significance flag indicates sig_coeff_flag.

[0200] In one embodiment, the sum of the numbers of significant coefficient flags, first transform coefficient level flags, parity level flags, and second transform coefficient level flags for quantized transform coefficients in the current block included in the residual information may be equal to or less than a predetermined threshold. In one example, the sum of the numbers of significant coefficient flags, first transform coefficient level flags, parity level flags, and second transform coefficient level flags for quantized transform coefficients associated with a current sub-block in the current block may be limited to be equal to or less than a predetermined threshold.

[0201] In one embodiment, the predetermined threshold is determined based on the size of the current block (or a current sub-block within the current block).

[0202] In one embodiment, the sum of the number of the significant coefficient flags, the number of the first transform coefficient level flags, and the number of the parity level flags included in the residual information is less than or equal to a third threshold, the number of the second transform coefficient level flags included in the residual information is less than or equal to a fourth threshold, and the predetermined threshold indicates the sum of the third threshold and the fourth threshold.

[0203] In one example, when the size of the current block or the current sub-block within the current block is 4x4, the third threshold value represents K, where K represents any one value from 0 to 48, and the fourth threshold value represents N, where N represents any one value from 0 to 16.

[0204] In another example, when the size of the current block or the current sub-block within the current block is 2×2, the third threshold value represents K, where K represents any one value from 0 to 12, and the fourth threshold value represents N, where N represents any one value from 0 to 4.

[0205] In one embodiment, when the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags and the number of second transform coefficient flags derived based on the 0th quantized transform coefficient to the nth quantized transform coefficient determined by the coefficient scanning order reaches the predetermined threshold, explicit signaling of the significant coefficient flag, the first transform coefficient level flag, the parity level flag and the second transform coefficient level is omitted for the n+1th quantized transform coefficient determined by the coefficient scanning order, and the value of the n+1th quantized transform coefficient is derived based on the value of coefficient level information included in the residual information.

[0206] For example, when the sum of the number of the sig_coeff_flag, the number of the rem_abs_gtx_flag[0], the number of the par_level_flag, and the number of the rem_abs_gt2_flag (or abs_level_gtx_flag[n][1]) derived based on the 0th quantized transform coefficient (or the first quantized transform coefficient) to the nth quantized transform coefficient (or the (n+1)th quantized transform coefficient) determined by the coefficient scan order reaches the predetermined threshold, For the first quantized transform coefficient, explicit signaling of sig_coeff_flag, rem_abs_gt1_flag (or abs_level_gtx_flag[n][0]), par_level_flag, abs_level_gtx_flag[n][1] and rem_abs_gt2_flag (or abs_level_gtx_flag[n][1]) is omitted, and the value of the n+1th quantized transform coefficient is derived based on the value of abs_remainder or dec_abs_level contained in the residual information.

[0207] In one embodiment, the significant coefficient flag, the first transform coefficient level flag, the parity level flag and the second transform coefficient level flag included in the residual information are coded based on a context, and the coefficient level information is coded based on a bypass.

[0208] According to the encoding device and the operating method of the encoding device of Figures 7 and 8, the encoding device derives residual samples for a current block (S700), derives quantized transform coefficients based on the residual samples for the current block (S710), and encodes residual information including information about the quantized transform coefficients (S720), where the residual information includes a parity level flag regarding the parity of a transform coefficient level for the quantized transform coefficients and a first transform coefficient level flag regarding whether the transform coefficient level is greater than a first threshold, and the step of encoding the residual information includes a step of deriving values ​​of the parity level flag and the first transform coefficient level based on the quantized transform coefficients, encoding the first transform coefficient level flag, and encoding the parity level flag, and the step of encoding the first transform coefficient level flag is characterized in that the step of encoding the first transform coefficient level flag is performed before the step of encoding the parity level flag. That is, according to the present disclosure, the decoding order of the parity level flag regarding the parity of the transform coefficient level for the quantized transform coefficient and the first transform coefficient level flag regarding whether the transform coefficient level is greater than the first threshold can be determined (or changed) to improve coding efficiency.

[0209] FIG. 9 is a flowchart showing the operation of the decoding device according to one embodiment, and FIG. 10 is a block diagram showing the configuration of the decoding device according to one embodiment.

[0210] The steps disclosed in Fig. 9 are performed by the decoding device 300 disclosed in Fig. 3. More specifically, steps S900 and S910 are performed by the entropy decoding unit 310 disclosed in Fig. 3, step S920 is performed by the inverse quantization unit 321 and / or the inverse transform unit 322 disclosed in Fig. 3, and step S930 is performed by the addition unit 340 disclosed in Fig. 3. In addition, the operations of steps S900 to S930 are based on part of the content described above with reference to Figs. 4 to 6. Therefore, the description of specific content that overlaps with the content described above with reference to Figs. 3 to 6 will be omitted or simplified.

[0211] 10, a decoding device according to one embodiment includes an entropy decoding unit 310, an inverse quantization unit 321, an inverse transform unit 322, and an addition unit 340. However, in some cases, all of the components shown in FIG. 10 may not be essential components of the decoding device, and the decoding device may be realized with more or fewer components than those shown in FIG.

[0212] In one embodiment of the decoding device, the entropy decoding unit 310, the inverse quantization unit 321, the inverse transform unit 322, and the addition unit 340 may each be implemented by a separate chip, or at least two or more components may be implemented by one chip.

[0213] A decoding device according to an embodiment receives a bitstream including residual information (S900). More specifically, an entropy decoding unit 310 of the decoding device receives the bitstream including residual information.

[0214] According to an embodiment, a decoding device derives quantized transform coefficients for a current block based on residual information included in the bitstream (S910). More specifically, an entropy decoding unit 310 of the decoding device derives quantized transform coefficients for the current block based on residual information included in the bitstream.

[0215] According to an embodiment, a decoding device derives residual samples for a current block based on the quantized transform coefficients (S920). More specifically, an inverse quantization unit 321 of the decoding device derives transform coefficients from the quantized transform coefficients through an inverse quantization process, and an inverse transform unit 322 of the decoding device inversely transforms the transform coefficients to derive residual samples for the current block.

[0216] The decoding apparatus according to an embodiment generates a reconstructed picture based on the residual samples for the current block (S930). More specifically, the adder 340 of the decoding apparatus generates the reconstructed picture based on the residual samples for the current block.

[0217] In one embodiment, the residual information includes a parity level flag indicating a parity of a transform coefficient level for the quantized transform coefficient and a first transform coefficient level flag indicating whether the transform coefficient level is greater than a first threshold. In one example, the parity level flag indicates par_level_flag, the first transform coefficient level flag indicates rem_abs_gt1_flag or abs_level_gtx_flag[n][0], and the second transform coefficient level flag indicates rem_abs_gt2_flag or abs_gtx_flag[n][1].

[0218] In one embodiment, the step of deriving the quantized transform coefficients includes the steps of decoding the transform coefficient level flags and decoding the parity level flags, and deriving the quantized transform coefficients based on the values ​​of the decoded parity level flags and the decoded first transform coefficient level flags.

[0219] In one embodiment, the step of decoding the first transform coefficient level flag may be performed before the step of decoding the parity level flag, for example, the decoding device may perform decoding for rem_abs_gt1_flag or abs_level_gtx_flag[n][0] before decoding for par_level_flag.

[0220] In one embodiment, the residual information may further include a significance coefficient flag indicating whether the quantized transform coefficient is a significance coefficient other than 0 and a second transform coefficient level flag regarding whether the transform coefficient level of the quantized transform coefficient is greater than a second threshold. In one example, the significance coefficient flag indicates sig_coeff_flag.

[0221] In one embodiment, the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags for the quantized transform coefficients in the current block included in the residual information may be less than or equal to a predetermined threshold. In one example, the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags for the quantized transform coefficients associated with a current sub-block in the current block is limited to be less than or equal to a predetermined threshold.

[0222] In one embodiment, the predetermined threshold is determined based on the size of the current block (or a current sub-block within the current block).

[0223] In one embodiment, the sum of the number of the significant coefficient flags, the number of the first transform coefficient level flags, and the number of the parity level flags included in the residual information is less than or equal to a third threshold, the number of the second transform coefficient level flags included in the residual information is less than or equal to a fourth threshold, and the predetermined threshold indicates the sum of the third threshold and the fourth threshold.

[0224] In one example, when the size of the current block or the current sub-block within the current block is 4x4, the third threshold value represents K, where K represents any one value from 0 to 48, and the fourth threshold value represents N, where N represents any one value from 0 to 16.

[0225] In another example, when the size of the current block or the current sub-block within the current block is 2×2, the third threshold value represents K, where K represents any one of values ​​from 0 to 12, and the fourth threshold value represents N, where N represents any one of values ​​from 0 to 4.

[0226] In one embodiment, when the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags and the number of second transform coefficient level flags derived based on the 0th quantized transform coefficient to the nth quantized transform coefficient determined by the coefficient scanning order reaches the predetermined threshold, explicit signaling of the significant coefficient flag, the first transform coefficient level flag, the parity level flag and the second transform coefficient level flag is omitted for the n+1th quantized transform coefficient determined by the coefficient scanning order, and the value of the n+1th quantized transform coefficient is derived based on the value of coefficient level information included in the residual information.

[0227] For example, when the sum of the number of the sig_coeff_flag, the number of the rem_abs_gtx_flag (or the abs_level_gtx_flag[n][0]), the number of the par_level_flag, and the number of the rem_abs_gt2_flag (or the abs_level_gtx_flag[n][1]) derived based on the 0th quantized transform coefficient (or the 1st quantized transform coefficient) to the nth quantized transform coefficient (or the (n+1)th quantized transform coefficient) determined by the coefficient scan order reaches the predetermined threshold, For the (n+1)th quantized transform coefficient determined by the scan order, explicit signaling of sig_coeff_flag, rem_abs_gt1_flag (or abs_level_gtx_flag[n][0]), par_level_flag, abs_level_gtx_flag[n][1] and rem_abs_gt2_flag (or abs_level_gtx_flag[n][1]) is omitted, and the value of the (n+1)th quantized transform coefficient is derived based on the value of abs_remainder or dec_abs_level included in the residual information.

[0228] In one embodiment, the significant coefficient flag, the first transform coefficient level flag, the parity level flag and the second transform coefficient level flag included in the residual information are coded based on a context, and the coefficient level information is coded based on a bypass.

[0229] According to the decoding device and the operating method of the decoding device disclosed in Figures 9 and 10, the decoding device receives a bitstream including residual information (S900), derives quantized transform coefficients for a current block based on the residual information included in the bitstream (S910), derives residual samples for the current block based on the quantized transform coefficients (S920), and generates a reconstructed picture based on the residual samples for the current block (S930), wherein the residual information is parity-based regarding the parity of transform coefficient levels for the quantized transform coefficients. and a parity level flag indicating whether the transform coefficient level is greater than a first threshold, and a first transform coefficient level flag indicating whether the transform coefficient level is greater than a first threshold, and the step of deriving the quantized transform coefficients includes the steps of decoding the transform coefficient level flag, decoding the parity level flag, and deriving the quantized transform coefficients based on the value of the decoded parity level flag and the decoded first transform coefficient level flag, wherein the step of decoding the first transform coefficient level flag is performed before the step of decoding the parity level flag. That is, according to the present disclosure, the decoding order of the parity level flag indicating the parity of the transform coefficient level for the quantized transform coefficients and the first transform coefficient level flag indicating whether the transform coefficient level is greater than a first threshold can be determined (or changed) to improve coding efficiency.

[0230] In the above-described embodiments, the method is described based on a flowchart as a series of steps or blocks, but the present disclosure is not limited to the order of steps, and certain steps may occur in different steps and orders or simultaneously than those described above. Also, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included, or one or more steps of the flowchart may be deleted without affecting the scope of the present disclosure.

[0231] The above-described method according to the present disclosure may be implemented in the form of software, and the encoding device and / or decoding device according to the present disclosure may be included in an image processing device such as a TV, a computer, a smartphone, a set-top box, or a display device.

[0232] In the present disclosure, when an embodiment is implemented in software, the method described above can be realized with modules (processes, functions, etc.) that perform the functions described above. The modules can be stored in memory and executed by a processor. The memory can be located inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory can include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation can be stored on a digital storage medium.

[0233] In addition, the decoding device and encoding device to which the present disclosure is applied can be included in multimedia broadcast transmitting / receiving devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communications, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providing devices, OTT (Over-the-Top video) devices, Internet streaming service providing devices, three-dimensional (3D) video devices, VR (Virtual Reality) devices, AR (Augmented Reality) devices, image telephone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), medical video devices, etc., and can be used to process video signals and data signals. For example, OTT (Over-the-Top video) devices include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.

[0234] Furthermore, the processing method to which the present disclosure is applied can be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to the present invention can also be stored in a computer-readable recording medium (computer-readable storage medium). The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, Blu-ray Discs (BDs), Universal Serial Buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission via the Internet). The bitstream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0235] Furthermore, the embodiments of the present disclosure may be realized as a computer program product with program code, which may be stored on a computer-readable carrier when executed on a computer according to the embodiments of the present disclosure.

[0236] FIG. 11 shows an example of a content streaming system to which the inventions disclosed in this document can be applied.

[0237] As shown in FIG. 11, the content streaming system to which the present disclosure is applied includes an encoding server, a streaming server, a web server, a media storage device (repository), a user device, and a multimedia input device.

[0238] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or video camera into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or camcorder directly generates a bitstream, the encoding server may be omitted.

[0239] The bitstream is generated by an encoding method or a bitstream generation method to which the present disclosure is applied, and the stream server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0240] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary for informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. Here, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.

[0241] The streaming server receives content from a media storage device and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.

[0242] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), navigation systems, slate PCs, tablet PCs, ULTRABOOK (registered trademark), wearable devices (e.g., smartwatches, smart glasses, and HMDs (Head Mounted Displays)), digital TVs, desktop computers, and digital signage.

[0243] Each server in the content streaming system can be operated as a distributed server, and in this case, data received by each server can be processed in a distributed manner.

Claims

1. An image decoding method performed by a decoding device, comprising: receiving a bitstream including residual information and prediction information; deriving quantized transform coefficients for a current block based on the residual information included in the bitstream; deriving residual samples for the current block based on the quantized transform coefficients; deriving a predicted sample for the current block based on the prediction information; generating a reconstructed picture based on the residual samples for the current block and the prediction samples for the current block; the residual information includes a significant coefficient flag, a parity level flag, a first transform coefficient level flag, a second transform coefficient level flag, an abs_remainder syntax element, and a sign flag for quantized transform coefficients for the current block; the significant coefficient flag relates to whether the quantized transform coefficient is a significant coefficient other than zero; the parity level flag relates to a parity of a transform coefficient level for the quantized transform coefficient; the first transform coefficient level flag relates to whether the transform coefficient level is greater than a first threshold; the second transform coefficient level flag relates to whether the transform coefficient level is greater than a second threshold; The abs_remainder syntax element relates to the remaining value of the transform coefficient level, the sign flag relates to the sign of the quantized transform coefficient; The step of deriving the quantized transform coefficients comprises: determining a threshold value associated with a sum of a number of significant coefficient flags, a number of first transform coefficient level flags, a number of parity level flags, and a number of second transform coefficient level flags for the transform coefficients in the current block based on a size of the current block, the threshold is a single threshold for all of the significant coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level flag, and is not determined as a sum of two or more thresholds; decoding the significance coefficient flag, the first transform coefficient level flag, the parity level flag, the second transform coefficient level flag, the abs_remainder syntax element, and the sign flag; deriving the quantized transform coefficients based on the value of the decoded significance coefficient flag, the value of the decoded parity level flag, the value of the decoded first transform coefficient level flag, the value of the decoded second transform coefficient level flag, the value of the decoded abs_remainder syntax element, and the value of the decoded sign flag; the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags for the quantized transform coefficients in the current block is less than or equal to the threshold; The method, wherein the significant coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level flag are included in the residual information.

2. 2. The method of claim 1, wherein, when the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags derived based on the 0th quantized transform coefficient to the nth quantized transform coefficient determined by the coefficient scan order reaches the threshold, the value of the n+1th quantized transform coefficient is derived based on the value of coefficient level information included in the residual information.

3. the significant coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level flag included in the residual information are decoded based on a context; The method of claim 2 , wherein the coefficient level information is decoded based on a bypass.

4. An image encoding method performed by an encoding device, comprising: deriving a predicted sample for the current block; deriving residual samples for the current block based on the predicted samples for the current block; deriving quantized transform coefficients based on the residual samples for the current block; encoding residual information including information about the quantized transform coefficients and prediction information related to deriving the prediction samples for the current block; the residual information includes a significant coefficient flag, a parity level flag, a first transform coefficient level flag, a second transform coefficient level flag, an abs_remainder syntax element, and a sign flag for quantized transform coefficients for the current block; the significant coefficient flag relates to whether the quantized transform coefficient is a significant coefficient other than zero; the parity level flag relates to a parity of a transform coefficient level for the quantized transform coefficient; the first transform coefficient level flag relates to whether the transform coefficient level is greater than a first threshold; the second transform coefficient level flag relates to whether the transform coefficient level is greater than a second threshold; The abs_remainder syntax element relates to the remaining value of the transform coefficient level, the sign flag relates to the sign of the quantized transform coefficient; The step of encoding the residual information comprises: determining a threshold value associated with a sum of a number of significant coefficient flags, a number of first transform coefficient level flags, a number of parity level flags, and a number of second transform coefficient level flags for the transform coefficients in the current block based on a size of the current block, the threshold is a single threshold for all of the significant coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level flag, and is not determined as a sum of two or more thresholds; deriving a value of the significance flag, a value of the parity level flag, a value of the first transform coefficient level flag, a value of the second transform coefficient level flag, a value of the abs_remainder syntax element, and a value of the sign flag based on the quantized transform coefficients; encoding the significance coefficient flag, the first transform coefficient level flag, the parity level flag, the second transform coefficient level flag, the abs_remainder syntax element, and the sign flag; the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags for the quantized transform coefficients in the current block is less than or equal to the threshold; The method, wherein the significant coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level flag are included in the residual information.

5. 5. The method of claim 4, wherein, when the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags derived based on the 0th quantized transform coefficient to the nth quantized transform coefficient determined by the coefficient scan order reaches the threshold, the value of the n+1th quantized transform coefficient is derived based on the value of coefficient level information included in the residual information.

6. the significant coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level flag included in the residual information are encoded based on a context; The method of claim 5 , wherein the coefficient level information is encoded based on a bypass.

7. A method for transmitting data relating to an image, comprising: obtaining a bitstream relating to the image, the bitstream comprising: deriving a predicted sample for the current block; deriving residual samples for the current block based on the predicted samples for the current block; deriving quantized transform coefficients based on the residual samples for the current block; generating the bitstream by encoding residual information including information about the quantized transform coefficients and prediction information related to deriving the prediction samples for the current block; transmitting the data including the bitstream; the residual information includes a significant coefficient flag, a parity level flag, a first transform coefficient level flag, a second transform coefficient level flag, an abs_remainder syntax element, and a sign flag for quantized transform coefficients for the current block; the significant coefficient flag relates to whether the quantized transform coefficient is a significant coefficient other than zero; the parity level flag relates to a parity of a transform coefficient level for the quantized transform coefficient; the first transform coefficient level flag relates to whether the transform coefficient level is greater than a first threshold; the second transform coefficient level flag relates to whether the transform coefficient level is greater than a second threshold; The abs_remainder syntax element relates to the remaining value of the transform coefficient level, the sign flag relates to the sign of the quantized transform coefficient; The step of encoding the residual information comprises: determining a threshold value associated with a sum of a number of significant coefficient flags, a number of first transform coefficient level flags, a number of parity level flags, and a number of second transform coefficient level flags for the transform coefficients in the current block based on a size of the current block, the threshold is a single threshold for all of the significant coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level flag, and is not determined as a sum of two or more thresholds; deriving a value of the significance flag, a value of the parity level flag, a value of the first transform coefficient level flag, a value of the second transform coefficient level flag, a value of the abs_remainder syntax element, and a value of the sign flag based on the quantized transform coefficients; encoding the significance coefficient flag, the first transform coefficient level flag, the parity level flag, the second transform coefficient level flag, the abs_remainder syntax element, and the sign flag; the sum of the number of significant coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags for the quantized transform coefficients in the current block is less than or equal to the threshold; The method, wherein the significant coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level flag are included in the residual information.

Citation Information

Patent Citations

  • Throughput improvements for CABAC coefficient level coding

    JP2015507424A