Transform Coefficient Coding Method and Apparatus

By introducing a transform skip mechanism in image encoding, optimizing residual coding, the problem of high resolution and high-quality image/video data transmission is solved, and more efficient image/video compression and decoding is achieved.

CN113170131BActive Publication Date: 2025-06-20VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980072715.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-10-11
Filing Date
2019-10-07
Publication Date
2025-06-20
Estimated Expiration
2039-10-07

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress and transmit high resolution and high quality image/video data, especially with challenges in terms of transmission costs and storage costs.

Method used

By introducing a transform skip mechanism in image encoding, the methods and equipment for residual encoding are optimized to improve image decoding efficiency. The specific steps include receiving the bit stream, deriving the quantized transformation coefficient, deriving the residual sample, and generating a reconstructed picture.

Benefits of technology

Improves image/video compression efficiency, reduces transmission and storage costs, and is suitable for high resolution and high-quality image/video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113170131B_ABST
    Figure CN113170131B_ABST
Patent Text Reader

Abstract

The image decoding method according to this document includes the following steps: receiving a bitstream including residual information; deriving quantization transform coefficients of a current block based on the residual information included in the bitstream; deriving residual samples of the current block based on the quantization transform coefficients; and generating a restored picture based on the residual samples of the current block, wherein the residual information includes position information for indicating the position of the last non-zero transform coefficient in a transform block, and the position information can be derived if transform skip is not applied to the transform block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to image coding techniques, and more particularly, to methods and apparatuses for coding transform coefficients. Background Art

[0002] Nowadays, the demand for high-resolution and high-quality images / videos such as 4K, 8K or higher ultra-high definition (UHD) images / videos has been continuously increasing in various fields. As image / video data becomes higher in resolution and quality, the amount of information or bits transmitted increases compared to traditional image data. Therefore, when transmitting image data using a medium such as a traditional wired / wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] In addition, nowadays, the interest and demand for immersive media such as virtual reality (VR), augmented reality (AR) content or holograms are increasing, and the broadcast of images / videos with image characteristics different from those of real images such as game images is increasing.

[0004] Therefore, there is a need for an efficient image / video compression technique that can effectively compress, transmit or store, and reproduce the information of high-resolution and high-quality images / videos with various characteristics as described above. Summary of the Invention

[0005] Technical Objectives

[0006] The present disclosure provides methods and apparatuses for improving image coding efficiency.

[0007] The present disclosure also provides methods and apparatuses for improving the efficiency of residual coding.

[0008] The present disclosure also provides methods and apparatuses for improving the efficiency of residual coding according to whether transform skip is applied.

[0009] Technical Solutions

[0010] According to an embodiment of the present disclosure, there is provided an image decoding method performed by a decoding device, the method including: receiving a bitstream including residual information; deriving quantized transform coefficients of a current block based on the residual information included in the bitstream; deriving residual samples of the current block based on the quantized transform coefficients; and generating a reconstructed picture based on the residual samples of the current block, wherein the residual information includes position information related to the position of the last non-zero transform coefficient in a transform block, and wherein the position information is derived when transform skip is not applied to the transform block.

[0011] According to another embodiment of the present disclosure, there is provided an image encoding method performed by an encoding device, the method including: deriving residual samples of a current block; deriving quantized transform coefficients based on the residual samples of the current block; encoding residual information including information on the quantized transform coefficients, where the residual information includes position information related to the position of the last non-zero transform coefficient in a transform block, and where the position information is derived when transform skip is not applied to the transform block.

[0012] According to yet another embodiment of the present disclosure, an image decoding device for performing an image decoding method, the image decoding device including: an entropy decoder that receives a bitstream including residual information and derives quantized transform coefficients of a current block based on information included in the bitstream; an inverse transformer that derives residual samples of the current block based on the quantized transform coefficients; an adder that generates a reconstructed picture based on the residual samples of the current block, where the residual information includes position information related to the position of the last non-zero transform coefficient in a transform block, and where the position information is derived when transform skip is not applied to the transform block.

[0013] According to yet another embodiment of the present disclosure, there is provided an encoding device for performing image encoding. The encoding device includes: a subtractor that derives residual samples of a current block; a quantizer that derives quantized transform coefficients based on the residual samples of the current block; and an entropy encoder that encodes residual information including information on the quantized transform coefficients, where the residual information includes position information related to the position of the last non-zero transform coefficient in a transform block, and where the position information is derived when transform skip is not applied to the transform block.

[0014] According to yet another embodiment of the present disclosure, a digital storage medium can be provided, in which image data including encoded image information generated according to an image encoding method performed by an encoding device is stored.

[0015] According to yet another embodiment of the present disclosure, a digital storage medium can be provided, in which image data including encoded image information that causes a decoding device to perform an image decoding method is stored.

[0016] Technical Effects

[0017] According to embodiments of the present disclosure, the overall image / video compression efficiency can be improved.

[0018] According to embodiments of the present disclosure, the efficiency of residual encoding can be improved.

[0019] According to the present disclosure, the efficiency of transform coefficient encoding can be improved.

[0020] According to the present disclosure, the efficiency of residual encoding can be improved according to whether transform skip is applied. Brief Description of the Drawings

[0021] Figure 1 Schematically shows an example where the video / image encoding system of the present disclosure can be applied.

[0022] Figure 2 Is a diagram schematically illustrating the configuration of a video / image encoding device to which the present disclosure can be applied.

[0023] Figure 3 Is a diagram schematically illustrating the configuration of a video / image decoding device to which the present disclosure can be applied.

[0024] Figure 4 Is a diagram illustrating a block diagram of a CABAC encoding system according to an embodiment.

[0025] Figure 5 Is a diagram illustrating an example of transform coefficients in a 4×4 block.

[0026] Figure 6 Is a diagram illustrating a residual signal decoder according to an example of the present disclosure.

[0027] Figure 7 Is a control flowchart illustrating a method for parsing coded_block_flag according to an embodiment of the present disclosure.

[0028] Figure 8 Shows the execution of Figure 7 The coded_block_flag derivator for the parsing in.

[0029] Figure 9 Is a control flowchart illustrating a method for parsing coded_block_flag according to another embodiment of the present disclosure.

[0030] Figure 10 Shows the execution of Figure 9 The coded_block_flag derivator for the parsing in.

[0031] Figure 11 Is a control flowchart illustrating a method for deriving the position of the last significant coefficient according to an embodiment of the present disclosure.

[0032] Figure 12 Shows the execution of Figure 11 The last coefficient position derivator for the derivation in.

[0033] Figure 13 Is a control flowchart illustrating a method for deriving the position of the last significant coefficient according to another embodiment of the present disclosure.

[0034] Figure 14 Shows the execution of Figure 13 The last coefficient position derivator for the derivation in.

[0035] Figure 15 is a control flowchart illustrating the operation of an encoding device according to an embodiment of the present disclosure.

[0036] Figure 16 is a diagram illustrating the configuration of an encoding device according to an embodiment of the present disclosure.

[0037] Figure 17 is a control flowchart illustrating the operation of a decoding device according to an embodiment of the present disclosure.

[0038] Figure 18 is a diagram illustrating the configuration of a decoding device according to an embodiment of the present disclosure.

[0039] Figure 19 Illustratively shows a content streaming system structure diagram to which the present disclosure can be applied. Detailed Embodiments

[0040] Although the present disclosure may be susceptible to various modifications and include various embodiments, specific embodiments thereof have been shown by way of example in the drawings and will now be described in detail. However, this is not intended to limit the present disclosure to the specific embodiments disclosed herein. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the technical concept of the present disclosure. The singular forms may include the plural forms unless the context clearly indicates otherwise. Terms such as "including" and "comprising" are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and thus should not be construed as precluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.

[0041] In addition, for the convenience of describing different characteristic functions, each component in the drawings described herein is illustratively shown independently. However, it is not meant that each component is implemented by separate hardware or software. For example, any two or more of these components may be combined to form a single component, and any single component may be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of the patent right of the present disclosure as long as they do not depart from the essence of the present disclosure.

[0042] Hereinafter, the preferred embodiments of the present disclosure will be described in more detail with reference to the drawings. In addition, in the drawings, the same reference numerals are used for the same components, and repeated descriptions of the same components will be omitted.

[0043] Figure 1 Schematically presents an example of a video / image encoding system to which the present disclosure can be applied.

[0044] Reference Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may deliver encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.

[0045] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0046] The video source may obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generating device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capturing process may be replaced by a process of generating relevant data.

[0047] The encoding device may encode the input video / image. The encoding device may perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.

[0048] The transmitter may send the encoded video / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating a media file in a predetermined file format and may include elements for sending via a broadcast / communication network. The receiver may receive / extract the bitstream and send the received / extracted bitstream to the decoding device.

[0049] The decoding device may decode the video / image by performing a series of processes such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding device.

[0050] The renderer may render the decoded video / image. The rendered video / image may be displayed through a display.

[0051] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the Versatile Video Coding (VVC), Efficient Video Coding (EVC) standard, AOMedia Video 1 (AV1) standard, Audio Video Coding Standard 2 (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.).

[0052] This document presents various embodiments of video / image coding, and unless otherwise mentioned, these embodiments can be executed in combination with each other.

[0053] In this document, video can refer to a series of images over time. A picture generally refers to a unit representing an image in a specific time region, and a slice / tile is a unit that forms part of a picture in coding. A slice / tile can include one or more Coding Tree Units (CTUs). A picture can be composed of one or more slices / tiles. A picture can be composed of one or more tile groups. A tile group can include one or more tiles. A brick can represent a rectangular region of CTU rows within a tile in a picture. A tile can be divided into multiple bricks, each of which consists of one or more CTU rows within the tile. A tile that is not divided into multiple bricks can also be referred to as a brick. Brick scan is a specific sequential ordering of the CTUs of a divided picture: the CTUs can be sorted in CTU raster scan within a brick, the bricks within a tile can be sequentially sorted in raster scan of the tile's bricks, and the tiles in a picture can be sequentially sorted in raster scan of the picture's tiles. A tile is a rectangular region of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular region of CTUs with a height equal to the height of the picture and a width specified by a syntax element in the picture parameter set. A tile row is a rectangular region of CTUs with a height specified by a syntax element in the picture parameter set and a width equal to the width of the picture. Brick scan is a specific sequential ordering of the CTUs of a divided picture: the CTUs are sequentially sorted in CTU raster scan within a tile, while the tiles in a picture are sequentially sorted in raster scan of the picture's tiles. A slice includes an integer number of bricks of a picture that can be exclusively contained in a single NAL unit. A slice can be composed of multiple complete tiles or only of a consecutive sequence of complete bricks of a single tile. In this document, tile groups can be used interchangeably with slices. For example, in this document, a tile group / tile group header can be referred to as a slice / slice header.

[0054] A pixel or pel may mean the smallest unit that makes up a picture (or image). Additionally, the term "sample" may be used as a term corresponding to a pixel. A sample generally may represent a pixel or the value of a pixel, and may represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.

[0055] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. A unit may include one luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term unit may be used interchangeably with terms such as block or region. Generally, an M×N block may include an array of M columns and N rows of samples (sample array) or a set (or array) of transform coefficients.

[0056] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". Additionally, "A, B" may mean "A and / or B". Additionally, "A / B / C" may mean at least one of A, B, and / or C. Additionally, "A / B / C" may mean at least one of A, B, and / or C.

[0057] Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".

[0058] Figure 2 is a diagram schematically depicting the configuration of a video / image encoding device to which the present disclosure may be applied. Hereinafter, a so-called video encoding device may include an image encoding device.

[0059] Refer to Figure 2, the encoding device 200 includes an image splitter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image splitter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be constructed by at least one hardware component (e.g., an encoder chipset or a processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.

[0060] The image splitter 210 splits an input image (or picture or frame) input to the encoding device 200 into one or more processing units. For example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split according to a quadtree binary tree ternary tree (QTBTTT) structure starting from a coding tree unit (CTU) or a largest coding unit (LCU). For example, one coding unit may be divided into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure may be applied first, and then the binary tree structure and / or the ternary tree structure may be applied. Alternatively, the binary tree structure may be applied first. The encoding process according to this document may be performed based on the final coding unit that is no longer split. In this case, the largest coding unit may be used as the final coding unit based on the encoding efficiency according to the image characteristics, or if necessary, the coding unit may be recursively split into coding units with a deeper depth and the coding unit with the optimal size may be used as the final coding unit. Here, the encoding process may include the processes of prediction, transformation, and reconstruction described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be divided or split from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal according to the transformation coefficients.

[0061] In some cases, the terms unit and terms such as block or region may be used interchangeably. Conventionally, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. Samples may generally represent pixels or pixel values, may represent only the pixel / pixel values of the luminance component, or may represent only the pixel / pixel values of the chrominance component. Samples may be used as a term corresponding to a picture (or image) for a pixel or pel.

[0062] In the encoding device 200, a prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 is subtracted from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) in the encoder 200 may be referred to as the subtractor 231. The predictor may perform prediction on a block to be processed (hereinafter referred to as the "current block") and may generate a prediction block including prediction samples for the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As described subsequently in the description of each prediction mode, the predictor may generate various prediction-related information such as prediction mode information and send the generated information to the entropy encoder 240. Information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0063] The intra-frame predictor 222 may predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference samples may be located near the current block or may be separated from the current block. In intra-frame prediction, the prediction mode may include a variety of non-directional modes and a variety of directional modes. The non-directional modes may include, for example, the DC mode and the planar mode. Depending on the level of detail of the prediction direction, the directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the setting. The intra-frame predictor 222 may determine the prediction mode to be applied to the current block by using the prediction mode applied to neighboring blocks.

[0064] The inter - frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. Here, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter - frame predictor 221 can configure a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter - frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, the residual signal cannot be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of a neighboring block can be used as a motion vector prediction term, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0065] The predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra - frame prediction or inter - frame prediction to the prediction of a block, but also apply intra - frame prediction and inter - frame prediction simultaneously. This can be referred to as combined inter - frame and intra - frame prediction (CIIP). Additionally, the predictor can be based on the intra - block copy (IBC) prediction mode or the palette mode to predict a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games (e.g., screen content coding (SCC)). Although IBC basically performs prediction in the current picture, the way it is performed is similar to inter - frame prediction in that a reference block is derived in the current picture. That is, IBC can use at least one of the inter - frame prediction techniques described in this document. The palette mode can be regarded as an example of intra - frame coding or intra - frame prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information about the palette table and the palette index.

[0066] The prediction signal generated by a predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) may be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT means a transform obtained from a graph when relationship information between pixels is represented by a graph. CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform process may be applied to square pixel blocks of the same size, or may be applied to blocks of variable size rather than square blocks.

[0067] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, and entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the quantized transform coefficients of the block type into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Information about the transform coefficients can be generated. Entropy encoder 240 can perform various coding methods such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. Entropy encoder 240 can encode the information required for video / image reconstruction together with or separately from the quantized transform coefficients (e.g., the values of syntax elements, etc.). The encoded information (e.g., the encoded video / image information) can be sent or stored in the form of a bitstream in units of NAL (network abstraction layer). The video / image information can also include information about various parameter sets such as adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), video parameter set (VPS). Additionally, the video / image information can also include general constraint information. In this document, the information and / or syntax elements sent from the encoding device to / signaled to the decoding device can be included in the video / picture information. The video / image information can be encoded through the above encoding process and included in the bitstream. The bitstream can be transmitted through a network or can be stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that sends the signal output from entropy encoder 240 or a storage unit (not shown) that stores the signal can be included as an internal / external element of encoding device 200. Alternatively, the transmitter can be included in entropy encoder 240.

[0068] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients using the inverse quantizer 234 and the inverse transformer 235, a residual signal (residual block or residual samples) can be reconstructed. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If there is no residual for the block to be processed, such as in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current block and, as described below, can be used for inter-prediction of the next picture by filtering.

[0069] In addition, during picture encoding and / or reconstruction, a luminance mapping with chroma scaling (LMCS) can be applied.

[0070] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 270, particularly in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. As discussed in the subsequent description of each filtering method, the filter 260 can generate various information related to the filtering and send the generated information to the entropy encoder 240. The information about the filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0071] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-prediction unit 221. When inter-prediction is applied by the encoding device, prediction mismatch between the encoding device 200 and the decoding device can be avoided, and the encoding efficiency can be improved.

[0072] The DPB of the memory 270 can store the modified reconstructed picture for use as a reference picture in the inter-prediction unit 221. The memory 270 can store the motion information of the blocks from which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks in the already reconstructed pictures. The stored motion information can be sent to the inter-prediction unit 221 and used as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and can transfer the reconstructed samples to the intra-prediction unit 222.

[0073] Figure 3It is a schematic diagram illustrating the configuration of a video / image decoding device to which the present document can be applied.

[0074] Referring to Figure 3 , the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra predictor 331 and an inter predictor 332. The residual processor 320 may include an inverse quantizer 321 and an inverse transformator 321. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be constructed of hardware components (e.g., a decoder chipset or a processor). Additionally, the memory 360 may include a decoded picture buffer (DPB), or may be constructed of a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0075] When receiving a bitstream including video / image information, the decoding device 300 may reconstruct an image corresponding to the processing of the video / image information in the Figure 2 encoding device accordingly. For example, the decoding device 300 may derive units / blocks based on information related to block segmentation obtained from the bitstream. The decoding device 300 may perform decoding using the processor applied in the encoding device. Thus, the decoding processor may be, for example, an encoding unit, and the encoding unit may be divided according to a quadtree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a largest coding unit. One or more transform units may be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproduction device.

[0076] The decoding device 300 may receive in the form of a bitstream from Figure 2The signal output by the encoding device, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). Additionally, the video / image information can also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. The signaled / received information and / or syntax elements described subsequently in this document can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 decodes the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, or CABAC, and outputs the values of the syntax elements required for image reconstruction and the quantization values of the transform coefficients for the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to the respective syntax elements in the bitstream, determine the context model using the decoding target syntax element information, the decoding information of the decoding target block, or the information of the symbols / bins decoded in the previous stage, perform arithmetic decoding on the bins by predicting the generation probability of the bins according to the determined context model, and generate the symbols corresponding to the values of each syntax element. In this case, the CABAC entropy decoding method can update the context model using the information of the symbols / bins decoded by the context model for the next symbol / bin after determining the context model. Among the information decoded in the entropy decoder 310, the information related to prediction can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantized transform coefficients) for which entropy decoding has been performed in the entropy decoder 310 and the associated parameter information can be input to the residual processor 320. The residual processor 320 can derive the residual signal (residual block, residual sample, residual sample array). Additionally, the information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) that receives the signal output by the encoding device can also be configured as an internal / external component of the decoding device 300, and the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to this document can be referred to as a video / image / picture encoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of an inverse quantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0077] The inverse quantizer 321 can inverse-quantize the quantized transform coefficients and output the transform coefficients. The inverse quantizer 321 can rearrange the quantized transform coefficients into the form of a two-dimensional block. In this case, the rearrangement can be performed based on the coefficient scan order executed in the encoding device. The inverse quantizer 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.

[0078] The inverse transformer 322 inverse-transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0079] The predictor can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and can determine a specific intra / inter prediction mode.

[0080] The predictor 320 can generate a prediction signal based on various prediction methods described below. For example, the predictor can apply not only intra prediction or inter prediction to the prediction of a block, but also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). Additionally, the predictor can perform prediction on the block based on the intra-block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode can be used for content image / video coding such as games (e.g., screen content coding (SCC)). Although IBC basically performs prediction in the current picture, the way it is performed is similar to inter prediction in that a reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information about the palette table and the palette index.

[0081] The intra predictor 331 can predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the reference samples can be located near the current block or can be separated from the current block. In intra prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra predictor 331 can determine the prediction mode to be applied to the current block by using the prediction mode applied to the neighboring blocks.

[0082] The inter-frame predictor 332 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, adjacent blocks may include spatial adjacent blocks present in the current picture and temporal adjacent blocks present in the reference picture. For example, the inter-frame predictor 332 may configure a motion information candidate list based on adjacent blocks, and derive a motion vector and / or a reference picture index of the current block based on the received candidate selection information. The inter-frame prediction may be performed based on various prediction modes, and information about the prediction may include information indicating the mode of the inter-frame prediction for the current block.

[0083] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (prediction block, prediction sample array) output from a predictor (inter-frame predictor 332 or intra-frame predictor 331). If there is no residual for the block to be processed, such as in the case of applying the skip mode, the prediction block may be used as the reconstructed block.

[0084] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, may be output through filtering as described below, or may be used for inter-frame prediction of the next picture.

[0085] In addition, a luminance mapping with chroma scaling (LMCS) may be applied in the picture decoding process.

[0086] The filter 350 may improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 may generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and may store the modified reconstructed picture in the memory 360, especially in the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0087] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter - predictor 332. The memory 360 can store the motion information of the blocks from which the motion information in the current picture is derived (or decoded) and / or the motion information of the blocks in the already - reconstructed pictures. The stored motion information can be sent to the inter - predictor 260 to be used as the motion information of spatially - adjacent blocks or temporally - adjacent blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and transfer the reconstructed samples to the intra - predictor 331.

[0088] In the present disclosure, the embodiments described in the filter 260, the inter - predictor 221, and the intra - predictor 222 of the encoding device 200 can be applied identically or correspondingly to the filter 350, the inter - predictor 332, and the intra - predictor 331 of the decoding device 300.

[0089] As described above, prediction is performed to improve the compression efficiency when performing video encoding. Accordingly, a prediction block including prediction samples for a current block as an encoding target block can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived identically in the encoding device and the decoding device, and the encoding device can improve the image encoding efficiency by signaling to the decoding device not the original sample values of the original block itself but the information about the residual between the original block and the prediction block (residual information). The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and generate a reconstructed picture including the reconstructed block.

[0090] The residual information can be generated through a transformation process and a quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transformation process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients so that it can signal the associated residual information to the decoding device (through a bitstream). Here, the residual information can include value information, position information, transformation technique, transformation kernel, quantization parameter, etc. of the quantized transform coefficients. The decoding device can perform a quantization / inverse - quantization process based on the residual information and derive residual samples (or a residual sample block). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can derive a residual block by performing inverse - quantization / inverse - transformation on the quantized transform coefficients for use as a reference for inter - prediction of the next picture, and can generate a reconstructed picture based on this.

[0091] Figure 4A block diagram of context - adaptive binary arithmetic coding (CABAC) for encoding a single syntax element is shown, as a figure that illustrates a block diagram of a CABAC encoding system according to an embodiment.

[0092] In the case where the input signal is a non - binarized syntax element, the encoding process of CABAC first transforms the input signal into a binarized value through binarization. In the case where the input signal is already a binarized value, the input signal bypasses binarization and does not undergo binarization, and is input to the encoding engine. Here, each binary number 0 or 1 that constitutes a binary value is called a bin. For example, in the case where the binarized binary string is "110", each of 1, 1, and 0 is called a bin. The bin of a syntax element can be the value of the syntax element.

[0093] The binarized bin is input to a regular encoding engine or a bypass encoding engine.

[0094] The regular encoding engine assigns a context model reflecting probability values to the corresponding bin and encodes the bin according to the assigned context model. After encoding each bin, the regular encoding engine can update the probability model of the bin. The bin encoded in this way is called a context - encoded bin.

[0095] The bypass encoding engine omits the process of estimating the probability of the input bin and the process of updating the probability model that has been applied to the bin after encoding. The bypass encoding engine encodes the bin input to it while applying a uniform probability distribution instead of assigning a context, thereby improving the encoding speed. The bin encoded in this way is called a bypass bin.

[0096] Entropy coding can determine whether to perform encoding through the regular encoding engine or through the bypass encoding engine, and switch the encoding path. Entropy decoding performs the same process as the encoding process in the reverse order.

[0097] In addition, in an embodiment, (quantized) transform coefficients are encoded and / or decoded based on syntax elements such as transform_skip_flag, last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix, coded_sub_block_flag, sig_coeff_flag, par_level_flag, rem_abs_gt1_flag, rem_abs_gt2_flag, abs_remainder, coeff_sign_flag, mts_idx, etc. Table 1 below shows syntax elements related to residual data encoding according to an example.

[0098] [Table 1]

[0099]

[0100]

[0101]

[0102] The transform_skip_flag indicates whether to skip the transform of an associated block. The associated block can be a coded block (CB) or a transform block (TB). Regarding transform (and quantization) and residual coding processing, CB and TB can be used interchangeably. For example, as described above, the residual samples of a CB can be derived, and (quantized) transform coefficients can be derived by transforming and quantizing the residual samples. Also, through the residual coding process, information (e.g., syntax elements) can be generated to efficiently signal the position, size, or sign, etc. of the (quantized) transform coefficients. Quantized transform coefficients can be abbreviated as transform coefficients. Generally, when the CB is not larger than the maximum TB, the size of the CB can be the same as the size of the TB, and in this case, the target block to be transformed (and quantized) and residual-coded can be referred to as a CB or a TB. In addition, when the CB is larger than the maximum TB, the target block to be transformed (and quantized) and residual-coded can be referred to as a TB. Although the syntax elements related to residual coding are described below by way of example as being signaled in units of transform blocks (TBs), as described above, TBs and coded blocks (CBs) can be used interchangeably.

[0103] In one embodiment, based on the syntax elements last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix, the (x, y) position information of the last non-zero transform coefficient in a transform block can be encoded. More specifically, last_sig_coeff_x_prefix indicates the prefix of the column position of the last valid coefficient in the scan order in the transform block, and last_sig_coeff_y_prefix indicates the prefix of the row position of the last valid coefficient in the scan order in the transform block; last_sig_coeff_x_suffix indicates the suffix of the column position of the last valid coefficient in the scan order in the transform block; and last_sig_coeff_y_suffix indicates the suffix of the row position of the last valid coefficient in the scan order in the transform block. Here, the valid coefficient can be a non-zero coefficient. The scan order can be a right-up diagonal scan order. Alternatively, the scan order can be a horizontal scan order or a vertical scan order. The scan order can be determined based on whether intra / inter prediction is applied to the target block (CB or CB including TB) and / or a specific intra / inter prediction mode.

[0104] Next, after dividing the transform block into 4×4 sub-blocks, one syntax element of coded_sub_block_flag can be used for each 4×4 sub-block to indicate whether there is a non-zero coefficient in the current sub-block.

[0105] If the value of coded_sub_block_flag is 0, there is no more information to send, so the encoding process for the current sub-block can be terminated. On the contrary, if the value of coded_sub_block_flag is 1, the encoding process for sig_coeff_flag can continue. Since the sub-block including the last non-zero coefficient does not need to encode coded_sub_block_flag, and the sub-block including the DC information of the transform block has a high probability of including non-zero coefficients, it can be assumed that coded_sub_block_flag has a value of 1 without being encoded.

[0106] If it is determined that there are non-zero coefficients in the current sub-block because the value of coded_sub_block_flag is 1, on the contrary, the sig_coeff_flag with a binary value can be encoded according to the scan order. The 1-bit syntax element sig_coeff_flag can be encoded for each coefficient according to the scan order. If the value of the transform coefficient at the current scan position is not 0, the value of sig_coeff_flag can be 1. Here, in the case of a sub-block including the last non-zero coefficient, since it is not necessary to encode sig_coeff_flag for the last non-zero coefficient, the encoding process for sig_coeff_flag can be omitted. Only when sig_coeff_flag is 1, horizontal information encoding can be performed, and four syntax elements can be used in the horizontal information encoding process. More specifically, each sig_coeff_flag[xC][yC] can indicate whether the horizontal (value) of the corresponding transform coefficient at each transform coefficient position (xC, yC) in the current TB is non-zero. In an embodiment, sig_coeff_flag can correspond to an example of a valid coefficient flag that indicates whether a quantized transform coefficient is a non-zero valid coefficient.

[0107] The remaining horizontal values after encoding sig_coeff_flag can be the same as those in Equation 1 below. That is, the syntax element remAbsLevel indicating the horizontal value to be encoded can be as shown in Equation 1 below. Here, coeff refers to the actual transform coefficient value.

[0108] [Equation 1]

[0109] remAbsLevel = |coeff| - 1

[0110] Through par_level_flag, the least significant bit (LSB) value of remAbsLevel written in Equation 1 can be encoded as shown in Equation 2 below. Here, par_level_flag[n] can indicate the parity of the transform coefficient level (value) at scan position n. After encoding par_leve_flag, the transform coefficient level value remAbsLevel to be encoded can be updated as shown in Equation 3 below.

[0111] [Equation 2]

[0112] par_level_flag = remAbsLevel & 1

[0113] [Equation 3]

[0114] remAbsLevel’ = remAbsLevel >> 1

[0115] The rem_abs_gt1_flag can indicate whether remAbsLevel' at the corresponding scan position n is greater than 1, and the rem_abs_gt2_flag can indicate whether remAbsLevel' at the corresponding scan position n is greater than 2. Encoding of abs_remainder can be performed only when rem_abs_gt2_flag is 1. When summarizing the relationship between the actual transform coefficient value coeff and each syntax element, for example, it can be Equation 4 below, and Table 2 below shows an example related to Equation 4. Additionally, a 1-bit sign coeff_sign_flag can be used to encode the sign of each coefficient. |coeff| can indicate the transform coefficient level (value) and can be represented as the AbsLevel of the transform coefficient.

[0116] [Equation 4]

[0117] |coeff| = sig_coeff_flag + par_level_flag + 2 * (rem_abs_gt1_flag + rem_abs_gt2_flag + abs_remainder)

[0118] In an embodiment, par_level_flag indicates an example of a parity level flag regarding the parity of the transform coefficient level of a quantized transform coefficient, rem_abs_gt1_flag indicates an example of a first transform coefficient level flag regarding whether the transform coefficient level is greater than a first threshold, and rem_abs_gt2_flag can indicate an example of a second transform coefficient level flag regarding whether the transform coefficient level is greater than a second threshold.

[0119] Additionally, in another embodiment, rem_abs_gt2_flag can be referred to as rem_abs_gt3_flag, and in another embodiment, rem_abs_gt1_flag and rem_abs_gt2_flag can be represented based on abs_level_gtx_flag[n][j]. abs_level_gtx_flag[n][j] can be a flag indicating whether the absolute value of the absolute value of the transform coefficient level at scan position n (or the transform coefficient level shifted right by 1) is greater than (j << 1) + 1. In one example, rem_abs_gt1_flag can perform the same and / or similar functions as abs_level_gtx_flag[n][0], and rem_abs_gt2_flag can perform the same and / or similar functions as abs_level_gtx_flag[n][1]. That is, abs_level_gtx_flag[n][0] can correspond to an example of the first transform coefficient level flag, while abs_level_gtx_flag[n][1] can correspond to an example of the second transform coefficient level flag. (j << 1) + 1 can be replaced with a predetermined threshold (e.g., a first threshold and a second threshold) according to the situation.

[0120] [Table 2]

[0121] |coeff| sig_coeff_flag par_level_flag rem_abs_gt1_flag rem_abs_gt2_flag abs_remainder 0 0 1 1 0 0 2 1 1 0 3 1 0 1 0 4 1 1 1 0 5 1 0 1 1 0 6 1 1 1 1 0 7 1 0 1 1 1 8 1 1 1 1 1 9 1 0 1 1 2 10 1 1 1 1 2 11 1 0 1 1 3 ... ... ... ... ... ...

[0122] Figure 5 is a diagram illustrating examples of transform coefficients in a 4×4 block.

[0123] Figure 5 The 4×4 block of shows examples of quantization coefficients. Figure 5 The box shown in can be a 4×4 sub-block of a 4×4 transform block, or an 8×8, 16×16, 32×32, or 64×64 transform block. Figure 5 The 4×4 block of can be a luminance block or a chrominance block. Figure 5 The coding results of the anti-diagonal scan coefficients of can be shown, for example, in Table 3. In Table 3, scan_pos indicates the position of the coefficient according to the anti-diagonal scan. scan_pos 15 is the coefficient that is first scanned in the 4×4 block, i.e., the coefficient in the lower right corner, while scan_pos 0 is the coefficient that is scanned last, i.e., the coefficient in the upper left corner. Additionally, in one embodiment, scan_pos can be referred to as the scan position. For example, scan_pos 0 can be referred to as scan position 0.

[0124] [Table 3]

[0125] scan_pos 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 0 coefficients 0 0 0 0 1 -1 0 2 0 3 -2 -3 4 6 -7 10 sig_coeff_flag 0 0 0 0 1 1 0 1 0 1 1 1 1 1 1 1 par_level_flag 0 0 1 0 1 0 1 1 0 1 rem_abs_gt1_flag 0 0 0 1 0 1 1 1 1 1 rem_abs_gt2_flag 0 0 0 1 1 1 abs_remainder 0 1 2 ceoff_sign_flag 0 1 0 0 1 1 0 0 1 0

[0126] In addition, as described with reference to Table 1, before encoding the residual signal and the special residual signal, it is first sent whether to apply the transform to the corresponding block. By representing the correlation between the residual signals in the transform domain, data compression is achieved and sent to the decoding device. If the correlation between the residual signals is insufficient, data compression may not be fully performed. In this case, the transform process including complex computational processing can be omitted, and the residual signal in the pixel domain (spatial domain) can be sent to the decoding device.

[0127] Since the residual signal in the pixel domain that has not undergone transformation has characteristics different from those of the general residual signal in the transform domain (distribution of the residual signal, absolute level of each residual signal, etc.), a residual signal encoding method for efficiently sending such a signal to the decoding device according to an example of the present disclosure will be proposed below.

[0128] Figure 6 FIG. is an illustration of a residual signal decoder according to an example of the present disclosure.

[0129] As shown in the figure, a transform application flag indicating whether to apply a transform to the corresponding transform block and information about the binarized code for encoding can be input to the residual signal decoder 600, and a decoded residual signal can be output from the residual signal decoder 600.

[0130] The flag regarding whether to apply a transform can be represented as transform_skip_flag, and the binarized code for encoding can be input to the residual signal decoder 600 through Figure 6 binarization processing.

[0131] The transform skip flag is sent in units of transform blocks, and in Table 1, the flag regarding whether to perform a transform is limited to a specific block size (only when the size of the transform block is 4×4 or smaller, the condition for parsing the transform skip flag is included). However, in the present embodiment, the size of the block for determining whether to parse the transform skip flag can be configured differently. The sizes of Log2TbWidth and log2TbHeight are determined as variables wN and hN, and wN and hN can be selected as one of the following.

[0132] [Equation 5]

[0133] wN = {2, 3, 4, 5}

[0134] wH = {2, 3, 4, 5}

[0135] The syntax elements to which Equation 5 can be applied are as follows.

[0136] [Table 4]

[0137]

[0138] As described above, a method for decoding a residual signal can be determined according to a transform skip flag. By the proposed method, the complexity in entropy decoding processing can be reduced and the coding efficiency can be improved by efficiently processing signals having different statistical characteristics from each other.

[0139] As described in Table 1 and the above embodiments, before encoding a residual signal (residual or transform coefficient), first, whether to apply the transform of the corresponding block is sent to the decoding device. By representing the correlation between residual signals in the transform domain, the data is compressed and sent to the decoding device. If the correlation between residual signals is insufficient, data compression may not be fully performed. In this case, the transform processing including complex calculation processing can be omitted, and the residual signal in the pixel domain (spatial domain) can be sent to the decoding device. The residual signal in the pixel domain as the skipped transform has different characteristics (distribution of the residual signal, absolute level of each residual signal, etc.) from the residual signal in the general transform domain.

[0140] Therefore, in the present embodiment, a residual signal encoding method for efficiently sending such a signal to the decoding device is proposed. More specifically, we propose a method for parsing coded_block_flag according to transform_skip_flag.

[0141] Generally, as shown in Table 1, for efficient transmission of residual coefficients, first, the presence or absence of the residual signal of the lower block is sent, and this information is sent as a syntax element called coded_block_flag. If coded_block_flag is 1, it means that there is a residual signal in the corresponding sub-block, and the signal is reconstructed through the residual coefficient decoding process. If coded_block_flag is 0, it means that there is no residual coefficient in the corresponding sub-block, so the parsing of the residual signal of this sub-block is no longer performed, but the next sub-block is parsed.

[0142] In addition, in the case of the residual coefficient in the pixel domain, different from the residual in the transform domain, since the correlation between data is very low and has randomness, sending coded_block_flag may send unnecessary redundant information. Therefore, in the present embodiment, a method for parsing coded_block_flag only when transform_skip_flag is 0 is proposed.

[0143] Examples of syntax elements for describing the proposed method are shown in Table 5 below. In Table 5, “&&transform_skip_flag[x0][y0][cIdx]” is specified as a condition for parsing the coded_block_flag.

[0144] [Table 5]

[0145]

[0146]

[0147]

[0148]

[0149] Figure 7 is a control flow diagram illustrating a method for parsing coded_block_flag according to an embodiment of the present disclosure, and Figure 8 shows the coded_block_flag derivator that performs the Figure 7 parsing therein.

[0150] As shown, a transform application skip flag (e.g., transform_skip_flag) indicating whether a transform has been applied and a bitstream including the flag are input to the coded_block_flag derivator 800, and the coded_block_flag can be output based on these.

[0151] The coded_block_flag derivator 800 can determine whether to apply transform skip to a transform block (S700) based on whether the transform_skip_flag is 0, the parsed transform_skip_flag value, or the transform_skip_flag parsing.

[0152] If the transform_skip_flag is 0, since transform skip is not applied and the transform is applied to the transform block, the coded_block_flag derivator 800 parses the coded_block_flag as described above (S710).

[0153] On the other hand, if the transform_skip_flag is 1 instead of 0, since transform skip is applied and the transform block is not transformed, the coded_block_flag is not parsed as described above.

[0154] According to an embodiment, as described in Table 1 and the above embodiments, before encoding the residual signal (residual or transform coefficient), first, whether to apply the transform of the corresponding block is sent to the decoding device. By representing the correlation between the residual signals in the transform domain, the data is compressed and sent to the decoding device. If the correlation between the residual signals is insufficient, data compression may not be performed sufficiently. In this case, the transform process including complex computational processing can be omitted, and the residual signal in the pixel domain (spatial domain) can be sent to the decoding device.

[0155] The residual signal in the pixel domain that skips the transform has different characteristics (distribution of the residual signal, absolute level of each residual signal, etc.) from the residual signal in the general transform domain. Therefore, in the present embodiment, a residual signal encoding method for efficiently sending such a signal to the decoding device is proposed. More specifically, we propose a method for parsing coded_block_flag according to transform_skip_flag.

[0156] Generally, as shown in Table 1, for efficient transmission of the residual coefficients, first, the presence or absence of the residual signal of the lower block is sent, and this information is sent as a syntax element called coded_block_flag. If coded_block_flag is 1, it means that there is a residual signal in the corresponding sub-block, and the signal is reconstructed through the residual coefficient decoding process. If coded_block_flag is 0, it means that there is no residual coefficient in the corresponding sub-block, so the parsing of the residual signal for this sub-block is no longer performed, but the next sub-block is parsed.

[0157] In addition, in the case of the residual coefficients in the pixel domain, different from the residuals in the transform domain, since the correlation between the data is very low and has randomness, sending coded_block_flag may send unnecessary redundant information. In particular, in the case of an intra-prediction block, as the distance from the left and upper sides of the block increases, the size of the residual increases. If the transform is not performed, these characteristics will be reflected as they are, so in a large block, information like the transformed residual is not implicit in the upper left corner of the block, while the residual still exists in the lower right corner. Therefore, sending coded_block_flag for all lower blocks sends redundant information.

[0158] Therefore, in the present embodiment, a method for parsing coded_block_flag is proposed only when the size of the target block (transform block or sub-block) for residual coding is greater than a certain size and transform_skip_flag is 0. In one example, the maximum width and maximum height of the target block (i.e., the width threshold and height threshold of the transform block) can be represented by log2ThWSize and log2ThHSize respectively, and these values are each one of 2, 3, 4, 5, 6, 7, and 8. However, in this document, the values of the size of the transform block are not limited to the above specific values.

[0159] An example of the syntax elements for describing the proposed method is shown in Table 6 below. In Table 6, it is specified that "(log2TbWidth <= wN) && (log2TbHeight <= hN)", and the parsing condition for coded_sub_block_flag[xS][yS] is specified as "&&transform_skip_flag[x0][y0][cIdx] && log2TbWidth < log2ThWSize && log2TbHeight < log2ThSizeH".

[0160] [Table 6]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166] Figure 9 is a control flow diagram illustrating a method for parsing coded_block_flag according to another embodiment of the present disclosure, and Figure 10 shows the execution of Figure 9 the coded_block_flag deriver that performs the parsing in

[0167] As shown in the figure, a transform application skip flag (such as transform_skip_flag) indicating whether a transform has been applied, the size of the transform block, and the bitstream including them are input into the coded_block_flag deriver 1000, and coded_block_flag can be output based on these.

[0168] The coded_block_flag derivator 1000 can determine whether to apply transform skip to a transform block (S900) based on whether transform_skip_flag is 0, the parsed transform_skip_flag value, or transform_skip_flag parsing.

[0169] If transform_skip_flag is 0, since transform skip is not applied and a transform is applied to the transform block, the coded_block_flag derivator 1000 can determine whether the width of the transform block (log2TbWidth) is less than a predetermined threshold (log2ThWSize) (S910).

[0170] If the width of the transform block (log2TbWidth) is less than the predetermined threshold (log2ThWSize), it can be determined whether the height of the transform block (log2TbHeight) is less than the predetermined threshold (log2ThHSize) (S920).

[0171] As described above, only when the width of the transform block (log2TbWidth) is less than the predetermined threshold (log2ThWSize) and the height of the transform block (log2TbHeight) is also less than the predetermined threshold (log2ThHSize), the coded_block_flag is parsed as described above (S930).

[0172] On the other hand, if transform_skip_flag is 1 instead of 0, since transform skip is applied and no transform is performed on the transform block, and the transform block width (log2TbWidth) is greater than the predetermined threshold (log2ThWSize) or the transform block height (log2TbHeight) is greater than the predetermined threshold (log2ThHSize), the coded_block_flag derivator 1000 does not parse the coded_block_flag, as described above.

[0173] According to an embodiment, as described in Table 1 and the above embodiments, whether to apply the transform of the corresponding block before encoding the residual signal (residual or transform coefficient) is first sent to the decoding device. By representing the correlation between the residual signals in the transform domain, the data is compressed and sent to the decoding device. If the correlation between the residual signals is insufficient, data compression may not be fully achieved. In this case, the transform process including complex calculation processing can be omitted, and the residual signal in the pixel domain (spatial domain) can be sent to the decoding device.

[0174] The residual signal in the pixel domain as a skip transformation has different characteristics from those in the general transform domain (distribution of the residual signal, absolute level of each residual signal, etc.). Therefore, in the present embodiment, a residual signal encoding method for efficiently transmitting such a signal to a decoding device is proposed. More specifically, we propose a method of parsing context elements indicating the positions of the last coefficients, such as last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix, according to the transform_skip_flag.

[0175] Generally, as shown in Table 1, for efficient transmission of residual coefficients, first, the presence or absence of the residual signal of the lower block is transmitted. To parse this information, information about the position where the non-zero residual coefficient first appears is transmitted to the decoding device starting from the lower right of the transform block. Such position information is encoded in the context elements of last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix and is transmitted to the decoding device. The residual coefficients are decoded based on the last coefficient position derived from these context elements according to the scan order, and the residual coefficients from the lower right position of the transform block to the last coefficient position can be inferred as 0.

[0176] In addition, in the case of the residual coefficients in the pixel domain, different from the residuals in the transform domain, since the correlation between data is very low and has randomness, transmitting the coded_block_flag may transmit unnecessary redundant information. In particular, in the case of an intra prediction block, as the distance from the left and upper sides of the block increases, the magnitude of the residual increases. If no transformation is performed, these characteristics will be reflected as they are. Therefore, in a large block, information such as the transformed residual is not implicit in the upper left corner of the block, and the residual still exists in the lower right corner. Thus, transmitting the last transform coefficient position information for all lower blocks will transmit redundant information.

[0177] Therefore, in the present embodiment, a method of parsing context elements indicating the positions of the last coefficients, such as last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix, is proposed only when the transform_skip_flag of the transform block is 0.

[0178] Figure 11 is a control flowchart illustrating a method of deriving a last valid coefficient position according to an embodiment of the present disclosure, and Figure 12 shows the execution of Figure 11 the last coefficient position deriver in

[0179] As shown, a transform application skip flag (e.g., transform_skip_flag) indicating whether a transform has been applied, and a bitstream including the flag are input to the last coefficient position deriver 1200, and the last coefficient position information can be output based on these.

[0180] The last coefficient position deriver 1200 can determine whether to apply transform skip to a transform block (S1100) based on whether the transform_skip_flag is 0, the parsed transform_skip_flag value, or the transform_skip_flag parsing.

[0181] If the transform_skip_flag is 0, since transform skip is not applied and the transform is applied to the transform block, as described above, the last coefficient position deriver 1200 parses the last coefficient position information (S1110).

[0182] On the other hand, if the transform_skip_flag is 1 instead of 0, since transform skip is applied and the transform block is not transformed, the last coefficient position information is not parsed as described above.

[0183] According to an embodiment, as described in Table 1 and the above embodiments, whether to apply the transform of the corresponding block before encoding the residual signal (residual or transform coefficient) is first sent to the decoding device. By representing the correlation between the residual signals in the transform domain, the data is compressed and sent to the decoding device. If the correlation between the residual signals is insufficient, data compression may not be performed sufficiently. In this case, the transform process including complex calculation processing can be omitted, and the residual signal in the pixel domain (spatial domain) can be sent to the decoding device.

[0184] The residual signal in the pixel domain as a skip transform has characteristics different from those of the residual signal in the general transform domain (distribution of the residual signal, absolute level of each residual signal, etc.). Therefore, in the present embodiment, a residual signal encoding method for efficiently transmitting such a signal to a decoding device is proposed. We propose a method of parsing context elements indicating the positions of the last coefficients, such as last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix, according to the transform_skip_flag and the size of the transform block.

[0185] Generally, as shown in Table 1, for efficient transmission of the residual coefficients, first, the presence or absence of the residual signal of the lower block is transmitted. To parse this information, the information about the position where the non-zero residual coefficients first appear is transmitted to the decoding device starting from the lower right of the transform block. Such position information is encoded in the context elements of last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix and is transmitted to the decoding device. The residual coefficients are decoded based on the last coefficient position derived from these context elements according to the scan order, and the residual coefficients from the lower right position of the transform block to the last coefficient position can be inferred as 0.

[0186] In addition, in the case of the residual coefficients in the pixel domain, different from the residuals in the transform domain, since the correlation between data is very low and has randomness, transmitting the coded_block_flag may transmit unnecessary redundant information. In particular, in the case of an intra-prediction block, as the distance from the left and upper sides of the block increases, the size of the residual increases. If no transform is performed, these characteristics will be reflected as they are. Therefore, in a large block, information such as the transformed residual is not implicit in the upper left corner of the block, and the residual still exists in the lower right corner. Thus, transmitting the last transform coefficient position information for all lower blocks transmits redundant information.

[0187] Therefore, in the present embodiment, a method is proposed to parse context elements indicating the positions of the last coefficients, such as last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix, only when the size of the target block (transform block or sub-block) for residual coding is equal to or greater than a certain size and transform_skip_flag is 0. In one example, the predetermined maximum width and maximum height of the target block (i.e., the width threshold and height threshold of the transform block) can be represented as log2ThWSize and log2ThHSize, and these values are each one of 2, 3, 4, 5, 6, 7, and 8. However, in this document, the values of the size of the transform block are not limited to the above specific values.

[0188] Figure 13 is a control flowchart illustrating a method for deriving the position of the last significant coefficient according to another embodiment of the present disclosure, and Figure 14 shows the execution of Figure 13 the last coefficient position derivator in

[0189] As shown in the figure, a transform application skip flag (e.g., transform_skip_flag) indicating whether a transform has been applied, the size of the transform block, and the bitstream including them are input to the last coefficient position derivator 1400, and the last coefficient position information can be output based on this information.

[0190] The last coefficient position derivator 1400 can determine whether to apply transform skip to the transform block based on whether transform_skip_flag is 0, the parsed transform_skip_flag value, or the transform_skip_flag parsing (S1300).

[0191] If transform_skip_flag is 0, since transform skip is not applied and the transform is applied to the transform block, the last coefficient position derivator 1400 can determine whether the width (log2TbWidth) of the transform block is less than a predetermined threshold (log2ThWSize) (S1310).

[0192] If the width (log2TbWidth) of the transform block is less than the predetermined threshold (log2ThWSize), it can be determined whether the height (log2TbHeight) of the transform block is less than the predetermined threshold (log2ThHSize) (S1320).

[0193] As described above, the last coefficient position information (S1330) is parsed as described above only when the width of the transform block (log2TbWidth) is less than a predetermined threshold (log2ThWSize) and the height of the transform block (log2TbHeight) is also less than a predetermined threshold (log2ThHSize).

[0194] On the other hand, if transform_skip_flag is 1 instead of 0, since transform skip is applied, the transform block is not transformed. If the transform block width (log2TbWidth) is greater than a predetermined threshold (log2ThWSize) or the transform block height (log2TbHeight) is greater than a predetermined threshold (log2ThHSize), then as described above, the last coefficient position deriver 1400 does not parse the last coefficient position information.

[0195] In addition, as described above, the syntax elements rem_abs_gt1_flag and rem_abs_gt2_flag can be represented based on abs_level_gtx_flag[n][j], and can also be represented as abs_rem_gt1_flag and abs_rem_gt2_flag or abs_rem_gtx_flag.

[0196] As described above, according to an embodiment of the present disclosure, different residual coding schemes (i.e., residual syntax) can be applied depending on whether transform skip is applied to residual coding.

[0197] For example, information about the position of the last significant coefficient (last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix) is encoded only when transform skip is applied to the transform block, is sent to the decoding device, and is parsed.

[0198] For example, depending on whether transform skip is applied, the signaling order of the flag (coeff_sign_flag) regarding the sign of the transform coefficient can be different. When transform skip is not applied, coeff_sign_flag is signaled after abs_remainder, while when transform skip is applied, coeff_sign_flag can be signaled before rem_abs_gt1_flag.

[0199] In addition, for example, the parsing of rem_abs_gt1_flag, rem_abs_gt2_flag (i.e., rem_abs_gtx_flag) and the parsing loop of abs_remainder can vary depending on whether transform skipping is applied.

[0200] In addition, context syntax elements encoded by context-based arithmetic coding may include: a significant coefficient flag (sig_coeff_flag) indicating whether a quantized transform coefficient is a non-zero significant coefficient, a parity level flag (par_level_flag) regarding the parity of the transform coefficient level of the quantized transform coefficient, a first transform coefficient level flag (rem_abs_gt1_flag) regarding whether the transform coefficient level is greater than a first threshold, and a second transform coefficient level flag (rem_abs_gt2_flag) regarding whether the transform coefficient level of the quantized transform coefficient is greater than a second threshold. In this case, the decoding of the first transform coefficient level flag may be performed before the decoding of the parity level flag.

[0201] Tables 7 to 9 show context elements according to the above examples.

[0202] [Table 7]

[0203]

[0204] [Table 8]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210] [Table 9]

[0211]

[0212]

[0213] Table 7 shows the branching of residual coding according to the value of transform_skip_flag, that is, different syntax elements are used for the residuals. Additionally, Table 8 shows the residual coding when the value of transform_skip_flag is 0 (i.e., when the transform is applied), and Table 9 shows the residual coding when the value of transform_skip_flag is 1 (i.e., when the transform is not applied).

[0214] In Tables 8 and 9, par_level_flag can be expressed as Equation 6 below.

[0215] [Equation 6]

[0216] par_level_flag = coeff & 1

[0217] Additionally, in Tables 8 and 9, since par_level_flag is parsed, that is, decoded after abs_level_gtx_flag, rem_abs_gt1_flag can indicate whether the transform coefficient at the corresponding scan position n is greater than 1, and rem_abs_gt2_flag can indicate whether the transform coefficient at the corresponding scan position n is greater than 3. That is, rem_abs_gt2_flag in Table 1 can be expressed as rem_abs_gt3_flag in Tables 8 and 9.

[0218] When changing Equations 2 to 3 as described above, in the cases of Tables 8 and 9 below, Equation 4 can be changed as follows.

[0219] [Equation 7]

[0220] |coeff| = sig_coeff_flag + par_level_flag + rem_abs_gt1_flag + 2 * (rem_abs_gt2_flag + abs_remainder

[0221] Figure 15 is a control flowchart illustrating the operation of an encoding device according to an embodiment of the present disclosure, and Figure 16 is a diagram showing the configuration of an encoding device according to an embodiment of the present disclosure.

[0222] According to Figure 15 and Figure 16 The encoding device can perform operations corresponding to the decoding device according to Figure 17 and Figure 18 Therefore, it will be described later in Figure 17 and Figure 18The operations of the decoding device described in can be similarly applied to the encoding device according to Figure 15 and Figure 16 .

[0223] In Figure 15 , each step disclosed can be performed by the encoding device 200 disclosed in Figure 2 . More specifically, S1500 can be performed by the subtractor 231 disclosed in Figure 2 , S1510 can be performed by the quantizer 233 disclosed in Figure 2 , and S1520 can be performed by the entropy encoder 240. Additionally, the operations according to S1500 to S1520 are based on some of the content described above in Figures 4 to 14 . Therefore, specific descriptions overlapping with those described above in Figure 2 and Figures 4 to 14 will be omitted or simplified.

[0224] As Figure 15 shows, the encoding device according to an embodiment may include a subtractor 231, a transformer 232, a quantizer 233, and an entropy encoder 240. However, in some cases, not all components as Figure 15 shown are necessary components of the encoding device, and the encoding device may be implemented with more or fewer components than Figure 15 shown.

[0225] In the encoding device according to an embodiment, the subtractor 231, the transformer 232, the quantizer 233, and the entropy encoder 240 are each implemented as separate chips, or at least two components may also be implemented through a single chip.

[0226] The encoding device according to an embodiment may derive the residual samples of the current block (S1500). More specifically, the subtractor 231 of the encoding device may derive the residual samples of the current block.

[0227] The encoding device according to an embodiment may derive the quantized transform coefficients based on the residual samples of the current block (S1510). More specifically, the quantizer 233 of the encoding device may derive the quantized transform coefficients based on the residual samples of the current block.

[0228] The encoding device according to an embodiment may encode the residual information including information about the quantized transform coefficients (S1520). More specifically, the entropy encoder 240 of the encoding device may encode the residual information including information about the quantized transform coefficients.

[0229] In an embodiment, the residual information may include information indicating the positions of the last significant coefficients, such as last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix. More specifically, last_sig_coeff_x_prefix represents the prefix of the column position of the last significant coefficient in the scan order in the transform block, while last_sig_coeff_y_prefix represents the prefix of the row position of the last significant coefficient in the scan order in the transform block, and last_sig_coeff_x_suffix represents the suffix of the column position of the last significant coefficient in the scan order in the transform block, while last_sig_coeff_y_suffix represents the suffix of the row position of the last significant coefficient in the scan order in the transform block. Here, the significant coefficient may represent a non-zero coefficient. The scan order may be a right-upper diagonal scan order. Alternatively, the scan order may be a horizontal scan order or a vertical scan order. According to an example, the information indicating the positions of the last significant coefficients may be encoded and decoded only when the transform skip is not applied to the transform block.

[0230] According to an example, when the transform skip is not applied to the transform block, the information indicating the positions of the last significant coefficients may be encoded and decoded based on the size of the transform block. For example, the information indicating the positions of the last significant coefficients may be encoded and decoded only when the width and height of the transform block are less than a predetermined threshold.

[0231] According to another example, the residual information may include coded_sub_block_flag indicating whether a sub-block includes the last non-zero coefficient, and the coded_sub_block_flag is encoded and decoded only when the transform skip is not applied to the transform block.

[0232] According to an example, when the transform skip is not applied to the transform block, the coded_sub_block_flag may be encoded and decoded based on the size of the transform block. For example, the coded_sub_block_flag may be encoded and decoded only when the width and height of the transform block are less than a predetermined threshold.

[0233] In one embodiment, the residual information includes a parity level flag regarding the parity of the transform coefficient level of the quantized transform coefficient and a first transform coefficient level flag regarding whether the transform coefficient level is greater than a first threshold. In one example, the parity level flag indicates par_level_flag, the first transform coefficient level flag indicates rem_abs_gt1_flag or abs_level_gtx_flag[n][0], and the second transform coefficient level flag indicates rem_abs_gt2_flag or abs_level_gtx_flag[n][1].

[0234] In an embodiment, the encoding of the residual information may include: deriving the value of the parity level flag and the value of the first transform coefficient level flag based on the quantized transform coefficient; and encoding the first transform coefficient level and encoding the parity level flag.

[0235] In an embodiment, the encoding of the first transform coefficient level flag may be performed before the encoding of the parity level flag. For example, the encoding device may perform encoding on rem_abs_gt1_flag or abs_level_gtx_flag[n][0] before encoding par_level_flag.

[0236] In an embodiment, the residual information may include: a valid coefficient flag indicating whether the quantized transform coefficient is a non-zero valid coefficient, and a second transform coefficient level flag indicating whether the transform coefficient level of the quantized transform coefficient is greater than a second threshold. In one example, the valid coefficient flag may be sig_coeff_flag.

[0237] The residual information may include context syntax elements encoded based on context, and the context syntax elements may include a valid coefficient flag, a parity level flag, a first transform coefficient level flag, and a second transform coefficient level flag.

[0238] In an embodiment, the step of deriving the quantized transform coefficient may encode the context syntax elements based on context and based on a predetermined maximum value of the context syntax elements.

[0239] In other words, the sum of the number of valid coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags of the quantized transform coefficients in the current block included in the residual information may be less than or equal to the predetermined maximum value.

[0240] The maximum value is the sum of the number of valid coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags of the quantized transform coefficients related to the current sub-block in the current block.

[0241] In an embodiment, the maximum value may be determined for each transform block. The current block may be a sub-block within a transform block that is a transform unit, and encoding of quantized transform coefficients may be performed for each sub-block. When encoding quantized transform coefficients for each sub-block, context syntax elements among residual information may be encoded based on the maximum value determined for each transform block.

[0242] In an embodiment, the threshold may be determined based on the size of the current block (or the current sub-block within the current block). If the current block is a transform block, the threshold may be determined based on the size of the transform block.

[0243] In an embodiment, when the sum of the number of valid coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags derived based on the 0th quantized transform coefficient to the nth quantized transform coefficient determined by the coefficient scan order reaches a predetermined maximum value, for the (n + 1)th quantized transform coefficient determined by the coefficient scan order, explicit signaling of the valid coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level flag may be omitted, and the value of the (n + 1)th quantized transform coefficient may be derived based on the value of the coefficient level information included in the residual information.

[0244] For example, when the sum of the number of sig_coeff_flag, the number of rem_abs_gt1_flag (or abs_level_gtx_flag[n][0]), the number of par_level_flags, and the number of rem_abs_gt2_flag (or abs_level_gtx_flag[n][1]) derived based on the 0th quantized transform coefficient (or the first quantized transform coefficient) to the nth quantized transform coefficient (or the nth quantized transform coefficient) determined by the coefficient scan order reaches a predetermined maximum value, for the (n + 1)th quantized transform coefficient determined by the coefficient scan order, explicit signaling of sig_coeff_flag, rem_abs_gt1_flag (or abs_level_gtx_flag[n][0]), par_level_flag, abs_level_gtx_flag[n][1], and rem_abs_gt2_flag (or abs_level_gtx_flag[n][1]) may be omitted, and the value of the (n + 1)th quantized transform coefficient may be derived based on the value of abs_remainder or dec_abs_level included in the residual information.

[0245] In an embodiment, the valid coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level flag included in the residual information may be context - encoded, and the coefficient level information may be bypass - encoded.

[0246] Figure 17 is a control flowchart illustrating the operation of a decoding device according to an embodiment of the present disclosure, and Figure 18 is a diagram showing the configuration of a decoding device according to an embodiment of the present disclosure.

[0247] In Figure 17 each step disclosed in Figure 3 may be performed by the decoding device 300 disclosed in Figure 3 More specifically, S1700 and S1710 may be performed by the entropy decoder 310 disclosed in Figure 3 and S1720 may be performed by the de - quantization unit 321 and / or the inverse transformer 322 disclosed in Figure 3 In addition, S1730 may be performed by the adder 340 disclosed in Figures 4 to 14 In addition, the operations according to S1700 to S1730 are based on some of the content described above in Figures 3 to 14 Therefore, detailed descriptions overlapping with those described above in

[0248] As Figure 18 shown, a decoding device according to an embodiment may include an entropy decoder 310, a de - quantization unit 321, an inverse transformer 322, and an adder 340. However, in some cases, not all components shown in Figure 18 may be essential components of the decoding device, and the decoding device may be implemented with more or fewer components than those shown in Figure 18 shown.

[0249] In a decoding device according to an embodiment, the entropy decoder 310, the de - quantization unit 321, the inverse transformer 322, and the adder 340 are each implemented as a separate chip, or at least two components may also be implemented through a single chip.

[0250] A decoding device according to an embodiment may receive a bitstream including residual information (S1700). More specifically, the entropy decoder 310 of the decoding device may receive a bitstream including residual information.

[0251] A decoding device according to an embodiment may derive the quantized transform coefficients of the current block based on the residual information included in the bitstream (S1710). More specifically, the entropy decoder 310 of the decoding device may derive the quantized transform coefficients of the current block based on the residual information included in the bitstream.

[0252] The decoding device according to an embodiment may derive the residual samples of the current block based on the quantized transform coefficients (S1720). More specifically, the dequantizer 321 of the decoding device may derive the transform coefficients from the quantized transform coefficients based on a dequantization process, and the inverse transformer 322 of the decoding device may derive the residual samples of the current block by performing an inverse transform on the transform coefficients.

[0253] The decoding device according to an embodiment may generate a reconstructed picture based on the residual samples of the current block (S1730). More specifically, the adder 340 of the decoding device may generate a reconstructed picture based on the residual samples of the current block.

[0254] In an embodiment, the residual information may include information indicating the positions of the last significant coefficients such as last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix. More specifically, last_sig_coeff_x_prefix represents the prefix of the column position of the last significant coefficient in the scan order in the transform block, and last_sig_coeff_y_prefix represents the prefix of the row position of the last significant coefficient in the scan order in the transform block, and last_sig_coeff_x_suffix represents the suffix of the column position of the last significant coefficient in the scan order in the transform block, and last_sig_coeff_y_suffix represents the suffix of the row position of the last significant coefficient in the scan order in the transform block. Here, the significant coefficient may represent a non-zero coefficient. The scan order may be a right-up diagonal scan order. Alternatively, the scan order may be a horizontal scan order or a vertical scan order. According to an example, the information indicating the positions of the last significant coefficients may be encoded and decoded only when the transform skip is not applied to the transform block.

[0255] According to an example, when the transform skip is not applied to the transform block, the information indicating the positions of the last significant coefficients may be encoded and decoded based on the size of the transform block. For example, the information indicating the positions of the last significant coefficients may be encoded and decoded only when the width and height of the transform block are less than a predetermined threshold.

[0256] According to another example, the residual information may include coded_sub_block_flag indicating whether the sub-block includes the last non-zero coefficient, and the coded_sub_block_flag may be encoded and decoded only when the transform skip is not applied to the transform block.

[0257] According to the example, when the transform skip is not applied to the transform block, the coded_sub_block_flag can be encoded and decoded based on the size of the transform block. For example, the coded_sub_block_flag can be encoded and decoded only when the width and height of the transform block are less than a predetermined threshold.

[0258] In one embodiment, the residual information includes a parity level flag regarding the parity of the transform coefficient levels of the quantized transform coefficients and a first transform coefficient level flag regarding whether the transform coefficient level is greater than a first threshold. In an example, the parity level flag indicates par_level_flag, the first transform coefficient level flag indicates rem_abs_gt1_flag or abs_level_gtx_flag[n][0], and the second transform coefficient level flag indicates rem_abs_gt2_flag or abs_level_gtx_flag[n][1].

[0259] In an embodiment, the step of deriving the residual information may include: deriving the value of the parity level flag and the value of the first transform coefficient level flag based on the quantized transform coefficients, and decoding the first transform coefficient level and decoding the parity level flag.

[0260] In an embodiment, the decoding of the first transform coefficient level flag may be performed before the decoding of the parity level flag. For example, the decoding device may perform the decoding of rem_abs_gt1_flag or abs_level_gtx_flag[n][0] before decoding par_level_flag.

[0261] In an embodiment, the residual information may include a valid coefficient flag indicating whether the quantized transform coefficient is a non-zero valid coefficient and a second transform coefficient level flag indicating whether the transform coefficient level of the quantized transform coefficient is greater than a second threshold. In an example, the valid coefficient flag may be sig_coeff_flag.

[0262] The residual information may include context-coded context syntax elements, and the context syntax elements may include a valid coefficient flag, a parity level flag, a first transform coefficient level flag, and a second transform coefficient level flag.

[0263] In an embodiment, the step of deriving the quantized transform coefficients may decode the context syntax elements based on the context and based on a predetermined maximum value of the context syntax elements.

[0264] In other words, the sum of the number of valid coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags of the quantization transform coefficients in the current block included in the residual information can be less than or equal to a predetermined maximum value.

[0265] This maximum value is the sum of the number of valid coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags of the quantization transform coefficients related to the current sub-block in the current block.

[0266] In an embodiment, the maximum value can be determined in units of transform blocks. The current block can be a sub-block within a transform block that is a transform unit, and the quantization transform coefficients are encoded in units of sub-blocks. When encoding the quantization transform coefficients in units of sub-blocks, the context syntax elements in the residual information can be decoded based on the maximum value determined in units of transform blocks.

[0267] In an embodiment, the threshold can be determined based on the size of the current block (or the current sub-block within the current block). If the current block is a transform block, the threshold can be determined based on the size of the transform block.

[0268] In an embodiment, when the sum of the number of valid coefficient flags, the number of first transform coefficient level flags, the number of parity level flags, and the number of second transform coefficient level flags derived from the 0th quantization transform coefficient to the nth quantization transform coefficient determined by the coefficient scan order reaches the predetermined maximum value, for the (n + 1)th quantization transform coefficient determined by the coefficient scan order, the explicit signaling of the valid coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level can be omitted, and the value of the (n + 1)th quantization transform coefficient can be derived based on the value of the coefficient level information included in the residual information.

[0269] For example, when the sum of the number of sig_coeff_flag, the number of rem_abs_gt1_flag (or abs_level_gtx_flag[n][0]), the number of par_level_flags, and the number of rem_abs_gt2_flag (or abs_level_gtx_flag[n][1]) derived from the 0th quantization transform coefficient (or the first quantization transform coefficient) to the nth quantization transform coefficient (or the nth quantization transform coefficient) determined by the coefficient scan order reaches a predetermined maximum value, for the (n + 1)th quantization transform coefficient determined by the coefficient scan order, the explicit signaling of sig_coeff_flag, rem_abs_gt1_flag (or abs_level_gtx_flag[n][0]), par_level_flag, abs_level_gtx_flag[n][1], and rem_abs_gt2_flag (or abs_level_gtx_flag[n][1]) can be omitted, and the value of the (n + 1)th quantization transform coefficient can be derived based on the value of abs_remainder or dec_abs_level included in the residual information.

[0270] In an embodiment, the significant coefficient flag, the first transform coefficient level flag, the parity level flag, and the second transform coefficient level flag included in the residual information may be context - based coded, and the coefficient level information may be bypass - based coded.

[0271] Although in the above - mentioned embodiment, the method is described based on a flowchart having a series of steps or block diagrams, the present disclosure is not limited to the order of the above - mentioned steps or block diagrams, and specific steps may occur simultaneously with other steps or occur in an order different from the above - mentioned order. In addition, those of ordinary skill in the art can understand that the steps shown in the flowchart are not exhaustive, and another step may be incorporated or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.

[0272] The foregoing method according to the present disclosure may be implemented in software form, and the encoding device and / or decoding device according to the present disclosure may be included in a device for performing image processing such as a television, a computer, a smart phone, a set - top box, and a display device.

[0273] When the embodiments in the present disclosure are implemented in software, the above - mentioned methods can be implemented as modules (processes or functions, etc.) for performing the above - mentioned functions. These modules can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor in various well - known ways. The processor can include an application - specific integrated circuit (ASIC), different chip sets, logic circuits, and / or a data processor. The memory can include a read - only memory (ROM), a random - access memory (RAM), a flash memory, a memory card, a storage medium, and / or another storage device. That is, the embodiments described in the present disclosure can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, the information for the embodiments (e.g., information about instructions) or algorithms can be stored in a digital storage medium.

[0274] In addition, the decoding device and the encoding device applying the present disclosure can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real - time communication device (e.g., video communication), a mobile streaming device, a storage medium, a portable camera, a video - on - demand (VoD) service providing device, an over - the - top (OTT) video device, an Internet streaming service providing device, a three - dimensional (3D) video device, a virtual reality device, an augmented reality device, a video - phone video device, a vehicle terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an aircraft terminal, a ship terminal, etc.), and a medical video device, and can be used to process video signals or data signals. For example, an over - the - top (OTT) video device can include a game console, a Blu - ray player, an Internet - access TV, a home theater system, a smart phone, a tablet PC, and a digital video recorder (DVR), etc.

[0275] In addition, the processing method of the present disclosure can be generated in the form of a program executed by a computer and stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes various storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). Additionally, a bitstream generated by an encoding method can be stored in a computer-readable recording medium, or the generated bitstream can be transmitted via a wired or wireless communication network.

[0276] In addition, embodiments of the present disclosure can be implemented as a computer program product by program code, and the program code can be executed in a computer by embodiments of the present disclosure. The program code can be stored on a computer-readable carrier.

[0277] Figure 19 An example of a content streaming system to which the content of this document can be applied is shown.

[0278] Refer to Figure 19 , a content streaming system applying the content of this document generally can include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0279] The encoding server is used to compress the content input from a multimedia input device (e.g., a smart phone, a camera, and a portable video camera, etc.) into digital data to generate a bitstream and send it to the streaming server. As another example, in the case where the multimedia input device (e.g., a smart phone, a camera, and a portable video camera, etc.) directly generates a bitstream, the encoding server can be omitted.

[0280] A bitstream can be generated by applying the encoding method or bitstream generation method of this document. And, the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.

[0281] The streaming server sends multimedia data to the user device via the web server based on the user's request, and the web server serves as a tool to notify the user of what services are available. When the user requests a service that the user wants, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this regard, the content streaming system can include a separate control server, and in this case, the control server is used to control the commands / responses between various devices in the content streaming system.

[0282] The streaming server can receive content from a media storage and / or an encoding server. For example, in the case of receiving content from an encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to smoothly provide the streaming service.

[0283] For example, the user device can include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a slate PC, a tablet PC, a superbook, a wearable device (e.g., a watch-type terminal (smart watch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, or a digital signage, etc.

[0284] Each server in the content streaming system can operate as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.

Claims

1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Receive a bitstream including residual information; Derive quantization transform coefficients of a current block based on the residual information included in the bitstream; Derive residual samples of the current block based on the quantization transform coefficients; and generate a reconstructed picture based on the residual samples of the current block, wherein the residual information includes position information related to positions of last non-zero transform coefficients in a transform block, wherein, based on a value of a transform skip flag related to whether transform skip is applied to the transform block being equal to 0 and when a size of the transform block is equal to or greater than a certain value, the position information is parsed; wherein, based on the value of the transform skip flag being equal to 1, parsing of the position information is skipped.

2. The image decoding method according to claim 1, wherein, The position information includes information related to a prefix of a column position of a last valid coefficient in a scan order in the transform block and information related to a prefix of a row position of the last valid coefficient in the scan order in the transform block and information related to a suffix of the column position of the last valid coefficient in the scan order in the transform block and information related to a suffix of the row position of the last valid coefficient in the scan order in the transform block.

3. The image decoding method according to claim 1, wherein, The residual information includes a valid coefficient flag related to whether the quantization transform coefficient is a non-zero valid coefficient, a parity level flag regarding parity of a transform coefficient level of the quantization transform coefficient, a first transform coefficient level flag related to whether the transform coefficient level is greater than a first threshold, and a second transform coefficient level flag related to whether the transform coefficient level of the quantization transform coefficient is greater than a second threshold, and wherein the step of deriving the quantization transform coefficients includes the following steps: Decode the first transform coefficient level flag and decode the parity level flag; and Derive the quantization transform coefficients based on the value of the decoded parity level flag and the value of the decoded first transform coefficient level flag, wherein decoding of the first transform coefficient level flag is performed before decoding of the parity level flag.

4. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: Derive residual samples of a current block; Derive quantization transform coefficients based on the residual samples of the current block; and Encode residual information including information about the quantization transform coefficients, wherein the residual information includes position information related to positions of last non-zero transform coefficients in a transform block, wherein, based on a value of a transform skip flag related to whether transform skip is applied to the transform block being equal to 0 and the size of the transform block being equal to or greater than a certain value, the position information is encoded; wherein, based on the value of the transform skip flag being equal to 1, encoding of the position information is skipped.

5. The image encoding method according to claim 4, wherein, The position information includes information related to a prefix of a column position of a last valid coefficient in a scan order in the transform block, information related to a prefix of a row position of the last valid coefficient in the scan order in the transform block, information related to a suffix of the column position of the last valid coefficient in the scan order in the transform block, and information related to a suffix of the row position of the last valid coefficient in the scan order in the transform block.

6. The image encoding method according to claim 4, wherein, The residual information includes a valid coefficient flag related to whether the quantized transform coefficient is a non-zero valid coefficient, a parity level flag regarding parity of a transform coefficient level of the quantized transform coefficient, a first transform coefficient level flag related to whether the transform coefficient level is greater than a first threshold, and a second transform coefficient level flag related to whether the transform coefficient level of the quantized transform coefficient is greater than a second threshold, and wherein, the step of deriving the quantized transform coefficient includes the following steps: encoding the first transform coefficient level flag and encoding the parity level flag; and deriving the quantized transform coefficient based on a value of the encoded parity level flag and a value of the encoded first transform coefficient level flag, wherein, the encoding of the first transform coefficient level flag is performed before the encoding of the parity level flag.

7. A non - transitory computer - readable storage medium storing encoded information generated by a video encoding method, the method comprising the following steps: Derive residual samples of a current block; derive quantized transform coefficients based on the residual samples of the current block; and encode residual information including information about the quantized transform coefficients, wherein, the residual information includes position information related to a position of a last non-zero transform coefficient in a transform block, wherein, based on a value of a transform skip flag related to whether transform skip is applied to the transform block being equal to 0 and a size of the transform block being equal to or greater than a certain value, the position information is encoded; wherein, based on the value of the transform skip flag being equal to 1, the encoding of the position information is skipped.

Citation Information

Patent Citations

  • Method and device for coding residual signal in video coding system

    KR1020180048739A

  • Coding significant coefficient information in transform skip mode

    US20130114730A1