Image decoding method and apparatus therefor

Through the image decoding method and device, an entropy decoder, a predictor and a residual processor are used to generate reconstructed samples. Combined with a simplified residual data coding structure and a syntax element-priority coding order, the high cost problem of high-resolution and high-quality images is solved, the coding efficiency is improved and the transmission and storage costs are reduced.

CN120602645APending Publication Date: 2025-09-05LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510886847.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-08-31
Filing Date
2020-08-31
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The transmission and storage costs of high-resolution and high-quality images are high, and efficient image compression technology is required to reduce the amount of information and compilation complexity.

Method used

The invention discloses an image decoding method and apparatus, which utilizes an entropy decoder, a predictor, a residual processor and an adder to generate reconstructed samples, and combines a simplified residual data coding structure and a syntax element-prioritized coding order to improve residual coding efficiency.

Benefits of technology

The complexity of syntax element compilation in bypass compilation is reduced, the overall residual compilation efficiency is improved, and the transmission and storage costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602645A_ABST
    Figure CN120602645A_ABST
Patent Text Reader

Abstract

The invention provides an image decoding method and an apparatus therefor. An image decoding method performed by a decoding device according to the present invention comprises the steps of: acquiring image data including residual data and prediction mode data for a current block via a bitstream; deriving a prediction mode of the current block based on the prediction mode data; deriving a prediction sample for the current block based on the prediction mode; deriving a residual sample for the current block based on the residual data; and generating a reconstructed sample for the current block based on the prediction sample and the residual sample.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with application number 202080068623.6 (PCT / KR2020 / 011614) filed on March 30, 2022, application date on August 31, 2020, and titled "Image decoding method and device thereof". Technical Field

[0002] The present disclosure relates to an image coding technology, and more particularly, to an image coding method and apparatus thereof, in which, in an image coding system, in residual data coded according to TSRC of a current block, subsequent residual data is coded according to a simplified residual data coding structure when all bins for maximum available context coding of the current block are used. Background Art

[0003] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition), is growing across various fields. Because image data has high resolution and quality, the amount of information, or bits, required for transmission increases compared to conventional image data. Consequently, when image data is transmitted using media such as conventional wired / wireless broadband lines or stored using existing storage media, transmission and storage costs increase.

[0004] Therefore, there is a need for efficient image compression technology for effectively transmitting, storing, and reproducing information of high-resolution and high-quality images. Summary of the Invention

[0005] Technical issues

[0006] The present disclosure provides a method and apparatus for improving image coding efficiency.

[0007] The present disclosure also provides a method and apparatus for improving residual coding efficiency.

[0008] Technical Solution

[0009] According to an embodiment of the present document, an image decoding method performed by a decoding device is provided. The method includes: obtaining image information including residual information and prediction mode information of a current block through a bitstream, deriving a prediction mode of the current block based on the prediction mode information, deriving prediction samples of the current block based on the prediction mode, deriving residual samples of the current block based on the residual information, and generating reconstructed samples of the current block based on the prediction samples and the residual samples.

[0010] According to another embodiment of the present document, a decoding device for performing image decoding is provided. The decoding device includes: an entropy decoder configured to obtain image information including residual information and prediction mode information of a current block through a bitstream; a predictor configured to derive a prediction mode of the current block based on the prediction mode information and derive prediction samples of the current block based on the prediction mode; a residual processor configured to derive residual samples of the current block based on the residual information; and an adder configured to generate reconstructed samples of the current block based on the prediction samples and the residual samples.

[0011] According to another embodiment of the present document, an image encoding method performed by an encoding device is provided. The method includes: deriving prediction samples of a current block based on inter-frame prediction or intra-frame prediction, deriving residual samples of the current block based on the prediction samples, deriving transform coefficients of the current block based on the residual samples, and encoding image information including prediction mode information of the current block and residual information for the transform coefficients.

[0012] According to another embodiment of the present document, an image encoding device is provided. The encoding device includes: a predictor configured to derive prediction samples of a current block based on inter-frame prediction or intra-frame prediction; a residual processor configured to derive residual samples of the current block based on the prediction samples and derive transform coefficients of the current block based on the residual samples; and an entropy encoder configured to encode image information including prediction mode information of the current block and residual information for the transform coefficients.

[0013] According to another embodiment of the present document, a non-transitory computer-readable storage medium storing a bitstream including image information for executing an image decoding method is provided. In the non-transitory computer-readable storage medium, the image decoding method includes: obtaining image information including residual information and prediction mode information of a current block through the bitstream; deriving a prediction mode for the current block based on the prediction mode information; deriving prediction samples for the current block based on the prediction mode; deriving residual samples for the current block based on the residual information; and generating reconstructed samples for the current block based on the prediction samples and the residual samples.

[0014] Beneficial effects

[0015] According to the present disclosure, the efficiency of residual coding can be improved.

[0016] According to the present disclosure, when the maximum number of context-coded bins for the current block is consumed in TSRC, syntax elements according to a simplified residual data coding structure can be signaled, and by this, the coding complexity of the bypass-coded syntax elements is reduced, and the overall residual coding efficiency can be improved.

[0017] According to the present disclosure, as the coding order of bypass-coded syntax elements, an order in which syntax elements are prioritized may be used, and by this, the coding efficiency of bypass-coded syntax elements may be improved, and the overall residual coding efficiency may be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 An example of a video / image coding device to which an embodiment of the present disclosure can be applied is briefly illustrated.

[0019] Figure 2 is a schematic diagram illustrating a configuration of a video / image encoding device to which an embodiment of the present disclosure can be applied.

[0020] Figure 3 is a schematic diagram illustrating a configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied.

[0021] Figure 4 An example of a video / image encoding method based on intra-frame prediction is illustrated.

[0022] Figure 5 An example of a video / image encoding method based on intra-frame prediction is illustrated.

[0023] Figure 6 The intra prediction process is schematically illustrated.

[0024] Figure 7 An example of a video / image encoding method based on inter-frame prediction is illustrated.

[0025] Figure 8 An example of a video / image decoding method based on inter-frame prediction is illustrated.

[0026] Figure 9 The inter-frame prediction process is schematically illustrated.

[0027] Figure 10 Context-Adaptive Binary Arithmetic Coding (CABAC) for encoding syntax elements is exemplarily shown.

[0028] Figure 11 is a diagram showing exemplary transform coefficients within a 4×4 block.

[0029] Figure 12 Illustrated is an example in which syntax elements are compiled in TSRC.

[0030] Figure 13 Another example is illustrated in which syntax elements are coded in TSRC.

[0031] Figure 14 Another example is illustrated in which syntax elements are coded in TSRC.

[0032] Figure 15 The diagram illustrates an example of coding bypass-coded syntax elements in a TSRC in a syntax element-first coding order rather than a coefficient position-first coding order.

[0033] Figure 16 The diagram illustrates an example of coding bypass-coded syntax elements in a TSRC in a syntax element-first coding order rather than a coefficient position-first coding order.

[0034] Figure 17 Illustrated is an example in which syntax elements are coded in a simplified residual data coding structure.

[0035] Figure 18 An example of coding bypass-coded syntax elements in a simplified residual data coding structure in a syntax element-first coding order rather than a coefficient position-first coding order is illustrated.

[0036] Figure 19a and 19b An embodiment in which syntax elements are coded in a simplified residual data coding structure is illustrated.

[0037] Figure 20a and 20b An example of coding bypass-coded syntax elements in a simplified residual data coding structure in a syntax element-first coding order rather than a coefficient position-first coding order is illustrated.

[0038] Figure 21 An embodiment in which syntax elements are coded in a simplified residual data coding structure is illustrated.

[0039] Figure 22 An example of coding bypass-coded syntax elements in a simplified residual data coding structure in a syntax element-first coding order rather than a coefficient position-first coding order is illustrated.

[0040] Figure 23 The following briefly illustrates an image encoding method performed by an encoding device according to the present disclosure.

[0041] Figure 24 A coding apparatus for performing an image coding method according to the present disclosure is briefly illustrated.

[0042] Figure 25 The following briefly illustrates an image decoding method performed by a decoding device according to the present disclosure.

[0043] Figure 26 A decoding device for performing an image decoding method according to the present disclosure is briefly illustrated.

[0044] Figure 27The present invention is a structural diagram of a content streaming system to which the present invention is applied. DETAILED DESCRIPTION

[0045] The present disclosure can be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, the embodiments are not intended to limit the present disclosure. The terms used in the following description are only used to describe specific embodiments and are not intended to limit the present disclosure. As long as it is clearly understood in different ways, singular expressions include plural expressions. Terms such as "including" and "having" are intended to indicate the presence of features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and it should be understood that the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof is not excluded.

[0046] Furthermore, the elements in the drawings described in this disclosure are drawn independently for the purpose of conveniently explaining different specific functions, and do not imply that these elements are embodied by independent hardware or independent software. For example, two or more elements in the drawings may be combined to form a single element, or an element may be divided into multiple elements. Embodiments of combining elements and / or dividing elements belong to the present disclosure and do not depart from the concepts of the present disclosure.

[0047] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In addition, throughout the drawings, like reference numerals are used to indicate like elements, and the same description of like elements will be omitted.

[0048] Figure 1 An example of a video / image coding device to which an embodiment of the present disclosure can be applied is briefly illustrated.

[0049] refer to Figure 1 The video / image coding system may include a first device (source device) and a second device (receiving device). The source device may send coded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or a network.

[0050] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0051] A video source can acquire video / images through a process that captures, synthesizes, or generates video / images. A video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet computer, or a smartphone, and may (electronically) generate video / images. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process that generates relevant data.

[0052] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, and quantization to achieve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0053] The transmitter can transmit the encoded image / image information or data in the form of a bitstream, output in the form of a file or stream, via a digital storage medium or network to a receiver of a receiving device. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating a media file in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received bitstream to a decoding device.

[0054] The decoding device may decode the video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.

[0055] The renderer may render the decoded video / image, and the rendered video / image may be displayed on a display.

[0056] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed in this disclosure can be applied to methods disclosed in Versatile Video Coding (VVC), EVC (Essential Video Coding) standards, AOMedia Video 1 (AV1) standards, the second generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).

[0057] The present disclosure presents various embodiments of video / image coding, and unless otherwise mentioned, the embodiments may be performed in combination with each other.

[0058] In the present disclosure, video may refer to a series of images over time. Generally, a picture refers to a unit representing an image in a specific time zone, and a sub-picture / slice / tile is a unit that constitutes part of a picture in coding. A sub-picture / slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more sub-pictures / slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A brick may represent a rectangular area of ​​CTU rows within a tile in a picture. A tile may be partitioned into multiple tiles, each tile consisting of one or more CTU rows within the tile. A tile that is not partitioned into multiple tiles may also be referred to as a tile. Tile scanning may sort the CTUs of a partitioned picture into a specific order, wherein CTUs are sorted consecutively within a tile by a CTU raster scan, tiles within a tile are sorted consecutively by a raster scan of the tiles of the tile, and tiles within a picture are sorted consecutively by a raster scan of the tiles of the picture. Additionally, a sub-picture can represent a rectangular area of ​​one or more slices within a picture. That is, a sub-picture contains one or more slices that collectively cover a rectangular area of ​​a picture. A tile is a rectangular area of ​​a CTU within a specific tile column and tile row in a picture. A tile column is a rectangular area of ​​a CTU whose height equals the height of the picture and whose width is specified by syntax elements in the picture parameter set. A tile row is a rectangular area of ​​a CTU whose height is specified by syntax elements in the picture parameter set and whose width equals the width of the picture. Tile scan is a specific sequential ordering of CTUs that partition a picture, where CTUs can be ordered consecutively within a tile using a CTU raster scan, while tiles within a picture can be ordered consecutively using a raster scan of the tiles of a picture. A slice comprises an integer number of tiles of a picture that can be exclusively contained in a single NAL unit. A slice can consist of multiple complete tiles or only a contiguous sequence of complete tiles of a tile. In this disclosure, tile group and slice may be used interchangeably. For example, in this disclosure, a tile group / tile group header may be referred to as a slice / slice header.

[0059] A pixel or picture element (pel) can represent the smallest unit that makes up a picture (or image). "Sample" can also be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.

[0060] A unit can represent the basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to that region. A unit can include a luma block and two chroma (e.g., CB, CR) blocks. In some cases, the term "unit" can be used interchangeably with terms such as "block" or "region." In general, an M×N block can include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients.

[0061] In this specification, "A or B" may mean "only A," "only B," or "A and B." In other words, in this specification, "A or B" may be interpreted as "A and / or B." For example, "A, B, or C" herein means "only A," "only B," "only C," or "any one and any combination of A, B, and C."

[0062] As used herein, a slash ( / ) or a comma may mean "and / or." For example, "A / B" may mean "A and / or B." Thus, "A / B" may mean "only A," "only B," or "A and B." For example, "A, B, C" may mean "A, B, or C."

[0063] In this specification, “at least one of A and B” may mean “only A,” “only B,” or “both A and B.” In addition, in this specification, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted as being the same as “at least one of A and B.”

[0064] In addition, in this specification, "at least one of A, B, and C" means "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".

[0065] Furthermore, parentheses used in this specification may refer to "for example." Specifically, when "prediction (intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-frame prediction"; "intra-frame prediction" may be provided as an example of "prediction." Furthermore, even when "prediction (i.e., intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction."

[0066] In this specification, technical features described separately in one drawing may be implemented separately or simultaneously.

[0067] The following figures are created to explain specific examples of this specification. Since the names of specific devices or the names of specific signals / messages / fields described in the figures are presented by way of example, the technical features of this specification are not limited to the specific names used in the following figures.

[0068] Figure 2 is a schematic diagram illustrating a configuration of a video / image encoding device to which an embodiment of the present disclosure can be applied. Hereinafter, a video encoding device may include an image encoding device.

[0069] refer to Figure 2 The encoding device 200 includes an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. Depending on the embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be composed of at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) or may be composed of a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.

[0070] The image splitter 210 can split the input image (or picture, or frame) input to the encoding device 200 into one or more processors. For example, a processor may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) based on a quadtree-binary-ternary-tree (QTBTTT) structure. For example, a coding unit may be split into multiple coding units of increasing depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, followed by the binary tree structure and / or the ternary structure. Alternatively, the binary tree structure may be applied first. The coding process according to the present disclosure may be performed based on the final coding unit that is no longer split. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics. Alternatively, if necessary, the coding unit may be recursively split into coding units of increasing depth, and the coding unit of the optimal size may be used as the final coding unit. The coding process may include prediction, transformation, and reconstruction, which will be described later. As another example, the processor may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be split or partitioned from the final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0071] In some cases, the term "unit" can be used interchangeably with terms such as "block" or "region." In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or pixel value, and can represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. Sample can be used as a term corresponding to a picture (or image) of pixels or picture elements.

[0072] In the encoding device 200, the prediction signal (prediction block, prediction sample array) output by the inter-frame predictor 221 or the intra-frame predictor 222 is subtracted from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the transformer 232. In this case, as shown in the figure, the unit in the encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as the subtractor 231. The predictor can perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction on a per-block or CU basis. As described later in the description of each prediction mode, the predictor can generate various information related to the prediction, such as prediction mode information, and transmit the generated information to the entropy encoder 240. The prediction information can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0073] The intra-frame predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or can be far away from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.

[0074] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. It can also include information about the inter-frame prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture containing the reference block and the reference picture containing the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture containing temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 may use the motion information of a neighboring block as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of a neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling a motion vector difference.

[0075] The predictor 220 can generate prediction signals based on various prediction methods described below. For example, the predictor can predict a block using not only intra prediction or inter prediction, but also both intra and inter prediction simultaneously. This can be referred to as combined intra-frame prediction (CIIP). Furthermore, the predictor can predict blocks based on intra block copy (IBC) prediction mode or palette mode. IBC prediction mode or palette mode can be used for content image / video coding for games, such as screen content coding (SCC). IBC essentially performs prediction within the current picture, but can be performed similarly to inter prediction because the reference block is derived from the current picture. That is, IBC can use at least one of the inter prediction techniques described in this disclosure. Palette mode can be considered an example of intra coding or intra prediction. When palette mode is applied, sample values ​​within the picture can be signaled based on information about the palette table and palette index.

[0076] The prediction signal generated by the predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstructed signal or a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size, or it can be applied to blocks of variable size other than square.

[0077] The quantizer 233 quantizes the transform coefficients and transmits them to the entropy encoder 240. The entropy encoder 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs a bitstream. This information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 rearranges the block-type quantized transform coefficients into a one-dimensional vector based on the coefficient scanning order and generates information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. This information about the transform coefficients can be generated. The entropy encoder 240 can implement various encoding methods, such as Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Besides the quantized transform coefficients, the entropy encoder 240 can encode information required for video / image reconstruction (e.g., syntax element values) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the bitstream in units of Network Abstraction Layers (NALs). The video / image information may also include information regarding various parameter sets, such as the Adaptation Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Furthermore, the video / image information may include general constraint information. In the present disclosure, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the aforementioned encoding process and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoder 240 and / or a storage unit (not shown) for storing the signal may be included as internal / external components of the encoding device 200. Alternatively, the transmitter may be included in the entropy encoder 240.

[0078] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients using the dequantizer 234 and the inverse transformer 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If the block to be processed has no residual (such as when skip mode is applied), the prediction block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture through filtering as described below.

[0079] Furthermore, during picture encoding and / or reconstruction, luma mapping and chroma scaling (LMCS) may be applied.

[0080] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and store the modified reconstructed image in the memory 270 (specifically, the DPB of the memory 270). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240, as described later in the description of various filtering methods. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.

[0081] The modified reconstructed picture sent to the memory 270 may be used as a reference picture in the inter-frame predictor 221. When inter-frame prediction is applied by the encoding apparatus, prediction mismatch between the encoding apparatus 200 and the decoding apparatus may be avoided, and coding efficiency may be improved.

[0082] The DPB of the memory 270 can store modified reconstructed pictures used as reference pictures in the inter-frame predictor 221. The memory 270 can store motion information of blocks from which motion information in the current picture was derived (or encoded) and / or motion information of reconstructed blocks in the picture. This stored motion information can be sent to the inter-frame predictor 221 and used as motion information for spatially or temporally neighboring blocks. The memory 270 can also store reconstructed samples of reconstructed blocks in the current picture and transmit the reconstructed samples to the intra-frame predictor 222.

[0083] Figure 3 is a schematic diagram illustrating a configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied.

[0084] refer to Figure 3 The decoding apparatus 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. Depending on the embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured by hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0085] When a bit stream including video / image information is input, the decoding apparatus 300 can be used with Figure 2 The image is reconstructed accordingly to the processing of the video / image information in the encoding device. For example, the decoding device 300 can derive the unit / block based on the block segmentation related information obtained from the bit stream. The decoding device 300 can perform decoding using the processor applied in the encoding device. Therefore, the decoding processor can be, for example, a coding unit, and the coding unit can be divided from the coding tree unit or the maximum coding unit according to the quadtree structure, the binary tree structure and / or the ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.

[0086] The decoding device 300 may receive the data in the form of a bit stream from Figure 2The received signal is output by the encoding device, and the entropy decoder 310 can decode the received signal. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as the Adaptation Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the parameter set information and / or general constraint information. The signaled / received information and / or syntax elements described later in this disclosure can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 decodes the information in the bitstream based on a coding method such as Exponential Golomb coding, CAVLC, or CABAC, and outputs the syntax elements required for image reconstruction and the quantized values ​​of the residual transform coefficients. More specifically, the CABAC entropy decoding method receives a bin (bin of bits) corresponding to each syntax element in the bitstream, uses information about the target syntax element to be decoded, information about the decoded target block to be decoded, or information about a symbol / bin decoded at a previous stage to determine a context model. It then arithmetically decodes the bin by predicting the probability of occurrence based on the determined context model, generating a symbol corresponding to the value of each syntax element. After determining the context model, the CABAC entropy decoding method updates the context model by applying the decoded symbol / bin information to the context model for the next symbol / bin. Information related to prediction within the information decoded by the entropy decoder 310 is provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and residual values ​​(i.e., quantized transform coefficients and related parameter information) entropy-decoded within the entropy decoder 310 are input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). Furthermore, information related to filtering within the information decoded by the entropy decoder 310 is provided to the filter 350. In addition, a receiver (not shown) for receiving a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiver may be a component of the entropy decoder 310. Furthermore, the decoding device according to the present disclosure may be referred to as a video / image / picture decoding device, and the decoding device may be categorized as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0087] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The dequantizer 321 can dequantize the quantized transform coefficients by using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.

[0088] The inverse transformer 322 performs an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0089] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310, and may determine a specific intra / inter prediction mode.

[0090] The predictor 330 can generate prediction signals based on various prediction methods described below. For example, the predictor can predict a block using not only intra prediction or inter prediction, but also both intra and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). Furthermore, the predictor can predict blocks based on intra block copy (IBC) prediction mode or palette mode. IBC prediction mode or palette mode can be used for content image / video coding for games, such as screen content coding (SCC). IBC essentially performs prediction within the current picture, but can be performed similarly to inter prediction because reference blocks are derived from the current picture. That is, IBC can use at least one of the inter prediction techniques described in this disclosure. Palette mode can be considered an example of intra coding or intra prediction. When palette mode is applied, sample values ​​within the picture can be signaled based on information about the palette table and palette index.

[0091] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples may be located near the current block or may be located far away from the current block. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 may determine the prediction mode to be applied to the current block by using the prediction modes applied to the neighboring blocks.

[0092] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. It can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and prediction information can include information indicating the inter-frame prediction mode for the current block.

[0093] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331). If the block to be processed has no residual (for example, when skip mode is applied), the prediction block can be used as the reconstructed block.

[0094] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, may be output through filtering as described below, or may be used for inter-frame prediction of the next picture.

[0095] Additionally, luma mapping and chroma scaling (LMCS) can be applied during picture decoding.

[0096] The filter 350 can improve the subjective and objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and store the modified reconstructed image in the memory 360 (specifically, the DPB of the memory 360). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.

[0097] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block from which the motion information in the current picture was derived (or decoded) and / or the motion information of the reconstructed block in the picture. The stored motion information can be sent to the inter-frame predictor 332 to be used as motion information for spatially or temporally neighboring blocks. The memory 360 can also store reconstructed samples of the reconstructed block in the current picture and transmit the reconstructed samples to the intra-frame predictor 331.

[0098] In the present disclosure, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may be the same as the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300 or may be applied to correspond to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300, respectively. The same contents can also be applied to the inter-frame predictor 332 and the intra-frame predictor 331.

[0099] In the present disclosure, at least one of quantization / inverse quantization and / or transform / inverse transform may be omitted. When quantization / inverse quantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or, for uniformity of expression, may still be referred to as a transform coefficient.

[0100] In this disclosure, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and this information may be signaled using residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficients), and scaled transform coefficients may be derived by inversely transforming (scaling) the transform coefficients. Residual samples may be derived based on inversely transforming (transforming) the scaled transform coefficients. This may also be applied / expressed in other parts of this disclosure.

[0101] As mentioned above, when performing video coding, prediction is performed to improve compression efficiency. This allows a prediction block to be generated, including prediction samples of the current block, as the block to be coded (i.e., the coding target block). Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived in the same manner in both the encoding and decoding devices, and the encoding device can signal information regarding the residual difference between the original block and the prediction block (residual information) to the decoding device, rather than the original sample values ​​of the original block, thereby improving image coding efficiency. Based on the residual information, the decoding device can derive a residual block including residual samples, add the residual block to the prediction block to generate a reconstructed block including reconstructed samples, and generate a reconstructed picture including the reconstructed block.

[0102] Residual information can be generated through a transform and quantization process. For example, the encoding device may derive a residual block between the original block and the predicted block, perform a transform process on the residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization process on the transform coefficients to derive quantized transform coefficients, and signal the relevant residual information to the decoding device (via the bitstream). This residual information may include information on the values ​​of the quantized transform coefficients, their locations, the transform technique, the transform kernel, and the values ​​of the quantization parameters. Based on this residual information, the decoding device may perform a dequantization / inverse transform process and derive residual samples (or residual blocks). The decoding device may generate a reconstructed picture based on the predicted block and the residual block. Furthermore, the encoding device may dequantize / inverse transform the quantized transform coefficients to derive a residual block for use as a reference for inter-frame prediction of subsequent reference pictures, and generate a reconstructed picture based on this residual block.

[0103] Intra-frame prediction may refer to generating prediction samples for the current block based on reference samples in the picture to which the current block belongs (hereinafter referred to as the current picture). When intra-frame prediction is applied to the current block, neighboring reference samples to be used for intra-frame prediction of the current block may be derived. Neighboring reference samples for the current block may include samples adjacent to the left boundary of the current block (nWxnH), a total of 2xnH samples adjacent to the lower left of the current block, samples adjacent to the upper boundary of the current block, a total of 2xnW samples adjacent to the upper right, and samples adjacent to the upper left of the current block. Alternatively, the neighboring reference samples for the current block may include multiple columns of upper neighboring samples and multiple rows of left neighboring samples. Furthermore, the neighboring reference samples for the current block may include a total of nH samples adjacent to the right boundary of the current block (nWxnH), a total of nW samples adjacent to the lower boundary of the current block, and samples adjacent to the lower right of the current block.

[0104] However, some neighboring reference samples of the current block have not yet been decoded or may be unavailable. In this case, the decoder can construct neighboring reference samples to be used for prediction by replacing unavailable samples with available samples. Alternatively, the neighboring reference samples to be used for prediction can be configured by interpolation of available samples.

[0105] When deriving adjacent reference samples, (i) the prediction sample can be derived based on an average or interpolation of adjacent reference samples of the current block, or (ii) the prediction sample can be derived based on a reference sample existing in a specific (prediction) direction relative to the prediction sample in the adjacent reference samples of the current block. Case (i) can be referred to as a non-directional mode or a non-angular mode, and case (ii) can be referred to as a directional mode or an angular mode.

[0106] Alternatively, prediction samples can be generated by interpolating the first neighboring sample located in the prediction direction of the intra prediction mode of the current block and the second neighboring sample located in the opposite direction of the prediction direction among the neighboring reference samples. This is referred to as linear interpolation intra prediction (LIP). Furthermore, a linear model (LM) can be used to generate chroma prediction samples based on luma samples. This is referred to as LM mode or chroma component LM (CCLM) mode.

[0107] In addition, a temporary prediction sample of the current block is derived based on the filtered adjacent reference samples, and the prediction sample of the current block can also be derived by weighted summing the temporary prediction sample with at least one reference sample derived according to the intra prediction mode among the existing adjacent reference samples (i.e., unfiltered adjacent reference samples). The above situation can be called position-dependent intra prediction (PDPC).

[0108] In addition, a reference sample line with the highest prediction accuracy is selected from multiple reference sample lines adjacent to the current block, and a prediction sample is derived using reference samples in the selected line located in the prediction direction. In this case, intra-frame prediction encoding can be performed by indicating (signaling) the reference sample line to be used to the decoding device. This situation can be referred to as multi-reference line intra-frame prediction or MRL-based intra-frame prediction.

[0109] In addition, the current block is divided into vertical or horizontal sub-partitions and intra prediction is performed based on the same intra prediction mode, but adjacent reference samples can be derived and used in units of sub-partitions. That is, in this case, the intra prediction mode of the current block is also applicable to the sub-partitions, but in some cases, the intra prediction performance can be improved by deriving and using adjacent reference samples in units of sub-partitions. This prediction method can be called intra prediction based on intra sub-partitions (ISP).

[0110] The above-mentioned intra-frame prediction method can be referred to as an intra-frame prediction type to distinguish it from the intra-frame prediction mode. The intra-frame prediction type can be referred to by various terms, such as intra-frame prediction technology or additional intra-frame prediction mode. For example, the intra-frame prediction type (or additional intra-frame prediction mode, etc.) may include at least one of the above-mentioned LIP, PDPC, MRL, and ISP. The general intra-frame prediction method that excludes specific intra-frame prediction types such as LIP, PDPC, MRL, and ISP may be referred to as a normal intra-frame prediction type. When the above-mentioned specific intra-frame prediction type is not applied, the normal intra-frame prediction type may generally be applied, and prediction may be performed based on the above-mentioned intra-frame prediction mode. At the same time, if necessary, post-processing filtering may be performed on the derived prediction samples.

[0111] Specifically, the intra prediction process may include an intra prediction mode / type determination step, an adjacent reference sample derivation step, and a prediction sample derivation step based on the intra prediction mode / type. In addition, if necessary, a post-filtering step may be performed on the derived prediction samples.

[0112] Figure 4 An example of a video / image encoding method based on intra-frame prediction is illustrated.

[0113] refer to Figure 4 , the encoding device performs intra prediction on the current block (S400). The encoding device derives the intra prediction mode / type of the current block, derives the adjacent reference samples of the current block, and generates prediction samples in the current block based on the intra prediction mode / type and the adjacent reference samples. Here, the intra prediction mode / type determination, adjacent reference sample derivation, and prediction sample generation processes can be performed simultaneously, or one process can be performed before the other. The encoding device can determine a mode / type to be applied to the current block from a plurality of intra prediction modes / types. The encoding device can compare the RD costs of the intra prediction modes / types and determine the optimal intra prediction mode / type for the current block.

[0114] At the same time, the encoding device may perform a prediction sample filtering process. Prediction sample filtering may be referred to as post-filtering. Some or all prediction samples may be filtered by the prediction sample filtering process. In some cases, the prediction sample filtering process may be omitted.

[0115] The encoding apparatus generates residual samples of the current block based on the (filtered) prediction samples (S410). The encoding apparatus may compare the prediction samples with the original samples of the current block based on phase and derive the residual samples.

[0116] The encoding device may encode image information including information regarding intra-frame prediction (prediction information) and residual information regarding residual samples (S420). The prediction information may include intra-frame prediction mode information and intra-frame prediction type information. The encoding device may output the encoded image information in the form of a bitstream. The output bitstream may be transmitted to a decoding device via a storage medium or a network.

[0117] The residual information may include a residual coding syntax described later. The encoding device may transform / quantize the residual samples to derive quantized transform coefficients. The residual information may include information about the quantized transform coefficients.

[0118] At the same time, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and a reconstructed block). To this end, the encoding device can derive (modified) residual samples by performing inverse quantization / inverse transformation on the quantized transform coefficients again. The reason for performing inverse quantization / inverse transformation on the residual samples after transforming / quantizing them in this manner is to derive residual samples that are identical to the residual samples derived in the decoding device described above. The encoding device can generate a reconstructed block including reconstructed samples for the current block based on the predicted samples and the (modified) residual samples. A reconstructed picture for the current picture can be generated based on the reconstructed block. As described above, the in-loop filtering process can be further applied to the reconstructed picture.

[0119] Figure 5 An example of a video / image encoding method based on intra-frame prediction is illustrated.

[0120] The decoding device may perform operations corresponding to those performed by the encoding device.

[0121] Prediction information and residual information can be obtained from the bitstream. Residual samples of the current block can be derived based on the residual information. Specifically, transform coefficients can be derived by performing inverse quantization based on quantized transform coefficients derived from the residual information, and residual samples of the current block can be derived by performing inverse transform on the transform coefficients.

[0122] Specifically, the decoding device may derive the intra-frame prediction mode / type of the current block based on the received prediction information (intra-frame prediction mode / type information) (S500). The decoding device may derive neighboring reference samples of the current block (S510). The decoding device generates prediction samples in the current block based on the intra-frame prediction mode / type and the neighboring reference samples (S520). In this case, the decoding device may perform a prediction sample filtering process. Prediction sample filtering may be referred to as post-filtering. Some or all prediction samples may be filtered by the prediction sample filtering process. In some cases, the prediction sample filtering process may be omitted.

[0123] The decoding device generates residual samples for the current block based on the received residual information (S530). The decoding device may generate reconstructed samples for the current block based on the predicted samples and the residual samples, and may derive a reconstructed block including the reconstructed samples (S540). A reconstructed picture of the current picture may be generated based on the reconstructed block. As described above, the in-loop filtering process may be further applied to the reconstructed picture.

[0124] The intra-frame prediction mode information may include, for example, flag information (e.g., intra_luma_mpm_flag) indicating whether MPM (most probable mode) is applied to the current block or whether the residual mode is applied, and when MPM is applied to the current block, the prediction mode information may further include index information (e.g., intra_luma_mpm_idx) indicating one of the intra-frame prediction mode candidates (MPM candidate). The intra-frame prediction mode candidate (MPM candidate) may be composed of an MPM candidate list or an MPM list. In addition, when MPM is not applied to the current block, the intra-frame prediction mode information includes residual mode information (e.g., intra_luma_mpm_remainder) indicating one of the remaining intra-frame prediction modes other than the intra-frame prediction mode candidate (MPM candidate). The decoding device may determine the intra-frame prediction mode of the current block based on the intra-frame prediction mode information.

[0125] In addition, the intra-frame prediction type information can be implemented in various forms. For example, the intra-frame prediction type information may include intra-frame prediction type index information indicating one of the intra-frame prediction types. As another example, the intra-frame prediction type information may include at least one of the following: reference sample line information indicating whether MRL is applied to the current block and, if applied, which reference sample line is used (e.g., intra_luma_ref_idx), ISP flag information indicating whether ISP is applied to the current block (e.g., intra_subpartitions_mode_flag), ISP type information indicating the split type of the sub-partition when ISP is applied (e.g., intra_subpartitions_split_flag), flag information indicating whether PDPC is applied, or flag information indicating whether LIP is applied. In addition, the intra-frame prediction type information may include a MIP flag indicating whether matrix-based intra-frame prediction (MIP) is applied to the current block.

[0126] The intra-frame prediction mode information and / or the intra-frame prediction type information may be encoded / decoded using the coding method described in the present disclosure. For example, the intra-frame prediction mode information and / or the intra-frame prediction type information may be encoded / decoded using entropy coding (e.g., CABAC, CAVLC).

[0127] Figure 6 The intra prediction process is schematically illustrated.

[0128] refer to Figure 6As described above, the intra-frame prediction process may include the steps of determining an intra-frame prediction mode / type, deriving adjacent reference samples, and performing intra-frame prediction (generating prediction samples). The intra-frame prediction process may be performed by the encoding device and decoding device described above. In the present disclosure, a coding device may include an encoding device and / or a decoding device.

[0129] refer to Figure 6 , the coding device determines the intra prediction mode / type S600.

[0130] The encoding device may determine the intra-frame prediction mode / type applied to the current block from the various intra-frame prediction modes / types described above, and may generate prediction-related information. The prediction-related information may include intra-frame prediction mode information indicating the intra-frame prediction mode applied to the current block and / or intra-frame prediction type information indicating the intra-frame prediction type applied to the current block. The decoding device may determine the intra-frame prediction mode / type applied to the current block based on the prediction-related information.

[0131] The intra-frame prediction mode information may include, for example, flag information (e.g., intra_luma_mpm_flag) indicating whether the most probable mode (MPM) is applied to the current block or the residual mode is applied, and when the MPM is applied to the current block, the prediction mode information may further include index information (e.g., intra_luma_mpm_idx) indicating one of the intra-frame prediction mode candidates (MPM candidates). The intra-frame prediction mode candidates (MPM candidates) may be composed of an MPM candidate list or an MPM list. In addition, when the MPM is not applied to the current block, the intra-frame prediction mode information may further include residual mode information (e.g., intra_luma_mpm_remainder) indicating one of the remaining intra-frame prediction modes other than the intra-frame prediction mode candidates (MPM candidates). The decoding device may determine the intra-frame prediction mode of the current block based on the intra-frame prediction mode information.

[0132] In addition, the intra-frame prediction type information can be implemented in various forms. For example, the intra-frame prediction type information may include intra-frame prediction type index information indicating one of the intra-frame prediction types. As another example, the intra-frame prediction type information may include at least one of the following: reference sample line information indicating whether MRL is applied to the current block and, if applied, which reference sample line is used (e.g., intra_luma_ref_idx), ISP flag information indicating whether ISP is applied to the current block (e.g., intra_subpartitions_mode_flag), ISP type information indicating the split type of the sub-partition when ISP is applied (e.g., intra_subpartitions_split_flag), flag information indicating whether PDPC is applied, or flag information indicating whether LIP is applied. In addition, the intra-frame prediction type information may include a MIP flag indicating whether matrix-based intra-frame prediction (MIP) is applied to the current block.

[0133] For example, when applying intra prediction, the intra prediction mode applied to the current block may be determined using the intra prediction modes of neighboring blocks. For example, the coding device may select one of the most probable mode (MPM) candidates in an MPM list derived based on additional candidate modes and / or the intra prediction modes of neighboring blocks of the current block (e.g., left and / or above neighboring blocks), or select one of the remaining intra prediction modes not included in the MPM candidates (and planar mode) based on MPM residual information (residual intra prediction mode information). The MPM list can be configured to include or exclude planar mode as a candidate. For example, if the MPM list includes planar mode as a candidate, the MPM list may have six candidates, and if it does not, the MPM list may have five candidates. If the MPM list does not include planar mode as a candidate, a non-planar flag (e.g., intra_luma_not_planar_flag) indicating whether the intra prediction mode of the current block is not planar mode may be signaled. For example, the MPM flag may be signaled first, and when the value of the MPM flag is 1, the MPM index and the non-planar flag may be signaled. Furthermore, when the value of the non-planar flag is 1, the MPM index may be signaled. Here, the fact that the MPM list is configured not to include the planar mode as a candidate is that the planar mode is always considered to be an MPM, rather than being considered not to be an MPM. Therefore, the flag (non-planar flag) is signaled first to check whether it is the planar mode.

[0134] For example, it can be indicated based on an MPM flag (e.g., intra_luma_mpm_flag) whether the intra prediction mode applied to the current block is among the MPM candidates (and planar mode) or among the remaining modes. An MPM flag with a value of 1 can indicate that the intra prediction mode of the current block is within the MPM candidates (and planar mode), while an MPM flag with a value of 0 can indicate that the intra prediction mode of the current block is not within the MPM candidates (and planar mode). A non-planar flag with a value of 0 (e.g., intra_luma_not_planar_flag) can indicate that the intra prediction mode of the current block is planar mode, while a non-planar flag with a value of 1 can indicate that the intra prediction mode of the current block is not planar mode. The MPM index can be signaled in the form of the mpm_idx or intra_luma_mpm_idx syntax element, and the remaining intra prediction mode information can be signaled in the form of the rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, the remaining intra prediction mode information may indicate one of the remaining intra prediction modes that is not included in the MPM candidates (and planar mode) among all intra prediction modes by indexing in order of the prediction mode number. The intra prediction mode may be an intra prediction mode of a luminance component (sample). Hereinafter, the intra prediction mode information may include at least one of an MPM flag (e.g., intra_luma_mpm_flag), a non-planar flag (e.g., intra_luma_not_planar_flag), an MPM index (e.g., mpm_idx or intra_luma_mpm_idx), or the remaining intra prediction mode information (rem_intra_luma_luma_mpm_mode or intra_luma_mpminder). In the present disclosure, the MPM list may be referred to by a variety of terms, such as an MPM candidate list and candModeList.

[0135] When a MIP is applied to a current block, a separate mpm flag (eg, intra_mip_mpm_flag), an mpm index (eg, intra_mip_mpm_idx), and remaining intra prediction mode information (eg, intra_mip_mpm_remainder) for the MIP may be signaled, and a non-planar flag may not be signaled.

[0136] In other words, generally speaking, when performing block segmentation of an image, the current block to be coded and the neighboring blocks have similar image characteristics. Therefore, there is a high probability that the current block and the neighboring blocks have the same or similar intra-frame prediction mode. Therefore, the encoder can use the intra-frame prediction mode of the neighboring block to encode the intra-frame prediction mode of the current block.

[0137] The coding device can construct a most probable mode (MPM) list for the current block. The MPM list can be referred to as an MPM candidate list. Here, MPM can refer to a mode used to improve coding efficiency by considering the similarity between the current block and neighboring blocks during intra-frame prediction mode coding. As described above, the MPM list can be constructed to include the planar mode or to exclude the planar mode. For example, when the MPM list includes the planar mode, the number of candidates in the MPM list can be 6. When the MPM list does not include the planar mode, the number of candidates in the MPM list can be 5.

[0138] The encoding device can perform prediction based on various intra-frame prediction modes, and can determine the optimal intra-frame prediction mode based on rate-distortion optimization (RDO) thereof. In this case, the encoding device can determine the optimal intra-frame prediction mode by using only the MPM candidates and planar mode configured in the MPM list, or by further using the remaining intra-frame prediction modes and the MPM candidates and planar mode configured in the MPM list. Specifically, for example, if the intra-frame prediction type of the current block is a specific type other than the normal intra-frame prediction type (such as LIP, MRL or ISP), the encoding device can determine the optimal intra-frame prediction mode by considering only the MPM candidates and planar mode as intra-frame prediction mode candidates for the current block. That is, in this case, the intra-frame prediction mode of the current block can be determined only from the MPM candidates and planar mode, and in this case, encoding / signaling of the mpm flag can be not performed. In this case, the decoding device can infer that the mpm flag is 1 without separately signaling the mpm flag.

[0139] Meanwhile, typically, when the intra prediction mode of the current block is not planar mode but is one of the MPM candidates in the MPM list, the encoding device generates an mpm index (mpm idx) indicating one of the MPM candidates. When the intra prediction mode of the current block is not included in the MPM list, the encoding device generates MPM residual information (remaining intra prediction mode information) indicating the same mode as the intra prediction mode of the current block among the remaining intra prediction modes not included in the MPM list (and planar mode). The MPM residual information may include, for example, the intra_luma_mpm_remainder syntax element.

[0140] The decoding device obtains intra-frame prediction mode information from the bitstream. As described above, the intra-frame prediction mode information may include at least one of an MPM flag, a non-planar flag, an MPM index, and MPM residual information (residual intra-frame prediction mode information). The decoding device may construct an MPM list. The construction of the MPM list is similar to the MPM list constructed by the encoding device. That is, the MPM list may include intra-frame prediction modes of neighboring blocks, or may further include a specific intra-frame prediction mode according to a predetermined method.

[0141] The decoding device can determine the intra prediction mode of the current block based on the MPM list and the intra prediction mode information. For example, when the value of the MPM flag is 1, the decoding device can derive the planar mode as the intra prediction mode of the current block (based on the non-planar flag), or derive the candidate indicated by the MPM index among the MPM candidates in the MPM list as the intra prediction mode of the current block. Here, the MPM candidate may refer to only the candidates included in the MPM list, or may include not only the candidates included in the MPM list but also the planar mode applicable when the value of the MPM flag is 1.

[0142] For another example, when the value of the MPM flag is 0, the decoding device may derive the intra prediction mode indicated by the remaining intra prediction mode information (which may be referred to as mpm residual information) among the remaining intra prediction modes not included in the MPM list and the planar mode as the intra prediction mode of the current block. At the same time, as another example, when the intra prediction type of the current block is a specific type (such as LIP, MRL, or ISP, etc.), the decoding device may derive the candidate indicated by the MPM flag in the planar mode or the MPM list as the intra prediction mode of the current block without parsing / decoding / checking the MPM flag.

[0143] The coding device derives neighboring reference samples for the current block (S610). When intra prediction is applied to the current block, neighboring reference samples to be used for intra prediction of the current block may be derived. The neighboring reference samples of the current block may include samples adjacent to the left boundary of the current block (nWxnH), a total of 2xnH samples adjacent to the lower left of the current block, samples adjacent to the upper boundary of the current block, a total of 2xnW samples adjacent to the upper right, and samples adjacent to the upper left of the current block. Alternatively, the neighboring reference samples of the current block may include multiple columns of upper neighboring samples and multiple rows of left neighboring samples. Furthermore, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block (nWxnH), a total of nW samples adjacent to the lower boundary of the current block, and samples adjacent to the lower right of the current block.

[0144] On the other hand, when MRL is applied (ie, when the value of the MRL index is greater than 0), the neighboring reference samples may be located in lines 1 to 2 instead of line 0 adjacent to the current block on the left / upper side, and in this case, the number of neighboring reference samples can be further increased. At the same time, when ISP is applied, the neighboring reference samples can be derived in units of sub-partitions.

[0145] The coding device derives prediction samples by performing intra prediction on the current block (S620). The coding device may derive the prediction samples based on the intra prediction mode / type and neighboring samples. The coding device may derive reference samples based on the intra prediction mode of the current block among neighboring reference samples of the current block, and may derive the prediction samples of the current block based on the reference samples.

[0146] When inter-frame prediction is applied, the predictor of the encoding / decoding device can derive prediction samples by performing inter-frame prediction on a block-by-block basis. When performing prediction on the current block, inter-frame prediction can be applied. That is, the predictor of the encoding / decoding device (more specifically, the inter-frame predictor) can derive prediction samples by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction may refer to a prediction derived using a method that depends on data elements (e.g., sample values ​​or motion information) of one or more pictures other than the current picture. When inter-frame prediction is applied to the current block, the prediction block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector in a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information for the current block can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When inter-frame prediction is used, neighboring blocks may include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture comprising the reference block and the reference picture comprising the temporally neighboring block may be the same or different. Temporally neighboring blocks may be referred to as collocated reference blocks, collocated CUs (ColCUs), etc., and the reference picture comprising the temporally neighboring blocks may be referred to as collocated pictures (ColPics). For example, a motion information candidate list may be configured based on the neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be the same as that of the selected neighboring block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and the motion vector difference may be signaled. In this case, the motion vector of the current block can be derived by summing the motion vector predictor and the motion vector difference.

[0147] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), the motion information may further include L0 motion information and / or L1 motion information. The L0-direction motion vector may be referred to as the L0 motion vector or MVL0, and the L1-direction motion vector may be referred to as the L1 motion vector or MVL1. Prediction based on the L0 motion vector may be referred to as L0 prediction, prediction based on the L1 motion vector may be referred to as L1 prediction, and prediction based on both the L0 and L1 motion vectors may be referred to as bi-prediction. Here, the L0 motion vector may indicate the motion vector associated with reference picture list L0, and the L1 motion vector may indicate the motion vector associated with reference picture list L1. Reference picture list L0 may include pictures preceding the current picture in output order, and reference picture list L1 may include pictures following the current picture in output order as reference pictures. The preceding picture may be referred to as a forward (reference) picture, and the subsequent picture may be referred to as a backward (reference) picture. Reference picture list L0 may further include pictures following the current picture in output order as reference pictures. In this case, the previous picture in reference picture list L0 may be indexed first, followed by the subsequent picture. Reference picture list L1 may further include pictures preceding the current picture in output order as reference pictures. In this case, the subsequent picture in reference picture list L1 may be indexed first, followed by the previous picture. Here, the output order may correspond to the picture order count (POC) order.

[0148] The video / image encoding process based on inter-frame prediction may schematically include, for example, the following contents.

[0149] Figure 7 An example of a video / image encoding method based on inter-frame prediction is illustrated.

[0150] The encoding device performs inter-frame prediction on the current block (S700). The encoding device may derive the inter-frame prediction mode and motion information for the current block, and generate prediction samples for the current block. The inter-frame prediction mode determination process, motion information derivation process, and prediction sample generation process may be performed simultaneously, and any one process may be performed earlier than the other. For example, the inter-frame prediction unit of the encoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit may determine the prediction mode for the current block, the motion information derivation unit may derive the motion information for the current block, and the prediction sample derivation unit may derive the prediction samples for the current block. For example, the inter-frame prediction unit of the encoding device may use motion estimation to search for blocks similar to the current block in a predetermined area (search area) of a reference picture and derive a reference block whose difference with the current block is minimal or equal to or less than a predetermined criterion. Based on this, a reference picture index indicating the reference picture in which the reference block is located may be derived, and a motion vector may be derived based on the positional difference between the reference block and the current block. The encoding device may determine which prediction mode to apply to the current block from among various prediction modes. The encoding apparatus may compare RD costs of various prediction modes and determine an optimal prediction mode for the current block.

[0151] For example, when skip mode or merge mode is applied to the current block, the encoding device may configure a merge candidate list (described below) and derive a reference block whose difference with the current block is the smallest or equal to or less than a predetermined criterion, from among the reference blocks indicated by the merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the decoding device. The motion information of the current block may be derived using the motion information of the selected merge candidate.

[0152] As another example, when (A) MVP mode is applied to the current block, the encoding device may configure an (A) MVP candidate list (described below) and use the motion vector of a selected motion vector predictor (MVP) candidate from among the motion vector predictor (MVP) candidates included in the (A) MVP candidate list as the MVP for the current block. In this case, for example, the motion vector of a reference block derived through motion estimation may be used as the motion vector for the current block, and the MVP candidate with the smallest motion vector difference from the motion vector of the current block may be selected. A motion vector difference (MVD) may be derived by subtracting the MVP from the motion vector of the current block. In this case, information regarding the MVD may be signaled to the decoding device. Furthermore, when (A) MVP mode is applied, the value of a reference picture index may be configured as reference picture index information and separately signaled to the decoding device.

[0153] The encoding apparatus may derive residual samples based on the prediction samples ( S710 ). The encoding apparatus may derive residual samples by comparing original samples and prediction samples of the current block.

[0154] The encoding device encodes image information including prediction information and residual information (S720). The encoding device can output the encoded image information in the form of a bitstream. The prediction information may include information about prediction mode information (e.g., a skip flag, a merge flag, or a mode index, etc.) and information about motion information as information related to the prediction process. The information about the motion information may include candidate selection information (e.g., a merge index, an mvp flag, or an mvp index), which is information for deriving a motion vector. In addition, the information about the motion information may include information about MVD and / or reference picture index information. In addition, the information about the motion information may include information indicating whether L0 prediction, L1 prediction, or dual prediction is applied. The residual information is information about the residual samples. The residual information may include information about the quantized transform coefficients used for the residual samples.

[0155] The output bitstream may be stored in a (digital) storage medium and transmitted to the decoding device, or transmitted to the decoding device via a network.

[0156] At the same time, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and residual samples. This is to derive a prediction result identical to the prediction result performed by the decoding device, thereby improving coding efficiency. Therefore, the encoding device can store the reconstructed picture (or reconstructed samples or reconstructed blocks) in a memory and use the reconstructed picture as a reference picture. As described above, the in-loop filtering process can further be applied to the reconstructed picture.

[0157] The video / image decoding process based on inter-frame prediction may schematically include, for example, the following contents.

[0158] Figure 8 An example of a video / image decoding method based on inter-frame prediction is illustrated.

[0159] refer to Figure 8 , the decoding device may perform an operation corresponding to the operation performed by the encoding device. The decoding device may perform prediction on the current block based on the received prediction information and derive a prediction sample.

[0160] Specifically, the decoding apparatus may determine a prediction mode of the current block based on the received prediction information (S800).The decoding apparatus may determine which inter prediction mode to apply to the current block based on prediction mode information in the prediction information.

[0161] For example, whether to apply merge mode or (A) MVP mode to the current block may be determined based on the merge flag. Alternatively, one of various inter-frame prediction mode candidates may be selected based on the mode index. Inter-frame prediction mode candidates may include skip mode, merge mode, and / or (A) MVP mode, or may include various inter-frame prediction modes described below.

[0162] The decoding apparatus derives motion information for the current block based on the determined inter-frame prediction mode (S810). For example, when skip mode or merge mode is applied to the current block, the decoding apparatus may configure a merge candidate list (described below) and select a merge candidate from among the merge candidates included in the merge candidate list. The selection may be performed based on selection information (merge index). The motion information of the selected merge candidate may be used to derive motion information for the current block. The motion information of the selected merge candidate may be used as the motion information for the current block.

[0163] As another example, when the (A) MVP mode is applied to the current block, the decoding device may configure the (A) MVP candidate list (described below) and use the motion vector of a motion vector predictor (MVP) candidate selected from among the motion vector predictor (MVP) candidates included in the (A) MVP candidate list as the MVP for the current block. The selection may be performed based on selection information (MVP flag or MVP index). In this case, the MVD of the current block may be derived based on information regarding the MVD, and the motion vector of the current block may be derived based on the MVP and MVD of the current block. Furthermore, the reference picture index of the current block may be derived based on reference picture index information. The picture indicated by the reference picture index in the reference picture list for the current block may be derived as the reference picture referenced for inter-frame prediction of the current block.

[0164] At the same time, as described below, the motion information of the current block can be derived without the candidate list configuration, and in this case, the motion information of the current block can be derived according to the process disclosed in the prediction mode. In this case, the candidate list configuration can be omitted.

[0165] The decoding apparatus may generate prediction samples for the current block based on the motion information of the current block (S820). In this case, a reference picture may be derived based on a reference picture index of the current block, and prediction samples of the current block may be derived using samples of the reference block indicated by the motion vector of the current block on the reference picture. In this case, in some cases, a prediction sample filtering process may be further performed for all or some of the prediction samples of the current block.

[0166] For example, the inter-frame prediction unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit and a prediction sample derivation unit, and the prediction mode determination unit may determine the prediction mode of the current block based on the received prediction mode information, the motion information derivation unit may derive the motion information (motion vector and / or reference picture index) of the current block based on the information about the received motion information, and the prediction sample derivation unit may derive the prediction sample of the current block.

[0167] The decoding apparatus generates residual samples for the current block based on the received residual information (S830). The decoding apparatus may generate reconstructed samples for the current block based on the predicted samples and the residual samples, and generate a reconstructed picture based on the generated reconstructed samples (S840). Thereafter, as described above, the in-loop filtering process may be further applied to the reconstructed picture.

[0168] Figure 9 The inter-frame prediction process is schematically illustrated.

[0169] refer to Figure 9 As described above, the inter-frame prediction process may include an inter-frame prediction mode determination step, a motion information derivation step based on the determined prediction mode, and a prediction process (prediction sample generation step) based on the derived motion information. The inter-frame prediction process may be performed by the encoding device and decoding device described above. In this context, a coding device may include an encoding device and / or a decoding device.

[0170] refer to Figure 9 The coding apparatus determines an inter-frame prediction mode for the current block (S900). Various inter-frame prediction modes can be used to predict the current block in the picture. For example, various modes can be used, such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, merge with MVD (MMVD) mode, and historical motion vector prediction (HMVP) mode. Decoder-side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bi-prediction with CU-level weights (BCW), and bidirectional optical flow (BDOF) can further be used as additional modes. Affine mode may also be referred to as affine motion prediction mode. MVP mode may also be referred to as advanced motion vector prediction (AMVP) mode. In this document, some modes and / or motion information candidates derived from some modes may also be included in one of the motion information-related candidates in other modes. For example, an HMVP candidate may be added to the merge candidate for merge / skip mode, or to the MVP candidate for MVP mode. If the HMVP candidate is used as a motion information candidate for merge mode or skip mode, the HMVP candidate may be referred to as an HMVP merge candidate.

[0171] Prediction mode information indicating the inter-frame prediction mode of the current block can be signaled from the encoding device to the decoding device. In this case, the prediction mode information can be included in the bitstream and received by the decoding device. The prediction mode information may include index information indicating one of a plurality of candidate modes. Alternatively, the inter-frame prediction mode may be indicated by hierarchical signaling of flag information. In this case, the prediction mode information may include one or more flags. For example, whether the skip mode is applied may be indicated by signaling a skip flag, and when the skip mode is not applied, whether the merge mode is applied may be indicated by signaling a merge flag, and when the merge mode is not applied, the MVP mode may be indicated or a flag for additional distinction may be further signaled. The affine mode may be signaled as an independent mode, or as a subordinate mode with respect to the merge mode or the MVP mode. For example, the affine mode may include an affine merge mode and an affine MVP mode.

[0172] The coding apparatus derives motion information of a current block (S910). The motion information may be derived based on an inter prediction mode.

[0173] The coding device can use the motion information of the current block to perform inter-frame prediction. The encoding device can derive optimal motion information for the current block through a motion estimation process. For example, the encoding device can use the original block in the original picture of the current block to search for a similar reference block with high correlation in units of fractional pixels within a predetermined search range in the reference picture, and derive motion information from the searched reference block. Block similarity can be derived based on the difference in phase-based sample values. For example, block similarity can be calculated based on the sum of absolute differences (SAD) between the current block (or a template of the current block) and a reference block (or a template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD in the search area. The derived motion information can be signaled to the decoding device according to various methods based on the inter-frame prediction mode.

[0174] The coding apparatus performs inter prediction based on the motion information of the current block (S920). The coding apparatus may derive (one or more) prediction samples of the current block based on the motion information. The current block including the prediction samples may be referred to as a prediction block.

[0175] At the same time, as described above, the encoding device can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). In addition, the decoding device can decode the information in the bitstream based on a coding method such as exponential Golomb, CAVLC, or CABAC, and output the values ​​of syntax elements required for image reconstruction and the quantized values ​​of the transform coefficients related to the residual.

[0176] For example, the above-mentioned compilation method may be performed as follows.

[0177] Figure 10 Context-Adaptive Binary Arithmetic Coding (CABAC) for encoding syntax elements is exemplified. For example, in the CABAC coding process, when the input signal is a syntax element rather than a binary value, the encoding device can convert the input signal into a binary value by binarizing the value of the input signal. In addition, when the input signal is already a binary value (i.e., when the value of the input signal is a binary value), binarization may not be performed and may be bypassed. Here, each binary number 0 or 1 that constitutes a binary value may be referred to as a bin. For example, if the binary string after binarization is 110, each of 1, 1, and 0 is referred to as a bin. The bin used for a syntax element may indicate the value of the syntax element.

[0178] Afterwards, the binarized bins of the syntax elements may be input to a regular coding engine or a bypass coding engine. The regular coding engine of the encoding device may assign a context model reflecting the probability value to the corresponding bin, and may encode the corresponding bin based on the assigned context model. After encoding each bin, the regular coding engine of the encoding device may update the context model for the corresponding bin. The encoded bins described above may be referred to as context-coded bins.

[0179] Meanwhile, when the binarized bins of syntax elements are input to the bypass coding engine, they can be coded as follows. For example, the bypass coding engine of the encoding device omits the process of estimating the probability of the input bins and the process of updating the probability model applied to the bins after coding. When bypass coding is applied, the encoding device can encode the input bins by applying a uniform probability distribution instead of assigning a context model, thereby increasing the coding speed. The coded bins described above can be referred to as bypass bins.

[0180] Entropy decoding may refer to a process of performing the same process as the above-described entropy encoding in reverse order.

[0181] For example, when decoding a syntax element based on a context model, the decoding device may receive a bin corresponding to the syntax element through a bitstream, determine the context model using the syntax element and decoded information of a decoding target block or a neighboring block, or information about a previously decoded symbol / bin, and derive the value of the syntax element by predicting the probability of occurrence of the received bin based on the determined context model and performing arithmetic decoding on the bin. The determined context model may then be used to update the context model for the next decoded bin.

[0182] Furthermore, for example, when bypass decoding a syntax element, the decoding apparatus may receive a bin corresponding to the syntax element through a bitstream and decode the input bin by applying a uniform probability distribution. In this case, the decoding apparatus may omit the process of deriving a context model for the syntax element and the process of updating the context model applied to the bin after decoding.

[0183] As described above, the residual samples can be derived as quantized transform coefficients through the transformation and quantization process. The quantized transform coefficients can also be referred to as transform coefficients. In this case, the transform coefficients in the block can be signaled in the form of residual information. The residual information may include residual coding syntax. That is, the encoding device can configure the residual coding syntax using the residual information, encode the residual coding syntax, and output it in the form of a bitstream, and the decoding device can decode the residual coding syntax from the bitstream and derive the residual (quantized) transform coefficients. The residual coding syntax may include syntax elements indicating whether the transform is applied to the corresponding block, the position of the last significant transform coefficient in the block, whether there is a significant transform coefficient in the sub-block, the size / sign of the significant transform coefficient, etc., as described later.

[0184] For example, transform coefficients (i.e., residual information) may be encoded and / or decoded (quantized) based on syntax elements such as transform_skip_flag, last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix, coded_sub_block_flag, sig_coeff_flag, par_level_flag, abs_level_gt1_flag, abs_level_gt3_flag, abs_remaind, coeff_sign_flag, dec_abs_level, and mts_idx. Syntax elements related to residual data encoding / decoding may be represented as shown in the following table.

[0185] [Table 1]

[0186]

[0187]

[0188]

[0189]

[0190] The transform_skip_flag flag indicates whether the transform is skipped for the associated block. The transform_skip_flag can be a syntax element that represents a transform skip flag. The associated block can be a coding block (CB) or a transform block (TB). With respect to the transform (and quantization) and residual coding processes, the terms "CB" and "TB" can be used interchangeably. For example, as described above, residual samples can be derived for the CB, and transform coefficients can be derived (quantized) through the transform and quantization of the residual samples. Furthermore, the residual coding process can generate and signal information (e.g., syntax elements) that effectively indicate the position, magnitude, sign, etc. of the (quantized) transform coefficients. Quantized transform coefficients can be simply referred to as transform coefficients. Generally, when the CB is not larger than the maximum TB, the size of the CB can be the same as the size of the TB. In this case, the target block to be transformed (and quantized) and residual coded can be referred to as either the CB or the TB. Meanwhile, when the CB is larger than the maximum TB, the target block to be transformed (and quantized) and residual coded can be referred to as the TB. Hereinafter, signaling of syntax elements related to residual coding in units of transform blocks (TBs) will be described, but this is an example and TBs may be used interchangeably with coding blocks (CBs) as described above.

[0191] Meanwhile, the syntax elements signaled after the transform skip flag is signaled may be the same as the syntax elements disclosed in the following Table 2, and a detailed description about the syntax elements is described below.

[0192] [Table 2]

[0193]

[0194]

[0195]

[0196] [Table 3]

[0197]

[0198]

[0199]

[0200]

[0201]

[0202]

[0203] [Table 4]

[0204]

[0205]

[0206]

[0207] According to this embodiment, as shown in Table 2, residual coding can be divided according to the value of the transform_skip_flag syntax element of the transform skip flag. That is, different syntax elements can be used for residual coding based on the value of the transform skip flag (based on whether the transform is skipped). Residual coding used when transform skipping is not applied (i.e., when the transform is applied) can be referred to as regular residual coding (RRC), while residual coding used when transform skipping is applied (i.e., when the transform is not applied) can be referred to as transform-skipped residual coding (TSRC). Furthermore, regular residual coding can be referred to as normal residual coding. Furthermore, regular residual coding can be referred to as a regular residual coding syntax structure, while transform-skipped residual coding can be referred to as a transform-skipped residual coding syntax structure. Table 3 above shows the syntax elements for residual coding when the value of transform_skip_flag is 0 (i.e., when the transform is applied), while Table 4 above shows the syntax elements for residual coding when the value of transform_skip_flag is 1 (i.e., when the transform is not applied).

[0208] Specifically, for example, a transform skip flag indicating whether the transform of a transform block is skipped may be parsed, and it may be determined whether the transform skip flag is 1. If the value of the transform skip flag is 0, as shown in Table 3, the syntax elements last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix, sb_coded_flag, sig_coeff_flag, abs_level_gtx_flag, par_level_flag, abs_remainder, coeff_sign_flag, and / or dec_abs_level for the residual coefficient of the transform block may be parsed, and the residual coefficient may be derived based on the syntax elements. In this case, the syntax elements may be parsed sequentially, and the parsing order may be changed. In addition, abs_level_gtx_flag may indicate abs_level_gt1_flag and / or abs_level_gt3_flag. For example, abs_level_gtx_flag[n][0] may be an example of a first transform coefficient level flag (abs_level_gt1_flag), and abs_level_gtx_flag[n][1] may be an example of a second transform coefficient level flag (abs_level_gt3_flag).

[0209] Referring to Table 3 above, last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix, sb_coded_flag, sig_coeff_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag, abs_remainder, coeff_sign_flag, and / or dec_abs_level may be encoded / decoded. Meanwhile, sb_coded_flag may be expressed as coded_sub_block_flag.

[0210] In an embodiment, the encoding device may encode the (x, y) position information of the last non-zero transform coefficient in the transform block based on the syntax elements last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix. More specifically, last_sig_coeff_x_prefix represents the prefix of the column position of the last significant coefficient in the scan order within the transform block, last_sig_coeff_y_prefix represents the prefix of the row position of the last significant coefficient in the scan order within the transform block, last_sig_coeff_x_suffix represents the suffix of the column position of the last significant coefficient in the scan order within the transform block, and last_sig_coeff_y_suffix represents the suffix of the row position of the last significant coefficient in the scan order within the transform block. Here, the significant coefficient may represent a non-zero coefficient. In addition, the scan order may be a right diagonal scan order. Alternatively, the scanning order may be a horizontal scanning order or a vertical scanning order.The scanning order may be determined based on whether intra / inter prediction is applied to a target block (CB or CB including TB) and / or a specific intra / inter prediction mode.

[0211] Thereafter, the encoding apparatus may divide the transform block into 4×4 sub-blocks, and then use a 1-bit syntax element coded_sub_block_flag to indicate, for each 4×4 sub-block, whether there is a non-zero coefficient in the current sub-block.

[0212] If the value of coded_sub_block_flag is 0, there is no more information to be transmitted, so the encoding device can terminate the encoding process of the current subblock. Conversely, if the value of coded_sub_block_flag is 1, the encoding device can continuously perform the encoding process on sig_coeff_flag. Since the subblock including the last non-zero coefficient does not need to encode coded_sub_block_flag, and the subblock including DC information of the transform block has a high probability of including a non-zero coefficient, coded_sub_block_flag may not be coded and its value may be assumed to be 1.

[0213] If the value of coded_sub_block_flag is 1, and it is therefore determined that there is a non-zero coefficient in the current sub-block, the encoding device may encode a sig_coeff_flag having a binary value according to the reverse scanning order. The encoding device may encode a 1-bit syntax element sig_coeff_flag for each transform coefficient according to the scanning order. If the value of the transform coefficient at the current scanning position is not 0, the value of sig_coeff_flag may be 1. Here, in the case of a sub-block including the last non-zero coefficient, it is not necessary to encode sig_coeff_flag for the last non-zero coefficient, and thus the coding process of the sub-block may be omitted. Level information coding may be performed only when sig_coeff_flag is 1, and four syntax elements may be used in the level information coding process. More specifically, each sig_coeff_flag[xC][yC] may indicate whether the level (value) of the corresponding transform coefficient at each transform coefficient position (xC, yC) in the current TB is non-zero. In an embodiment, sig_coeff_flag may correspond to an example of a syntax element of a significant coefficient flag indicating whether a quantized transform coefficient is a non-zero significant coefficient.

[0214] The level value remaining after encoding sig_coeff_flag can be derived as shown in the following equation. That is, the syntax element remAbsLevel indicating the level value to be encoded can be derived from the following equation.

[0215] [Equation 1]

[0216]

[0217] Here, coeff represents the actual transform coefficient value.

[0218] In addition, abs_level_gt1_flag may indicate whether the remAbsLevel of the corresponding scanning position (n) is greater than 1. For example, when the value of abs_level_gt1_flag is 0, the absolute value of the transform coefficient of the corresponding position may be 1. In addition, when the value of abs_level_gt1_flag is 1, remAbsLevel indicating a level value to be encoded later may be updated as shown in the following equation.

[0219] [Equation 2]

[0220]

[0221] In addition, a least significant coefficient (LSB) value of remAbsLevel described in the above Equation 2 may be encoded through par_level_flag as in the following Equation 3.

[0222] [Equation 3]

[0223]

[0224] Here, par_level_flag[n] may indicate the parity of the transform coefficient level (value) at the scanning position n.

[0225] The transform coefficient level value remAbsLevel to be encoded after performing par_level_flag encoding may be updated as shown in the following equation as follows.

[0226] [Equation 4]

[0227]

[0228] abs_level_gt3_flag may indicate whether the remAbsLevel of the corresponding scanning position (n) is greater than 3. Encoding of abs_remainder may be performed only if rem_abs_gt3_flag is equal to 1. The relationship between the actual transform coefficient value coeff and each syntax element may be as shown below in the following equation.

[0229] [Equation 5]

[0230]

[0231] In addition, the following table indicates examples related to the above-mentioned Equation 5.

[0232] [Table 5]

[0233]

[0234] Here, |coeff| indicates a transform coefficient level (value), and may also be indicated as AbsLevel for the transform coefficient. In addition, the sign of each coefficient may be encoded by using coeff_sign_flag which is a 1-bit symbol.

[0235] Furthermore, if the transform skip flag has a value of 1, as shown in Table 4, the syntax elements sb_coded_flag, sig_coeff_flag, coeff_sign_flag, abs_level_gtx_flag, par_level_flag, and / or abs_remainder for the residual coefficients of the transform block may be parsed, and the residual coefficients may be derived based on the syntax elements. In this case, the syntax elements may be parsed sequentially, and the parsing order may be changed. Furthermore, abs_level_gtx_flag may represent abs_level_gt1_flag, abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and / or abs_level_gt9_flag. For example, abs_level_gtx_flag[n][j] may be a flag indicating whether the absolute value or level (value) of the transform coefficient at scan position n is greater than (j<<1)+1. The condition (j<<1)+1 may optionally be replaced with a specific threshold value, such as a first threshold value, a second threshold value, or the like.

[0236] At the same time, CABAC provides high performance, but disadvantageously has poor throughput performance. This is caused by CABAC's conventional compilation engine. Conventional encoding (i.e., compilation by CABAC's conventional compilation engine) shows high data dependency because it uses probability states and ranges updated by compilation of previous bins, and reading probability intervals and determining the current state takes a lot of time. CABAC's throughput problem can be solved by limiting the number of bins for context coding. For example, as shown in Table 1 or Table 3 above, the sum of bins used to represent sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag can be limited to the number of bins that depends on the size of the corresponding block. Furthermore, for example, as shown in Table 4 above, the sum of bins that can be used to represent sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, abs_level_gt9_flag may be limited to the number of bins depending on the size of the corresponding block. For example, if the corresponding block is a 4×4 block, the sum of bins for sig_coeff_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag, or sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, abs_level_gt9_flag may be limited to 32 (or 28 in example), and if the corresponding block is a 2×2 block, the sum of bins for sig_coeff_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag may be limited to 8 (or 7 in example). The limited number of bins may be represented by remBinsPass1 or RemCcbs. Alternatively, for example, for higher CABAC throughput, the number of context-coded bins can be limited for a block (CB or TB) including a coding target CG. In other words, the number of context-coded bins can be limited per block (CB or TB).For example, when the size of the current block is 16×16, the number of bins used for context coding of the current block can be limited to 1.75 times the number of pixels of the current block, ie, 448, regardless of the current CG.

[0237] In this case, if a limited number of bins of all context coding are used when coding the context element, the encoding device can binarize the remaining coefficients by the method of binarizing the coefficients as described below, instead of using context coding, and bypass coding can be performed. In other words, for example, if the number of bins of context coding for 4×4 CG coding is 32 (or 28 in the example), or if the number of bins of context coding for 2×2 CG coding is 8 (or 7 in the example), sig_coeff_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag coded with the bins of context coding can no longer be coded and can be directly coded into dec_abs_level. Alternatively, for example, when the number of context-coded bins for 4×4 block coding is 1.75 times the number of pixels of the entire block, that is, when limited to 28, sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag coded as context-coded bins may no longer be coded and may be directly coded as dec_abs_level, as shown in Table 6 below.

[0238] [Table 6]

[0239]

[0240] The value |coeff| can be derived based on dec_abs_level. In this case, the transform coefficient value (ie, |coeff|) can be derived as shown in the following equation.

[0241] [Equation 6]

[0242]

[0243] In addition, coeff_sign_flag may indicate the sign of the transform coefficient level at the corresponding scanning position n. That is, coeff_sign_flag may indicate the sign of the transform coefficient at the corresponding scanning position n.

[0244] Figure 11 An example of transform coefficients in a 4x4 block is shown.

[0245] Figure 11The 4×4 block represents an example of quantized coefficients. Figure 11 The block can be a 4×4 transform block, or a 4×4 sub-block of an 8×8, 16×16, 32×32, or 64×64 transform block. Figure 11 A 4×4 block can represent a luma block or a chroma block.

[0246] At the same time, as described above, when the input signal is not a binary value but a syntax element, the encoding device can convert the input signal into a binary value by binarizing the value of the input signal. In addition, the decoding device can decode the syntax element to derive the binary value (e.g., binarization bin) of the syntax element, and can debinarize the binarized value to derive the value of the syntax element. The binarization process can be performed as a truncated Rice (TR) binarization process, a k-order exponential Golomb (EGk) binarization process, a limited k-order exponential Golomb (limited EGk), a fixed length (FL) binarization process, etc. In addition, the debinarization process can refer to a process performed based on the TR binarization process, the EGk binarization process, or the FL binarization process to derive the value of the syntax element.

[0247] For example, the TR binarization process can be performed as follows.

[0248] The input of the TR binarization process may be the cMax and cRiceParam for the syntax element and a request for TR binarization. Additionally, the output of the TR binarization process may be the TR binarization for symbolVal, which is the value corresponding to the bin string.

[0249] Specifically, for example, if a suffix bin string exists for a syntax element, the TR bin string for the syntax element may be a concatenation of the prefix bin string and the suffix bin string. If a suffix bin string does not exist, the TR bin string for the syntax element may be the prefix bin string. For example, the prefix bin string may be derived as follows.

[0250] The prefix value of symbolVal for a syntax element may be derived as shown in the following equation.

[0251] [Equation 7]

[0252]

[0253] Here, prefixVal may represent a prefix value of symbolVal. A prefix of a TR bin string of a syntax element (ie, a prefix bin string) may be derived as described below.

[0254] For example, if prefixVal is less than cMax>>cRiceParam, the prefix bin string may be a bit string of length prefixVal+1, indexed by binIdx. That is, if prefixVal is less than cMax>>cRiceParam, the prefix bin string may be a bit string of length prefixVal+1, indicated by binIdx. A bin with a binIdx less than prefixVal may be equal to 1. Additionally, a bin with the same binIdx as prefixVal may be equal to 0.

[0255] For example, the bin string derived by unary binarization of prefixVal may be as shown in the following table.

[0256] [Table 7]

[0257]

[0258] Meanwhile, if prefixVal is not less than cMax>>cRiceParam, the prefix bin string may be a bit string having a length of cMax>>cRiceParam and all bits being 1.

[0259] In addition, if cMax is greater than symbolVal and if cRiceParam is greater than 0, a bin suffix bin string of the TR bin string may exist. For example, the suffix bin string may be derived as follows.

[0260] The suffix value of symbolVal for a syntax element can be derived as shown in the following equation.

[0261] [Equation 8]

[0262]

[0263] Here, suffixVal may represent the suffix value of symbolVal.

[0264] The suffix of the TR bin string (ie, the suffix bin string) may be derived based on the FL binarization process for suffixVal, and the value cMax of suffixVal is (1 << cRiceParam)-1.

[0265] Meanwhile, if the value of the input parameter (ie, cRiceParam) is 0, TR binarization may be exact truncated unary binarization, and may always use the same value cMax as the possible maximum value of the syntax element to be decoded.

[0266] In addition, for example, the EGk binarization process may be performed as follows: The syntax element coded with ue(v) may be a syntax element subjected to Exponential Golomb coding.

[0267] For example, a 0-order Exponential Golomb (EG0) binarization process may be performed as follows.

[0268] The parsing process for a syntax element may start by reading the bits including the first non-zero bit starting at the current position of the bitstream and counting the number of leading bits equal to 0. The process may be represented as shown in the following table.

[0269] [Table 8]

[0270]

[0271] Additionally, the variable "codeNum" may be derived as shown in the following equation.

[0272] [Equation 9]

[0273]

[0274] Here, the value returned from read_bits(leadingZeroBits), that is, the value indicated by read_bits(leadingZeroBits), may be interpreted as a binary representation of an unsigned integer with the most significant bit recorded first.

[0275] The structure of the Exponential-Golomb code in which the bit string is divided into "prefix" bits and "suffix" bits can be represented as shown in the following table.

[0276] [Table 9]

[0277]

[0278] The "prefix" bits may be bits parsed as described above to calculate leadingZeroBits, and may be represented by 0 or 1 of the bit string in Table 9. That is, the bit string disclosed by 0 or 1 in Table 9 above may represent the prefix bit string. The "suffix" bits may be bits parsed in the calculation of codeNum, and may be represented by xi in Table 9 above. That is, the bit string disclosed as xi in Table 9 above may represent the suffix bit string. Here, i may be a value in the range of LeadingZeroBits-1. In addition, each xi may be equal to 0 or 1.

[0279] The bit string assigned to CodeNum may be as shown in the following table.

[0280] [Table 10]

[0281]

[0282] If the descriptor of the syntax element is ue(v), that is, if the syntax element is coded with ue(v), the value of the syntax element may be equal to codeNum.

[0283] In addition, for example, the EGk binarization process can be performed as follows.

[0284] The input to the EGk binarization process may be a request for EGk binarization. Additionally, the output of the EGk binarization process may be the EGk binarization for symbolVal (ie, the value corresponding to the bin string).

[0285] The bit string used for the EGk binarization process of symbolVal can be derived as follows.

[0286] [Table 11]

[0287]

[0288] Referring to Table 11 above, each call to put(X) adds a binary value X to the end of the bin string. Here, X can be 0 or 1.

[0289] In addition, for example, the limited EGk binarization process can be performed as follows.

[0290] The input to the finite EGk binarization process may be a request for finite EGk binarization, a rice parameter riceParam, a variable log2TransformRange representing the base-2 logarithm of the maximum value, and a variable maxPreExtLen representing the maximum prefix extension length. Additionally, the output of the finite EGk binarization process may be a finite EGk binarization for symbolVal, which is a value corresponding to an empty string.

[0291] The bit string for the finite EGk binarization process of symbolVal can be derived as follows.

[0292] [Table 12]

[0293]

[0294] In addition, for example, the FL binarization process can be performed as follows.

[0295] The input to the FL binarization process may be a request for FL binarization and cMax for a syntax element. Additionally, the output of the FL binarization process may be the FL binarization for symbolVal which is the value corresponding to the bin string.

[0296] FL binarization can be configured using a bit string for symbolVal with a fixed bit length. Here, the fixed-length bit string can be an unsigned integer bit string. That is, the bit string for symbolVal, which serves as the symbol value, can be derived through FL binarization, and the bit length (i.e., the number of bits) of the bit string can be fixed.

[0297] For example, the fixed length can be derived as shown in the following equation.

[0298] [Equation 10]

[0299]

[0300] The index of the bin used for FL binarization may be a method of using values ​​that increase sequentially from the most significant bit to the least significant bit. For example, the bin index associated with the most significant bit may be binIdx=0.

[0301] Meanwhile, for example, a binarization process for the syntax element abs_remainder in the residual information may be performed as follows.

[0302] The input to the binarization process of abs_remainder may be a request for binarization of the syntax element abs_remainder[n], the color component cIdx, and the luma position (x0, y0). The luma position (x0, y0) may indicate the top left sample of the current luma transform block based on the top left luma sample of the picture.

[0303] An output of the binarization process for abs_remainder may be a binarization of abs_remainder (ie, a binarized bin string of abs_remainder). A usable bin string for abs_remainder may be derived through the binarization process.

[0304] The Rice parameter cRiceParam for abs_remainer[n] may be derived through a Rice parameter derivation process performed by inputting a color component cIdx and a luma position (x0, y0), a current coefficient scan position (xC, yC), log2TbWidth (which is the base 2 logarithm of the width of the transform block), and log2TbHeight (which is the base 2 logarithm of the height of the transform block). A detailed description of the Rice parameter derivation process will be described later.

[0305] In addition, for example, cMax for abs_remainder[n] currently to be compiled can be derived based on the Rice parameter cRiceParam. cMax can be derived as shown in the following equation.

[0306] [Equation 11]

[0307]

[0308] Meanwhile, the binarization for abs_remainder (ie, the bin string for abs_remainder) can be the concatenation of the prefix bin string and the suffix bin string if there is a suffix bin string. In addition, if there is no suffix bin string, the bin string for abs_remainder can be the prefix bin string.

[0309] For example, the prefix bin string may be derived as follows.

[0310] The prefix value prefixVal of abs_remainder[n] can be derived as shown in the following equation.

[0311] [Equation 12]

[0312]

[0313] The prefix of the bin string of abs_remainder[n] (ie, prefix bin string) can be derived through the TR binarization process for prefixVal, where cMax and cRiceParam are used as input.

[0314] If the prefix bin string is identical to a bit string with all bits being 1 and a bit length of 6, then the suffix bin string of the bin string of abs_remainder[n] may exist and can be derived as described below.

[0315] The Rice parameter derivation process for dec_abs_level[n] can be as follows.

[0316] The input of the Rice parameter derivation process can be the color component index cIdx, the luma position (x0, y0), the current coefficient scan position (xC, yC), log2TbWidth as the base 2 logarithm of the width of the transform block, and log2TbHeight as the base 2 logarithm of the height of the transform block. The luma position (x0, y0) can indicate the upper left sample of the current luma transform block based on the upper left luma sample of the picture. In addition, the output of the Rice parameter derivation process can be the Rice parameter cRiceParam.

[0317] For example, the variable locSumAbs may be derived based on the array AbsLevel[x][y] of a transform block with a given component index cIdx and an upper left luma position (x0, y0) similar to the pseudo code disclosed in the following table.

[0318] [Table 13]

[0319]

[0320] Then, based on the given variable locSumAbs, the Rice parameter cRiceParam can be derived as shown in the following table.

[0321] [Table 14]

[0322]

[0323] Additionally, for example, during the Rice parameter derivation for abs_remainder[n], baseLevel can be set to 4.

[0324] Alternatively, for example, the Rice parameter cRiceParam can be determined based on whether the transform skip is applied to the current block. That is, if the transform is not applied to the current TB including the current CG, in other words, if the transform skip is applied to the current TB including the current CG, then the Rice parameter cRiceParam can be derived as 1.

[0325] Furthermore, the suffix value suffixVal of abs_remainder may be derived as shown in the following equation.

[0326] [Equation 13]

[0327]

[0328] The suffix bin string of the bin string of abs_remainder can be derived through a finite EGk binarization process for suffixVal, where k is set to cRiceParam+1, riceParam is set to cRiceParam, log2TransformRange is set to 15, and maxPreExtLen is set to 11.

[0329] Meanwhile, for example, a binarization process for the syntax element dec_abs_level in the residual information may be performed as follows.

[0330] The input to the binarization process for dec_abs_level may be a request for binarization of the syntax element dec_abs_level[n], the color component cIdx, the luma position (x0, y0), the current coefficient scan position (xC, yC), log2TbWidth as the base 2 logarithm of the width of the transform block, and log2TbHeight as the base 2 logarithm of the height of the transform block. The luma position (x0, y0) may indicate the top left sample of the current luma transform block based on the top left luma sample of the picture.

[0331] An output of the binarization process for dec_abs_level may be a binarization of dec_abs_level (ie, a binarized bin string of dec_abs_level). A usable bin string for dec_abs_level may be derived through the binarization process.

[0332] The Rice parameter cRiceParam for dec_abs_level[n] can be derived by a Rice parameter derivation process performed using the color component cIdx, the luma position (x0, y0), the current coefficient scan position (xC, yC), log2TbWidth as the base 2 logarithm of the width of the transform block, and log2TbHeight as the base 2 logarithm of the height of the transform block. Hereinafter, the Rice parameter derivation process will be described in detail.

[0333] In addition, for example, cMax for dec_abs_level[n] can be derived based on the Rice parameter cRiceParam. cMax can be derived as shown in the following table.

[0334] [Equation 14]

[0335]

[0336] Meanwhile, the binarization for dec_abs_level[n] (i.e., the bin string for dec_abs_level[n]) may be a concatenation of the prefix bin string and the suffix bin string if there is a suffix bin string. In addition, if there is no suffix bin string, the bin string for dec_abs_level[n] may be a prefix bin string.

[0337] For example, the prefix bin string may be derived as follows.

[0338] The prefix value prefixVal of dec_abs_level[n] can be derived as shown in the following equation.

[0339] [Equation 15]

[0340]

[0341] The prefix of the bin string of dec_abs_level[n] (ie, prefix bin string) may be derived through the TR binarization process for prefixVal, where cMax and cRiceParam are used as inputs.

[0342] If the prefix bin string is identical to a bit string with all bits set to 1 and a bit length of 6, the suffix bin string of the bin string of dec_abs_level[n] may exist and may be derived as described below.

[0343] The Rice parameter derivation process for dec_abs_level[n] can be as follows.

[0344] The input of the Rice parameter derivation process can be the color component index cIdx, the luma position (x0, y0), the current coefficient scan position (xC, yC), log2TbWidth as the base 2 logarithm of the width of the transform block, and log2TbHeight as the base 2 logarithm of the height of the transform block. The luma position (x0, y0) can indicate the upper left sample of the current luma transform block based on the upper left luma sample of the picture. In addition, the output of the Rice parameter derivation process can be the Rice parameter cRiceParam.

[0345] For example, the variable locSumAbs may be derived based on the array AbsLevel[x][y] of a transform block with a given component index cIdx and an upper left luma position (x0, y0) similar to the pseudo code disclosed in the following table.

[0346] [Table 15]

[0347]

[0348] Then, based on the given variable locSumAbs, the Rice parameter cRiceParam can be derived as shown in the following table.

[0349] [Table 16]

[0350]

[0351] In addition, for example, in the Rice parameter derivation process for dec_abs_level[n], baseLevel can be set to 0, and ZeroPos[n] can be derived as follows.

[0352] [Equation 16]

[0353]

[0354] In addition, the suffix value suffixVal of dec_abs_level[n] can be derived as shown in the following equation.

[0355] [Equation 17]

[0356]

[0357] The suffix bin string of the bin string of dec_abs_level[n] can be derived through a finite EGk binarization process for suffixVal, where k is set to cRiceParam+1, truncSuffixLen is set to 15, and maxPreExtLen is set to 11.

[0358] Meanwhile, RRC and TSRC may have the following differences.

[0359] - For example, in TSRC, the Rice parameter for the syntax element abs_remainder[] may be derived as 1. The Rice parameter cRiceParam of the syntax element abs_remainder[] in RRC may be derived based on LastAbsRemainder and lastRiceParam as described above, but the Rice parameter cRiceParam of the syntax element abs_remainder[] in TSRC may be derived as 1. That is, for example, when transform skipping is applied to the current block (e.g., the current TB), the Rice parameter cRiceParam of abs_remainder[] of the TSRC for the current block may be derived as 1.

[0360] - In addition, for example, referring to Table 3 and Table 4, in RRC, abs_level_gtx_flag[n][0] and / or abs_level_gtx_flag[n][1] may be signaled, but in TSRC, abs_level_gtx_flag[n][0], abs_level_gtx_flag[n][1], abs_level_gtx_flag[n][2], abs_level_gtx_flag[n][3], and abs_level_gtx_flag[n][4] may be signaled. Here, abs_level_gtx_flag[n][0] may be expressed as abs_level_gt1_flag or a first coefficient level flag, abs_level_gtx_flag[n][1] may be expressed as abs_level_gt3_flag or a second coefficient level flag, abs_level_gtx_flag[n][2] may be expressed as abs_level_gt5_flag or a third coefficient level flag, abs_level_gtx_flag[n][3] may be expressed as abs_level_gt7_flag or a fourth coefficient level flag, and abs_level_gtx_flag[n][4] may be expressed as abs_level_gt9_flag or a fifth coefficient level flag. Specifically, the first coefficient level flag may be a flag indicating whether the coefficient level is greater than a first threshold value (e.g., 1), the second coefficient level flag may be a flag indicating whether the coefficient level is greater than a second threshold value (e.g., 3), the third coefficient level flag may be a flag indicating whether the coefficient level is greater than a third threshold value (e.g., 5), the fourth coefficient level flag may be a flag indicating whether the coefficient level is greater than a fourth threshold value (e.g., 7), and the fifth coefficient level flag may be a flag indicating whether the coefficient level is greater than a fifth threshold value (e.g., 9). As described above, in TSRC, compared to RRC, abs_level_gtx_flag[n][0], abs_level_gtx_flag[n][1], and abs_level_gtx_flag[n][2], abs_level_gtx_flag[n][3], and abs_level_gtx_flag[n][4] may also be included.

[0361] - Furthermore, for example, in RRC, the syntax element coeff_sign_flag may be bypass coded, but in TSRC, the syntax element coeff_sign_flag may be bypass coded or context coded.

[0362] Meanwhile, in residual data coding of a transform skip applied block (i.e., a transform skip block), when a syntax element is bypass coded, the present disclosure proposes a method for coding bypass coded bins by grouping the bypass coded bins for each syntax element.

[0363] As a proposed embodiment, the number of context-coded bins available for transform-skipped residual coding (i.e., the above-mentioned transform-skipped residual coding (TSRC)) in one TU can be limited to a specific threshold, and in a case where all context-coded bins available for the TU are consumed and then, syntax elements for the TU are coded as bypassed bins, a method is proposed in which a coding order in which syntax elements are prioritized is used instead of a conventional coding order in which coefficient positions are prioritized.

[0364] Specifically, for example, in a conventional TSRC, coding of the syntax elements sig_coeff_flag, coeff_sign_flag, abs_level_gtx_flag[n][0], par_level_flag, abs_level_gtx_flag[n][1], abs_level_gtx_flag[n][2], abs_level_gtx_flag[n][3], abs_level_gtx_flag[n][4], and / or abs_remainder is included. The above syntax elements may be coded in the order shown in the following figures.

[0365] Figure 12 This figure shows an example of coding syntax elements in TSRC.

[0366] Meanwhile, in the present disclosure, a layer may mean a group / unit in which syntax elements are continuously compiled in a single repetitive sentence, and may be described with the same meaning in other elements below. Figure 12 In the example, “sig” may represent sig_coeff_flag, “sign” may represent coeff_sign_flag, gt0 may represent abs_level_gtx_flag[n][0], “par” may represent par_level_flag, gt1 may represent abs_level_gtx_flag[n][1], gt2 may represent abs_level_gtx_flag[n][2], gt3 may represent abs_level_gtx_flag[n][3], gt4 may represent abs_level_gtx_flag[n][4], and “rem” may represent abs_remainder.

[0367] For example, according to Figure 12 In the case of the TSRC of the present embodiment shown in FIG, the syntax elements can be coded in the order of priority of the position of the coefficients in a single layer. That is, for example, referring to Figure 12 , in the first layer, sig_coeff_flag, coeff_sign_flag, abs_level_gtx_flag[n][0], and par_level_flag for a specific coefficient (e.g., Coeff0) may be coded, and sig_coeff_flag, coeff_sign_flag, abs_level_gtx_flag[n][0], and par_level_flag for the next coefficient (e.g., Coeff1) may be coded. Later, for example, in the second layer, abs_level_gtx_flag[n][1], abs_level_gtx_flag[n][2], abs_level_gtx_flag[n][3], and abs_level_gtx_flag[n][4] for a specific coefficient (e.g., Coeff0) are coded, and abs_level_gtx_flag[n][1], abs_level_gtx_flag[n][2], abs_level_gtx_flag[n][3], and abs_level_gtx_flag[n][4] for the next coefficient (e.g., Coeff1) are coded. Subsequently, in the third layer, all coefficients for the subblock (e.g., from Coeff0 to Coeff1) may be coded. n-1 ) to compile abs_remainder.

[0368] Meanwhile, in TSRC of the VVC standard, as described above, the maximum number of bins available for context coding is limited to a specific threshold value for residual data coding (for example, RemCcbs or MaxCcbs shown in Table 4), and the specific threshold value may be derived based on the number of samples included in the transform block or the width and / or height of the transform block, etc. For example, the specific threshold value may be derived as represented by the following equation.

[0369] [Equation 18]

[0370] MaxCcbs = c × horizontal size of transform block × vertical size of transform block

[0371] Here, "c" can represent any real value. In the present disclosure, the value of c is not limited to a specific value. For example, c can have an integer value such as 2 or a decimal value such as 1.5, 1.75, or 1.25. In addition, for example, a threshold value that limits the maximum number of available context coded bins can also be derived based on whether the transform block is a chroma block, the number of samples included in the transform block, and the width and / or height of the transform block. In addition, the threshold value (RemCcbs) can be initialized in units of transform blocks, and the threshold value can be reduced by as much as the number of context coded bins for coding of syntax elements used for residual data coding.

[0372] Meanwhile, in TSRC, the syntax elements sig_coeff_flag, coeff_sign_flag, abs_level_gtx_flag[n][0], par_level_flag, abs_level_gtx_flag[n][1], abs_level_gtx_flag[n][2], abs_level_gtx_flag[n][3], abs_level_gtx_flag[n][4], and / or abs_remainder may be context coded but may also be bypass coded.

[0373] For example, ctxInc for the above syntax elements may be allocated as shown in the following table.

[0374] [Table 17]

[0375]

[0376] As represented in Table 17, when the syntax element sig_coeff_flag, coeff_sign_flag, abs_level_gtx_flag[n][0], par_level_flag, abs_level_gtx_flag[n][1], abs_level_gtx_flag[n][2], abs_level_gtx_flag[n][3], or abs_level_gtx_flag[n][4] is coded, if the threshold (e.g., RemCcbs or MaxCcbs) is greater than 0, the syntax element may be coded as a context-coded bin, and if the threshold is less than or equal to 0, the syntax element may be coded as a bypass bin using a uniform probability distribution.

[0377] Figure 13 FIGURE 1 illustrates another example of coding syntax elements in TSRC. For example, Figure 13It can be illustrated that the threshold value becomes zero after coding the coeff_sign_flag for Coeff0 as a context-coded bin when coding any sub-block / coefficient group in the transform block. In this case, refer to Figure 13 , the syntax elements coded after coeff_sign_flag for Coeff0 including abs_level_gtx_flag[n][0] for Coeff0 can be coded as bypass bins because the remaining threshold is zero, that is, no remaining context coding bins are available. Figure 12 The compilation order of the embodiment shown can be maintained without any changes in Figure 13 In the embodiment shown.

[0378] Figure 14 Another example of coding syntax elements in TSRC is shown. Figure 14 It can be illustrated that in coding any subblock / coefficient group in the transform block, the threshold value becomes zero after coding abs_level_gtx_flag[n][2] for Coeff0 as a context coded bin. In this case, refer to Figure 14 , syntax elements coded after abs_level_gtx_flag[n][2] for Coeff0 including abs_level_gtx_flag[n][3] for Coeff0 can be coded as bypassed bins because the remaining threshold is zero, that is, no remaining context coding bins are available. Figure 12 The compilation order of the embodiment shown in can be maintained without any changes in Figure 14 In the embodiment shown.

[0379] As described above, in the process of coding the TSRC for a block in TSRC, when all available remaining context coded bins of the block are used, subsequent syntax elements can be bypass coded. However, in conventional TSRC, when all available remaining context coded bins are used and syntax elements are coded as bypass bins, as described above, syntax elements can be coded in a coding order with coefficient position priority. Therefore, the present disclosure proposes an embodiment of coding syntax elements with bypass coding in a coding order with syntax element priority instead of the existing coding order with coefficient position priority. Through this, the advantages of a bypass coding engine with high throughput can be maximized, and the residual data coding efficiency of an image can be improved.

[0380] Specifically, for example, the entropy encoder / decoder may include a binarization unit, a conventional coding engine, and a bypass coding engine. For example, the value of a syntax element may be input to the binarization unit. The binarization unit may transform the value of the syntax element into a bin string and output the bin string. Here, a bin string may refer to a binary sequence or a binary code consisting of one or more bins. A bin may refer to the value (0 or 1) of each digit constituting the binary sequence (or binary code) when the value of a symbol and / or syntax element is represented as a binary sequence (or binary code) through binarization.

[0381] Later, the binarized signal (bin string) can be input to a regular compilation engine or a bypass compilation engine. The regular compilation engine can assign a context reflecting the probability value of the corresponding bin and compile the corresponding bin based on the assigned context. The regular compilation engine can perform compilation on each bin and then update the probability and / or context of the bin. The bin compiled by using the regular compilation engine can be called a context-compiled bin.

[0382] Furthermore, the bypass compilation engine can bypass the process of estimating the probability for an input bin and updating the probability applied to the bin after compiling the probability. In bypass mode, context is not assigned based on the input bin, but the input bin is simply compiled, which can improve throughput. For example, in bypass mode, the compilation process can be performed by applying a uniform probability distribution (0.5). A bin compiled using the bypass compilation engine can be referred to as a bypass-compiled bin or a bypass bin.

[0383] Generally, bypass mode has better throughput performance than context-coded bins. To compile one context-coded bin, one or more processing cycles may be required. However, for the bypass coding engine, only one cycle is required to compile n bypass-coded bins. Here, n can be greater than 1. In order to improve the throughput of entropy coding, it may be beneficial to change the coding order (i.e., grouping) so that the bypass-coded bins are compiled continuously. In particular, in the case of grouping the bypass-coded bins, it may be beneficial to group the bypass-coded bins for each syntax element from the perspective of throughput and hardware complexity.

[0384] Figure 15 The diagram shows an example of coding the bypass-coded syntax elements in the TSRC in a syntax element-first coding order rather than a coefficient position-first coding order. Figure 15It can be illustrated that in the coding of any subblock / coefficient group in the transform block, the threshold value becomes zero after the coeff_sign_flag for Coeff0 is coded as a bin for context coding. In this case, refer to Figure 15 , the syntax elements coded after coeff_sign_flag for Coeff0 including abs_level_gtx_flag[n][0] for Coeff0 can be coded as bypass bins because the remaining threshold is 0, that is, no remaining context coding bins are available. For example, refer to Figure 15 , the remaining context elements (abs_level_gtx_flag[n][0] and par_level_flag) for the Coeff0 context coding in the first layer can be bypassed according to the existing coding order, but later, for coefficients Coeff1 to Coeff n-1 The syntax elements of can be coded in the order of syntax elements. In other words, later, the coefficients Coeff1 to Coeff n-1 The sig_coeff_flag of the sig_coeff_flag is bypass coded, and then the coefficients Coeff1 to Coeff n-1 Next, the coeff_sign_flag for coefficients Coeff1 to Coeff n-1 abs_level_gtx_flag[n][0] is bypass-coded, and then, the coefficients Coeff1 to Coeff n-1 par_level_flag for bypass compilation. In addition, for example, refer to Figure 15 , for coding syntax elements in subsequent layers after the first layer, a suggested coding order in which syntax elements are prioritized may be maintained.

[0385] Figure 16 The diagram shows an example of coding the bypass-coded syntax elements in the TSRC in a syntax element-first coding order rather than a coefficient position-first coding order. Figure 16 It can be illustrated that in the coding of any subblock / coefficient group in the transform block, the threshold value becomes zero after abs_level_gtx_flag[n][2] for Coeff0 is coded as a context coded bin. In this case, refer to Figure 16, syntax elements coded after abs_level_gtx_flag[n][2] for Coeff0 including abs_level_gtx_flag[n][3] for Coeff0 can be coded as bypass bins because the remaining threshold is 0, that is, there is no remaining context coded bin available. For example, refer to Figure 16 , as the remaining context elements (abs_level_gtx_flag[n][3] and abs_level_gtx_flag[n][4]) for context coding of Coeff0 in the second layer can be bypass coded according to the existing coding order, but later, for coefficients Coeff1 to Coeff n-1 The syntax elements of can be coded in syntax element order. In other words, later, for coefficients Coeff1 to Coeff n-1 abs_level_gtx_flag[n][1] can be bypassed continuously and subsequently used for coefficients Coeff1 to Coeff n-1 The abs_level_gtx_flag[n][2] of can be bypassed continuously. Next, the coefficients Coeff1 to Coeff n-1 The abs_level_gtx_flag[n][3] of the gtx_level_gtx_flag[n][3] is bypass-coded, and then the coefficients Coeff1 to Coeff n-1 abs_level_gtx_flag[n][4] is bypassed for compilation. In addition, for example, refer to Figure 16 , even for coding the syntax element in a subsequent layer after the second layer, the suggested coding order that takes precedence over the syntax element can be maintained.

[0386] In addition, the present disclosure proposes a method for coding bypass-coded bins by grouping the bypass-coded bins of each syntax element when the syntax elements are bypass-coded in a simplified residual data coding structure. Under certain conditions, there are advantages in coding performance such as lossless coding or near-lossless coding, and the simplified residual coding structure can be used for a single coding block or transform block. In this case, by using the method proposed in the present disclosure, a syntax element-first coding order can be used instead of the existing coefficient position-first coding order.

[0387] In addition, because the number of context coding bins available for residual coding in a TU may be limited to a specific threshold, in the case where all the context coding bins available for the TU are consumed and the syntax elements for the TU are subsequently coded as bypass bins, one embodiment of the present disclosure proposes a method of using a syntax element-first coding order instead of the existing coefficient position-first coding order. At the same time, the existing simplified residual data coding structure may be the same as those shown in the following figures and descriptions for the figures.

[0388] Figure 17 An example in which syntax elements are coded in a simplified residual data coding structure is shown. Figure 17 , the existing simplified residual data coding may include the coding of syntax elements sig_coeff_flag, coeff_sign_flag and abs_remainder. Figure 17 The syntax element coding order of the simplified residual data coding structure with the existing coding order may be illustrated, where coefficient position takes precedence for one sub-block / coefficient group / transform block / coding block.

[0389] For example, reference Figure 17, in the first layer, sig_coeff_flag and coeff_sign_flag for a specific coefficient can be coded, sig_coeff_flag and coeff_sign_flag for the next coefficient after the specific coefficient can be coded, and sig_coeff_flag and coeff_sign_flag for the coefficients up to the last coefficient position in the scan order can be coded. Later, in the second layer, the coding of abs_remainder for all coefficients in the subblock can be performed in scan order. Here, in the case where the value of the coefficient of the position is zero, the value of sig_coeff_flag can be zero, and in the case where the value of the coefficient of the position is non-zero, the value of sig_coeff_flag can be 1. In addition, coeff_sign_flag can indicate the sign of the coefficient of the position. For example, in the case where the coefficient of the position is zero, that is, in the case where the sig_coeff_flag for the coefficient is zero, the coeff_sign_flag for the coefficient may not be coded. Further, when a coefficient is non-zero and negative, the coeff_sign_flag value for the coefficient may be 1 (or zero), and when the coefficient is non-zero and positive, the coeff_sign_flag value for the coefficient may be zero (or 1). Alternatively, when a coefficient is negative regardless of sig_coeff_flag for the coefficient, the coeff_sign_flag value for the coefficient may be 1 (or zero), and when the coefficient is positive or zero, the coeff_sign_flag value for the coefficient may be zero (or 1). Alternatively, when a coefficient is positive regardless of sig_coeff_flag for the coefficient, the coeff_sign_flag value for the coefficient may be 1 (or zero), and when the coefficient is negative or zero, the coeff_sign_flag value for the coefficient may be zero (or 1).

[0390] Meanwhile, the present disclosure proposes a method for coding and grouping each syntax element to generate advantages in CABAC throughput and hardware complexity in a simplified residual data coding structure.

[0391] Figure 18 The diagram shows an example of coding the syntax elements that are bypass-coded in the simplified residual data coding structure in a syntax element-first coding order rather than a coefficient position-first coding order. Figure 18 As shown in the first layer, it is possible to go from sig_coeff_flag of Coeff0 to Coeff n-1(the last coefficient in the scan order) performs bypass coding continuously, and then, it is possible to bypass coding from the coeff_sign_flag of Coeff0 to the coeff n-1 Later, you can bypass coding from abs_remainder of Coeff0 to abs_remainder of Coeff n-1 (last coefficient in scan order) abs_remainder continuously performs bypass compilation.

[0392] At the same time, the simplified residual data coding structure can have the same Figure 17 For example, the simplified residual data coding structure shown in the figure below can be compiled.

[0393] Figure 19a and 19b FIGURE 1 illustrates an embodiment in which syntax elements are coded in a simplified residual data coding structure. Figure 19a and 19b , simplified residual data coding may include coding of syntax elements dec_abs_level and coeff_sign_flag. Figure 19a and 19b The syntax element coding order of the simplified residual data coding structure with the existing coding order can be illustrated, where for one sub-block / coefficient group / transform block / coding block, the coefficient position is prioritized. Figure 19a , in a layer, dec_abs_level and coeff_sign_flag for a specific coefficient may be coded, dec_abs_level and coeff_sign_flag for the next coefficient after the specific coefficient may be coded, and dec_abs_level and coeff_sign_flag for coefficients up to the last coefficient position in the scan order may be coded. In addition, for example, referring to Figure 19b , in a layer, coeff_sign_flag and dec_abs_level for a specific coefficient may be coded, coeff_sign_flag and dec_abs_level for the next coefficient after the specific coefficient may be coded, and coeff_sign_flag and dec_abs_level for coefficients up to the last coefficient position in the scanning order may be coded.

[0394] Here, if the value of the coefficient at the position is zero, the value of dec_abs_level may be zero, and if the value of the coefficient at the position is non-zero, the value of dec_abs_level may be the absolute value of the coefficient. Furthermore, coeff_sign_flag may indicate the sign of the coefficient at the position. For example, if the coefficient at the position is zero, that is, if the dec_abs_level for the coefficient is zero, the coeff_sign_flag for the coefficient may not be coded. Furthermore, if the coefficient is non-zero and negative, the coeff_sign_flag value for the coefficient may be 1 (or zero), and if the coefficient is non-zero and positive, the coeff_sign_flag value for the coefficient may be zero (or 1). Alternatively, if the coefficient is negative, regardless of the dec_abs_level for the coefficient, the coeff_sign_flag value for the coefficient may be 1 (or zero), and if the coefficient is positive or zero, the coeff_sign_flag value for the coefficient may be zero (or 1). Alternatively, the coeff_sign_flag value for a coefficient may be 1 (or zero) when the coefficient is positive regardless of the dec_abs_level for the coefficient, and may be zero (or 1) when the coefficient is negative or zero.

[0395] At the same time, a simplified residual data coding structure can be used when specific conditions in the above RRC or TSRC are met. For example, when the current block is losslessly coded or nearly losslessly coded, or when all available context coding bins are consumed for the current block, the residual data can be coded in the simplified residual data coding structure.

[0396] For example, the present disclosure proposes a method for coding syntax elements for coefficients of a current block in a simplified residual data coding structure when all available context coded bins for the current block are consumed in the TSRC for the current block.

[0397] Specifically, for example, syntax elements based on the TSRC for the current block can be parsed. In this case, the maximum number of bins that can be used for context coding of the current block can be derived. If all of the maximum number of context-coded bins for the current block are used to code syntax elements for a transform coefficient preceding the current transform coefficient in scan order, syntax elements for the current transform coefficient and subsequent transform coefficients of the current transform coefficient in scan order can be coded in a simplified residual data coding structure. Therefore, the syntax elements for the current transform coefficient and subsequent transform coefficients of the current transform coefficient in scan order can include a sign flag and coefficient level information for the transform coefficient. The decoding device can derive the transform coefficient based on the syntax elements for the transform coefficient coded in the simplified residual data coding structure. For example, the coefficient level information can indicate the absolute value of the coefficient level of the transform coefficient. Furthermore, the sign flag can indicate the sign of the current transform coefficient. The decoding device can derive the coefficient level of the transform coefficient based on the coefficient level information and the sign of the transform coefficient based on the sign flag.

[0398] Meanwhile, in order to generate advantages in CABAC throughput and hardware complexity, the present disclosure proposes a method for coding and grouping syntax elements that are bypass-coded in the simplified residual data coding structure shown in FIG. 19 for each syntax element.

[0399] Figure 20a and 20b The diagram shows an example of coding the syntax elements bypass-coded in the simplified residual data coding structure in a syntax element-first coding order rather than a coefficient position-first coding order. Figure 20a As shown in , in the layer, it is possible to go from dec_abs_level of Coeff0 to Coeff n-1 The bypass coding is performed continuously from the dec_abs_level of Coeff0 to the dec_abs_level of Coeff0 (the last coefficient in the scan order), and then, from the coeff_sign_flag of Coeff0 to the dec_abs_level of Coeff0 n-1 Alternatively, for example, according to this embodiment, as Figure 20b As shown in the layer, it is possible to go from coeff_sign_flag of Coeff0 to Coeff n-1 The bypass coding is performed continuously from the coeff_sign_flag of Coeff0 (the last coefficient in the scan order), and then, from the dec_abs_level of Coeff0 to the dec_abs_level of Coeff n-1According to this embodiment, the bypass coded bins can be grouped and coded continuously for each syntax element, which can improve the throughput of entropy coding and reduce hardware complexity.

[0400] At the same time, the simplified residual data coding structure can have the same Figure 17 、 Figure 19a as well as Figure 19b For example, the simplified residual data coding structure shown in the figure below can be compiled.

[0401] Figure 21 FIGURE 1 illustrates an embodiment in which syntax elements are coded in a simplified residual data coding structure. Figure 21 , simplified residual data coding may include coding of syntax elements sig_coeff_flag, coeff_sign_flag, abs_level_gtx_flag[n][0], par_level_flag, and abs_remainder. Figure 21 The syntax element coding order of the simplified residual data coding structure with the existing coding order can be illustrated, where for one sub-block / coefficient group / transform block / coding block, the coefficient position is prioritized. Figure 21 In the first layer, sig_coeff_flag, coeff_sign_flag, abs_level_gtx_flag[n][0], and par_level_flag for a specific coefficient may be coded, sig_coeff_flag, coeff_sign_flag, abs_level_gtx_flag[n][0], and par_level_flag for the next coefficient after the specific coefficient may be coded, and sig_coeff_flag, coeff_sign_flag, abs_level_gtx_flag[n][0], and par_level_flag for coefficients up to the last coefficient position in the scan order may be coded. Later, in the second layer, abs_remainder coding for all coefficients in the subblock can be performed in the scan order.

[0402] At the same time, in order to produce advantages in terms of CABAC throughput and hardware complexity, the present disclosure proposes a method for Figure 21 The method for coding and grouping syntax elements of bypass coding in the simplified residual data coding structure shown in FIG.

[0403] Figure 22 The diagram shows an example of coding the syntax elements bypass-coded in the simplified residual data coding structure in a syntax element-first coding order rather than a coefficient position-first coding order. Figure 22 As shown in the first layer, it is possible to go from sig_coeff_flag of Coeff0 to Coeff n-1 (last coefficient in scan order) is bypass coded continuously, followed by coeff_sign_flag from Coeff0 to Coeff n-1 (last coefficient in scan order) is executed consecutively, followed by abs_level_gtx_flag[n][0] from Coeff0 to Coeff n-1 (last coefficient in scan order) and then from the par_level_flag of Coeff0 to Coeff n-1 (last coefficient in scan order) par_level_flag is executed continuously. Later, in the second layer, it is possible to go from abs_remainder of Coeff0 to abs_remainder of Coeff n-1 According to this embodiment, the bypass coded bins can be grouped and coded continuously for each syntax element, and the throughput of entropy coding can be improved and the hardware complexity can be reduced.

[0404] Figure 23 An image encoding method performed by the encoding device according to the present disclosure is briefly shown. Figure 23 The method proposed in Figure 2 Specifically, for example, Figure 23 Step S2300 shown in FIG. 2 may be performed by a predictor of an encoding device. Figure 23 Steps S2310 and S2320 shown in FIG. 2 may be performed by a residual processor of an encoding device, and Figure 23 Step S2330 shown in the figure can be performed by the entropy encoder of the encoding device. In addition, although not shown, the process of generating the reconstructed samples and the reconstructed picture for the current block based on the residual samples and the prediction samples for the current block can also be performed by the adder of the encoding device.

[0405] The encoding device derives prediction samples for the current block based on inter-frame prediction or intra-frame prediction (step S2300). The encoding device may derive prediction samples for the current block based on the prediction mode. In this case, various prediction methods disclosed in this disclosure, such as inter-frame prediction or intra-frame prediction, may be applied.

[0406] For example, the encoding device may determine whether to perform inter prediction or intra prediction on the current block and determine a specific inter prediction mode or a specific intra prediction mode based on the RD cost. According to the determined mode, the encoding device may derive a prediction sample for the current block.

[0407] The encoding apparatus derives residual samples of the current block based on the prediction samples (step S2310). For example, the encoding apparatus may derive the residual samples of the current block by subtracting the original samples and the prediction samples for the current block.

[0408] The encoding device derives transform coefficients for the current block based on the residual samples (step S2320). For example, the encoding device may derive transform coefficients for the current block based on the residual samples. For example, the encoding device may determine whether a transform is applied to the current block. That is, the encoding device may determine whether a transform is applied to the residual samples of the current block. The encoding device may determine whether a transform is applied to the current block based on coding efficiency. For example, the encoding device may determine that a transform is not applied to the current block. A block to which a transform is not applied may be referred to as a transform skip block. That is, for example, the current block may be a transform skip block.

[0409] When a transform is not applied to the current block, that is, when a transform is not applied to the residual samples, the encoding device may derive the derived residual samples as transform coefficients. Furthermore, when a transform is applied to the current block, that is, when a transform is applied to the residual samples, the encoding device may derive transform coefficients by performing a transform on the residual samples. The current block may include multiple subblocks or coefficient groups (CGs). Furthermore, the size of the subblocks of the current block may be 4x4 or 2x2. That is, the subblocks of the current block may include a maximum of 16 non-zero transform coefficients or a maximum of 4 non-zero transform coefficients.

[0410] Here, the current block may be a coding block (CB) or a transform block (TB). In addition, the transform coefficient may also be expressed as a residual coefficient.

[0411] The encoding apparatus encodes image information including prediction mode information and residual information for transform coefficients of the current block (step S2330). The encoding apparatus may encode image information including prediction mode information indicating the prediction mode of the current block and residual information for transform coefficients. For example, the encoding apparatus may generate and encode prediction-related information for the current block. The prediction-related information may include prediction mode information.

[0412] Furthermore, for example, the residual information may include syntax elements for transform skip residual coding (TSRC) of transform coefficients of the current block. For example, the encoding device may generate and encode syntax elements for transform skip residual coding (TSRC) of transform coefficients of the current block. For example, the residual information may include syntax elements for a first residual data coding structure according to TSRC and syntax elements for a second residual data coding structure according to TSRC.

[0413] For example, the encoding apparatus may generate and encode syntax elements for the first to nth transform coefficients.The residual information of the current block may include syntax elements for the first to nth transform coefficients of the current block.

[0414] For example, the residual information may include syntax elements for the first to nth transform coefficients of the current block. Here, for example, the syntax elements may be syntax elements of a first residual data coding structure according to transform skip residual coding (TSRC). The syntax elements according to the first residual data coding structure may include syntax elements for context coding of the transform coefficients and / or syntax elements for bypass coding. The syntax elements according to the first residual data coding structure may include syntax elements such as sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, par_level_flag, abs_level_gtX_flag, abs_remainder, and / or coeff_sign_flag.

[0415] For example, a syntax element for context coding of a transform coefficient may include a significant coefficient flag indicating whether the transform coefficient is a non-zero transform coefficient, a sign flag indicating a sign of the transform coefficient, a first coefficient level flag indicating whether the coefficient level of the transform coefficient is greater than a first threshold, and / or a parity level flag indicating parity of the coefficient level of the transform coefficient. Furthermore, for example, the syntax element for context coding may include a second coefficient level flag indicating whether the coefficient level of the transform coefficient is greater than a second threshold, a third coefficient level flag indicating whether the coefficient level of the transform coefficient is greater than a third threshold, a fourth coefficient level flag indicating whether the coefficient level of the transform coefficient is greater than a fourth threshold, and / or a fifth coefficient level flag indicating whether the coefficient level of the transform coefficient is greater than a fifth threshold. Here, the significant coefficient flag may be sig_coeff_flag, the sign flag may be ceff_sign_flag, the first coefficient level flag may be abs_level_gt1_flag, and the parity level flag may be par_level_flag. Furthermore, the second coefficient level flag may be abs_level_gt3_flag or abs_level_gtx_flag, the third coefficient level flag may be abs_level_gt5_flag or abs_level_gtx_flag, the fourth coefficient level flag may be abs_level_gt7_flag or abs_level_gtx_flag, and the fifth coefficient level flag may be abs_level_gt9_flag or abs_level_gtx_flag.

[0416] In addition, for example, the syntax elements for bypass coding of transform coefficients may include coefficient level information for the value (or coefficient level) of the transform coefficient and / or a sign flag indicating the sign of the transform coefficient. The coefficient level information may be abs_remainder and / or dec_abs_level, and the sign flag may be ceff_sign_flag.

[0417] In addition, the number of syntax elements for context coding of the first transform coefficient to the nth transform coefficient can be the same as the maximum number of bins for context coding of the current block. That is, for example, all the bins for context coding of the current block can be used as bins for syntax elements for context coding of the first transform coefficient to the nth transform coefficient. For example, the maximum number of bins for context coding of the current block can be derived based on the width and height of the current block. For example, the maximum number of bins for context coding of the current block can be derived as a value of the number of samples of the current block multiplied by a specific value. Here, the number of samples can be derived as a value multiplied by the width and height of the current block. In addition, the specific value can have an integer value such as 2 or a decimal value such as 1.5, 1.75, or 1.25.

[0418] At the same time, for example, the syntax elements according to the first residual data coding structure can be coded in a coding order according to the coefficient position. The coding order according to the coefficient position may be a scanning order of the transform coefficients. For example, the scanning order may be a raster scan order. For example, the raster scan order may represent an order of scanning from the top row downwards and scanning from left to right in each row. For example, the syntax elements according to the first residual data coding structure can be coded in an order from the syntax element for the first transform coefficient to the syntax element for the nth transform coefficient. In addition, for example, when the current block includes a plurality of sub-blocks or coefficient groups (CGs), the plurality of sub-blocks or coefficient groups can be coded in a scanning order, and the syntax elements for the transform coefficients in each sub-block or coefficient group can be coded in a scanning order.

[0419] At the same time, for example, the residual information may include a transform skip flag for the current block. The transform skip flag may indicate whether a transform is applied to the current block. That is, the transform skip flag may indicate whether a transform is applied to the transform coefficients of the current block. The syntax element indicating the transform skip flag may be the above-mentioned transform_skip_flag. For example, when the value of the transform skip flag is 0, the transform skip flag may indicate that a transform is not applied to the current block, and when the value of the transform skip flag is 1, the transform skip flag may indicate that a transform is applied to the current block. For example, when the current block is a transform skip block, the value of the transform skip flag for the current block may be 1.

[0420] In addition, the encoding apparatus may generate and encode syntax elements for an (n+1)th transform coefficient to a last transform coefficient of the current block.

[0421] For example, the residual information may include syntax elements for the n+1th transform coefficient to the last transform coefficient of the current block. Here, for example, the syntax elements may be syntax elements of the second residual data coding structure according to transform skip residual coding (TSRC). The syntax elements according to the second residual data coding structure may be referred to as syntax elements according to a simplified residual data coding structure. For example, when all context-coded bins for the current block are used as bins for syntax elements for context coding of the first transform coefficient to the nth transform coefficient, that is, for example, when the number of syntax elements for context coding of the first transform coefficient to the nth transform coefficient is equal to or greater than the maximum number of context-coded bins of the current block, the encoding device may generate and encode syntax elements for the n+1th transform coefficient to the last transform coefficient of the current block, which are syntax elements of the second residual data coding structure according to TSRC.

[0422] For example, syntax elements according to the second residual data coding structure may include syntax elements for bypass coding of transform coefficients. For example, syntax elements according to the second residual data coding structure may include a sign flag and coefficient level information for the transform coefficients. For example, syntax elements according to the second residual data coding structure may include coefficient level information for the absolute value of the coefficient level of the transform coefficient and a sign flag for the sign of the transform coefficient. Syntax elements for transform coefficients may be coded based on bypass. That is, residual syntax elements for transform coefficients may be coded based on a uniform probability distribution. For example, the coefficient level information may indicate the absolute value of the coefficient level of the transform coefficient. Additionally, the sign flag may indicate the sign of the transform coefficient. For example, when the sign flag has a value of 0, the sign flag may indicate that the coefficient level of the transform coefficient is positive, and when the sign flag has a value of 1, the sign flag may indicate that the coefficient level of the transform coefficient is negative. The coefficient level information may be the abs_remainder described above, and the sign flag may be the coeff_sign_flag described above.

[0423] At the same time, for example, the syntax elements according to the second residual data coding structure can be coded in a coding order according to the coefficient position. The coding order according to the coefficient position may be a scanning order of the transform coefficients. For example, the scanning order may be a raster scanning order. For example, the raster scanning order may represent an order of scanning from the top row downwards and from left to right in each row. For example, the syntax elements according to the second residual data coding structure can be coded in an order from the syntax elements for the n+1th transform coefficient to the last transform coefficient. In addition, for example, when the current block includes a plurality of sub-blocks or coefficient groups (CGs), the plurality of sub-blocks or coefficient groups can be coded in a scanning order, and the syntax elements for the transform coefficients in each sub-block or coefficient group can be coded in a scanning order.

[0424] Furthermore, for example, syntax elements according to the second residual data coding structure can be coded in a coding order according to the syntax elements. That is, for example, syntax elements according to the second residual data coding structure can be coded in a coding order in which the syntax elements are prioritized. For example, coefficient level information for the (n+1)th transform coefficient to the last transform coefficient can be coded, and thereafter, sign flags for the (n+1)th transform coefficient to the last transform coefficient can be coded. Specifically, for example, the coefficient level information can be coded in the order from the coefficient level information for the (n+1)th transform coefficient to the coefficient level information for the last transform coefficient, and later, the sign flags can be coded in the order from the sign flag for the (n+1)th transform coefficient to the sign flag for the last transform coefficient.

[0425] For example, the encoding device may encode image information including prediction mode information and residual information and output the image information in a bitstream format. The bitstream may be transmitted to the decoding device via a network or a storage medium.

[0426] At the same time, the bit stream can be sent to the decoding device via a network or a (digital) storage medium. Here, the network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0427] Figure 24 The following briefly illustrates an encoding device for executing the image encoding method according to the present disclosure. Figure 23 The method proposed in Figure 24 In particular, for example, Figure 24 The predictor of the encoding device shown in FIG can perform Figure 23 In step S2300 shown in FIG. Figure 24 The residual processor of the encoding device shown in FIG can perform Figure 23 , and Figure 24 The entropy encoder of the encoding device shown can perform Figure 23 In addition, although not shown in the drawings, the process of generating the reconstructed samples and the reconstructed picture for the current block may also be performed by the adder of the encoding device based on the residual samples and the prediction samples for the current block.

[0428] Figure 25 The following briefly illustrates an image decoding method performed by a decoding device according to the present disclosure. Figure 25 The method proposed in Figure 3 Specifically, for example, Figure 25Step S2500 shown in FIG. 1 may be performed by an entropy decoder of a decoding device. Figure 25 Steps S2510 and S2520 shown in FIG. 2 may be performed by a predictor of a decoding device. Figure 25 Step S2530 shown in FIG can be performed by the residual processor of the decoding device, and Figure 25 Step S2540 shown in can be performed by an adder of the decoding device.

[0429] The decoding apparatus obtains image information including residual information and prediction mode information for the current block (step S2500). The decoding apparatus may obtain image information including residual information and prediction mode information for the current block. For example, the image information may include prediction mode information for the current block. For example, the image information may include prediction-related information for the current block, and the prediction-related information may include prediction mode information. The prediction mode information may indicate whether inter-frame prediction or intra-frame prediction is applied to the current block.

[0430] In addition, for example, the image information may include residual information for the current block. The decoding apparatus may obtain the image information including the residual information for the current block.

[0431] The residual information may include syntax elements for the transform coefficients of the current block. Here, the current block may be a coding block (CB) or a transform block (TB). In addition, the transform coefficients may also be represented as residual coefficients. Furthermore, for example, the current block may be a transform skip block.

[0432] For example, the residual information may include syntax elements for the first to nth transform coefficients of the current block and syntax elements for the n+1th to last transform coefficients of the current block. In this case, for example, the number of syntax elements used for context coding of the first to nth transform coefficients may be the same as the maximum number of context coding bins for the current block. That is, for example, all of the context coding bins used for the current block may be used as bins for syntax elements used for context coding of the first to nth transform coefficients. For example, the maximum number of context coding bins for the current block may be derived based on the width and height of the current block. For example, the maximum number of context coding bins for the current block may be derived as the number of samples of the current block multiplied by a specific value. Here, the number of samples may be derived as the value multiplied by the width and height of the current block. Furthermore, the specific value may have an integer value such as 2 or a decimal value such as 1.5, 1.75, or 1.25.

[0433] Specifically, for example, the decoding device may obtain syntax elements for syntax elements for the first transform coefficient to the nth transform coefficient of the current block. For example, the residual information may include syntax elements for syntax elements for the first transform coefficient to the nth transform coefficient of the current block. Here, the syntax elements may be syntax elements of a first residual data coding structure according to transform skip residual coding (TSRC). The syntax elements according to the first residual data coding structure may include syntax elements for context coding of transform coefficients and / or syntax elements for bypass coding. The syntax elements according to the first residual data coding structure may include syntax elements such as sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, par_level_flag, abs_level_gtX_flag, abs_remainder, and / or coeff_sign_flag.

[0434] For example, a syntax element for context coding of a transform coefficient may include a significant coefficient flag indicating whether the transform coefficient is a non-zero transform coefficient, a sign flag indicating a sign of the transform coefficient, a first coefficient level flag indicating whether the coefficient level for the transform coefficient is greater than a first threshold, and / or a parity level flag indicating the parity of the coefficient level for the transform coefficient. Furthermore, for example, the syntax element for context coding may include a second coefficient level flag indicating whether the coefficient level for the transform coefficient is greater than a second threshold, a third coefficient level flag indicating whether the coefficient level for the transform coefficient is greater than a third threshold, a fourth coefficient level flag indicating whether the coefficient level for the transform coefficient is greater than a fourth threshold, and / or a fifth coefficient level flag indicating whether the coefficient level for the transform coefficient is greater than a fifth threshold. Here, the significant coefficient flag may be sig_coeff_flag, the sign flag may be ceff_sign_flag, the first coefficient level flag may be abs_level_gt1_flag, and the parity level flag may be par_level_flag. Furthermore, the second coefficient level flag may be abs_level_gt3_flag or abs_level_gtx_flag, the third coefficient level flag may be abs_level_gt5_flag or abs_level_gtx_flag, the fourth coefficient level flag may be abs_level_gt7_flag or abs_level_gtx_flag, and the fifth coefficient level flag may be abs_level_gt9_flag or abs_level_gtx_flag.

[0435] In addition, for example, the syntax elements for bypass coding of transform coefficients may include coefficient level information for the value (or coefficient level) of the transform coefficient and / or a sign flag indicating the sign of the transform coefficient. The coefficient level information may be abs_remainder and / or dec_abs_level, and the sign flag may be ceff_sign_flag.

[0436] At the same time, for example, the syntax elements according to the first residual data coding structure can be encoded in a coding order according to the coefficient position. The coding order according to the coefficient position may be a scanning order of the transform coefficients. For example, the scanning order may be a raster scanning order. For example, the raster scanning order may represent an order of scanning from the top row downwards and scanning from left to right in each row. For example, the syntax elements according to the first residual data coding structure can be coded in an order from the syntax elements of the first transform coefficient to the syntax elements of the nth transform coefficient. In addition, for example, when the current block includes a plurality of sub-blocks or coefficient groups (CGs), the plurality of sub-blocks or coefficient groups can be coded in a scanning order, and the syntax elements for the transform coefficients in each sub-block or coefficient group can be coded in a scanning order.

[0437] In addition, for example, the residual information may include a transform skip flag for the current block. The transform skip flag may indicate whether a transform is applied to the current block. That is, the transform skip flag may indicate whether a transform is applied to the transform coefficients of the current block. The syntax element indicating the transform skip flag may be the above-mentioned transform_skip_flag. For example, when the value of the transform skip flag is 0, the transform skip flag may indicate that a transform is not applied to the current block, and when the value of the transform skip flag is 1, the transform skip flag may indicate that a transform is applied to the current block. For example, when the current block is a transform skip block, the value of the transform skip flag for the current block may be 1.

[0438] In addition, for example, the decoding device may obtain syntax elements for the n+1th transform coefficient to the last transform coefficient of the current block. For example, the residual information may include syntax elements for the n+1th transform coefficient to the last transform coefficient of the current block. That is, for example, the syntax elements may be syntax elements of the second residual data coding structure according to transform skip residual coding (TSRC). The syntax elements according to the second residual data coding structure may be referred to as syntax elements according to a simplified residual data coding structure. For example, when all the bins used for context coding of the current block are used as bins for syntax elements for context coding of the first transform coefficient to the nth transform coefficient, that is, when the number of syntax elements for context coding of the first transform coefficient to the nth transform coefficient is equal to or greater than the maximum number of bins for context coding of the current block, the decoding device may obtain syntax elements for the n+1th transform coefficient to the last transform coefficient of the current block, which are syntax elements of the second residual data coding structure according to TSRC.

[0439] For example, syntax elements according to the second residual data coding structure may include syntax elements for bypass coding of transform coefficients. For example, syntax elements according to the second residual data coding structure may include a sign flag and coefficient level information for the transform coefficients. For example, syntax elements according to the second residual data coding structure may include coefficient level information for the absolute value of the coefficient level of the transform coefficient and a sign flag for the sign of the transform coefficient. Syntax elements for transform coefficients may be coded based on bypass. That is, residual syntax elements for transform coefficients may be coded based on a uniform probability distribution. For example, the coefficient level information may indicate the absolute value of the coefficient level of the transform coefficient. Additionally, the sign flag may indicate the sign of the transform coefficient. For example, when the sign flag has a value of 0, the sign flag may indicate that the coefficient level of the transform coefficient is positive, and when the sign flag has a value of 1, the sign flag may indicate that the coefficient level of the transform coefficient is negative. The coefficient level information may be the abs_remainder described above, and the sign flag may be the coeff_sign_flag described above.

[0440] At the same time, for example, the syntax elements according to the second residual data coding structure can be coded in a coding order according to the coefficient position. The coding order according to the coefficient position may be a scanning order of the transform coefficients. For example, the scanning order may be a raster scanning order. For example, the raster scanning order may represent an order of scanning from the top row downwards and from left to right in each row. For example, the syntax elements according to the second residual data coding structure can be coded in an order from the syntax elements for the n+1th transform coefficient to the last transform coefficient. In addition, for example, when the current block includes a plurality of sub-blocks or coefficient groups (CGs), the plurality of sub-blocks or coefficient groups can be coded in a scanning order, and the syntax elements of the transform coefficients in each sub-block or coefficient group can be coded in a scanning order.

[0441] Furthermore, for example, syntax elements according to the second residual data coding structure can be coded in a coding order according to the syntax elements. That is, for example, syntax elements according to the second residual data coding structure can be coded in a coding order in which the syntax elements are prioritized. For example, coefficient level information for the (n+1)th transform coefficient to the last transform coefficient can be coded, and thereafter, sign flags for the (n+1)th transform coefficient to the last transform coefficient can be coded. Specifically, for example, the coefficient level information can be coded in the order from the coefficient level information for the (n+1)th transform coefficient to the coefficient level information for the last transform coefficient, and later, the sign flags can be coded in the order from the sign flag for the (n+1)th transform coefficient to the sign flag for the last transform coefficient.

[0442] The decoding apparatus derives a prediction mode for the current block based on the prediction mode information (step S2510 ). The decoding apparatus may determine whether to apply inter prediction or intra prediction to the current block based on the prediction mode information, and perform prediction based on the determination.

[0443] The decoding apparatus derives prediction samples for the current block based on the prediction mode (step S2520). For example, the decoding apparatus may derive the prediction mode applied to the current block based on the prediction mode information and derive the prediction samples for the current block based on the prediction mode. For example, when inter-frame prediction is applied to the current block, the decoding apparatus may derive motion information for the current block based on the prediction mode information included in the image information and derive the prediction samples for the current block based on the motion information. Furthermore, for example, when intra-frame prediction is applied to the current block, the decoding apparatus may derive reference samples based on neighboring samples of the current block and derive the prediction samples for the current block based on the intra-frame prediction mode. The reference samples for the current block may include an upper reference sample and a left reference sample of the current block. For example, when the size of the current block is NxN, the x component of the upper left sample position of the current block is 0, and the y component of the upper left sample position of the current block is 0, the left reference sample may be p[-1][0] to p[-1][2N-1], and the upper reference sample may be p[0][-1] to p[2N-1][-1].

[0444] The decoding apparatus derives residual samples of the current block based on the residual information (step S2530). The decoding apparatus may derive residual samples of the current block based on the residual information.

[0445] For example, the decoding apparatus may derive syntax elements according to the first residual data coding structure and syntax elements according to the second residual data coding structure.

[0446] For example, the decoding apparatus may derive first to nth transform coefficients of the current block based on syntax elements according to the first residual data coding structure.

[0447] In addition, for example, the decoding device may derive the (n+1)th transform coefficient to the last transform coefficient based on the syntax elements according to the second residual data coding structure. For example, the coefficient level for the transform coefficient may be derived as a value indicated by the coefficient level information, and the sign of the transform coefficient may be derived as a sign indicated by the sign flag. In this case, for example, the transform coefficient can be derived without performing level mapping.

[0448] Later, for example, the decoding device may derive residual samples for the current block based on the transform coefficients. In one example, when the transform skip flag indicates that a transform is not applied to the current block, that is, when the transform skip flag has a value of 1, the decoding device may derive the transform coefficients as the residual samples for the current block. Alternatively, for example, when the transform skip flag indicates that a transform is not applied to the current block, that is, when the transform skip flag has a value of 1, the decoding device may derive residual samples for the current block by dequantizing the transform coefficients. Alternatively, for example, when the transform skip flag indicates that a transform is applied to the current block, that is, when the transform skip flag has a value of zero, the decoding device may derive residual samples for the current block by inversely transforming the transform coefficients. Alternatively, for example, when the transform skip flag indicates that a transform is applied to the current block, that is, when the transform skip flag has a value of zero, the decoding device may derive residual samples for the current block by dequantizing the transform coefficients and inversely transforming the dequantized transform coefficients.

[0449] The decoding apparatus generates reconstructed samples of the current block based on the predicted samples and the residual samples (step S2540). For example, the decoding apparatus may generate reconstructed samples and / or a reconstructed picture of the current block based on the predicted samples and the residual samples. For example, the decoding apparatus may generate the reconstructed samples by adding the predicted samples and the residual samples.

[0450] Later, as occasion demands, in order to improve subjective / objective image quality, the decoding apparatus may apply a loop filtering process, such as deblocking filtering, SAO and / or ALF process, to the reconstructed picture as described above.

[0451] Figure 26 A decoding device for performing the image decoding method according to the present disclosure is briefly illustrated. Figure 25 The method proposed in Figure 26 In particular, for example, Figure 26 The entropy decoder of the decoding device shown in FIG can perform Figure 25 In step S2500 shown in FIG. Figure 26 The predictor of the decoding device shown in FIG can perform Figure 25 Steps S2510 and S2520 shown in FIG. Figure 26 The residual processor of the decoding device shown in FIG can perform Figure 25 , and Figure 26 The adder of the decoding device shown in FIG. 1 can perform Figure 25 Step S2540 is shown.

[0452] According to the present disclosure described above, the efficiency of residual coding can be improved.

[0453] In addition, according to the present disclosure, when the maximum number of context coding bins for the current block is consumed in TSRC, syntax elements according to a simplified residual data coding structure can be signaled, and by this, the coding complexity of the bypass-coded syntax elements is reduced and the overall residual coding efficiency can be improved.

[0454] In addition, according to the present disclosure, as the coding order of bypass-coded syntax elements, an order in which syntax elements are prioritized can be used, and through this, the coding efficiency of bypass-coded syntax elements can be improved, and the overall residual coding efficiency can be improved.

[0455] In the above embodiments, the method is described based on a flow chart with a series of steps or boxes. The present disclosure is not limited to the order of the above steps or boxes. Some steps or boxes can be performed in an order different from the above-mentioned other steps or boxes or performed simultaneously. In addition, it will be understood by those skilled in the art that the steps shown in the flow chart are not exclusive and may also include other steps, or one or more steps in the flow chart may be deleted without affecting the scope of the present disclosure.

[0456] The embodiments described in this specification can be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information about instructions) or algorithms used for implementation can be stored in a digital storage medium.

[0457] In addition, the decoding device and encoding device to which the present disclosure is applied may be included in the following devices: multimedia broadcast transmission / reception devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, portable cameras, VoD service providers, over-the-top (OTT) video devices, Internet streaming service providers, three-dimensional (3D) video devices, teleconferencing video devices, transportation user devices (e.g., vehicle user devices, aircraft user devices, and ship user devices), and medical video equipment, and the decoding device and encoding device to which the present disclosure is applied may be used to process video signals or data signals. For example, over-the-top (OTT) video devices may include game consoles, Blu-ray players, Internet-connected televisions, home theater systems, smartphones, tablet computers, digital video recorders (DVRs), and the like.

[0458] In addition, the processing method of the present invention can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. The multimedia data with a data structure according to the present invention can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices in which computer-readable data is stored. Computer-readable recording media may include, for example, BD, Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmission via the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired / wireless communication network.

[0459] In addition, the embodiments of the present disclosure may be implemented using a computer program product according to a program code, and the program code may be executed in a computer through the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.

[0460] Figure 27 The present invention is a structural diagram of a content streaming system to which the present disclosure is applied.

[0461] A content streaming system to which the embodiments of this document are applied may mainly include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.

[0462] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, or camcorder directly generates a bitstream, the encoding server can be omitted.

[0463] A bitstream may be generated by an encoding method or a bitstream generating method to which an embodiment of the present disclosure is applied, and a streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0464] The streaming server transmits multimedia data to user devices via a network server based on user requests, and the network server serves as an intermediary for notifying users of services. When a user requests a desired service from the network server, the network server delivers the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands and responses between devices within the content streaming system.

[0465] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a stable streaming service, the streaming server can store the bitstream for a predetermined period of time.

[0466] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, and head-mounted displays), digital TVs, desktop computers, and digital signage. Each server within the content streaming system may operate as a distributed server, in which case data received from each server may be distributed.

[0467] The claims described in this disclosure can be combined in various ways. For example, the technical features of the method claims of this disclosure can be combined to implement an apparatus, and the technical features of the apparatus claims of this disclosure can be combined to implement a method. Furthermore, the technical features of the method claims of this disclosure can be combined with the technical features of the apparatus claims of this disclosure to implement an apparatus, and the technical features of the method claims of this disclosure can be combined with the technical features of the apparatus claims of this disclosure to implement a method.

Claims

1. A method for decoding an image, the method comprising: Obtaining image information including residual information and prediction mode information of a current block through a bitstream; deriving a prediction mode of the current block based on the prediction mode information; deriving a prediction sample of the current block based on the prediction mode; deriving residual samples of the current block based on the residual information; as well as generating a reconstructed sample of the current block based on the predicted sample and the residual sample, The residual information includes syntax elements for the first transform coefficient to the nth transform coefficient of the current block and syntax elements for the n+1th transform coefficient to the last transform coefficient, The number of syntax elements used for context coding of the first transform coefficient to the nth transform coefficient is equal to the maximum number of bins of context coding of the current block, The maximum number of bins for the context coding of the current block is derived based on the width of the current block and the height of the current block, wherein the syntax elements for the first transform coefficient to the nth transform coefficient are syntax elements of a first residual data coding structure according to transform skip residual coding (TSRC), wherein the syntax elements for the n+1th transform coefficient to the last transform coefficient are syntax elements of the second residual data coding structure according to the TSRC, The syntax elements according to the second residual data coding structure include sign flags and coefficient level information for transform coefficients, wherein the syntax elements according to the first residual data coding structure are context-coded, wherein the syntax elements according to the second residual data coding structure are bypass coded, and Wherein, after all syntax elements of previous transform coefficients of the target transform coefficient of the current block are coded and before the syntax elements of the target transform coefficient are coded, it is determined whether the number of context coded bins used is equal to the maximum number of context coded bins.

2. A method for encoding an image, the method comprising: Based on inter-frame prediction or intra-frame prediction, derive prediction samples of the current block; deriving residual samples of the current block based on the predicted samples; deriving transform coefficients of the current block based on the residual samples; as well as encoding image information including prediction mode information of the current block and residual information for the transform coefficients, The residual information includes syntax elements for the first transform coefficient to the nth transform coefficient of the current block and syntax elements for the n+1th transform coefficient to the last transform coefficient, The number of syntax elements used for context coding of the first transform coefficient to the nth transform coefficient is equal to the maximum number of bins of context coding of the current block, The maximum number of bins for the context coding of the current block is derived based on the width of the current block and the height of the current block, wherein the syntax elements for the first transform coefficient to the nth transform coefficient are syntax elements of a first residual data coding structure according to transform skip residual coding (TSRC), wherein the syntax elements for the n+1th transform coefficient to the last transform coefficient are syntax elements of the second residual data coding structure according to the TSRC, The syntax elements according to the second residual data coding structure include sign flags and coefficient level information for transform coefficients, wherein the syntax elements according to the first residual data coding structure are context-coded, wherein the syntax elements according to the second residual data coding structure are bypass coded, and Wherein, after all syntax elements of previous transform coefficients of the target transform coefficient of the current block are coded and before the syntax elements of the target transform coefficient are coded, it is determined whether the number of context coded bins used is equal to the maximum number of context coded bins.

3. A method for transmitting image data, the method comprising: Obtaining a bitstream of image information including prediction mode information of a current block and residual information of transform coefficients for the current block; as well as sending data of the bitstream including the image information, wherein the image information includes the prediction mode information and the residual information, wherein the transform coefficients are derived based on residual samples of the current block, and the residual samples are derived based on predicted samples of the current block, wherein the prediction sample is derived based on inter-frame prediction or intra-frame prediction, and the prediction mode information indicates a prediction mode for the prediction sample, The residual information includes syntax elements for the first transform coefficient to the nth transform coefficient of the current block and syntax elements for the n+1th transform coefficient to the last transform coefficient, The number of syntax elements used for context coding of the first transform coefficient to the nth transform coefficient is equal to the maximum number of bins of context coding of the current block, The maximum number of bins for the context coding of the current block is derived based on the width of the current block and the height of the current block, wherein the syntax elements for the first transform coefficient to the nth transform coefficient are syntax elements of a first residual data coding structure according to transform skip residual coding (TSRC), wherein the syntax elements for the n+1th transform coefficient to the last transform coefficient are syntax elements of the second residual data coding structure according to the TSRC, The syntax elements according to the second residual data coding structure include sign flags and coefficient level information for transform coefficients, wherein the syntax elements according to the first residual data coding structure are context-coded, wherein the syntax elements according to the second residual data coding structure are bypass coded, and Wherein, after all syntax elements of previous transform coefficients of the target transform coefficient of the current block are coded and before the syntax elements of the target transform coefficient are coded, it is determined whether the number of context coded bins used is equal to the maximum number of context coded bins.