Image decoding method associated with residual coding, and device therefor
The method improves image coding efficiency by utilizing prediction-related information, dependent quantization flags, and TSRC flags in the image decoding process, effectively reducing transmission and storage costs for high-resolution images.
Patent Information
- Application Number
- JP2025047087
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-02-05
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2041-02-05
AI Technical Summary
The increasing demand for high-resolution and high-quality images has led to a need for more efficient image coding technologies to reduce transmission and storage costs.
The proposed method involves an image decoding technique that uses prediction-related information, dependent quantization flags, and Transform Skip Residual Coding (TSRC) flags to efficiently encode and decode image blocks, thereby improving coding efficiency.
This approach enhances the efficiency of image coding and residual coding, reducing the amount of bits required for transmission and storage while maintaining image quality.
Smart Images

Figure 2025089406000001_ABST
Abstract
Description
Technical Field
[0001] This document relates to image coding technology, and more particularly, to an image decoding method and apparatus for coding flag information indicating whether TSRC is enabled when coding residual data of a current block in an image coding system.
Background Art
[0002] Recently, the demand for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to the existing image data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image data using an existing recording medium, the transmission cost and storage cost increase.
[0003] Accordingly, in order to effectively transmit, store, and reproduce information of high-resolution and high-quality images, a highly efficient image compression technology is required.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The technical problem of this document is to provide a method and apparatus for increasing image coding efficiency.
[0005] Another technical problem of this document is to provide a method and apparatus for increasing the efficiency of residual coding.
Means for Solving the Problems
[0006] According to one embodiment of the present document, an image decoding method executed by a decoding device is provided. The method includes steps of: obtaining prediction-related information for a current block; obtaining a dependent quantization enable flag regarding whether dependent quantization is enabled; obtaining a TSRC enable flag regarding whether Transform Skip Residual Coding (TSRC) is enabled based on the dependent quantization enable flag; obtaining residual information of a syntax of residual coding for the current block derived based on the TSRC enable flag; deriving motion information of the current block based on the prediction-related information; deriving a prediction sample of the current block based on the motion information; deriving a residual sample of the current block based on the residual information; and generating a reconstructed picture based on the prediction sample and the residual sample.
[0007] According to another embodiment of the present document, a decoding device that executes image decoding is provided. The decoding device acquires prediction-related information for a current block, acquires a dependent quantization available flag regarding whether dependent quantization can be used, acquires a TSRC available flag regarding whether Transform Skip Residual Coding (TSRC) can be used based on the dependent quantization available flag, and an entropy decoding (decoding) unit that acquires residual information of the syntax of the residual coding for the current block derived based on the TSRC available flag, a prediction unit that derives motion information of the current block based on the prediction-related information and derives a prediction sample of the current block based on the motion information, a residual processing unit that derives a residual sample of the current block based on the residual information, and an addition unit that generates a restored picture based on the prediction sample and the residual sample.
[0008] According to still another embodiment of the present document, there is provided a video encoding method executed by an encoding apparatus. The method includes steps of deriving a prediction sample of a current block based on inter prediction, encoding (encoding) prediction-related information for the current block, encoding a dependent quantization available flag regarding whether dependent quantization can be used, encoding a TSRC available flag regarding whether Transform Skip Residual Coding (TSRC) can be used based on the dependent quantization available flag, determining a syntax of residual coding for the current block based on the TSRC available flag, encoding residual information of the determined syntax of residual coding for the current block, and generating a bitstream having the prediction-related information, the dependent quantization available flag, the TSRC available flag, and the residual information.
[0009] According to yet another embodiment of the present document, a video encoding apparatus is provided. The encoding apparatus includes a prediction unit that derives prediction samples of a current block based on inter prediction, encodes prediction-related information for the current block, encodes a dependent quantization available flag indicating whether dependent quantization can be used, encodes a TSRC available flag indicating whether Transform Skip Residual Coding (TSRC) can be used based on the dependent quantization available flag, determines a syntax of residual coding for the current block based on the TSRC available flag, encodes residual information of the determined syntax of residual coding for the current block, and an entropy encoding (encoding) unit that generates a bitstream having the prediction-related information, the dependent quantization available flag, the TSRC available flag, and the residual information.
[0010] According to yet another embodiment of the present document, there is provided a computer-readable digital storage medium storing a bitstream having image information that causes an image decoding method to be performed. The computer-readable digital storage medium, wherein the image decoding method includes: obtaining prediction-related information for a current block; obtaining a dependent quantization available flag regarding whether dependent quantization is available; obtaining a TSRC available flag regarding whether Transform Skip Residual Coding (TSRC) is available based on the dependent quantization available flag; obtaining residual information of the syntax of the residual coding for the current block derived based on the TSRC available flag; deriving motion information of the current block based on the prediction-related information; deriving a prediction sample of the current block based on the motion information; deriving a residual sample of the current block based on the residual information; and generating a reconstructed picture based on the prediction sample and the residual sample.
Advantages of the Invention
[0011] According to the present document, the efficiency of image coding performed based on inter prediction and / or residual coding can be improved.
[0012] According to the present document, the efficiency of residual coding can be increased.
[0013] According to this document, a signaling relationship between a dependent quantization available flag and a TSRC available flag is set, so that when dependent quantization is not available, the TSRC available flag can be signaled. Through this, when TSRC is not available and RRC syntax is coded for a conversion skip block, dependent quantization is not used, thereby improving coding efficiency, reducing the amount of bits to be coded, and improving the overall residual coding efficiency.
[0014] According to this document, the TSRC available flag can be signaled only when dependent quantization is not used. Through this, it is ensured that coding RRC syntax for a conversion skip block and using dependent quantization do not overlap and are not executed repeatedly, so that the TSRC available flag can be coded more effectively (efficiently), reducing the amount of bits and improving the overall residual coding efficiency.
Brief Description of Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Embodiments for Carrying Out the Invention
[0016] This document can be modified in various ways, can have various embodiments, and specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are merely used to describe specific embodiments and are not used with the intention of limiting the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the existence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof, etc. is not precluded in advance.
[0017] On the other hand, each configuration on the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each configuration is realized by separate hardware or separate software. For example, among each configuration, two or more configurations can be combined to form one configuration, and one configuration can also be divided into a plurality of configurations. Embodiments in which each configuration is integrated and / or separated are included in the scope of rights of this document as long as they do not deviate from the essence of this document.
[0018] Hereinafter, with reference to the accompanying drawings, preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals will be used for the same components on the drawings, and overlapping descriptions for the same components can be omitted.
[0019] FIG. 1 schematically shows an example of a video / image coding system to which an embodiment of this document can be applied.
[0020] As shown in FIG. 1, a video / image coding system can include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device via a digital recording medium or a network in file or streaming form.
[0021] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be referred to as a video / image encoding device, and the decoding device can be referred to as a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can include a display unit, and the display unit can also be composed of a separate device or an external component.
[0022] The video source can obtain video / images through processes such as video / image capture, synthesis, or generation (processing, process). The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer or the like, and in this case, the video / image capture process can be replaced during the process of generating related data.
[0023] The encoding device can encode input video / images. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0024] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital recording medium or a network in file or streaming form. The digital recording medium can include various recording media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the above bitstream and transmit it to the decoding device.
[0025] The decoding device can decode a video / image by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0026] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0027] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the VVC (Versatile Video Coding) standard, EVC (Essential Video Coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd Generation of Audio Video Coding Standard), or next-generation video / image coding standards (such as H.267 or H.268, etc.).
[0028] This document presents various embodiments related to video / image coding, and unless otherwise noted, the above embodiments can also be executed in combination with each other.
[0029] In this document, video may mean a collection of a series of images over time. Picture generally means a unit indicating one image in a specific time period, and subpicture / slice / tile is a unit that constitutes a part of a picture in coding. A subpicture / slice / tile may include one or more CTUs (Coding Tree Units). One picture may be composed of one or more subpictures / slices / tiles. One picture may be composed of one or more groups of tiles. One tile group may include one or more tiles. A brick represents a rectangular region of CTU rows within a tile in a picture. A tile may be partitioned into multiple bricks, each of which consisting of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may be also referred to as a brick.A brick scan shows a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a brick, bricks within a tile are ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. Also, a subpicture may represent a rectangular region of one or more slices within a picture. That is, a subpicture contains one or more slices that collectively cover a rectangular region of a picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture.The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture. A tile scan is a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture.A slice includes an integer number of bricks of a picture that maybe exclusively contained in a single NAL unit. A slice may consists of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile. In this document, tile group and slice may be used interchangeably. For example, in this document, tile group / tile group header may be called slice / slice header.
[0030] A pixel or pel can mean the smallest unit that makes up a picture (or image). Also, the term "sample" can be used as the term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and can also indicate only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.
[0031] A unit can indicate the basic unit of image processing. A unit can include at least one of a specific area of a picture and information related to the area. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can include a sample (or, sample array) consisting of M columns and N rows, or a set (or, array) of transform coefficients.
[0032] As used herein, "A or B" can mean "only A", "only B", or "both A and B". In other words, as used herein, "A or B" can be understood as "A and / or B". For example, as used herein, "A, B or C" can mean "only A", "only B", "only C", or "any combination of A, B and C".
[0033] The slashes ( / ) and commas used herein can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B or C".
[0034] In this specification, "(at least one of) A and B" can mean "only A", "only B", or "both A and B". Also, in this specification, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".
[0035] Also, in this specification, "(at least one of) A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Also, "(at least one of) A, B, or C" and "(at least one of) A, B, and / or C" can mean "(at least one of) A, B, and C".
[0036] Also, the parentheses used in this specification can mean "for example". Specifically, when it is shown as "prediction (intra prediction)", "intra prediction" can be proposed as an example of "prediction". In other words, "prediction" in this specification is not limited to "intra prediction", and "intra prediction" can be proposed as an example of "prediction". Also, when it is shown as "prediction (that is, intra prediction)", "intra prediction" can be proposed as an example of "prediction".
[0037] The technical features separately described within one drawing in this specification may be realized separately or simultaneously.
[0038] The following drawings are created to illustrate a specific example of this specification. Since the names of specific devices and the names of specific signals / messages / fields described in the drawings are presented exemplarily, the technical features of this specification are not limited to the specific names used in the following drawings.
[0039] FIG. 2 is a diagram schematically explaining the configuration of a video / image encoding device to which an embodiment of this document can be applied. Hereinafter, the video encoding device can include an image encoding device.
[0040] As shown in FIG. 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (Decoded Picture Buffer) and can be configured by a digital recording medium. The above hardware components can further include the memory 270 as an internal / external component.
[0041] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-Tree Binary-Tree Ternary-Tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quadtree (quad-tree) structure, a binary tree (binary-tree) structure, and / or a ternary tree (ternary) structure. In this case, for example, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied later. Alternatively, the binary tree structure can also be applied first. The coding procedure according to this document can be executed based on the final coding unit that cannot be further divided. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be immediately used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth, and the coding unit with an optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transformation unit (TU: Transform Unit). In this case, the prediction unit and the transformation unit can each be divided or partitioned from the aforementioned final coding unit.The prediction unit is a unit of sample prediction, and the conversion unit is a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.
[0042] The term "unit" can, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or can also represent only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel for one picture (or image).
[0043] The encoding device 200 can subtract a prediction signal (predicted block, predicted sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoder 200 can be called the subtraction unit 231. The prediction unit can perform a prediction on a block to be processed (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit them to the entropy encoding unit 240 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0044] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block according to the prediction mode, or can also be located remotely. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of fineness of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used according to the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.
[0045] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The above motion information can include a motion vector and a reference picture index. The above motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the above reference block and the reference picture including the above temporal neighboring block may be the same or different. The above temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (colCU), and the reference picture including the above temporal neighboring block can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidates are used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (Motion Vector Prediction, MVP) mode, the motion vector of an adjacent block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0046] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, for the prediction of one block, the prediction unit can not only apply intra prediction or inter prediction, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the Intra Block Copy (IBC) prediction mode or the palette mode for the prediction of a block. The above IBC prediction mode or palette mode can be used for content image / video coding such as games, for example, like SCC (Screen Content Coding). IBC basically performs prediction within the current picture, but can be executed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information regarding the palette table and palette index.
[0047] The prediction signal generated through the above prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when expressing the relationship information between pixels in a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square, and can also be applied to a non-square, variable-size block.
[0048] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 233 can reorder the block-form quantized transform coefficients in a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can execute various encoding methods such as, for example, exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), CABAC (Context-Adaptive Binary Arithmetic Coding). In addition to the quantized transform coefficients, the entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (such as the values of syntax elements). The encoded information (such as the encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (Network Abstraction Layer) units. The video / image information can further include information regarding various parameter sets such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. Information and / or syntax elements transmitted / signaled from the encoding device to the decoding device in this document can be included in the video / image information. The video / image information can be encoded through the above-described encoding procedure and included in the above bitstream.The above bitstream can be transmitted via a network or stored in a digital recording medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital recording medium can include various recording media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 240 can be configured as internal / external elements of the encoding device 200, or the transmitting unit can also be included in the entropy encoding unit 240.
[0049] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be restored by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.
[0050] On the other hand, LMCS (Luma Mapping with Chroma Scaling) can also be applied in the picture encoding and / or restoration process.
[0051] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0052] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. Through this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device 300 when inter prediction is applied, and can also improve the encoding efficiency.
[0053] The memory 270 DPB can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the block where the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 222.
[0054] FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding apparatus to which an embodiment of the present document can be applied.
[0055] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an inter-predictor 331 and an intra-predictor 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 322. The entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above can be configured by one hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Also, the memory 360 can include a DPB (Decoded Picture Buffer) and can also be configured by a digital recording medium. The above hardware component can further include the memory 360 as an internal / external component.
[0056] When a bitstream including video / image information is input, the decoding device 300 can restore an image corresponding to the process in which the video / image information was processed by the encoding device in FIG. 2. For example, the decoding device 300 can derive units / blocks based on the block division related information obtained from the above bitstream. The decoding device 300 can execute decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding is, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding device 300 can be reproduced via a playback device.
[0057] The decoding device 300 can receive the signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the above bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The above video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the above video / image information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the above parameter set and / or the above general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the above decoding procedure and obtained from the above bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values of syntax elements necessary for image restoration, quantized values of transform coefficients regarding residuals, and the like. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, as well as decoding information of surrounding and blocks to be decoded, or information of symbols / bins decoded in previous steps, predicts the occurrence probability of a bin based on the determined context model, and can execute arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction units (inter prediction unit 332 and intra prediction unit 331), and the residual values for which entropy decoding is performed by the entropy decoding unit 310, that is, the quantized transform coefficients and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit is a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / image / picture decoding device, and the above decoding device can also be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The above information decoder can include the above entropy decoding unit 310, and the above sample decoder can include at least one of the above inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.
[0058] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the above reordering can be performed based on the coefficient scan order executed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (for example, quantization step size information) to obtain transform coefficients.
[0059] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).
[0060] The prediction unit can perform a prediction on the current block and generate a predicted block that includes a prediction sample for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.
[0061] The prediction unit 320 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the Intra Block Copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (Screen Content Coding). IBC basically performs prediction within the current picture, but can be executed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be included in and signaled in the video / image information.
[0062] The Intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block according to the prediction mode, or can be located remotely. In intra prediction, the prediction mode can include a plurality of non - directional modes and a plurality of directional modes. The Intra prediction unit 331 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.
[0063] The Inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted from the inter prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between the adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the Inter prediction unit 332 can construct a motion information candidate list based on the adjacent blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be executed based on various prediction modes, and the information related to the prediction can include information indicating the mode of inter prediction for the current block.
[0064] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block can be used as the restored block.
[0065] The adder 340 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next block to be processed in the current picture, can be output after filtering as described later, or can also be used for inter prediction of the next picture.
[0066] On the other hand, LMCS (Luma Mapping with Chroma Scaling) can also be applied in the picture decoding process.
[0067] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0068] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the block from which the motion information in the current picture was derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 331.
[0069] In this specification, the embodiments described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding device 200 can be applied in the same or corresponding manner to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding device 300, respectively.
[0070] In this document, at least one of quantization / inverse quantization and / or transform / inverse transform can be omitted. When the quantization / inverse quantization is omitted, the quantized transform coefficients can be called transform coefficients. When the transform / inverse transform is omitted, the transform coefficients can be called coefficients or residual coefficients, or can still be called transform coefficients for the sake of consistency of expression.
[0071] In this document, the quantized transform coefficients and the transform coefficients can each be referred to as a transform coefficient and a scaled transform coefficient, respectively. In this case, the residual information can include information regarding the transform coefficient(s), and the information regarding the transform coefficient(s) can be signaled via a residual coding syntax. The transform coefficient can be derived based on the residual information (or the information regarding the transform coefficient(s)), and the scaled transform coefficient can be derived via an inverse transform (scaling) with respect to the transform coefficient. The residual sample can be derived based on an inverse transform (transformation) with respect to the scaled transform coefficient. This can be similarly applied / expressed in other parts of this document.
[0072] As described above, in performing video coding, prediction is performed to increase the compression efficiency. Through this, a predicted block including prediction samples for the current block, which is a block to be coded, can be generated. Here, the predicted block includes prediction samples in a spatial domain (domain) (or pixel domain). The predicted block is derived identically in an encoding device and a decoding device, and the encoding device can increase the image coding efficiency by signaling information (residual information) regarding a residual between the original block and the predicted block, which is not the original sample value of the original block, to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, and can generate a restored block including restored samples by combining the residual block and the predicted block, and can generate a restored picture including the restored block.
[0073] The residual information can be generated through the conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, perform a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, perform a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and signal the relevant residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can perform an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. The encoding device can further inverse quantize / inverse transform the quantized conversion coefficients to derive a residual block for reference in the inter-prediction of subsequent pictures, and generate a restored picture based on this.
[0074] When inter prediction is applied, the prediction unit of the encoding device / decoding device can derive prediction samples by performing inter prediction in block units. Inter prediction can be a prediction derived in a manner that is dependent on data elements (ex., sample values or motion information) of picture(s) other than the current picture (Inter prediction can be a prediction derived in a manner that is dependent on data elements(ex., sample values or motion information) of picture(s) other than the current picture). When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The above motion information can include a motion vector and a reference picture index. The above motion information can further include inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When inter prediction is applied, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the above reference block and the reference picture including the above temporal neighboring block can be the same or different. The above temporal neighboring block can be called by names such as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the above temporal neighboring block can also be called a collocated picture (colPic).For example, a motion information candidate list can be configured based on neighboring blocks of a current block, and in order to derive the motion vector and / or reference picture index of the current block, flag or index information indicating which candidate is selected (used) can be signaled. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the motion information of the current block can be the same as that of the selected neighboring block. In the case of skip mode, unlike merge mode, the residual signal can be not transmitted. In the case of Motion Vector Prediction (MVP) mode, the motion vector of the selected neighboring block is used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived by using the sum of the motion vector predictor and the motion vector difference.
[0075] The above motion information can include L0 motion information and / or L1 motion information depending on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). The motion vector in the L0 direction can be called the L0 motion vector or MVL0, and the motion vector in the L1 direction can be called the L1 motion vector or MVL1. The prediction based on the L0 motion vector can be called L0 prediction, the prediction based on the L1 motion vector can be called L1 prediction, and the prediction based on both the above L0 motion vector and the above L1 motion vector can be called the pair (Bi) prediction. Here, the L0 motion vector can represent the motion vector associated (correlated) with the reference picture list L0 (L0), and the L1 motion vector can represent the motion vector associated with the reference picture list L1 (L1). The reference picture list L0 can include, as reference pictures, pictures that are earlier (previous) in output order than the above current picture, and the reference picture list L1 can include pictures that are later in output order than the above current picture. The above previous picture can be called the forward (reference) picture, and the above later picture can be called the backward (reference) picture. The reference picture list L0 can further include, as reference pictures, pictures that are later in output order than the above current picture. In this case, the above previous picture can be indexed first within the reference picture list L0, and the above later picture can be indexed next. The reference picture list L1 can further include, as reference pictures, pictures that are earlier in output order than the above current picture. In this case, the above later picture can be indexed first within the reference picture list 1 (L1), and the above previous picture can be indexed next. Here, the output order can correspond to the POC (Picture Order Count) order.
[0076] The video / image encoding procedure based on inter prediction can generally include, for example, the following.
[0077] FIG. 4 shows an example of an inter-prediction based video / image encoding method.
[0078] The encoding device performs inter-prediction on the current block (S400). The encoding device can derive the inter-prediction mode and motion information of the current block and generate a prediction sample of the current block. Here, the inter-prediction mode determination, motion information derivation, and prediction sample generation procedures can be performed simultaneously, or any one of the procedures can be performed prior to the other procedures. For example, the inter-prediction unit of the encoding device can include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit determines the prediction mode for the current block, the motion information derivation unit derives the motion information of the current block, and the prediction sample derivation unit can derive the prediction sample of the current block. For example, the inter-prediction unit of the encoding device can search for a block similar to the current block within a certain region (search region) of the reference picture via motion estimation, and derive a reference block whose difference from the current block is the minimum or below a certain criterion. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine the mode to be applied to the current block among various prediction modes. The encoding device can compare the RD cost for the various prediction modes and determine the optimal prediction mode for the current block.
[0079] For example, when the skip mode or the merge mode is applied to the current block, the encoding device constructs a merge candidate list described later, and among the reference blocks pointed to by the merge candidates included in the merge candidate list, the current block and a reference block whose difference from the current block is the smallest or below a certain criterion can be derived. In this case, a merge candidate associated with the derived reference block is selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.
[0080] As another example, when the (A)MVP mode is applied to the current block, the encoding device constructs an (A)MVP candidate list described later, and among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list, the motion vector of the selected mvp candidate can be used as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the above-described motion estimation can be used as the motion vector of the current block, and among the mvp candidates, the mvp candidate having the motion vector with the smallest difference from the motion vector of the current block can be the selected mvp candidate. An MVD (Motion Vector Difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In this case, information regarding the MVD can be signaled to the decoding device. Also, when the (A)MVP mode is applied, the value of the reference picture index can be configured with reference picture index information and separately signaled to the decoding device.
[0081] The encoding device can derive a residual sample based on the prediction sample (S410). The encoding device can derive the residual sample by comparing the original sample of the current block with the prediction sample.
[0082] The encoding device encodes image information including prediction information and residual information (S420). The encoding device can output the encoded image information in the form of a bitstream. The prediction information can include prediction mode information (e.g., skip flag, merge flag or mode index, etc.) and information related to motion information as information related to the prediction procedure. The information related to the motion information can include candidate selection information (e.g., merge index, mvp flag or mvp index) which is information for deriving a motion vector. Also, the information related to the motion information can include information related to the above-mentioned MVD and / or reference picture index information. Also, the information related to the motion information can include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information related to the residual samples. The residual information can include information related to the quantized transform coefficients for the residual samples.
[0083] The output bitstream can be stored in a (digital) recording medium and transmitted to the decoding device, or can also be transmitted to the decoding device via a network.
[0084] On the other hand, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is to derive the same prediction result in the encoding device as that performed in the decoding device, and through this, the coding efficiency can be improved. Therefore, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in the memory and utilize it as a reference picture for inter prediction. As described above, further in-loop filtering procedures and the like can be applied to the reconstructed picture.
[0085] The video / image decoding procedure based on inter prediction can generally include, for example, the following.
[0086] FIG. 5 shows an example of an inter-prediction based video / image decoding method.
[0087] As shown in FIG. 5, the decoding device can perform operations corresponding to the operations performed by the encoding device. The decoding device can perform prediction on the current block based on the received prediction information and derive a prediction sample.
[0088] Specifically, the decoding device can determine a prediction mode for the current block based on the received prediction information (S500). The decoding device can determine which inter-prediction mode is applied to the current block based on the prediction mode information in the prediction information.
[0089] For example, based on the merge flag, it can be determined whether the merge mode is applied to the current block or whether the (A)MVP mode is determined. Alternatively, based on the mode index, one of various inter-prediction mode candidates can be selected. The inter-prediction mode candidates can include a skip mode, a merge mode, and / or the (A)MVP mode, or can include various inter-prediction modes described later.
[0090] The decoding device derives motion information of the current block based on the determined inter-prediction mode (S510). For example, when the skip mode or the merge mode is applied to the current block, the decoding device constructs a merge candidate list described later and can select one merge candidate from among the merge candidates included in the merge candidate list. The selection can be performed based on the selection information (merge index) described above. The motion information of the selected merge candidate can be used to derive the motion information of the current block. The motion information of the selected merge candidate can be used as the motion information of the current block.
[0091] As another example, when the (A)MVP mode is applied to the current block, the decoding device constructs an (A)MVP candidate list described below, and among the mvp (motion vector predictor) candidates included in the (A)MVP candidate list, the motion vector of the selected mvp candidate can be used as the mvp of the current block. The above selection can be performed based on the above-described selection information (mvp flag or mvp index). In this case, the MVD of the current block can be derived based on the information regarding the MVD, and the motion vector of the current block can be derived based on the mvp of the current block and the MVD. Also, the reference picture index of the current block can be derived based on the reference picture index information. The picture pointed to by the reference picture index within the reference picture list regarding the current block can be derived as the reference picture to be referred to for the inter prediction of the current block.
[0092] On the other hand, as will be described later, the motion information of the current block can be derived without constructing a candidate list, and in this case, the motion information of the current block can be derived by the procedure disclosed in the prediction mode described later. In this case, the candidate list construction as described above can be omitted.
[0093] The decoding device can generate a prediction sample for the current block based on the motion information of the current block (S520). In this case, the reference picture is derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived using the sample of the reference block pointed to by the motion vector of the current block on the reference picture. In this case, as will be described later, in some cases, a prediction sample filtering procedure can be further performed on all or part of the prediction samples of the current block.
[0094] For example, the inter prediction unit of the decoding device can include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. Based on the prediction mode information received by the prediction mode determination unit, it determines the prediction mode for the current block, and based on the information related to the motion information received by the motion information derivation unit, it derives the motion information (such as motion vectors and / or reference picture indexes, etc.) of the current block. The prediction sample derivation unit can derive the prediction sample of the current block.
[0095] The decoding device generates a residual sample for the current block based on the received residual information (S530). The decoding device can generate a restored sample for the current block based on the prediction sample and the residual sample, and generate a restored picture based on this (S540). As described above, an in-loop filtering procedure or the like can be further applied to the restored picture.
[0096] FIG. 6 exemplarily shows the inter prediction procedure.
[0097] Referring to FIG. 6, as described above, the inter prediction procedure can include an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction execution (prediction sample generation) step based on the derived motion information. As described above, the inter prediction procedure can be performed by the encoding device and the decoding device. In this document, the coding device can include the encoding device and / or the decoding device.
[0098] As shown in FIG. 6, the coding device determines an inter prediction mode for the current block (S600). For the prediction of the current block within a picture, various inter prediction modes can be used. For example, various modes such as a merge mode, a skip mode, an MVP (Motion Vector Prediction) mode, an Affine mode, a sub-block merge mode, an MMVD (Merge with MVD) mode, etc. can be used. A DMVR (Decoder side Motion Vector Refinement) mode, an AMVR (Adaptive Motion Vector Resolution) mode, Bi-prediction with CU-level Weight (BCW), Bi-Directional Optical Flow (BDOF), etc. can be used additionally or alternatively as accompanying modes. The Affine mode can also be called an affine motion prediction mode. The MVP mode can also be called an AMVP (Advanced Motion Vector Prediction) mode. In this document, some modes and / or motion information candidates derived by some modes can also be included as one of the motion information related candidates of other modes. For example, an HMVP candidate can be added as a merge candidate of the above merge / skip mode, or can be added as an mvp candidate of the above MVP mode. When the above HMVP candidate is used as a motion information candidate of the above merge mode or skip mode, the above HMVP candidate can be called an HMVP merge candidate.
[0099] Prediction mode information indicating an inter prediction mode of a current block can be signaled from an encoding device to a decoding device. The prediction mode information can be included in a bitstream and received by the decoding device. The prediction mode information can include index information indicating one of a plurality of candidate modes. Alternatively, the inter prediction mode can also be indicated via hierarchical signaling of flag information. In this case, the prediction mode information can include one or more flags. For example, a skip flag is signaled to indicate whether skip mode application is possible, and when skip mode is not applied, a merge flag is signaled to indicate whether merge mode application is possible. When merge mode is not applied, it is also possible to signal whether MVP mode is applied or to further signal a flag for additional classification. The affine mode can be signaled as an independent mode, or can also be signaled as a mode subordinate to a merge mode or an MVP mode, etc. For example, the affine mode can include an affine merge mode and an affine MVP mode.
[0100] The coding device derives motion information for the current block (S610). The motion information derivation can be derived based on the inter prediction mode.
[0101] The coding device can perform inter prediction using the motion information of the current block. The encoding device can derive optimal motion information for the current block through a motion estimation procedure. For example, the encoding device can search for a highly correlated and similar reference block within a determined search range in the reference picture in fractional pixel units using the original block in the original picture for the current block, and derive motion information through this. The similarity of the blocks can be derived based on the difference in phase-based sample values. For example, the similarity of the blocks can be calculated based on the SAD between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, the motion information can be derived based on the reference block with the smallest SAD within the search space (search area). The derived motion information can be signaled to the decoding device in various ways based on the inter prediction mode.
[0102] The coding device performs inter prediction based on the motion information for the above-mentioned current block (S620). The coding device can derive prediction sample(s) for the current block based on the motion information. The current block including the prediction sample can be called a predicted block.
[0103] On the other hand, as described above, the encoding device can execute various encoding methods such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), CABAC (Context-Adaptive Binary Arithmetic Coding), etc. Also, the decoding device can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC (decode), and output values of syntax elements necessary for image restoration, quantized values of transform coefficients regarding residuals, etc.
[0104] For example, the coding method described above can be performed as described in the content to be described later.
[0105] FIG. 7 exemplarily shows CABAC (Context-Adaptive Binary Arithmetic Coding) for encoding a syntax element. For example, in the encoding process of CABAC, when the input signal is a syntax element that is not a binary value, the value of the input signal can be binarized to convert the input signal into a binary value. Also, when the input signal is already a binary value (i.e., when the value of the input signal is a binary value), binarization is not performed and it can be bypassed. Here, each binary number 0 or 1 that constitutes a binary value can be called a bin. For example, when the binary string after binarization is 110, each of 1, 1, and 0 is called one bin. The bin (one or more) for one syntax element can represent the value of the syntax element.
[0106] Thereafter, the binarized bins of the syntax element and the like can be input as a regular coding engine or a bypass coding engine. The regular coding engine of the encoding device can assign a context model that reflects a probability value to the bin, and can encode the bin based on the assigned context model. The regular coding engine of the encoding device can update the context model for the bin after performing encoding (encoding) for each bin. The bin encoded as described above can be represented as a context-coded bin.
[0107] On the one hand, when binary bins of the above syntax elements are input to the bypass encoding engine, they can be coded as follows. For example, the bypass encoding engine of the encoding device omits the procedure of estimating the probability for the input bin and the procedure of updating the probability model applied to the bin after encoding. When bypass encoding is applied, the encoding device can apply a uniform probability distribution instead of assigning a context model to encode the input bin, thereby improving the encoding speed. The bin encoded as described above can be represented as a bypass bin.
[0108] Entropy decoding can be represented as a process of performing the same process as the above-described entropy encoding in reverse order.
[0109] For example, when a syntax element is decoded based on a context model, the decoding device can receive the bin corresponding to the syntax element via a bitstream, and use the syntax element, the decoding information of the block to be decoded or the peripheral block, or the information of the symbol / bin decoded in the previous step to determine the context model. The generated probability of the received bin can be predicted by the determined context model, and the arithmetic decoding of the bin can be performed to derive the value of the syntax element. Thereafter, the context model of the bin to be decoded next can be updated in the determined context model.
[0110] Also, for example, when a syntax element is bypass decoded, the decoding device can receive a bin corresponding to the syntax element via a bitstream and can decode the input bin by applying a uniform probability distribution. In this case, the procedure for deriving the context model of the syntax element and the procedure for updating the context model applied to the bin after decoding can be omitted.
[0111] As described above, residual samples and the like can be derived as quantized transform coefficients and the like that have undergone the conversion and quantization processes. The quantized transform coefficients and the like can also be referred to as transform coefficients and the like. In this case, the transform coefficients and the like within a block can be signaled in the form of residual information. The residual information can include a residual coding syntax. That is, the encoding device can construct a residual coding syntax as residual information, encode this, and output it in the form of a bitstream, and the decoding device can decode the residual coding syntax from the bitstream and derive residual (quantized) transform coefficients and the like. The residual coding syntax can include syntax elements such as those representing whether a transform has been applied to the block, where the position of the last valid transform coefficient within the block is, whether there are valid transform coefficients within a sub-block, and what the magnitude / symbol of the valid transform coefficients are, as will be described later.
[0112] For example, the syntax elements related to the encoding / decoding of residual data can be represented as in the following table.
[0113]
Table 1-1
[0114]
Table 1-2
[0115]
Table 1-3
[0116] The transform_skip_flag indicates whether transformation is skipped in an associated block. The above transform_skip_flag can be a syntax element of the transform skip flag. The above associated block can be a CB (Coding Block) or a TB (Transform Block). For the transformation (and quantization) and residual coding procedures, the CB and the TB can be used interchangeably. For example, as described above, residual samples etc. are derived for a CB, and (quantized) transform coefficients etc. can be derived through transformation and quantization of the above residual samples etc., and information (e.g., syntax elements etc.) for efficiently representing the position, magnitude, sign, etc. of the above (quantized) transform coefficients etc. can be generated and signaled through the residual coding procedure. The quantized transform coefficients etc. can be simply called transform coefficients etc. Generally, when the CB is not larger than the maximum TB, the size of the CB can be the same as the size of the TB, and in this case, the block to be transformed (and quantized) and residual coded can be called a CB or a TB. On the other hand, when the CB is larger than the maximum TB, the block to be transformed (and quantized) and residual coded can be called a TB. Hereinafter, it will be described that syntax elements etc. related to residual coding are signaled in units of the transform block TB, but this is an example, and as described above, the above TB can be used interchangeably with the coding block CB.
[0117] On the one hand, the syntax element signaled after the above conversion skip flag is signaled is the same as the syntax element disclosed in Table 2 and / or Table 3 described below, and the specific description of the above syntax element is as described below.
[0118]
Table 2-1
[0119]
Table 2-2
[0120]
Table 2-3
[0121]
Table 2-4
[0122]
Table 2-5
[0123]
Table 2-6
[0124]
Table 3-1
[0125]
Table 3-2
[0126]
Table 3-3
[0127] According to this embodiment, as shown in Table 1, residual coding can be branched according to the value of the syntax element transform_skip_flag of the conversion skip flag. That is, different syntax elements can be used for residual coding based on the value of the conversion skip flag (based on whether conversion skip can be used). The residual coding used when conversion skip is not applied (i.e., when conversion is applied) can be called regular residual coding (RRC), and the residual coding when conversion skip is applied (i.e., when conversion is not applied) can be called transform skip residual coding (TSRC). Also, the above regular residual coding can also be called general residual coding. Also, the above regular residual coding can also be called general residual coding. Also, the above regular residual coding can be called the syntax structure of regular residual coding, and the above transform skip residual coding can be called the syntax structure of transform skip residual coding. Table 2 above can represent the syntax elements of residual coding when the value of transform_skip_flag is 0, that is, when conversion is applied, and Table 3 can represent the syntax elements of residual coding when the value of transform_skip_flag is 1, that is, when conversion is not applied.
[0128] Specifically, for example, a conversion skip flag indicating whether conversion skip of a conversion block is available can be parsed, and it can be determined whether the conversion skip flag is 1. When the value of the conversion skip flag is 0, as shown in Table 2, syntax elements last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix, sb_coded_flag, sig_coeff_flag, abs_level_gtx_flag, par_level_flag, abs_remainder, coeff_sign_flag and / or dec_abs_level regarding the residual coefficients of the conversion block can be parsed, and the residual coefficients can be derived based on the syntax elements. In this case, the syntax elements may be parsed sequentially, or the parsing order may be changed. Also, the abs_level_gtx_flag may represent abs_level_gt1_flag and / or abs_level_gt3_flag. For example, abs_level_gtx_flag[n][0] may be an example of the first conversion coefficient level flag (abs_level_gt1_flag), and the abs_level_gtx_flag[n][1] may be an example of the second conversion coefficient level flag (abs_level_gt3_flag).
[0129] Referring to Table 2 mentioned above, last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix, sb_coded_flag, sig_coeff_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag, abs_remainder, coeff_sign_flag, and / or dec_abs_level can be encoded / decoded. On the other hand, the sb_coded_flag may also be represented as coded_sub_block_flag.
[0130] In one embodiment, the encoding device can encode the (x, y) position information of the last non-zero transform coefficient in the transform block based on the syntax elements last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix. More specifically, the last_sig_coeff_x_prefix represents the prefix of the column position of the last significant coefficient in the scanning order within the transform block, the last_sig_coeff_y_prefix represents the prefix of the row position of the last significant coefficient in the scanning order within the transform block, the last_sig_coeff_x_suffix represents the suffix of the column position of the last significant coefficient in the scanning order within the transform block, and the last_sig_coeff_y_suffix represents the suffix of the row position of the last significant coefficient in the scanning order within the transform block. Here, the significant coefficient can represent the non-zero coefficient. Also, the scanning order can be the upper right diagonal scanning order. Alternatively, the scanning order can be the horizontal scanning order or the vertical scanning order. The scanning order can be determined based on whether intra / inter prediction is applied to the target block (CB or CB including TB) and / or the specific intra / inter prediction mode.
[0131] Next, after the encoding device divides the above conversion block into 4×4 sub-blocks (such as), for each 4×4 sub-block, it can use a 1-bit syntax element coded_sub_block_flag to indicate whether there are non-zero coefficients in the current sub-block.
[0132] If the value of coded_sub_block_flag is 0, since there is no more information to be transmitted, the encoding device can terminate the encoding process for the current sub-block. Conversely, if the value of coded_sub_block_flag is 1, the encoding device can continue the encoding process for sig_coeff_flag. The sub-block containing the last non-zero coefficient does not require encoding for coded_sub_block_flag, and the sub-block containing the DC information of the conversion block is likely to contain non-zero coefficients. Therefore, coded_sub_block_flag can be assumed to have a value of 1 without being encoded.
[0133] When it is determined that the value of coded_sub_block_flag is 1 and there are non-zero coefficients in the current sub-block, the encoding device can encode sig_coeff_flag having binary values according to the reverse scan order (the order scanned in reverse, reverse scanning order). The encoding device can encode the 1-bit syntax element sig_coeff_flag for each transform coefficient according to the scan order. If the value of the transform coefficient at the current scan position is not 0, the value of sig_coeff_flag can be 1. Here, in the case of a sub-block including the last non-zero coefficient, since sig_coeff_flag does not need to be encoded for the last non-zero coefficient, the encoding process for the above sub-block can be omitted. Level information encoding can be performed only when sig_coeff_flag is 1, and four syntax elements or the like can be used in the level information encoding process. More specifically, each sig_coeff_flag[xC][yC] can represent whether the level (value) of the transform coefficient at each transform coefficient position (xC, yC) in the current TB is non-zero. In one embodiment, the above sig_coeff_flag can correspond to an example of a syntax element of a valid coefficient flag indicating whether the quantized transform coefficient is a valid coefficient that is not 0.
[0134] The remaining level values after encoding for sig_coeff_flag can be derived as shown in the following formula. That is, the syntax element remAbsLevel representing the level value to be encoded can be derived as shown in the following formula.
[0135] <Equation 1>
Number
[0136] Here, coeff means the actual transform coefficient value.
[0137] Also, the abs_level_gt1_flag may indicate whether the remAbsLevel at the scan position (n) is greater than 1. For example, if the value of the abs_level_gt1_flag is 0, the absolute value of the conversion coefficient at that position may be 1. Also, if the value of the abs_level_gt1_flag is 1, the remAbsLevel representing the level value to be encoded hereinafter can be updated as shown in the following equation.
[0138] <Equation 2>
Number
[0139] Also, the least significant coefficient (LSB) value of the remAbsLevel described in the above Equation 2 can be encoded as shown in the following Equation 3 via the par_level_flag.
[0140] <Equation 3>
Number
[0141] Here, par_level_flag[n] may represent the parity of the conversion coefficient level (value) at the scan position n.
[0142] After encoding the par_leve_flag, the level value remAbsLevel of the conversion coefficient to be encoded can be updated as shown in the following equation.
[0143] <Equation 4>
Number
[0144] The abs_level_gt3_flag may indicate whether the remAbsLevel at the scan position (n) is greater than 3. Encoding for abs_remainder can be performed only when abs_level_gt3_flag is 1. The relationship between the actual conversion coefficient value coeff and each syntax element is as follows.
[0145] <Equation 5>
Number
[0146] Also, the following table shows an example related to Equation 5 above.
[0147]
Table 4
[0148] Here, |coeff| represents the level (value) of the conversion coefficient, and may also be displayed as AbsLevel for the conversion coefficient. Also, the sign of each coefficient can be encoded using the 1-bit symbol coeff_sign_flag.
[0149] Also, for example, when the value of the above conversion skip flag is 1, as shown in Table 3, the syntax elements sb_coded_flag, sig_coeff_flag, coeff_sign_flag, abs_level_gtx_flag, par_level_flag, and / or abs_remainder regarding the residual coefficients of the conversion block can be parsed, and the residual coefficients can be derived based on the above syntax elements. In this case, the above syntax elements may be parsed sequentially, or the parsing order may be changed. Also, the above abs_level_gtx_flag may represent abs_level_gt1_flag, abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and / or abs_level_gt9_flag. For example, abs_level_gtx_flag[n][j] may be a flag indicating whether the absolute value or level (value) of the conversion coefficient is greater than (j << 1) + 1 at the scan position n. The above (j << 1) + 1 may be replaced by a predetermined threshold value such as a first threshold (critical) value, a second threshold value, etc. in some cases.
[0150] On the one hand, CABAC provides high performance but has the drawback of poor throughput performance. This is due to the normalization encoding engine of CABAC. Normalization encoding (i.e., encoding through the normalization encoding engine of CABAC) shows high data dependency because it uses the probability state and range updated through the encoding of previous bins, and it can take a lot of time to read the probability interval to determine the current state. The throughput problem of CABAC can be solved by restricting the number of context-coded bins. For example, as shown in Table 2 above, the sum of the bins used to represent sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag can be restricted to a number based on the size of the block. Also, for example, as shown in Table 3 above, the sum of the bins used to represent sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and abs_level_gt9_flag can be restricted to a number based on the size of the block. As an example, when the block is a 4x4 size block, the sum of the bins for the above sig_coeff_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag or sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, abs_level_gt9_flag can be restricted to 32 (or, for example, 28). When the block is a 2x2 size block, the sum of the bins for the above sig_coeff_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag can be restricted to 8 (or, for example, 7).The limited number of the bins can be represented by remBinsPass1 or RemCcbs. Alternatively, as an example, for higher CABAC throughput, the number of context coded bins can be limited for a block (CB or TB) including the CG to be coded. In other words, the number of context coded bins can be limited in units of blocks (CB or TB). For example, if the size of the current block is 16x16, regardless of the current CG, the number of context coded bins for the current block can be limited to 1.75 times the number of pixels of the current block, i.e., 448.
[0151] In this case, when the encoding device uses all the limited number of context coded bins for encoding context elements, the remaining coefficients can be binary-coded through the binary-coding method for the coefficients described below without using context coding, and bypass coding can be performed. In other words, for example, when the number of context coded bins coded for a 4x4 CG is 32 (or, for example, 28), or when the number of context coded bins coded for a 2x2 CG is 8 (or, for example, 7), sig_coeff_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag that are to be coded with more context coded bins may not be coded and can be immediately coded with dec_abs_level. Alternatively, for example, when the number of context coded bins coded for a 4x4 block is limited to 1.75 times the number of pixels of the entire block, i.e., 28, sig_coeff_flag, abs_level_gt1_flag, par_level_flag, abs_level_gt3_flag that are to be coded with more context coded bins may not be coded and can be immediately coded with dec_abs_level as shown in Table 5 described below.
[0152]
Table 5
[0153] The |coeff| value can be derived based on the dec_abs_level. In this case, the |coeff|, which is the conversion coefficient value, can be derived as in the following formula.
[0154] <Equation 6>
Equation
[0155] Also, the above coeff_sign_flag indicates the sign of the conversion coefficient level at the scan position (n). That is, the above coeff_sign_flag indicates the sign of the conversion coefficient at the scan position (n).
[0156] FIG. 8 is a diagram showing an example of conversion coefficients within a 4×4 block.
[0157] The 4×4 block in FIG. 8 shows an example of quantized coefficients. The block shown in FIG. 8 can be a 4×4 conversion block or a 4×4 sub-block of an 8×8, 16×16, 32×32, or 64×64 conversion block. The 4×4 block in FIG. 8 represents a luma block or a chroma block.
[0158] On the one hand, as described above, when the input signal is a syntax element that is not binary, the encoding device can binarize the value of the input signal to convert the input signal into a binary value. Further, the decoding device can decode the syntax element to derive the binarized value of the syntax element (i.e., the binarized bin), and can inverse-binarize the binarized value to derive the value of the syntax element. The binarization process can be performed by, for example, a Truncated Rice (TR) binarization process, a k-th order Exp-Golomb (EGk) binarization process, a Limited k-th order Exp-Golomb (Limited EGk), or a Fixed-Length (FL) binarization process, which will be described later. Also, the inverse-binarization process can be performed based on the TR binarization process, the EGk binarization process, or the FL binarization process, and can represent the process of deriving the value of the syntax element.
[0159] For example, the TR binarization process can be performed as follows.
[0160] The input of the TR binarization process can be the requirements for TR binarization and cMax and cRiceParam for the syntax element. Also, the output of the TR binarization process can be the TR binarization for the value symbolVal corresponding to the bin string.
[0161] Specifically, as an example, when there is a suffix bit string for a syntax element, the TR bit string for the syntax element can be a concatenation of a prefix bit string and the suffix bit string, and when there is no such suffix bit string, the TR bit string for the syntax element can be the prefix bit string. For example, the prefix bit string can be derived as described later.
[0162] The prefix value of the symbolVal for the syntax element can be derived as in the following formula.
[0163] <Equation 7> [Number]
[0164] Here, prefixVal can represent the prefix value of the symbolVal. The prefix of the TR bit string of the syntax element (i.e., the prefix bit string) can be derived as described later.
[0165] For example, when the prefixVal is smaller than cMax >> cRiceParam, the prefix bit string can be a bit string of length prefixVal + 1 indexed by binIdx. That is, when the prefixVal is smaller than cMax >> cRiceParam, the prefix bit string can be a bit string of prefixVal + 1 bits pointed to by binIdx. The bins for binIdx smaller than prefixVal can be the same as 1. Also, the bin for binIdx equal to prefixVal can be the same as 0.
[0166] For example, the bin string derived by unary binarization for the above prefixVal can be as shown in the following table.
[0167] [Table 6]
[0168] On the other hand, when the above prefixVal is not smaller than cMax >> cRiceParam, the prefix bit string can be a bit string with a length of cMax >> cRiceParam and all bits being 1.
[0169] Also, when cMax is larger than symbolVal and cRiceParam is larger than 0, a suffix bit string of the TR bin string may exist. For example, the above suffix bit string can be derived as described later.
[0170] The suffix value of the above symbolVal for the above syntax element can be derived as in the following formula.
[0171] <Formula 8> [Number]
[0172] Here, suffixVal can represent the suffix value of the above symbolVal.
[0173] The suffix of the TR bin string (i.e., the suffix bit string) can be derived based on the FL binarization process for suffixVal where the cMax value is (1 << cRiceParam) - 1.
[0174] On the one hand, if the value of cRiceParam, which is an input parameter, is 0, the above TR binarization can be exactly the truncated unary binarization, and the same cMax value as the possible maximum value of the syntax element that is always decoded can be used.
[0175] Also, for example, the above EGk binarization process can be performed as follows. The syntax element coded by ue(v) can be the syntax element coded by Exp-Golomb coding.
[0176] As an example, the 0-th order Exp-Golomb (EG0) binarization process can be performed as follows.
[0177] The parsing process for the above syntax element can start from the current position of the bitstream, read the bits including the first non-zero bit, and start by counting the number of leading bits such as 0. The above process can be represented as shown in the following table.
[0178]
Table 7
[0179] Also, the variable codeNum can be derived as in the following formula.
[0180] <Formula 9>
Number
[0181] Here, the value returned by read_bits(leadingZeroBits), i.e., the value represented by read_bits(leadingZeroBits), can be interpreted as the binary representation of an unsigned integer for the most significant bit (the leftmost bit) recorded first.
[0182] The structure of the Exp - Golomb code that separates the bit string into "prefix" bits and "suffix" bits can be represented as in the following table.
[0183] [Table 8]
[0184] The "prefix" bits are the bits parsed as described above for the leadingZeroBits calculation and can be shown as 0 or 1 of the bit string in Table 8. That is, the bit string starting from 0 or 1 in Table 8 above can indicate the prefix bit string. The "suffix" bits are the bits parsed in the calculation of codeNum and are shown as xi in Table 8 above. That is, the bit string starting from xi in Table 8 above can indicate the suffix bit string. Here, i can be a value in the range from 0 to LeadingZeroBits - 1. Also, each xi can be the same as 0 or 1.
[0185] The bit string assigned to the above codeNum is as follows in the following table.
[0186] [Table 9]
[0187] When the descriptor of the syntax element is ue(v), that is, when the syntax element is coded with ue(v), the value of the syntax element can be the same as codeNum.
[0188] Also, for example, the above EGk evolution process can be performed as follows.
[0189] The input of the above EGk evolution process can be a requirement for EGk evolution. Also, the output of the above EGk evolution process can be the EGk evolution for the value symbolVal corresponding to the bit string.
[0190] The bit string of the EGk evolution process for symbolVal can be derived as follows.
[0191] [Table 10]
[0192] Referring to Table 10 described above, the binary value X can be added to the end of the bit string via each call of put(x). Here, x can be 0 or 1.
[0193] Also, for example, the above Limited EGk evolution process can be performed as follows.
[0194] The input of the above Limited EGk binary evolution process can be a requirement for Limited EGk binary evolution, the rice parameter riceParam, a variable log2TransformRange representing the binary logarithm of the maximum value, and a variable maxPreExtLen representing the maximum prefix extension length. Also, the output of the above Limited EGk binary evolution process can be the Limited EGk binary evolution for the value symbolVal corresponding to the bit string.
[0195] The bit string of the Limited EGk binary evolution process for symbolVal can be derived as follows.
[0196]
Table 11
[0197] Also, for example, the above FL binary evolution process can be performed as follows.
[0198] The input of the above FL binary evolution process can be a requirement for FL binary evolution and cMax for the above syntax element. Also, the output of the above FL binary evolution process can be the FL binary evolution for the value symbolVal corresponding to the bit string.
[0199] FL binary evolution can be configured using a bit string having a number of bits that is the fixed length of the symbol value symbolVal. Here, the above fixed-length bits can be an unsigned integer bitstring. That is, a bit string for the symbol value symbolVal can be derived by FL binary evolution, and the bit length (i.e., the number of bits) of the above bit string can be a fixed length.
[0200] For example, the above fixed length can be derived as in the following mathematical formula.
[0201] <Equation 10>
Number
[0202] Indexing such as bins for FL binary evolution can use a method where the values increase in order from the most significant bit to the least significant bit. For example, the bin index related to the above most significant bit can be binIdx = 0.
[0203] On the other hand, for example, among the above residual information, the binary evolution process for the syntax element abs_remainder can be performed as follows.
[0204] The input to the binary evolution process for the above abs_remainder can be the requirement for the binary evolution of the syntax element abs_remainder[n], the hue component cIdx, and the luma position (x0, y0). The above luma position (x0, y0) can refer to the upper left sample of the current luma transform block based on the upper left luma sample of the picture.
[0205] The output of the binary evolution process for the above abs_remainder can be the binary evolution of the above abs_remainder (i.e., the binary string of the above abs_remainder after binary evolution). The available bin string for the above abs_remainder can be derived by the above binary evolution process.
[0206] The Rice parameter cRiceParam for the above abs_remainder[n] can be derived through the process of deriving the Rice parameter that is executed with the above hue component cIdx, luma position (x0, y0), current coefficient scan position (xC, yC), binary logarithm log2TbWidth of the width of the transform block, and binary logarithm log2TbHeight of the height of the transform block as inputs. A specific explanation regarding the process of deriving the above Rice parameter will be described later.
[0207] Also, for example, cMax for the currently coded abs_remainder[n] can be derived based on the above Rice parameter cRiceParam. The above cMax can be derived as follows.
[0208] <Equation 11>
Number
[0209] On the other hand, the binary conversion for the above abs_remainder, that is, the bin string for the above abs_remainder, can be the concatenation of a prefix bin string and a suffix bin string if the suffix bin string exists. Also, if the above suffix bin string does not exist, the above bin string for the above abs_remainder can be the above prefix bin string.
[0210] For example, the above prefix bin string can be derived as described later.
[0211] The prefix value prefixVal of the above abs_remainder[n] can be derived as follows.
[0212] <Equation 12>
Number
[0213] The prefix of the bin string of the above abs_remainder[n] (i.e., the prefix bin string) can be derived through the TR binary evolution process for the above prefixVal using the above cMax and the above cRiceParam as inputs.
[0214] When the prefix bin string is the same as a bit string with all bits being 1 and a bit length of 6, a suffix bin string of the bin string of the above abs_remainder[n] may exist and can be derived as described below.
[0215] The derivation process of the Rice parameter for the above abs_remainder[n] is as follows.
[0216] The inputs to the derivation process of the Rice parameter can be the hue component index cIdx, the luma position (x0, y0), the current coefficient scan position (xC, yC), the binary logarithm log2TbWidth of the width of the transform block, and the binary logarithm log2TbHeight of the height of the transform block. The above luma position (x0, y0) can refer to the upper left sample of the current luma transform block with reference to the upper left luma sample of the picture. Also, the output of the derivation process of the Rice parameter can be the above Rice parameter cRiceParam.
[0217] For example, based on a given component index cIdx and an array AbsLevel[x][y] for a transform block having the above upper left luma position (x0, y0), the variable locSumAbs can be derived as in the pseudo code disclosed in the following table.
[0218] [Table 12]
[0219] After that, based on the given variable locSumAbs, the Rice parameter cRiceParam can be derived as shown in the following table.
[0220]
Table 13
[0221] Also, for example, in the process of deriving the Rice parameter for abs_remainder[n], baseLevel can be set to 4.
[0222] Alternatively, for example, based on whether the transform skip of the current block is available, the Rice parameter cRiceParam can be determined. That is, when no transform is applied to the current TB including the current CG, in other words, when a transform skip is applied to the current TB including the current CG, the Rice parameter cRiceParam can be derived to be 1.
[0223] Also, the suffix value suffixVal of the above abs_remainder can be derived as shown in the following formula.
[0224] <Formula 13>
Equation
[0225] The suffix bit string of the above abs_remainder's bin string can be derived by the Limited EGk binary evolution process for the above suffixVal where k is set to cRiceParam + 1, riceParam is set to cRiceParam, log2TransformRange is set to 15, and maxPreExtLen is set to 11.
[0226] On one hand, for example, among the above residual information, the binary process for the syntax element dec_abs_level can be performed as follows.
[0227] The input of the binary process for the dec_abs_level can be the requirement for the binary of the syntax element dec_abs_level[n], the hue component (colour component) cIdx, the luma position (x0, y0), the current coefficient scan position (xC, yC), the binary logarithm of the width of the transform block, i.e., log2TbWidth, and the binary logarithm of the height of the transform block, i.e., log2TbHeight. The above luma position (x0, y0) can refer to the upper left sample of the current luma transform block based on the upper left luma sample of the picture.
[0228] The output of the binary process for the dec_abs_level can be the binary of the dec_abs_level (i.e., the binary bit string of the dec_abs_level). A usable bit string for the dec_abs_level and the like can be derived by the above binary process.
[0229] The Rice parameter cRiceParam for the dec_abs_level[n] can be derived through a Rice parameter derivation process that takes the above hue component cIdx, luma position (x0, y0), current coefficient scan position (xC, yC), the binary logarithm of the width of the transform block, i.e., log2TbWidth, and the binary logarithm of the height of the transform block, i.e., log2TbHeight as inputs. A specific description of the above Rice parameter derivation process will be given later.
[0230] Also, for example, cMax for the dec_abs_level[n] can be derived based on the above Rice parameter cRiceParam. The cMax can be derived as in the following formula.
[0231] <Equation 14>
Number
[0232] On the other hand, the binary evolution for the above dec_abs_level[n], that is, the bit string for the above dec_abs_level[n] can be the concatenation of the prefix bit string and the suffix bit string if the suffix bit string exists. Also, if the suffix bit string does not exist, the bit string for the above dec_abs_level[n] can be the above prefix bit string.
[0233] For example, the above prefix bit string can be derived as described later.
[0234] The prefix value prefixVal of the above dec_abs_level[n] can be derived as in the following equation.
[0235] <Equation 15>
Number
[0236] The prefix of the bit string of the above dec_abs_level[n] (i.e., the prefix bit string) can be derived by the TR binary evolution process for the above prefixVal using the above cMax and the above cRiceParam as inputs.
[0237] If the above prefix bit string is the same as the bit string in which all bits are 1 and the bit length is 6, the suffix bit string of the bit string of the above dec_abs_level[n] may exist and can be derived as described later.
[0238] The process of deriving the Rice parameter for the above dec_abs_level[n] can be as follows.
[0239] The input of the above Rice parameter derivation process can be the hue component index (colour component index) cIdx, the luma position (x0, y0), the current coefficient scan position (xC, yC), the binary logarithm of the width of the transform block, log2TbWidth, and the binary logarithm of the height of the transform block, log2TbHeight. The above luma position (x0, y0) can refer to the top-left sample of the current luma transform block based on the top-left luma sample of the picture. Also, the output of the above Rice parameter derivation process can be the above Rice parameter cRiceParam.
[0240] For example, based on a given component index cIdx and the array AbsLevel[x][y] for a transform block having the above top-left luma position (x0, y0), the variable locSumAbs can be derived as in the pseudo code disclosed in the following table.
[0241] [Table 14]
[0242] After that, based on the given variable locSumAbs, the above Rice parameter cRiceParam can be derived as in the following table.
[0243] [Table 15]
[0244] Also, for example, in the process of deriving the Rice parameter for dec_abs_level[n], baseLevel can be set to 0, and the above ZeroPos[n] can be derived as follows in the following formula.
[0245] <Equation 16>
Number
[0246] Also, the suffix value suffixVal of the above dec_abs_level[n] can be derived as follows in the following formula.
[0247] <Equation 17>
Number
[0248] The suffix bit string of the above bin string of dec_abs_level[n] can be derived through the Limited EGk binary evolution process for the above suffixVal where k is set to cRiceParam + 1, truncSuffixLen is set to 15, and maxPreExtLen is set to 11.
[0249] On the other hand, the above-mentioned RRC and TSRC may have the following differences.
[0250] - For example, the Rice parameter cRiceParam of the syntax elements abs_remainder[] and dec_abs_level[] in RRC can be derived based on the above locSumAbs, look-up table, and / or baseLevel as described above, but the Rice parameter cRiceParam of the syntax element abs_remainder[] in TSRC can be derived as 1. That is, for example, when transform skip is applied to the current block (e.g., the current TB), the Rice parameter cRiceParam for abs_remainder[] of TSRC for the current block can be derived as 1.
[0251] - Also, for example, referring to Table 3 and Table 4, in RRC, abs_level_gtx_flag[n][0] and / or abs_level_gtx_flag[n][1] can be signaled, while in TSRC, abs_level_gtx_flag[n][0], abs_level_gtx_flag[n][1], abs_level_gtx_flag[n][2], abs_level_gtx_flag[n][3], and abs_level_gtx_flag[n][4] can be signaled. Here, the above abs_level_gtx_flag[n][0] can be represented as abs_level_gt1_flag or the first coefficient level flag, the above abs_level_gtx_flag[n][1] can be represented as abs_level_gt3_flag or the second coefficient level flag, the above abs_level_gtx_flag[n][2] can be represented as abs_level_gt5_flag or the third coefficient level flag, the above abs_level_gtx_flag[n][3] can be represented as abs_level_gt7_flag or the fourth coefficient level flag, and the above abs_level_gtx_flag[n][4] can be represented as abs_level_gt9_flag or the fifth coefficient level flag. Specifically, the above first coefficient level flag is a flag regarding whether the coefficient level is greater than a first threshold value (e.g., 1), the above second coefficient level flag is a flag regarding whether the coefficient level is greater than a second threshold value (e.g., 3), the above third coefficient level flag is a flag regarding whether the coefficient level is greater than a third threshold value (e.g., 5), the above fourth coefficient level flag is a flag regarding whether the coefficient level is greater than a fourth threshold value (e.g., 7), and the above fifth coefficient level flag is a flag regarding whether the coefficient level is greater than a fifth threshold value (e.g., 9). As described above, TSRC can further include abs_level_gtx_flag[n][2], abs_level_gtx_flag[n][3], and abs_level_gtx_flag[n][4] together with abs_level_gtx_flag[n][0] and abs_level_gtx_flag[n][1] compared to RRC.
[0252] - Also, for example, in RRC, the syntax element coeff_sign_flag can be bypass-coded, while in TSRC, the syntax element coeff_sign_flag can be either bypass-coded or context-coded.
[0253] Also, for the quantization process of residual samples, dependent quantization can be proposed. Dependent quantization can represent a manner in which the set of restored values allowed for the current transform coefficient depends on the values of the transform coefficients (the values of the transform coefficient levels) that precede the current transform coefficient in the restoration order. That is, for example, dependent quantization can be realized by (a) the restoration levels defining two other scalar quantizers and (b) defining a process for switching between the scalar quantizers. The above dependent quantization can have the effect that the allowed restoration vectors are denser in the N-dimensional vector space compared to the existing independent scalar quantization. Here, the above N can represent the number of transform coefficients of the transform block.
[0254] FIG. 9 exemplarily shows the scalar quantizers used in dependent quantization. Referring to FIG. 9, the positions of the available restoration levels can be specified as the size △ of the quantization step. Referring to FIG. 9, the scalar quantizers can be represented as Q0 and Q1. The scalar quantizers used can be derived without being explicitly signaled in the bitstream. For example, the quantizer used for the current transform coefficient can be determined by the parity of the transform coefficient levels that precede the current transform coefficient in the coding / restoration order.
[0255] FIG. 10 exemplarily shows the state transition and the selection of quantizers for dependent quantization.
[0256] Referring to FIG. 10, the switching between the two scalar quantizers (Q0 and Q1) can be realized by a state machine having four states. The four states can have four different values (0, 1, 2, 3). In the coding / decoding order, the state for the current transform coefficient can be determined by the parity of the transform coefficient levels before the current transform coefficient.
[0257] For example, when the inverse quantization process for a transform block is started, the state for the dependent quantization can be set to 0. Thereafter, the transform coefficients for the transform block can be restored in the scan order (i.e., the same order as the entropy decoded ones). For example, after the current transform coefficient is restored, as shown in FIG. 10, the state for the dependent quantization can be updated. Based on the updated state, the inverse quantization process for the transform coefficients restored after the current transform coefficient in the scan order can be executed. k shown in FIG. 10 can represent the value of the transform coefficient, i.e., the level value of the transform coefficient. For example, when the current state is 0, if k (the value of the current transform coefficient) & 1 is 0, the state can be updated to 0, and if k & 1 is 1, the state can be updated to 2. Also, for example, when the current state is 1, if k & 1 is 0, the state can be updated to 2, and if k & 1 is 1, the state can be updated to 0. Also, for example, when the current state is 2, if k & 1 is 0, the state can be updated to 1, and if k & 1 is 1, the state can be updated to 3. Also, for example, when the current state is 3, if k & 1 is 0, the state can be updated to 3, and if k & 1 is 1, the state can be updated to 1. Referring to FIG. 10, when the state is one of 0 and 1, the scalar quantizer used in the inverse quantization process can be Q0, and when the state is one of 2 and 3, the scalar quantizer used in the inverse quantization process can be Q1. The transform coefficients can be inverse quantized based on the quantization parameters for the restored levels of the transform coefficients by the scalar quantizer for the current state.
[0258] On the one hand, this document proposes embodiments related to residual data coding. The embodiments described in this document may be combined with each other. As described above, the methods of residual data coding may include Regular Residual Coding (RRC) and Transform Skip Residual Coding (TSRC).
[0259] Among the two methods described above, the method of residual data coding for the current block can be determined based on the values of transform_skip_flag and sh_ts_residual_coding_disabled_flag, as shown in Table 1. Here, the syntax element sh_ts_residual_coding_disabled_flag may indicate whether the above TSRC is available. Therefore, even when the transform_skip_flag indicates that transform skip is performed, and when sh_ts_residual_coding_disabled_flag indicates that the above TSRC is not available, the syntax element by RRC can be signaled for the transform skip block. That is, when the value of transform_skip_flag is 0 or the value of slice_ts_residual_coding_disabled_flag is 1, RRC can be used, and in other cases, TSRC can be used.
[0260] Using the above slice_ts_residual_coding_disabled_flag, high coding efficiency can be obtained in specific applications (e.g., lossless coding, etc.). However, in existing video / image coding standards, no constraints are proposed for the case where the aforementioned dependent quantization and the above slice_ts_residual_coding_disabled_flag are used together. That is, when dependent quantization is activated at a higher level (e.g., SPS (Sequence Parameter Set) syntax / VPS (Video Parameter Set) syntax / DPS (Decoding Parameter Set) syntax / picture header syntax / slice header syntax, etc.) or a lower level (CU / TU), and the above slice_ts_residual_coding_disabled_flag is 1, a value depending on the state of dependent quantization in the RRC performs an unnecessary operation (i.e., an operation by dependent quantization), resulting in a degradation of coding performance, or due to incorrect settings in the encoding device, an unintentional loss of coding performance may occur. Therefore, in the present embodiment, a proposal is made to set the dependency / constraints between the two technologies to prevent the occurrence of unintentional coding losses or malfunctions when dependent quantization and residual coding when slice_ts_residual_coding_disabled_flag = 1 (i.e., coding the residual samples of the transform skip block within the current slice by the RRC) are used together.
[0261] As one embodiment, this document proposes a method in which slice_ts_residual_coding_disabled_flag depends on ph_dep_quant_enabled_flag. For example, the syntax elements proposed in this embodiment are as follows in the following table.
[0262]
Table 16
[0263] According to this embodiment, the slice_ts_residual_coding_disabled_flag can be signaled when the value of the ph_dep_quant_enabled_flag is 0. Here, the ph_dep_quant_enabled_flag can indicate whether dependent quantization can be used. For example, when the value of the ph_dep_quant_enabled_flag is 1, the ph_dep_quant_enabled_flag can indicate that dependent quantization can be used, and when the value of the ph_dep_quant_enabled_flag is 0, the ph_dep_quant_enabled_flag can indicate that dependent quantization cannot be used.
[0264] Therefore, according to this embodiment, the slice_ts_residual_coding_disabled_flag can be signaled only when the above-mentioned dependent quantization is not available. When the above-mentioned dependent quantization is available and the slice_ts_residual_coding_disabled_flag is not signaled, the slice_ts_residual_coding_disabled_flag can be considered (inferred) as 0. On the other hand, the ph_dep_quant_enabled_flag and the slice_ts_residual_coding_disabled_flag can be signaled in the picture header syntax and / or the slice header syntax, or in other higher-level syntaxes (High Level Syntax, HLS) that are not the picture header syntax and the slice header syntax (for example, SPS syntax / VPS syntax / DPS syntax, etc.) or lower levels (CU / TU). When the ph_dep_quant_enabled_flag is signaled in a syntax other than the picture header syntax, it may be called by another name. For example, the ph_dep_quant_enabled_flag may also be represented as sh_dep_quant_enabled_flag, sh_dep_quant_used_flag, or sps_dep_quant_enabled_flag.
[0265] Furthermore, this document proposes another embodiment that sets the dependency / constraint between dependent quantization and residual coding when slice_ts_residual_coding_disabled_flag = 1 (i.e., coding the residual samples of the transform skip blocks within the current slice with RRC). For example, in this embodiment, when the value of slice_ts_residual_coding_disabled_flag is 1, in order to prevent dependent quantization and residual coding when slice_ts_residual_coding_disabled_flag = 1 (i.e., coding the residual samples of the transform skip blocks within the current slice with RRC) from causing unintended coding loss or malfunction together, a proposal is made that the state of the above-mentioned dependent quantization is not used in the coding of the level values of the transform coefficients. The syntax of the residual coding according to this embodiment is as follows.
[0266] [Table 17-1]
[0267] [Table 17-2]
[0268] [Table 17-3]
[0269] [Table 17-4]
[0270] [Table 17-5]
[0271]
Table 17-6
[0272] Referring to the aforementioned Table 17, when the value of ph_dep_quant_enabled_flag is 1 and the value of slice_ts_residual_coding_disabled_flag is 0, QState can be derived, and based on the above QState, the value of the transform coefficient (transform coefficient level) can be derived. For example, referring to Table 17, the above transform coefficient level TransCoeffLevel[x0][y0][cIdx][xC][yC] can be derived by (2*AbsLevel[xC][yC]-(QState>1?1:0))*(1-2*coeff_sign_flag[n]). Here, AbsLevel[xC][yC] can be the absolute value of the transform coefficient derived based on the syntax element of the transform coefficient, coeff_sign_flag[n] can be the syntax element of the sign flag representing the sign of the transform coefficient, and (QState>1?1:0) can represent that when the value of state QState is greater than 1, that is, when the value of state QState is 2 or 3, it is 1, and when the value of state QState is 1 or less, that is, when the value of state QState is 0 or 1, it is 0.
[0273] Also, referring to the aforementioned Table 17, when the value of slice_ts_residual_coding_disabled_flag is 1, the value of the transform coefficient (transform coefficient level) can be derived without using the above QState. For example, referring to Table 17, the above transform coefficient level TransCoeffLevel[x0][y0][cIdx][xC][yC] can be derived by AbsLevel[xC][yC]*(1-2*coeff_sign_flag[n]). Here, AbsLevel[xC][yC] can be the absolute value of the transform coefficient derived based on the syntax element of the transform coefficient, and coeff_sign_flag[n] can be the syntax element of the sign flag representing the sign of the transform coefficient.
[0274] Also, according to this embodiment, when the value of slice_ts_residual_coding_disabled_flag is 1, in the coding of the level value of the transform coefficient, the state of the dependent quantization described above may not be used, and the update of the state may not be executed either. For example, the syntax of the residual coding according to this embodiment is as shown in the following table.
[0275]
Table 18-1
[0276]
Table 18-2
[0277]
Table 18-3
[0278]
Table 18-4
[0279]
Table 18-5
[0280]
Table 18-6
[0281] Referring to Table 18 above, when the value of ph_dep_quant_enabled_flag is 1 and the value of slice_ts_residual_coding_disabled_flag is 0, the QState can be updated. For example, when the value of ph_dep_quant_enabled_flag is 1 and the value of slice_ts_residual_coding_disabled_flag is 0, the QState can be updated to QStateTransTable[QState][AbsLevelPass1[xC][yC]&1] or QStateTransTable[QState][AbsLevel[xC][yC]&1]. Also, when the value of slice_ts_residual_coding_disabled_flag is 1, the process of updating the QState may not be executed.
[0282] Referring to Table 18 above, when the value of ph_dep_quant_enabled_flag is 1 and the value of slice_ts_residual_coding_disabled_flag is 0, the value of the transform coefficient (transform coefficient level) can be derived based on the above QState. For example, referring to Table 18, the above transform coefficient level TransCoeffLevel[x0][y0][cIdx][xC][yC] can be derived by (2*AbsLevel[xC][yC]-(QState>1?1:0))*(1-2*coeff_sign_flag[n]). Here, AbsLevel[xC][yC] can be the absolute value of the transform coefficient derived based on the syntax element of the transform coefficient, coeff_sign_flag[n] can be the syntax element of the sign flag representing the sign of the transform coefficient, and (QState>1?1:0) can represent that when the value of the state QState is greater than 1, that is, when the value of the state QState is 2 or 3, it is 1, and when the value of the state QState is 1 or less, that is, when the value of the state QState is 0 or 1, it is 0.
[0283] Also, referring to Table 18 above, when the value of slice_ts_residual_coding_disabled_flag is 1, the value of the transform coefficient (transform coefficient level) can be derived without using the above QState. For example, referring to Table 18, the above transform coefficient level TransCoeffLevel[x0][y0][cIdx][xC][yC] can be derived as AbsLevel[xC][yC]*(1 - 2*coeff_sign_flag[n]). Here, AbsLevel[xC][yC] can be the absolute value of the transform coefficient derived based on the syntax element of the transform coefficient, and coeff_sign_flag[n] can be the syntax element of the sign flag representing the sign of the transform coefficient.
[0284] Also, this document proposes another embodiment for setting the dependency / constraint between the dependent quantization and the residual coding when slice_ts_residual_coding_disabled_flag = 1 (i.e., coding the residual samples of the transform skip block within the current slice with RRC). For example, this embodiment proposes adding a constraint using transform_skip_flag to the process of updating the state of the dependent quantization in RRC or deriving the value of the transform coefficient (transform coefficient level) depending on the state. That is, this embodiment proposes a solution to prevent the process of updating the state of the dependent quantization in RRC and / or deriving the value of the transform coefficient (transform coefficient level) depending on the state from being used based on the above transform_skip_flag. The syntax of the residual coding according to this embodiment is as follows.
[0285]
Table 19-1
[0286]
Table 19-2
[0287]
Table 19-3
[0288]
Table 19-4
[0289]
Table 19-5
[0290]
Table 19-6
[0291] Referring to Table 19 above, when the value of ph_dep_quant_enabled_flag is 1 and the value of transform_skip_flag is 0, QState can be updated. For example, when the value of ph_dep_quant_enabled_flag is 1 and the value of transform_skip_flag is 0, QState can be updated to QStateTransTable[QState][AbsLevelPass1[xC][yC]&1] or QStateTransTable[QState][AbsLevel[xC][yC]&1]. Also, when the value of transform_skip_flag is 1, the process of updating QState may not be executed.
[0292] Also, referring to Table 19 above, when the value of ph_dep_quant_enabled_flag is 1 and the value of transform_skip_flag is 0, QState can be derived, and based on the above QState, the value of the transform coefficient (transform coefficient level) can be derived. For example, referring to Table 19, the above transform coefficient level TransCoeffLevel[x0][y0][cIdx][xC][yC] can be derived by (2 * AbsLevel[xC][yC] - (QState > 1? 1 : 0)) * (1 - 2 * coeff_sign_flag[n]). Here, AbsLevel[xC][yC] can be the absolute value of the transform coefficient derived based on the syntax element of the transform coefficient, coeff_sign_flag[n] can be the syntax element of the sign flag representing the sign of the transform coefficient, and (QState > 1? 1 : 0) can represent that when the value of state QState is greater than 1, that is, when the value of state QState is 2 or 3, it is 1, and when the value of state QState is 1 or less, that is, when the value of state QState is 0 or 1, it is 0.
[0293] Also, referring to Table 19 above, when the value of transform_skip_flag is 1, the value of the transform coefficient (transform coefficient level) can be derived without using the above QState. Therefore, when residual data by RRC is coded for the transform skip block, the value of the transform coefficient can be derived without using Qstate. For example, referring to Table 19, the above transform coefficient level TransCoeffLevel[x0][y0][cIdx][xC][yC] can be derived by AbsLevel[xC][yC] * (1 - 2 * coeff_sign_flag[n]). Here, AbsLevel[xC][yC] can be the absolute value of the transform coefficient derived based on the syntax element of the transform coefficient, and coeff_sign_flag[n] can be the syntax element of the sign flag representing the sign of the transform coefficient.
[0294] On the one hand, as described above, the information (syntax elements) in the syntax table disclosed in this document can be included in the image / video information, configured / encoded by the encoding device, and transmitted to the decoding device in the form of a bitstream. The decoding device can parse / decode the information (syntax elements) in the syntax table. The decoding device can execute a block / image / video restoration procedure based on the decoded information.
[0295] FIG. 11 schematically shows an image encoding method by the encoding device according to this document. The method disclosed in FIG. 11 can be executed by the encoding device disclosed in FIG. 2. Specifically, for example, S1100 in FIG. 11 can be executed by the prediction unit of the encoding device, and S1110 to S1160 in FIG. 11 can be executed by the entropy encoding unit of the encoding device. Also, although not shown, the process of deriving the residual sample for the current block based on the original sample and the predicted sample for the current block can be executed by the subtraction unit of the encoding device, and the process of generating the restored sample and the restored picture for the current block based on the residual sample and the predicted sample for the current block can be executed by the addition unit of the encoding device.
[0296] The encoding device derives a predicted sample of the current block based on inter prediction (S1100). For example, the encoding device may derive an inter prediction mode and motion information of the current block, and may generate a predicted sample of the current block. Here, the determination of the inter prediction mode, the derivation of the motion information, and the generation procedure of the predicted sample may be executed simultaneously as described above, or any one of the procedures may be executed prior to the other procedures. For example, the encoding device may search for a block similar to the current block within a certain area (search area) of the reference picture through motion estimation, and may derive a reference block whose difference from the current block is minimum or below a certain criterion. Based on this, the encoding device may derive an index of the reference picture indicating the reference picture where the reference block is located, and may derive a motion vector based on the positional difference between the reference block and the current block. The encoding device can determine the inter prediction mode applied to the current block among various inter prediction modes. For example, the encoding device can compare the RD costs for the various inter prediction modes and determine the optimal inter prediction mode for the current block.
[0297] For example, the encoding device may configure a motion information candidate list for the current block, and may derive a reference block whose difference from the current block is minimum or below a certain criterion among the reference blocks pointed to by the motion information candidates included in the motion information candidate list. In this case, the motion information candidate related to the derived reference block may be selected, and the motion information of the current block may be derived based on the motion information of the selected motion information candidate.
[0298] The encoding device encodes prediction-related information for the current block (S1110). The image information may include prediction-related information for the current block. For example, the prediction-related information is information related to the prediction procedure and may include prediction mode information and information related to the motion information of the current block. The information related to the motion information may include index information of motion information candidates that is information for deriving a motion vector. Also, for example, the information related to the motion information may include information related to the aforementioned MVD (Motion Vector Difference, MVD) and / or index information of a reference picture.
[0299] The encoding device encodes a dependent quantization available flag regarding whether dependent quantization can be used (S1120). The encoding device can encode a dependent quantization available flag regarding whether dependent quantization can be used. The image information may include the dependent quantization available flag. For example, the encoding device can determine whether dependent quantization can be used for blocks of pictures within a sequence, and can encode a dependent quantization available flag regarding whether dependent quantization can be used. For example, the dependent quantization available flag may be a flag regarding whether dependent quantization can be used. For example, the dependent quantization available flag may indicate whether dependent quantization can be used. That is, for example, the dependent quantization available flag may indicate whether dependent quantization can be used for blocks of pictures within a sequence. For example, the dependent quantization available flag may indicate whether there can be a dependent quantization use flag indicating whether dependent quantization is used for the current slice. For example, the dependent quantization available flag with a value of 1 may indicate that the dependent quantization can be used, and the dependent quantization available flag with a value of 0 may indicate that the dependent quantization cannot be used. Also, for example, the dependent quantization available flag may be signaled by SPS syntax or slice header syntax, etc. The syntax element of the dependent quantization available flag may be the aforementioned sps_dep_quant_enabled_flag. The sps_dep_quant_enabled_flag may be called sh_dep_quant_enabled_flag, sh_dep_quant_used_flag, or ph_dep_quant_enabled_flag.
[0300] The encoding device encodes a TSRC available flag indicating whether Transform Skip Residual Coding (TSRC) can be used based on the dependent quantization available flag (S1130). The image information may include the TSRC available flag.
[0301] For example, the encoding device can encode the TSRC available flag based on the dependent quantization available flag. For example, the TSRC available flag can be encoded based on the dependent quantization available flag having a value of 0. That is, for example, when the value of the dependent quantization available flag is 0 (i.e., when the dependent quantization available flag indicates that dependent quantization is not available), the TSRC available flag can be encoded. In other words, for example, when the value of the dependent quantization available flag is 0 (i.e., when the dependent quantization available flag indicates that dependent quantization is not available), the TSRC available flag can be signaled. Also, for example, when the value of the dependent quantization available flag is 1, the TSRC available flag may not be encoded, and the value of the TSRC available flag can be derived as 0 at the decoding device. That is, for example, when the value of the dependent quantization available flag is 1 (e.g., when dependent quantization is applied (or used) to the current block), the TSRC available flag may not be signaled, and the value of the TSRC available flag can be derived as 0 at the decoding device. Therefore, for example, when dependent quantization is not available for the current block, the TSRC available flag can be signaled (or encoded). When dependent quantization is available for the current block, the TSRC available flag may not be signaled (or encoded), and the value of the TSRC available flag can be derived as 0 at the decoding device. Here, the current block can be a Coding Block (CB) or a Transform Block (TB).
[0302] Here, for example, the above-mentioned TSRC available flag may be a flag regarding whether TSRC is available. That is, for example, the above-mentioned TSRC available flag may be a flag indicating whether TSRC is available for blocks within a slice. For example, the above-mentioned TSRC available flag with a value of 1 may indicate that the above-mentioned TSRC is not available, and the above-mentioned TSRC available flag with a value of 0 may indicate that the above-mentioned TSRC is available. Also, for example, the above-mentioned TSRC available flag may be signaled in the Slice Header syntax. The syntax element of the above-mentioned TSRC available flag may be the aforementioned sh_ts_residual_coding_disabled_flag.
[0303] The encoding device determines the syntax of residual coding for the current block based on the above-mentioned TSRC available flag (S1140). The encoding device can determine the syntax of residual coding for the current block based on the above-mentioned TSRC available flag. For example, the encoding device can determine the syntax of residual coding for the current block as one of the syntax of Regular Residual Coding (RRC) and the syntax of Transform Skip Residual Coding (TSRC) based on the above-mentioned TSRC available flag. The RRC syntax may represent the syntax by RRC, and the TSRC syntax may represent the syntax by TSRC.
[0304] For example, based on the above-mentioned TSRC available flag with a value of 1, the syntax of the residual coding for the current block can be determined as the syntax of Regular Residual Coding (RRC). In this case, for example, a transform skip flag regarding whether the transform skip of the current block is available can be encoded, and the value of the transform skip flag can be 1. For example, the above image information may include a transform skip flag for the current block. The transform skip flag can indicate whether the transform skip of the current block is available. That is, the transform skip flag can indicate whether a transform is applied to the transform coefficients of the current block. The syntax element representing the transform skip flag can be the aforementioned transform_skip_flag. For example, when the value of the transform skip flag is 1, the transform skip flag can indicate that no transform is applied to the current block (i.e., transform is skipped), and when the value of the transform skip flag is 0, the transform skip flag can indicate that a transform is applied to the current block. For example, when the current block is a transform skip block, the value of the transform skip flag for the current block can be 1.
[0305] Also, for example, based on the above-mentioned TSRC available flag having a value of 0, the syntax of the residual coding for the current block can be determined as the syntax of Transform Skip Residual Coding (TSRC). Also, for example, a transform skip flag regarding whether transform skip of the current block is available can be encoded, and based on the above-mentioned transform skip flag having a value of 1 and the above-mentioned TSRC available flag having a value of 0, the syntax of the residual coding for the current block can be determined as the syntax of Transform Skip Residual Coding (TSRC). Also, for example, a transform skip flag regarding whether transform skip of the current block is available can be encoded, and based on the above-mentioned transform skip flag having a value of 0 and the above-mentioned TSRC available flag having a value of 0, the syntax of the residual coding for the current block can be determined as the syntax of Regular Residual Coding (RRC).
[0306] The encoding device encodes the residual information of the determined residual coding syntax for the current block (S1150). The encoding device can derive a residual sample for the current block and can encode the residual information of the determined residual coding syntax for the residual sample of the current block. The image information may include residual information.
[0307] For example, the encoding device can derive a residual sample for the current block through subtraction of the original sample and the prediction sample for the current block.
[0308] Thereafter, for example, the encoding device may derive the transform coefficients of the current block based on the residual samples. For example, the encoding device can determine whether transformation is applicable to the current block. That is, the encoding device can determine whether transformation is applicable to the residual samples of the current block. The encoding device can determine whether it is possible to apply transformation to the current block in consideration of coding efficiency. For example, the encoding device can determine that transformation is not applicable to the current block. The block to which the transformation is not applied may be represented as a transform skip block. That is, for example, the current block may be a transform skip block.
[0309] When transformation is not applied to the current block, that is, when transformation is not applied to the residual samples, the encoding device may derive the derived residual samples as the transform coefficients. Also, when transformation is applied to the current block, that is, when transformation is applied to the residual samples, the encoding device may perform transformation on the residual samples and derive the transform coefficients. The current block may include a plurality of sub-blocks or coefficient groups (CG). Also, the size of the sub-blocks of the current block may be 4x4 size or 2x2 size. That is, the sub-blocks of the current block may include a maximum of 16 non-zero transform coefficients or a maximum of 4 non-zero transform coefficients. Here, the current block may be a coding block (CB) or a transform block (TB). Also, the transform coefficient may also be represented as a residual coefficient.
[0310] On the one hand, the encoding device can determine whether dependent quantization is applied to the current block. For example, when the dependent quantization is applied to the current block, the encoding device can perform the dependent quantization process on the transform coefficients to derive the transform coefficients of the current block. For example, when the dependent quantization is applied to the current block, the encoding device can update the state (Qstate) for the dependent quantization based on the coefficient level of the transform coefficient immediately before the current transform coefficient in the scan order, and can derive the coefficient level of the current transform coefficient based on the updated state and the syntax element regarding the current transform coefficient, quantize the derived coefficient level, and derive the current transform coefficient. For example, the current transform coefficient can be quantized based on the quantization parameter for the restoration level of the current transform coefficient by a scalar quantizer for the updated state.
[0311] For example, when the syntax of the residual coding for the current block is determined as the RRC syntax, the encoding device can encode the residual information of the RRC syntax for the current block. For example, the residual information of the RRC syntax may include the syntax elements disclosed in Table 2 above.
[0312] For example, the residual information of the RRC syntax may include syntax elements regarding the transform coefficients of the current block. Here, the transform coefficient can also be expressed as a residual coefficient.
[0313] For example, the syntax elements may include syntax elements such as last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix, sb_coded_flag, sig_coeff_flag, par_level_flag, abs_level_gtX_flag (e.g., abs_level_gtx_flag[n][0] and / or abs_level_gtx_flag[n][1]), abs_remainder, dec_abs_level, and / or coeff_sign_flag.
[0314] Specifically, for example, the syntax elements may include position information representing the position of the last non-zero transform coefficient in the array of residual coefficients of the current block. That is, the syntax elements may include position information representing the position of the last non-zero transform coefficient in the scanning order of the current block. The position information may include information representing the prefix of the column position of the last non-zero transform coefficient, information representing the prefix of the row position of the last non-zero transform coefficient, information representing the suffix of the column position of the last non-zero transform coefficient, and information representing the suffix of the row position of the last non-zero transform coefficient. The syntax elements related to the position information may be last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix. On the other hand, the non-zero transform coefficient may also be called a significant coefficient.
[0315] Also, for example, the syntax element may include an encoded sub-block flag indicating whether the current sub-block of the current block includes non-zero transform coefficients, a valid coefficient flag indicating whether the transform coefficients of the current block are non-zero transform coefficients, a first coefficient level flag regarding whether the coefficient level for the transform coefficient is greater than a first threshold, a parity level flag regarding the parity of the coefficient level, and / or a second coefficient level flag regarding whether the coefficient level of the transform coefficient is greater than a second threshold. Here, the encoded sub-block flag may be sb_coded_flag or coded_sub_block_flag, the valid coefficient flag may be sig_coeff_flag, the first coefficient level flag may be abs_level_gt1_flag or abs_level_gtx_flag, the parity level flag may be par_level_flag, and the second coefficient level flag may be abs_level_gt3_flag or abs_level_gtx_flag.
[0316] Also, for example, the syntax element may include coefficient value related information regarding the value of the transform coefficient of the current block. The coefficient value related information may be abs_remainder and / or dec_abs_level.
[0317] Also, for example, the syntax element may include a sign flag representing the sign of the transform coefficient. The sign flag may be coeff_sign_flag.
[0318] For example, when the syntax of the residual coding for the current block is determined as the TSRC syntax, the encoding device can encode the residual information of the TSRC syntax for the current block. For example, the residual information of the TSRC syntax may include the syntax elements disclosed in Table 3 above.
[0319] For example, the residual information of the TSRC syntax may include syntax elements related to the transform coefficients of the current block. Here, the transform coefficient may also be expressed as a residual coefficient.
[0320] For example, the syntax elements may include context-coded syntax elements and / or bypass-coded syntax elements for the transform coefficients. The syntax elements may include syntax elements such as sig_coeff_flag, coeff_sign_flag, par_level_flag, abs_level_gtX_flag (e.g., abs_level_gtx_flag[n][0], abs_level_gtx_flag[n][1], abs_level_gtx_flag[n][2], abs_level_gtx_flag[n][3] and / or abs_level_gtx_flag[n][4]), abs_remainder and / or coeff_sign_flag.
[0321] For example, the context-coded syntax element for the conversion coefficient may include a valid coefficient flag indicating whether the conversion coefficient is a non-zero conversion coefficient, a sign flag indicating the sign of the conversion coefficient, a first coefficient level flag regarding whether the coefficient level of the conversion coefficient is greater than a first threshold value, and / or a parity level flag regarding the parity of the coefficient level of the conversion coefficient. Also, for example, the context-coded syntax element may include a second coefficient level flag regarding whether the coefficient level of the conversion coefficient is greater than a second threshold value, a third coefficient level flag regarding whether the coefficient level of the conversion coefficient is greater than a third threshold value, a fourth coefficient level flag regarding whether the coefficient level of the conversion coefficient is greater than a fourth threshold value, and / or a fifth coefficient level flag regarding whether the coefficient level of the conversion coefficient is greater than a fifth threshold value. Here, the valid coefficient flag may be sig_coeff_flag, the sign flag may be ceff_sign_flag, the first coefficient level flag may be abs_level_gt1_flag, and the parity level flag may be par_level_flag. Also, the second coefficient level flag may be abs_level_gt3_flag or abs_level_gtx_flag, the third coefficient level flag may be abs_level_gt5_flag or abs_level_gtx_flag, the fourth coefficient level flag may be abs_level_gt7_flag or abs_level_gtx_flag, and the fifth coefficient level flag may be abs_level_gt9_flag or abs_level_gtx_flag.
[0322] Also, for example, the bypass-coded syntax element for the above conversion coefficient may include coefficient level information regarding the value (or coefficient level) of the above conversion coefficient and / or a sign flag representing the sign of the above conversion coefficient. The above coefficient level information may be abs_remainder and / or dec_abs_level, and the above sign flag may be ceff_sign_flag.
[0323] The encoding device generates a bitstream including the above prediction-related information, the above dependent quantization available flag, the above TSRC available flag, and the above residual information (S1160). For example, the encoding device can output image information including the above prediction-related information, the above dependent quantization available flag, the above TSRC available flag, and the above residual information as a bitstream. The above bitstream may include the above prediction-related information, the above dependent quantization available flag, the above TSRC available flag, and the above residual information.
[0324] On the other hand, the above bitstream can be transmitted to the decoding device via a network or a (digital) storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0325] FIG. 12 schematically shows an encoding apparatus that performs the image encoding method according to this document. The method disclosed in FIG. 11 can be executed by the encoding apparatus disclosed in FIG. 12. Specifically, for example, the prediction unit of the encoding apparatus in FIG. 12 can execute S1100 in FIG. 11, and the entropy encoding unit of the encoding apparatus in FIG. 12 can execute S1110 to S1160 in FIG. 11. Also, although not shown, the process of deriving the residual sample for the current block based on the original sample and the predicted sample for the current block can be executed by the subtraction unit of the encoding apparatus, and the process of generating the restored sample and the restored picture for the current block based on the residual sample and the predicted sample for the current block can be executed by the addition unit of the encoding apparatus.
[0326] FIG. 13 schematically shows an image decoding method by the decoding apparatus according to this document. The method disclosed in FIG. 13 can be executed by the decoding apparatus disclosed in FIG. 3. Specifically, for example, S1300 to S1330 in FIG. 13 can be executed by the entropy decoding unit of the decoding apparatus, (S1300 to) S1340 to S1350 in FIG. 13 can be executed by the prediction unit of the decoding apparatus, S1360 in FIG. 13 can be executed by the residual processing unit of the decoding apparatus, and S1370 can be executed by the addition unit of the decoding apparatus.
[0327] The decoding device acquires prediction-related information for the current block (S1300). The decoding device can acquire the prediction-related information for the current block via a bitstream. For example, the image information may include prediction-related information for the current block. For example, the prediction-related information may include prediction mode information for the current block. The decoding device can determine which inter prediction mode is applied to the current block based on the prediction mode information. For example, the inter prediction mode may include a skip mode, a merge mode, and / or an (A)MVP mode, or may include various inter prediction modes described above.
[0328] The decoding device acquires a dependent quantization usable flag regarding whether dependent quantization is usable (S1310). The decoding device can acquire image information including the dependent quantization usable flag via a bitstream. The image information may include the dependent quantization usable flag. For example, the dependent quantization usable flag may be a flag regarding whether dependent quantization is usable. For example, the dependent quantization usable flag may indicate whether dependent quantization is usable. That is, for example, the dependent quantization usable flag may indicate whether dependent quantization is usable for blocks of pictures within a sequence. For example, the dependent quantization usable flag may indicate whether there can be a dependent quantization usage flag indicating whether dependent quantization is used for the current slice. For example, the dependent quantization usable flag having a value of 1 may indicate that the dependent quantization is usable, and the dependent quantization usable flag having a value of 0 may indicate that the dependent quantization is not usable. Also, for example, the dependent quantization usable flag may be signaled by an SPS syntax or a slice header syntax or the like. The syntax element of the dependent quantization usable flag may be the aforementioned sps_dep_quant_enabled_flag. The sps_dep_quant_enabled_flag may be called sh_dep_quant_enabled_flag, sh_dep_quant_used_flag, or ph_dep_quant_enabled_flag.
[0329] The decoding device acquires a TSRC usable flag regarding whether Transform Skip Residual Coding (TSRC) is usable based on the dependent quantization usable flag (S1320). The image information may include the TSRC usable flag.
[0330] For example, the decoding device can obtain the TSRC available flag based on the above-mentioned dependent quantization available flag. For example, the TSRC available flag can be obtained based on the dependent quantization available flag whose value is 0. That is, for example, when the value of the dependent quantization available flag is 0 (i.e., when the dependent quantization available flag indicates that dependent quantization is not available), the TSRC available flag can be obtained. In other words, for example, when the value of the dependent quantization available flag is 0 (i.e., when the dependent quantization available flag indicates that dependent quantization is not available), the TSRC available flag can be signaled. Also, for example, when the value of the dependent quantization available flag is 1, the TSRC available flag may not be obtained, and the value of the TSRC available flag can be derived as 0. That is, for example, when the value of the dependent quantization available flag is 1 (e.g., when dependent quantization is applied (or used) to the current block), the TSRC available flag may not be signaled, and the value of the TSRC available flag can be derived as 0. Therefore, for example, when dependent quantization is not available for the current block, the TSRC available flag can be signaled (or obtained). When dependent quantization is available for the current block, the TSRC available flag may not be signaled (or obtained), and the value of the TSRC available flag can be derived as 0. Here, the current block can be a coding block (CB) or a transform block (TB).
[0331] Here, for example, the above-mentioned TSRC available flag may be a flag regarding whether the TSRC is available. That is, for example, the above-mentioned TSRC available flag may be a flag indicating whether the TSRC is available for blocks within a slice. For example, the above-mentioned TSRC available flag with a value of 1 may indicate that the above-mentioned TSRC is not available, and the above-mentioned TSRC available flag with a value of 0 may indicate that the above-mentioned TSRC is available. Also, for example, the above-mentioned TSRC available flag may be signaled in a Slice Header syntax. The syntax element of the above-mentioned TSRC available flag may be the aforementioned sh_ts_residual_coding_disabled_flag.
[0332] The decoding device acquires residual information of the syntax of residual coding for the current block derived based on the above-mentioned TSRC available flag (S1330). The decoding device may derive one of the syntax of Regular Residual Coding (RRC) and the syntax of TSRC as the syntax of residual coding for the current block based on the above-mentioned TSRC available flag, and may acquire the residual information of the derived syntax of residual coding.
[0333] For example, the decoding device may derive the syntax of residual coding for the current block based on the above-mentioned TSRC available flag. For example, the decoding device may derive the syntax of residual coding for the current block as one of the syntax of Regular Residual Coding (RRC) and the syntax of Transform Skip Residual Coding (TSRC) based on the above-mentioned TSRC available flag. The RRC syntax may represent the syntax by RRC, and the TSRC syntax may represent the syntax by TSRC.
[0334] For example, based on the above TSRC usable flag with a value of 1, the syntax of the residual coding for the current block can be derived as the syntax of Regular Residual Coding (RRC). In this case, for example, a transform skip flag regarding whether the transform skip of the current block is usable can be obtained, and the value of the transform skip flag can be 1. For example, the above image information may include a transform skip flag for the current block. The transform skip flag may indicate whether the transform skip of the current block is usable. That is, the transform skip flag may indicate whether a transform is applied to the transform coefficients of the current block. The syntax element representing the transform skip flag may be the aforementioned transform_skip_flag. For example, when the value of the transform skip flag is 1, the transform skip flag may indicate that no transform is applied to the current block (i.e., transform is skipped), and when the value of the transform skip flag is 0, the transform skip flag may indicate that a transform is applied to the current block. For example, when the current block is a transform skip block, the value of the transform skip flag for the current block can be 1.
[0335] Also, for example, based on the above TSRC available flag with a value of 0, the syntax of the residual coding for the current block can be derived as the syntax of Transform Skip Residual Coding (TSRC). Also, for example, a transform skip flag regarding whether the transform skip of the current block is available can be obtained, and based on the transform skip flag with a value of 1 and the TSRC available flag with a value of 0, the syntax of the residual coding for the current block can be derived as the syntax of Transform Skip Residual Coding (TSRC). Also, for example, a transform skip flag regarding whether the transform skip of the current block is available can be obtained, and based on the transform skip flag with a value of 0 and the TSRC available flag with a value of 0, the syntax of the residual coding for the current block can be derived as the syntax of Regular Residual Coding (RRC).
[0336] Thereafter, for example, the decoding device can obtain the residual information of the derived residual coding syntax for the current block. The image information may include the residual information.
[0337] For example, when the syntax of the residual coding for the current block is derived as the RRC syntax, the decoding device can obtain the residual information of the RRC syntax for the current block. For example, the residual information of the RRC syntax may include the syntax elements disclosed in Table 2 above.
[0338] For example, the residual information of the RRC syntax may include syntax elements regarding the transform coefficients of the current block. Here, the transform coefficient can also be expressed as a residual coefficient.
[0339] For example, the syntax element may include syntax elements such as last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix, sb_coded_flag, sig_coeff_flag, par_level_flag, abs_level_gtX_flag (e.g., abs_level_gtx_flag[n][0] and / or abs_level_gtx_flag[n][1]), abs_remainder, dec_abs_level, and / or coeff_sign_flag.
[0340] Specifically, for example, the syntax element may include position information representing the position of the last non-zero transform coefficient in the array of residual coefficients of the current block. That is, the syntax element may include position information representing the position of the last non-zero transform coefficient in the scanning order of the current block. The position information may include information representing the prefix of the column position of the last non-zero transform coefficient, information representing the prefix of the row position of the last non-zero transform coefficient, information representing the suffix of the column position of the last non-zero transform coefficient, and information representing the suffix of the row position of the last non-zero transform coefficient. The syntax elements related to the position information may be last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, last_sig_coeff_y_suffix. On the other hand, the non-zero transform coefficient may also be called a significant coefficient.
[0341] Also, for example, the syntax element may include an encoded sub-block flag indicating whether the current sub-block of the current block contains non-zero transform coefficients, a valid coefficient flag indicating whether the transform coefficient of the current block is a non-zero transform coefficient, a first coefficient level flag regarding whether the coefficient level for the transform coefficient is greater than a first threshold, a parity level flag regarding the parity of the coefficient level, and / or a second coefficient level flag regarding whether the coefficient level of the transform coefficient is greater than a second threshold. Here, the encoded sub-block flag may be sb_coded_flag or coded_sub_block_flag, the valid coefficient flag may be sig_coeff_flag, the first coefficient level flag may be abs_level_gt1_flag or abs_level_gtx_flag, the parity level flag may be par_level_flag, and the second coefficient level flag may be abs_level_gt3_flag or abs_level_gtx_flag.
[0342] Also, for example, the syntax element may include coefficient value related information regarding the value of the transform coefficient of the current block. The coefficient value related information may be abs_remainder and / or dec_abs_level.
[0343] Also, for example, the syntax element may include a sign flag representing the sign of the transform coefficient. The sign flag may be coeff_sign_flag.
[0344] For example, when the syntax of the residual coding for the current block is derived as the TSRC syntax, the decoding device may obtain the residual information of the TSRC syntax for the current block. For example, the residual information of the TSRC syntax may include the syntax elements disclosed in Table 3 above.
[0345] For example, the residual information of the above TSRC syntax may include syntax elements related to the transform coefficients of the current block. Here, the transform coefficient may also be expressed as a residual coefficient.
[0346] For example, the above syntax elements may include context-coded syntax elements and / or bypass-coded syntax elements for the transform coefficients. The above syntax elements may include syntax elements such as sig_coeff_flag, coeff_sign_flag, par_level_flag, abs_level_gtX_flag (e.g., abs_level_gtx_flag[n][0], abs_level_gtx_flag[n][1], abs_level_gtx_flag[n][2], abs_level_gtx_flag[n][3] and / or abs_level_gtx_flag[n][4]), abs_remainder and / or coeff_sign_flag.
[0347] For example, the context-coded syntax element for the transform coefficient may include a valid coefficient flag indicating whether the transform coefficient is a non-zero transform coefficient, a sign flag representing the sign of the transform coefficient, a first coefficient level flag regarding whether the coefficient level of the transform coefficient is greater than a first threshold value, and / or a parity level flag regarding the parity of the coefficient level of the transform coefficient. Also, for example, the context-coded syntax element may include a second coefficient level flag regarding whether the coefficient level of the transform coefficient is greater than a second threshold value, a third coefficient level flag regarding whether the coefficient level of the transform coefficient is greater than a third threshold value, a fourth coefficient level flag regarding whether the coefficient level of the transform coefficient is greater than a fourth threshold value, and / or a fifth coefficient level flag regarding whether the coefficient level of the transform coefficient is greater than a fifth threshold value. Here, the valid coefficient flag may be sig_coeff_flag, the sign flag may be ceff_sign_flag, the first coefficient level flag may be abs_level_gt1_flag, and the parity level flag may be par_level_flag. Also, the second coefficient level flag may be abs_level_gt3_flag or abs_level_gtx_flag, the third coefficient level flag may be abs_level_gt5_flag or abs_level_gtx_flag, the fourth coefficient level flag may be abs_level_gt7_flag or abs_level_gtx_flag, and the fifth coefficient level flag may be abs_level_gt9_flag or abs_level_gtx_flag.
[0348] Also, for example, the syntax element coded by bypass coding for the above conversion coefficient may include coefficient level information regarding the value (or coefficient level) of the above conversion coefficient and / or a sign flag representing the sign for the above conversion coefficient. The above coefficient level information may be abs_remainder and / or dec_abs_level, and the above sign flag may be ceff_sign_flag.
[0349] The decoding device derives the motion information of the current block based on the above prediction-related information (S1340). For example, the decoding device may derive the motion information of the current block based on the inter prediction mode determined based on the above prediction-related information. For example, the decoding device may configure a motion information candidate list for the current block, select one motion information candidate in the motion information candidate list based on the index information of the motion information candidate included in the above prediction-related information, and derive the motion information of the current block based on the selected motion information candidate.
[0350] The decoding device derives the predicted sample of the current block based on the above motion information (S1350). For example, the decoding device may derive the reference picture of the current block based on the index of the reference picture of the current block, and may derive the predicted sample of the current block based on the sample of the reference block pointed to by the motion vector of the current block on the above reference picture. The above motion information may include the index of the above reference picture of the current block and the above motion vector.
[0351] The decoding device derives the residual sample of the current block based on the above residual information (S1360). For example, the decoding device may derive the conversion coefficient of the current block based on the above residual information, and may derive the residual sample of the current block based on the conversion coefficient.
[0352] For example, the decoding device may derive the transform coefficients of the current block based on the syntax elements of the residual information. Thereafter, the decoding device may derive the residual samples of the current block based on the transform coefficients. As an example, when it is derived based on the transform skip flag that no transform is applied to the current block, that is, when the value of the transform skip flag is 1, the decoding device may derive the transform coefficients as the residual samples of the current block. Alternatively, for example, when it is derived based on the transform skip flag that no transform is applied to the current block, that is, when the value of the transform skip flag is 1, the decoding device may inverse-quantize the transform coefficients and derive the residual samples of the current block. Alternatively, for example, when it is derived based on the transform skip flag that a transform is applied to the current block, that is, when the value of the transform skip flag is 0, the decoding device may inverse-transform the transform coefficients and derive the residual samples of the current block. Alternatively, for example, when it is derived based on the transform skip flag that a transform is applied to the current block, that is, when the value of the transform skip flag is 0, the decoding device may inverse-quantize the transform coefficients and inverse-transform the inverse-quantized transform coefficients to derive the residual samples of the current block.
[0353] On one hand, for example, based on the above-mentioned dependent quantization available flag, it can be determined whether the dependent quantization is applied to the current block. For example, when the value of the above-mentioned dependent quantization available flag is 1 (i.e., when the above-mentioned dependent quantization available flag indicates that the dependent quantization is available), the dependent quantization can be applied to the current block. For example, when the dependent quantization is applied to the current block, the decoding device can perform the dependent quantization process on the above-mentioned transform coefficients and derive the residual samples of the current block. That is, for example, when the dependent quantization is applied to the current block, the decoding device can derive the residual samples of the current block based on the dependent quantization of the above-mentioned transform coefficients. For example, when the dependent quantization is applied to the current block, the decoding device can update the state (Qstate) for the dependent quantization based on the coefficient level of the transform coefficient immediately before the current transform coefficient in the scan order, and can derive the coefficient level of the current transform coefficient based on the updated state and the syntax element related to the current transform coefficient, and can inverse quantize the derived coefficient level to derive the residual samples. For example, the above-mentioned current transform coefficient can be inverse quantized based on the quantization parameter for the restoration level of the current transform coefficient in the scalar quantizer for the updated state. Here, the above-mentioned restoration level can be derived based on the syntax element related to the current transform coefficient.
[0354] Also, for example, when the dependent quantization is not applied to the current block, the decoding device can derive the coefficient level of the transform coefficient based on the syntax element related to the transform coefficient of the current block, and can inverse quantize the coefficient level to derive the residual samples. That is, for example, when the dependent quantization is not applied to the current block, the decoding device may not execute the process of updating the state (Qstate) based on the coefficient level of the transform coefficient immediately before the current transform coefficient in the scan order.
[0355] The decoding device generates a restored picture based on the prediction sample and the residual sample (S1370). For example, the decoding device can generate a restored sample and / or a restored picture of the current block based on the prediction sample and the residual sample. For example, the decoding device can generate the restored sample through addition of the prediction sample and the residual sample.
[0356] As described above, hereinafter, if necessary, in-loop filtering procedures such as deblock filtering, SAO, and / or ALF procedures can be applied to the restored picture to improve subjective / objective picture quality.
[0357] FIG. 14 schematically shows a decoding device that performs the image decoding method according to this document. The method disclosed in FIG. 13 can be executed by the decoding device disclosed in FIG. 14. Specifically, for example, the entropy decoding unit of the decoding device in FIG. 14 can execute S1300 to S1330 in FIG. 13, the prediction unit of the decoding device in FIG. 14 can execute S1340 to S1350 in FIG. 13, the residual processing unit of the decoding device in FIG. 14 can execute S1360 in FIG. 13, and the addition unit of the decoding device in FIG. 14 can execute S1370 in FIG. 13.
[0358] According to the foregoing document, the efficiency of residual coding can be increased.
[0359] Also, according to this document, a signaling relationship between the dependent quantization usable flag and the TSRC usable flag can be set, and when the dependent quantization is not usable, the TSRC usable flag can be signaled, through which when TSRC is not usable and the RRC syntax is coded for the transform skip block, the dependent quantization is not used, so that the coding efficiency can be improved, the amount of bits to be coded can be reduced, and the overall residual coding efficiency can be improved.
[0360] Also, according to this document, the TSRC availability flag can be signaled only when dependent quantization is not used, so that the RRC syntax is not coded for the transform skip block and the use of dependent quantization do not overlap and execute, and the TSRC availability flag can be coded more effectively (efficiently) to reduce the number of bits, thereby improving the overall residual coding efficiency.
[0361] In the foregoing embodiments, the method is described based on a flowchart in a series of steps or blocks, but this document is not limited to the order of the steps, and a certain step can occur in a different step and different order or simultaneously from the foregoing. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps are included, or one or more steps of the flowchart can be deleted without affecting the scope of this document.
[0362] The embodiments described in this document can be realized and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each drawing can be realized and executed on a computer, processor, microprocessor, controller, or chip. In this case, the information for realization (for example, information on instructions) or algorithms can be stored on a digital recording medium.
[0363] In addition, the decoding device and the encoding device to which the embodiments of this document are applied can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a recording medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT video (Over The Top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a videophone video device, a transportation means terminal (e.g., a vehicle terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, etc., and can be used to process video signals or data signals. For example, as an OTT video (Over The Top video) device, it can be equipped with a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recoder), etc.
[0364] In addition, the processing method to which the embodiments of this document are applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having the data structure according to this document can also be stored in a computer-readable recording medium. The above computer-readable recording medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The above computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial (Universal Serial) Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. In addition, the above computer-readable recording medium includes a medium realized in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0365] In addition, the embodiments of this document can be realized by a computer program product with program code, and the above program code can be executed by a computer according to the embodiments of this document. The above program code can be stored on a carrier readable by a computer.
[0366] FIG. 15 exemplarily shows a content streaming system structure diagram to which the embodiments of this document are applied.
[0367] The content streaming system to which the embodiments of this document are applied can generally include an encoding server, a streaming server, a web server, a media storage device (repository, storage), a user device, and a multimedia input device.
[0368] The above encoding server compresses the content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and plays the role of transmitting this to the above streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate a bitstream, the above encoding server can be omitted.
[0369] The above bitstream can be generated by an encoding method or a bitstream generation method to which the embodiments of this document are applied, and the above streaming server can temporarily store the above bitstream in the process of transmitting or receiving the above bitstream.
[0370] The above streaming server transmits multimedia data to a user device based on a user request via a web server, and the above web server plays the role of a medium for informing the user of what services are available. When the user requests a desired service from the above web server, the above web server transmits this to the streaming server, and the above streaming server transmits multimedia data to the user. At this time, the above content streaming system can include another control server, and in this case, the above control server plays the role of controlling commands / responses between each device within the above content streaming system.
[0371] The above streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from the above encoding server, the above content can be received in real time. In this case, in order to provide a smooth streaming service, the above streaming server can store the above bitstream for a certain period of time.
[0372] Examples of the above user devices include mobile phones, smartphones, laptop computers, digital broadcast terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), navigation devices, slate PCs, tablet PCs, ULTRABOOKs (registered trademark), wearable devices (e.g., smartwatches (watch-type terminals), smart glasses (glass-type terminals), HMDs (Head Mounted Displays)), digital TVs, desktop computers, digital signatures (signid), etc. Each server in the above content streaming system can be operated as a distributed server. In this case, the data received by each server can be distributedly processed.
[0373] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and realized as a device, and the technical features of the device claims in this specification can be combined and realized as a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and realized as a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and realized as a method.
Claims
1. An image decoding method performed by a decoding device, comprising: obtaining a dependent quantization enable flag regarding whether dependent quantization is enabled; obtaining a TSRC disable flag according to the dependent quantization enable flag, the TSRC disable flag being related to whether TSRC is enabled; obtaining residual information of a residual coding syntax for a current block derived based on the TSRC invalid flag; deriving residual samples for the current block based on the residual information; generating a reconstructed picture based on the residual samples; The image decoding method, wherein the TSRC disable flag is obtained based on the dependent quantization enable flag having a value of 0.
2. 1. An image encoding method performed by an encoding device, comprising: encoding a dependent quantization enabled flag regarding whether dependent quantization is enabled; encoding a TSRC disable flag regarding whether TSRC is available based on the dependent quantization available flag; determining a residual coding syntax for the current block based on the TSRC invalid flag; encoding residual information of the determined residual coding syntax for the current block; generating a bitstream including the dependent quantization enabled flag, the TSRC disabled flag, and the residual information; A method for encoding an image, wherein the TSRC disabled flag is encoded based on the dependent quantization enabled flag having a value of 0.
3. 1. A method of transmission of data including a bitstream of image information, comprising: Obtaining the bitstream of the image information including residual information, the bitstream comprising: encoding a dependent quantization enabled flag regarding whether dependent quantization is enabled; encoding a TSRC disable flag regarding whether TSRC is available based on the dependent quantization available flag; determining a residual coding syntax for the current block based on the TSRC invalid flag; encoding the residual information in the determined residual coding syntax for the current block; generating a bitstream including the dependent quantization enabled flag, the TSRC disabled flag, and the residual information; transmitting the data including the bitstream of the image information including the residual information; A method of transmission, wherein the TSRC disabled flag is encoded based on the dependent quantization enabled flag having a value of 0.
Citation Information
Patent Citations
Image decoding method and apparatus for residual coding
JP7495565B2
Method and apparatus for processing video signals using reduced transform
US20190387241A1
Binarization in transform SKIP residual coding
WO2020263922A1