Image encoding / decoding method and apparatus for signaling chroma component prediction information depending on whether palette mode is applied, and bitstream transmission method
The image encoding/decoding method optimizes chroma component prediction based on palette mode to enhance encoding/decoding efficiency, addressing the increased costs associated with high-resolution images.
Patent Information
- Application Number
- JP2022504309
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-21
- Filing Date
- 2020-07-21
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2040-07-21
AI Technical Summary
The increasing demand for high-resolution, high-quality images leads to a significant increase in transmission and storage costs due to the higher amount of information required, necessitating highly efficient image compression techniques.
An image encoding/decoding method and apparatus that signals chroma component prediction information based on whether a palette mode is applied, allowing for improved encoding/decoding efficiency by optimizing the transmission of bitstreams.
The method enhances encoding/decoding efficiency by optimizing the transmission of bitstreams, enabling cost-effective storage and reproduction of high-resolution, high-quality images.
Smart Images

Figure 0007813696000003 
Figure 0007813696000004 
Figure 0007813696000005
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly to an image encoding / decoding method and apparatus that signal chroma component prediction information depending on whether a palette mode is applied, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure. [Background technology]
[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As image data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.
[0003] This requires highly efficient image compression techniques for effectively transmitting, storing, and reproducing high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]
[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus that improves encoding / decoding efficiency by signaling chroma component prediction information depending on whether or not palette mode is applied.
[0006] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0007] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0008] Another object of the present disclosure is to provide a recording medium storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image.
[0009] The technical problems to be solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not described above will be clearly understood by a person having ordinary skill in the technical field to which the present disclosure pertains from the following description. [Means for solving the problem]
[0010] An image decoding method performed by an image decoding device according to one aspect of the present disclosure may include the steps of: dividing an image to determine a current block; identifying whether a palette mode is applied to the current block based on a palette mode flag obtained from a bitstream; obtaining palette mode encoding information for the current block from the bitstream based on the tree type of the current block and whether a palette mode is applied to the current block; and, if a palette mode is not applied to the current block, obtaining chroma component prediction information for the current block from the bitstream.
[0011] In addition, an image decoding device according to one aspect of the present disclosure includes a memory and at least one processor, wherein the at least one processor divides an image to determine a current block, identifies whether a palette mode is applied to the current block based on a palette mode flag obtained from a bitstream, obtains palette mode encoding information for the current block from the bitstream based on a tree type of the current block and whether a palette mode is applied to the current block, and, if a palette mode is not applied to the current block, obtains chroma component prediction information for the current block from the bitstream.
[0012] In addition, an image encoding method performed by an image encoding device according to one aspect of the present disclosure may include the steps of dividing an image to determine a current block, determining a prediction mode of the current block, encoding a palette mode flag indicating whether the prediction mode of the current block is palette mode or not based on whether the prediction mode of the current block is palette mode or not, encoding palette mode encoding information obtained by encoding the current block in palette mode based on the tree type of the current block and whether the prediction mode of the current block is palette mode or not, and if the prediction mode of the current block is not palette mode, encoding chroma component prediction information of the current block.
[0013] Furthermore, a transmission method according to one aspect of the present disclosure can transmit a bitstream generated by the image encoding device or image encoding method of the present disclosure.
[0014] Furthermore, a computer-readable recording medium according to an aspect of the present disclosure can store a bitstream generated by the image encoding method or image encoding device of the present disclosure.
[0015] The features described above in this brief summary of the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and are not intended to limit the scope of the present disclosure. [Effects of the Invention]
[0016] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0017] Furthermore, according to the present disclosure, an image encoding / decoding method and apparatus can be provided that can improve encoding / decoding efficiency by signaling chroma component prediction information depending on whether or not palette mode is applied.
[0018] The present disclosure also provides a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0019] Furthermore, according to the present disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure can be provided.
[0020] Furthermore, according to the present disclosure, it is possible to provide a recording medium that stores a bitstream that is received by the image decoding device according to the present disclosure, decoded, and used to restore an image.
[0021] The effects obtained by the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those having ordinary skill in the art to which the present disclosure pertains from the following description. [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a diagram illustrating a video coding system to which embodiments of the present disclosure can be applied; [Figure 2] 1 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied. [Figure 3] FIG. 1 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied. [Figure 4]FIG. 2 is a diagram illustrating a division structure of an image according to an embodiment. [Figure 5] FIG. 10 is a diagram showing an example of block division types using a multi-type tree structure. [Figure 6] FIG. 1 illustrates a signaling mechanism for block partition information in a quadtree with nested multi-type tree structure according to the present disclosure. [Figure 7] FIG. 10 illustrates an embodiment in which a CTU is divided into multiple CUs. [Figure 8] FIG. 10 is a diagram illustrating an example of a redundant division pattern. [Figure 9] FIG. 10 illustrates a syntax for chroma format signaling according to one embodiment. [Figure 10] FIG. 10 is a diagram illustrating a chroma format classification table according to one embodiment. [Figure 11] FIG. 2 illustrates horizontal and vertical scanning according to one embodiment. [Figure 12] FIG. 10 illustrates syntax for palette mode according to one embodiment. [Figure 13] FIG. 10 illustrates syntax for palette mode according to one embodiment. [Figure 14] FIG. 10 illustrates syntax for palette mode according to one embodiment. [Figure 15] FIG. 10 illustrates syntax for palette mode according to one embodiment. [Figure 16] FIG. 10 illustrates syntax for palette mode according to one embodiment. [Figure 17] FIG. 10 illustrates syntax for palette mode according to one embodiment. [Figure 18] FIG. 10 illustrates syntax for palette mode according to one embodiment. [Figure 19] FIG. 10 illustrates syntax for palette mode according to one embodiment. [Figure 20]FIG. 10 illustrates a mathematical formula for determining PredictorPaletteEntries and CurrentPaletteEntries according to one embodiment. [Figure 21] FIG. 10 illustrates a modified coding unit syntax according to one embodiment. [Figure 22] 1 is a flowchart illustrating a method for signaling predetermined chroma intra prediction information according to one embodiment. [Figure 23] 10 is a flowchart illustrating a method for a decoding device to obtain chroma prediction information according to an embodiment. [Figure 24] 1 is a flowchart illustrating a method for encoding an image by an encoding device according to an embodiment. [Figure 25] 10 is a flowchart illustrating a method for decoding an image by a decoding device according to an embodiment. [Figure 26] FIG. 1 illustrates a content streaming system to which an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE INVENTION
[0023] The present disclosure will be described in detail below with reference to the accompanying drawings, so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.
[0024] In describing the embodiments of the present disclosure, if it is determined that a detailed description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, in the drawings, parts that are not related to the description of the present disclosure will be omitted, and similar parts will be designated by similar reference numerals.
[0025] In this disclosure, when a component is referred to as being "coupled," "coupled," or "connected" to another component, this includes not only a direct connection, but also an indirect connection where another component exists between them. Furthermore, when a component is referred to as "including" or "having" another component, this does not mean that the other component is excluded, but that the component can further include the other component, unless otherwise specified.
[0026] In this disclosure, terms such as "first" and "second" are used only to distinguish one component from another component, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be called a second component in another embodiment, and similarly, a second component in one embodiment may be called a first component in another embodiment.
[0027] In this disclosure, components that are distinguished from one another are used to clearly describe the characteristics of each component and do not necessarily mean that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not otherwise specified, such integrated or distributed embodiments are also included within the scope of this disclosure.
[0028] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also within the scope of this disclosure. Furthermore, an embodiment including other components in addition to the components described in various embodiments is also within the scope of this disclosure.
[0029] The present disclosure relates to image encoding and decoding, and terms used in this disclosure may have their ordinary meaning in the technical field to which the present disclosure belongs unless they are newly defined in this disclosure.
[0030] In this disclosure, a "picture" generally refers to a unit representing any one image in a specific time period, and a slice / tile is a coding unit constituting a part of a picture, and one picture may be composed of one or more slices / tiles. Furthermore, a slice / tile may include one or more coding tree units (CTUs).
[0031] In this disclosure, "pixel" or "pel" may refer to the smallest unit constituting one picture (or image). Also, "sample" may be used as a term corresponding to pixel. A sample may generally indicate a pixel or a pixel value, may indicate only a pixel / pixel value of a luma component, or may indicate only a pixel / pixel value of a chroma component.
[0032] In this disclosure, the term "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. The term "unit" may be used interchangeably with terms such as "sample array," "block," or "area," depending on the situation. In general, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0033] In the present disclosure, a "current block" may refer to any one of a "current coding block," a "current coding unit," a "block to be coded," a "block to be decoded," or a "block to be processed." When prediction is performed, a "current block" may refer to a "current predicted block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, a "current block" may refer to a "current transformed block" or a "block to be transformed." When filtering is performed, a "current block" may refer to a "block to be filtered."
[0034] Furthermore, in this disclosure, "current block" may mean "luma block of the current block" unless explicitly stated as a chroma block. "Chroma block of the current block" may be expressed explicitly including the explicit description of a chroma block, such as "chroma block" or "current chroma block."
[0035] In the present disclosure, " / " and "," can be interpreted as "and / or." For example, "A / B" and "A, B" can be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."
[0036] In this disclosure, "or" can be interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." Alternatively, in this disclosure, "or" can mean "additionally or alternatively."
[0037] Video Coding System Overview
[0038] FIG. 1 is a diagram illustrating a video coding system according to this disclosure.
[0039] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.
[0040] An encoding device 10 according to an embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment may include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The reception unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, which may be configured as a separate device or an external component.
[0041] The video source generation unit 11 can acquire video / images through a video / image capture, synthesis, or generation process. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated via a computer, etc. In this case, the video / image capture process can be replaced with a process in which related data is generated.
[0042] The encoder 12 may encode the input video / image. The encoder 12 may perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoder 12 may output the encoded data (encoded video / image information) in a bitstream format.
[0043] The transmitter 13 may transmit the encoded video / image information or data output in a bitstream format to the receiver 21 of the decoding device 20 in a file or streaming format via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. The transmitter 13 may include elements for generating a media file in a predetermined file format and elements for transmitting via a broadcasting / communication network. The receiver 21 may extract / receive the bitstream from the storage medium or network and transmit it to the decoder 22.
[0044] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding unit 12.
[0045] The rendering unit 23 can render the decoded video / images, and the rendered video / images can be displayed via the display unit.
[0046] Overview of the image encoding device
[0047] FIG. 2 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied.
[0048] 2, the image encoding device 100 may include an image division unit 110, a subtraction unit 115, a transform unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transform unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 may be collectively referred to as a "prediction unit." The transform unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transform unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.
[0049] Depending on the embodiment, all or at least some of the components constituting the image encoding device 100 may be realized by a single hardware component (e.g., an encoder or a processor). Also, the memory 170 may include a decoded picture buffer (DPB) and may be realized by a digital storage medium.
[0050] The image division unit 110 may divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing units may be called coding units (CUs). The coding units may be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) using a QT / BT / TT (quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit may be divided into multiple coding units at deeper depths based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. To divide the coding units, the quad-tree structure may be applied first, and then the binary-tree structure and / or the ternary-tree structure may be applied later. The coding procedure according to the present disclosure may be performed based on the final coding unit that is not further divided. The maximum coding unit may be used as the final coding unit, or a lower-depth coding unit obtained by dividing the maximum coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or reconstruction, which will be described later. As another example, a processing unit of the coding procedure may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0051] The prediction unit (inter prediction unit 180 or intra prediction unit 185) may perform prediction on a current block (current block) to generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block or CU. The prediction unit may generate various information related to prediction of the current block and transmit it to the entropy coding unit 190. The prediction information may be coded by the entropy coding unit 190 and output in a bitstream format.
[0052] The intra prediction unit 185 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block according to the intra prediction mode and / or intra prediction technique. The intra prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of precision of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0053] The inter prediction unit 180 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation between the motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 180 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter predictor 180 may use motion information of neighboring blocks as motion information for the current block. In the case of skip mode, unlike in merge mode, a residual signal may not be transmitted.In the case of a motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.
[0054] The predictor may generate a prediction signal based on various prediction methods and / or prediction techniques, which will be described later. For example, the predictor may apply intra prediction or inter prediction to predict the current block, or may simultaneously apply intra prediction and inter prediction. A prediction method that simultaneously applies intra prediction and inter prediction to predict the current block may be referred to as combined inter and intra prediction (CIIP). The predictor may also perform intra block copy (IBC) to predict the current block. Intra block copy can be used for content image / video coding, such as screen content coding (SCC), for games. IBC is a method of predicting a current block using an already reconstructed reference block in a current picture that is located a predetermined distance away from the current block. When IBC is applied, the position of the reference block in the current picture may be coded as a vector (block vector) corresponding to the predetermined distance. IBC is essentially performed within the current picture, but may be similar to inter prediction in that a reference block is derived within the current picture. That is, the IBC may use at least one of the inter prediction techniques described in this disclosure.
[0055] The prediction signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 may subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal may be transmitted to the conversion unit 120.
[0056] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph representing inter-pixel relationship information. The CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or to non-square blocks of variable size.
[0057] The quantization unit 130 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy coding unit 190. The entropy coding unit 190 may encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream format. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 130 may rearrange the quantized transform coefficients in a block format into a one-dimensional vector format based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector format.
[0058] The entropy coding unit 190 may perform various coding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy coding unit 190 may also code information necessary for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized transform coefficients. The coded information (e.g., coded video / image information) may be transmitted or stored in a bitstream format in network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The signaling information, transmitted information and / or syntax elements mentioned in this disclosure may be encoded through the above-described encoding procedure and included in the bitstream.
[0059] The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) that transmits and / or a storing unit (not shown) that stores the signal output from the entropy encoding unit 190 may be provided as an internal / external element of the image encoding device 100, or the transmitting unit may be provided as a component of the entropy encoding unit 190.
[0060] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150.
[0061] The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the current block to be processed, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The adder 155 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as will be described later.
[0062] The filtering unit 160 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 160 may generate various information related to filtering and transmit it to the entropy coding unit 190, as will be described later in connection with each filtering method. The filtering information may be coded by the entropy coding unit 190 and output in a bitstream format.
[0063] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 can avoid a prediction mismatch between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.
[0064] The DPB in the memory 170 may store modified reconstructed pictures for use as reference pictures in the inter predictor 180. The memory 170 may store motion information of blocks from which motion information in the current picture is derived (or coded) and / or motion information of already reconstructed intra-picture blocks. The stored motion information may be transmitted to the inter predictor 180 to be used as motion information of spatially surrounding blocks or temporally surrounding blocks. The memory 170 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 185.
[0065] Overview of the image decoding device
[0066] FIG. 3 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied.
[0067] 3, the image decoding apparatus 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit." The inverse quantization unit 220 and the inverse transform unit 230 may be included in a residual processing unit.
[0068] Depending on the embodiment, all or at least some of the components constituting the image decoding device 200 may be realized by a single hardware component (e.g., a decoder or a processor). Also, the memory 170 may include a DPB and may be realized by a digital storage medium.
[0069] The image decoding device 200, which receives a bitstream including video / image information, can reconstruct an image by performing a process corresponding to the process performed by the image encoding device 100 of FIG. 1. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).
[0070] The image decoding apparatus 200 may receive a signal output from the image encoding apparatus of FIG. 2 in a bitstream format. The received signal may be decoded via an entropy decoding unit 210. For example, the entropy decoding unit 210 may parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The image decoding apparatus may further use the information on the parameter sets and / or the general constraint information to decode an image. The signaling information, received information, and / or syntax elements referred to in the present disclosure may be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using information on the syntax element to be decoded and decoded information on neighboring blocks and the block to be decoded, or information on symbols / bins decoded in a previous step, predicts the occurrence probability of the bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. After determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and residual values entropy decoded by the entropy decoding unit 210, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, information related to filtering may be provided to the filtering unit 240. Meanwhile, a receiving unit (not shown) for receiving a signal output from the image encoding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.
[0071] Meanwhile, the image decoding apparatus according to the present disclosure may be referred to as a video / image / picture decoding apparatus. The image decoding apparatus may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265.
[0072] The inverse quantization unit 220 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 220 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the image encoding device. The inverse quantization unit 220 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0073] The inverse transform unit 230 can inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0074] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 210, and may determine a specific intra / inter prediction mode (prediction technique).
[0075] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described below, as described in the description of the prediction unit of the image encoding device 100.
[0076] The intra predictor 265 may predict the current block by referring to samples in the current picture. The description of the intra predictor 185 may also be applied to the intra predictor 265.
[0077] The inter prediction unit 260 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on correlations between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 260 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes (techniques), and the prediction information may include information indicating the inter prediction mode (technique) for the current block.
[0078] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the current block to be processed, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The description of the adder 155 may also be applied to the adder 235. The adder 235 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, and may also be used for inter prediction of the next picture after undergoing filtering, as will be described later.
[0079] The filtering unit 240 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 250, specifically, in a DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0080] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter predictor 260. The memory 250 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 260 to be used as motion information of a spatially surrounding block or a temporally surrounding block. The memory 250 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 265.
[0081] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180 and intra prediction unit 185 of the image encoding device 100 can also be applied in a similar or corresponding manner to the filtering unit 240, inter prediction unit 260 and intra prediction unit 265 of the image decoding device 200, respectively.
[0082] Image Segmentation Overview
[0083] The video / image coding method according to the present disclosure may be performed based on the following image partition structure. Specifically, procedures such as prediction, residual processing (e.g., (inverse) transform, (inverse) quantization), syntax element coding, and filtering, which will be described later, may be performed based on CTUs and CUs (and / or TUs and PUs) derived based on the image partition structure. An image may be divided into blocks, and the block partition procedure may be performed by the image partitioning unit 110 of the encoding device described above. Partition-related information may be coded by the entropy coding unit 190 and transmitted to the decoding device in the form of a bitstream. The entropy decoding unit 210 of the decoding device may derive a block partition structure of the current picture based on the partition-related information obtained from the bitstream, and perform a series of procedures for image decoding (e.g., prediction, residual processing, block / picture reconstruction, in-loop filtering, etc.) based on the block partition structure.
[0084] A picture can be divided into a sequence of coding tree units (CTUs). FIG. 4 shows an example of dividing a picture into CTUs. A CTU can correspond to a coding tree block (CTB). Alternatively, a CTU can include a coding tree block of luma samples and two coding tree blocks of corresponding chroma samples. For example, for a picture containing three sample arrays, a CTU can include an N×N block of luma samples and two corresponding blocks of chroma samples.
[0085] Overview of CTU division
[0086] As described above, a coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) according to a QT / BT / TT (quad-tree / binary-tree / ternary-tree) structure. For example, a CTU can be first divided into a quad-tree structure. Then, the leaf nodes of the quad-tree structure can be further divided according to a multi-type tree structure.
[0087] Quadtree division refers to dividing the current CU (or CTU) into four equal parts. By quadtree division, the current CU can be divided into four CUs with the same width and height. If the current CU is not further divided into a quadtree structure, the current CU corresponds to a leaf node of the quadtree structure. A CU corresponding to a leaf node of the quadtree structure is not further divided and can be used as the final coding unit described above. Alternatively, a CU corresponding to a leaf node of the quadtree structure can be further divided into four parts according to a multi-type tree structure.
[0088] 5 is a diagram showing the types of division of blocks using a multi-type tree structure. Division using a multi-type tree structure can include two divisions using a binary tree structure and two divisions using a ternary tree structure.
[0089] The two divisions based on the binary tree structure can include vertical binary splitting (SPLIT_BT_VER) and horizontal binary splitting (SPLIT_BT_HOR). Vertical binary splitting (SPLIT_BT_VER) refers to a division that divides the current CU into two equal parts vertically. As shown in FIG. 4, vertical binary splitting can generate two CUs with the same height as the current CU and half the width of the current CU. Horizontal binary splitting (SPLIT_BT_HOR) refers to a division that divides the current CU into two equal parts horizontally. As shown in FIG. 5, horizontal binary splitting can generate two CUs with the same height as the current CU and half the width of the current CU.
[0090] The two divisions based on the ternary tree structure can include vertical ternary splitting (SPLIT_TT_VER) and horizontal ternary splitting (SPLIT_TT_HOR). Vertical ternary splitting (SPLIT_TT_VER) divides the current CU vertically at a ratio of 1:2:1. As shown in FIG. 5, vertical ternary splitting can generate two CUs each having the same height as the current CU and a width equal to one-quarter of the current CU's width, and a CU each having the same height as the current CU and a width equal to half the current CU's width. Horizontal ternary splitting (SPLIT_TT_HOR) divides the current CU horizontally at a ratio of 1:2:1. As shown in FIG. 4, horizontal ternary splitting can generate two CUs each having a height equal to one-quarter of the current CU's height and a width equal to the current CU's width.
[0091] FIG. 6 is a diagram illustrating a signaling mechanism for block partition information in a quadtree with nested multi-type tree structure according to the present disclosure.
[0092] Here, the CTU is treated as the root node of the quadtree, and the CTU is first split into a quadtree structure. Information (e.g., qt_split_flag) indicating whether quadtree splitting is performed on the current CU (CTU or quadtree node (QT_node)) can be signaled. For example, if qt_split_flag is a first value (e.g., '1'), the current CU can be split into a quadtree. Also, if qt_split_flag is a second value (e.g., '0'), the current CU is not split into a quadtree but becomes a quadtree leaf node (QT_leaf_node). The leaf nodes of each quadtree can then be further split into a multitype tree structure. That is, the leaf nodes of the quadtree can become multitype tree nodes (MTT_node). In a multi-type tree structure, a first flag (e.g., mtt_split_cu_flag) may be signaled to indicate whether the current node is to be further split. If the node is to be further split (e.g., if the first flag is 1), a second flag (e.g., mtt_split_cu_vertical_flag) may be signaled to indicate the splitting direction. For example, if the second flag is 1, the splitting direction may be vertical, and if the second flag is 0, the splitting direction may be horizontal. Then, a third flag (e.g., mtt_split_cu_binary_flag) may be signaled to indicate whether the splitting type is a binary splitting type or a ternary splitting type. For example, if the third flag is 1, the splitting type may be a binary splitting type, and if the third flag is 0, the splitting type may be a ternary splitting type. The nodes of a multitype tree obtained by binary or ternary splitting can be further partitioned into a multitype tree structure, but the nodes of a multitype tree cannot be partitioned into a quadtree structure.If the first flag is 0, the corresponding node of the multitype tree is not further divided and becomes a leaf node (MTT_leaf_node) of the multitype tree. The CU corresponding to the leaf node of the multitype tree can be used as the final coding unit described above.
[0093] Based on the above mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree splitting mode (MttSplitMode) of the CU can be derived as shown in Table 1. In the following description, the multi-type tree splitting mode can be abbreviated as multi-tree split type or split type.
[0094] [Table 1]
[0095] FIG. 7 illustrates an example in which a CTU is divided into multiple CUs by applying a multi-type tree after applying a quadtree. In FIG. 7, a bold block edge 710 indicates quadtree division, and the remaining edge 720 indicates multi-type tree division. A CU may correspond to a coding block (CB). In one embodiment, a CU may include a coding block of luma samples and two coding blocks of chroma samples corresponding to the luma samples. The CB or TB size of a chroma component (sample) may be derived based on the CB or TB size according to the component ratio according to the color format of a picture / image (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). If the color format is 4:4:4, the CB / TB size of the chroma component may be set to be the same as the CB / TB size of the luma component. If the color format is 4:2:2, the width of the chroma components CB / TB can be set to half the width of the luma components CB / TB, and the height of the chroma components CB / TB can be set to the height of the luma components CB / TB. If the color format is 4:2:0, the width of the chroma components CB / TB can be set to half the width of the luma components CB / TB, and the height of the chroma components CB / TB can be set to half the height of the luma components CB / TB.
[0096] In one embodiment, when the size of the CTU is 128 based on the luma sample unit, the size of the CU can range from 128 x 128, which is the same size as the CTU, to 4 x 4. In one embodiment, in the case of a 4:2:0 color format (or chroma format), the chroma CB size can range from 64 x 64 to 2 x 2.
[0097] Meanwhile, in one embodiment, the CU size and the TU size may be the same, or multiple TUs may exist within a CU region. The TU size may generally refer to the luma component (sample) TB (Transform Block) size.
[0098] The TU size may be derived based on a preset maximum allowable TB size (maxTbSize). For example, if the CU size is larger than the maxTbSize, multiple TUs (TBs) having the maxTbSize may be derived from the CU, and transform / inverse transform may be performed in units of the TUs (TBs). For example, the maximum allowable luma TB size may be 64x64, and the maximum allowable chroma TB size may be 32x32. If the width or height of a CB divided by the tree structure is larger than the maximum transform width or height, the CB may be automatically (or implicitly) divided until the horizontal and vertical TB size constraints are satisfied.
[0099] Also, for example, when intra prediction is applied, the intra prediction mode / type may be derived in units of the CU (or CB), and the procedure of deriving neighboring reference samples and generating predicted samples may be performed in units of TU (or TB). In this case, one or more TUs (or TBs) may exist within one CU (or CB) region, and in this case, the multiple TUs (or TBs) may share the same intra prediction mode / type.
[0100] Meanwhile, for a quadtree coding tree scheme with a multitype tree, the following parameters can be signaled from the encoding device to the decoding device as SPS syntax elements. For example, at least one of CTUsize, a parameter indicating the size of the root node of a quadtree, MinQTSize, a parameter indicating the minimum allowable size of a leaf node of a quadtree, MaxBTSize, a parameter indicating the maximum allowable size of a root node of a binary tree, MaxTTSize, a parameter indicating the maximum allowable size of a root node of a ternary tree, MaxMttDepth, a parameter indicating the maximum allowed hierarchy depth of a multitype tree split from a leaf node of a quadtree, MinBtSize, a parameter indicating the minimum allowable leaf node size of a binary tree, and MinTtSize, a parameter indicating the minimum allowable leaf node size of a ternary tree, can be signaled.
[0101] In one embodiment using the 4:2:0 chroma format, the CTU size may be set to a 128x128 luma block and two 64x64 chroma blocks corresponding to the luma block. In this case, MinQTSize may be set to 16x16, MaxBtSize may be set to 128x128, MaxTtSzie may be set to 64x64, MinBtSize and MinTtSize may be set to 4x4, and MaxMttDepth may be set to 4. Quadtree partitioning may be applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes may be called leaf QT nodes. The quadtree leaf nodes may have a size from 16x16 (e.g., the MinQTSize) to 128x128 (e.g., the CTU size). If the leaf QT node is 128x128, it may not be further partitioned into a binary tree or a ternary tree. This is because even if the partitioning is performed in this case, it would exceed MaxBtSize and MaxTtszie (e.g., 64 x 64). In other cases, the leaf QT node can be further partitioned into a multitype tree. Thus, the leaf QT node is the root node for the multitype tree, and the leaf QT node can have a multitype tree depth (mttDepth) value of 0. If the multitype tree depth reaches MaxMttDepth (e.g., 4), no further subdivisions can be considered. If the width of the multitype tree node is equal to MinBtSize and equal to or less than 2 x MinTtSize, no further horizontal subdivisions can be considered. If the height of the multitype tree node is equal to MinBtSize and equal to or less than 2 x MinTtSize, no further vertical subdivisions can be considered. If no subdivision is considered in this way, the encoding device can omit signaling of the subdivision information. In such cases, the decoding device can guide the subdivision information to a predetermined value.
[0102] Meanwhile, one CTU may include a coding block of luma samples (hereinafter referred to as a "luma block") and two coding blocks of corresponding chroma samples (hereinafter referred to as "chroma blocks"). The above-mentioned coding tree scheme may be applied equally to the luma blocks and chroma blocks of the current CU, or may be applied separately. Specifically, the luma blocks and chroma blocks in one CTU may be divided into the same block tree structure, which may be referred to as a single tree (SINGLE_TREE). Alternatively, the luma blocks and chroma blocks in one CTU may be divided into separate block tree structures, which may be referred to as a dual tree (DUAL_TREE). In other words, when a CTU is divided into a dual tree, a block tree structure for the luma blocks and a block tree structure for the chroma blocks may exist separately. In this case, the block tree structure for the luma block may be referred to as a dual tree luma (DUAL_TREE_LUMA), and the block tree structure for the chroma block may be referred to as a dual tree chroma (DUAL_TREE_CHROMA). For P and B slices / tile groups, the luma block and the chroma block in one CTU may be restricted to have the same coding tree structure. However, for I slices / tile groups, the luma block and the chroma block may have separate block tree structures. If a separate block tree structure is applied, the luma coding tree block (CTB) may be divided into CUs based on a specific coding tree structure, and the chroma CTB may be divided into chroma CUs based on another coding tree structure. That is, this may mean that a CU in an I slice / tile group to which a separate block tree structure is applied may be composed of a coding block of a luma component or a coding block of two chroma components, and a CU in a P or B slice / tile group may be composed of blocks of three color components (a luma component and two chroma components).
[0103] Although the quadtree coding tree structure with a multi-type tree has been described above, the structure in which a CU is divided is not limited to this. For example, the BT structure and the TT structure may be interpreted as concepts included in a Multiple Partitioning Tree (MPT) structure, and a CU may be interpreted as being divided by a QT structure and an MPT structure. In an example in which a CU is divided by a QT structure and an MPT structure, the division structure may be determined by signaling a syntax element (e.g., MPT_split_type) containing information on whether a leaf node of the QT structure is divided into several blocks and a syntax element (e.g., MPT_split_mode) containing information on whether the leaf node of the QT structure is divided vertically or horizontally.
[0104] In another example, CUs may be divided in a manner different from that of the QT structure, BT structure, or TT structure. That is, unlike the QT structure in which lower-depth CUs are divided into 1 / 4 the size of higher-depth CUs, the BT structure in which lower-depth CUs are divided into 1 / 2 the size of higher-depth CUs, or the TT structure in which lower-depth CUs are divided into 1 / 4 or 1 / 2 the size of higher-depth CUs, lower-depth CUs may be divided into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 the size of higher-depth CUs, as the case may be, and the manner in which CUs are divided is not limited thereto.
[0105] In this way, the quadtree coding block structure with the multi-type tree can provide a very flexible block partitioning structure. Meanwhile, due to the partitioning types supported by the multi-type tree, different partitioning patterns can potentially result in the same coding block structure in some cases. By limiting the occurrence of such redundant partitioning patterns, the encoding device and the decoding device can reduce the amount of data for partitioning information.
[0106] For example, FIG. 8 exemplarily illustrates redundant partitioning patterns that may occur in binary tree partitioning and ternary tree partitioning. As shown in FIG. 8, consecutive binary partitions 810 and 820 in one direction at two-step levels have the same coding block structure as the binary partitioning for the center partition after ternary partitioning. In this case, binary tree partitioning for the center blocks 830 and 840 of the ternary tree partitioning can be prohibited. This prohibition can be applied to all CUs of a picture. When such a specific partitioning is prohibited, the signaling of the corresponding syntax element can be modified to reflect this prohibition, thereby reducing the number of bits signaled for the partitioning. For example, as in the example shown in FIG. 8, when binary tree partitioning for the center block of a CU is prohibited, the mtt_split_cu_binary_flag syntax element, which indicates whether the partitioning is binary or ternary, is not signaled and can be set to 0 by the decoder.
[0107] Chroma Format Overview
[0108] The following describes a chroma format. An image can be encoded with encoding data including a luma component (e.g., Y) array and two chroma component (e.g., Cb, Cr) arrays. For example, one pixel of the encoded image can include a luma sample and a chroma sample. The term chroma format can be used to indicate the configuration format of the luma sample and the chroma sample, and the chroma format is sometimes called a color format.
[0109] In one embodiment, an image can be encoded in various chroma formats, such as monochrome, 4:2:0, 4:2:2, and 4:4:4. In monochrome sampling, there can be one sample array, which can be a luma array. In 4:2:0 sampling, there can be one luma sample array and two chroma sample arrays, each of which can be half the height and half the width of the luma array. In 4:2:2 sampling, there can be one luma sample array and two chroma sample arrays, each of which can be the same height and half the width of the luma array. In 4:4:4 sampling, there can be one luma sample array and two chroma sample arrays, each of which can be the same height and width as the luma array.
[0110] For example, in the case of 4:2:0 sampling, the position of a chroma sample can be located at the bottom of the corresponding luma sample. In the case of 4:2:2 sampling, the chroma sample can be located overlapping the position of the corresponding luma sample. In the case of 4:4:4 sampling, both the luma sample and the chroma sample can be located overlapping.
[0111] The chroma format used in the encoding device and the decoding device may be predetermined. Alternatively, the chroma format may be signaled from the encoding device to the decoding device for adaptive use in the encoding device and the decoding device. In one embodiment, the chroma format may be signaled based on at least one of chroma_format_idc and separate_colour_plane_flag. At least one of chroma_format_idc and separate_colour_plane_flag may be signaled via a higher-level syntax such as DPS, VPS, SPS, or PPS. For example, chroma_format_idc and separate_colour_plane_flag may be included in the SPS syntax as shown in FIG. 9.
[0112] Meanwhile, FIG. 10 shows an example of chroma format classification using signaling of chroma_format_idc and separate_colour_plane_flag. chroma_format_idc may be information indicating the chroma format applied to the coded image. separate_colour_plane_flag may indicate whether the color array is processed separately in a specific chroma format. For example, a first value (e.g., 0) of chroma_format_idc may indicate monochrome sampling. A second value (e.g., 1) of chroma_format_idc may indicate 4:2:0 sampling. A third value (e.g., 2) of chroma_format_idc may indicate 4:2:2 sampling. A fourth value (e.g., 3) of chroma_format_idc may indicate 4:4:4 sampling.
[0113] For 4:4:4 sampling, the following applies depending on the value of separate_colour_plane_flag: If the value of separate_colour_plane_flag is the first value (e.g., 0), each of the two chroma arrays can have the same height and width as the luma array. In this case, the value of ChromaArrayType, which indicates the type of chroma sample array, can be set to the same as chroma_format_idc. If the value of separate_colour_plane_flag is the second value (e.g., 1), the luma, Cb, and Cr sample arrays can be processed separately, thereby being processed in the same way as a monochrome sampled picture. In this case, ChromaArrayType can be set to 0.
[0114] Intra prediction for chroma blocks
[0115] When intra prediction is performed on a current block, prediction can be performed on the luma component block (luma block) and prediction can be performed on the chroma component block (chroma block) of the current block, and in this case, the intra prediction mode for the chroma block can be set separately from the intra prediction mode for the luma block.
[0116] For example, the intra-prediction mode for a chroma block may be indicated based on intra-chroma prediction mode information, which may be signaled in the form of an intra_chroma_pred_mode syntax element. For example, the intra-chroma prediction mode information may indicate one of a planar mode, a DC mode, a vertical mode, a horizontal mode, a derived mode (DM), and a cross-component linear model (CCLM) mode. Here, the planar mode may indicate intra-prediction mode 0, the DC mode may indicate intra-prediction mode 1, the vertical mode may indicate intra-prediction mode 26, and the horizontal mode may indicate intra-prediction mode 10. DM may also be referred to as a direct mode. CCLM may also be referred to as a linear model (LM). The CCLM mode may include one of L_CCLM, T_CCLM, and LT_CCLM.
[0117] Meanwhile, DM and CCLM are dependent intra prediction modes that predict a chroma block using information of a luma block. DM may indicate a mode in which the same intra prediction mode as the intra prediction mode for the luma component is applied as the intra prediction mode for the chroma component. CCLM may indicate an intra prediction mode in which, in generating a prediction block for a chroma block, reconstructed samples of a luma block are subsampled, and then CCLM parameters α and β are applied to the subsampled samples to use the derived samples as prediction samples for the chroma block.
[0118] CCLM (Cross-component linear model) mode
[0119] As described above, the CCLM mode can be applied to a chroma block. The CCLM mode is an intra prediction mode that uses correlation between a luma block and a chroma block corresponding to the luma block, and is performed by deriving a linear model based on surrounding samples of the luma block and surrounding samples of the chroma block. Then, predicted samples of the chroma block can be derived based on the derived linear model and reconstructed samples of the luma block.
[0120] Specifically, when a CCLM mode is applied to a current chroma block, parameters for a linear model may be derived based on neighboring samples used for intra prediction of the current chroma block and neighboring samples used for intra prediction of the current luma block. For example, the linear model for CCLM may be expressed based on the following equation:
[0121]
number
[0122] where pred c (i, j) may represent a predicted sample at the (i, j) coordinate of the current chroma block in the current CU. L (i, j) may represent a reconstructed sample at the (i, j) coordinate of the current luma block in the CU. For example, rec L '(i,j) may represent a down-sampled reconstruction sample of the current luma block. The linear model coefficients α and β may be signaled or derived from neighboring samples.
[0123] Palette Mode Overview
[0124] Palette mode will now be described. According to an embodiment, an encoding device may encode an image using the palette mode, and a decoding device may decode the image using the palette mode in a corresponding manner. The palette mode may be referred to as a palette coding mode, an intra palette mode, an intra palette coding mode, etc. The palette mode may be considered as a type of intra coding mode and may also be considered as one of the intra prediction methods. However, similar to the above-described skip mode, a separate residual value for the corresponding block may not be signaled.
[0125] In one embodiment, palette mode can be used to improve coding efficiency when encoding screen content, which is a computer-generated image containing a significant amount of text and graphics. Typically, local areas of an image generated by screen content are separated by sharp edges and represented with a small number of colors. To take advantage of this characteristic, palette mode can represent samples for a block with an index that points to a color entry in a palette table.
[0126] To apply the palette mode, information for a palette table can be signaled. In one embodiment, the palette table can include index values corresponding to each color. To signal the index values, palette index prediction information can be signaled. The palette index prediction information can include index values for at least a portion of a palette index map. The palette index map can map pixels of the video data to color indices in the palette table.
[0127] The palette index prediction information may include run value information. For at least a portion of the palette index map, the run value information may be information that associates run values with index values. One run value may be associated with an escape color index. The palette index map may be generated from the palette index prediction information. For example, at least a portion of the palette index map may be generated by determining whether to adjust the index value of the palette index prediction information based on the final index value.
[0128] The current block in the current picture can be coded or reconstructed according to a palette index map. When the palette mode is applied, pixel values in the current coding unit can be represented by a small set of representative color values. Such a set can be named a palette. For pixels having values close to the palette colors, a palette index can be signaled. For pixels having values that do not belong to the palette (outliers), the pixel can be represented by an escape symbol, and the quantized pixel value can be directly signaled. In this specification, a pixel or pixel value can be described by a sample.
[0129] To decode a block coded in palette mode, a decoding device can decode palette colors and indices. The palette colors can be described as a palette table and coded using a palette table coding tool. An escape flag can be signaled for each coding unit. The escape flag can indicate whether an escape symbol exists in the current coding unit. If an escape symbol exists, the palette table is incremented by one unit (e.g., index unit), and the last index can be designated as escape mode. The palette indices of all pixels for one coding unit can constitute a palette index map and can be coded using a palette index map coding tool.
[0130] For example, to encode the palette table, a palette predictor can be maintained. The palette predictor can be initialized at each slice start point. For example, the palette predictor can be reset to 0. For each entry of the palette predictor, a reuse flag can be signaled to indicate whether it is currently part of the palette. The reuse flag can be signaled using run-length coding of 0 values.
[0131] Then, the numbers for the new palette entries can be signaled using a zeroth-order exponential-Golomb code. Finally, the component values for the new palette entries can be signaled. After encoding the current coding unit, the palette predictor can be updated with the current palette, and entries from the previous palette predictor that are not reused in the current palette can be appended to the end of the new palette predictor (until the maximum allowed size is reached); this can be called palette stuffing.
[0132] For example, to encode a palette index map, the indices can be coded using horizontal or vertical scanning. The scan order can be signaled via the bitstream using a parameter, palette_transpose_flag, indicating the scan direction. For example, if horizontal scanning is applied to scan the indices for samples in the current coding unit, palette_transpose_flag can have a first value (e.g., 0), and if vertical scanning is applied, palette_transpose_flag can have a second value (e.g., 1). Figure 11 shows an example of horizontal scanning and vertical scanning according to one embodiment.
[0133] Also, in one embodiment, the palette index can be coded using the "INDEX" and "COPY_ABOVE" modes. The two modes can be signaled using one flag, except when the palette index mode is signaled for the top row when horizontal scanning is used, when the palette index mode is signaled for the leftmost column when vertical scanning is used, and when the previous mode is "COPY_ABOVE".
[0134] In "INDEX" mode, the palette index can be signaled explicitly. For "INDEX" and "COPY_ABOVE" modes, the same mode can be used to signal a run value indicating the number of coded pixels.
[0135] The coding order for the index map can be set as follows: First, the number of index values for a coding unit can be signaled. This can be done after signaling the actual index values for the entire coding unit using truncated binary coding. Both the number of indexes and the index values can be coded in bypass mode, which allows bypass bins associated with the index to be grouped. Then, the palette mode (INDEX or COPY_ABOVE) and run value can be signaled in an interleaved manner.
[0136] Finally, component escape values corresponding to escape samples for the entire coding unit can be grouped together and coded in bypass mode. An additional syntax element, last_run_type_flag, can be signaled after signaling the index value. By using last_run_type_flag together with the index number, it is possible to omit signaling the run value corresponding to the last run in the block.
[0137] In one embodiment, a dual-tree type that performs independent coding unit partitioning for the luma and chroma components can be used for an I slice. Palette mode can be applied to the luma and chroma components individually or together. If a dual-tree is not applied, palette mode can be applied to all of the Y, Cb, and Cr components.
[0138] In one embodiment, signaling of syntax elements for palette mode can be coded and signaled as shown in Figures 12 to 19. Figures 12 and 13 show the successive syntax in a coding unit (CU) for palette mode, and Figures 14 to 19 show the successive syntax for palette mode.
[0139] Each syntax element will be described below. A palette mode flag, pred_mode_plt_flag, may indicate whether palette mode is applied to the current coding unit. For example, a first value (e.g., 0) of pred_mode_plt_flag may indicate that palette mode is not applied to the current coding unit. A second value (e.g., 1) of pred_mode_plt_flag may indicate that palette mode is applied to the current coding unit. If pred_mode_plt_flag is not obtained from the bitstream, the value of pred_mode_plt_flag may be determined to be the first value.
[0140] The parameter PredictorPaletteSize[startComp] can indicate the size of the predictor palette for startComp, which is the first color component in the current palette table.
[0141] The parameter PalettePredictorEntryReuseFlags[i] may be information indicating whether an entry is reused. For example, a first value (e.g., 0) of PalettePredictorEntryReuseFlags[i] may indicate that the i-th entry of the predictor palette is not an entry of the current palette, and a second value (e.g., 1) may indicate that the i-th entry of the predictor palette can be reused in the current palette. For use of PalettePredictorEntryReuseFlags[i], the initial value may be set to 0.
[0142] The parameter palette_predictor_run can indicate the number of zeros that exist before non-zero entries in the array PalettePredictorEntryReuseFlags.
[0143] The parameter num_signalled_palette_entries may indicate the number of entries in the current palette that are explicitly signaled for the first color component startComp of the current palette table. If num_signalled_palette_entries is not obtained from the bitstream, the value of num_signalled_palette_entries may be determined to be 0.
[0144] The parameter CurrentPaletteSize[startComp] indicates the size of the current palette for the first color component startComp in the current palette table. This can be calculated using the following formula: The value of CurrentPaletteSize[startComp] can range from 0 to palette_max_size.
[0145] [Number 2] CurrentPaletteSize[startComp]=NumPredictedPaletteEntries+num_signalled_palette_entries
[0146] The parameter new_palette_entries[cIdx][i] may indicate the value of the i-th signaled palette entry for color component cIdx.
[0147] The parameter PredictorPaletteEntries[cIdx][i] may indicate the i-th element in the predictor palette for color component cIdx.
[0148] The parameter CurrentPaletteEntries[cIdx][i] can indicate the i-th element in the current palette for the color component cIdx. PredictorPaletteEntries and CurrentPaletteEntries can be generated using the formulas shown in FIG.
[0149] The parameter palette_escape_val_present_flag may indicate the presence or absence of an escape coded sample. For example, a first value (e.g., 0) of palette_escape_val_present_flag may indicate that no escape coded sample exists for the current coding unit, and a second value (e.g., 1) of palette_escape_val_present_flag may indicate that the current coding unit includes at least one escape coded sample. If palette_escape_val_present_flag is not obtained from the bitstream, the value of palette_escape_val_present_flag may be determined to be 1.
[0150] The parameter MaxPaletteIndex may indicate the maximum available value of the palette index for the current coding unit. The value of MaxPaletteIndex may be determined as CurrentPaletteSize[startComp]+palette_escape_val_present_flag.
[0151] The parameter num_palette_indices_minus1 may indicate the number of palette indices that are explicitly or implicitly signaled for the current block. For example, a value of num_palette_indices_minus1 plus 1 may indicate the number of palette indices that are explicitly or implicitly signaled for the current block. If num_palette_indices_minus1 is not included in the bitstream, the value of num_palette_indices_minus1 may be determined to be 0.
[0152] The parameter palette_idx_idc may be an index indicator for the palette table CurrentPaletteEntries. The value of palette_idx_idc may have a value from 0 to MaxPaletteIndex for the first index of the block, and a value from 0 to MaxPaletteIndex-1 for the remaining indexes of the block. If the value of palette_idx_idc is not obtained from the bitstream, the value of palette_idx_idc may be determined to be 0.
[0153] The parameter PaletteIndexIdc[i] may be an array that stores the value of the i-th palette_idx_idc that is explicitly or implicitly signaled. The values of all elements of PaletteIndexIdc[i] may be initialized to 0.
[0154] The parameter copy_above_indices_for_final_run_flag may indicate information indicating whether to copy previous indices for the final run, where a first value (e.g., 0) may indicate that the palette index at the end of the current coding unit is explicitly or implicitly signaled via the bitstream, and a second value (e.g., 1) may indicate that the palette index at the end of the current coding unit is explicitly or implicitly signaled via the bitstream. If copy_above_indices_for_final_run_flag is not obtained from the bitstream, the value of copy_above_indices_for_final_run_flag may be determined to be 0.
[0155] The parameter palette_transpose_flag may be information indicating a scanning method used to scan indices for pixels of the current coding unit. For example, a first value (e.g., 0) of palette_transpose_flag may indicate that horizontal scanning is applied to scan indices for pixels of the current coding unit, and a second value (e.g., 1) of palette_transpose_flag may indicate that vertical scanning is applied to scan indices for pixels of the current coding unit. If palette_transpose_flag is not acquired from the bitstream, the value of palette_transpose_flag may be determined to be 0.
[0156] A first value (e.g., 0) of the parameter copy_above_palette_indices_flag may indicate that an indicator for the palette index of a sample is obtained or derived from an encoded value of the bitstream. A second value (e.g., 1) of copy_above_palette_indices_flag may indicate that the palette index is the same as that of the neighboring sample. For example, the neighboring sample may be a sample located in the same position as the current sample in the left column of the current sample when vertical scanning is currently used. Alternatively, the neighboring sample may be a sample located in the same position as the current sample in the row above the current sample when horizontal scanning is currently used.
[0157] The first value (e.g., 0) of the parameter CopyAboveIndicesFlag[xC][yC] may indicate that the palette index is obtained explicitly or implicitly from the bitstream. The second value (e.g., 1) may indicate that the palette index is generated by copying the palette index of the left column if vertical scanning is currently used, or by copying the palette index of the upper row if horizontal scanning is currently used. Here, xC and yC are coordinate indicators that indicate the relative position of the current sample from the top-left sample of the current picture. The value of PaletteIndexMap[xC][yC] may range from 0 to (MaxPaletteIndex-1).
[0158] The parameters PaletteIndexMap[xC][yC] indicate palette indices, and may indicate, for example, indexes into the array represented by CurrentPaletteEntries. The array indices xC and yC are coordinate indicators that indicate the coordinates of the current sample relative to the top-left sample of the current picture, as described above. PaletteIndexMap[xC][yC] may have values from 0 to (MaxPaletteIndex-1).
[0159] The parameter PaletteRun can indicate the number of consecutive positions having the same palette index when the value of CopyAboveIndicesFlag[xC][yC] is 0. On the other hand, when the value of CopyAboveIndicesFlag[xC][yC] is 1, PaletteRun can indicate the number of consecutive positions having the same palette index as the palette index at a position in the upper row when the current scan direction is horizontal scan, or the same palette index as the palette index at a position in the left column when the current scan direction is vertical scan.
[0160] The parameter PaletteMaxRun may indicate the maximum available value of PaletteRun. The value of PaletteMaxRun may be an integer greater than 0.
[0161] The parameter palette_run_prefix can indicate the prefix portion used in binarizing the PaletteRun.
[0162] The parameter palette_run_suffix may indicate the suffix portion used in the binarization of PaletteRun. If palette_run_suffix is not obtained from the bitstream, its value may be determined to be 0.
[0163] The value of PaletteRun can be determined as follows: For example, if the value of palette_run_prefix is less than 2, it can be calculated as follows:
[0164] [Number 3] PaletteRun=palette_run_prefix
[0165] On the other hand, if the value of palette_run_prefix is 2 or more, it can be calculated as follows:
[0166] [Number 4]
[0167] PrefixOffset=1<<(palette_run_prefix-1)
[0168] PaletteRun=PrefixOffset+palette_run_suffix
[0169] The parameter palette_escape_val may indicate the quantized escape coded sample value for the component. The parameter PaletteEscapeVal[cIdx][xC][yC] may indicate the escape value of the sample whose PaletteIndexMap[xC][yC] value is (MaxPaletteIndex-1) and whose palette_escape_val_present_flag value is 1. Here, cIdx may indicate a color component. The array indicators xC and yC may be position indicators that represent the position of the current sample as a relative distance from the top-left sample of the current picture, as described above.
[0170] Chroma prediction mode signaling when palette mode is applied
[0171] A method for signaling chroma prediction mode information when palette mode is applied will now be described. In one embodiment, chroma prediction coding such as CCLM may not be applied to a coding unit (or coding block) to which palette mode is applied. In addition, intra_chroma_pred_mode may not be signaled for a coding unit to which palette mode is applied.
[0172] If it is determined that CCLM is available for a chroma component, an on / off flag for it can be signaled. In one embodiment, the availability of CCLM can be determined using sps_palette_enabled_flag or sps_plt_enabled_flag, and the on / off flag signaling of CCLM can be signaled using sps_palette_enabled_flag or sps_plt_enabled_flag.
[0173] Meanwhile, in the example of Fig. 13, the signaling of CCLM information (e.g., cclm_mode_flag) does not take into account whether palette mode is applied to the coding unit. For example, the example of Fig. 13 illustrates an embodiment in which predetermined chroma prediction information (e.g., cclm_mode_flag, intra_chroma_pred_mode) is signaled when palette mode is not applied to the current coding unit or when the coding unit is not dual-tree chroma.
[0174] In such a case, signaling predetermined chroma prediction information for a palette-coded block using a single-tree structure may result in unnecessary syntax signaling. Furthermore, signaling predetermined chroma prediction information in this manner may result in chroma intra prediction using CCLM or DM mode even though the chroma components in a coding unit are coded in palette mode.
[0175] To solve this problem, the syntax for a coding unit may be modified as shown in Figure 21. Figure 21 shows the syntax of a coding unit indicating that predetermined chrominance intra prediction information 2120 is obtained from the bitstream when the value of pred_mode_plt_flag 2110, a parameter indicating whether palette mode is applied to the current coding unit, is a first value (e.g., 0) indicating that palette mode is not applied. Furthermore, the syntax of Figure 21 indicates that predetermined chrominance intra prediction information 2120 is not obtained from the bitstream when the value of pred_mode_plt_flag 2110 is a second value (e.g., 1) indicating that palette mode is applied.
[0176] As in the embodiment of Figure 21, certain chroma prediction information (e.g., cclm_mode_flag, intra_chroma_pred_mode) can be signaled depending on whether palette mode is applied to the current coding unit.
[0177] Hereinafter, signaling of predetermined chrominance intra prediction information according to the syntax of Figure 21 will be described with reference to Figure 22. According to an embodiment, an encoding device or a decoding device may determine whether palette mode is applied to a current coding unit (e.g., a coding block) (S2210). For example, the decoding device may determine whether palette mode is applied to the current coding unit according to the value of pred_mode_plt_flag.
[0178] Next, if the palette mode is applied to the current coding unit, the encoding device may encode the coding unit in the palette mode, and the decoding device may decode the coding unit in the palette mode. Accordingly, the encoding device or the decoding device may not signal certain chroma prediction information (S2220). For example, the encoding device may not need to encode certain chroma prediction information (e.g., cclm_mode_flag, intra_chroma_pred_mode), and the decoding device may not need to obtain certain chroma prediction information from the bitstream.
[0179] Next, if palette mode is not applied to the current coding unit, the encoding device or decoding device may signal the predetermined chroma prediction information. In one embodiment, the encoding device or decoding device determines whether CCLM mode is available for the current coding unit (S2230), and if available, may signal CCLM parameters (S2240), or if not available, may signal intra_chroma_pred_mode parameters (S2250).
[0180] In contrast, the step of the decoding device obtaining chroma prediction information will be described in more detail with reference to FIG. 23. When palette mode is not applied to the current coding unit (S2310), the decoding device may determine whether CCLM mode is available for the current coding unit (S2320). For example, the decoding device may determine that CCLM mode is not available for the current coding unit if a parameter sps_cclm_enabled_flag indicating availability of CCLM mode signaled in the sequence parameter set has a first value (e.g., 0) indicating that CCLM mode is not available. Alternatively, when sps_cclm_enabled_flag has a second value (e.g., 1) indicating that CCLM mode is available, the decoding device may determine that CCLM mode is available for the current coding unit if a slice type parameter sh_slice_type transmitted via a slice header indicates that the current slice type is not an I-slice or the size of the luma component of the current block is smaller than 64. Alternatively, the decoding device can determine that CCLM mode is available for the current coding unit if sps_cclm_enabled_flag has a second value (e.g., 1) indicating that CCLM mode is available, and if the sps_qtbtt_dual_tree_intra_flag parameter signaled in the sequence parameter set does not indicate that each CTU (coding tree unit) included in the I slice is divided into luma component blocks of size 64x64 and the development CTU becomes a header node of the dual tree.
[0181] If the CCLM mode is available, the decoding device may determine whether the CCLM mode is applied to the current coding unit (S2330). For example, the decoding device may acquire a cclm_mode_flag parameter from the bitstream. The parameter cclm_mode_flag may indicate whether the CCLM mode is applied. A first value (e.g., 0) of cclm_mode_flag may indicate that the CCLM mode is not applied. A second value (e.g., 1) of cclm_mode_flag may indicate that one of the CCLM modes T_CCLM, L_CCLM, and LT_CCLM may be applied. If the value of cclm_mode_flag is not acquired from the bitstream, the value of cclm_mode_flag may be determined to be 0.
[0182] When the CCLM mode is applied (e.g., cclm_mode_flag==1), the decoding device may obtain a parameter cclm_mode_idx from the bitstream (S2340). The parameter cclm_mode_idx may indicate an index indicating the CCLM mode to be used to decode the chroma components of the current coding unit, among T_CCLM, L_CCLM, and LT_CCLM.
[0183] On the other hand, if the CCLM mode is unavailable or not applicable (e.g., cclm_mode_flag==0), the decoding device may obtain the parameter intra_chroma_pred_mode from the bitstream (S2350). As described above, the parameter intra_chroma_pred_mode may indicate the intra prediction mode used to decode the chroma components of the current coding unit. For example, the parameter intra_chroma_pred_mode may indicate any one of the planar mode, DC mode, vertical mode, horizontal mode, and derived mode (DM).
[0184] Encoding method
[0185] Hereinafter, a method for encoding by an encoding device according to an embodiment using the above-described method will be described with reference to Fig. 24. The encoding device according to an embodiment includes a memory and at least one processor, and the at least one processor can perform the following encoding method.
[0186] First, the encoding device may divide an image to determine a current block (S2410). For example, the encoding device may divide an image to determine a current block as described above with reference to FIGS. 4 to 6. In the division process according to one embodiment, image division information may be coded, and the coded image division information may be generated as a bitstream.
[0187] Next, the encoding device may determine the prediction mode of the current block (S2420). Next, the encoding device may encode a palette mode flag (e.g., pred_mode_plt_flag) indicating whether the prediction mode of the current block is palette mode based on whether the prediction mode of the current block is palette mode (S2430). The encoded palette mode flag may be generated as a bitstream.
[0188] Next, the encoding device may encode palette mode encoding information obtained by encoding the current block in palette mode based on the tree type of the current block and whether the prediction mode of the current block is palette mode (S2440). For example, if it is determined that the palette mode is to be applied, the encoding device may generate a bitstream by generating palette mode encoding information using the palette_coding() syntax, as described with reference to FIGS.
[0189] Meanwhile, encoding palette mode encoding information for the current block may include encoding palette mode encoding information for a luma component of the current block. For example, if the tree type of the current block is a single tree type or a dual tree luma type and a palette mode is applied to the current block, information for palette mode prediction for the luma component of the current block may be encoded, and a bitstream may be generated using the encoded information.
[0190] In this case, the palette mode coding information for the luma component of the current block may be coded based on the size of the luma component block of the current block. For example, as in the embodiment of Figure 13, the coding apparatus may generate the palette mode coding information as a bitstream according to the palette_coding() syntax defined based on the width (e.g., cbWidth) and height (e.g., cbHeight) of the luma component block of the current block.
[0191] In addition, encoding palette mode encoding information for the current block may further include encoding palette mode encoding information for chroma components of the current block. For example, if a palette mode is applied to the current block and the tree type of the current block is a dual-tree chroma type, information for palette mode prediction for chroma components of the current block may be encoded, and a bitstream may be generated using the encoded information.
[0192] In this case, the palette mode coding information for the chroma components of the current block may be coded based on the size of the chroma component blocks of the current block. For example, as in the embodiment of FIG. 13, the coding device may generate the palette mode coding information as a bitstream according to the palette_coding() syntax defined based on the width (e.g., cbWidth / subWidthC) and height (e.g., cbHeight / subHeightC) of the chroma component blocks of the current block. Here, subWidthC and subHeightC may be the ratios of the height and width of the chroma component blocks with respect to the luma component block. In one embodiment, subWidthC and subHeightC may be determined based on chroma_format_idc and separate_cour_plane_flag as shown in FIG. 10.
[0193] Next, if the prediction mode of the current block is not the palette mode, the encoding device may encode chroma component prediction information of the current block (S2450). The chroma component prediction information may be information for CCLM (Cross-component linear model) prediction (e.g., cclm_mode_flag, cclm_mode_idx) or chroma component intra-prediction information (e.g., intra_chroma_pred_mode).
[0194] Meanwhile, if the palette mode is applied to the current block, the encoding apparatus may not encode the chroma component prediction information.
[0195] More specifically, the information for CCLM prediction may include a CCLM flag (e.g., cclm_mode_flag) indicating whether CCLM prediction is performed and a CCLM mode index (e.g., cclm_mode_idx) indicating the mode of CCLM prediction. The CCLM flag may be coded and generated as the bitstream if CCLM prediction is available for the current block. The CCLM mode index may be coded and generated as the bitstream if the CCLM flag indicates that CCLM prediction is performed.
[0196] On the other hand, if the CCLM flag indicates that the CCLM prediction is not performed, the chroma component intra-prediction information (eg, intra_chroma_pred_mode) may be coded and generated as the bitstream.
[0197] Decryption method
[0198] A method for performing decoding by a decoding device according to an embodiment using the above-described method will now be described with reference to Figure 25. The decoding device according to an embodiment includes a memory and at least one processor, and the at least one processor can perform the following decoding method.
[0199] First, the decoding apparatus may divide an image to determine a current block (S2510). For example, the decoding apparatus may divide an image to determine a current block as described above with reference to Figures 4 to 6. In the division process according to one embodiment, image division information obtained from a bitstream may be used.
[0200] Next, the decoding apparatus can determine whether the palette mode is applied to the current block based on a palette mode flag (eg, pred_mode_plt_flag) obtained from the bitstream (S2520).
[0201] Next, the decoding apparatus may obtain palette mode coding information for the current block from the bitstream based on the tree type of the current block and whether the palette mode is applied to the current block (S2530). For example, if it is determined that the palette mode is applied, the decoding apparatus may obtain palette mode coding information from the bitstream using the palette_coding() syntax as described with reference to FIGS.
[0202] Meanwhile, obtaining palette mode coding information for the current block may include obtaining palette mode coding information for a luma component of the current block. For example, if the tree type of the current block is a single tree type or a dual tree luma type and a palette mode is applied to the current block, information for palette mode prediction for the luma component of the current block may be obtained from a bitstream.
[0203] In this case, palette mode coding information for the luma component of the current block may be obtained based on the size of the luma component block of the current block. For example, as in the embodiment of Figure 13, the palette_coding() syntax may be performed based on the width (e.g., cbWidth) and height (e.g., cbHeight) of the luma component block of the current block.
[0204] Meanwhile, the step of obtaining palette mode encoding information for the current block may further include obtaining palette mode encoding information for chroma components of the current block. For example, if a palette mode is applied to the current block and the tree type of the current block is a dual-tree chroma type, information for palette mode prediction for chroma components of the current block may be obtained from the bitstream.
[0205] In this case, palette mode coding information for the chroma components of the current block may be obtained based on the size of the chroma component blocks of the current block. For example, as in the embodiment of FIG. 13, the palette_coding() syntax may be performed based on the width (e.g., cbWidth / subWidthC) and height (e.g., cbHeight / subHeightC) of the chroma component blocks of the current block. subWidthC and subHeightC may be the ratios of the height and width of the chroma component blocks to the luma component block. In one embodiment, subWidthC and subHeightC may be determined based on chroma_format_idc and separate_cour_plane_flag, as shown in FIG. 10.
[0206] Next, if palette mode is not applied to the current block, the decoding apparatus may acquire chroma component prediction information of the current block from the bitstream (S2540). For example, if palette mode is not applied, the decoding apparatus may acquire CCLM prediction information (e.g., cclm_mode_flag, cclm_mode_idx) or chroma component intra-prediction information (e.g., intra_chroma_pred_mode) from the bitstream, as described with reference to FIG. 23. On the other hand, if palette mode is applied to the current block, the decoding apparatus may not acquire the chroma component prediction information from the bitstream.
[0207] More specifically, the information for CCLM prediction may include a CCLM flag (e.g., cclm_mode_flag) indicating whether CCLM prediction is performed and a CCLM mode index (e.g., cclm_mode_idx) indicating the mode of CCLM prediction. The CCLM flag may be obtained from the bitstream if CCLM prediction is available for the current block. The CCLM mode index may be obtained from the bitstream if the CCLM flag indicates that CCLM prediction is performed.
[0208] On the other hand, if the CCLM flag indicates that the CCLM prediction is not performed, the chroma component intra-prediction information (eg, intra_chroma_pred_mode) can be obtained from the bitstream.
[0209] Application example
[0210] Although the exemplary method of the present disclosure is expressed as a series of operations for clarity of explanation, this is not intended to limit the order in which the steps are performed, and the steps may be performed simultaneously or in a different order if necessary. To achieve the method according to the present disclosure, the steps illustrated may include other steps, or some steps may be omitted and the remaining steps may be included, or some steps may be omitted and additional other steps may be included.
[0211] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform the operation (step) to check the execution conditions and circumstances of the operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation to check whether the predetermined condition is satisfied.
[0212] The various embodiments of the present disclosure are not intended to enumerate all possible combinations, but are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.
[0213] Additionally, various embodiments of the present disclosure may be implemented using hardware, firmware, software, or a combination thereof, etc. In the case of a hardware implementation, the implementation may be using one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.
[0214] In addition, an image decoding apparatus and an image encoding apparatus to which an embodiment of the present disclosure is applied may be included in a multimedia broadcast transmitting / receiving apparatus, a mobile communication terminal, a home cinema video apparatus, a digital cinema video apparatus, a surveillance camera, a video conversation apparatus, a real-time communication apparatus such as video communication, a mobile streaming apparatus, a storage medium, a camcorder, a video on demand (VoD) service providing apparatus, an over-the-top (OTT) video apparatus, an internet streaming service providing apparatus, a three-dimensional (3D) video apparatus, an image telephone video apparatus, a medical video apparatus, etc., and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video apparatus may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0215] FIG. 26 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.
[0216] As shown in FIG. 26, a content streaming system to which an embodiment of the present disclosure is applied can broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0217] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or video camera directly generates a bitstream, the encoding server can be omitted.
[0218] The bitstream can be generated by an image encoding method and / or image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0219] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as an intermediary for informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which may control commands and responses between devices in the content streaming system.
[0220] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time to provide a smooth streaming service.
[0221] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation system, a slate PC, a tablet PC, an ultrabook, a wearable device such as a smartwatch, smart glass, a head mounted display (HMD), a digital TV, a desktop computer, and digital signage.
[0222] Each server in the content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0223] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands can be stored and executed on a device or computer. [Industrial Applicability]
[0224] The embodiments of the present disclosure can be used to encode / decode images.
Claims
1. An image decoding method performed by an image decoding device, comprising: determining a current block by dividing the image; identifying whether a palette mode is applied to the current block based on a palette mode flag obtained from the bitstream; determining whether a tree type of the current block is a dual tree chroma type and whether the palette mode is applied to the current block; determining whether the palette mode is applied to the current block based on the tree type of the current block being not the dual tree chroma type or the palette mode not being applied to the current block; and obtaining chroma component prediction information of the current block from the bitstream based on the fact that the palette mode is not applied to the current block.
2. An image coding method performed by an image coding device, comprising: determining a current block by dividing the image; determining a prediction mode of the current block; encoding a palette mode flag indicating whether the prediction mode of the current block is the palette mode based on whether the prediction mode of the current block is the palette mode; determining whether a tree type of the current block is a dual tree chroma type and whether the palette mode is applied to the current block; determining whether the palette mode is applied to the current block based on the tree type of the current block being not the dual tree chroma type or the palette mode not being applied to the current block; encoding chroma component prediction information of the current block based on the prediction mode of the current block not being the palette mode.
3. A method for transmitting a bitstream generated by an image coding method, comprising: determining a current block by dividing the image; determining a prediction mode of the current block; determining whether a tree type of the current block is a dual tree chroma type and whether a palette mode is applied to the current block; determining whether the palette mode is applied to the current block based on the tree type of the current block being not the dual tree chroma type or the palette mode not being applied to the current block; generating the bitstream by encoding at least one of a palette mode flag indicating whether the prediction mode of the current block is a palette mode and chroma component prediction information of the current block; The chroma component prediction information is encoded based on the prediction mode of the current block being not the palette mode.
Citation Information
Patent Citations
Method and apparatus of palette mode coding for colour video data
WO2017206805A1
Signaling in transform SKIP mode
WO2020223612A1