Image coding / decoding method, apparatus, and method for transmitting a bitstream that determine the prediction mode of a chroma block by referring to the luma sample position.
The image encoding/decoding method addresses high-resolution image efficiency challenges by determining chroma block prediction modes based on fewer luma sample positions, enhancing encoding/decoding efficiency and reducing transmission/storage costs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-03-27
AI Technical Summary
The increasing demand for high-resolution, high-quality images leads to higher transmission and storage costs due to increased information bits, necessitating a more efficient image compression technique.
An image encoding/decoding method that determines the prediction mode of a chroma block by referring to a smaller number of luma sample positions, utilizing matrix-based intra-prediction modes and palette modes, and includes a method for transmitting and storing the generated bitstream.
Improves encoding/decoding efficiency by reducing the number of luma sample positions required for prediction, thereby optimizing transmission and storage of high-resolution images.
Smart Images

Figure 2026054566000001_ABST
Abstract
Description
Technical Field
[0007] ,
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding / decoding method for determining an intra prediction mode of a chroma block, an apparatus, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure.
Background Art
[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.
[0003] Thereby, a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution, high-quality images is required.
Summary of the Invention
Problems to be Solved by the Invention
[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus that improve encoding / decoding efficiency by determining a prediction mode of a chroma block by referring to a smaller number of luma sample positions.
[0006] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0007] Furthermore, this disclosure aims to provide a recording medium that stores a bitstream generated by the image encoding method or apparatus according to this disclosure.
[0008] Furthermore, this disclosure aims to provide a recording medium that stores a bitstream received by the image decoding device provided herein, decoded, and used for image restoration.
[0009] The technical problems that this disclosure seeks to solve are not limited to those described above, and other technical problems not mentioned above will be clearly understood by a person with ordinary skill in the art to which this disclosure pertains from the following description. [Means for solving the problem]
[0010] An image decoding method performed by an image decoding apparatus according to one aspect of the present disclosure includes the steps of: dividing an image to identify a current chroma block; identifying whether a matrix-based intra-prediction mode is applied to a first luma sample position corresponding to the current chroma block; if the matrix-based intra-prediction mode is not applied, identifying whether a predetermined prediction mode is applied to a second luma sample position corresponding to the current chroma block; and if the predetermined prediction mode is not applied, determining a candidate intra-prediction mode for the current chroma block based on an intra-prediction mode applied to a third luma sample position corresponding to the current chroma block. The predetermined prediction mode may be an IBC (Intra Block Copy) mode or a palette mode.
[0011] The first luma sample position can be determined based on at least one of the width and height of the luma block corresponding to the current chroma block. The first luma sample position can be determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block. The first luma sample position may be the same position as the third luma sample position.
[0012] The second luma sample position can be determined based on at least one of the width and height of the luma block corresponding to the current chroma block. The second luma sample position can be determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block. The second luma sample position may be the same position as the third luma sample position.
[0013] Alternatively, the first luma sample position, the second luma sample position, and the third luma sample position may be the same as each other.
[0014] The first luma sample position may be the center position of the luma block corresponding to the current chromat block. The x-component position of the first luma sample position can be determined by adding half the width of the luma block to the x-component position of the upper left sample of the luma block corresponding to the current chromat block, and the y-component position of the first luma sample position can be determined by adding half the height of the luma block to the y-component position of the upper left sample of the luma block corresponding to the current chromat block.
[0015] The first luma sample position, the second luma sample position, and the third luma sample position can be determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block, respectively.
[0016] Furthermore, an image decoding apparatus according to one aspect of the present disclosure includes a memory and at least one processor, the at least one processor, which divides an image to identify a current chroma block, identifies whether a matrix-based intra-prediction mode is applied to a first luma sample position corresponding to the current chroma block, identifies whether a predetermined prediction mode is applied to a second luma sample position corresponding to the current chroma block if the matrix-based intra-prediction mode is not applied, and determines a candidate intra-prediction mode for the current chroma block based on an intra-prediction mode applied to a third luma sample position corresponding to the current chroma block if the predetermined prediction mode is not applied.
[0017] Furthermore, an image encoding method performed by an image encoding apparatus according to one aspect of the present disclosure may include the steps of: dividing an image to identify a current chroma block; identifying whether a matrix-based intra-prediction mode is applied to a first luma sample position corresponding to the current chroma block; if the matrix-based intra-prediction mode is not applied, identifying whether a predetermined prediction mode is applied to a second luma sample position corresponding to the current chroma block; and if the predetermined prediction mode is not applied, determining a candidate intra-prediction mode for the current chroma block based on an intra-prediction mode applied to a third luma sample position corresponding to the current chroma block.
[0018] Furthermore, a transmission method according to one aspect of the present disclosure can transmit a bitstream generated by an image encoding device or image encoding method of the present disclosure.
[0019] Furthermore, a computer-readable recording medium according to one aspect of the present disclosure can store a bitstream generated by an image encoding method or image encoding apparatus of the present disclosure.
[0020] The features described above, which are a brief summary of this disclosure, are merely illustrative examples of the detailed description of this disclosure described below and do not limit the scope of this disclosure.
Advantages of the Invention
[0021] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0022] Also, according to the present disclosure, an image encoding / decoding method and apparatus that improve encoding / decoding efficiency by determining a prediction mode of a chroma block by referring to a smaller number of luma sample positions can be provided.
[0023] Also, according to the present disclosure, a method of transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure can be provided.
[0024] Also, according to the present disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure can be provided.
[0025] Also, according to the present disclosure, a recording medium storing a bitstream received by the image decoding apparatus according to the present disclosure, decoded, and used for restoring an image can be provided.
[0026] The effects obtained in the present disclosure are not limited to the above-described effects, and other effects not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description.
Brief Description of the Drawings
[0027] [Figure 1] A diagram schematically showing a video coding system to which an embodiment according to the present disclosure can be applied. [Figure 2] A diagram schematically showing an image encoding apparatus to which an embodiment according to the present disclosure can be applied. [Figure 3] A diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure can be applied. [Figure 4]This figure shows the image division structure according to one embodiment. [Figure 5] This figure shows one example of a block division type using a multi-type tree structure. [Figure 6] This figure illustrates the signaling mechanism for block partitioning information in a quadtree with nested multi-type trees as described in this disclosure. [Figure 7] This figure shows one embodiment in which a CTU is divided into multiple CUs. [Figure 8] This figure shows one example of a redundant partitioning pattern. [Figure 9] This figure shows the positional relationship between luma samples and chroma samples determined by a chroma format according to one embodiment. [Figure 10] This figure shows the positional relationship between luma samples and chroma samples determined by a chroma format according to one embodiment. [Figure 11] This figure shows the positional relationship between luma samples and chroma samples determined by a chroma format according to one embodiment. [Figure 12] This figure shows the syntax for chroma format signaling according to one embodiment. [Figure 13] This figure shows a chroma format classification table according to one example. [Figure 14] This figure illustrates a directional intra-prediction mode according to one embodiment. [Figure 15] This figure illustrates a directional intra-prediction mode according to one embodiment. [Figure 16] This is a reference diagram illustrating the MIP mode according to one embodiment. [Figure 17] This is a reference diagram illustrating the MIP mode according to one embodiment. [Figure 18] This figure shows an example of horizontal scanning and vertical scanning according to one embodiment. [Figure 19]This figure shows a method for determining the intra-prediction mode of a chromablock according to one embodiment. [Figure 20] This figure shows a reference table for determining the Chroma Intra prediction mode. [Figure 21] This figure shows a reference table for determining the Chroma Intra prediction mode. [Figure 22] This is a flowchart illustrating a method for determining the Lumaintra prediction mode information according to one embodiment. [Figure 23] This is a flowchart illustrating a method for determining the Lumaintra prediction mode information according to one embodiment. [Figure 24] This is a flowchart illustrating a method for determining the Lumaintra prediction mode information according to one embodiment. [Figure 25] This is a flowchart illustrating a method for determining the Lumaintra prediction mode information according to one embodiment. [Figure 26] This is a flowchart illustrating a method by which an encoding device and a decoding device according to one embodiment perform encoding and decoding according to one embodiment. [Figure 27] This figure illustrates a content streaming system to which the embodiments of this disclosure can be applied. [Modes for carrying out the invention]
[0028] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings, so that they can be easily implemented by a person with ordinary skill in the art to which the present disclosure pertains. However, the present disclosure can be implemented in a variety of different forms and is not limited to the embodiments described herein.
[0029] In describing embodiments of this disclosure, if it is determined that a specific description of a known configuration or function would obscure the gist of this disclosure, such detailed description will be omitted. In the drawings, parts unrelated to the description of this disclosure will be omitted, and similar parts will be denoted by the same reference numerals.
[0030] In this disclosure, when one component is described as being “connected,” “joined,” or “linked” to another component, this can include not only direct connections but also indirect connections where another component exists between them. Furthermore, when one component is described as “containing” or “having” another component, this means, unless otherwise stated to the contrary, that it may include another component rather than excluding it.
[0031] In this disclosure, terms such as "first," "second," etc., are used solely for the purpose of distinguishing one component from another, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, the first component of one embodiment may be called the second component in another embodiment, and similarly, the second component of one embodiment may be called the first component in another embodiment.
[0032] In this disclosure, components that are distinguished from each other are used to clearly describe their respective characteristics and do not necessarily mean that the components are separate. In other words, multiple components may be integrated to constitute a single hardware or software unit, or a single component may be distributed to constitute multiple hardware or software units. Therefore, such integrated or distributed embodiments are also included in the scope of this disclosure, without needing to be specifically mentioned.
[0033] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, embodiments consisting of a subset of the components described in one embodiment are also included in the scope of this disclosure. Furthermore, embodiments that include additional components in addition to the components described in various embodiments are also included in the scope of this disclosure.
[0034] This disclosure relates to the encoding and decoding of images, and the terms used in this disclosure may have their ordinary meanings in the art to which this disclosure pertains, unless otherwise defined herein.
[0035] In this disclosure, "picture" generally means a unit representing any one image within a specific time period, and "slice / tile" is an encoding unit that constitutes part of a picture, and a single picture can consist of one or more slices / tiles. Furthermore, a slice / tile may contain one or more CTUs (coding tree units).
[0036] In this disclosure, “pixel” or “pel” may mean the smallest unit that constitutes a picture (or image). The term “sample” may also be used as a counterpart to pixel. A sample may generally represent a pixel or a pixel value, or it may represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.
[0037] In this disclosure, “unit” can refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with that region. A unit may be used interchangeably with terms such as “sample array,” “block,” or “area,” as it may be used. Generally, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0038] In this disclosure, “current block” can mean any one of the following: “current coding block,” “current coding unit,” “block to encode,” “block to decode,” or “block to process.” If prediction is performed, “current block” can mean “current prediction block” or “block to predict.” If transformation (inverse transformation) / quantization (inverse quantization) is performed, “current block” can mean “current transformation block” or “block to transform.” If filtering is performed, “current block” can mean “block to filter.”
[0039] Furthermore, in this disclosure, “current block” may mean “chroma block of the current block” unless there is an explicit mention of chroma block. “Chroma block of the current block” may be expressed explicitly as “chroma block” or “current chroma block,” including an explicit mention of chroma block.
[0040] In this disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."
[0041] In this disclosure, “or” may be interpreted as “and / or.” For example, “A or B” may mean 1) “A” only, 2) “B” only, or 3) “A and B.” Alternatively, in this disclosure, “or” may mean “additionally or alternatively.”
[0042] Overview of the video coding system
[0043] Figure 1 shows the video coding system according to this disclosure.
[0044] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 via a digital storage medium or network in file or streaming format.
[0045] An encoding device 10 according to one embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to one embodiment may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be called a video / image encoding unit, and the decoding unit 22 may be called a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may also include a display unit, which may be configured as a separate device or external component.
[0046] The video source generation unit 11 can acquire video / images through processes such as video / image capture, synthesis, or generation. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. The video / image generation device may include, for example, a computer, tablet, and smartphone, and may generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by a process in which the relevant data is generated.
[0047] The encoding unit 12 can encode the input video / image. The encoding unit 12 can perform a series of steps such as prediction, transformation, and quantization for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in bitstream format.
[0048] The transmission unit 13 can transmit encoded video / image information or data, output in bitstream format, to the receiving unit 21 of the decoding device 20 via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. The transmission unit 13 may include elements for generating media files via a predetermined file format and elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.
[0049] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction, corresponding to the operation of the encoding unit 12.
[0050] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0051] Overview of Image Encoding Devices
[0052] Figure 2 is a schematic diagram showing an image encoding device to which the embodiments of this disclosure can be applied.
[0053] As shown in Figure 2, the image coding device 100 may include an image splitting unit 110, a subtraction unit 115, a transformation unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transformation unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter-prediction unit 180, an intra-prediction unit 185, and an entropy coding unit 190. The inter-prediction unit 180 and the intra-prediction unit 185 can together be called the "prediction unit". The transformation unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transformation unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.
[0054] All or at least some of the multiple components constituting the image encoding device 100 can be implemented by a single hardware component (e.g., an encoder or processor) depending on the embodiment. Furthermore, the memory 170 may include a DPB (decoded picture buffer) and can be implemented by a digital storage medium.
[0055] The image splitting unit 110 can split an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. For example, the processing units may be called coding units (CUs). Coding units can be obtained by recursively splitting a coding tree unit (CTU) or the largest coding unit (LCU) using a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, a single coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure and / or a ternary-tree structure. For the splitting of coding units, a quad-tree structure may be applied first, followed by a binary-tree structure and / or a ternary-tree structure. Based on the final coding unit that cannot be further split, the coding procedure according to this disclosure can be performed. The largest coding unit can be used as the final coding unit, or a lower-depth coding unit obtained by dividing the largest coding unit can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or restoration, as described later. As another example, the processing units of the coding procedure may be prediction units (PU) or transformation units (TU). The prediction unit and the transformation unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.
[0056] The prediction unit (inter-prediction unit 180 or intra-prediction unit 185) can make predictions for the block to be processed (current block) and generate a predicted block that includes prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block or on a CU basis. The prediction unit can generate various information regarding the prediction of the current block and transmit it to the entropy coding unit 190. The prediction information can be encoded by the entropy coding unit 190 and output in bitstream format.
[0057] The intra-prediction unit 185 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) or at a distance from the current block, according to the intra-prediction mode and / or intra-prediction technique. The intra-prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of detail of the prediction direction. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.
[0058] The interprediction unit 180 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, collocated CU (colCU), etc. The reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 180 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 180 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, the residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of the surrounding block is used as the motion vector predictor, and the motion vector of the current block can be signaled by encoding the motion vector difference and an indicator for the motion vector predictor. The motion vector difference can represent the difference between the motion vector of the current block and the motion vector predictor.
[0059] The prediction unit can generate a prediction signal based on various prediction methods and / or techniques described later. For example, the prediction unit can apply intra-prediction or inter-prediction to predict the current block, and can also apply intra-prediction and inter-prediction simultaneously. A prediction method that applies intra-prediction and inter-prediction simultaneously to predict the current block can be called CIIP (combined inter and intra prediction). The prediction unit can also perform intra-block copy (IBC) to predict the current block. Intra-block copy can be used for content image / video coding such as in games, for example, in SCC (screen content coding). IBC is a method of predicting the current block using a reference block that has already been restored in the current picture at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed similarly to inter-prediction in that it derives the reference block within the current picture. In other words, IBC can use at least one of the interpretation techniques described in this disclosure.
[0060] The predicted signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. The subtraction unit 115 can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal output from the prediction unit (predicted block, predicted sample array) from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit 120.
[0061] The transformation unit 120 can generate transformation coefficients by applying transformation techniques to the residual signal. For example, the transformation techniques may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph, where the relationship information between pixels is represented by a graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels. The transformation process can be applied to pixel blocks of the same size and square shape, or to non-square, variable-sized blocks.
[0062] The quantization unit 130 can quantize the conversion coefficients and transmit them to the entropy coding unit 190. The entropy coding unit 190 can encode the quantized signal (information about the quantized conversion coefficients) and output it in bitstream format. The information about the quantized conversion coefficients can be called residual information. The quantization unit 130 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector format based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector format of the quantized conversion coefficients.
[0063] The entropy coding unit 190 can perform various coding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy coding unit 190 can also encode information necessary for video / image restoration (e.g., the values of syntax elements) together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream format in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). The video / image information may also further include general constraint information. The signaling information, transmitted information and / or syntax elements referred to in this disclosure may be encoded via the encoding procedure described above and included in the bitstream.
[0064] The bitstream can be transmitted over a network or stored on a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. A transmission unit (not shown) for transmitting the signal output from the entropy encoding unit 190 and / or a storage unit (not shown) for storing it may be provided as internal / external elements of the image encoding device 100, or the transmission unit may be provided as a component of the entropy encoding unit 190.
[0065] The quantized conversion coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be reconstructed.
[0066] The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 180 or the intra-prediction unit 185. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 155 may be called the reconstruction unit or the reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.
[0067] The filtering unit 160 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 160 can generate various filtering-related information, as will be described later in the explanation of each filtering method, and transmit it to the entropy coding unit 190. The filtering-related information can be encoded by the entropy coding unit 190 and output in bitstream format.
[0068] The corrected restored picture transmitted to memory 170 can be used as a reference picture in the interpretation unit 180. When interpretation is applied via this, the image encoding device 100 can avoid prediction mismatches between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.
[0069] The DPB in memory 170 can store the modified restored picture for use as a reference picture in the inter-prediction unit 180. Memory 170 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 180 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 170 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 185.
[0070] Overview of the image decoding device
[0071] Figure 3 is a schematic diagram showing an image decoding apparatus to which the embodiments of this disclosure can be applied.
[0072] As shown in Figure 3, the image decoding device 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an additive unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265. The inter-prediction unit 260 and the intra-prediction unit 265 can together be called the "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in the residual processing unit.
[0073] All or at least some of the multiple components constituting the image decoding device 200 can be implemented by a single hardware component (e.g., a decoder or processor) according to the embodiment. Furthermore, the memory 170 may include a DPB and can be implemented by a digital storage medium.
[0074] An image decoding device 200, upon receiving a bitstream containing video / image information, can restore the image by executing a process corresponding to the process performed in the image encoding device 100 in Figure 1. For example, the image decoding device 200 can perform decoding using the processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The restored image signal decoded and output via the image decoding device 200 can then be reproduced via a playback device (not shown).
[0075] The image decoding device 200 can receive the signal output from the image encoding device 2 in bitstream format. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The image decoding device may further use the parameter set information and / or the general constraint information to decode the image. The signaling information, received information, and / or syntax elements referred to in this disclosure can be obtained from the bitstream by decoding via the decoding procedure. For example, the entropy decoding unit 210 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image reconstruction and the quantized values of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding blocks and the blocks to be decoded, or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction unit (inter-prediction unit 260 and intra-prediction unit 265), and the residual values that have undergone entropy decoding in the entropy decoding unit 210, i.e., quantized conversion coefficients and related parameter information, can be input to the inverse quantization unit 220. In addition, of the information decoded by the entropy decoding unit 210, information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives signals output from the image coding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.
[0076] On the other hand, the image decoding device according to this disclosure may be called a video / image / picture decoding device. The image decoding device may also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265.
[0077] The inverse quantization unit 220 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 220 can rearrange the quantized transformation coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed by the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) to obtain the transformation coefficients.
[0078] The inverse conversion unit 230 can inversely convert the conversion coefficients to obtain residual signals (residual blocks, residual sample arrays).
[0079] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 210, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode (prediction technique).
[0080] As described in the explanation of the prediction unit of the image coding device 100, the prediction unit can generate prediction signals based on various prediction methods (techniques) described later.
[0081] The intra-prediction unit 265 can predict the current block by referring to the samples in the current picture. The description of the intra-prediction unit 185 can also be applied to the intra-prediction unit 265.
[0082] The interprediction unit 260 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on a reference picture. In this case, to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in block, sub-block, or sample units based on the correlation of motion information between surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In interprediction, surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the interprediction unit 260 can construct a motion information candidate list based on surrounding blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes (techniques), and the prediction information may include information indicating the mode (technique) of interprediction for the current block.
[0083] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 260 and / or intra-prediction unit 265). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 can also be applied to the adder 235. The adder 235 can be called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.
[0084] The filtering unit 240 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 250, specifically in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.
[0085] The restored picture stored (modified) in the DPB of memory 250 can be used as a reference picture in the inter-prediction unit 260. Memory 250 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 250 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 265.
[0086] In this specification, the embodiments described for the filtering unit 160, inter-prediction unit 180, and intra-prediction unit 185 of the image coding device 100 can be applied similarly or in a corresponding manner to the filtering unit 240, inter-prediction unit 260, and intra-prediction unit 265 of the image decoding device 200, respectively.
[0087] Overview of image segmentation
[0088] The video / image coding method according to this disclosure can be performed based on the following image segmentation structure. Specifically, procedures such as prediction, residual processing (inverse transformation, inverse quantization, etc.), syntax element coding, and filtering, described later, can be performed based on CTU, CU (and / or TU, PU) derived from the image segmentation structure. The image can be segmented into blocks, and the block segmentation procedure can be performed in the image segmentation unit 110 of the encoding device described above. Segmentation-related information can be encoded in the entropy encoding unit 190 and transmitted to the decoding device in bitstream format. The entropy decoding unit 210 of the decoding device can derive the block segmentation structure of the current picture based on the segmentation-related information obtained from the bitstream, and perform a series of procedures for image decoding (e.g., prediction, residual processing, block / picture reconstruction, in-loop filtering, etc.) based on this.
[0089] A picture can be divided into a sequence of coding tree units (CTUs). Figure 4 shows an example of a picture being divided into CTUs. A CTU can correspond to a coding tree block (CTB). Alternatively, a CTU can contain two coding tree blocks: one for a luma sample and one for a corresponding chroma sample. For example, for a picture containing three sample arrays, the CTU can contain an N×N block for the luma sample and two corresponding blocks for the chroma sample.
[0090] Overview of CTU division
[0091] As mentioned above, coding units can be obtained by recursively partitioning a coding tree unit (CTU) or maximum coding unit (LCU) using QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structures. For example, a CTU can first be partitioned into a quadtree structure. Then, the leaf nodes of the quadtree structure can be further partitioned into a multi-type tree structure.
[0092] A quadtree partition means dividing the current CU (or CTU) into four equal parts. Through a quadtree partition, the current CU can be divided into four CUs of the same width and height. If the current CU is not further divided into a quadtree structure, it corresponds to a leaf node in the quadtree structure. A CU that corresponds to a leaf node in a quadtree structure is not further divided and can be used as the final coding unit as described above. Alternatively, a CU that corresponds to a leaf node in a quadtree structure can be further divided by a multi-type tree structure.
[0093] Figure 5 shows the types of block partitioning using a multi-type tree structure. Partitioning using a multi-type tree structure can include two partitions using a binary tree structure and two partitions using a ternary tree structure.
[0094] The two types of partitioning using a binary tree structure include vertical binary splitting (SPLIT_BT_VER) and horizontal binary splitting (SPLIT_BT_HOR). Vertical binary splitting (SPLIT_BT_VER) means splitting the current CU vertically into two equal parts. As shown in Figure 4, vertical binary splitting can generate two CUs that have the same height as the current CU and half the width of the current CU. Horizontal binary splitting (SPLIT_BT_HOR) means splitting the current CU horizontally into two equal parts. As shown in Figure 5, horizontal binary splitting can generate two CUs that have half the height of the current CU and the same width as the current CU.
[0095] Two types of partitioning using a ternary structure are vertical ternary splitting (SPLIT_TT_VER) and horizontal ternary splitting (SPLIT_TT_HOR). Vertical ternary splitting (SPLIT_TT_VER) divides the current CU vertically in a 1:2:1 ratio. As shown in Figure 5, vertical ternary splitting can produce two CUs with the same height as the current CU and a width of 1 / 4 of the current CU's width, and one CU with the same height as the current CU and a width of half the current CU's width. Horizontal ternary splitting (SPLIT_TT_HOR) divides the current CU horizontally in a 1:2:1 ratio. As shown in Figure 4, horizontal ternary splitting can produce two CUs with the same height as the current CU and a width of 1 / 4 of the current CU's width, and one CU with the same height as the current CU and a width of half the current CU's width.
[0096] Figure 6 illustrates the signaling mechanism for block partitioning information in a quadtree with nested multi-type tree structure according to this disclosure.
[0097] Here, the CTU is treated as the root node of the quadtree, and the CTU is the first node to be split into a quadtree structure. Information (e.g., qt_split_flag) indicating whether or not to split the quadtree can be signaled to the current CU (CTU or quadtree node (QT_node)). For example, if qt_split_flag is the first value (e.g., "1"), the current CU can be split into a quadtree. If qt_split_flag is the second value (e.g., "0"), the current CU will not be split into a quadtree and will become a leaf node (QT_leaf_node) of the quadtree. Each leaf node of the quadtree can subsequently be further split into a multitype tree structure. In other words, a leaf node of a quadtree can become a node (MTT_node) of a multitype tree. In a multi-type tree structure, a first flag (e.g., mtt_split_cu_flag) may be signaled to indicate whether the current node will be further split. If the node is to be further split (e.g., the first flag is 1), a second flag (e.g., mtt_split_cu_verticla_flag) may be signaled to indicate the splitting direction. For example, if the second flag is 1, the splitting direction is vertical, and if the second flag is 0, the splitting direction is horizontal. Subsequently, a third flag (e.g., mtt_split_cu_binary_flag) may be signaled to indicate whether the splitting type is binary or ternary. For example, if the third flag is 1, the splitting type is binary, and if the third flag is 0, the splitting type is ternary. Nodes in a multitype tree obtained by binary partitioning or ternary partitioning can be further partitioned into a multitype tree structure. However, nodes in a multitype tree cannot be partitioned into a quadtree structure.If the first flag is 0, the corresponding node in the multitype tree is not further subdivided and becomes a leaf node (MTT_leaf_node) of the multitype tree. A CU corresponding to a leaf node in the multitype tree can be used as the final coding unit as described above.
[0098] Based on the aforementioned mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree splitting mode (MttSplitMode) of the CU can be derived as shown in Table 1. In the following description, the multi-tree splitting mode may be abbreviated as multi-tree splitting type or splitting type.
[0099] [Table 1]
[0100] Figure 7 shows an example where a CTU is divided into multiple CUs by applying a multitype tree after a quadtree. In Figure 7, the bold block edge 710 represents the quadtree division, and the remaining edge 720 represents the multitype tree division. A CU can correspond to a coding lock (CB). In one embodiment, a CU may include two coding blocks: a coding block for a luma sample and a coding block for a chroma sample corresponding to the luma sample. The chroma component (sample) CB or TB size can be derived based on the luma component (sample) CB or TB size according to the component ratio of the picture / image color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). If the color format is 4:4:4, the chroma component CB / TB size can be set to be the same as the luma component CB / TB size. If the color format is 4:2:2, the width of the chroma component CB / TB can be set to half the width of the luma component CB / TB, and the height of the chroma component CB / TB can be set to the height of the luma component CB / TB. If the color format is 4:2:0, the width of the chroma component CB / TB can be set to half the width of the luma component CB / TB, and the height of the chroma component CB / TB can be set to half the height of the luma component CB / TB.
[0101] In one embodiment, when the size of the CTU is 128 based on the luma sample unit, the size of the CU can range from 128×128, which is the same size as the CTU, to 4×4. In one embodiment, when the color format is 4:2:0 (or chroma format), the chroma CB size can range from 64×64 to 2×2.
[0102] On the other hand, in one embodiment, the CU size and TU size can be the same. Alternatively, multiple TUs can exist within the CU region. The TU size generally refers to the Luma component (sample) TB (Transform Block) size.
[0103] The TU size can be derived based on a preset value, the maximum allowable TB size (maxTbSize). For example, if the CU size is larger than the maxTbSize, multiple TUs (TBs) with the maxTbSize can be derived from the CU, and conversion / inverse conversion can be performed in units of the TUs (TBs). For example, the maximum allowable lumen TB size may be 64×64, and the maximum allowable chromen TB size may be 32×32. If the width or height of a CB divided by the tree structure is larger than the maximum conversion width or height, the CB can be automatically (or implicitly) divided until the horizontal and vertical TB size limits are satisfied.
[0104] Furthermore, for example, when intra-prediction is applied, the intra-prediction mode / type is derived on a CU (or CB) basis, and the peripheral reference sample derivation and prediction sample generation procedures can be performed on a TU (or TB) basis. In this case, one or more TUs (or TBs) can exist within a single CU (or CB) region, and in this case, the multiple TUs (or TBs) can share the same intra-prediction mode / type.
[0105] On the other hand, for a quadtree coding tree scheme with multitype trees, the following parameters can be signaled from the encoder to the decoder as SPS syntax elements. For example, at least one of the following can be signaled: CTUsize, which indicates the size of the root node of the quadtree; MinQTSize, which indicates the minimum allowed size of the leaf nodes of the quadtree; MaxBTSize, which indicates the maximum allowed size of the root node of the binary tree; MaxTTSize, which indicates the maximum allowed size of the root node of the ternary tree; MaxMttDepth, which indicates the maximum allowed hierarchy depth of the multitype trees that are split from the leaf nodes of the quadtree; MinBtSize, which indicates the minimum allowed leaf node size of the binary tree; and MinTtSize, which indicates the minimum allowed leaf node size of the ternary tree.
[0106] In one embodiment using the 4:2:0 chroma format, the CTU size can be set to a 128x128 chroma block and two corresponding 64x64 chroma blocks. In this case, MinQTSize can be set to 16x16, MaxBtSize to 128x128, MaxTtSzie to 64x64, MinBtSize and MinTtSize to 4x4, and MaxMttDepth to 4. Quadritree partitioning can be applied to the CTU to generate leaf nodes of the quadritree. Leaf nodes of the quadritree can be called leaf QT nodes. Leaf nodes of the quadritree can have a size of 16x16 (e.g., the MinQTSize) to 128x128 (e.g., the CTU size). If a leaf QT node is 128x128, it may not be further divided into a binary / ternary tree. This is because even if partitioned in this case, it would exceed MaxBtsize and MaxTtszie (e.g., 64x64). Otherwise, a leaf QT node can be further partitioned into a multitype tree. Thus, a leaf QT node is the root node for a multitype tree, and a leaf QT node can have a multitype tree depth (mttDepth) value of 0. If the multitype tree depth reaches MaxMttdepth (e.g., 4), further additional partitioning may not be considered. If the width of a multitype tree node is the same as MinBtSize and equal to or less than 2xMinTtSize, further additional horizontal partitioning may not be considered. If the height of a multitype tree node is the same as MinBtSize and equal to or less than 2xMinTtSize, further additional vertical partitioning may not be considered. When partitioning is not considered in this way, the encoding device can omit signaling of partitioning information. In such cases, the decoding device can induce the partitioning information to a predetermined value.
[0107] On the other hand, a single CTU can include a coding block for a luma sample (hereinafter referred to as a "luma block") and two coding blocks for corresponding chroma samples (hereinafter referred to as "chroma blocks"). The coding tree scheme described above can be applied similarly to the luma blocks and chroma blocks of a CU, or it can be applied separately. Specifically, luma blocks and chroma blocks within a single CTU can be divided into the same block tree structure, in which case the tree structure can be represented as a single tree (SINGLE_TREE). Alternatively, luma blocks and chroma blocks within a single CTU can be divided into separate block tree structures, in which case the tree structure can be represented as a dual tree (DUAL_TREE). In other words, when a CTU is divided into a dual tree, the block tree structure for luma blocks and the block tree structure for chroma blocks can exist separately. In this case, the block tree structure for a luma block can be called a dual-tree luma (DUAL_TREE_LUMA), and the block tree structure for a chroma block can be called a dual-tree chroma (DUAL_TREE_CHROMA). For P and B slice / tile groups, luma blocks and chroma blocks within a single CTU can be restricted to having the same coding tree structure. However, for I slice / tile groups, luma blocks and chroma blocks can have separate block tree structures from each other. If separate block tree structures are applied, a luma CTB (Coding Tree Block) can be divided into CUs based on a specific coding tree structure, and a chroma CTB can be divided into chroma CUs based on a different coding tree structure. That is, a CU within an I slice / tile group to which a separate block tree structure is applied can consist of a coding block for a luma component or a coding block for two chroma components, while a CU in a P or B slice / tile group can consist of a block for three color components (a luma component and two chroma components).
[0108] In the above, a quadtree coding tree structure with a multitype tree was described, but the structures in which a CU is split are not limited to this. For example, BT structures and TT structures can be interpreted as concepts included in multiple partitioning tree (MPT) structures, and a CU can be interpreted as being split by QT structures and MPT structures. In one example of a CU being split by QT and MPT structures, the split structure can be determined by signaling a syntax element (e.g., MPT_split_type) containing information about how the leaf nodes of the QT structure are split into several blocks, and a syntax element (e.g., MPT_split_mode) containing information about whether the leaf nodes of the QT structure are split vertically or horizontally.
[0109] In another example, the CU can be divided in a way different from the QT, BT, or TT structures. That is, unlike the QT structure which divides the lower-depth CU into quarters the size of the upper-depth CU, or the BT structure which divides the lower-depth CU into half the size of the upper-depth CU, or the TT structure which divides the lower-depth CU into quarters or half the size of the upper-depth CU, the lower-depth CU can, depending on the case, be divided into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 the size of the upper-depth CU, and the way in which the CU is divided is not limited to this.
[0110] Thus, the quadtree coding block structure with the multitype tree can provide a highly flexible block partition structure. On the other hand, due to the partition types supported by the multitype tree, different partition patterns may, in some cases, lead to potentially identical coding block structures. By limiting the occurrence of such redundant partition patterns, the encoding and decoding devices can reduce the amount of data in the partition information.
[0111] For example, Figure 8 illustrates redundant partition patterns that can occur in binary and ternary tree partitions. As shown in Figure 8, a 2-step level unidirectional consecutive binary partition 810 and 820 has the same coding block structure as a binary partition on the center partition after a ternary partition. In such a case, a binary tree partition on the center blocks 830 and 840 of the ternary partition can be prohibited. Such prohibitions can be applied to the CU of all pictures. When such a particular partition is prohibited, the signaling of the corresponding syntax element can be modified to reflect this prohibition, thereby reducing the number of bits signaled for the partition. For example, if a binary tree partition on the center block of a CU is prohibited, as in the example shown in Figure 8, the mtt_split_cu_binary_flag syntax element, which indicates whether the partition is a binary or ternary partition, is not signaled, and its value can be induced to 0 by the decoder.
[0112] Overview of Chroma Format
[0113] The following describes chroma formats. Images can be encoded with encoded data that includes a luma component (e.g., Y) array and two chroma component (e.g., Cb, Cr) arrays. For example, one pixel in an encoded image can contain a luma sample and a chroma sample. A chroma format can be used to indicate the configuration format of luma and chroma samples, and a chroma format is sometimes called a color format.
[0114] In one embodiment, the image can be encoded in various chroma formats such as monochrome, 4:2:0, 4:2:2, and 4:4:4. In monochrome sampling, there may be one sample array, which may be a luma array. In 4:2:0 sampling, there may be one luma sample array and two chroma sample arrays, each of which may have half the height and half the width of the luma array. In 4:2:2 sampling, there may be one luma sample array and two chroma sample arrays, each of which may have the same height as the luma array and half the width of the luma array. In 4:4:4 sampling, there may be one luma sample array and two chroma sample arrays, each of which may have the same height and width as the luma array.
[0115] Figure 9 shows the relative positions of a luma sample and a chroma sample in one example using 4:2:0 sampling. Figure 10 shows the relative positions of a luma sample and a chroma sample in one example using 4:2:2 sampling. Figure 11 shows the relative positions of a luma sample and a chroma sample in one example using 4:4:4 sampling. As shown in Figure 9, in the case of 4:2:0 sampling, the position of the chroma sample can be located at the lower end of the corresponding luma sample. As shown in Figure 10, in the case of 4:2:2 sampling, the chroma sample can be located overlapping the position of the corresponding luma sample. As shown in Figure 11, in the case of 4:4:4 sampling, both the luma sample and the chroma sample can be located in overlapping positions.
[0116] The chroma format used in the encoding and decoding devices can be predetermined. Alternatively, the chroma format can be signaled from the encoding device to the decoding device for adaptive use in the encoding and decoding devices. In one embodiment, the chroma format can be signaled based on at least one of chroma_format_idc and separate_colour_plane_flag. At least one of chroma_format_idc and separate_colour_plane_flag can be signaled via a higher-level syntax such as DPS, VPS, SPS, or PPS. For example, chroma_format_idc and separate_colour_plane_flag can be included in the SPS syntax as shown in Figure 12.
[0117] On the other hand, Figure 13 shows an example of chroma format classification utilizing the signaling of chroma_format_idc and separate_colour_plane_flag. chroma_format_idc can be information indicating the chroma format applied to the encoded image. separate_colour_plane_flag can indicate whether the color array is processed separately in a particular chroma format. For example, the first value of chroma_format_idc (e.g., 0) can indicate monochrome sampling. The second value of chroma_format_idc (e.g., 1) can indicate 4:2:0 sampling. The third value of chroma_format_idc (e.g., 2) can indicate 4:2:2 sampling. The fourth value of chroma_format_idc (e.g., 3) can indicate 4:4:4 sampling.
[0118] In 4:4:4 sampling, the following applies based on the value of separate_colour_plane_flag: If the value of separate_colour_plane_flag is the first value (e.g., 0), each of the two chroma arrays can have the same height and width as the luma array. In this case, the value of ChromaArrayType, which indicates the type of chroma sample array, can be set to be the same as chroma_format_idc. If the value of separate_colour_plane_flag is the second value (e.g., 1), the luma, Cb, and Cr sample arrays can be processed separately, similar to monochrome-sampled pictures. In this case, ChromaArrayType can be set to 0.
[0119] Overview of Intra Prediction Mode
[0120] The intra-prediction modes will be described in more detail below. Figure 14 shows an intra-prediction direction according to one embodiment. To capture an arbitrary edge direction presented in a natural video, the intra-prediction mode can include two non-directional intra-prediction modes and 65 directional intra-prediction modes, as shown in Figure 14. The non-directional intra-prediction modes can include a Planar intra-prediction mode and a DC intra-prediction mode, and the directional intra-prediction modes can include intra-prediction modes 2 through 66.
[0121] On the other hand, the intra-prediction mode may further include a CCLM (cross-component linear model) mode for chroma samples, in addition to the intra-prediction mode described above. The CCLM mode can be divided into L_CCLM, T_CCLM, and LT_CCLM depending on whether the left sample, the upper sample, or both are considered for LM parameter derivation, and can be applied only to chroma components. For example, the intra-prediction mode can be indexed according to the intra-prediction mode value, as shown in the table below.
[0122] [Table 2]
[0123] Figure 15 shows the intra-prediction direction according to another embodiment. Here, the dashed direction indicates the wide-angle mode, which is applied only to blocks that are not squares. As shown in Figure 15, in order to capture any edge direction presented in a natural video, the intra-prediction mode according to one embodiment can include 93 directional intra-prediction modes along with two non-directional intra-prediction modes. The non-directional intra-prediction modes can include Planar prediction modes and DC prediction modes. The directional intra-prediction modes can include intra-prediction modes numbered 2 to 80 and -1 to -14, as indicated by the arrows in Figure 15. The Planar prediction mode can be denoted as INTRA_PLANAR, and the DC prediction mode can be denoted as INTRA_DC. The directional intra-prediction modes can be denoted as INTRA_ANGULAR-14 to INTRA_ANGULAR-1 and INTRA_ANGULAR2 to INTRA_ANGULAR80. On the other hand, the intra prediction type (or additional intra prediction mode, etc.) may include at least one of the above-mentioned LIP, PDPC, MRL, ISP, and MIP. The intra prediction type can be indicated based on intra prediction type information, and the intra prediction type information can be implemented in various forms. For example, the intra prediction type information may include intra prediction type index information that indicates one of the intra prediction types. As another example, the intra prediction type information may include reference sample line information (e.g., intra_luma_ref_idx) indicating whether the MRL is applied to the current block and, if so, which reference sample line is used; ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block; ISP type information (e.g., intra_subpartitions_split_flag) indicating the subpartition splitting type when the ISP is applied; flag information indicating whether PDPC is applied or flag information indicating whether LIP is applied; and MIP flag information indicating whether MIP is applied.
[0124] The intra-prediction mode information and / or the intra-prediction type information can be encoded / decoded via the coding methods described herein. For example, the intra-prediction mode information and / or the intra-prediction type information can be encoded / decoded via entropy coding (e.g., CABAC, CAVLC) based on truncated (rice) binary code.
[0125] When intraprediction is performed on the current block, predictions can be made for the luma component block (luma block) and the chroma component block (chroma block) of the current block. In this case, the intraprediction mode for the chroma block can be set separately from the intraprediction mode for the luma block.
[0126] For example, the intra-prediction mode for a chroma block can be specified based on intra-chroma prediction mode information, which can be signaled in the form of an intra_chroma_pred_mode syntax element. As an example, the intra-chroma prediction mode information can refer to one of the following: Planar mode, DC mode, vertical mode, horizontal mode, DM (Derived Mode), or CCLM mode. Here, Planar mode can refer to intra-prediction mode 0, DC mode to intra-prediction mode 1, vertical mode to intra-prediction mode 26, and horizontal mode to intra-prediction mode 10. DM is sometimes called direct mode. CCLM is sometimes called LM.
[0127] On the other hand, DM and CCLM are dependent intra-prediction modes that predict the chroma block using information from the luma block. DM can indicate a mode where the same intra-prediction mode used for the luma component is applied as the intra-prediction mode for the chroma component. CCLM can indicate an intra-prediction mode where, in the process of generating a predicted block for the chroma block, the reconstructed sample of the luma block is subsampled, and then the sample derived by applying CCLM parameters α and β to the subsampled sample is used as the predicted sample for the chroma block.
[0128] Overview of MIP Mode
[0129] Matrix-based intra-prediction (MIP) is sometimes called ALWIP (affine linear weighted intra-prediction) mode, LWIP (linear weighted intra-prediction) mode, or MWIP (matrix weighted intra-prediction) mode. Intra-prediction modes that are not matrix-based can be defined as non-matrix-based prediction modes. For example, non-matrix-based prediction modes can refer to non-directional intra-prediction and directional intra-prediction. In the following, the terms intra-prediction mode and general intra-prediction mode will be used interchangeably to refer to non-matrix-based prediction modes. Hereafter, matrix-based prediction will be referred to as MIP mode.
[0130] When the MIP mode is applied to the current block, predicted samples for the current block can be derived by i) using the surrounding reference samples that have undergone an averaging step, ii) performing a matrix-vector-multiplication step, and iii) further performing horizontal / vertical interpolation steps as necessary.
[0131] The averaging step can be performed by averaging the values of the surrounding samples. The averaging procedure can be performed as shown in Figure 16(a) by taking the average of each boundary surface to generate a total of 4 samples (2 on the top and 2 on the left) if the current block width and depth are 4 in pixels, or as shown in Figure 16(b) by taking the average of each boundary surface to generate a total of 8 samples (4 on the top and 4 on the left).
[0132] The matrix-vector product step can be performed by multiplying the averaged samples by the matrix vector and then adding the offset vector, thereby generating a predicted signal for the subsampled pixel set of the original block. The sizes of the matrix and offset vector can be determined according to the current block width and depth.
[0133] The horizontal / vertical interpolation step is a step that generates a prediction signal of the original block size from a subsampled prediction signal. As shown in Figure 17, a prediction signal of the original block size can be generated by performing vertical and horizontal interpolation using the subsampled prediction signal and peripheral pixel values. Figure 17 shows one example in which MIP prediction is performed on an 8x8 block. In the case of an 8x8 block, a total of 8 averaged samples can be generated, as shown in Figure 16(b). By multiplying the 8 averaged samples by a matrix vector and adding an offset vector, 16 sample values can be generated at even coordinate positions, as shown in Figure 17(a). Subsequently, as shown in Figure 17(b), vertical interpolation can be performed using the average value of the upper samples of the current block. Then, as shown in Figure 17(c), horizontal interpolation can be performed using the left-hand samples of the current block.
[0134] The intra-prediction mode used for the MIP mode can be configured differently from the intra-prediction modes used for the LIP, PDPC, MRL, ISP intra-prediction, and normal intra-prediction described above. The intra-prediction mode for the MIP mode may be called the MIP intra-prediction mode, MIP prediction mode, or MIP mode. For example, depending on the intra-prediction mode for the MIP, the matrix and offset used in the matrix-vector product may be set to be different. Here, the matrix may be called the (MIP) weight matrix, and the offset may be called the (MIP) offset vector or (MIP) bias vector.
[0135] The intra prediction type information described above may include an MIP flag (e.g., intra_mip_flag) indicating whether or not MIP mode is applied to the current block. intra_mip_flag[x0][y0] can indicate whether the current block was predicted in MIP mode. For example, the first value of intra_mip_flag[x0][y0] (e.g., 0) can indicate that the current block was not predicted in MIP mode. The second value of intra_mip_flag[x0][y0] (e.g., 1) can indicate that the current block was predicted in MIP mode.
[0136] If intra_mip_flag[x0][y0] has a second value (e.g., 1), further information about the MIP mode can be obtained from the bitstream. For example, the syntax elements intra_mip_mpm_flag[x0][y0], intra_mip_mpm_idx[x0][y0], and intra_mip_mpm_remainder[x0][y0], which indicate the MIP mode of the current block, can be further obtained from the bitstream. If a MIP prediction mode is applied to the current block, an MPM list for the MIP can be constructed, and the intra_mip_mpm_flag can indicate whether the MIP mode for the current block exists in the MPM list for the MIP (or among the MPM candidates). The intra_mip_mpm_idx can indicate the index of a candidate in the MPM list to be used as the MIP prediction mode for the current block, if a MIP prediction mode for the current block exists in the MPM list for the MIP (i.e., the value of intra_mip_mpm_flag is 1). The intra_mip_mpm_remainder can indicate the MIP prediction mode for the current block if a MIP prediction mode for the current block does not exist in the MPM list for the MIP (i.e., the value of intra_mip_mpm_flag is 0), and can indicate any one of the overall MIP prediction modes, or indicate any one of the remaining modes from the overall MIP prediction modes, excluding the candidate modes in the MPM list for the MIP, as the MIP prediction mode for the current block.
[0137] On the other hand, if intra_mip_flag[x0][y0] has a first value (for example, 0), no information for MIP is obtained from the bitstream, and intra prediction information other than MIP can be obtained from the bitstream. In one embodiment, intra_luma_mpm_flag[x0][y0], which indicates whether or not an MPM list for general intra prediction is generated, can be obtained from the bitstream.
[0138] If an intra-planar mode is currently applied to a block, an MPM list can be constructed for it, and intra_luma_mpm_flag can indicate whether or not an intra-planar mode for the current block exists in the MPM list (or among the MPM candidates). For example, a first value of intra_luma_mpm_flag (e.g., 0) can indicate that an intra-planar mode for the current block does not exist in the MPM list. A second value of intra_luma_mpm_flag (e.g., 1) can indicate that an intra-planar mode for the current block exists in the MPM list. If the value of intra_luma_mpm_flag is 1, the intra_luma_not_planar_flag can be obtained from the bitstream.
[0139] The `intra_luma_not_planar_flag` can indicate whether the current block's intra-prediction mode is planar mode. For example, a first value of `intra_luma_not_planar_flag` (e.g., 0) can indicate that the current block's intra-prediction mode is planar mode. A second value of `intra_luma_not_planar_flag` (e.g., 1) can indicate that the current block's intra-prediction mode is not planar mode.
[0140] intra_luma_mpm_idx can be parsed and coded if intra_luma_not_planar_flag is "true" (i.e., value 1). In one embodiment, Planar mode can always be a candidate in the MPM list, however, Planar mode can be excluded from the MPM list by signaling intra_luma_not_planar_flag first as described above, in which case an MPM list unified with the various intra prediction types described above (general intra prediction, MRL, ISP, LIP, etc.) can be constructed. In this case, the number of candidates in the MPM list can be reduced to 5. intra_luma_mpm_idx can indicate the candidate to be used as the intra prediction mode for the current block from among the candidates included in the MPM list from which Planar mode has been excluded.
[0141] On the other hand, if the value of intra_luma_mpm_flag is 0, the intra_luma_mpm_remainder can be parsed / coded. The intra_luma_mpm_remainder can either specify one of the overall intra prediction modes as the intra prediction mode for the current block, or specify one of the remaining modes in the MPM list, excluding the candidate modes, as the intra prediction mode for the current block.
[0142] Overview of Palette Mode
[0143] The following describes the palette mode (PLT mode). An encoding device according to one embodiment can encode an image using palette mode, and a decoding device can decode an image using palette mode in a corresponding manner. Palette mode can also be called palette encoding mode, intra-palette mode, intra-palette encoding mode, etc. Palette mode can be considered as a type of intra-coding mode, and can also be considered as one of the intra-prediction methods. However, similar to the skip mode described above, a separate residual value for the block in question may not be signaled.
[0144] In one embodiment, palette mode can be used to improve encoding efficiency when encoding screen content, which is a computer-generated image containing a significant amount of text and graphics. Generally, local areas of images generated in screen content are separated by sharp edges and represented by a small number of colors. To take advantage of these characteristics, palette mode can represent a sample for a block with an index that points to a color entry in a palette table.
[0145] To apply palette mode, information can be signaled to the palette table. In one embodiment, the palette table may include index values corresponding to each color. To signal the index values, palette index prediction information can be signaled. The palette index prediction information may include index values for at least a portion of the palette index map. The palette index map can map pixels of video data to color indices in the palette table.
[0146] Palette index prediction information may include run value information. For at least a portion of the palette index map, the run value information may be information that associates run values with index values. One run value may be associated with an escape color index. A palette index map can be generated from the palette index prediction information. For example, at least a portion of the palette index map can be generated by deciding whether or not to adjust the index values of the palette index prediction information based on the last index value.
[0147] The current block in the current picture can be encoded or restored according to a palette index map. When palette mode is applied, pixel values in the current encoding unit can be represented by a small set of representative color values. Such a set can be named a palette. For pixels with values close to palette colors, a palette index can be signaled. For pixels with values that do not belong to (are outside) the palette, the pixel is represented by an escape symbol, and the quantized pixel value can be directly signaled. In this specification, pixels or pixel values can be illustrated by examples.
[0148] To decode a block encoded in palette mode, the decoder can decode the palette color and index. The palette color can be described as a palette table and encoded using a palette table coding tool. An escape flag can be signaled for each encoding unit. The escape flag can indicate whether or not an escape symbol exists in the current encoding unit. If an escape symbol exists, the palette table is incremented by one unit (e.g., an index unit), and the last index can be specified as escape mode. The palette indices of all pixels for a single encoding unit can form a palette index map and can be encoded using a palette index map coding tool.
[0149] For example, a palette predictor can be maintained to encode the palette table. The palette predictor can be initialized at each slice start point. For example, the palette predictor can be reset to 0. For each entry in the palette predictor, a reuse flag can be signaled to indicate whether or not it is currently part of the palette. The reuse flag can be signaled using run-length coding with a value of 0.
[0150] Subsequently, numbers for new palette entries can be signaled using zero-order exponential Golomb codes. Finally, component values for new palette entries can be signaled. After encoding the currently encoded units, palette predictors can be updated using the current palette, and entries from previous palette predictors that are not reused in the current palette can be appended to the end of the new palette predictor (until the maximum allowed size is reached), which can be called palette stuffing.
[0151] For example, to encode a palette index map, the index can be encoded using horizontal or vertical scanning. The scan order can be signaled via the bitstream using a parameter palette_transpose_flag that indicates the scan direction. For example, if a horizontal scan is applied to scan the index for a sample in the current encoding unit, palette_transpose_flag can have a first value (e.g., 0), and if a vertical scan is applied, palette_transpose_flag can have a second value (e.g., 1). Figure 18 shows an example of horizontal and vertical scanning according to one embodiment.
[0152] In one embodiment, the palette index can be encoded using "INDEX" mode and "COPY_ABOVE" mode. When a horizontal scan is used, the mode of the palette index is signaled to the top row; when a vertical scan is used, the mode of the palette index is signaled to the leftmost column; and the two modes can be signaled using a single flag, except when the preceding mode is "COPY_ABOVE".
[0153] In "INDEX" mode, the palette index can be explicitly signaled. For both "INDEX" mode and "COPY_ABOVE" mode, the run value indicating the number of encoded pixels can be signaled using the same mode.
[0154] The encoding order for an index map can be set as follows: First, the number of index values for an encoding unit can be signaled. This can be done after the actual index values for the overall encoding unit are signaled using truncated binary coding. Both the number of indices and the index values can be encoded in bypass mode. This allows for grouping of bypass bins associated with the index. Then, palette mode (INDEX or COPY_ABOVE) and run values can be signaled in an interleaved manner.
[0155] Finally, component escape values corresponding to escape samples for the overall coding unit can be grouped together and coded in bypass mode. An additional syntex element, last_run_type_flag, can be signaled after the index value has been signaled. By using last_run_type_flag together with the index number, signaling of the run value corresponding to the last run in the block can be omitted.
[0156] In one embodiment, a dual-tree type can be used for an I-slice, which performs independent coding unit partitioning for the luminal and chroma components. Palette mode can be applied to the luminal and chroma components individually or together. If dual-tree is not applied, palette mode can be applied to all Y, Cb, and Cr components.
[0157] Overview of IBC (Intra Block Copy) mode
[0158] IBC prediction can be performed in the prediction unit of an image encoding / decoding device. IBC prediction can be simply referred to as "IBC". The IBC can be used for content image / video coding such as games, for example, as in SCC (screen content coding). The IBC basically performs prediction within the current picture, but can be performed similarly to interpretation in that it derives a reference block within the current picture. In other words, the IBC can use at least one of the interpretation techniques described in this disclosure. For example, the IBC can use at least one of the motion information (motion vector) derivation methods described above. At least one of the interpretation techniques can also be used with some modification taking into consideration the IBC prediction. The IBC can reference the current picture. Therefore, it can also be called CPR (current picture referencing).
[0159] For IBC, the image encoding device can perform block matching BM to derive the optimal block vector (or motion vector) for the current block (e.g., CU). The derived block vector (or motion vector) can be signaled to the image decoding device via a bitstream using a method similar to the motion information (motion vector) signaling in interpretation described above. The image decoding device can derive a reference block for the current block in the current picture via the signaled block vector (motion vector), thereby deriving a prediction signal (predicted block or predicted sample) for the current block. Here, the block vector (or motion vector) can indicate the displacement from the current block to the reference block located in an already restored region within the current picture. Therefore, the block vector (or motion vector) can also be called a displacement vector. Hereinafter, the motion vector in IBC can correspond to the block vector or the displacement vector. The motion vector of the current block can include a motion vector for the luma component (luma motion vector) or a motion vector for the chroma component (chroma motion vector). For example, the chroma motion vector for an IBC-coded CU can also be integer-sample-level (i.e., integer precision). The chroma motion vector can also be clipped to integer-sample levels. As mentioned above, IBC can employ at least one of the interpretation techniques, and for example, the chroma motion vector can be encoded / decoded using the merge mode or MVP mode described above.
[0160] When merge mode is applied to a Luma IBC block, the merge candidate list for the Luma IBC block can be constructed in the same way as the merge candidate list in inter-mode. However, in the case of Luma IBC blocks, time-peripheral blocks do not necessarily have to be used as merge candidates.
[0161] When MVP mode is applied to a Luma IBC block, the MVP candidate list for the Luma IBC block can be configured in the same way as the MVP candidate list in inter-mode. However, in the case of Luma IBC blocks, time candidate blocks do not need to be used as MVP candidates.
[0162] The IBC derives the reference block from the already restored area within the current picture. To reduce memory consumption and the complexity of the image decoding device, only predefined areas within the already restored area of the current picture can be referenced. These predefined areas may include the current CTU containing the current block. By restricting the referenced restored area to the predefined area in this way, the IBC mode can be implemented in hardware using local on-chip memory.
[0163] An image encoding device performing IBC can search the previously defined region to determine the reference block with the smallest RD cost, and derive a motion vector (block vector) based on the positions of the reference block and the current block.
[0164] Whether or not to apply IBC to the current block can be signaled at the CU level as IBC execution information. Information regarding the signaling method for the current block's motion vector (IBC MVP mode or IBC skip / merge mode) can also be signaled. IBC execution information can be used to determine the prediction mode of the current block. Therefore, IBC execution information can be included in the information regarding the prediction mode of the current block.
[0165] In IBC skip / merge mode, the merge candidate index is signaled and can be used to indicate which block vector is currently used to predict the luma block from among the block vectors included in the merge candidate list. In this case, the merge candidate list can include peripheral blocks encoded by IBC. The merge candidate list can be configured to include spatial merge candidates but not temporal merge candidates. Furthermore, the merge candidate list can include HMVP (History-based motion vector predictor) candidates and / or pairwise candidates.
[0166] In IBC MVP mode, block vector difference values can be encoded in the same way as the motion vector difference values in intermode described above. The block vector prediction method, similar to intermode's MVP mode, can be constructed and used by configuring an MVP candidate list containing two candidates as predictors. One of the two candidates can be derived from the left peripheral block, and the other from the upper peripheral block. In this case, a candidate can only be derived from the peripheral block if the left or upper peripheral block is encoded with IBC. If the left or upper peripheral block is not available, for example, if it is not encoded with IBC, a default block vector can be included in the MVP candidate list as a predictor. Also, similar to intermode's MVP mode, information (e.g., a flag) to indicate one of the two block vector predictors is signaled and used as candidate selection information. The MVP candidate list can include an HMVP candidate and / or a zero motion vector as the default block vector.
[0167] The aforementioned HMVP candidates are sometimes called history-based MVP candidates. MVP candidates, merge candidates, or block vector candidates previously used in the encoding / decoding of the current block can be stored in the HMVP list as HMVP candidates. Thereafter, if the current block's merge candidate list or MVP candidate list does not contain the maximum number of candidates, candidates stored in the HMVP list can be added to the current block's merge candidate list or MVP candidate list as HMVP candidates.
[0168] The pairwise candidates mentioned above refer to candidates that are derived by selecting two candidates in a predetermined order from among the candidates already included in the current block merge candidate list, and then averaging the two selected candidates.
[0169] Intra prediction for chromablock
[0170] When intraprediction is performed on the current block, predictions can be made for the luma component block (luma block) and the chroma component block (chroma block) of the current block. In this case, the intraprediction mode for the chroma block can be set separately from the intraprediction mode for the luma block.
[0171] The intra-prediction mode for a chroma block can be determined based on the intra-prediction mode of the corresponding luma block. Figure 19 shows a method for determining the intra-prediction mode of a chroma block according to one embodiment.
[0172] Referring to Figure 19, a method by which an encoding / decoding device according to one embodiment determines the intra-prediction mode of a chroma block will be described. The decoding device will be described below, and this description can be directly applied to the encoding device.
[0173] According to the following explanation, the intra-prediction mode IntraPredModeC[xCb][yCb] for the chroma block can be induced, and the following parameters can be used in this process.
[0174] - Luminous sample coordinates (xCb, yCb) indicating the relative position of the upper-left sample of the current chroma block to the position of the upper-left luminous sample of the current picture.
[0175] - cbWidth indicates the width of the current coding block in Luma samples.
[0176] - cbHeight indicates the current coding block height in luma samples.
[0177] Furthermore, the following explanation can be used when the current slice containing the chroma block is an I-slice and a luma-chroma dual-tree partitioning structure is applied. However, the following explanation is not limited to the above examples. For example, the following explanation can be applied regardless of whether the current slice is an I-slice or not, and furthermore, the following explanation can be applied in common even when a dual-tree partitioning structure is not used.
[0178] First, the decoding device can determine lumaIntraPredMode information (e.g., lumaIntraPredMode) based on the prediction mode of the luma block currently corresponding to the chroma block (S1910). A detailed explanation of this step will be given later.
[0179] Next, the decoding device can determine the chroma intra-prediction mode information based on the luma intra-prediction mode information and additional information (S1920). In one embodiment, the decoding device can determine the chroma intra-prediction mode based on the cclm_mode_flag parameter indicating whether or not a CCLM prediction mode obtained from the bitstream is applied, the cclm_mode_idx parameter indicating the CCLM mode type applied when a CCLM prediction mode is applied, the intra_chroma_pred_mode parameter indicating the intra-prediction mode type applied to the chroma sample, and the luma intra-prediction mode information (e.g., lumaIntraPredMode) and the table in Figure 20.
[0180] As an example, intra_chroma_pred_mode can refer to one of the following modes: Planar mode, DC mode, vertical mode, horizontal mode, DM (Derived Mode), or CCLM (Cross-component linear model) mode. Here, Planar mode can refer to intra-prediction mode 0, DC mode to intra-prediction mode 1, vertical mode to intra-prediction mode 26, and horizontal mode to intra-prediction mode 10. DM is sometimes called direct mode. CCLM is sometimes called LM (linear model). CCLM mode can include one of L_CCLM, T_CCLM, or LT_CCLM.
[0181] On the other hand, DM and CCLM are dependent intra-prediction modes that predict chroma blocks using information from luma blocks. DM can indicate a mode in which the same intra-prediction mode as the intra-prediction mode of the luma block currently corresponding to the chroma block is applied as the intra-prediction mode for the chroma block. In DM mode, the intra-prediction mode of the current chroma block can be determined by the intra-prediction mode indicated by the luma intra-prediction mode information.
[0182] Furthermore, the CCLM can exhibit an intra-prediction mode in which, during the process of generating a prediction block for a chroma block, the reconstructed sample of the chroma block is subsampled, and then the sample derived by applying CCLM parameters α and β to the subsampled sample is used as the prediction sample for the chroma block.
[0183] Next, the decoding device can map the chroma intra prediction mode based on the chroma format (S1930). In one embodiment, the decoding device can map the chroma intra prediction mode X determined according to the table in Figure 20 to a new chroma intra prediction mode Y based on the table in Figure 21, but only when the chroma format is 4:2:2 (for example, when the value of chroma_format_idc is 2). For example, if the value of the chroma intra prediction mode determined according to the table in Figure 20 is 16, this can be mapped to chroma intra prediction mode 14 according to the mapping table in Figure 21.
[0184] Determination of Lumaintra prediction mode information
[0185] The following describes in more detail step S1910, which determines the lumaIntraPredMode information. Figure 22 is a flowchart illustrating how the decoding device determines the lumaIntraPredMode information.
[0186] First, the decoding device can determine whether or not MIP mode is applied to the luma block currently corresponding to the chroma block (S2210). If MIP mode is applied to the luma block currently corresponding to the chroma block, the decoding device can set the value of the luma intra prediction mode information to INTRA_PLANAR mode (S2220).
[0187] On the other hand, if MIP mode is not applied to the luma block currently corresponding to the chroma block, the decoding device can determine whether IBC mode or PLT mode is applied to the luma block currently corresponding to the chroma block (S2230).
[0188] Currently, when IBC mode or PLT mode is applied to the luma block corresponding to the chroma block, the decoding device can set the value of the luma intra prediction mode information to INTRA_DC mode (S2240).
[0189] On the other hand, if IBC mode or PLT mode is not applied to the luma block currently corresponding to the chroma block, the decoding device can set the value of the luma intra prediction mode information to the intra prediction mode of the luma block currently corresponding to the chroma block (S2250).
[0190] In the process of determining the Luma Intra prediction mode information, as shown in Figure 22, various methods can be applied to identify the Luma Block corresponding to the current Chroma Block. As mentioned above, the upper-left sample position of the current Chroma Sample can be expressed by the relative coordinates of the Luma Sample separated from the position of the upper-left Luma Sample of the current Picture. Below, we will explain how to identify the Luma Block corresponding to the Chroma Block under these circumstances.
[0191] Figure 23 shows a first embodiment in which the luma sample position is referenced to identify the prediction mode of the luma block corresponding to the chroma block. Steps S2310 to S2350 below correspond to steps S2210 to S2250 in Figure 22 described above, so only the differences will be explained.
[0192] In the first embodiment shown in Figure 23, a first luma sample position corresponding to the current upper-left sample position of the chroma block and a second luma sample position determined based on the current upper-left sample position of the chroma block and the current width and height of the luma block are referenced. Here, the first luma sample position may be (xCb, yCb). The second luma sample position may be (xCb + cbWidth / 2, yCb + cbHeight / 2).
[0193] More specifically, in step S2310, the first luma sample position can be identified to determine whether MIP mode has been applied to the luma block corresponding to the current chroma block. More specifically, it can be determined whether the value of the parameter intra_mip_flag[xCb][yCb], which indicates whether MIP mode has been applied, is 1 or not, identified at the first luma sample position. The first value of intra_mip_flag[xCb][yCb] (for example, 1) can indicate that MIP mode is applied at the [xCb][yCb] sample position.
[0194] Furthermore, in step S2330, the first luma sample position can be identified to determine whether IBC or palette mode has been applied to the luma block currently corresponding to the chroma block. More specifically, it can be determined whether the value of the prediction mode parameter CuPredMode[0][xCb][yCb] of the luma block identified by the first luma sample position is a value indicating IBC mode (e.g., MODE_IBC) or a value indicating palette mode (e.g., MODE_PLT).
[0195] Furthermore, in step S2350, the second luma sample position can be identified in order to set the luma intra prediction mode information to the intra prediction mode information of the luma block corresponding to the current chroma block. More specifically, the value of the parameter IntraPredModeY[xCb+cbWidth / 2][yCb+cbHeight / 2], which indicates the luma intra prediction mode identified at the second luma sample position, can be identified.
[0196] However, in this first embodiment, both the first and second luma sample positions are considered to determine the intra-prediction mode of the chroma block. This reduces the encoding and decoding complexity compared to considering only one luma sample position compared to considering two sample positions.
[0197] The second and third embodiments, which consider only one luma sample position, are described below. Figures 24 and 25 show the second and third embodiments, which refer to the luma sample position to identify the prediction mode of the luma block currently corresponding to the chroma block. The following steps S2410 to S2450 and S2510 to S2550 correspond to the steps S2210 to S2250 in Figure 22 described above, so only the differences will be explained.
[0198] In the second embodiment shown in Figure 24, the second luma sample position, determined based on the current upper-left sample position of the chroma block and the current width and height of the luma block, can be referenced. Here, the second luma sample position may be (xCb + cbWidth / 2, yCb + cbHeight / 2). Thus, in the second embodiment, the predicted value of the luma sample corresponding to the center position of the luma block corresponding to the current chroma block (xCb + cbWidth / 2, yCb + cbHeight / 2) can be identified. In one embodiment, if the luma block has an even number of columns and rows, the center position can indicate the position of the lower-right sample among the four central samples of the luma block.
[0199] More specifically, in step S2410, the second luma sample position can be identified to determine whether MIP mode has been applied to the luma block corresponding to the current chroma block. More precisely, it can be determined whether the value of the parameter intra_mip_flag[xCb+cbWidth / 2][yCb+cbHeight / 2], which indicates whether MIP mode has been applied as identified by the second luma sample position, is 1 or not.
[0200] Furthermore, in step S2430, the second luma sample position can be identified to determine whether IBC or palette mode has been applied to the luma block currently corresponding to the chroma block. More specifically, it can be determined whether the value of the prediction mode parameter CuPredMode[0][xCb+cbWidth / 2][yCb+cbHeight / 2] of the luma block identified by the second luma sample position is a value indicating IBC mode (e.g., MODE_IBC) or a value indicating palette mode (e.g., MODE_PLT).
[0201] Furthermore, in step S2450, the second luma sample position can be identified in order to set the luma intra-prediction mode information to the intra-mode information of the luma block corresponding to the current chroma block. More specifically, the value of the parameter IntraPredModeY[xCb+cbWidth / 2][yCb+cbHeight / 2], which indicates the luma intra-prediction mode identified at the second luma sample position, can be identified.
[0202] In the third embodiment shown in Figure 25, the first luma sample position corresponding to the upper left sample position of the current chroma block can be referenced. Here, the first luma sample position may be (xCb, yCb). Thus, in the third embodiment, the predicted value of the luma sample corresponding to the upper left sample position (xCb, yCb) of the luma block corresponding to the current chroma block can be identified. In one embodiment, if the luma block has an even number of columns and rows,
[0203] More specifically, in step S2510, the first luma sample position can be identified to determine whether MIP mode has been applied to the luma block corresponding to the current chroma block. More specifically, it can be determined whether the value of the parameter intra_mip_flag[xCb][yCb], which indicates whether MIP mode has been applied as identified at the first luma sample position, is 1 or not.
[0204] Furthermore, in step S2530, the first luma sample position can be identified to determine whether IBC or palette mode has been applied to the luma block currently corresponding to the chroma block. More specifically, it can be determined whether the value of the prediction mode parameter CuPredMode[0][xCb][yCb] of the luma block identified by the first luma sample position is a value indicating IBC mode (e.g., MODE_IBC) or a value indicating palette mode (e.g., MODE_PLT).
[0205] Furthermore, in step S2550, the first luma sample position can be identified in order to set the luma intra prediction mode information to the intra prediction mode information of the luma block corresponding to the current chroma block. More specifically, the values of the parameters IntraPredModeY[xCb][yCb], which indicate the luma intra prediction mode identified at the first luma sample position, can be identified.
[0206] Encoding and Decoding Methods
[0207] The following describes, with reference to Figure 26, how an encoding device according to one embodiment performs encoding and how a decoding device performs decoding using the method described above. The encoding device according to one embodiment includes a memory and at least one processor, and the following methods can be performed by the at least one processor. Similarly, the decoding device according to one embodiment also includes a memory and at least one processor, and the following methods can be performed by the at least one processor. For the sake of explanation, the operation of the decoding device will be described below, but the following description can also be applied to the encoding device.
[0208] First, the decoding device can divide the image to identify the current chroma block (S2610). Next, the decoding device can determine whether or not a matrix-based intra-prediction mode is applied to the first luma sample position corresponding to the current chroma block (S2620). Here, the first luma sample position can be determined based on at least one of the width and height of the luma block corresponding to the current chroma block. For example, the first luma sample position can be determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block.
[0209] Next, the decoding device can determine whether a predetermined prediction mode is applied to the second luma sample position corresponding to the current chroma block, unless the matrix-based intra-prediction mode is applied (S2630). Here, the second luma sample position can be determined based on at least one of the width and height of the luma block corresponding to the current chroma block. For example, the second luma sample position can be determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block.
[0210] Next, if the predetermined prediction mode is not applied, the decoding device can determine a candidate intra-prediction mode for the current chroma block based on the intra-prediction mode applied to the third chroma sample position corresponding to the current chroma block (S2640). Here, the predetermined prediction mode may be IBC (Intra Block Copy) mode or palette mode.
[0211] On the other hand, the first luma sample position may be the same as the third luma sample position. Or, the second luma sample position may be the same as the third luma sample position. Or, the first luma sample position, the second luma sample position, and the third luma sample position may be the same as each other.
[0212] Alternatively, the first luma sample position may be the center position of the luma block corresponding to the current chromat block. For example, the x-component position of the first luma sample position can be determined by adding half the width of the luma block to the x-component position of the upper left sample of the luma block corresponding to the current chromat block, and the y-component position of the first luma sample position can be determined by adding half the height of the luma block to the y-component position of the upper left sample of the luma block corresponding to the current chromat block.
[0213] Alternatively, the first luma sample position, the second luma sample position, and the third luma sample position can be determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block, respectively.
[0214] Application Examples
[0215] The exemplary methods in this disclosure are presented as a series of actions for clarity of explanation, but this is not intended to restrict the order in which the steps are performed, and each step may be performed simultaneously or in a different order, if necessary. To implement the methods according to this disclosure, the exemplary steps may be further varied, including the remaining steps with some exceptions, or including additional steps with some exceptions.
[0216] In this disclosure, an image encoding device or image decoding device that performs a predetermined operation (step) may perform an operation (step) to confirm the conditions or status of the execution of said operation (step). For example, if it is stated that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or image decoding device may perform an operation to confirm whether or not the predetermined condition is satisfied, and then perform the predetermined operation.
[0217] The various embodiments of this disclosure are not intended to list all possible combinations, but rather to illustrate representative aspects of this disclosure. The matters described in the various embodiments may be applied independently or in combination of two or more.
[0218] Furthermore, various embodiments of this disclosure can be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, it can be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.
[0219] Furthermore, the image decoding and image encoding devices to which the embodiments of this disclosure are applied can be included in multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video conferencing equipment, real-time communication equipment such as video communications, mobile streaming equipment, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video equipment, internet streaming service providers, 3D video equipment, image-phone video equipment, and medical video equipment, and can be used to process video signals or data signals. For example, over-the-top (OTT) video equipment can include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0220] Figure 27 illustrates a content streaming system to which the embodiments of this disclosure can be applied.
[0221] As shown in Figure 27, the content streaming system to which the embodiments of this disclosure are applied may broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.
[0222] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmitting this bitstream to the streaming server. In other cases, if a multimedia input device such as a smartphone, camera, or video camera directly generates the bitstream, the encoding server can be omitted.
[0223] The bitstream can be generated by an image encoding method and / or image encoding apparatus to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0224] The streaming server transmits multimedia data to the user's device based on the user's request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server can play a role in controlling the commands and responses between the devices within the content streaming system.
[0225] The streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0226] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices such as smartwatches, smart glasses, HMDs (head-mounted displays), digital TVs, desktop computers, and digital signage.
[0227] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received from each server can be processed in a distributed manner.
[0228] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that enable the operation of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands etc. are stored and can be executed on a device or computer. [Industrial applicability]
[0229] The embodiments described herein can be used for encoding / decoding images.
Claims
1. An image decoding method performed by an image decoding device, The current step is to identify the chroma block by dividing the image, The steps include determining the tree type of the current chroma block, Based on the fact that the tree type of the current chroma block is a dual tree, the step of determining whether a matrix-based intra-prediction mode is applied to the first chroma sample position corresponding to the current chroma block, Based on the fact that the matrix-based intra-prediction mode is not applied, the step of determining whether a predetermined prediction mode is applied to the second luma sample position corresponding to the current chroma block, Based on the fact that the predetermined prediction mode is not applied, the steps include determining a candidate intra-prediction mode for the current chroma block based on the intra-prediction mode applied to the third luma sample position corresponding to the current chroma block, The steps include determining the current intra-prediction mode of the chroma block based on the intra-prediction mode candidates and additional information, The step of mapping the current chroma block's intra-prediction mode based on the chroma format is included, The additional information includes information indicating whether a CCLM (cross component linear model) prediction mode is applied, information indicating the CCLM mode type, and information indicating the intra-prediction mode type currently applied to the chroma block. The first luma sample position, the second luma sample position, and the third luma sample position are the same. The predetermined prediction mode is a palette mode in an image decoding method.
2. The image decoding method according to claim 1, wherein the first luma sample position is determined based on at least one of the width and height of the luma block corresponding to the current chroma block.
3. The image decoding method according to claim 1, wherein the first luma sample position is determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block.
4. The image decoding method according to claim 1, wherein the second luma sample position is determined based on at least one of the width and height of the luma block corresponding to the current chroma block.
5. The image decoding method according to claim 1, wherein the second luma sample position is determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block.
6. The image decoding method according to claim 1, wherein the first luma sample position is the central position of the luma block corresponding to the current chroma block.
7. The x-component position of the first luma sample position is determined by adding half the width of the luma block to the x-component position of the upper left sample of the luma block corresponding to the current chromatic block. The image decoding method according to claim 1, wherein the y-component position of the first luma sample position is determined by adding half the height of the luma block to the y-component position of the upper left sample of the luma block corresponding to the current chroma block.
8. The image decoding method according to claim 1, wherein the first luma sample position, the second luma sample position, and the third luma sample position are determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block, respectively.
9. An image encoding method performed by an image encoding device, The current step is to identify the chroma block by dividing the image, Based on the fact that the tree type of the current chroma block is a dual tree, the step of determining whether a matrix-based intra-prediction mode is applied to the first chroma sample position corresponding to the current chroma block, Based on the fact that the matrix-based intra-prediction mode is not applied, the step of determining whether a predetermined prediction mode is applied to the second luma sample position corresponding to the current chroma block, The step of determining a candidate intra-prediction mode for the current chroma block based on the fact that the predetermined prediction mode is not applied, based on the intra-prediction mode applied to the third luma sample position corresponding to the current chroma block, is included. The current intra-prediction mode of the chromablock is determined based on the intra-prediction mode candidates and additional information. The current intra-prediction mode of the chroma block is mapped based on the chroma format. The additional information includes information indicating whether a CCLM (cross component linear model) prediction mode is applied, information indicating the CCLM mode type, and information indicating the intra-prediction mode type currently applied to the chroma block. The first luma sample position, the second luma sample position, and the third luma sample position are the same. An image coding method in which the predetermined prediction mode is a palette mode.
10. A method for transmitting a bitstream, The current step is to identify the chroma block by dividing the image, Based on the fact that the tree type of the current chroma block is a dual tree, the step of determining whether a matrix-based intra-prediction mode is applied to the first chroma sample position corresponding to the current chroma block, Based on the fact that the matrix-based intra-prediction mode is not applied, the step of determining whether a predetermined prediction mode is applied to the second luma sample position corresponding to the current chroma block, Based on the fact that the predetermined prediction mode is not applied, the steps include determining a candidate intra-prediction mode for the current chroma block based on the intra-prediction mode applied to the third luma sample position corresponding to the current chroma block, The steps include encoding additional information to generate the bitstream, The step of transmitting data including the bitstream, The current intra-prediction mode of the chromablock is determined based on the intra-prediction mode candidates and the additional information. The current intra-prediction mode of the chroma block is mapped based on the chroma format. The additional information includes information indicating whether a CCLM (cross component linear model) prediction mode is applied, information indicating the CCLM mode type, and information indicating the intra-prediction mode type currently applied to the chroma block. The first luma sample position, the second luma sample position, and the third luma sample position are the same. The predetermined prediction mode is a palette mode, in this method.