Image encoding / decoding method and apparatus for determining prediction mode of chroma block by referring to luma sample position, and method for transmitting bitstream

The image encoding/decoding method improves efficiency by determining the prediction mode of a chroma block with reference to fewer luma sample positions, effectively addressing the challenge of handling high-resolution images.

JP2025083533AActive Publication Date: 2025-05-30NOKIA TECHNOLOGIES OY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025042093
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-08-14
Filing Date
2025-03-17
Publication Date
2025-05-30
Estimated Expiration
2040-08-14

AI Technical Summary

Technical Problem

There is a need for an image encoding/decoding method and apparatus that improves encoding/decoding efficiency, particularly by determining a prediction mode of a chroma block with reference to a smaller number of luma sample positions, to effectively handle high-resolution, high-quality images.

Method used

The method involves dividing an image to identify a current chroma block and determining its intra prediction mode by checking if a matrix-based intra prediction mode is applied to a first luma sample position. If not, it checks if a predetermined prediction mode, such as IBC or palette mode, is applied to a second luma sample position. If neither is applied, it determines an intra prediction mode candidate for the chroma block based on the intra prediction mode applied to a third luma sample position.

Benefits of technology

This approach enhances encoding/decoding efficiency by reducing the number of luma sample positions considered for determining the chroma block's prediction mode, thereby improving the handling of high-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025083533000001_ABST
    Figure 2025083533000001_ABST
Patent Text Reader

Abstract

To provide an image encoding / decoding method and apparatus.SOLUTION: The invention provides an image encoding / decoding method and apparatus. An image decoding method practiced by an image decoding apparatus may comprise: identifying a current chroma block by splitting an image; identifying whether a matrix-based intra prediction mode is applied to a first luma sample position corresponding to the current chroma block; if the matrix-based intra prediction mode is not applied, then identifying whether a predetermined prediction mode is applied to a second luma sample position corresponding to the current chroma block; and, if the predetermined prediction mode is not applied, then determining an intra prediction mode candidate of the current chroma block based on an intra prediction mode applying to a third luma sample position corresponding to the current chroma block.SELECTED DRAWING: Figure 26
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding / decoding method and apparatus for determining an intra prediction mode of a chroma block, and a method for transmitting a bitstream generated by the image encoding method / apparatus of the present disclosure.

Background Art

[0002] Recently, demands for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, have been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.

[0003] Accordingly, there is a need for a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution, high-quality images.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus that improve encoding / decoding efficiency by determining a prediction mode of a chroma block with reference to a smaller number of luma sample positions.

[0006] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0007] In addition, an object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0008] In addition, an object of the present disclosure is to provide a recording medium storing a bitstream received by an image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.

[0009] The technical problems to be solved in the present disclosure are not limited to the above-described technical problems, and other technical problems not described above will be clearly understood by those having ordinary knowledge in the technical field to which the present disclosure pertains from the following description.

Means for Solving the Problems

[0010] An image decoding method performed by an image decoding apparatus according to an aspect of the present disclosure may include: dividing an image to identify a current chroma block; identifying whether a matrix-based intra prediction mode is applied to a first luma sample position corresponding to the current chroma block; if the matrix-based intra prediction mode is not applied, identifying whether a predetermined prediction mode is applied to a second luma sample position corresponding to the current chroma block; and if the predetermined prediction mode is not applied, determining an intra prediction mode candidate for the current chroma block based on an intra prediction mode applied to a third luma sample position corresponding to the current chroma block. The predetermined prediction mode may be an IBC (Intra Block Copy) mode or a palette mode.

[0011] The first luma sample position may be determined based on at least one of the width and height of the luma block corresponding to the current chroma block. The first luma sample position may be determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block. The first luma sample position may be the same position as the third luma sample position.

[0012] The second luma sample position can be determined based on at least one of the width and height of the luma block corresponding to the current chroma block. The second luma sample position can be determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block. The second luma sample position can be the same position as the third luma sample position.

[0013] Alternatively, the first luma sample position, the second luma sample position, and the third luma sample position can be the same position as each other.

[0014] The first luma sample position can be the center position of the luma block corresponding to the current chroma block. The x - component position of the first luma sample position is determined by adding half of the width of the luma block to the x - component position of the upper left sample of the luma block corresponding to the current chroma block, and the y - component position of the first luma sample position can be determined by adding half of the height of the luma block to the y - component position of the upper left sample of the luma block corresponding to the current chroma block.

[0015] The first luma sample position, the second luma sample position, and the third luma sample position can each be determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block.

[0016] Also, an image decoding apparatus according to an aspect of the present disclosure includes a memory and at least one processor. The at least one processor divides an image to identify a current chroma block, identifies whether a matrix-based intra prediction mode is applied to a first luma sample position corresponding to the current chroma block, if the matrix-based intra prediction mode is not applied, identifies whether a predetermined prediction mode is applied to a second luma sample position corresponding to the current chroma block, and if the predetermined prediction mode is not applied, can determine an intra prediction mode candidate of the current chroma block based on an intra prediction mode applied to a third luma sample position corresponding to the current chroma block.

[0017] Also, an image encoding method performed by an image encoding apparatus according to an aspect of the present disclosure can include steps of dividing an image to identify a current chroma block, identifying whether a matrix-based intra prediction mode is applied to a first luma sample position corresponding to the current chroma block, if the matrix-based intra prediction mode is not applied, identifying whether a predetermined prediction mode is applied to a second luma sample position corresponding to the current chroma block, and if the predetermined prediction mode is not applied, determining an intra prediction mode candidate of the current chroma block based on an intra prediction mode applied to a third luma sample position corresponding to the current chroma block.

[0018] Also, a transmission method according to an aspect of the present disclosure can transmit a bitstream generated by the image encoding apparatus or the image encoding method of the present disclosure.

[0019] Also, a computer-readable recording medium according to an aspect of the present disclosure can store a bitstream generated by the image encoding method or the image encoding apparatus of the present disclosure.

[0020] The features briefly summarized and described above about the present disclosure are merely exemplary aspects of the detailed description of the present disclosure to be described later, and do not limit the scope of the present disclosure. [Advantages of the Invention]

[0021] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0022] Also, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus for improving encoding / decoding efficiency by determining a prediction mode of a chroma block with reference to a smaller number of luma sample positions.

[0023] Also, according to the present disclosure, it is possible to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0024] Also, according to the present disclosure, it is possible to provide a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0025] Also, according to the present disclosure, it is possible to provide a recording medium storing a bitstream received by the image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.

[0026] The effects obtained in the present disclosure are not limited to the above-described effects, and other effects not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description. [Brief Description of the Drawings]

[0027]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Best Mode for Carrying Out the Invention

[0028] Hereinafter, with reference to the accompanying drawings, embodiments of the present disclosure will be described in detail so that those having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement them. However, the present disclosure can be realized in various different forms and is not limited to the embodiments described herein.

[0029] In describing the embodiments of the present disclosure, when it is determined that a specific description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. And in the drawings, parts not related to the description of the present disclosure are omitted, and the same reference numerals are given to the same parts.

[0030] In the present disclosure, when a component is "connected", "coupled", or "joined" to another component, this can include not only a direct connection relationship but also an indirect connection relationship in which another component exists between them. Also, when a component "includes" or "has" another component, this means that, unless otherwise stated to the contrary, it does not exclude other components but can further include other components.

[0031] In the present disclosure, terms such as "first", "second", etc. are used only for the purpose of distinguishing one component from another and do not limit the order or importance between components, etc., unless otherwise specifically mentioned. Therefore, within the scope of the present disclosure, the first component of one embodiment may be referred to as the second component in another embodiment, and similarly, the second component of one embodiment may be referred to as the first component in another embodiment.

[0032] In the present disclosure, components that are distinguished from each other are for clearly explaining their respective features and do not necessarily mean that the components are separated. That is, a plurality of components may be integrated and configured as one hardware or software unit, or one component may be distributed and configured as a plurality of hardware or software units. Therefore, without further mention, such integrated or distributed embodiments are also included in the scope of the present disclosure.

[0033] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, embodiments constituted by a subset of the components described in one embodiment are also included in the scope of the present disclosure. Also, embodiments that further include other components in addition to the components described in various embodiments are included in the scope of the present disclosure.

[0034] The present disclosure relates to image encoding and decoding, and the terms used in the present disclosure can have the ordinary meaning in the technical field to which the present disclosure belongs, unless newly defined in the present disclosure.

[0035] In the present disclosure, "picture" generally means a unit indicating any one image in a specific time period, and a slice / tile is an encoding unit constituting a part of a picture, and one picture can be composed of one or more slices / tiles. Further, a slice / tile can include one or more CTUs (coding tree units).

[0036] In the present disclosure, "pixel" or "pel" can mean the smallest unit constituting one picture (or image). Further, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and can also indicate only the pixel / pixel value of the luma component, or can also indicate only the pixel / pixel value of the chroma component.

[0037] In the present disclosure, "unit" can indicate a basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. A unit can be used interchangeably with terms such as "sample array", "block", or "area" as the case may be. In general, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0038] In the present disclosure, "current block" can mean any one of "current coding block", "current coating unit", "block to be coded", "block to be decoded", or "block to be processed". When prediction is performed, "current block" can mean "current prediction block" or "block to be predicted". When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" can mean "current transformation block" or "block to be transformed". When filtering is performed, "current block" can mean "block to be filtered".

[0039] Also, in the present disclosure, "current block" can mean "luma block of the current block" unless explicitly stated as a chroma block. "Chroma block of the current block" can be explicitly expressed including an explicit description of a chroma block such as "chroma block" or "current chroma block".

[0040] In the present disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C".

[0041] In the present disclosure, "or" can be interpreted as "and / or". For example, "A or B" can mean 1) only "A", 2) only "B", or 3) "A and B". Or, in the present disclosure, "or" can mean "additionally or alternatively".

[0042] Overview of Video Coding System

[0043] FIG. 1 is a diagram showing a video coding system according to the present disclosure.

[0044] A video coding system according to an embodiment can include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.

[0045] The encoding device 10 according to an embodiment can include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. The decoding device 20 according to an embodiment can include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 can be called a video / image encoding unit, and the decoding unit 22 can be called a video / image decoding unit. The transmission unit 13 can be included in the encoding unit 12. The reception unit 21 can be included in the decoding unit 22. The rendering unit 23 can also include a display unit, and the display unit can be configured as a separate device or an external component.

[0046] The video source generation unit 11 can obtain video / images through processes such as the capture, synthesis, or generation of video / images. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer or the like, and in this case, the video / image capture process can be replaced by a process in which related data is generated.

[0047] The encoding unit 12 can encode the input video / image. The encoding unit 12 can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0048] The transmission unit 13 can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit 21 of the decoding device 20 via a digital storage medium or network in a file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray (registered trademark), HDD, SSD, etc. The transmission unit 13 can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.

[0049] The decoding unit 22 can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operations of the encoding unit 12.

[0050] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0051] Overview of Image Encoding Device

[0052] Figure 2 is a diagram schematically showing an image encoding apparatus to which an embodiment according to the present disclosure can be applied.

[0053] As shown in FIG. 2, the image encoding apparatus 100 can include an image dividing unit 110, a subtraction unit 115, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse conversion unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 can be collectively referred to as a “prediction unit”. The conversion unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse conversion unit 150 can be included in a residual processing unit. The residual processing unit can further include the subtraction unit 115.

[0054] All or at least a part of the plurality of components constituting the image encoding apparatus 100 can be realized by one hardware component (for example, an encoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB (decoded picture buffer) and can be realized by a digital storage medium.

[0055] The image segmentation unit 110 can divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). The coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) according to a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units at a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the division of the coding unit, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied later. Based on the final coding unit that cannot be further divided, the coding procedure according to the present disclosure can be performed. The largest coding unit can be used as the final coding unit, and the coding units at a lower depth obtained by dividing the largest coding unit can also be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and / or restoration described later. As another example, the processing unit of the coding procedure can be a prediction unit (PU: Prediction Unit) or a transformation unit (TU: Transform Unit). The prediction unit and the transformation unit can be divided or partitioned from the final coding unit respectively. The prediction unit can be a unit of sample prediction, and the transformation unit can be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.

[0056] The prediction unit (inter prediction unit 180 or intra prediction unit 185) can perform a prediction on a processing target block (current block) and generate a predicted block that includes prediction samples for the current block. The prediction unit can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various information related to the prediction of the current block and transmit it to the entropy encoding unit 190. The information related to the prediction can be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0057] The intra prediction unit 185 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block or at a distance according to the intra prediction mode and / or intra prediction technique. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of fineness of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used based on the setting. The intra prediction unit 185 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the neighboring blocks.

[0058] The inter prediction unit 180 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block can be called by names such as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block can be called a collocated picture (colPic). For example, the inter prediction unit 180 can construct a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 180 can use the motion information of neighboring blocks as the motion information of the current block. In the case of the skip mode, different from the merge mode, the residual signal cannot be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vectors of neighboring blocks are used as motion vector predictors, and the motion vector difference and the indicator for the motion vector predictor are encoded to signal the motion vector of the current block. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.

[0059] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit can apply intra prediction or inter prediction for the prediction of the current block, and can also apply intra prediction and inter prediction simultaneously. The prediction method of applying intra prediction and inter prediction simultaneously for the prediction of the current block can be called CIIP (combined inter and intra prediction). In addition, the prediction unit can also perform intra block copy (IBC) for the prediction of the current block. Intra block copy can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC is a method of predicting the current block using a restored reference block within the current picture at a position separated from the current block by a predetermined distance. When IBC is applied, the position of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed in the same manner as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.

[0060] The prediction signal generated by the prediction unit can be used to generate a restored signal or can be used to generate a residual signal. The subtraction unit 115 can subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal can be transmitted to the conversion unit 120.

[0061] The conversion unit 120 can apply a conversion technique to the residual signal to generate conversion coefficients (transform coefficients). For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when the relationship information between pixels is represented by a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. The conversion process can also be applied to a pixel block having the same size of a square, or can be applied to a block of variable size that is not square.

[0062] The quantization unit 130 can quantize the transform coefficients and transmit them to the entropy encoding unit 190. The entropy encoding unit 190 can encode the quantized signal (information regarding the quantized transform coefficients) and output it in the form of a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 130 can reorder the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0063] The entropy encoding unit 190 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 190 can also encode, together or separately, information necessary for video / image restoration (such as the values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (such as the encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information can further include information regarding various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The signaling information, the transmitted information, and / or the syntax elements referred to in the present disclosure can be encoded via the above-described encoding procedure and included in the bitstream.

[0064] The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray (registered trademark), HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 can be provided as internal / external elements of the image encoding apparatus 100, or the transmission unit can also be provided as a component of the entropy encoding unit 190.

[0065] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be restored.

[0066] The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 155 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and can also be used for inter prediction of the next picture after passing through filtering as described later.

[0067] The filtering unit 160 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 160 can generate various information related to filtering as described later in the description of each filtering method and transmit it to the entropy encoding unit 190. The information related to filtering can be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0068] The modified restored picture transmitted to the memory 170 can be used as a reference picture by the inter prediction unit 180. When inter prediction is applied through this, the image encoding apparatus 100 can avoid prediction mismatches between the image encoding apparatus 100 and the image decoding apparatus, and can also improve the encoding efficiency.

[0069] The DPB in the memory 170 can store the modified restored picture for use as a reference picture by the inter prediction unit 180. The memory 170 can store the motion information of the block in which the motion information in the current picture has been derived (or encoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 180 for utilization as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 170 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 185.

[0070] Overview of Image Decoding Device

[0071] FIG. 3 is a diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure can be applied.

[0072] As shown in FIG. 3, the image decoding apparatus 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 can be collectively referred to as a "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in a residual processing unit.

[0073] All or at least a part of the plurality of components constituting the image decoding apparatus 200 can be realized by one hardware component (for example, a decoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB and can be realized by a digital storage medium.

[0074] The image decoding apparatus 200 that has received a bitstream including video / image information can execute a process corresponding to the process performed by the image encoding apparatus 100 of FIG. 1 to restore an image. For example, the image decoding apparatus 200 can perform decoding using the processing unit applied in the image encoding apparatus. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. Then, the restored image signal decoded and output via the image decoding apparatus 200 can be reproduced via a reproducing apparatus (not shown).

[0075] The image decoding device 200 can receive the signal output from the image encoding device of FIG. 2 in the form of a bitstream. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. The image decoding device can further use the information regarding the parameter set and / or the general constraint information for decoding the image. The signaling information, the received information, and / or the syntax elements referred to in the present disclosure can be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration, the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the information of the surrounding blocks and the decoded information of the block to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Of the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and the residual values that have undergone entropy decoding in the entropy decoding unit 210, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 220. Also, of the information decoded by the entropy decoding unit 210, the information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives the signal output from the image encoding device can be further provided as an internal / external element of the image decoding device 200, or the receiving unit can also be provided as a component of the entropy decoding unit 210.

[0076] On the other hand, the image decoding device according to the present disclosure can be called a video / image / picture decoding device. The image decoding device can also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 210, and the sample decoder can include at least one of the inverse quantization unit 220, the inverse transform unit 230, the addition unit 235, the filtering unit 240, the memory 250, the inter prediction unit 260, and the intra prediction unit 265.

[0077] In the inverse quantization unit 220, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 220 can reorder the quantized transform coefficients in a two-dimensional block format. In this case, the reordering can be performed based on the coefficient scan order performed in the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain transform coefficients.

[0078] In the inverse conversion unit 230, the conversion coefficient can be inversely converted to obtain a residual signal (residual block, residual sample array).

[0079] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode (prediction technique).

[0080] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described below, which is the same as described in the explanation of the prediction unit of the image encoding device 100.

[0081] The intra prediction unit 265 can predict the current block by referring to samples within the current picture. The explanation of the intra prediction unit 185 can be similarly applied to the intra prediction unit 265.

[0082] The inter prediction unit 260 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the surrounding blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the surrounding blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 260 can construct a motion information candidate list based on the surrounding blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information regarding the prediction can include information indicating the mode (technique) of the inter prediction for the current block.

[0083] The adder 235 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block. The description of the adder 155 can be similarly applied to the adder 235. The adder 235 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after passing through filtering as described later.

[0084] The filtering unit 240 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and can store the modified restored picture in the memory 250, specifically in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.

[0085] The (corrected) restored picture stored in the DPB of the memory 250 can be used as a reference picture in the inter prediction unit 260. The memory 250 can store the motion information of the blocks for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of spatial neighboring blocks or the motion information of temporal neighboring blocks. The memory 250 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 265.

[0086] In this specification, the embodiments described in the filtering unit 160, inter prediction unit 180, and intra prediction unit 185 of the image encoding apparatus 100 can be similarly or correspondingly applied to the filtering unit 240, inter prediction unit 260, and intra prediction unit 265 of the image decoding apparatus 200, respectively.

[0087] Overview of Image Segmentation

[0088] The video / image coding method according to the present disclosure can be performed based on the following image segmentation structure. Specifically, procedures such as prediction, residual processing ((inverse) transformation, (inverse) quantization, etc.), syntax element coding, and filtering described later can be performed based on CTUs, CUs (and / or TUs, PUs) derived based on the image segmentation structure. The image can be divided in block units, and the block division procedure can be performed in the image division unit 110 of the encoding apparatus described above. The division-related information can be encoded by the entropy encoding unit 190 and transmitted to the decoding apparatus in the form of a bitstream. The entropy decoding unit 210 of the decoding apparatus can derive the block division structure of the current picture based on the division-related information obtained from the bitstream, and based on this, perform a series of procedures for image decoding (for example, prediction, residual processing, block / picture restoration, in-loop filtering, etc.).

[0089] A picture can be divided into a sequence of coding tree units (CTUs). FIG. 4 shows an example of a picture divided into CTUs. A CTU can correspond to a coding tree block (CTB). Alternatively, a CTU can include a coding tree block of luma samples and two coding tree blocks of corresponding chroma samples. For example, for a picture including three sample arrays, a CTU can include an N×N block of luma samples and two corresponding blocks of chroma samples.

[0090] Overview of CTU Segmentation

[0091] As described above, a coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) according to a QT / BT / TT (Quad-tree / Binary-tree / Ternary-tree) structure. For example, a CTU can first be divided into a quadtree structure. Then, the leaf nodes of the quadtree structure can be further divided according to a multi-type tree structure.

[0092] The division by a quadtree means dividing the current CU (or CTU) into four equal parts. By the division by a quadtree, the current CU can be divided into four CUs having the same width and the same height. If the current CU is not further divided into a quadtree structure, the current CU corresponds to a leaf node of the quadtree structure. The CU corresponding to the leaf node of the quadtree structure is not further divided and can be used as the aforementioned final coding unit. Alternatively, the CU corresponding to the leaf node of the quadtree structure can be further divided according to a multi-type tree structure.

[0093] FIG. 5 is a diagram showing the division types of blocks according to a multi-type tree structure. The division according to a multi-type tree structure can include two divisions according to a binary-tree structure and two divisions according to a ternary-tree structure.

[0094] The two splits by the binary tree structure can include a vertical binary splitting (SPLIT_BT_VER) and a horizontal binary splitting (SPLIT_BT_HOR). The vertical binary splitting (SPLIT_BT_VER) means splitting the current CU vertically into two equal parts. As shown in FIG. 4, two CUs can be generated by the vertical binary splitting, having the same height as the current CU and a width that is half of the width of the current CU. The horizontal binary splitting (SPLIT_BT_HOR) means splitting the current CU horizontally into two equal parts. As shown in FIG. 5, two CUs can be generated by the horizontal binary splitting, having a height that is half of the height of the current CU and the same width as the current CU.

[0095] The two splits by the ternary tree structure can include a vertical ternary splitting (SPLIT_TT_VER) and a horizontal ternary splitting (SPLIT_TT_HOR). The vertical ternary splitting (SPLIT_TT_VER) splits the current CU vertically in a 1:2:1 ratio. As shown in FIG. 5, two CUs having the same height as the current CU and a width that is 1 / 4 of the width of the current CU and one CU having the same height as the current CU and a width that is half of the width of the current CU can be generated by the vertical ternary splitting. The horizontal ternary splitting SPLIT_TT_HOR splits the current CU horizontally in a 1:2:1 ratio. As shown in FIG. 4, two CUs having a height that is 1 / 4 of the height of the current CU and the same width as the current CU and one CU having a height that is half of the height of the current CU and the same width as the current CU can be generated by the horizontal ternary splitting.

[0096] FIG. 6 is a diagram illustrating a signaling mechanism for block splitting information in a quadtree with nested multi-type tree structure according to the present disclosure.

[0097] Here, the CTU is treated as the root node of the quad-tree, and the CTU is first divided into the quad-tree structure. Information (e.g., qt_split_flag) indicating whether to perform quad-tree splitting on the current CU (CTU or a node of the quad-tree (QT_node)) can be signaled. For example, if the qt_split_flag is the first value (e.g., "1"), the current CU can be divided into the quad-tree. Also, if the qt_split_flag is the second value (e.g., "0"), the current CU is not divided into the quad-tree and becomes a leaf node of the quad-tree (QT_leaf_node). Each leaf node of the quad-tree can then be further divided into a multi-type tree structure. That is, the leaf node of the quad-tree can become a node of the multi-type tree (MTT_node). In the multi-type tree structure, to indicate whether the current node is further divided, a first flag (a first flag, e.g., mtt_split_cu_flag) can be signaled. If the node is further divided (e.g., when the first flag is 1), to indicate the splitting direction, a second flag (a second flag, e.g., mtt_split_cu_verticla_flag) can be signaled. For example, when the second flag is 1, the splitting direction is the vertical direction, and when the second flag is 0, the splitting direction can be the horizontal direction. Then, to indicate whether the splitting type is a binary splitting type or a ternary splitting type, a third flag (a third flag, e.g., mtt_split_cu_binary_flag) can be signaled. For example, when the third flag is 1, the splitting type is the binary splitting type, and when the third flag is 0, the splitting type can be the ternary splitting type. The nodes of the multi-type tree obtained by binary splitting or ternary splitting can be further partitioned into the multi-type tree structure. However, the nodes of the multi-type tree cannot be partitioned into the quad-tree structure.When the first flag is 0, the corresponding node of the multi-type tree is not further split and becomes a leaf node (MTT_leaf_node) of the multi-type tree. The CU corresponding to the leaf node of the multi-type tree can be used as the aforementioned final coding unit.

[0098] Based on the aforementioned mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree splitting mode (MttSplitMode) of the CU can be derived as shown in Table 1. In the following description, the multi-tree splitting mode can be abbreviated as the multi-tree splitting type or splitting type.

[0099]

Table 1

[0100] FIG. 7 shows an example in which a CTU is divided into multiple CUs by applying a multi-type tree after applying a quadtree. In FIG. 7, the thick block edge 710 indicates a quadtree division, and the remaining edges 720 indicate a multi-type tree division. A CU can correspond to a coding block (CB). In one embodiment, a CU can include a coding block of luma samples and two coding blocks of chroma samples corresponding to the luma samples. The chroma component (sample) CB or TB size can be derived based on the luma component (sample) CB or TB size according to the component ratio by the color format (chroma format) of the picture / image, such as 4:4:4, 4:2:2, 4:2:0, etc. When the color format is 4:4:4, the chroma component CB / TB size can be set to be the same as the luma component CB / TB size. When the color format is 4:2:2, the width of the chroma component CB / TB can be set to half of the width of the luma component CB / TB, and the height of the chroma component CB / TB can be set to the height of the luma component CB / TB. When the color format is 4:2:0, the width of the chroma component CB / TB can be set to half of the width of the luma component CB / TB, and the height of the chroma component CB / TB can be set to half of the height of the luma component CB / TB.

[0101] In one embodiment, when the size of the CTU is 128 based on the luma sample unit, the size of the CU can have a size ranging from 128×128, which is the same size as the CTU, to 4×4. In one embodiment, in the case of a 4:2:0 color format (or chroma format), the chroma CB size can have a size ranging from 64×64 to 2×2.

[0102] On the other hand, in one embodiment, the CU size and the TU size can be the same. Or, a plurality of TUs can exist within the CU region. The TU size generally can indicate the luma component (sample) TB (Transform Block) size.

[0103] The TU size can be derived based on a preset maximum allowable TB size (maxTbSize). For example, when the CU size is larger than the maxTbSize, a plurality of TUs (TBs) with the maxTbSize can be derived from the CU, and conversion / inverse conversion can be performed in units of the TU (TB). For example, the maximum allowable luma TB size can be 64×64, and the maximum allowable chroma TB size can be 32×32. If the width or height of the CB divided by the tree structure is larger than the maximum conversion width or height, the CB can be automatically (or implicitly) divided until it satisfies the TB size limitations in the horizontal and vertical directions.

[0104] Also, for example, when intra prediction is applied, the intra prediction mode / type is derived in units of the CU (or CB), and the surrounding reference sample derivation and prediction sample generation procedures can be performed in units of the TU (or TB). In this case, one or more TUs (or TBs) can exist within one CU (or CB) region, and in this case, the plurality of TUs (or TBs) can share the same intra prediction mode / type.

[0105] For a quadtree coding tree scheme with a multi-type tree, the following parameters can be signaled from an encoder to a decoder as SPS syntax elements. For example, CTUsize, which is a parameter indicating the size of the root node of the quadtree; MinQTSize, which is a parameter indicating the minimum allowable size of the leaf nodes of the quadtree; MaxBTSize, which is a parameter indicating the maximum allowable size of the root node of the binary tree; MaxTTSize, which is a parameter indicating the maximum allowable size of the root node of the ternary tree; MaxMttDepth, which is a parameter indicating the maximum allowed hierarchy depth of the multi-type tree split from the leaf nodes of the quadtree; MinBtSize, which is a parameter indicating the minimum allowable leaf node size of the binary tree; and MinTtSize, which is a parameter indicating the minimum allowable leaf node size of the ternary tree. At least one of these can be signaled.

[0106] In one embodiment using the 4:2:0 chroma format, the CTU size can be set to 128×128 luma blocks and two 64×64 chroma blocks corresponding to the luma blocks. In this case, MinQTSize can be set to 16×16, MaxBtSize can be set to 128×128, MaxTtSzie can be set to 64×64, MinBtSize and MinTtSize can be set to 4×4, and MaxMttDepth can be set to 4. Quad-tree partitioning can be applied to the CTU to generate leaf nodes of the quad-tree. The leaf nodes of the quad-tree can be called leaf QT nodes. The leaf nodes of the quad-tree can have a size from 16×16 (e.g., the MinQTSize) to 128×128 (e.g., the CTU size). If the leaf QT node is 128×128, it cannot be further divided into binary / ternary trees. This is because dividing in this case would exceed MaxBtsize and MaxTtszie (e.g., 64×64). In other cases, the leaf QT node can be further divided into multi-type trees. Thus, the leaf QT node is the root node for the multi-type tree, and the leaf QT node can have a multi-type tree depth (mttDepth) of 0. If the multi-type tree depth reaches MaxMttdepth (e.g., 4), no further additional division can be considered. If the width of the multi-type tree node is the same as MinBtSize and is the same as or smaller than 2xMinTtSize, no further additional horizontal division can be considered. If the height of the multi-type tree node is the same as MinBtSize and is the same as or smaller than 2xMinTtSize, no further additional vertical division can be considered. When such division is not considered, the encoding device can omit signaling of the division information. In such a case, the decoding device can derive the division information to a predetermined value.

[0107] On one hand, one CTU can include a coding block of luma samples (hereinafter referred to as "luma block") and two coding blocks of corresponding chroma samples (hereinafter referred to as "chroma blocks"). The coding tree scheme described above can be applied to the luma block and chroma block of the current CU in the same way, or can be applied separately. Specifically, the luma block and chroma block within one CTU can be divided into the same block tree structure, and the tree structure in this case can be represented as SINGLE_TREE. Or, the luma block and chroma block within one CTU can be divided into individual block tree structures, and the tree structure in this case can be represented as DUAL_TREE. That is, when the CTU is divided into a dual tree, the block tree structure for the luma block and the block tree structure for the chroma block can exist separately. At this time, the block tree structure for the luma block can be called DUAL_TREE_LUMA, and the block tree structure for the chroma block can be called DUAL_TREE_CHROMA. For P and B slices / tile groups, the luma block and chroma block within one CTU can be restricted to have the same coding tree structure. However, for I slices / tile groups, the luma block and chroma block can have individual block tree structures with respect to each other. If the individual block tree structure is applied, the luma CTB (Coding Tree Block) can be divided into CUs based on a specific coding tree structure, and the chroma CTB can be divided into chroma CUs based on another coding tree structure. That is, the CUs within the I slice / tile group to which the individual block tree structure is applied are composed of a coding block of the luma component or two coding blocks of the chroma components, which can mean that the CUs of the P or B slice / tile group can be composed of blocks of three color components (luma component and two chroma components).

[0108] In the above, the quadtree coding tree structure with a multi-type tree has been described, but the structure in which the CU is divided is not limited to this. For example, the BT structure and the TT structure can be interpreted as concepts included in a Multiple Partitioning Tree (MPT) structure, and it can be interpreted that the CU is divided by the QT structure and the MPT structure. In an example where the CU is divided by the QT structure and the MPT structure, a syntax element (e.g., MPT_split_type) including information on how many blocks the leaf node of the QT structure is divided into and a syntax element (e.g., MPT_split_mode) including information on whether the leaf node of the QT structure is divided in the vertical or horizontal direction are signaled, so that the division structure can be determined.

[0109] In another example, the CU can be divided in a way different from the QT structure, the BT structure, or the TT structure. That is, unlike the case where the CU at a lower depth is divided into 1 / 4 the size of the CU at a higher depth by the QT structure, or the CU at a lower depth is divided into 1 / 2 the size of the CU at a higher depth by the BT structure, or the CU at a lower depth is divided into 1 / 4 or 1 / 2 the size of the CU at a higher depth by the TT structure, the CU at a lower depth can, in some cases, be divided into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 the size of the CU at a higher depth, and the way the CU is divided is not limited to this.

[0110] Thus, the quadtree coding block structure with the multi-type tree can provide a very flexible block division structure. On the other hand, due to the division types supported by the multi-type tree, in some cases, different division patterns can potentially lead to the same coding block structure result. The encoding device and the decoding device can reduce the data amount of the division information by restricting the occurrence of such redundant division patterns.

[0111] For example, FIG. 8 exemplarily shows redundant division patterns that may occur in binary tree division and ternary tree division. As shown in FIG. 8, the consecutive binary divisions 810 and 820 in one direction at the 2-step level have the same coding block structure as the binary division for the center partition after ternary division. In such a case, the binary division for the center blocks 830, 840 of the ternary tree division can be prohibited. Such a prohibition can be applied to the CUs of all pictures. When such a specific division is prohibited, the signaling of the corresponding syntax element can be modified to reflect the case where it is prohibited in this way, thereby reducing the number of bits signaled for the division. For example, as in the example shown in FIG. 8, when the binary division for the center block of the CU is prohibited, the mtt_split_cu_binary_flag syntax element indicating whether the division is a binary division or a ternary division is not signaled, and its value can be induced by the decoder to be 0.

[0112] Overview of Chroma Format

[0113] The chroma format will be described below. An image can be encoded with encoded data including a luma component (e.g., Y) array and two chroma component (e.g., Cb, Cr) arrays. For example, one pixel of the encoded image can include a luma sample and a chroma sample. A chroma format can be used to indicate the composition format of the luma sample and the chroma sample, and the chroma format is sometimes called a color format.

[0114] In one embodiment, the image can be encoded in various chroma formats such as monochrome, 4:2:0, 4:2:2, 4:4:4. In monochrome sampling, one sample array can exist, and the sample array can be a luma array. In 4:2:0 sampling, one luma sample array and two chroma sample arrays can exist, and each of the two chroma arrays can have a height that is half of the luma array and a width that can also be half of the luma array. In 4:2:2 sampling, one luma sample array and two chroma sample arrays can exist, and each of the two chroma arrays can have the same height as the luma array and a width that can be half of the luma array. In 4:4:4 sampling, one luma sample array and two chroma sample arrays can exist, and each of the two chroma arrays can have the same height and width as the luma array.

[0115] FIG. 9 is a diagram showing the relative positions of luma samples and chroma samples according to one embodiment of 4:2:0 sampling. FIG. 10 is a diagram showing the relative positions of luma samples and chroma samples according to one embodiment of 4:2:2 sampling. FIG. 11 is a diagram showing the relative positions of luma samples and chroma samples according to one embodiment of 4:4:4 sampling. As shown in FIG. 9, in the case of 4:2:0 sampling, the position of the chroma sample can be located at the lower end of the corresponding luma sample. As shown in FIG. 10, in the case of 4:2:2 sampling, the chroma sample can be located overlapping the position of the corresponding luma sample. As shown in FIG. 11, in the case of 4:4:4 sampling, both the luma sample and the chroma sample can be located in an overlapping position.

[0116] The chroma format used in the symbolization device and the decoding device can also be predetermined. Or, in order to be adaptively used in the symbolization device and the decoding device, the chroma format can also be signaled from the symbolization device to the decoding device. In one embodiment, the chroma format can be signaled based on at least one of chroma_format_idc and separate_colour_plane_flag. At least one of chroma_format_idc and separate_colour_plane_flag can be signaled via a higher-level syntax such as DPS, VPS, SPS, or PPS. For example, chroma_format_idc and separate_colour_plane_flag can be included in the SPS syntax as shown in FIG. 12.

[0117] On the other hand, FIG. 13 shows an embodiment of chroma format classification utilizing the signaling of chroma_format_idc and separate_colour_plane_flag. chroma_format_idc can be information indicating the chroma format applied to the encoded image. separate_colour_plane_flag can indicate whether the color array is separated and processed in a specific chroma format. For example, the first value of chroma_format_idc (e.g., 0) can indicate monochrome sampling. The second value of chroma_format_idc (e.g., 1) can indicate 4:2:0 sampling. The third value of chroma_format_idc (e.g., 2) can indicate 4:2:2 sampling. The fourth value of chroma_format_idc (e.g., 3) can indicate 4:4:4 sampling.

[0118] In 4:4:4 sampling, the following content can be applied based on the value of separate_colour_plane_flag. When the value of separate_colour_plane_flag is the first value (e.g., 0), each of the two chroma arrays can have the same height and the same width as the luma array. In such a case, the value of ChromaArrayType indicating the type of chroma sample array can be set the same as chroma_format_idc. If the value of separate_colour_plane_flag is the second value (e.g., 1), the luma, Cb, and Cr sample arrays can be processed separately so that they can be processed in the same way as monochrome sampled pictures respectively. At this time, ChromaArrayType can be set to 0.

[0119] Overview of Intra Prediction Mode

[0120] Hereinafter, the intra prediction mode will be described in more detail. FIG. 14 is a diagram showing the intra prediction direction according to an embodiment. In order to capture any edge direction presented in a natural video, as shown in FIG. 14, the intra prediction mode can include two non-directional intra prediction modes and 65 directional intra prediction modes. The non-directional intra prediction modes can include the Planar intra prediction mode and the DC intra prediction mode, and the directional intra prediction modes can include the intra prediction modes numbered from 2 to 66.

[0121] On the one hand, in addition to the intra prediction modes described above, the intra prediction mode can further include a CCLM (cross-component linear model) mode for chroma samples. The CCLM mode can be divided into L_CCLM, T_CCLM, and LT_CCLM depending on whether the left samples, the upper samples, or both are considered for deriving the LM parameters, and can be applied only to the chroma components. For example, as shown in the following table, the intra prediction mode can be indexed according to the intra prediction mode value.

[0122]

Table 2

[0123] FIG. 15 is a diagram showing an intra prediction direction according to another embodiment. Here, the broken line direction indicates a wide-angle mode that is applied only to blocks that are not squares. As shown in FIG. 15, in order to capture any edge direction presented in a natural video, an intra prediction mode according to an embodiment can include 93 directional intra prediction modes together with two non-directional intra prediction modes. The non-directional intra prediction modes can include a Planar prediction mode and a DC prediction mode. The directional intra prediction modes can include intra prediction modes consisting of No. 2 to No. 80 and No. -1 to No. -14, as indicated by the arrows in FIG. 15. The Planar prediction mode can be denoted as INTRA_PLANAR, and the DC prediction mode can be denoted as INTRA_DC. And the directional intra prediction modes can be denoted as INTRA_ANGULAR-14 to INTRA_ANGULAR-1 and INTRA_ANGULAR2 to INTRA_ANGULAR80. On the other hand, the intra prediction type (or additional intra prediction mode, etc.) can include at least one of the above-described LIP, PDPC, MRL, ISP, and MIP. The intra prediction type can be indicated based on intra prediction type information, and the intra prediction type information can be realized in various forms. As an example, the intra prediction type information can include intra prediction type index information indicating one of the intra prediction types. As another example, the intra prediction type information can include reference sample line information (for example, intra_luma_ref_idx) indicating whether the MRL is applied to the current block and, if applied, which reference sample line is used, ISP flag information (for example, intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block, ISP type information (for example, intra_subpartitions_split_flag) indicating the split type of the subpartition when the ISP is applied, flag information indicating whether PDPC is applied or flag information indicating whether LIP is applied, and at least one of MIP flag information indicating whether MIP is applied.

[0124] The intra prediction mode information and / or the intra prediction type information can be encoded / decoded through the coding method described in the present disclosure. For example, the intra prediction mode information and / or the intra prediction type information can be encoded / decoded through entropy coding (e.g., CABAC, CAVLC) based on truncated (rice) binary code.

[0125] When intra prediction is performed on the current block, prediction can be performed on the luma component block (luma block) of the current block and on the chroma component block (chroma block). In this case, the intra prediction mode for the chroma block can be set independently of the intra prediction mode for the luma block.

[0126] For example, the intra prediction mode for the chroma block can be indicated based on intra chroma prediction mode information, and the intra chroma prediction mode information can be signaled in the form of an intra_chroma_pred_mode syntax element. As an example, the intra chroma prediction mode information can indicate any one of Planar mode, DC mode, vertical mode, horizontal mode, DM (Derived Mode), and CCLM mode. Here, the Planar mode can indicate the 0th intra prediction mode, the DC mode can indicate the 1st intra prediction mode, the vertical mode can indicate the 26th intra prediction mode, and the horizontal mode can indicate the 10th intra prediction mode, respectively. DM may also be called direct mode. CCLM may also be called LM.

[0127] On one hand, DM and CCLM are dependent intra prediction modes that predict chroma blocks using luma block information. The DM can indicate a mode in which the same intra prediction mode for the luma component is applied as the intra prediction mode for the chroma component. Also, the CCLM can indicate an intra prediction mode in which, after subsampling the restored samples of the luma block in the process of generating a prediction block for the chroma block, the samples derived by applying CCLM parameters α and β to the subsampled samples are used as the prediction samples of the chroma block.

[0128] Overview of MIP Mode

[0129] The matrix-based intra prediction mode (MIP) may also be called the ALWIP (affine linear weighted intra prediction) mode, the LWIP (linear weighted intra prediction) mode, or the MWIP (matrix weighted intra prediction) mode. An intra prediction mode that is not a matrix-based prediction can be defined as a non-matrix-based prediction mode. For example, the non-matrix-based prediction mode can indicate non-directional intra prediction and directional intra prediction. Hereinafter, the terms intra prediction mode or general intra prediction mode are used interchangeably as terms for indicating the non-matrix-based prediction mode. Hereinafter, matrix-based prediction is described by referring to it as the MIP mode.

[0130] When the MIP mode is applied to the current block, prediction samples for the current block can be derived by: i) using the neighboring reference samples for which an averaging step has been performed, ii) performing a matrix-vector-multiplication step, and iii) further performing a horizontal / vertical interpolation step as necessary.

[0131] The averaging step can be performed by averaging the values of the surrounding samples. As shown in FIG. 16(a), if the width of the current block and the width are both 4 in pixel units, the averaging procedure can be performed by taking the average of each boundary surface and generating a total of 4 samples, i.e., 2 upper samples and 2 left samples. As shown in FIG. 16(b), if the width of the current block and the width are not both 4 in pixel units, the averaging procedure can be performed by taking the average of each boundary surface and generating a total of 8 samples, i.e., 4 upper samples and 4 left samples.

[0132] The matrix-vector multiplication step can be performed by multiplying the averaged samples by a matrix-vector and then adding an offset vector, as a result of which a prediction signal for the subsampled pixel set of the original block can be generated. The sizes of the matrix and the offset vector can be determined according to the width of the current block.

[0133] The horizontal / vertical interpolation step is a step of generating a prediction signal of the original block size from the subsampled prediction signal. As shown in FIG. 17, a prediction signal of the original block size can be generated by performing vertical and horizontal interpolation using the subsampled prediction signal and the surrounding pixel values. FIG. 17 shows an example in which MIP prediction is performed for an 8×8 block. In the case of an 8×8 block, as shown in FIG. 16(b), a total of 8 averaged samples can be generated. By multiplying the 8 averaged samples by a matrix-vector and adding an offset vector, 16 sample values can be generated at even coordinate positions as shown in FIG. 17(a). Then, as shown in FIG. 17(b), vertical interpolation can be performed using the average value of the upper samples of the current block. Then, as shown in FIG. 17(c), horizontal interpolation can be performed using the left samples of the current block.

[0134] The intra prediction modes used for the MIP mode can be configured to be different from the LIP, PDPC, MRL, ISP intra predictions described above and the intra prediction modes used in normal intra prediction. The intra prediction mode for the MIP mode may be referred to as the MIP intra prediction mode, the MIP prediction mode, or the MIP mode. For example, according to the intra prediction mode for the MIP, the matrix and offset used in the matrix vector product can be set to be different. Here, the matrix may be referred to as the (MIP) weight matrix, and the offset may be referred to as the (MIP) offset vector or the (MIP) bias vector.

[0135] The above-described intra prediction type information can include an MIP flag (e.g., intra_mip_flag) indicating whether the MIP mode is applied to the current block. intra_mip_flag[x0][y0] can indicate whether the current block is predicted in the MIP mode. For example, the first value (e.g., 0) of intra_mip_flag[x0][y0] can indicate that the current block is not predicted in the MIP mode. The second value (e.g., 1) of intra_mip_flag[x0][y0] can indicate that the current block is predicted in the MIP mode.

[0136] When intra_mip_flag[x0][y0] has a second value (e.g., 1), information for the MIP mode can be further obtained from the bitstream. For example, the syntax elements intra_mip_mpm_flag[x0][y0], intra_mip_mpm_idx[x0][y0], and intra_mip_mpm_remainder[x0][y0], which are information indicating the MIP mode of the current block, can be further obtained from the bitstream. When the MIP prediction mode is applied to the current block, an MPM list for MIP can be configured, and the intra_mip_mpm_flag can indicate whether the MIP mode for the current block exists in the MPM list for MIP (or among the MPM candidates). The intra_mip_mpm_idx can indicate the index of the candidate used as the MIP prediction mode of the current block among the candidates in the MPM list when the MIP prediction mode for the current block exists in the MPM list for MIP (i.e., when the value of intra_mip_mpm_flag is 1). Intra_mip_mpm_remainder can indicate the MIP prediction mode of the current block when the MIP prediction mode for the current block does not exist in the MPM list for MIP (i.e., when the value of intra_mip_mpm_flag is 0), and can indicate any one of the overall MIP prediction modes, or can indicate any one of the remaining modes excluding the candidate modes in the MPM list for MIP among the overall MIP prediction modes as the MIP prediction mode of the current block.

[0137] On the other hand, when intra_mip_flag[x0][y0] has a first value (e.g., 0), information for MIP is not obtained from the bitstream, and intra prediction information other than MIP can be obtained from the bitstream. In one embodiment, intra_luma_mpm_flag[x0][y0], which indicates whether an MPM list for general intra prediction is generated, can be obtained from the bitstream.

[0138] When an intra prediction mode is applied to the current block, an MPM list can be configured for that purpose, and the intra_luma_mpm_flag can indicate whether the intra prediction mode for the current block exists in the MPM list (or exists among the MPM candidates). For example, the first value of the intra_luma_mpm_flag (e.g., 0) can indicate that the intra prediction mode for the current block does not exist in the MPM list. The second value of the intra_luma_mpm_flag (e.g., 1) can indicate that the intra prediction mode for the current block exists in the MPM list. When the intra_luma_mpm_flag value is 1, the intra_luma_not_planar_flag can be obtained from the bitstream.

[0139] The intra_luma_not_planar_flag can indicate whether the intra prediction mode of the current block is the planar mode. For example, the first value of the intra_luma_not_planar_flag (e.g., 0) can indicate that the intra prediction mode of the current block is the Planar mode. The second value of the intra_luma_not_planar_flag (e.g., 1) can indicate that the intra prediction mode of the current block is not the planar mode.

[0140] The intra_luma_mpm_idx can be parsed and coded when the intra_luma_not_planar_flag is "true" (i.e., value 1). In one embodiment, the Planar mode can always enter as a candidate in the MPM list. However, by signaling the intra_luma_not_planar_flag first as described above, the Planar mode can be excluded from the MPM list. In this case, an MPM list unified with various intra prediction types (general intra prediction, MRL, ISP, LIP, etc.) described above can be constructed. In this case, the number of candidates in the MPM list can be reduced to 5. The intra_luma_mpm_idx can indicate a candidate used as the intra prediction mode of the current block among the candidates included in the MPM list from which the Planar mode is excluded.

[0141] On the other hand, when the value of intra_luma_mpm_flag is 0, the intra_luma_mpm_remainder can be parsed / coded. The intra_luma_mpm_remainder can indicate any one of the overall intra prediction modes as the intra prediction mode of the current block, or any one of the remaining modes excluding the candidate modes in the MPM list as the intra prediction mode of the current block.

[0142] Overview of Palette Mode

[0143] The palette mode (Pallette mode, PLT mode) will be described below. An encoding device according to an embodiment can encode an image using the palette mode, and a decoding device can decode the image using the palette mode in a corresponding manner. The palette mode can be called, for example, the palette encoding mode, the intra-palette mode, the intra-palette encoding mode, etc. The palette mode can be regarded as a type of intra encoding mode and can also be regarded as any one of the intra prediction methods. However, similar to the skip mode described above, a separate residual value for the block may not be signaled.

[0144] In one embodiment, the palette mode can be used to improve the encoding efficiency when encoding screen content, which is an image generated by a computer containing a significant amount of text and graphics. Generally, the local regions of an image generated from screen content are separated by sharp edges and are represented by a small number of colors. To utilize such characteristics, the palette mode can represent samples for one block with indices indicating color entries in the palette table.

[0145] To apply the palette mode, information about the palette table can be signaled. In one embodiment, the palette table can include index values corresponding to each color. To signal the index values, palette index prediction information can be signaled. The palette index prediction information can include index values for at least a part of the palette index map. The palette index map can map the pixels of the video data to the color indices of the palette table.

[0146] The palette index prediction information can include run value information. For at least a part of the palette index map, the run value information can be information that associates a run value with an index value. One run value can be associated with an escape color index. The palette index map can be generated from the palette index prediction information. For example, at least a part of the palette index map can be generated by determining whether to adjust the index value of the palette index prediction information based on the last index value.

[0147] The current block in the current picture can be encoded or decoded according to the palette index map. When the palette mode is applied, the pixel values in the current coding unit can be represented by a small set of representative color values. Such a set can be named a palette. For pixels having values close to the palette colors, a palette index can be signaled. For pixels having values not belonging to (out-of-range) the palette, the pixel is denoted by an escape symbol, and the quantized pixel value can be directly signaled. In this specification, a pixel or a pixel value can be described by a sample.

[0148] To decode a block encoded in palette mode, the decoder can decode the palette color and index. The palette color can be described as a palette table and can be encoded using a palette table coding tool. An escape flag can be signaled for each coding unit. The escape flag can indicate whether an escape symbol exists in the current coding unit. If an escape symbol exists, the palette table can be incremented by one unit (e.g., index unit), and the last index can be designated as the escape mode. The palette indices of all pixels for one coding unit can form a palette index map and can be encoded using a palette index map coding tool.

[0149] For example, a palette predictor can be maintained to encode the palette table. The palette predictor can be initialized at each slice start point. For example, the palette predictor can be reset to 0. For each entry of the palette predictor, a reuse flag indicating whether it is a part of the current palette can be signaled. The reuse flag can be signaled using run-length coding with a value of 0.

[0150] Then, the number for the new palette entry can be signaled using a truncated exponential Golomb code. Finally, the component values for the new palette entry can be signaled. After encoding the current coding unit, the palette predictor can be updated using the current palette, and the entries from the previous palette predictor that are not reused in the current palette can be appended to the end of the new palette predictor (until the maximum allowed size is reached), which can be called palette stuffing.

[0151] For example, to encode a palette index map, the index can be encoded using a horizontal or vertical scan. The scan order can be signaled via a bitstream using a parameter palette_transpose_flag that indicates the scan direction. For example, if a horizontal scan is applied to scan the index for samples in the current coding unit, palette_transpose_flag can have a first value (e.g., 0), and if a vertical scan is applied, palette_transpose_flag can have a second value (e.g., 1). FIG. 18 shows examples of horizontal and vertical scans according to an embodiment.

[0152] Also, in one embodiment, the palette index can be encoded using an "INDEX" mode and a "COPY_ABOVE" mode. When a horizontal scan is used, if the mode of the palette index is signaled for the topmost row, when a vertical scan is used and the mode of the palette index is signaled for the leftmost column, and except when the previous mode is "COPY_ABOVE", the two modes can be signaled using one flag.

[0153] In the "INDEX" mode, the palette index can be explicitly signaled. For the "INDEX" mode and the "COPY_ABOVE" mode, a run value indicating the number of encoded pixels can be signaled using the same mode.

[0154] The coding order for the index map can be set as follows. First, the number of index values for the coding unit can be signaled. This can be done after signaling the actual index values for the entire coding unit using truncated binary coding. Both the number of indexes and the index values can be coded in bypass mode. Thereby, the bypass bins associated with the indexes can be grouped. Then, the palette mode (INDEX or COPY_ABOVE) and the run values can be signaled in an interleaved manner.

[0155] Finally, the component escape values corresponding to the escape samples for the entire coding unit can be grouped together and coded in bypass mode. An additional syntax element, the last_run_type_flag, can be signaled after signaling the index values. By using the last_run_type_flag together with the number of indexes, signaling of the run value corresponding to the last run in the block can be omitted.

[0156] In one embodiment, a dual-tree type that performs independent coding unit partitioning for the luma and chroma components can be used for I slices. The palette mode can be applied to the luma and chroma components respectively or together. If the dual-tree is not applied, the palette mode can be applied to all of the Y, Cb, and Cr components.

[0157] Overview of IBC (Intra Block Copy) Mode

[0158] IBC prediction can be performed in the prediction unit of an image encoding device / image decoding device. IBC prediction can be simply called "IBC". The IBC can be used for content image / video coding of games, etc., such as SCC (screen content coding). The IBC basically performs prediction within the current picture, but can be performed in the same manner as inter prediction in terms of deriving a reference block within the current picture. That is, the IBC can use at least one of the inter prediction techniques described in the present disclosure. For example, in IBC, at least one of the above-described motion information (motion vector) derivation methods can be used. At least one of the inter prediction techniques can also be used with some modifications in consideration of the IBC prediction. The IBC can refer to the current picture. Therefore, it can also be called CPR (current picture referencing).

[0159] For IBC, the image encoding device can perform block matching BM to derive an optimal block vector (or motion vector) for the current block (e.g., CU). The derived block vector (or motion vector) can be signaled to the image decoding device via a bitstream using a method similar to the motion information (motion vector) signaling in the aforementioned inter prediction. The image decoding device can derive a reference block for the current block within the current picture via the signaled block vector (motion vector), and thereby derive a prediction signal (predicted block or predicted sample) for the current block. Here, the block vector (or motion vector) can indicate the displacement from the current block to the reference block located in the already restored area within the current picture. Therefore, the block vector (or motion vector) can also be called a displacement vector. Hereinafter, the motion vector in IBC can correspond to the block vector or the displacement vector. The motion vector of the current block can include a motion vector for the luma component (luma motion vector) or a motion vector for the chroma component (chroma motion vector). For example, the luma motion vector for an IBC-coded CU can also be in integer sample units (i.e., integer precision). The chroma motion vector can also be clipped in integer sample units. As described above, IBC can use at least one of the inter prediction techniques. For example, the luma motion vector can be encoded / decoded using the aforementioned merge mode or MVP mode.

[0160] When the merge mode is applied to the luma IBC block, the merge candidate list for the luma IBC block can be configured in the same manner as the merge candidate list in the inter mode. However, in the case of the luma IBC block, temporal neighboring blocks may not be used as merge candidates.

[0161] When the MVP mode is applied to a Luma IBC block, the mvp candidate list for the Luma IBC block can be configured in the same way as the mvp candidate list in the inter mode. However, in the case of a Luma IBC block, a temporal candidate block may not be used as an mvp candidate.

[0162] IBC derives a reference block from the already reconstructed area within the current picture. At this time, in order to reduce the memory consumption and the complexity of the image decoder, only the predefined area among the already reconstructed areas within the current picture can be referenced. The predefined area can include the current CTU containing the current block. Thus, by restricting the referenceable reconstructed area to the predefined area, the IBC mode can be realized hardware-wise using local on-chip memory.

[0163] The image encoding device that executes IBC can search the predefined area to determine the reference block with the smallest RD cost, and derive a motion vector (block vector) based on the positions of the reference block and the current block.

[0164] Whether to apply IBC to the current block can be signaled as IBC execution information at the CU level. Information regarding the signaling method of the motion vector of the current block (IBC MVP mode or IBC skip / merge mode) can be signaled. The IBC execution information can be used to determine the prediction mode of the current block. Therefore, the IBC execution information can be included in the information regarding the prediction mode of the current block.

[0165] In the case of IBC skip / merge mode, among the block vectors in which the merge candidate index is signaled and included in the merge candidate list, it can be used to indicate the block vector used for predicting the current luma block. At this time, the merge candidate list can include the surrounding blocks encoded by IBC. The merge candidate list can include spatial merge candidates and can be configured not to include temporal merge candidates. Also, the merge candidate list can further include HMVP (Histrory-based motion vector predictor) candidates and / or pairwise candidates.

[0166] In the case of IBC MVP mode, the block vector difference value can be encoded in the same way as the motion vector difference value in the aforementioned inter mode. The block vector prediction method can be used by constructing an mvp candidate list that includes two candidates as predictors, similar to the MVP mode in the inter mode. Either one of the two candidates can be derived from the left surrounding block, and the other one can be derived from the upper surrounding block. At this time, candidates can be derived from the surrounding block only when the left or upper surrounding block is encoded by IBC. If the left or upper surrounding block is not available, for example, if it is not encoded by IBC, the default block vector can be included in the mvp candidate list as a predictor. Also, information (e.g., a flag) for indicating either one of the two block vector predictors is signaled as candidate selection information and is used in the same way as the MVP mode in the inter mode. The mvp candidate list can include HMVP candidates and / or zero motion vectors as default block vectors.

[0167] The HMVP candidate may also be called a history-based MVP candidate. The MVP candidate, merge candidate, or block vector candidate used before the encoding / decoding of the current block can be stored in the HMVP list as an HMVP candidate. Thereafter, if the merge candidate list or mvp candidate list of the current block does not contain the maximum number of candidates, the candidates stored in the HMVP list can be added to the merge candidate list or mvp candidate list of the current block as HMVP candidates.

[0168] The pairwise candidate means a candidate derived by selecting two candidates from among the candidates already included in the merge candidate list of the current block in a predetermined order and averaging the two selected candidates.

[0169] Intra Prediction for Chroma Block

[0170] When intra prediction is performed on the current block, prediction can be performed on the luma component block (luma block) of the current block and prediction can be performed on the chroma component block (chroma block). In this case, the intra prediction mode for the chroma block can be set separately from the intra prediction mode for the luma block.

[0171] The intra prediction mode for the chroma block can be determined based on the intra prediction mode of the luma block corresponding to the chroma block. FIG. 19 is a diagram showing a method for determining the intra prediction mode of a chroma block according to an embodiment.

[0172] Referring to FIG. 19, a method for a coding / decoding apparatus according to an embodiment to determine the intra prediction mode of a chroma block will be described. Hereinafter, the decoding apparatus will be described, and the description can be directly applied to the coding apparatus.

[0173] According to the following description, the IntraPredModeC[xCb][yCb] for the chroma block can be derived, and the following parameters can be used in this process.

[0174] - The luma sample coordinates (xCb, yCb) indicating the relative position of the upper left sample of the current chroma block with respect to the position of the upper left luma sample of the current picture

[0175] - cbWidth indicating the width of the current coding block in luma sample units

[0176] - cbHeight indicating the height of the current coding block in luma sample units

[0177] Also, the following description can be used when the current slice including the current chroma block is an I slice and the luma-chroma dual-tree partitioning structure is applied. However, the following description is not limited to the above examples. For example, regardless of whether the current slice is an I slice or not, the following description can be applied, and furthermore, the following description can be commonly applied even when it is not a dual-tree partitioning structure.

[0178] First, the decoding device can determine luma intra prediction mode information (e.g., lumaIntraPredMode) based on the prediction mode of the luma block corresponding to the current chroma block (S1910). A detailed description of this step will be given later.

[0179] Next, the decoding device can determine the chroma intra prediction mode information based on the luma intra prediction mode information and the additional information (S1920). In one embodiment, the decoding device determines whether the CCLM prediction mode obtained from the bitstream is applied, a cclm_mode_flag parameter indicating whether the CCLM prediction mode is applied, a cclm_mode_idx parameter indicating the CCLM mode type applied when the CCLM prediction mode is applied, an intra_chroma_pred_mode parameter indicating the intra prediction mode type applied to the chroma samples, and the chroma intra prediction mode can be determined based on the luma intra prediction mode information (e.g., lumaIntraPredMode) and the table in FIG. 20.

[0180] As an example, the intra_chroma_pred_mode can refer to any one of the Planar mode, DC mode, vertical mode, horizontal mode, DM (Derived Mode), and CCLM (Cross-component linear model) mode. Here, the Planar mode can indicate the 0th intra prediction mode, the DC mode can indicate the 1st intra prediction mode, the vertical mode can indicate the 26th intra prediction mode, and the horizontal mode can indicate the 10th intra prediction mode. DM may also be called the direct mode. CCLM may also be called the LM (linear model). The CCLM mode can include any one of L_CCLM, T_CCLM, and LT_CCLM.

[0181] On the other hand, DM and CCLM are dependent intra prediction modes that predict the chroma block using the information of the luma block. DM can indicate a mode in which the same intra prediction mode as the intra prediction mode of the luma block corresponding to the current chroma block is applied as the intra prediction mode for the chroma block. In the DM mode, the intra prediction mode of the current chroma block can be determined as the intra prediction mode indicated by the luma intra prediction mode information.

[0182] Also, in the process of generating a prediction block for a chroma block, after subsampling the restored samples of the luma block, the CCLM can indicate an intra prediction mode in which samples derived by applying CCLM parameters α and β to the subsampled samples are used as the prediction samples of the chroma block.

[0183] Next, the decoding device can map the chroma intra prediction mode based on the chroma format (S1930). The decoding device according to one embodiment can map the chroma intra prediction mode X determined according to the table in FIG. 20 to a new chroma intra prediction mode Y based on the table in FIG. 21 only when the chroma format is 4:2:2 (for example, when the value of chroma_format_idc is 2). For example, if the value of the chroma intra prediction mode determined according to the table in FIG. 20 is 16, this can be mapped to chroma intra prediction mode 14 according to the mapping table in FIG. 21.

[0184] Determination of Luma Intra Prediction Mode Information

[0185] Hereinafter, step S1910 of determining luma intra prediction mode information (for example, lumaIntraPredMode) will be described in more detail. FIG. 22 is a flowchart for explaining how the decoding device determines luma intra prediction mode information.

[0186] First, the decoding device can identify whether the MIP mode is applied to the luma block corresponding to the current chroma block (S2210). If the MIP mode is applied to the luma block corresponding to the current chroma block, the decoding device can set the value of the luma intra prediction mode information to the INTRA_PLANAR mode (S2220).

[0187] On the other hand, when the MIP mode is not applied to the luma block corresponding to the current chroma block, the decoding device can identify whether the IBC mode or the PLT mode is applied to the luma block corresponding to the current chroma block (S2230).

[0188] When the IBC mode or the PLT mode is applied to the luma block corresponding to the current chroma block, the decoding device can set the value of the luma intra prediction mode information to the INTRA_DC mode (S2240).

[0189] On the other hand, when the IBC mode or the PLT mode is not applied to the luma block corresponding to the current chroma block, the decoding device can set the value of the luma intra prediction mode information to the intra prediction mode of the luma block corresponding to the current chroma block (S2250).

[0190] In the determination process of the luma intra prediction mode information as shown in FIG. 22, various methods can be applied to identify the luma block corresponding to the current chroma block. As described above, the upper left sample position of the current chroma sample can be expressed by the relative coordinates of the luma sample separated from the position of the upper left luma sample of the current picture. Hereinafter, in such a background, a method for identifying the luma block corresponding to the chroma block will be described.

[0191] FIG. 23 is a diagram showing a first embodiment of referring to the luma sample position to identify the prediction mode of the luma block corresponding to the current chroma block. Hereinafter, since steps S2310 to S2350 correspond to steps S2210 to S2250 of FIG. 22 described above, only the differences will be described.

[0192] In the first embodiment of FIG. 23, a first luma sample position corresponding to the upper left sample position of the current chroma block and a second luma sample position determined based on the upper left sample position of the current chroma block and the width and height of the current luma block are referred to. Here, the first luma sample position may be (xCb, yCb). And the second luma sample position may be (xCb + cbWidth / 2, yCb + cbHeight / 2).

[0193] More specifically, in step S2310, the first luma sample position can be identified to determine whether the MIP mode is applied to the luma block corresponding to the current chroma block. More specifically, it can be identified whether the value of the parameter intra_mip_flag[xCb][yCb] indicating the presence or absence of application of the MIP mode identified at the first luma sample position is 1. The first value (for example, 1) of intra_mip_flag[xCb][yCb] can indicate that the MIP mode is applied at the [xCb][yCb] sample position.

[0194] Also, in step S2330, the first luma sample position can be identified to determine whether the IBC or palette mode is applied to the luma block corresponding to the current chroma block. More specifically, it can be identified whether the value of the prediction mode parameter CuPredMode[0][xCb][yCb] of the luma block identified at the first luma sample position is a value indicating the IBC mode (for example, MODE_IBC) or a value indicating the palette mode (for example, MODE_PLT).

[0195] Also, in step S2350, the second luma sample position can be identified to set the luma intra prediction mode information to the intra prediction mode information of the luma block corresponding to the current chroma block. More specifically, the value of the parameter IntraPredModeY[xCb + cbWidth / 2][yCb + cbHeight / 2] indicating the luma intra prediction mode identified at the second luma sample position can be identified.

[0196] However, in the case of such a first embodiment, both the first luma sample position and the second luma sample position are considered to determine the current chroma block's intra prediction mode. Thereby, if only one luma sample position is considered, the encoding and decoding complexity can be reduced compared to the case where two sample positions are considered.

[0197] Hereinafter, a second embodiment and a third embodiment that consider only one luma sample position will be described. FIGS. 24 and 25 are diagrams showing the second embodiment and the third embodiment that refer to the luma sample position to identify the prediction mode of the luma block corresponding to the current chroma block. Hereinafter, steps S2410 to S2450 and steps S2510 to S2550 correspond to steps S2210 to S2250 of FIG. 22 described above, and only the differences will be described.

[0198] In the second embodiment of FIG. 24, the second luma sample position determined based on the upper left sample position of the current chroma block and the width and height of the current luma block can be referred to. Here, the second luma sample position can be (xCb + cbWidth / 2, yCb + cbHeight / 2). Thereby, in the second embodiment, the predicted value of the luma sample corresponding to the center position (xCb + cbWidth / 2, yCb + cbHeight / 2) of the luma block corresponding to the current chroma block can be identified. In one embodiment, when the luma block has an even number of columns and rows, the center position can indicate the position of the lower right sample among the four central samples of the luma block.

[0199] More specifically, in step S2410, the second luma sample position can be identified to determine whether the MIP mode is applied to the luma block corresponding to the current chroma block. More specifically, it can be determined whether the value of the parameter intra_mip_flag[xCb + cbWidth / 2][yCb + cbHeight / 2] indicating the presence or absence of the application of the MIP mode identified at the second luma sample position is 1.

[0200] Also, in step S2430, in order to identify whether IBC or the palette mode is applied to the luma block corresponding to the current chroma block, the second luma sample position can be identified. More specifically, it can be identified whether the value of the prediction mode parameter CuPredMode[0][xCb + cbWidth / 2][yCb + cbHeight / 2] of the luma block identified at the second luma sample position is a value indicating the IBC mode (e.g., MODE_IBC) or a value indicating the palette mode (e.g., MODE_PLT).

[0201] Also, in step S2450, in order to set the luma intra prediction mode information to the intra mode information of the luma block corresponding to the current chroma block, the second luma sample position can be identified. More specifically, the value of the parameter IntraPredModeY[xCb + cbWidth / 2][yCb + cbHeight / 2] indicating the luma intra prediction mode identified at the second luma sample position can be identified.

[0202] In the third embodiment of FIG. 25, the first luma sample position corresponding to the upper left sample position of the current chroma block can be referred to. Here, the first luma sample position can be (xCb, yCb). Thereby, in the third embodiment, the predicted value of the luma sample corresponding to the upper left sample position (xCb, yCb) of the luma block corresponding to the current chroma block can be identified. In one embodiment, when the luma block has an even number of columns and rows,

[0203] More specifically, in step S2510, in order to identify whether the MIP mode is applied to the luma block corresponding to the current chroma block, the first luma sample position can be identified. More specifically, it can be identified whether the value of the parameter intra_mip_flag[xCb][yCb] indicating the presence or absence of application of the MIP mode identified at the first luma sample position is 1.

[0204] Also, in step S2530, in order to identify whether IBC or the palette mode is applied to the luma block corresponding to the current chroma block, the first luma sample position can be identified. More specifically, it can be identified whether the value of the prediction mode parameter CuPredMode[0][xCb][yCb] of the luma block identified at the first luma sample position is a value indicating the IBC mode (e.g., MODE_IBC) or a value indicating the palette mode (e.g., MODE_PLT).

[0205] Also, in step S2550, in order to set the luma intra prediction mode information to the intra prediction mode information of the luma block corresponding to the current chroma block, the first luma sample position can be identified. More specifically, the value of the parameter IntraPredModeY[xCb][yCb] indicating the luma intra prediction mode identified at the first luma sample position can be identified.

[0206] Encoding and Decoding Method

[0207] Hereinafter, with reference to FIG. 26, a method for encoding by an encoding apparatus according to an embodiment using the above-described method and a method for decoding by a decoding apparatus will be described. An encoding apparatus according to an embodiment includes a memory and at least one processor, and the following method can be performed by the at least one processor. Also, a decoding apparatus according to an embodiment includes a memory and at least one processor, and the following method can be performed by the at least one processor. For convenience of explanation, hereinafter, the operation of the decoding apparatus will be described, but the following explanation can be similarly performed in the encoding apparatus.

[0208] First, the decoding device can divide an image and identify a current chroma block (S2610). Next, the decoding device can identify whether a matrix-based intra prediction mode is applied to a first luma sample position corresponding to the current chroma block (S2620). Here, the first luma sample position can be determined based on at least one of the width and height of the luma block corresponding to the current chroma block. For example, the first luma sample position can be determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block.

[0209] Next, if the matrix-based intra prediction mode is not applied, the decoding device can identify whether a predetermined prediction mode is applied to a second luma sample position corresponding to the current chroma block (S2630). Here, the second luma sample position can be determined based on at least one of the width and height of the luma block corresponding to the current chroma block. For example, the second luma sample position can be determined based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block.

[0210] Next, if the predetermined prediction mode is not applied, the decoding device can determine an intra prediction mode candidate for the current chroma block based on the intra prediction mode applied to a third luma sample position corresponding to the current chroma block (S2640). Here, the predetermined prediction mode can be an IBC (Intra Block Copy) mode or a palette mode.

[0211] On the other hand, the first luma sample position can be the same position as the third luma sample position. Or, the second luma sample position can be the same position as the third luma sample position. Or, the first luma sample position, the second luma sample position, and the third luma sample position can be the same position as each other.

[0212] Alternatively, the first luma sample position may be the center position of the luma block corresponding to the current chroma block. For example, the x-component position of the first luma sample position can be determined by adding half of the width of the luma block to the x-component position of the upper left sample of the luma block corresponding to the current chroma block, and the y-component position of the first luma sample position can be determined by adding half of the height of the luma block to the y-component position of the upper left sample of the luma block corresponding to the current chroma block.

[0213] Alternatively, the first luma sample position, the second luma sample position, and the third luma sample position can be determined respectively based on the upper left sample position of the luma block corresponding to the current chroma block, the width of the luma block, and the height of the luma block.

[0214] Application Example

[0215] The exemplary method of the present disclosure is presented in a series of operations for clarity of explanation, but this is not intended to limit the order in which the steps are performed. If necessary, each step can also be performed simultaneously or in a different order. To implement the method according to the present disclosure, it can also include additional other steps in addition to the exemplary steps, or include the remaining steps except for some steps, or include additional other steps except for some steps.

[0216] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) of checking the execution conditions and situations of the operation (step). For example, when it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation of checking whether the predetermined condition is satisfied.

[0217] The various embodiments of the present disclosure do not enumerate all possible combinations, but are for explaining representative aspects of the present disclosure. The matters described in the various embodiments may be applied independently or in combination of two or more.

[0218] In addition, the various embodiments of the present disclosure can be implemented by hardware, firmware, software, or a combination thereof. In the case of implementation by hardware, it can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.

[0219] In addition, the image decoding device and the image encoding device to which the embodiments of the present disclosure are applied can be included in multimedia broadcast transmission / reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conferencing devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, picture phone video devices, and medical video devices, etc., and can be used for processing video signals or data signals. For example, over-the-top (OTT) video devices can include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0220] FIG. 27 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.

[0221] As shown in FIG. 27, a content streaming system to which an embodiment of the present disclosure is applied can generally include an encoding server, a streaming server, a Web server, a media storage, a user device, and a multimedia input device.

[0222] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream, and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a video camera directly generates a bitstream, the encoding server can be omitted.

[0223] The bitstream can be generated by an image encoding method and / or an image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0224] The streaming server transmits multimedia data to a user device based on a user's request via a Web server, and the Web server can serve as a medium for informing the user of what services are available. When the user requests a desired service from the Web server, the Web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system can include a separate control server, and in this case, the control server can play a role in controlling commands / responses between each device in the content streaming system.

[0225] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0226] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device, for example, a smartwatch, smart glass, an HMD (head mounted display), a digital TV, a desktop computer, a digital signage, and the like.

[0227] Each server in the content streaming system can be operated as a distributed server. In this case, the data received from each server can be processed distributively.

[0228] The scope of the present disclosure includes software or machine-executable commands (e.g., an operating system, an application, firmware, a program, etc.) that enable the operations of various example methods to be executed on a device or a computer, and a non-transitory computer-readable medium on which such software or commands are stored and can be executed on a device or a computer.

Industrial Applicability

[0229] Examples according to the present disclosure are available for encoding / decoding images.

Claims

1. An image decoding method performed by an image decoding device, comprising: Segmenting the image to identify a current chroma block; identifying whether a matrix-based intra prediction mode is applied to a first luma sample position corresponding to the current chroma block; identifying whether a predetermined prediction mode is applied to a second luma sample position corresponding to the current chroma block based on the fact that the matrix-based intra prediction mode is not applied; determining a candidate intra prediction mode for the current chroma block based on an intra prediction mode applied to a third luma sample position corresponding to the current chroma block based on the predetermined prediction mode not being applied; determining an intra prediction mode of the current chroma block based on the intra prediction mode candidates and additional information; The additional information includes information indicating whether a cross component linear model (CCLM) prediction mode is applied, information indicating a CCLM mode type, and information indicating an intra prediction mode type applied to the current chroma block, The image decoding method, wherein the first luma sample position, the second luma sample position, and the third luma sample position are the same.

2. The image decoding method of claim 1 , wherein the first luma sample position is determined based on at least one of a width and a height of a luma block corresponding to the current chroma block.

3. The image decoding method of claim 1 , wherein the first luma sample position is determined based on a top left sample position of a luma block corresponding to the current chroma block, a width of the luma block, and a height of the luma block.

4. The image decoding method of claim 1 , wherein the second luma sample position is determined based on at least one of a width and a height of a luma block corresponding to the current chroma block.

5. The image decoding method of claim 1 , wherein the second luma sample position is determined based on a top left sample position of a luma block corresponding to the current chroma block, a width of the luma block, and a height of the luma block.

6. The image decoding method of claim 2 , wherein the first luma sample position is a center position of a luma block corresponding to the current chroma block.

7. an x-component position of the first luma sample position is determined by adding half a width of the luma block to an x-component position of a top-left sample of a luma block corresponding to the current chroma block; 2. The image decoding method of claim 1, wherein a y-component position of the first luma sample position is determined by adding half the height of the luma block to a y-component position of the top-left sample of the luma block corresponding to the current chroma block.

8. 2. The image decoding method of claim 1, wherein the first luma sample position, the second luma sample position, and the third luma sample position are determined based on a top left sample position of a luma block corresponding to the current chroma block, a width of the luma block, and a height of the luma block, respectively.

9. An image coding method performed by an image coding device, comprising: Segmenting the image to identify a current chroma block; identifying whether a matrix-based intra prediction mode is applied to a first luma sample position corresponding to the current chroma block; identifying whether a predetermined prediction mode is applied to a second luma sample position corresponding to the current chroma block based on the fact that the matrix-based intra prediction mode is not applied; determining a candidate intra prediction mode for the current chroma block based on an intra prediction mode applied to a third luma sample position corresponding to the current chroma block based on the predetermined prediction mode not being applied; an intra prediction mode for the current chroma block is determined based on the intra prediction mode candidates and the side information; The additional information includes information indicating whether a cross component linear model (CCLM) prediction mode is applied, information indicating a CCLM mode type, and information indicating an intra prediction mode type applied to the current chroma block, 11. A method for encoding an image, wherein the first luma sample position, the second luma sample position, and the third luma sample position are the same.

10. 1. A method for transmitting a bitstream, comprising: Segmenting the image to identify a current chroma block; identifying whether a matrix-based intra prediction mode is applied to a first luma sample position corresponding to the current chroma block; identifying whether a predetermined prediction mode is applied to a second luma sample position corresponding to the current chroma block based on the fact that the matrix-based intra prediction mode is not applied; determining a candidate intra prediction mode for the current chroma block based on an intra prediction mode applied to a third luma sample position corresponding to the current chroma block based on the predetermined prediction mode not being applied; encoding side information to generate the bitstream; transmitting data including the bitstream; an intra prediction mode of the current chroma block is determined based on the intra prediction mode candidates and the side information; The additional information includes information indicating whether a cross component linear model (CCLM) prediction mode is applied, information indicating a CCLM mode type, and information indicating an intra prediction mode type applied to the current chroma block, 11. A method for transmitting a bitstream, wherein the first luma sample position, the second luma sample position, and the third luma sample position are the same.

Citation Information

Patent Citations

  • Using luma information for chroma prediction with separate luma-chroma framework in video coding

    US20170272748A1

  • Intra video coding using a decoupled tree structure

    US20180048889A1