Image encoding / decoding method, and apparatus for transmitting compressed video data
By employing methods for encoding and decoding chroma block division using luma block information and dual tree structures, the method addresses the challenge of high data volumes in high-resolution video, reducing transmission and storage costs while improving prediction efficiency.
Patent Information
- Application Number
- PCT/KR2025/008795
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-24
- Filing Date
- 2025-06-24
- Publication Date
- 2026-01-02
AI Technical Summary
The increasing demand for high-resolution and high-quality video content, particularly stereoscopic video, leads to higher data volumes, resulting in increased transmission and storage costs due to existing image compression technologies.
A method for encoding and decoding block division information of chroma blocks using block division information of luma blocks, and deriving intra prediction modes under a dual tree structure, optimizing data compression and prediction efficiency.
Reduces the amount of data to be encoded and decoded, enhancing prediction efficiency by leveraging the block division and intra prediction methods.
Smart Images

Figure KR2025008795_02012026_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device for transmitting compressed video data
[0001] The present disclosure relates to a video signal processing method and device.
[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, is increasing across various application fields. As image data becomes higher in resolution and quality, the relative amount of data increases compared to conventional image data. Therefore, transmitting image data using existing media such as wired and wireless broadband lines or storing it using existing storage media leads to increased transmission and storage costs. To address these issues arising from the increasing resolution and quality of image data, high-efficiency image compression technologies can be utilized.
[0003] There are various technologies for image compression, such as inter-picture prediction technology that predicts pixel values included in the current picture from pictures before or after the current picture, intra-picture prediction technology that predicts pixel values included in the current picture using pixel information in the current picture, and entropy encoding technology that assigns short codes to values with high frequency of appearance and long codes to values with low frequency of appearance. Using these image compression technologies, image data can be effectively compressed and transmitted or stored.
[0004] Meanwhile, as demand for high-resolution video grows, so does the demand for stereoscopic video content as a new video service. Discussions are underway on video compression technologies to effectively deliver high-resolution and ultra-high-resolution stereoscopic video content.
[0005] The present disclosure aims to provide a method for encoding / decoding block division information of a chroma block using block division information of a luma block, and a device therefor.
[0006] The present disclosure aims to provide a method for deriving an intra prediction mode of a chroma block under a dual tree structure and a device therefor.
[0007] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.
[0008] A video decoding method according to the present disclosure may include a step of determining a tree type; and a step of determining whether to split the chroma block based on first block split information of the chroma block. At this time, when the tree type indicates a dual tree type, the first block split information of the chroma block is determined based on a first prediction flag decoded from a bitstream, and the first prediction flag indicates whether the first block split information of the chroma block matches a first prediction value, and the first prediction value may be set to a value of the first block split information of a luma block corresponding to the chroma block.
[0009] In the video decoding method according to the present disclosure, when the first block division information of the luma block indicates that the luma block is divided and the first prediction flag indicates that the first block division information of the chroma block matches the first prediction value, whether to apply QT (Quad Tree)-based division to the chroma block is determined based on the second block division information of the chroma block, and the second block division information of the chroma block is determined based on a second prediction flag that indicates whether the second block division information of the chroma block matches the second prediction value, and the second prediction value can be set to a value of the second division information of the luma block.
[0010] In the video decoding method according to the present disclosure, when the first block division information of the luma block indicates that the luma block is not divided and the first prediction flag indicates that the first block division information of the chroma block does not match the first prediction value, second block division information indicating whether to apply QT-based division to the chroma block can be decoded from the bitstream.
[0011] In the video decoding method according to the present disclosure, whether the first prediction flag is decoded can be determined based on a result of comparing the split depth of the chroma block with a threshold value.
[0012] In the video decoding method according to the present disclosure, the first prediction flag can be decoded only when the division depth of the chroma block is smaller than the threshold.
[0013] In the video decoding method according to the present disclosure, when the division depth of the chroma block is not less than the threshold, the first block division information of the chroma block can be directly decoded from the bitstream.
[0014] In the video decoding method according to the present disclosure, when the difference between the maximum segmentation depth of the luma component and the segmentation depth of the chroma block is greater than a threshold, decoding of the first prediction flag is omitted, and the first block segmentation information of the chroma block can be inferred to indicate that the chroma block is segmented.
[0015] In the image decoding method according to the present disclosure, the tree type is determined in units of predefined blocks, and the predefined blocks may have a size smaller than a coding tree block.
[0016] In the image decoding method according to the present disclosure, a single tree type can be fixedly applied to a coding block having a size larger than the above-defined block.
[0017] In the image decoding method according to the present disclosure, the predefined block may be a coding block having a splitting depth having a predefined value.
[0018] In the video decoding method according to the present disclosure, when the tree type is the dual tree type, DM (Direct Mode) can be fixedly applied when deriving the intra prediction mode of the chroma block.
[0019] In the video decoding method according to the present disclosure, when there are a plurality of luma blocks corresponding to the chroma block, the intra prediction mode of the chroma block can be derived by referring to a luma block including a sample at a predefined position among the luma blocks.
[0020] In the video decoding method according to the present disclosure, when there are a plurality of luma blocks corresponding to the chroma block, the intra prediction mode with the highest occurrence frequency for the luma blocks can be derived as the intra prediction mode of the chroma block.
[0021] A video encoding method according to the present disclosure may include a step of encoding information indicating a tree type; and a step of determining the first block division information of the chroma block indicating whether to divide the chroma block. At this time, when the tree type is a dual tree type, a first prediction flag may be encoded in a bitstream instead of the first block division information of the chroma block, and the first prediction flag may indicate whether the first block division information of the chroma block matches a first prediction value, and the first prediction value may be set to a value of the first block division information of a luma block corresponding to the chroma block.
[0022] According to the present disclosure, a computer-readable recording medium having recorded thereon a command for storing / transmitting a bitstream generated by an image encoding method can be provided.
[0023] According to the present disclosure, a computer-readable recording medium having recorded thereon a command for performing an image decoding method or an image encoding method can be provided.
[0024] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.
[0025] According to the present disclosure, there is an effect of reducing the amount of data to be encoded / decoded by encoding / decoding block division information of a chroma block using block division information of a luma block.
[0026] According to the present disclosure, by providing a method for deriving an intra prediction mode of a chroma block under a dual tree structure, there is an effect of reducing the amount of data to be encoded / decoded and increasing prediction efficiency.
[0027] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.
[0028] FIG. 1 is a block diagram illustrating an image encoding device according to an embodiment of the present disclosure.
[0029] FIG. 2 is a block diagram showing an image decoding device according to an embodiment of the present disclosure.
[0030] Figure 3 illustrates an example in which decryption is performed in empty units.
[0031] Figure 4 shows a decryption method based on a general coding engine.
[0032] Figure 5 shows an example in which the variable ivlCurrRange is updated identically to the variable ivlMpsRange.
[0033] Figure 6 is a flowchart showing the renormalization process.
[0034] Figure 7 shows a decryption process based on a bypass coding engine.
[0035] Figure 8 is a schematic diagram of a typical video signal.
[0036] Figures 9 to 12 illustrate a block division method according to the present disclosure.
[0037] Figure 13 is a diagram for explaining an example in which block division information for determining the block division structure of a block is encoded / decoded.
[0038] Figure 14 illustrates a tree structure for luma components and chroma components under a single tree type.
[0039] Figure 15 illustrates a tree structure for luma components and chroma components under a dual tree type.
[0040] Figure 16 illustrates an example in which block division information of a chroma block is encoded / decoded with reference to division information of a luma block.
[0041] Figure 17 is a flowchart showing a method of encoding / decoding block division information according to tree type.
[0042] FIG. 18 illustrates an image encoding / decoding method performed by an image encoding / decoding device according to the present disclosure.
[0043] FIG. 19 illustrates an example of multiple intra prediction modes according to the present disclosure.
[0044] Figure 20 shows an example of an extended directional mode.
[0045] Figure 21 is a diagram for explaining a luma block referenced to derive an intra prediction mode of a chroma block.
[0046] Figure 22 is a diagram illustrating an example in which chroma blocks share one intra prediction mode.
[0047] Figure 23 shows an example where segmented chroma blocks share one intra prediction mode.
[0048] The present disclosure may be modified in various ways and encompasses numerous embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.
[0049] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0050] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0051] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0052] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the attached drawings. Hereinafter, identical components in the drawings will be designated by the same reference numerals, and redundant descriptions of identical components will be omitted.
[0053] FIG. 1 is a block diagram illustrating an image encoding device according to an embodiment of the present disclosure.
[0054] Referring to FIG. 1, a video encoding device (100) may include a picture segmentation unit (110), a prediction unit (120, 125), a transformation unit (130), a quantization unit (135), a reordering unit (160), an entropy encoding unit (165), an inverse quantization unit (140), an inverse transformation unit (145), a filter unit (150), and a memory (155).
[0055] Each component shown in Fig. 1 is independently depicted to represent different characteristic functions in the video encoding device, and does not mean that each component is composed of separate hardware or a single software component. That is, each component is listed and included as a separate component for convenience of explanation, and at least two components among each component may be combined to form a single component, or one component may be divided into multiple components to perform a function, and such integrated and separate embodiments of each component are also included in the scope of the present disclosure as long as they do not deviate from the essence of the present disclosure.
[0056] Additionally, some components may not be essential components that perform the essential functions of the present disclosure, but may be optional components merely used to enhance performance. The present disclosure may be implemented by including only components essential to implementing the essence of the present disclosure, excluding components used solely for performance enhancement. A structure that includes only essential components, excluding optional components used solely for performance enhancement, is also within the scope of the present disclosure.
[0057] The picture splitting unit (110) can split the input picture into at least one processing unit. At this time, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The picture splitting unit (110) can split one picture into a combination of multiple coding units, prediction units, and transform units, and select one combination of coding units, prediction units, and transform units based on a predetermined criterion (e.g., a cost function) to encode the picture.
[0058] For example, a picture can be split into multiple coding units. A recursive tree structure such as a quad tree, a ternary tree, or a binary tree can be used to split a coding unit in a picture. A coding unit that is split into other coding units starting from an image or the largest coding unit as the root can be split into as many child nodes as the number of split coding units. A coding unit that cannot be split any further according to a certain restriction becomes a leaf node. For example, assuming that a quad tree split is applied to a coding unit, a coding unit can be split into at most four different coding units.
[0059] Hereinafter, in the embodiments of the present disclosure, the encoding unit may be used to mean a unit that performs encoding or may be used to mean a unit that performs decoding.
[0060] A prediction unit may be divided into at least one square or rectangular shape of the same size within a single coding unit, or may be divided such that one prediction unit among the divided prediction units within a single coding unit has a different shape and / or size from another prediction unit.
[0061] When predicting within a screen, the transformation unit and the prediction unit can be set to be the same. In this case, the encoding unit can be divided into multiple transformation units, and then intra-screen prediction can be performed for each transformation unit. The encoding unit can be divided in the horizontal direction or the vertical direction. The number of transformation units generated by dividing the encoding unit can be 2 or 4, depending on the size of the encoding unit. Alternatively, when the size of the transformation unit is small, multiple transformation units can be set as a single prediction unit.
[0062] The prediction unit (120, 125) may include an inter-prediction unit (120) that performs inter-prediction and an intra-prediction unit (125) that performs intra-prediction. It may be determined whether to use inter-prediction or intra-prediction for an encoding unit, and specific information (e.g., reference sample line, intra-prediction mode, motion vector, reference picture, etc.) according to each prediction method may be determined. At this time, the processing unit where prediction is performed and the processing unit where the prediction method and specific contents are determined may be different. For example, the prediction method and prediction mode, etc. are determined in the encoding unit, and the prediction may be performed in the prediction unit or the transformation unit. The residual value (residual block) between the generated prediction block and the original block may be input to the transformation unit (130). In addition, the prediction mode information, motion vector information, etc. used for prediction may be encoded together with the residual value in the entropy encoding unit (165) and transmitted to the decoding device. When using a specific encoding mode, it is also possible to encode the original block as is and transmit it to the decoding unit without generating a prediction block through the prediction unit (120, 125).
[0063] The inter-screen prediction unit (120) may predict a prediction unit based on information of at least one picture among the previous or subsequent pictures of the current picture, and in some cases, may predict a prediction unit based on information of a portion of an encoded region within the current picture. The inter-screen prediction unit (120) may include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.
[0064] The reference picture interpolation unit can receive reference picture information from the memory (155) and generate pixel information less than an integer pixel from the reference picture. In the case of luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 4 pixels. In the case of a chrominance signal, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 8 pixels.
[0065] The motion prediction unit can perform motion prediction based on a reference picture interpolated by the reference picture interpolation unit. Various methods can be used to derive a motion vector, such as FBMA (Full search-based Block Matching Algorithm), TSS (Three Step Search), and NTS (New Three-Step Search Algorithm). The motion vector can have a motion vector value of 1 / 2 or 1 / 4 pixel unit based on the interpolated pixel. The motion prediction unit can predict the current prediction unit by using different motion prediction methods. Various methods can be used as motion prediction methods, such as the Skip method, the Merge method, the AMVP (Advanced Motion Vector Prediction) method, and the Intra Block Copy method.
[0066] The on-screen prediction unit (125) can generate a prediction block based on reference pixel information, which is pixel information within the current picture. The reference pixel information can be derived from one selected from among a plurality of reference pixel lines. The Nth reference pixel line among the plurality of reference pixel lines can include left pixels having an x-axis difference of N from the upper left pixel within the current block and upper pixels having a y-axis difference of N from the upper left pixel. The number of reference pixel lines that the current block can select can be 1, 2, 3, or 4.
[0067] If the neighboring blocks of the current prediction unit are blocks that have performed inter-screen prediction and the reference pixel is a pixel that has performed inter-screen prediction, the reference pixel included in the block that has performed inter-screen prediction can be replaced with the reference pixel information of the neighboring block that has performed intra-screen prediction. That is, if the reference pixel is unavailable, the unavailable reference pixel information can be replaced with information from at least one of the available reference pixels.
[0068] In intra-screen prediction, the prediction mode can have a directional prediction mode that uses reference pixel information according to the prediction direction, and a non-directional mode that does not use directional information when performing prediction. The mode for predicting luminance information and the mode for predicting chrominance information can be different, and the intra-screen prediction mode information used to predict luminance information or the predicted luminance signal information can be utilized to predict chrominance information.
[0069] When performing intra-screen prediction, if the size of the prediction unit and the size of the transformation unit are the same, intra-screen prediction for the prediction unit can be performed based on the pixels on the left side of the prediction unit, the pixels on the upper left side, and the pixels on the upper side.
[0070] The on-screen prediction method can generate prediction blocks by applying a smoothing filter to reference pixels according to the prediction mode. Depending on the selected reference pixel line, whether or not the smoothing filter is applied can be determined.
[0071] In order to perform an intra-screen prediction method, the intra-screen prediction mode of the current prediction unit can be predicted from the intra-screen prediction modes of prediction units existing around the current prediction unit. When the prediction mode of the current prediction unit is predicted using mode information predicted from the surrounding prediction units, if the intra-screen prediction modes of the current prediction unit and the surrounding prediction units are the same, information indicating that the prediction modes of the current prediction unit and the surrounding prediction units are the same can be transmitted using predetermined flag information, and if the prediction modes of the current prediction unit and the surrounding prediction units are different, entropy encoding can be performed to encode the prediction mode information of the current block.
[0072] Additionally, a residual block containing residual value information, which is the difference between the prediction unit that performed the prediction based on the prediction unit generated in the prediction unit (120, 125) and the original block of the prediction unit, can be generated. The generated residual block can be input to the transformation unit (130).
[0073] In the transformation unit (130), the residual block including the residual value information of the prediction unit generated through the original block and the prediction unit (120, 125) can be transformed using a transformation method such as DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), or KLT. Whether to apply DCT, DST, or KLT to transform the residual block can be determined based on at least one of the size of the transformation unit, the shape of the transformation unit, the prediction mode of the prediction unit, or the prediction mode information within the screen of the prediction unit. Meanwhile, the transformation can be performed by separating the horizontal direction and the vertical direction.
[0074] After performing transformations in the horizontal and vertical directions, a secondary transformation can be performed. The secondary transformation may be in a form in which the horizontal and vertical directions are not separated. The secondary transformation can be performed on the transformation coefficients obtained by the primary transformation to generate final transformation coefficients. Meanwhile, the number of final transformation coefficients output by the secondary transformation may be smaller than the number of transformation coefficients input for the secondary transformation. Specifically, the secondary transformation can be performed using a reduced transformation matrix having different numbers of columns and rows.
[0075] The quantization unit (135) can quantize values converted to the frequency domain by the transformation unit (130). The quantization coefficients can vary depending on the block or the importance of the image. The values produced by the quantization unit (135) can be provided to the dequantization unit (140) and the reordering unit (160).
[0076] The rearrangement unit (160) can perform rearrangement of coefficient values for quantized residual values.
[0077] The reordering unit (160) can change a two-dimensional block-shaped coefficient into a one-dimensional vector form through a coefficient scanning method. For example, the reordering unit (160) can change the two-dimensional block-shaped coefficient into a one-dimensional vector form by scanning from the DC coefficient to the coefficient of the high-frequency region using a zig-zag scan method. Depending on the size of the conversion unit and the intra-screen prediction mode, a vertical scan that scans the two-dimensional block-shaped coefficient in the column direction, a horizontal scan that scans the two-dimensional block-shaped coefficient in the row direction, or a diagonal scan that scans the two-dimensional block-shaped coefficient in the diagonal direction may be used instead of the zig-zag scan. That is, depending on the size of the conversion unit and the intra-screen prediction mode, it is possible to determine which scan method among the zig-zag scan, the vertical scan, the horizontal scan, or the diagonal scan is to be used.
[0078] The entropy encoding unit (165) can perform entropy encoding based on the values produced by the rearrangement unit (160). Entropy encoding can use various encoding methods such as, for example, Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).
[0079] The entropy encoding unit (165) can encode various information such as residual value coefficient information of the encoding unit, block type information, prediction mode information, division unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information from the rearrangement unit (160) and the prediction unit (120, 125).
[0080] The entropy encoding unit (165) can entropy encode the coefficient values of the encoding unit input from the rearrangement unit (160).
[0081] The inverse quantization unit (140) and the inverse transformation unit (145) inversely quantize the values quantized in the quantization unit (135) and inversely transform the values transformed in the transformation unit (130). The residual values generated in the inverse quantization unit (140) and the inverse transformation unit (145) can be combined with the predicted prediction units predicted through the motion estimation unit, motion compensation unit, and intra-screen prediction unit included in the prediction unit (120, 125) to generate a reconstructed block.
[0082] The filter unit (150) may include at least one of a deblocking filter, an offset correction unit, and an ALF (Adaptive Loop Filter).
[0083] A deblocking filter can remove block distortion caused by boundaries between blocks in a reconstructed picture. To determine whether to perform deblocking, a deblocking filter can be applied to the current block based on the pixels contained in several columns or rows within the block. When applying a deblocking filter to a block, a strong filter or a weak filter can be applied depending on the required deblocking filtering strength. Furthermore, when applying a deblocking filter, horizontal and vertical filtering can be processed in parallel when performing vertical and horizontal filtering.
[0084] The offset correction unit can correct the offset from the original image on a pixel-by-pixel basis for an image that has undergone deblocking. To perform offset correction for a specific picture, the pixels contained in the image can be divided into a certain number of regions, the regions to be offset can be determined, and the offset can be applied to those regions. Alternatively, the offset can be applied by considering the edge information of each pixel.
[0085] Adaptive Loop Filtering (ALF) can be performed based on the comparison of the filtered restored image with the original image. After dividing the pixels included in the image into predetermined groups, a filter to be applied to each group can be determined, and filtering can be performed differentially for each group. Information regarding whether to apply ALF can be transmitted by luminance signal for each coding unit (CU), and the shape and filter coefficients of the ALF filter to be applied can vary depending on each block. Furthermore, an ALF filter of the same form (fixed form) can be applied regardless of the characteristics of the target block.
[0086] The memory (155) can store a restoration block or picture produced through the filter unit (150), and the stored restoration block or picture can be provided to the prediction unit (120, 125) when performing inter-screen prediction.
[0087] FIG. 2 is a block diagram showing an image decoding device according to an embodiment of the present disclosure.
[0088] Referring to FIG. 2, the image decoding device (200) may include an entropy decoding unit (210), a rearrangement unit (215), an inverse quantization unit (220), an inverse transformation unit (225), a prediction unit (230, 235), a filter unit (240), and a memory (245).
[0089] When a video bitstream is input to a video encoding device, the input bitstream can be decoded in the opposite procedure to that of the video encoding device.
[0090] The entropy decoding unit (210) can perform entropy decoding in a procedure opposite to that of the entropy encoding unit of the video encoding device. For example, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied in response to the method performed in the video encoding device.
[0091] The entropy decoding unit (210) can decode information related to intra-screen prediction and inter-screen prediction performed in the encoding device.
[0092] The reordering unit (215) can perform reordering based on the method in which the bitstream entropy-decoded by the entropy decoding unit (210) is reordered by the encoding unit. The coefficients expressed in the form of a one-dimensional vector can be reordered by restoring them back to coefficients in the form of a two-dimensional block. The reordering unit (215) can perform reordering by receiving information related to the coefficient scanning performed by the encoding unit and performing reverse scanning based on the scanning order performed by the corresponding encoding unit.
[0093] The dequantization unit (220) can perform dequantization based on the quantization parameters provided from the encoding device and the coefficient values of the rearranged block.
[0094] The inverse transform unit (225) can perform an inverse transform of the transform performed by the transform unit on the quantization result performed by the image encoding device. That is, at least one of an inverse transform of a secondary transform (secondary inverse transform) or an inverse transform for DCT, DST, and KLT (i.e., first inverse transform) can be performed. The inverse transform can be performed based on a transmission unit determined by the image encoding device. The inverse transform unit (225) of the image decoding device can determine a transform matrix for the second inverse transform or a transform technique (e.g., DCT, DST, KLT) for the first inverse transform according to a plurality of pieces of information such as a prediction method, the size and shape of the current block, the prediction mode, and the prediction direction within the screen. Alternatively, information for determining the transform matrix or the transform technique may be explicitly encoded and signaled.
[0095] The prediction unit (230, 235) can generate a prediction block based on prediction block generation related information provided from the entropy decoding unit (210) and previously decoded block or picture information provided from the memory (245).
[0096] As described above, when performing intra-screen prediction in the same manner as the operation in the video encoding device, if the size of the prediction unit and the size of the transformation unit are the same, intra-screen prediction for the prediction unit is performed based on the pixels on the left side of the prediction unit, the pixels on the upper left side, and the pixels on the upper side. However, when performing intra-screen prediction, if the size of the prediction unit and the size of the transformation unit are different, intra-screen prediction can be performed using reference pixels based on the transformation unit. In addition, intra-screen prediction using NxN division only for the minimum coding unit can be used.
[0097] The prediction unit (230, 235) may include a prediction unit determination unit, an inter-screen prediction unit, and an intra-screen prediction unit. The prediction unit determination unit may receive various information such as prediction unit information input from the entropy decoding unit (210), prediction mode information of an intra-screen prediction method, and motion prediction-related information of an inter-screen prediction method, and may distinguish a prediction unit from a current encoding unit and determine whether the prediction unit performs inter-screen prediction or intra-screen prediction. The inter-screen prediction unit (230) may perform inter-screen prediction on the current prediction unit based on information included in at least one of a previous picture or a subsequent picture of the current picture including the current prediction unit, using information necessary for inter-screen prediction of the current prediction unit provided from the video encoding device. Alternatively, inter-screen prediction may be performed based on information on a pre-restored portion of the current picture including the current prediction unit.
[0098] In order to perform inter-screen prediction, it is possible to determine whether the motion prediction method of the prediction unit included in the encoding unit is Skip Mode, Merge Mode, AMVP Mode, or Intra-screen Block Copy Mode based on the encoding unit.
[0099] The intra-screen prediction unit (235) can generate a prediction block based on pixel information within the current picture. If the prediction unit is a prediction unit that has performed intra-screen prediction, intra-screen prediction can be performed based on intra-screen prediction mode information of the prediction unit provided by the video encoding device. The intra-screen prediction unit (235) can include an AIS (Adaptive Intra Smoothing) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a part that performs filtering on the reference pixels of the current block, and can determine and apply whether to apply the filter according to the prediction mode of the current prediction unit. AIS filtering can be performed on the reference pixels of the current block using the prediction mode and AIS filter information of the prediction unit provided by the video encoding device. If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.
[0100] The reference pixel interpolation unit can generate a reference pixel of a pixel unit less than an integer value by interpolating the reference pixel when the prediction mode of the prediction unit is a prediction unit that performs intra-screen prediction based on the pixel value interpolated from the reference pixel. If the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter can generate a prediction block through filtering when the prediction mode of the current block is the DC mode.
[0101] The restored block or picture may be provided to a filter unit (240). The filter unit (240) may include a deblocking filter, an offset correction unit, and an ALF.
[0102] Information regarding whether a deblocking filter has been applied to a corresponding block or picture may be received from a video encoding device, and if a deblocking filter has been applied, information regarding whether a strong or weak filter has been applied. The deblocking filter of the video decoding device may receive information related to the deblocking filter provided by the video encoding device, and the video decoding device may perform deblocking filtering on the corresponding block.
[0103] The offset correction unit can perform offset correction on the restored image based on the type of offset correction applied to the image during encoding and offset value information.
[0104] ALF can be applied to an encoding unit based on information such as whether ALF is applied and ALF coefficient information provided from an encoding device. This ALF information can be provided by being included in a specific parameter set.
[0105] The memory (245) can store a restored picture or block so that it can be used as a reference picture or reference block, and can also provide the restored picture to an output unit.
[0106] As described above, in the following embodiments of the present disclosure, for convenience of explanation, the term coding unit is used as an encoding unit, but it may also be a unit that performs not only encoding but also decoding.
[0107] In addition, the current block represents a block to be encoded / decoded, and may represent a coding tree block (or coding tree unit), an encoding block (or encoding unit), a transform block (or transform unit), a prediction block (or prediction unit), or a block to which an in-loop filter is applied, depending on the encoding / decoding step. In this specification, a 'unit' represents a basic unit for performing a specific encoding / decoding process, and a 'block' may represent a pixel array of a predetermined size. Unless otherwise distinguished, 'block' and 'unit' may be used with the same meaning. For example, in the embodiment described below, an encoding block (coding block) and an encoding unit (coding unit) may be understood to have the same meaning.
[0108] Additionally, the encoding parameters for the current block may be commonly applied to multiple color components for the current block. For example, if the encoding mode of the current block is determined, prediction for the Y component block, the Cb component block, and the Cr component block may be performed based on the encoding mode.
[0109] Alternatively, depending on the color component to be encoded / decoded, the current block may mean a Y component block, a Cb component block, or a Cr component block.
[0110] Furthermore, we will refer to the picture that contains the current block as the current picture.
[0111] When generating a bitstream, the encoder can binarize the syntaxes. Binarization of the syntaxes can be based on CABAC (Context-based Arithmetic Binary Coding). At this time, encoding / decoding of the bitstream can be performed in units of bins. Specifically, the encoder performs encoding in units of bins to output bits, and the decoder receives bits and outputs bins through CABAC.
[0112] Meanwhile, a set of bins can be named an empty string. For example, if the value of the syntax merge_idx is 4, the value of the syntax merge_idx can be binarized to 1110. In this case, 1 and 0 each represent a bin, and 1110 represents an empty string. That is, the syntax merge_idx with a value of 4 can be represented as an empty string composed of 4 bins.
[0113] Each of the bins constituting the empty string can be identified by a bin index. Specifically, the indices can be sequentially assigned from the left to the right of the empty string. For example, if the empty string is 1110, the value of the bin assigned to index 0 can be 1, the value of the bin assigned to index 1 can be 1, the value of the bin assigned to index 2 can be 1, and the value of the bin assigned to index 3 can be 0.
[0114] Meanwhile, encoding / decoding for bins can be performed based on a general coding engine or through a bypass coding engine.
[0115] Figure 3 illustrates an example in which decryption is performed in empty units.
[0116] As shown in the example, depending on the value of the variable bypassFlag, it can be determined whether the decoding of the bin is performed through the general coding engine or the bypass coding engine. Here, the general coding engine may indicate a coding method using contextual information, and the bypass coding engine may indicate a coding method that does not use contextual information.
[0117] The variable bypassFlag is an internal variable defined in the encoder and decoder, which indicates whether the bin to be encoded / decoded is encoded through the bypass coding engine.
[0118] Meanwhile, whether to use the bypass coding engine can be determined for each syntax element or each bin of the syntax element. For example, when encoding / decoding residual coefficients, the value of the variable bypassFlag can be determined based on whether the number of bins encoded through probability encoding reaches a threshold (e.g., CCB (Context Coded Bin)). Alternatively, the value of the variable bypassFlag can be determined depending on the type of the syntax element.
[0119] Depending on the variable bypassFlag, the bin can be encoded / decoded using either a general coding engine or a bypass coding engine. Below, the bin encoding / decoding method will be described in detail.
[0120] To encode / decode bins using CABAC, initialization of the probability and coding engines can be performed.
[0121] The initial probability can be determined based on the slice type and / or the empty index. Accordingly, the initial probability value (initValue) can vary for each empty index. The initial probability value can be expressed in 6 bits.
[0122] Once the initial probability value (initValue) is determined, two probability state indices can be derived using the initial probability value. Equations 1 to 7 illustrate the process of deriving the first probability state index pStateIdx0 and the second probability state index pStateIdx1 using the initial probability value initValue.
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130] The two probability state indices are values that represent the probability that the value of the bin is 1 (i.e., the probability of occurrence of 1). That is, the larger the value of the probability state indices, the greater the probability that the value of the bin is 1.
[0131] The first probability state index and the second probability state index have different speeds at which probabilities are updated. For example, when bins with a value of 1 are continuously input, the first probability state index pStateIdx0 is updated to increase rapidly compared to the second probability state index pStateIdx1. In other words, the second probability state index pStateIdx1 is updated to increase relatively more gradually compared to the first probability state index pStateIdx0.
[0132] Finally, the probability of occurrence of 1 is determined by taking the average of the first probability state index pStateIdx0 and the second probability state index pStateIdx1. Meanwhile, referring to mathematical expressions 6 and 7, there is a 4-bit length difference between the first probability state index pStateIdx0 and the second probability state index pStateIdx1. Accordingly, when calculating the average between the first probability state index pStateIdx0 and the second probability state index pStateIdx1, the precision of the two variables can be adjusted to the same extent. For example, after performing an operation of shifting the first probability state index pStateIdx0 to the left by 4, the average between the shifted first probability state index and the second probability state index pStateIdx1 can be obtained.
[0133] The coding engine can operate based on the variables ivlCurrRange and ivlOffset. The ivlCurrRange variable can be initialized to a predefined value (e.g., 510). Conversely, the ivlOffset variable can be initialized based on information parsed from the bitstream (e.g., 9-bit information).
[0134] Figure 4 shows a decryption method based on a general coding engine.
[0135] To decrypt a single bin, a probability can be set. To this end, a variable pState representing a probability state can be derived. The variable pState can be derived by averaging the first probability state index pStateIdx0 and the second probability state index pStateIdx1. Furthermore, in order to adjust the precision of the two probability state indices to be the same, the first probability state index pStateIdx0 can be shifted to the left by 4, and then the variable pState can be derived. Meanwhile, the variable pState can be a positive integer represented by 15 bits.
[0136] The value with the highest probability of occurrence between 0 and 1 can be set as the Most Probable Symbol (MPS), and the value with the lowest probability of occurrence can be set as the Least Probable Symbol (LPS). Since the value of a bin is either 0 or 1, the sum of the occurrence probability of 0 and the occurrence probability of 1 can be 1.0.
[0137] Depending on the variable pState, it can be determined whether the value of MPS is 0 or 1. The variable valMps, which indicates whether MPS is 0 or 1, can be derived by the following mathematical expression 8.
[0138]
[0139] The variable pState is a positive integer represented by 15 bits. Accordingly, if the value of the variable pState is greater than 16383, valMps can be set to 1. This means that the probability of occurrence of 1 is higher than the probability of occurrence of 0.
[0140] On the other hand, if the value of variable pState is less than or equal to 16383, variable valMps can be set to 0. This means that the probability of occurrence of 0 is higher than the probability of occurrence of 1.
[0141] The variable ivlLpsRange represents the range of LPS. The variable ivlLpsRange can be derived by the following mathematical expressions 9 and 10.
[0142]
[0143]
[0144] The range of MPS, ivlMpsRange, can be derived by differentiating the variable ivlLpsRange from the variable ivlCurrRange.
[0145] As a result, the probability of occurrence of MPS within the ivlCurrRange range is P MPS and the probability of occurrence of LPS P LPS can be defined as in mathematical expression 11.
[0146]
[0147] At this time, the sum of the occurrence probability of MPS and the occurrence probability of LPS can be 1 (i.e., 100%). For example, let's assume that MPS is 1 (i.e., the value of valMPS is 1) and the value of ivlCurrRange is 200. P MPS and P LPS If are 140 and 60 respectively, the probability of occurrence of 1 (i.e., MPS) may be 70%, and the probability of occurrence of 0 (i.e., LPS) may be 30%.
[0148] Afterwards, the variable ivlOffset is derived from the bitstream, and the variable ivlCurrRange is updated. The variable ivlCurrRange can be updated to a value that is the difference between ivlLpsRange and ivlCurrRange, i.e., the same value as ivlMpsRange.
[0149] Figure 5 shows an example in which the variable ivlCurrRange is updated identically to the variable ivlMpsRange.
[0150] Then, compare the sizes of the variables ivlOffset and ivlCurrRange.
[0151] If the variable ivlOffset is greater than or equal to the variable ivlCurrRange, then ivlOffset can be determined to be within the range of the LPS (i.e., ivlLpsRange). Otherwise, the variable ivlOffset can be determined to be within the range of the MPS (i.e., ivlMpsRange).
[0152] Based on the above results, if the variable ivlOffset is determined to belong to the LPS interval, the value set to LPS can be output as the bin value (i.e., variable binVal). On the other hand, if ivlOffset belongs to the MPS interval, the value set to MPS can be output as the bin value (i.e., variable binVal).
[0153] If the variable ivlOffset belongs to the MPS interval, the value of the variable ivlCurrRange remains the same. On the other hand, if the variable ivlOffset belongs to the LPS interval, the variable ivlCurrRange can be updated to the variable ivlLpsRange.
[0154] Similarly, if the variable ivlOffset falls within the LPS interval, the value of the variable ivlOffset can also be updated.
[0155] After the bin values are determined, probability updates are performed. Specifically, the first probability state index pStateIdx0 and the second probability state index pStateIdx1, which indicate the probability of occurrence of 1, can be updated at different rates by the decrypted bin values (i.e., binVal) and a variable that controls the update rate.
[0156] Specifically, the first probability state index pStateIdx0 and the second probability state index pStateIdx1 can be updated at different rates by the first shifting variable shift0 and the second shifting variable shift1 that control the update rate.
[0157] Meanwhile, the first shifting index shift0 and the second shifting variable shift1 can be derived as in the following mathematical expressions 12 and 13.
[0158]
[0159]
[0160] Meanwhile, the variable shiftIdx may be predefined (i.e., a fixed value) in the encoder and decoder for each syntax element to be encoded / decoded.
[0161] The first probability state index pStateIdx0 and the second probability state index pStateIdx1 can be updated as in the following mathematical expressions 14 and 15.
[0162]
[0163]
[0164] After the probability update is performed, a renormalization process can be performed.
[0165] Figure 6 is a flowchart showing the renormalization process.
[0166] As in the example shown in Figure 6, the variable ivlCurrRange is compared with the predefined constant 256. If the variable ivlCurrRange is greater than or equal to 256, renormalization may not be performed.
[0167] Otherwise, updates to the variables ivlCurrRange and ivlOffset may be performed. In Figure 6, read_bits(1) indicates that one bit is read from the bitstream and output.
[0168] Figure 7 shows a decryption process based on a bypass coding engine.
[0169] As in the example illustrated in Fig. 7, the value of the bin (i.e., binVal) can be determined by determining the values of the variables ivlOffset and ivlCurrRange. If the value of the bin is 1, the variable ivlCurrRange can be updated with a value that is less than the variable ivlOffset. On the other hand, if the value of the bin is 0, the variable ivlCurrRange may not be updated.
[0170] In the bypass coding engine, probability information is not utilized. That is, when the bypass coding engine is applied, the probability of occurrence of 0 or 1 is not defined, and the bin values can be encoded / decoded. In other words, when the bypass coding engine is used, the probability of occurrence of 0 and the probability of occurrence of 1 can be set to the same value.
[0171] When a bypass coding engine is used, the number of bins and the number of bits appear to be the same.
[0172] Based on the above characteristics, the bypass coding engine is used for information for which probability settings are meaningless. Furthermore, the bypass coding engine's primary goal is to improve throughput, i.e., processing rate, rather than improving encoding / decoding efficiency through entropy coding.
[0173] Figure 8 is a schematic diagram of a typical video signal.
[0174] As in the example shown in Figure 8, a video can be defined as a set of Group of Pictures (GOPs).
[0175] A GOP can be set as a random access unit. Alternatively, intra-picture insertion within a GOP can prevent restoration errors within the GOP from propagating to the next GOP.
[0176] A GOP can represent a set of pictures.
[0177] A single picture can be divided into multiple regions. For example, a single picture can be divided into multiple sub-pictures, multiple slices, or multiple tiles.
[0178] The subpicture structure can be useful for viewport-based 360-degree VR video streaming. When a single picture is divided into multiple subpictures, the bitstream can be extracted and merged on a subpicture basis. Accordingly, the decoder can decode only the bitstreams of the subpictures required for rendering. A subpicture can be configured to include one or more slices.
[0179] If a picture is divided into multiple slides, data can be encapsulated into packets and signaled on a slice-by-slice basis.
[0180] When a single picture is divided into multiple tiles, encoding / decoding can be performed in parallel between the tiles.
[0181] In an encoder, the current picture can be divided into multiple reference blocks. Here, the reference blocks may be called Coding Tree Units (CTUs) or Coding Tree Blocks (CTBs). A CTU or CTB may be a coding block with the largest size.
[0182] Each of the above-described subpictures, slices and tiles may be a set of reference blocks.
[0183] Meanwhile, the size of the reference block may be predefined in the encoder and decoder. Alternatively, information related to the size of the reference block may be encoded and signaled to the decoder. The information may be encoded / decoded via an upper header. For example, the information may be encoded / decoded via a sequence parameter set or a picture header.
[0184] The reference block may be further divided into multiple blocks (i.e., multiple coding blocks) based on a tree structure partitioning. Here, the tree structure partitioning may include at least one of a quad tree partitioning, a binary tree partitioning, or a ternary tree partitioning.
[0185] As the reference block is divided, when the block to be encoded / decoded (i.e., the leaf node block) is finally determined, the encoder can encode the samples within the block through processes such as prediction, transformation, quantization, and entropy encoding. In addition, the decoder can reconstruct the samples within the block through processes such as entropy decoding, inverse quantization, inverse transformation, and prediction.
[0186] Figures 9 to 11 illustrate a block division method according to the present disclosure.
[0187] In the embodiments described below, a 'block' may represent any one of a coding block, a prediction block, or a transformation block as a target of encoding / decoding.
[0188] A single block can be divided into multiple blocks of various sizes and shapes through a tree structure. These divided blocks can then be further divided into multiple blocks of various sizes and shapes. This recursive division of a block can be defined as "tree-structure-based" division.
[0189] The above tree structure-based segmentation can be performed based on predetermined segmentation information. Here, the segmentation information may be encoded by an encoding device and transmitted through a bitstream, or may be derived from an encoding / decoding device. The segmentation information may include information indicating whether a block is to be segmented (hereinafter referred to as a segmentation flag). If the segmentation flag indicates segmentation of a block, the block is segmented and the encoding order is followed by moving to the next block. Here, the next block refers to the block to be encoded first among the segmented blocks. If the segmentation flag indicates that the block is not to be segmented, after encoding the encoding information of the block, the segmentation process is terminated or moving to the next block depending on whether a next block exists.
[0190] Partition information may include information about tree partitioning. Below, the tree partitioning method used for block partitioning is described.
[0191] The BT (Binary Tree) partitioning method divides a block into two. The blocks created through this partitioning can have the same size. Figure 9 illustrates an example of BT partitioning a block using the BT flag.
[0192] The BT flag can be used to determine whether a block should be split. For example, if the BT flag is 0, BT splitting is terminated. On the other hand, if the BT flag is 1, the block can be split into two blocks using the Dir flag, which indicates the splitting direction.
[0193] Additionally, the segmented blocks can be expressed as depth information. Figure 10 shows an example of depth information.
[0194] FIG. 10 (a) is an example showing the process of splitting a block (400) through BT splitting and the value of depth information (depth). Each time a block is split, the value of depth information can increase by 1. When a block of depth N is split into blocks of depth (N+1), the block of depth N is called the parent block of the blocks of depth (N+1). Conversely, the block of depth (N+1) is called the child block of the block of depth N. This can be applied equally to the tree structure described below. FIG. 10 (b) shows the final split shape when the block (400) is split as in (a) using BT.
[0195] The TT (Ternary-tree) partitioning method divides a block into three parts. The child blocks can have a ratio of 1:2:1. Figure 11 illustrates an example of TT partitioning a block using the TT flag.
[0196] The TT flag can be used to determine whether a block is to be split. For example, if the TT flag is 0, TT splitting is terminated. On the other hand, if the TT flag is 1, the block can be split into three horizontally or vertically using the Dir flag.
[0197] The QT (quad-tree) partitioning method divides a block into four blocks. The four child blocks can have the same size. Figure 12 illustrates an example of QT partitioning a block using the QT flag.
[0198] The QT flag can be used to determine whether a block is to be split. For example, if the QT flag is 0, QT splitting is terminated. On the other hand, if the QT flag is 1, the block can be split into four parts.
[0199] Figure 13 is a diagram for explaining an example in which block division information for determining the block division structure of a block is encoded / decoded.
[0200] Referring to FIG. 13, the block splitting information may include at least one of information indicating whether a block is split (e.g., split_flag), information indicating whether a block is QT split (e.g., QT_flag), information indicating a splitting direction of a block (e.g., DIR_flag), or information indicating whether a block is BT / TT split (e.g., BT_flag). In the embodiment described below, the block splitting information may refer to one of the above-listed information, two or more of the above-listed information, or all of the above-listed information. Whether the block splitting information indicates a single information element or multiple information elements will be clearly understood by those skilled in the art depending on the context.
[0201] The syntax split_flag indicates whether a block is split. For example, a value of 0 for the syntax split_flag indicates that the block is not split. Conversely, a value of 1 for the syntax split_flag indicates that the block is split.
[0202] If the value of the syntax split_flag is 1, the QT_flag indicating whether QT splitting is performed can be additionally encoded / decoded. For example, if the syntax QT_flag is 1, it indicates that QT splitting is performed. If QT splitting is performed, each block generated by QT splitting can be recursively split. That is, split_flag can be encoded / decoded for each block generated by QT splitting.
[0203] A syntax QT_flag of 0 indicates that QT segmentation is not performed. In this case, the DIR_flag indicating the direction of BT or TT segmentation may be further encoded / decoded. For example, a syntax DIR_flag of 1 indicates that the block is segmented horizontally, and a syntax DIR_flag of 0 indicates that the block is segmented vertically.
[0204] Afterwards, BT_flag can be additionally encoded / decoded. Syntax BT_flag can indicate whether the block is BT split or TT split. For example, syntax BT_flag being 1 indicates that BT splitting is applied to the block. Syntax BT_flag being 0 indicates that TT splitting is applied to the block. Each of the blocks generated by BT or TT splitting can be recursively split. That is, split_flag can be encoded / decoded for each of the blocks generated by BT splitting or TT splitting.
[0205] In Fig. 13, the encoding / decoding order of block splitting information is described as being split_flag, QT_flag, DIR_flag, and BT_flag. Unlike what is shown in Fig. 13, QT_flag may be encoded / decoded before split_flag.
[0206] As described above, through tree structure partitioning, a reference block can be partitioned into blocks of various sizes. Here, the tree structure can include at least one of QT, BT, and TT.
[0207] For color images, the tree types for the luma component (i.e., Y component) and chroma components (i.e., Cb and / or Cr components) can represent single tree types or dual tree types.
[0208] The single tree type indicates that the tree structures of the luma component and chroma component are set identically.
[0209] Figure 14 illustrates a tree structure for luma components and chroma components under a single tree type.
[0210] Since the tree structures of the luma component and the chroma component are identical, block segmentation information can be encoded and signaled only for the luma component.
[0211] The dual tree type indicates that the tree structures of the luma component and chroma component are independent.
[0212] Figure 15 illustrates a tree structure for luma components and chroma components under a dual tree type.
[0213] When a dual tree type is applied, information related to the tree structure can be encoded and signaled for each of the luma component and the chroma component.
[0214] Meanwhile, in the example illustrated in Fig. 15, the tree structures of the Cb component and the Cr component are illustrated as being identical. Accordingly, information related to the tree structure can be encoded and signaled for only one of the Cb component and the Cr component.
[0215] Unlike the illustrated example, the tree structures of the Cb and Cr components may be determined independently. In this case, information related to the tree structure may be encoded and signaled for each of the Cb and Cr components.
[0216] For convenience of explanation, in the embodiments described below, it is assumed that the tree structures of the Cb component and the Cr component are the same under the dual tree type.
[0217] The tree type can be determined on a per-reference block basis. That is, information indicating whether the tree type is a single tree type or a dual tree type can be encoded and signaled on a per-reference block basis. The information can be a 1-bit flag. For example, a flag of 0 can indicate that a single tree type is applied, and a flag of 1 can indicate that a dual tree type is applied.
[0218] Meanwhile, depending on the picture type or slice type, whether information indicating the tree type is encoded / decoded can be determined. Here, the picture type or slice type can indicate a type that only allows intra prediction (I type) or a type that allows inter prediction (P type or B type).
[0219] For example, information indicating a tree type may be encoded and signaled only if the picture type or slice type indicates a type that is inter-predictable.
[0220] If the picture type or slice type indicates a type that only allows intra prediction, encoding / decoding of information indicating the tree type may be omitted, and the tree type may be determined as a single tree type. That is, if the picture type or slice type indicates a type that only allows intra prediction, the dual tree type may not be available.
[0221] Alternatively, conversely, if the picture type or slice type indicates a type that only allows intra prediction, encoding / decoding of information indicating the tree type may be omitted, and the tree type may be determined as a dual tree type. That is, if the picture type or slice type indicates a type that only allows intra prediction, a single tree type may not be available.
[0222] Instead of encoding / decoding information indicating a tree type on a per-reference block basis, information indicating a tree type may be encoded / decoded on a per-predefined block basis. Here, the predefined block may be at least one of a coding block of a predetermined size or a coding block having a split depth of a predefined value. In this case, the predefined block may have a different size from the reference block.
[0223] For example, if the number of samples in a coding block is within a preset range, information indicating a tree type for the corresponding coding block can be encoded / decoded. Here, the preset range can be defined by at least one of the maximum or minimum values of the range interval.
[0224] On the other hand, if the number of samples in a coding block does not exist within a preset range, information indicating the tree type may not be encoded / decoded for the coding block.
[0225] Meanwhile, if the number of samples in a coding block does not exist within a preset range, a single tree type or a dual tree type can be fixedly used.
[0226] For example, if the preset range indicates 64 or less, a single tree type may be applied to coding blocks in which the number of samples is greater than 64. Accordingly, for coding blocks in which the number of samples is greater than 64, information related to the tree structure may be encoded and signaled only for the luma component.
[0227] When the number of samples in each coding block generated by dividing the above coding block is 64 or less, information indicating a tree type for each coding block can be encoded / decoded. For a coding block among the divided coding blocks whose tree type indicates a single tree type, information related to the tree structure can be encoded / decoded only for the luma component, and for a coding block indicating a dual tree type, information related to the tree structure can be encoded / decoded for each of the luma component and the chroma component.
[0228] As described above, information indicating a tree type can be encoded / decoded for each of the divided coding blocks.
[0229] Alternatively, information indicating the tree partition type may be encoded / decoded for only one of the partitioned coding blocks. The block in which the information indicating the tree partition type is encoded / decoded may be referred to as a representative coding block. The tree partition type of the representative coding block may be applied equally to coding blocks having the same partition depth.
[0230] Meanwhile, the representative coding block may represent the first coding block among the divided coding blocks or the coding block with the largest size among the divided coding blocks.
[0231] In the above example, the preset range is exemplified as being defined using a first threshold indicating a maximum value. Alternatively, the preset range can be defined using a second threshold indicating a minimum value. For example, the preset range can be set to be greater than or equal to the second threshold.
[0232] Alternatively, the preset range may be defined by a first threshold value and a second threshold value. That is, the preset range may be defined as a range below the first threshold value and above the second threshold value.
[0233] The threshold value may represent the number of samples in a block or the size of the block. Here, the size of the block may represent at least one of the width, height, or the product of the width and height of the block.
[0234] For convenience of explanation, it is assumed that a tree type is determined for each of the partitioned coding blocks.
[0235] When a dual tree type is applied, the tree structure of the luma component block and the tree structure of the chroma component block may be different.
[0236] Furthermore, when a dual tree type is applied, the tree structure available to the luma component block may be different from the tree structure available to the chroma component block. For example, block partitioning based on BT, TT, and QT may be applicable to the luma component block, while only block partitioning based on QT may be applicable to the chroma component block. In other words, block partitioning based on BT and TT may be applicable only to the luma component block, but not to the chroma component block.
[0237] Information indicating the tree structure available for each color component can be encoded and signaled. The information can be encoded / decoded via an upper header. Here, the upper header can indicate a slice, tile, sub-picture, picture, GOP (Group of Pictures), or sequence.
[0238] When a dual tree type is applied, at least one of the maximum split depth allowed for splitting based on a specific tree structure or the minimum block size allowed for splitting based on a specific tree structure may be different for each color component.
[0239] For example, the maximum split depth of a luma component block that is allowed to be split may be different from the maximum split depth of a chroma component block that is allowed to be split.
[0240] Accordingly, the size of the smallest block that cannot be divided any further (i.e., a leaf node block) may also differ between the luma component and the chroma component.
[0241] At least one of information indicating the maximum splitting depth allowed for splitting or the minimum block size allowed for splitting may be encoded and signaled for each color component. The information may be encoded and signaled via an upper header.
[0242] Meanwhile, instead of information indicating the minimum block size, information indicating the difference between the minimum block size at which division for a luma component block is allowed and the minimum block size at which division for a chroma component block is allowed may be encoded / decoded.
[0243] For example, for the luma component, information indicating the minimum block size for which segmentation is allowed can be encoded / decoded, and for the chroma component, differential information can be encoded / decoded.
[0244] If the tree type is a dual tree type, the block division information of the chroma block can also be encoded / decoded by referring to the division information of the luma block.
[0245] Figure 16 illustrates an example in which block division information of a chroma block is encoded / decoded with reference to division information of a luma block.
[0246] After setting the block division structure / block division information of the same-location luma block as the predicted value, information indicating whether the block division structure / block division information of the chroma block matches the predicted value can also be encoded / decoded. The above information can be a 1-bit flag, and the flag can be called a "prediction flag."
[0247] That is, instead of directly encoding / decoding block division information for chroma blocks, prediction flags for block division information can be encoded / decoded.
[0248] For convenience of explanation, it is assumed that a prediction flag of 1 (True) indicates that the block partition structure / block partition information of the chroma block is identical to the predicted value, and a prediction flag of 0 (False) indicates that the block partition structure / block partition information of the chroma block is not identical to the predicted value. However, it is also possible that the value of the prediction flag is set opposite to the above.
[0249] In the past, split_flag was used to determine whether a chroma block was split. On the other hand, in the present embodiment, instead of encoding / decoding split_flag as is, information indicating whether the split_flag of a chroma block matches a predicted value can be encoded / decoded.
[0250] That is, after checking whether the same-position luma block has been split, whether the luma block has been split can be set as a predicted value for the split_flag of the current block.
[0251] In the example shown in Fig. 16, since the same-position luma block is split, the prediction value for the split_flag of the chroma block can be set to 1.
[0252] Afterwards, a prediction flag indicating whether the split of the chroma block is identical to the predicted value can be encoded / decoded. Assuming that the chroma block is QT-split, the prediction flag for the split_flag of the chroma block can be set to 1 (True) and encoded / decoded.
[0253] If the splitting of the chroma block is not the same as the predicted value, the prediction flag for the split_flag of the chroma block is set to 0 (False) so that it can be encoded / decoded.
[0254] If the predicted value for syntax split_flag is 1 and the value of syntax split_flag of the current block (chroma block) is determined to be 1, the predicted flag for QT_flag of the chroma block can be encoded / decoded.
[0255] In Fig. 16, QT segmentation is applied to a luma block at the same location. Accordingly, the predicted value for the QT_flag of the chroma block can be set to 1.
[0256] Afterwards, the prediction flag for QT_flag, which indicates whether the QT splitting applied to the chroma block is the same as the predicted value, can be encoded / decoded. If it is assumed that QT splitting is applied to the chroma block, the prediction flag for QT_flag of the chroma block can be set to 1 (True) and encoded / decoded.
[0257] If the application of QT splitting to the chroma block is not the same as the predicted value, the prediction flag for QT_flag is set to 0 (False) and encoding / decoding can be performed.
[0258] Meanwhile, when a chroma block is QT-split, there is no need to encode / decode information for BT / TT splitting (e.g., DIR_flag indicating the splitting direction and BT_flag indicating whether BT / TT splitting is performed). Accordingly, when a chroma block is QT-split, the prediction flag for DIR_flag and the prediction flag for BT_flag indicating whether BT / TT splitting is performed are also not encoded / decoded.
[0259] On the other hand, if QT splitting is not applied to the chroma block, the splitting direction of the same-position luma block can be set as a predicted value for DIR_flag, and the predicted flag for DIR_flag can be encoded / decoded. In addition, whether or not the same-position luma block is split into BT / TT can be set as a predicted value for BT_flag, and the predicted flag for BT_flag can be encoded / decoded.
[0260] Assuming that the block split structure is determined by referring to the block split information in the order of split_flag, QT_flag, DIR_flag, and BT_flag, as in the example illustrated in Fig. 13, the prediction flag for the current syntax can be encoded / decoded only when the value of the prediction flag for the previous syntax is a preset specific value. Hereinafter, the explanation will be given assuming that the preset specific value is 1 (True).
[0261] If the value of the prediction flag for the previous syntax is not 1, the current syntax can be encoded / decoded instead of the prediction flag for the current syntax.
[0262] For example, if the value of the prediction flag for the syntax split_flag is 1, and the value of the syntax split_flag is determined to be 1, the prediction flag for the syntax QT_flag can be encoded / decoded instead of the syntax QT_flag. On the other hand, if the value of the prediction flag for the syntax split_flag is 0, the prediction flag for the QT_flag may not be encoded / decoded. On the other hand, even if the value of the prediction flag for the syntax split_flag is 0, if the value of the syntax split_flag is determined to be 1, QT_flag, DIR_flag and / or BT_flag must be additionally encoded / decoded to determine the block split structure of the chroma block. At this time, if the value of the prediction flag for the syntax split_flag was 0, instead of encoding / decoding the prediction flag of the syntax QT_flag, the syntax QT_flag itself can be encoded / decoded.
[0263] Similarly, if the value of the prediction flag for syntax QT_flag is 1 and the value of syntax QT_flag is determined to be 0, the prediction flag for syntax DIR_flag can be encoded / decoded instead of syntax DIR_flag. On the other hand, if the value of the prediction flag for syntax split_flag is 0 and the value of QT_flag is determined to be 0, the syntax DIR_flag itself can be encoded / decoded instead of the prediction flag for syntax DIR_flag.
[0264] Meanwhile, whether syntax BT_flag is encoded / decoded depends only on syntax QT_flag, not on syntax DIR_flag. Accordingly, regardless of how the value of syntax DIR _flag is determined, it is possible to determine whether to encode / decode the prediction flag for syntax BT_flag based on whether the prediction flag for syntax DIR_flag is 1 (True). For example, if the prediction flag for syntax DIR_flag is 1, the prediction flag for syntax BT_flag can be encoded / decoded instead of syntax BT_flag. On the other hand, if the prediction flag for syntax DIR_flag is 0, syntax BT_flag itself can be encoded / decoded instead of the prediction flag for syntax BT_flag.
[0265] Alternatively, instead of the prediction flag for DIR_flag, it may be determined whether to encode / decode the prediction flag for syntax BT_flag based on the value of the prediction flag for QT_flag and the value of QT_flag. For example, if the value of the prediction flag for syntax QT_flag is 1 and the value of syntax QT_flag is determined to be 0, the prediction flag for syntax BT_flag may be encoded / decoded instead of syntax BT_flag. On the other hand, if the value of the prediction flag for syntax split_flag is 0 and the value of QT_flag is determined to be 0, the syntax BT_flag itself may be encoded / decoded instead of the prediction flag for syntax BT_flag.
[0266] As another example, it is also possible to determine whether to encode / decode a prediction flag or directly encode / decode the block partition information based on a prediction value for the current block partition information (i.e., the current syntax). For example, if the prediction value for the current block partition information is equal to a predetermined specific value, the prediction flag may be encoded / decoded, and if the prediction value for the current block partition information is different from the predetermined specific value, the current block partition information may be directly encoded / decoded rather than the prediction flag.
[0267] As a result, for the luma block, the block division information can be encoded and signaled as is. On the other hand, for the chroma block corresponding to the luma block, instead of encoding / decoding the block division information as is, the block division information of the luma block can be set as a predicted value, and then a prediction flag indicating whether the block division information of the chroma block is identical to the predicted value can be encoded / decoded.
[0268] When encoding / decoding prediction flags, a buffer for storing accumulated probability information can be set. At this time, the buffer for the prediction flags can exist separately from the buffer for the block segmentation information. That is, pStateIdx0 and pStateIdx1 for the prediction flags can exist separately.
[0269] For example, when encoding split_flag, QT_flag, DIR_flag, and BT_flag using a general coding engine, the initial probability state is defined by initValue defined in Equations 1 and 2. In addition, the probability state can be updated every time the syntax is encoded / decoded.
[0270] Meanwhile, the prediction flag may be encoded / decoded using a probability state buffer that is different (i.e., separate) from the probability state buffers used by the block split information, split_flag, QT_flag, DIR_flag, and BT_flag. That is, the probability state buffer for the prediction flag may have an initial probability state defined by a separate initValue. In addition, the probability state of the corresponding probability state buffer may be updated only when encoding / decoding the prediction flag, and may not be updated when encoding / decoding the block split information (i.e., at least one of split_flag, QT_flag, DIR_flag, and BT_flag).
[0271] Instead of encoding / decoding the block segmentation information as is, a flag indicating whether to encode / decode the prediction flag for the block segmentation information can be encoded and signaled.
[0272] Alternatively, depending on the split depth of the current chroma block, it may be determined whether to encode / decode the prediction flow for the block split information instead of the block split information.
[0273] Meanwhile, to set the block partition information of the luma block as a predicted value, the block partition information of the luma block must be stored. To this end, block partition information of the luma component can be stored up to the leaf node.
[0274] Alternatively, for efficient buffer management, block segmentation information may be stored only up to a certain depth. For example, block segmentation information for a luma block may be stored only up to segmentation depth 2 or segmentation depth 3. Accordingly, encoding / decoding prediction flags instead of block segmentation information may be permitted only when the segmentation depth of a chroma block is below a certain segmentation depth.
[0275] Meanwhile, the split depth at which block split information is stored may be predefined in the encoder and decoder. Alternatively, information indicating the split depth at which block split information is stored may be encoded and signaled through an upper header.
[0276] Block splitting information of the current chroma block can also be derived by comparing the maximum depth of the luma block (i.e., the splitting depth of the leaf node) with the splitting depth of the current chroma block. For example, if the difference between the maximum depth of the luma block and the splitting depth of the current chroma block is greater than a threshold, the current chroma block can be determined to be split into multiple blocks. That is, if the difference is greater than the threshold, encoding / decoding of split_flag (or a prediction flag for split_flag) for the current chroma block can be omitted, and its value can be inferred to be 1.
[0277] Figure 17 is a flowchart showing a method of encoding / decoding block division information according to tree type.
[0278] If the tree type is a single tree type (S1710), block division information can be encoded / decoded only for the luma block (S1720).
[0279] On the other hand, if the tree type is a dual tree type (S1710), for a luma block, block division information may be encoded / decoded (S1730), and for a chroma block, a prediction flag may be encoded / decoded instead of block division information (S1740). Here, the prediction flag may indicate whether the block division information of the chroma block matches the prediction value, and the prediction value may be set to the block division information of the luma block.
[0280] Meanwhile, depending on the split depth of the chroma block, block split information may be directly encoded / decoded instead of the prediction flag.
[0281] Meanwhile, blocks divided according to the tree structure can be encoded / decoded by intra prediction or inter prediction.
[0282]
[0283] Below, we will describe in detail how to perform intra prediction on a block that is currently being encoded / decoded.
[0284] FIG. 18 illustrates an image encoding / decoding method performed by an image encoding / decoding device according to the present disclosure.
[0285] Referring to FIG. 18, a reference line for intra prediction of the current block can be determined (S1800).
[0286] The current block can use one or more of a plurality of pre-defined reference line candidates in the video encoding / decoding device as reference lines for intra prediction. Here, the plurality of pre-defined reference line candidates can include neighboring reference lines adjacent to the current block to be decoded and N non-neighboring reference lines that are 1 to N samples away from the boundary of the current block. N can be 1, 2, 3, or an integer greater than or equal to 1. For convenience of explanation, it is assumed hereafter that the plurality of reference line candidates available to the current block are composed of neighboring reference line candidates and three non-neighboring reference line candidates, but the present invention is not limited thereto. That is, it goes without saying that the plurality of reference line candidates available to the current block can include four or more non-neighboring reference line candidates.
[0287] An image encoding device can determine an optimal reference line candidate from among a plurality of reference line candidates and encode an index for specifying the optimal reference line candidate. An image decoding device can determine a reference line of a current block based on an index signaled through a bitstream. The index can specify any one of the plurality of reference line candidates. The reference line candidate specified by the index can be used as a reference line of the current block.
[0288] The number of indexes signaled to determine the reference line of the current block may be 1, 2, or more. For example, when the number of indexes signaled is 1, the current block can perform intra prediction using only a single reference line candidate specified by the signaled index among a plurality of reference line candidates. Alternatively, when the number of indexes signaled is 2 or more, the current block can perform intra prediction using a plurality of reference line candidates specified by a plurality of indexes among a plurality of reference line candidates.
[0289] Referring to FIG. 18, the intra prediction mode of the current block can be determined (S1810).
[0290] The intra prediction mode of the current block can be determined from among multiple intra prediction modes pre-defined in the video encoding / decoding device. The multiple pre-defined intra prediction modes will be described with reference to FIGS. 19 and 20.
[0291] FIG. 19 illustrates an example of multiple intra prediction modes according to the present disclosure.
[0292] Referring to FIG. 20, the multiple intra prediction modes pre-defined in the video encoding / decoding device may be configured as a non-directional mode and a directional mode. The non-directional mode may include at least one of a planar mode or a DC mode. The directional mode may include directional modes numbered 2 to 66.
[0293] The directional mode can be further expanded than that shown in Fig. 19. Fig. 20 shows an example in which the directional mode is expanded.
[0294] In Fig. 20, modes -1 to -14 and modes 67 to 80 are added as examples. These directional modes may be referred to as wide-angle intra prediction modes. Whether to use the wide-angle intra prediction mode may be determined depending on the shape of the current block. For example, if the current block is a non-square block whose width is greater than its height, some directional modes (e.g., 2 to 15) may be converted to wide-angle intra prediction modes between 67 and 80. On the other hand, if the current block is a non-square block whose height is greater than its width, some directional modes (e.g., 53 to 66) may be converted to wide-angle intra prediction modes between -1 and -14.
[0295] The range of available wide-angle intra prediction modes can be adaptively determined based on the width-to-height ratio of the current block. Table 1 shows the range of available wide-angle intra prediction modes based on the width-to-height ratio of the current block.
[0296] Available Wide Angle Intra Prediction Mode Ranges W / H = 1667~80 W / H = 867~78 W / H = 467~76 W / H = 267~74 W / H = 1 None W / H = 1 / 2-1~-8 W / H = 1 / 4-1~-10 W / H = 1 / 8-1~-12 W / H = 1 / 16-1~-14
[0297] Among the above multiple intra prediction modes, K candidate modes (most probable modes, MPMs) can be selected. A candidate list including the selected candidate modes can be generated. An index indicating one of the candidate modes in the candidate list can be signaled. The intra prediction mode of the current block can be determined based on the candidate mode indicated by the index. For example, the candidate mode indicated by the index can be set as the intra prediction mode of the current block. Alternatively, the intra prediction mode of the current block can be determined based on a value of the candidate mode indicated by the index and a predetermined difference value. The difference value can be defined as a difference between a value of the intra prediction mode of the current block and a value of the candidate mode indicated by the index. The difference value can be signaled through a bitstream. Alternatively, the difference value may be a pre-defined value in the video encoding / decoding device. Alternatively, the intra prediction mode of the current block may be determined based on a flag indicating whether a mode identical to the intra prediction mode of the current block exists in the candidate list. For example, when the flag has a first value, the intra prediction mode of the current block may be determined from the candidate list. In this case, an index indicating any one of a plurality of candidate modes belonging to the candidate list may be signaled. The candidate mode indicated by the index may be set as the intra prediction mode of the current block. On the other hand, when the flag has a second value, any one of the remaining intra prediction modes may be set as the intra prediction mode of the current block. The remaining intra prediction mode may mean a mode excluding a candidate mode belonging to the candidate list among the plurality of pre-defined intra prediction modes. When the flag has a second value, an index indicating any one of the remaining intra prediction modes may be signaled.The intra prediction mode indicated by the signaled index can be set as the intra prediction mode of the current block.
[0298] The intra prediction mode of a chroma block can be selected from among intra prediction mode candidates of multiple chroma blocks. To this end, index information indicating one of the intra prediction mode candidates of the chroma block can be explicitly encoded and signaled through the bitstream. Table 2 illustrates intra prediction mode candidates of the chroma block.
[0299] Intra prediction mode candidates for indexed chroma blocks Luma mode: 0 Luma mode: 50 Luma mode: 18 Luma mode: 1 Other 0 6 6 0 0 0 1 5 0 6 6 5 0 5 0 5 0 2 1 8 1 8 6 6 1 8 1 8 3 1 1 6 6 1 4 DM
[0300] In the example of Table 2, DM (Direct Mode) means setting the intra prediction mode of the luma block co-located with the chroma block to the intra prediction mode of the chroma block. Meanwhile, the luma block co-located with the chroma block can be determined based on the position of the upper left sample or the position of the center sample of the chroma block.
[0301] For example, if the intra prediction mode (luma mode) of the luma block is 0 (planar mode) and the index points to 2, the intra prediction mode of the chroma block can be determined as horizontal mode (18). For example, if the intra prediction mode (luma mode) of the luma block is 1 (DC mode) and the index points to 0, the intra prediction mode of the chroma block can be determined as planar mode (0).
[0302] Consequently, the intra prediction mode of the chroma block may also be set to one of the intra prediction modes illustrated in FIG. 19 or FIG. 20. The intra prediction mode of the current block may also be used to determine a reference line of the current block, in which case step S1810 may be performed before step S1800.
[0303] Meanwhile, in the present disclosure, a chroma block may represent at least one of a Cb component block or a Cr component block.
[0304] Referring to FIG. 18, intra prediction can be performed for the current block based on the reference line and intra prediction mode of the current block (S1820).
[0305] When the tree type is determined by a predefined block unit, only the predefined prediction mode can be applied to the chroma block. For example, when the tree type is a dual tree type, the DM (Direct Mode) mode can always be used when deriving the intra prediction mode of the chroma block. Accordingly, the encoding / decoding of the index information indicating one of the intra prediction mode candidates of the chroma block can be omitted.
[0306] That is, when the tree type is a dual tree type, the intra prediction mode of the chroma block can be derived based on the intra prediction mode of the luma block at the same location.
[0307] Figure 21 is a diagram for explaining a luma block referenced to derive an intra prediction mode of a chroma block.
[0308] Referring to FIG. 21, the area within the luma picture corresponding to the chroma block that is currently being encoded / decoded is exemplified as including luma blocks a to h.
[0309] If there are multiple luma blocks within a region corresponding to the current chroma block, an index indicating one of the multiple luma blocks can be encoded / decoded. The intra prediction mode of the luma block indicated by the index among the multiple luma blocks can be set to the intra prediction mode of the current chroma block.
[0310] Meanwhile, the index may indicate one of the luma blocks encoded / decoded using intra prediction among a plurality of luma blocks. That is, a luma block encoded / decoded using inter prediction may not be available for deriving the intra prediction mode of the current chroma block.
[0311] Alternatively, the intra prediction mode of a luma block at a preset position among multiple luma blocks may be set to the intra prediction mode of the current chroma block. Here, the preset position may be the center position or the upper left position of the current chroma block.
[0312] For example, if the preset position is the center position of the current chroma block, the intra prediction mode of the luma block g can be set to the intra prediction mode of the current chroma block.
[0313] For example, if the preset position is the upper left position of the current chroma block, the intra prediction mode of luma block a may be set to the intra prediction mode of the current chroma block.
[0314] Alternatively, the intra prediction mode of the current chroma block can be derived based on the priorities among multiple intra prediction modes for multiple luma blocks. That is, the intra prediction mode with the highest priority among the multiple intra prediction modes can be set as the intra prediction mode of the current chroma block.
[0315] Meanwhile, priorities among multiple intra prediction modes may be predefined in the encoder and decoder. For example, a non-directional intra prediction mode (e.g., a planar mode and / or a DC mode) may have a higher priority than a directional intra prediction mode. For example, among directional intra prediction modes, an intra prediction mode that can be predicted using integer position reference samples (e.g., at least one of a vertical mode, a horizontal mode, a top-left diagonal mode, a top-right diagonal mode, or a top-left diagonal mode) may have a higher priority than a directional intra prediction mode in which prediction is performed using fractional position reference samples.
[0316] Alternatively, priorities can be set based on the frequency of occurrence of intra prediction modes of multiple luma blocks. That is, the intra prediction mode with the highest frequency of occurrence can be set as the intra prediction mode of the current chroma block.
[0317] Here, the occurrence frequency of the intra prediction mode can indicate the number of blocks predicted by the intra prediction mode.
[0318] Alternatively, the frequency of occurrence of an intra prediction mode may indicate the number of samples predicted by the intra prediction mode. That is, the frequency of occurrence of an intra prediction mode used for a large luma block may have a greater value than the frequency of occurrence of an intra prediction mode used for a small luma block.
[0319] Alternatively, multiple intra predictions can be performed on the current chroma block based on multiple intra prediction modes derived from multiple luma blocks. The multiple prediction blocks can then be averaged or weighted to obtain the final prediction block of the current chroma block.
[0320] For example, among multiple intra prediction modes, N intra prediction modes with high priorities can be selected. Here, N is a natural number greater than or equal to 2, such as 2, 3, or 4.
[0321] If the tree type is a dual tree type, multiple chroma blocks may correspond to a single luma block. In this case, the multiple chroma blocks may share a single intra prediction mode.
[0322] Figure 22 is a diagram illustrating an example in which chroma blocks share one intra prediction mode.
[0323] In the example illustrated in Fig. 22, the areas corresponding to chroma blocks A to D within a luma picture belong to one luma block. Accordingly, chroma blocks A to D can share the intra prediction mode of the corresponding luma block.
[0324] Alternatively, when a dual tree type is applied, one intra prediction mode can be set to be commonly used for the split chroma blocks. That is, instead of applying a DM mode to each of the split chroma blocks, one intra prediction mode that is commonly applied to multiple chroma blocks can be derived.
[0325] Figure 23 shows an example where segmented chroma blocks share one intra prediction mode.
[0326] When the dual tree type is applied, one chroma block can be split into multiple chroma blocks through QT, TT or BT splitting.
[0327] Alternatively, as in the example illustrated in FIG. 23, when a dual tree type is applied, one chroma block can be set to be split only in the horizontal or vertical direction.
[0328] Figures 23 (a) and (b) illustrate examples in which a chroma block is divided into two or four sub-blocks through vertical division.
[0329] Additionally, (c) and (d) of FIG. 23 illustrate examples in which a chroma block is divided into two or four sub-blocks through horizontal division.
[0330] According to the example illustrated in Fig. 23, the segmented chroma blocks can share a single intra prediction mode. Accordingly, information indicating a single intra prediction mode shared by multiple chroma blocks can be encoded and signaled.
[0331] Meanwhile, a chroma block generated by vertical or horizontal division can be further divided into multiple chroma blocks. That is, a chroma block can be divided recursively.
[0332] At this time, only the division example shown in Fig. 23 can be applied to the division of chroma blocks.
[0333] Alternatively, the maximum split depth of a chroma block can be set to 1 to prevent the chroma blocks generated by the split from being split further.
[0334] Meanwhile, when the tree type is set in units of predefined blocks, the split depth of the chroma block may be newly counted from the predefined block. That is, the split depth of the predefined block may be 0, and the split depth of the chroma blocks generated by splitting the predefined block may be 1.
[0335]
[0336] Applying the embodiments described above, focusing on the decoding or encoding process, to the encoding or decoding process is within the scope of the present disclosure. Changing the embodiments described above, in a given order, to a different order is also within the scope of the present disclosure.
[0337] Although the above-described disclosure is described based on a series of steps or a flowchart, this does not limit the chronological order of the invention, and may be performed simultaneously or in a different order as needed. In addition, each component (e.g., unit, module, etc.) constituting the block diagram in the above-described disclosure may be implemented as a hardware device or software, or multiple components may be combined to be implemented as a single hardware device or software. For example, the hardware device may include at least one of a processor for performing calculations, a memory for storing data, a transmitter for transmitting data, and a receiver for receiving data.
[0338] The above-described disclosure may be implemented in the form of program commands that can be executed by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, etc., either singly or in combination.
[0339] In addition, according to the present disclosure, a computer-readable recording medium can be provided that stores a bitstream generated by the above-described encoding method. The bitstream can be transmitted by an encoding device, and a decoding device can receive the bitstream and decode an image.
[0340] Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memories. The hardware devices may be configured to operate as one or more software modules to perform processing according to the present disclosure, and vice versa.
[0341] Embodiments according to the present disclosure can be applied to electronic devices that support encoding / decoding of images.
Claims
1. Step of determining the tree type; and A step of determining whether to split the chroma block based on first block split information of the chroma block, If the tree type indicates a dual tree type, the first block split information of the chroma block is determined based on a first prediction flag decoded from the bitstream, The first prediction flag indicates whether the first block division information of the chroma block matches the first prediction value, A video decoding method, characterized in that the first prediction value is set to a value of first block division information of a luma block corresponding to the chroma block.
2. In paragraph 1, If the first block division information of the luma block indicates that the luma block is divided, and the first prediction flag indicates that the first block division information of the chroma block matches the first prediction value, Based on the second block division information of the chroma block, it is determined whether to apply QT (Quad Tree)-based division to the chroma block, The second block division information of the chroma block is determined based on a second prediction flag indicating whether the second block division information of the chroma block matches the second prediction value, An image decoding method, characterized in that the second prediction value is set to a value of the second segmentation information of the luma block.
3. In paragraph 1, If the first block division information of the luma block indicates that the luma block is not divided, and the first prediction flag indicates that the first block division information of the chroma block does not match the first prediction value, A video decoding method, characterized in that second block segmentation information indicating whether to apply QT-based segmentation to the chroma block is decoded from the bitstream.
4. In paragraph 1, A video decoding method, characterized in that whether the first prediction flag is decoded is determined based on a result of comparing the split depth of the chroma block with a threshold value.
5. In paragraph 4, A video decoding method, characterized in that the first prediction flag is decoded only when the division depth of the chroma block is less than the threshold.
6. In paragraph 5, A video decoding method, characterized in that when the division depth of the chroma block is not less than the threshold, the first block division information of the chroma block is directly decoded from the bitstream.
7. In paragraph 1, A video decoding method, characterized in that when the difference between the maximum segmentation depth of the luma component and the segmentation depth of the chroma block is greater than a threshold, decoding of the first prediction flag is omitted, and the first block segmentation information of the chroma block is inferred to indicate that the chroma block is segmented.
8. In paragraph 1, A video decoding method, characterized in that the above tree type is determined in units of predefined blocks, and the predefined blocks have a size smaller than a coding tree block.
9. In paragraph 8, A video decoding method characterized in that a single tree type is fixedly applied to a coding block having a size larger than the above-defined block.
10. In paragraph 8, A method for decoding an image, wherein the above-definition block is a coding block having a segmentation depth having a pre-defined value.
11. In paragraph 1, If the above tree type is the above dual tree type, A video decoding method characterized in that DM (Direct Mode) is fixedly applied when deriving the intra prediction mode of the above chroma block.
12. In paragraph 11, A video decoding method, characterized in that, when there are a plurality of luma blocks corresponding to the chroma block, the intra prediction mode of the chroma block is derived by referring to a luma block including a sample at a predefined position among the luma blocks.
13. In paragraph 11, A video decoding method, characterized in that, when there are multiple luma blocks corresponding to the chroma block, the intra prediction mode with the highest occurrence frequency for the luma blocks is derived as the intra prediction mode of the chroma block.
14. A step of encoding information indicating a tree type; and A step of determining the first block division information of the chroma block, which indicates whether to divide the chroma block, If the tree type is a dual tree type, the first prediction flag is encoded in the bitstream instead of the first block division information of the chroma block, The first prediction flag indicates whether the first block division information of the chroma block matches the first prediction value, A video encoding method, characterized in that the first prediction value is set to a value of first block division information of a luma block corresponding to the chroma block.
15. A processor that generates compressed video data; and In a device including a transmitter for transmitting the compressed video data, The process of generating the above compressed video is as follows: A step of encoding information indicating a tree type; and A step of determining the first block division information of the chroma block, which indicates whether to divide the chroma block, If the tree type is a dual tree type, the first prediction flag is encoded in the bitstream instead of the first block division information of the chroma block, The first prediction flag indicates whether the first block division information of the chroma block matches the first prediction value, A device for transmitting compressed video data, characterized in that the first prediction value is set to a value of first block division information of a luma block corresponding to the chroma block.
Citation Information
Patent Citations
Image encoding / decoding method and device using division restriction for chroma blocks, and bitstream transmission method
JP2022524441A
Method and apparatus for intra sub-partition coding mode
JP2023179591A
Composition for preventing or treating liver cancer comprising Artemisia iwayomogi Kitamura extract
KR1020210007460A
Multi-fuction holder
KR1020230036276A
Apparatus and method for video conferencing service
KR102575038B1