Intra prediction-based video encoding / decoding method and device
By determining a reference region and dividing intra prediction modes into candidate groups with offsets, the method addresses inefficiencies in block division and intra-prediction for high-resolution images, enhancing encoding/decoding efficiency and accuracy.
Patent Information
- Application Number
- JP2025100287
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-12-29
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-09
AI Technical Summary
Existing image compression techniques lack efficient methods for block division and intra-prediction mode derivation, particularly for high-resolution and high-quality images, leading to suboptimal encoding and decoding efficiency.
The method involves determining a reference region for intra prediction, dividing intra prediction modes into MPM candidate groups, and applying offsets based on block size, shape, and component type to derive accurate intra prediction modes, with chrominance blocks having separate reference-based prediction modes.
This approach enhances the efficiency and accuracy of intra-prediction encoding/decoding by adaptive block division and inter-component reference, improving prediction accuracy and efficiency.
Smart Images

Figure 2025131838000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image encoding / decoding method and apparatus. [Background technology]
[0002] Recently, the demand for high-resolution and high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has increased in various application fields, and as a result, highly efficient image compression techniques have been discussed.
[0003] There are various image compression techniques, such as inter-prediction techniques that predict pixel values contained in a current picture from pictures before or after the current picture, intra-prediction techniques that predict pixel values contained in a current picture using pixel information within the current picture, and entropy coding techniques that assign short codes to values that occur frequently and long codes to values that occur less frequently. These image compression techniques can be used to effectively compress image data for transmission or storage. Summary of the Invention [Problem to be solved by the invention]
[0004] An object of the present invention is to provide an efficient block division method and apparatus.
[0005] SUMMARY OF THE INVENTION An object of the present invention is to provide a method and apparatus for deriving an intra-prediction mode.
[0006] SUMMARY OF THE INVENTION An object of the present invention is to provide a method and apparatus for determining a reference region for intra prediction.
[0007] The present invention aims to provide a method and apparatus for intra prediction based on component type. [Means for solving the problem]
[0008] The image encoding / decoding method and apparatus of the present invention can determine a reference region for intra prediction of a current block, derive an intra prediction mode of the current block, and decode the current block based on the reference region and the intra prediction mode.
[0009] In the image encoding / decoding method and apparatus of the present invention, the intra prediction modes already defined in the encoding / decoding apparatus are divided into an MPM candidate group and a non-MPM candidate group, and the MPM candidate group may include at least one of a first candidate group or a second candidate group.
[0010] In the image encoding / decoding method and apparatus of the present invention, the intra prediction mode of the current block can be derived from either the first candidate group or the second candidate group.
[0011] In the image encoding / decoding method and apparatus of the present invention, the first candidate group may be configured with a default mode already defined in the decoding apparatus, and the second candidate group may be configured with a plurality of MPM candidates.
[0012] In the image encoding / decoding method and apparatus of the present invention, the default mode may be at least one of a planar mode, a DC mode, a vertical mode, a horizontal mode, a vertical mode, or a diagonal mode.
[0013] In the image encoding / decoding method and apparatus of the present invention, the multiple MPM candidates include at least one of the intra prediction modes of the surrounding blocks, a mode obtained by subtracting an n value from the intra prediction mode of the surrounding blocks, or a mode obtained by adding an n value to the intra prediction mode of the surrounding blocks, where n may represent a natural number of 1, 2 or more.
[0014] In the image encoding / decoding method and apparatus of the present invention, the multiple MPM candidates include at least one of a DC mode, a vertical mode, a horizontal mode, a mode obtained by subtracting or adding an m value to the vertical mode, or a mode obtained by subtracting or adding an m value to the horizontal mode, where m can be a natural number of 1, 2, 3, 4 or more.
[0015] In the image encoding / decoding method and apparatus of the present invention, the encoding device determines a candidate group to which the intra prediction mode of the current block belongs and encodes a flag identifying the candidate group, and the decoding device can select either the first candidate group or the second candidate group based on the flag signaled by the encoding device.
[0016] In the image encoding / decoding method and apparatus of the present invention, the induced intra prediction mode may be changed by applying a predetermined offset to the induced intra prediction mode.
[0017] In the image encoding / decoding method and apparatus of the present invention, the offset may be selectively applied based on at least one of the size, shape, partition information, intra prediction mode value, or component type of the current block.
[0018] In the image encoding / decoding method and apparatus of the present invention, the step of determining the reference area may include a step of searching for unavailable pixels belonging to the reference area, and a step of replacing the unavailable pixels with available pixels.
[0019] In the image encoding / decoding method and apparatus of the present invention, the available pixels may be determined based on a bit depth value, or may be pixels adjacent to at least one of the left, right, top, or bottom ends of the unavailable pixels. [Effects of the Invention]
[0020] The present invention can improve the efficiency of intra-prediction encoding / decoding by adaptive block division.
[0021] According to the present invention, prediction can be performed more accurately and efficiently by deriving an intra prediction mode based on a group of MPM candidates.
[0022] According to the present invention, for chrominance blocks, the inter-component reference-based prediction modes are defined as a separate group, thereby enabling more efficient intra prediction mode derivation for the chrominance blocks.
[0023] According to the present invention, the accuracy and efficiency of intra prediction can be improved by substituting predetermined available pixels for unavailable pixels in a reference region for intra prediction.
[0024] According to the present invention, it is possible to improve the efficiency of inter-frame prediction based on inter-component reference. [Brief explanation of the drawings]
[0025] [Figure 1] 1 is a block diagram illustrating an image encoding device according to an embodiment of the present invention. [Figure 2] 1 is a block diagram showing an image decoding apparatus according to an embodiment of the present invention; [Figure 3] FIG. 1 is a diagram showing a method for dividing a picture into a plurality of fragment regions as an embodiment to which the present invention is applied. [Figure 4] 1 is an exemplary diagram illustrating intra-prediction modes predefined in an image encoding / decoding device according to an embodiment of the present invention. [Figure 5] 1 is a diagram illustrating a method for decoding a current block based on intra prediction, according to an embodiment of the present invention; [Figure 6] FIG. 10 is a diagram illustrating a method for substituting unavailable pixels in a reference area as an embodiment to which the present invention is applied. [Figure 7]FIG. 1 is a diagram illustrating a method for changing / correcting an intra prediction mode as an embodiment to which the present invention is applied. [Figure 8] FIG. 1 is a diagram illustrating a prediction method based on inter-component reference as an embodiment to which the present invention is applied. [Figure 9] FIG. 10 is a diagram showing a method for configuring a reference region as an embodiment to which the present invention is applied. [Figure 10] 1 is a diagram illustrating an example of configuring an intra prediction mode set for each step according to an embodiment of the present invention. [Figure 11] FIG. 1 is a diagram illustrating a method for classifying intra prediction modes into a plurality of candidate groups, as an embodiment to which the present invention is applied. [Figure 12] 1 is an exemplary diagram showing a current block and its neighboring pixels according to an embodiment of the present invention; [Figure 13] FIG. 1 is a diagram illustrating a method for performing intra prediction in steps as an embodiment to which the present invention is applied. [Figure 14] 1 is an exemplary diagram of an arbitrary pixel for intra prediction as an embodiment to which the present invention is applied. [Figure 15] 1 is an exemplary diagram showing an embodiment to which the present invention is applied, in which an image is divided into a plurality of sub-regions based on an arbitrary pixel; DETAILED DESCRIPTION OF THE INVENTION
[0026] The image encoding / decoding method and apparatus of the present invention can determine a reference region for intra prediction of a current block, derive an intra prediction mode of the current block, and decode the current block based on the reference region and the intra prediction mode.
[0027] In the image encoding / decoding method and apparatus of the present invention, the intra prediction modes already defined in the encoding / decoding apparatus are divided into an MPM candidate group and a non-MPM candidate group, and the MPM candidate group may include at least one of a first candidate group or a second candidate group.
[0028] In the image encoding / decoding method and apparatus of the present invention, the intra prediction mode of the current block can be derived from either the first candidate group or the second candidate group.
[0029] In the image encoding / decoding method and apparatus of the present invention, the first candidate group may be configured with a default mode already defined in the decoding apparatus, and the second candidate group may be configured with a plurality of MPM candidates.
[0030] In the image encoding / decoding method and apparatus of the present invention, the default mode may be at least one of a planar mode, a DC mode, a vertical mode, a horizontal mode, a vertical mode, or a diagonal mode.
[0031] In the image encoding / decoding method and apparatus of the present invention, the multiple MPM candidates include at least one of the intra prediction modes of the surrounding blocks, a mode obtained by subtracting an n value from the intra prediction mode of the surrounding blocks, or a mode obtained by adding an n value to the intra prediction mode of the surrounding blocks, where n may represent a natural number of 1, 2 or more.
[0032] In the image encoding / decoding method and apparatus of the present invention, the multiple MPM candidates include at least one of a DC mode, a vertical mode, a horizontal mode, a mode obtained by subtracting or adding an m value to the vertical mode, or a mode obtained by subtracting or adding an m value to the horizontal mode, where m can be a natural number of 1, 2, 3, 4 or more.
[0033] In the image encoding / decoding method and apparatus of the present invention, the encoding device determines a candidate group to which the intra prediction mode of the current block belongs and encodes a flag identifying the candidate group, and the decoding device can select either the first candidate group or the second candidate group based on the flag signaled by the encoding device.
[0034] In the image encoding / decoding method and apparatus of the present invention, the induced intra prediction mode may be changed by applying a predetermined offset to the induced intra prediction mode.
[0035] In the image encoding / decoding method and apparatus of the present invention, the offset may be selectively applied based on at least one of the size, shape, partition information, intra prediction mode value, or component type of the current block.
[0036] In the image encoding / decoding method and apparatus of the present invention, the step of determining the reference area may include the steps of searching for unavailable pixels belonging to the reference area and replacing the unavailable pixels with available pixels.
[0037] In the image encoding / decoding method and apparatus of the present invention, the available pixels may be determined based on a bit depth value, or may be pixels adjacent to at least one of the left, right, top, or bottom ends of the unavailable pixels. [Mode for carrying out the invention]
[0038] Since the present invention can be modified in various ways and can have various embodiments, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, it should be understood that the present invention is not limited to the specific embodiments, but includes all modifications, equivalents, and alternatives within the spirit and technical scope of the present invention. In describing each drawing, like reference numerals are used to refer to like components.
[0039] The terms "first," "second," etc. may be used to describe various elements, but these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element can be termed a second element, and similarly, a second element can be termed a first element, without departing from the scope of the present invention. The term "and / or" includes a combination of two or more related listed items or any of two or more related listed items.
[0040] When a component is said to be "coupled" or "connected" to another component, it should be understood that the component may be directly coupled or connected to the other component, but there may be other components between them. Conversely, when a component is said to be "directly coupled" or "directly connected" to another component, it should be understood that there are no other components between them.
[0041] The terms used in the present invention are merely used to describe specific embodiments and do not limit the present invention. A singular expression includes a plural expression unless the context clearly indicates otherwise. In the present invention, the terms "comprise" or "have" and the like specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the presence or possibility of addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0042] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. The same reference numerals are used to designate the same components in the drawings, and duplicated descriptions of the same components will be omitted.
[0043] FIG. 1 is a block diagram showing an image encoding device according to an embodiment of the present invention.
[0044] Referring to FIG. 1, the image encoding device 100 may include a picture division unit 110, a prediction unit 120, 125, a transformation unit 130, a quantization unit 135, a realignment unit 160, an entropy encoding unit 165, an inverse quantization unit 140, an inverse transform unit 145, a filter 150, and a memory 155.
[0045] 1 are illustrated independently to illustrate different characteristic functions of the image encoding device, and do not mean that each component is composed of separate hardware or a single software component. That is, each component is included as a separate component for the sake of convenience of explanation, and at least two of the components may be combined to form a single component, or one component may be divided into multiple components to perform its function. Such integrated and separated embodiments of each component are also within the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0046] Furthermore, some components may not be essential components that perform essential functions in the present invention, but may be optional components simply for improving performance. The present invention can be realized by including only components that are essential for realizing the essence of the present invention, excluding components used simply for improving performance, and a structure including only essential components excluding optional components used simply for improving performance is also included in the scope of the present invention.
[0047] The picture division unit 110 can divide an input picture into at least one processing unit. At this time, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The picture division unit 110 can divide one picture into a plurality of combinations of coding units, prediction units, and transform units, and select one combination of coding units, prediction units, and transform units based on a predetermined criterion (e.g., a cost function) to code the picture.
[0048] For example, a picture can be divided into multiple coding units. A recursive tree structure, such as a quad tree structure, can be used to divide a picture into coding units. A coding unit that is divided into other coding units using an image, the largest coding unit, or a coding tree unit (CTU) as the root can be divided into child nodes equal to the number of divided coding units. A coding unit that is not further divided within a certain limit becomes a leaf node. In other words, assuming that only square division is possible for a coding unit, one coding unit can be divided into a maximum of four different coding units.
[0049] Hereinafter, in the embodiments of the present invention, the coding unit may be used to mean a unit for performing coding, or may be used to mean a unit for performing decoding.
[0050] The prediction units may be divided into at least one shape, such as a square or rectangle, of the same size within one coding unit, or may be divided so that one of the prediction units divided within one coding unit has a different shape and / or size from another prediction unit.
[0051] When generating a prediction unit for performing intra prediction based on a coding unit, if the coding unit is not the smallest coding unit, intra prediction can be performed without dividing the coding unit into a plurality of N×N prediction units.
[0052] The prediction units 120 and 125 may include an inter prediction unit 120 for performing inter prediction and an intra prediction unit 125 for performing intra prediction. It is possible to determine whether to use inter prediction or intra prediction for a prediction unit, and to determine specific information (e.g., intra prediction mode, motion vector, reference picture, etc.) according to each prediction method. Here, the processing unit in which prediction is performed may differ from the processing unit in which the prediction method and its specific contents are determined. For example, the prediction method and prediction mode may be determined in a prediction unit, and the prediction may be performed in a transform unit. Residual values (residual blocks) between the generated prediction block and the original block may be input to the transform unit 130. In addition, prediction mode information, motion vector information, etc. used for prediction may be coded by the entropy coding unit 165 together with the residual values and transmitted to the decoder. When a specific coding mode is used, it is also possible to directly code the original block and transmit it to the decoder without generating a prediction block via the prediction units 120 and 125.
[0053] The inter prediction unit 120 may predict a prediction unit based on information of at least one of a picture preceding or following the current picture, and in some cases, may predict a prediction unit based on information of a partial region in the current picture for which encoding has been completed. The inter prediction unit 120 may include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.
[0054] The reference picture interpolation unit receives reference picture information from memory 155 and can generate sub-integer pixel information from the reference picture. For luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate sub-integer pixel information in 1 / 4 pixel units. For color difference signals, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate sub-integer pixel information in 1 / 8 pixel units.
[0055] The motion prediction unit may perform motion prediction based on the reference picture interpolated by the reference picture interpolation unit. Various methods, such as a full search-based block matching algorithm (FBMA), a three-step search algorithm (TSS), and a new three-step search algorithm (NTS), may be used to calculate a motion vector. The motion vector may have a motion vector value in half or quarter pixel units based on the interpolated pixels. The motion prediction unit may predict the current prediction unit using different motion prediction methods. Various methods, such as a skip method, a merge method, an advanced motion vector prediction method (AMVP), and an intra block copy method, may be used as the motion prediction method.
[0056] The intra prediction unit 125 may generate a prediction unit based on reference pixel information surrounding a current block, which is pixel information within a current picture. When a neighboring block of the current prediction unit is an inter-predicted block and the reference pixel is an inter-predicted pixel, the reference pixel included in the inter-predicted block may be replaced with reference pixel information of a neighboring intra-predicted block. In other words, when reference pixels are unavailable, the unavailable reference pixel information may be replaced with at least one of the available reference pixels.
[0057] Prediction modes in intra prediction may include a directional prediction mode that uses reference pixel information according to a prediction direction, and a non-directional mode that does not use directional information when performing prediction. A mode for predicting luma information and a mode for predicting chroma information may be different from each other, and intra prediction mode information used for predicting luma information or predicted luma signal information may be used to predict chroma information.
[0058] When performing intra prediction, if the size of the prediction unit and the size of the transform unit are the same, intra prediction for the prediction unit can be performed based on the pixel located to the left of the prediction unit, the pixel located at the upper left corner, and the pixel located at the upper corner. However, when performing intra prediction, if the size of the prediction unit and the size of the transform unit are different, intra prediction can be performed using reference pixels based on the transform unit. In addition, intra prediction using NxN division can be used only for the minimum coding unit.
[0059] The intra prediction method may generate a predicted block after applying an adaptive intra smoothing (AIS) filter to reference pixels according to a prediction mode. The types of AIS filters applied to the reference pixels may differ. To perform the intra prediction method, the intra prediction mode of a current prediction unit may be predicted from the intra prediction mode of a prediction unit existing in the vicinity of the current prediction unit. When predicting the prediction mode of the current prediction unit using mode information predicted from the surrounding prediction units, if the intra prediction modes of the current prediction unit and the surrounding prediction units are the same, information indicating that the prediction modes of the current prediction unit and the surrounding prediction units are the same may be transmitted using predetermined flag information. If the prediction modes of the current prediction unit and the surrounding prediction units are different from each other, entropy coding may be performed to encode the prediction mode information of the current block.
[0060] In addition, a residual block including residual value information, which is a difference value between a prediction block predicted based on the prediction unit generated by the prediction units 120 and 125 and an original block of the prediction unit, can be generated. The generated residual block can be input to the conversion unit 130.
[0061] The transform unit 130 may transform a residual block including residual value information of the original block and the prediction unit generated through the prediction units 120 and 125 using a transform method such as a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or a KLT. Whether to apply the DCT, the DST, or the KLT to transform the residual block may be determined based on intra prediction mode information of the prediction unit used to generate the residual block.
[0062] The quantization unit 135 quantizes the values transformed into the frequency domain by the transformation unit 130. The quantization coefficients may vary depending on the block or the importance of the image. The values calculated by the quantization unit 135 may be provided to the inverse quantization unit 140 and the reordering unit 160.
[0063] The reordering unit 160 may reorder coefficient values for the quantized residual values.
[0064] The reordering unit 160 may convert two-dimensional block configuration coefficients into one-dimensional vector form using a coefficient scanning method. For example, the reordering unit 160 may convert two-dimensional block configuration coefficients into one-dimensional vector form by scanning from DC coefficients to high-frequency region coefficients using a zig-zag scan method. Depending on the size of the transform unit and the intra prediction mode, vertical scanning, which scans two-dimensional block configuration coefficients in a column direction, or horizontal scanning, which scans two-dimensional block configuration coefficients in a row direction, may be used instead of zig-zag scanning. That is, depending on the size of the transform unit and the intra prediction mode, it may be determined which scanning method to use among zig-zag scanning, vertical scanning, and horizontal scanning.
[0065] The entropy coding unit 165 may perform entropy coding based on the value calculated by the reordering unit 160. The entropy coding may use various coding methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).
[0066] The entropy coding unit 165 can encode various information such as residual value coefficient information and block type information of the coding unit from the realignment unit 160 and the prediction units 120 and 125, prediction mode information, division unit information, prediction unit information and transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information.
[0067] The entropy coding unit 165 can entropy code the coefficient values of the coding unit input from the reordering unit 160 .
[0068] The inverse quantization unit 140 and the inverse transform unit 145 inversely quantize the values quantized by the quantization unit 135 and inversely transform the values transformed by the transform unit 130. The residual values generated by the inverse quantization unit 140 and the inverse transform unit 145 can be combined with prediction units predicted through the motion estimation unit, motion compensation unit, and intra prediction unit included in the prediction units 120 and 125 to generate reconstructed blocks.
[0069] The filter unit 150 may include at least one of a deblocking filter, an offset correction unit, and an adaptive loop filter (ALF).
[0070] A deblocking filter can remove block artifacts caused by boundaries between blocks in a reconstructed picture. To determine whether to perform deblocking, it can be determined whether to apply a deblocking filter to a current block based on pixels included in several columns or rows included in the block. When applying a deblocking filter to a block, a strong filter or a weak filter can be applied depending on the required deblocking filtering strength. In addition, when vertical filtering and horizontal filtering are performed when applying a deblocking filter, horizontal filtering and vertical filtering can be processed in parallel.
[0071] The offset correction unit can correct the offset between the deblocked image and the original image on a pixel-by-pixel basis. To perform offset correction on a specific picture, the offset correction unit can divide the pixels included in the image into a certain number of regions, determine the regions to be offset, and apply the offset to the corresponding regions, or apply the offset by taking into account edge information of each pixel.
[0072] Adaptive Loop Filtering (ALF) can be performed based on the comparison between a filtered restored image and the original image. After dividing the pixels in an image into predetermined groups, a filter to be applied to each group is determined, and differential filtering can be performed for each group. Information related to whether to apply ALF can be transmitted for each coding unit of the luminance signal, and the shape and filter coefficients of the ALF filter applied to each block can vary. Alternatively, the same type (fixed type) of ALF filter can be applied regardless of the characteristics of the target block.
[0073] The memory 155 can store the reconstructed blocks or pictures calculated through the filter unit 150, and the stored reconstructed blocks or pictures can be provided to the prediction units 120 and 125 when performing inter prediction.
[0074] FIG. 2 is a block diagram showing an image decoding apparatus according to an embodiment of the present invention.
[0075] Referring to FIG. 2, the image decoder 200 may include an entropy decoding unit 210, a reordering unit 215, an inverse quantization unit 220, an inverse transform unit 225, prediction units 230 and 235, a filter unit 240, and a memory 245.
[0076] When an image bitstream is input from an image encoder, the input bitstream can be decoded in the reverse order of the image encoder.
[0077] The entropy decoding unit 210 can perform entropy decoding in the reverse order of the entropy encoding performed by the entropy encoding unit of the image encoder. For example, various methods such as Exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding) can be applied in accordance with the method used in the image encoder.
[0078] The entropy decoding unit 210 can decode information related to the intra-prediction and inter-prediction performed in the encoder.
[0079] The reordering unit 215 may reorder the bitstream entropy decoded by the entropy decoding unit 210 based on the reordering method used by the encoder. The reordering unit 215 may reconstruct coefficients expressed in a one-dimensional vector format into coefficients in a two-dimensional block format to perform the reordering. The reordering unit 215 may receive information related to coefficient scanning performed by the encoder and perform the reordering by scanning in the reverse order based on the scanning order performed by the encoder.
[0080] The inverse quantization unit 220 may perform inverse quantization based on the quantization parameter provided by the encoder and the coefficient values of the reordered blocks.
[0081] The inverse transform unit 225 may perform inverse transforms, i.e., inverse DCT, inverse DST, and inverse KLT, on the transforms, i.e., DCT, DST, and KLT, performed by the transform unit on the quantization result performed by the image encoder. The inverse transform may be performed based on a transmission unit determined by the image encoder. The inverse transform unit 225 of the image decoder may selectively perform a transform technique (e.g., DCT, DST, or KLT) depending on a plurality of pieces of information, such as a prediction method, a size of a current block, and a prediction direction.
[0082] The prediction units 230, 235 can generate a prediction block based on information related to prediction block generation provided from the entropy decoding unit 210 and previously decoded block or picture information provided from the memory 245.
[0083] As described above, when performing intra prediction, similar to the operation of an image encoder, if the size of the prediction unit and the size of the transform unit are the same, intra prediction for the prediction unit is performed based on the pixel located to the left of the prediction unit, the pixel located at the upper left corner, and the pixel located at the upper corner. However, when performing intra prediction, if the size of the prediction unit and the size of the transform unit are different from each other, intra prediction can be performed using reference pixels based on the transform unit. Also, intra prediction using NxN division can be used only for the minimum coding unit.
[0084] The prediction units 230 and 235 may include a prediction unit determination unit, an inter prediction unit, and an intra prediction unit. The prediction unit determination unit receives various information, such as prediction unit information input from the entropy decoding unit 210, prediction mode information for the intra prediction method, and motion prediction-related information for the inter prediction method, to classify prediction units in a current coding unit and determine whether the prediction unit performs inter prediction or intra prediction. The inter prediction unit 230 may perform inter prediction on the current prediction unit based on information included in at least one picture, either a previous picture or a subsequent picture of the current picture including the current prediction unit, using information necessary for inter prediction of the current prediction unit provided from the image encoding device. Alternatively, the inter prediction unit may perform inter prediction based on information of a partial region already restored within the current picture including the current prediction unit.
[0085] In order to perform inter prediction, it is possible to determine, based on the coding unit, whether the motion prediction method of the prediction unit included in the corresponding coding unit is skip mode, merge mode, AMVP mode, or intra block copy mode.
[0086] The intra prediction unit 235 may generate a prediction block based on pixel information within the current picture. If the prediction unit is a prediction unit for which intra prediction has been performed, the intra prediction may be performed based on intra prediction mode information of the prediction unit provided from the image encoder. The intra prediction unit 235 may include an adaptive intra smoothing (AIS) filter, a reference pixel interpolator, and a DC filter. The AIS filter is a part that performs filtering on reference pixels of the current block, and may determine whether to apply a filter depending on the prediction mode of the current prediction unit and apply the filter. The AIS filtering may be performed on reference pixels of the current block using the prediction mode of the prediction unit and AIS filter information provided from the image encoder. If the prediction mode of the current block is a mode in which AIS filtering is not performed, the AIS filter may not be applied.
[0087] The reference pixel interpolation unit may generate reference pixels in units of pixels less than an integer value by interpolating reference pixels when the prediction mode of the prediction unit is a prediction unit that performs intra prediction based on pixel values obtained by interpolating reference pixels. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating reference pixels, the reference pixels may not be interpolated. When the prediction mode of the current block is a DC mode, the DC filter may generate a prediction block through filtering.
[0088] The reconstructed block or picture may be provided to a filter unit 240. The filter unit 240 may include a deblocking filter, an offset correction unit, and an ALF.
[0089] The image encoder can provide information on whether a deblocking filter is applied to the block or picture, and information on whether a strong filter or a weak filter is applied if a deblocking filter is applied. The deblocking filter of the image decoder can receive deblocking filter-related information provided by the image encoder and perform deblocking filtering on the block.
[0090] The offset correction unit can perform offset correction on the restored image based on information such as the type and offset value of offset correction applied to the image during encoding.
[0091] The ALF can be applied to a coding unit based on ALF application information, ALF coefficient information, etc. provided from the encoding device. Such ALF information may be provided by being included in a specific parameter set.
[0092] The memory 245 can store the reconstructed pictures or blocks for use as reference pictures or blocks, and can provide the reconstructed pictures to an output.
[0093] In this specification, the terms coding unit, coding block, current block, etc. may be interpreted as having the same meaning. The embodiments described below may be implemented in corresponding units of the image coding device and / or the image decoding device.
[0094] FIG. 3 is a diagram showing a method for dividing a picture into a plurality of fragment areas as an embodiment to which the present invention is applied.
[0095] A picture can be divided into predetermined fragment regions, which can include at least one of a subpicture, a slice, a tile, a coding tree unit (CTU), and a coding unit (CU).
[0096] Referring to Figure 3, a picture 300 can include one or more sub-pictures, i.e., a picture can consist of one sub-picture or can be divided into multiple sub-pictures as shown in Figure 3.
[0097] Subpictures can be configured to have different division information depending on the encoding settings. (1) For example, subpictures can be obtained by a collective division method based on vertical or horizontal lines that cross the picture. (2) Alternatively, subpictures can be obtained by a partial division method based on the characteristic information of each subpicture (such as its position, size, and shape; the shape is assumed to be rectangular, as described below).
[0098] (1) In the former case, sub-picture division information can be configured based on vertical or horizontal lines that divide the sub-picture.
[0099] The line-based division can be performed using either an equal or non-uniform division method. When the equal division method is used, information regarding the number of divisions of each line can be generated, and when the non-uniform division method is used, information regarding the spacing (width or height) between each line can be generated. Depending on the encoding setting, either the equal or non-uniform division method can be used, and information regarding the method selection can be explicitly generated. The equal or non-uniform division method can be applied to both vertical and horizontal lines. Alternatively, different methods can be applied to vertical and horizontal lines. Based on the division information, information regarding the number of sub-pictures can be derived.
[0100] The inter-line spacing information may be coded in units of n samples, CTU size, (2*CTU size), (4*CTU size), etc., where n may be a natural number of 4, 8, 16, 32, 64, 128, 256, or more. The generated information may be signaled at at least one level of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), and a picture header (PH).
[0101] (2) In the latter case, subpicture division information can be constructed from subpicture position information (e.g., information indicating the top left, top right, bottom left, and bottom right positions of each subpicture), size information (e.g., information indicating width or height), and information on the number of subpictures.
[0102] Information specifying the number of sub-pictures (hereinafter referred to as "number information") is coded by a coding device, and a decoding device can determine the number of sub-pictures constituting one picture based on the coded number information. The number information can be signaled at at least one level among VPS, SPS, PPS, and PH. Alternatively, the number information of sub-pictures can be implicitly derived based on sub-picture division information (such as position and size information).
[0103] Information specifying the position of each subpicture (hereinafter referred to as position information) may include the x-coordinate or y-coordinate of a previously agreed-upon position of the subpicture. The previously agreed-upon position may be determined from among the upper left corner, upper right corner, lower left corner, and lower right corner of the subpicture. The position information is encoded by an encoding device, and a decoding device can determine the position of each subpicture based on the encoded position information. Here, the x-coordinate / y-coordinate may be expressed in units of n samples, CTU size, (2*CTUSize), (4*CTUSize), etc., where n may be 1, 2, 4, 8, 16, 32, 64, 128, 256, or a natural number greater than or equal to 1. For example, if the position information is encoded using the x-coordinate and y-coordinate of the upper left CTU of a subpicture, and the width and height in units of CTU (CtbSize) are 2 and 3, respectively, the position (upper left corner) of the subpicture may be determined as (2*CtbSize, 3*CtbSize).
[0104] Information specifying the size of each subpicture (hereinafter, referred to as size information) may include at least one of width information and height information of the subpicture. Here, the width / height information may be coded in units of n samples, CTU size, (2*CTU size), (4*CTU size), etc. Here, n may be 4, 8, 16, 32, 64, 128, 256, or a natural number greater than or equal to 4. For example, if the width information is coded in units of CTU size (CtbSize) and the width information is 6, the width of the subpicture may be determined to be (6*CtbSize).
[0105] The position information and size information may be restricted to be encoded / decoded only when the number of sub-pictures belonging to a picture is two or more. That is, if the number of sub-pictures based on the number information is greater than or equal to two, the position information and size information are signaled; otherwise, the sub-pictures may be set to the same position / size as the picture. However, even when the number of sub-pictures is two or more, the position information for the first sub-picture located at the upper left corner of the picture may not be signaled, and position information for the second sub-picture may be signaled instead. Also, at least one of the position information and size information for the last sub-picture of the picture may not be signaled.
[0106] 3, one subpicture can include one or more slices. That is, one subpicture can be composed of one slice or can be divided into multiple slices. A subpicture can be composed of multiple slices divided horizontally or multiple slices divided vertically.
[0107] Information specifying the number of slices belonging to one picture or subpicture (hereinafter referred to as "number information") is coded by a coding device, and a decoding device can determine the number of subpictures constituting one picture or subpicture based on the coded number information. The number information can be signaled at at least one level among VPS, SPS, PPS, and PH. However, the number information can be signaled only when at least one of the following conditions is satisfied: rectangular slices are allowed; or one subpicture is not composed of one slice.
[0108] Information specifying the size of each slice (hereinafter referred to as size information) may include at least one of width information and height information of the slice, where the width / height information may be coded in units of tiles or CTUs.
[0109] However, it may not be permitted for a slice to be divided across multiple sub-pictures, i.e., a sub-picture may be partitioned to entirely contain one or more slices, or the slices constituting a sub-picture may be restricted to being divided only in one of the horizontal or vertical directions.
[0110] Referring to FIG. 3, one subpicture or slice 310 may include one or more tiles. That is, one slice may consist of one tile or multiple tiles. However, without being limited thereto, one tile may also include multiple slices. For example, one slice may consist of a subset of multiple CTU rows belonging to the tile. In this case, information specifying the number of slices belonging to the tile (hereinafter, "number information") is encoded by the encoding device, and the decoding device may determine the number of slices constituting the tile based on the encoded number information. Information specifying the size of each slice (hereinafter, "size information") may include at least one of width information and height information of the slice. Here, the width / height information may be encoded in units of CTU size. However, if one slice consists of a subset of multiple CTU rows, only the height information of the slice may be signaled, and the width information may not be signaled. The number / size information may be signaled at at least one level among VPS, SPS, PPS, and PH.
[0111] At least one of the information regarding the number, position, and size is required only when a picture is divided into predetermined fragment regions. For example, the information may be signaled only when the picture is divided into multiple slices or tiles. For this purpose, a separate flag may be used to indicate whether the current picture is divided into multiple slices or tiles. The flag may be signaled at at least one level of VPS, SPS, PPS, and PH.
[0112] Referring to FIG. 3, one tile may be composed of multiple CTUs, and one CTU 320 (hereinafter referred to as a "first block") may be divided into multiple sub-blocks (hereinafter referred to as a "second block") by at least one vertical line or horizontal line. The vertical line and horizontal line may be one, two, or more. Hereinafter, the first block is not limited to a CTU, but may also be a coding block (CU) divided from a CTU, a prediction block (PU) which is a basic unit of predictive coding / decoding, or a transform block (TU) which is a basic unit of transform coding / decoding. The first block may be a square block or a non-square block.
[0113] The division of the first block can be performed based on a multi-tree such as a binary tree, a ternary tree, or a quad tree, as well as a quad tree.
[0114] Specifically, quadtree partitioning (QT) is a partitioning type that divides a first block into four second blocks. For example, if a first block of 2Nx2N is partitioned by QT partitioning, the first block can be quartered into four second blocks of size NxN. QT can be restricted to be applied only to square blocks, but it can also be applied to non-square blocks.
[0115] Binary tree partitioning (BT) is a partitioning type that divides a first block into two second blocks. BT can include horizontal binary trees (hereinafter referred to as horizontal BT) and vertical binary trees (hereinafter referred to as vertical BT). Horizontal BT is a partitioning type in which a first block is divided into two second blocks by a horizontal line. The division can be performed symmetrically or asymmetrically. For example, if a 2N x 2N first block is divided by horizontal BT, the first block can be divided into two second blocks with a height ratio of (a:b). Here, a and b can be the same value, or a can be larger or smaller than b. Vertical BT is a partitioning type in which a first block is divided into two second blocks by a vertical line. The division can be performed symmetrically or asymmetrically. For example, if a 2N x 2N first block is divided by vertical BT, the first block can be divided into two second blocks with a width ratio of (a:b). Here, a and b may be the same value, or a may be larger or smaller than b.
[0116] Ternary tree partitioning (TT) is a partitioning type that divides a first block into three second blocks. Similarly, TT can include horizontal ternary tree (hereinafter referred to as horizontal TT) and vertical ternary tree (hereinafter referred to as vertical TT). Horizontal TT is a partitioning type in which a first block is divided into three second blocks by two horizontal lines. For example, if a 2N x 2N first block is divided by horizontal TT, the first block can be divided into three second blocks with a height ratio of (a:b:c). Here, a, b, and c can be the same value. Alternatively, a and c can be the same, and b can be greater or less than a. For example, a and c can be 2, and b can be 1. Vertical TT is a partitioning type in which a first block is divided into three second blocks by two vertical lines. For example, if a 2N x 2N first block is divided by vertical TT, the first block can be divided into three second blocks with a width ratio of (a:b:c). Here, a, b, and c may be the same value or different values. Alternatively, a and c may be the same, and b may be larger or smaller than a. Alternatively, a and b may be the same, and c may be larger or smaller than a. Alternatively, b and c may be the same, and a may be larger or smaller than b. For example, a and c may be 2, and b may be 1.
[0117] The division may be performed based on division information signaled from the encoding device, which may include at least one of division type information, division direction information, and division ratio information.
[0118] The partition type information may specify any of the partition types predefined in the encoding / decoding device. The predefined partition types may include at least one of QT, Horizontal BT, Vertical BT, Horizontal TT, Vertical TT, and no split mode (No split). Alternatively, the partition type information may indicate whether QT, BT, or TT is applied. This may be coded in the form of a flag or index. For example, the partition type information may include at least one of a first flag indicating whether QT is applied or a second flag indicating whether BT or TT is applied. Depending on the second flag, either BT or TT can be selectively used. However, the first flag may be signaled only if the size of the first block is equal to or smaller than a predetermined threshold size. The threshold size may be a natural number greater than 64, 128, or more. If the size of the first block is greater than the threshold size, the first block may be forced to be split using only QT partitioning. Furthermore, the second flag may be signaled only if QT is not applied depending on the first flag.
[0119] The division direction information may indicate whether the division is performed horizontally or vertically in the case of BT or TT, and the division ratio information may indicate the ratio of the width and / or height of the second block in the case of BT or TT.
[0120] Assume that block 320 shown in FIG. 3 is a square block (hereinafter referred to as "first block") with a size of 8N x 8N and a partition depth of k. If the partition information of the first block indicates QT partitioning, the first block can be partitioned into four sub-blocks (hereinafter referred to as "second block"). The second block can have a size of 4N x 4N and a partition depth of (k+1).
[0121] The four second blocks can be further divided based on any one of QT, BT, TT, or non-division mode. For example, if the division information of the second block indicates Horizontal BT, the second block can be divided into two sub-blocks (hereinafter referred to as "third blocks"). In this case, the third block can have a size of 4N x 2N and a division depth of (k+2).
[0122] The third block can also be partitioned again based on any one of QT, BT, TT, or non-partition mode. For example, if the partition information of the third block represents Vertical BT, the third block can be bisected into two sub-blocks 321 and 322. In this case, the sub-blocks 321 and 322 can have a size of 2N×2N and a partition depth of (k+3). Alternatively, if the partition information of the third block represents Horizontal BT, the third block can be bisected into two sub-blocks 323 and 324. In this case, the sub-blocks 413 and 414 can have a size of 4N×N and a partition depth of (k+3).
[0123] The division may be performed independently of or in parallel with the surrounding blocks, or may be performed sequentially based on a predetermined priority.
[0124] The partition information of the current block to be divided may be determined dependently based on at least one of the partition information of the upper block of the current block or the partition information of the neighboring block. For example, if the second block is divided by horizontal BT partitioning and the upper third block is divided by vertical BT partitioning, the lower third block does not need to be divided by vertical BT partitioning. This is because if the lower third block is divided by vertical BT partitioning, the result is the same as if the second block were divided by QT partitioning. Therefore, the partition information (especially, the partition direction information) of the lower third block can be omitted from coding, and the decoding device can set the lower third block to be divided horizontally.
[0125] The upper block may refer to a block having a partition depth smaller than that of the current block. For example, if the partition depth of the current block is (k+2), the partition depth of the upper block may be (k+1). The neighboring block may be a block adjacent to the upper or left side of the current block. The neighboring block may be a block having the same partition depth as the current block.
[0126] The above-described division may be repeated up to the minimum unit of encoding / decoding. Once divided into minimum units, partition information for the corresponding block is not further signaled from the encoding device. The information for the minimum unit may include at least one of the size or shape of the minimum unit. The size of the minimum unit may be expressed as the block width, height, minimum or maximum value of the width and height, the sum of the width and height, the number of pixels, or the division depth. The information for the minimum unit may be signaled in units of at least one of a video sequence, a picture, a slice, or a block. Alternatively, the information for the minimum unit may be a value already specified in the encoding / decoding device. The information for the minimum unit may be signaled for each CU, PU, and TU. Information for one minimum unit may be equally applied to a CU, a PU, and a TU. The blocks in the embodiments described below may be obtained by the above-described block division.
[0127] According to an embodiment of the present invention, block partitioning can be obtained within a supportable range, and block partitioning configuration information for this purpose can be supported. For example, the largest coding unit (CTU), smallest coding unit, largest transform unit, smallest transform unit, and maximum partition depth k (e.g., coding / transform × Intra / Inter × QT / BT / TT, where k is 0, 1, 2, or more) of a size of m × n (e.g., m and n are natural numbers such as 2, 4, 8, 16, 32, 64, 128, etc.) per block may be included in the block partitioning configuration information. This information may be signaled at at least one level of the VPS, SPS, PPS, PH, and slice header.
[0128] In the case of some of the fragment areas (subpictures, slices, tiles, etc.) mentioned above, certain basic information (e.g., information on a sub-unit or basic unit of the fragment area, such as CTU or tile) may be required to partition / divide each fragment area (e.g., derive position and size information of the fragment area). In this case, this may be performed in the order of VPS-SPS-PPS, etc., but for simultaneous encoding / decoding, it may be necessary to provide the basic information at a level at which partition / division of each fragment area is supported.
[0129] For example, CTU information can be generated (fixedly generated) in SPS, and based on this, CTU information can be used (if divided) depending on whether or not a sub-picture (assuming processing in SPS) is divided. Alternatively, CTU information can be generated (if divided, or further generated) depending on whether or not a slice or tile (assuming processing in PPS) is divided, and based on this, division into slices and tiles can be performed.
[0130] In summary, the basis information used for the partitioning / segmentation of fragment images can occur at one level, or depending on the type of fragment image, the basis information used for the partitioning / segmentation can occur at two or more levels.
[0131] Regardless of the type of fragment image, the basic information (syntax or flags) referenced in the section of the fragment image may occur at one level. Alternatively, depending on the fragment image, the basic information referenced in each fragment image section may occur at multiple levels. In this case, even if the basic information occurs and exists at two or more levels, it can be set to have the same value or information, maintaining the same effect as occurring and existing at one level. However, in the case of the basic information, having the same value or information may be a basic setting. However, without being limited to this, variations having different values or information are also possible.
[0132] FIG. 4 is an example diagram illustrating intra prediction modes predefined in an image encoding / decoding device as an embodiment to which the present invention is applied.
[0133] 4, the previously defined intra prediction modes may be defined as a prediction mode candidate group consisting of 67 modes, specifically including 65 directional modes (Nos. 2 to 66) and two non-directional modes (DC, Planar). In this case, the directional modes may be classified into gradients (e.g., dy / dx) or angle information (Degree). All or some of the intra prediction modes described in the above example may be included in the prediction mode candidate group for the luma component or chroma component, and other additional modes may be included in the prediction mode candidate group.
[0134] In addition, a reconstructed block of another color space that has been encoded / decoded using the correlation between color spaces can be used to predict the current block, and a prediction mode that supports this can be included. For example, in the case of a chrominance component, a predicted block of the current block can be generated using a reconstructed block of a luminance component corresponding to the current block. That is, a predicted block can be generated based on a reconstructed block taking into account the correlation between color spaces.
[0135] The prediction mode candidates can be adaptively determined based on the encoding / decoding settings. The number of candidates can be increased to improve prediction accuracy, or decreased to reduce the bit amount according to the prediction mode.
[0136] For example, one of candidate groups such as candidate group A (67 candidates, including 65 directional modes and two non-directional modes), candidate group B (35 candidates, including 33 directional modes and two non-directional modes), and candidate group C (18 candidates, including 17 directional modes and one non-directional mode) can be selected, and the candidate group can be adaptively selected or determined depending on the size and shape of the block.
[0137] In addition, the prediction mode candidate group may be configured in various ways depending on the encoding / decoding settings. For example, as shown in FIG. 4, the prediction mode candidate group may be configured with equal distribution between modes, or the number of modes between mode 18 and mode 34 in FIG. 4 may be greater than the number of modes between mode 2 and mode 18. Or, the opposite is also possible. The candidate group may be configured adaptively depending on the shape of the block (i.e., square, non-square where the width is greater than the height, non-square where the height is greater than the width, etc.).
[0138] For example, if the width of the current block is greater than its height, some or all of the intra prediction modes belonging to Nos. 2 to 18 are not used and may be replaced with some or all of the intra prediction modes belonging to Nos. 67 to 80. On the other hand, if the width of the current block is smaller than its height, some or all of the intra prediction modes belonging to Nos. 50 to 66 are not used and may be replaced with some or all of the intra prediction modes belonging to Nos. -14 to -1.
[0139] Unless otherwise specified, the present invention will be described assuming that intra prediction is performed using a single predefined prediction mode candidate group (candidate group A) having uniform mode spacing. However, the main elements of the present invention can also be modified and applied to adaptive intra prediction settings as described above.
[0140] FIG. 5 illustrates a method for decoding a current block based on intra prediction as an embodiment to which the present invention is applied.
[0141] Referring to FIG. 5, a reference region for intra prediction of a current block can be determined (S500).
[0142] The reference region according to the present invention may be a region adjacent to at least one of the left side, top edge, top left edge, bottom left edge, or top right edge of the current block. Although not shown in Fig. 5, the reference region may further include a region adjacent to at least one of the right side, bottom right edge, or bottom edge of the current block. This may be selectively used based on the intra prediction mode, encoding / decoding order, scan order, etc. of the current block.
[0143] The encoding / decoding device may define a plurality of pixel lines available for intra prediction, the plurality of pixel lines including at least one of a first pixel line adjacent to the current block, a second pixel line adjacent to the first pixel line, a third pixel line adjacent to the second pixel line, or a fourth pixel line adjacent to the third pixel line.
[0144] For example, based on the encoding / decoding setting, the plurality of pixel lines may include all of the first to fourth pixel lines, or may include only the remaining pixel lines excluding the third pixel line, or may include only the first pixel line and the fourth pixel line, or may include only the first pixel line to the third pixel line.
[0145] The current block may select one or more pixel lines from the plurality of pixel lines and use the selected pixel lines as a reference region. The selection may be based on an index (refIdx) signaled by the encoding device. Alternatively, the selection may be based on predetermined encoding information. The encoding information may include at least one of the size, shape, and partition type of the current block, whether the intra prediction mode is a non-directional mode, whether the intra prediction mode is a horizontal mode, and the angle or component type of the intra prediction mode.
[0146] For example, if the intra prediction mode is planar mode or DC mode, only the first pixel line may be used. Alternatively, if the size of the current block is equal to or smaller than a predetermined threshold, only the first pixel line may be used. Here, the size may be expressed as either the width or height of the current block (e.g., a maximum value, a minimum value, etc.), the sum of the width and height, or the number of samples belonging to the current block. Alternatively, if the intra prediction mode is greater than a predetermined threshold angle (or smaller than a predetermined threshold angle), only the first pixel line may be used. The threshold angle may be the angle of the intra prediction mode corresponding to mode 2 or mode 66 from the above-mentioned candidate prediction modes.
[0147] Meanwhile, there may be a case where at least one pixel in the reference area is unavailable, in which case the unavailable pixel may be replaced with a previously determined default value or with an available pixel, as will be described in detail with reference to FIG.
[0148] Referring to FIG. 5, the intra prediction mode of the current block can be derived (S510).
[0149] The current block is a concept including a luminance block and a chrominance block, and the intra prediction mode can be determined for each of the luminance block and the chrominance block. Hereinafter, it is assumed that the intra prediction modes already defined in the decoding device are composed of a non-directional mode (Planar mode, DC mode) and 65 directional modes.
[0150] 1. Luminance block The previously defined intra prediction modes can be divided into an MPM candidate group and a non-MPM candidate group. The intra prediction mode of a current block can be derived selectively from either an MPM candidate group or a non-MPM candidate group. To this end, a flag (hereinafter, a first flag) indicating whether the intra prediction mode of the current block is derived from an MPM candidate group can be used. For example, if the first flag is a first value, the MPM candidate group can be used, and if the first flag is a second value, the non-MPM candidate group can be used.
[0151] Specifically, when the first flag is a first value, the intra prediction mode of the current block can be determined based on an MPM candidate group (candModeList) including at least one MPM candidate and an MPM index. The MPM candidate group may be information identifying any one of the MPM candidates belonging to the MPM candidate group. The MPM index can be signaled only when multiple MPM candidates belong to the MPM candidate group.
[0152] On the other hand, if the first flag is set to the second value (i.e., if there is no MPM candidate that is the same as the intra prediction mode of the current block in the MPM candidate group), the intra prediction mode of the current block can be determined based on the signaled residual mode information. The residual mode information can identify any one of the remaining modes excluding the MPM candidates.
[0153] A method for determining the MPM candidate group will be described below.
[0154] (Embodiment 1) The MPM candidate group may include at least one of intra-prediction modes modeA, (modeA-n), (modeA+n), or a default mode of neighboring blocks. The value of n may be an integer of 1, 2, 3, 4, or greater. The neighboring blocks may refer to blocks adjacent to the left and / or top of the current block. However, without being limited thereto, the neighboring blocks may also include at least one of blocks adjacent to the top left, bottom left, or top right. The default mode may be at least one of a planar mode, a DC mode, or a predetermined directional mode. The predetermined directional mode may include at least one of a horizontal mode (modeV), a vertical mode (modeH), (modeV-k), (modeV+k), (modeH-k), or (modeH+k), where k may be an integer of 1, 2, 3, 4, 5, or greater.
[0155] The MPM index may identify the same MPM as the intra prediction mode of the current block among the MPM candidates, that is, the MPM identified by the MPM index may be set as the intra prediction mode of the current block.
[0156] (Embodiment 2) The MPM candidate group may be divided into m candidate groups, where m may be an integer of 2, 3, 4, or greater. For ease of explanation, it is assumed below that the MPM candidate group is divided into a first candidate group and a second candidate group.
[0157] The encoding / decoding apparatus may select either a first candidate group or a second candidate group. The selection may be made based on a flag (hereinafter, a second flag) that specifies whether the intra prediction mode of the current block belongs to the first candidate group or the second candidate group. For example, if the second flag is a first value, the intra prediction mode of the current block may be derived from the first candidate group; otherwise, the intra prediction mode of the current block may be derived from the second candidate group.
[0158] Specifically, when a first candidate group is used based on the second flag, a first MPM index identifying one of a plurality of default modes belonging to the first candidate group may be signaled. The default mode corresponding to the signaled first MPM index may be set as the intra prediction mode of the current block. On the other hand, when the first candidate group is configured with one default mode, the first MPM index is not signaled, and the intra prediction mode of the current block may be set as the default mode of the first candidate group.
[0159] When a second candidate group is used based on the second flag, a second MPM index identifying one of a plurality of MPM candidates belonging to the second candidate group may be signaled. The MPM candidate corresponding to the signaled second MPM index may be set as the intra prediction mode of the current block. On the other hand, when the second candidate group is composed of one MPM candidate, the second MPM index is not signaled, and the intra prediction mode of the current block may be set to the MPM candidate of the second candidate group.
[0160] Meanwhile, the second flag can be signaled only when the first flag is a first value (condition 1). Furthermore, the second flag can be signaled only when the reference region of the current block is determined to be the first pixel line. When the current block refers to a non-adjacent pixel line, MPM candidates of the first candidate group can be restricted from being used. Conversely, when the intra prediction mode of the current block is derived from the first candidate group based on the second flag, the current block can be restricted to only refer to the first pixel line.
[0161] Also, the second flag can be signaled only if the current block does not perform sub-block-based intra prediction (condition 2). Conversely, if the current block performs sub-block-based intra prediction, the flag can be set to a second value in the decoding device without being signaled.
[0162] If either one of the above conditions 1 or 2 is satisfied, the second flag may be signaled, and if both conditions 1 and 2 are satisfied, the second flag may be set to be signaled.
[0163] The first candidate group may be configured with a predefined default mode. The default mode may be at least one of a directional mode or a non-directional mode. For example, the directional mode may include at least one of a vertical mode, a horizontal mode, or a diagonal mode. The non-directional mode may include at least one of a planar mode or a DC mode.
[0164] The first candidate group may consist of only r non-directional modes or directional modes, where r may be an integer of 1, 2, 3, 4, 5, or more. r may be a fixed value already assigned to the encoding / decoding device, or may be variably determined based on predetermined encoding parameters.
[0165] The second candidate group may include multiple MPM candidates. However, the second candidate group may be limited to not include the default mode belonging to the first candidate group. The number of MPM candidates may be two, three, four, five, six, or more. The number of MPM candidates may be a fixed value already set in the encoding / decoding device or may be variably determined based on encoding parameters. The MPM candidates may be derived based on the intra-prediction modes of neighboring blocks adjacent to the current block. The neighboring blocks may be blocks adjacent to at least one of the left, top, upper left, lower left, or upper right corners of the current block.
[0166] Specifically, the MPM candidate can be determined by considering whether the intra prediction mode (candIntraPredModeA) of the left block and the intra prediction mode (candIntraPredModeB) of the top block are the same, and whether candIntraPredModeA and candIntraPredModeB are non-directional modes.
[0167] [CASE 1] For example, if candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is not a non-directional mode, the MPM candidates for the current block may include at least one of candIntraPredModeA, (candIntraPredModeA-n), (candIntraPredModeA+n), or a non-directional mode. Here, n may be an integer of 1, 2, or more. The non-directional mode may include at least one of a planar mode or a DC mode. As an example, the MPM candidates for the current block may be determined as shown in Table 1 below. The index in Table 1 specifies the position or priority of the MPM candidate, but is not limited to this.
[0168] [Table 1]
[0169] [CASE 2] Alternatively, if candIntraPredModeA and candIntraPredModeB are not identical and neither candIntraPredModeA nor candIntraPredModeB is a non-directional mode, the MPM candidates for the current block may include at least one of candIntraPredModeA, candIntraPredModeB, (maxAB-n), (maxAB+n), (minAB-n), (minAB+n), or a non-directional mode. Here, maxAB and minAB refer to the maximum and minimum values of candIntraPredModeA and candIntraPredModeB, respectively, and n may be an integer of 1, 2, or greater. The non-directional mode may include at least one of planar mode or DC mode. As an example, based on the difference value (D) between candIntraPredModeA and candIntraPredModeB, the candidate modes of the second candidate group may be determined as shown in Table 2 below. The index in Table 2 identifies the position or priority of the MPM candidate, but is not limited thereto.
[0170] [Table 2]
[0171] In Table 2 above, one of the MPM candidates is derived based on minAB and the other is derived based on maxAB. However, this is not limited to this, and the MPM candidate can be derived based on maxAB regardless of minAB, and conversely, the MPM candidate can be derived based on minAB regardless of maxAB.
[0172] [CASE 3] When candIntraPredModeA and candIntraPredModeB are not identical and only one of candIntraPredModeA and candIntraPredModeB is a non-directional mode, the MPM candidates for the current block may include at least one of maxAB, (maxAB-n), (maxAB+n), or a non-directional mode. Here, maxAB refers to the maximum value of candIntraPredModeA and candIntraPredModeB, and n may be an integer of 1, 2, or greater. The non-directional mode may include at least one of a planar mode or a DC mode. As an example, the MPM candidates for the current block may be determined as shown in Table 3 below. The index in Table 3 specifies the position or priority of the MPM candidate, but is not limited thereto.
[0173] [Table 3]
[0174] [CASE 4] When candIntraPredModeA and candIntraPredModeB are not identical and both candIntraPredModeA and candIntraPredModeB are non-directional modes, the MPM candidates for the current block may include at least one of a non-directional mode, a vertical mode, a horizontal mode, (vertical mode-m), (vertical mode+m), (horizontal mode-m), or (horizontal mode+m). Here, m may be an integer of 1, 2, 3, 4, or greater. The non-directional mode may include at least one of a planar mode or a DC mode. As an example, the MPM candidates for the current block may be determined as shown in Table 5 below. The index in Table 4 specifies the position or priority of the MPM candidate, but is not limited to this. For example, index 1 may be assigned to the horizontal mode, or the largest index may be assigned. The MPM candidates may also include at least one of a diagonal mode (e.g., mode 2, mode 34, mode 66), (diagonal mode-m), or (diagonal mode+m).
[0175] [Table 4]
[0176] The intra prediction mode (IntraPredMode) decoded through the above process can be changed / corrected based on a predetermined offset, which will be described in detail with reference to FIG.
[0177] 2. Color difference block The predefined intra prediction modes for chrominance blocks can be divided into a first group and a second group, where the first group is composed of inter-component reference-based prediction modes, and the second group is composed of all or some of the predefined intra prediction modes.
[0178] The intra prediction mode of the chrominance block may be derived selectively using either the first group or the second group, and the selection may be made based on a predetermined third flag. The third flag may indicate whether the intra prediction mode of the chrominance block is derived based on the first group or the second group.
[0179] For example, if the third flag is a first value, the intra prediction mode of the chrominance block may be determined to be one of one or more inter-component reference-based prediction modes belonging to a first group, which will be described in detail with reference to FIG.
[0180] On the other hand, if the third flag is a second value, the intra prediction mode of the chrominance block may be determined to be one of a plurality of intra prediction modes belonging to group 2. For example, group 2 may be defined as shown in Table 5, and the intra prediction mode of the chrominance block may be derived based on information (intra_chroma_pred_mode) signaled by the encoding device and the intra prediction mode (IntraPredModeY) of the luma block.
[0181] [Table 5]
[0182] According to Table 5, the intra prediction mode of the chrominance block can be determined based on the signaled information and the intra prediction mode of the luma block. The mode numbers listed in Table 5 correspond to the mode numbers in FIG. 4. For example, if the signaled information intra_chroma_pred_mode has a value of 0, the intra prediction mode of the chrominance block can be determined to be the diagonal mode (66) or the planar mode (0) according to the intra prediction mode of the luma block. Alternatively, if the signaled information intra_chroma_pred_mode has a value of 4, the intra prediction mode of the chrominance block can be set to be the same as the intra prediction mode of the luma block. Meanwhile, the intra prediction mode (IntraPredModeY) of the luma block can be the intra prediction mode of a sub-block including a specific position within the luma block. Here, the specific position within the luma block can correspond to a central position within the chrominance block.
[0183] However, a sub-block in a luminance block corresponding to the center position of a chrominance block may be unavailable. Here, "unavailable" may mean that the sub-block is not coded in intra mode. For example, if the sub-block does not have an intra prediction mode, such as when the sub-block is coded in inter mode or current picture reference mode, the sub-block may be determined to be unavailable. In this case, the intra prediction mode (IntraPredModeY) of the luminance block may be set to a mode previously assigned to the encoding / decoding apparatus. Here, the previously assigned mode may be any one of planar mode, DC mode, vertical mode, or horizontal mode.
[0184] Referring to FIG. 5, the current block can be decoded based on the reference region for intra prediction and the intra prediction mode (S520).
[0185] The decoding of the current block may be performed in units of sub-blocks of the current block. To this end, the current block may be divided into a plurality of sub-blocks. Here, the current block may correspond to a leaf node. The leaf node may refer to a coding block that is not further divided into smaller coding blocks. That is, the leaf node may refer to a block that is not further divided through the above-mentioned tree-based block division.
[0186] The division may be performed based on the size of the current block (embodiment 1).
[0187] For example, if the size of the current block is smaller than a predetermined threshold size, the current block can be divided into two vertically or horizontally. Conversely, if the size of the current block is equal to or larger than the threshold size, the current block can be divided into four vertically or horizontally. The threshold size may be signaled by the encoding device or may be a fixed value predefined in the decoding device. For example, the threshold size may be expressed as N×M, where N and M may be 4, 8, 16, or more. N and M may be set to be the same or different from each other.
[0188] Alternatively, if the size of the current block is smaller than a predetermined threshold size, the current block is non-split; otherwise, the current block can be split into two or four.
[0189] The division can be performed based on the shape of the current block (embodiment 2).
[0190] For example, if the shape of the current block is square, the current block may be divided into four, otherwise the current block may be divided into two. Conversely, if the shape of the current block is square, the current block may be divided into two, otherwise the current block may be divided into four.
[0191] Alternatively, if the shape of the current block is square, the current block may be divided into halves or quarters, otherwise the current block may be left undivided. Conversely, if the shape of the current block is square, the current block may be left undivided, otherwise the current block may be divided into halves or quarters.
[0192] The division may be performed by selectively applying either the first or second embodiment described above, or may be performed by combining the first and second embodiments.
[0193] The bisection may be bisection in either the vertical or horizontal direction, and the quadrant may include bisection in either the vertical or horizontal direction, or bisection in both the vertical and horizontal directions.
[0194] In the above embodiment, the current block is divided into two or four parts, but is not limited thereto, and the current block may be divided into three parts vertically or horizontally, in which case the width or height ratio may be (1:1:2), (1:2:1), or (2:1:1).
[0195] Information regarding whether to divide into sub-block units, whether to divide into four, the division 'num', the number of divisions, etc. may be signaled from the encoding device or variably determined by the decoding device based on predetermined encoding parameters. Here, the encoding parameters may refer to the size / shape of the block, the division type (quarters, two, three), the intra prediction mode, the range / position of neighboring pixels for intra prediction, the component type (e.g., luma, chroma), the maximum / minimum size of the transform block, the transform type (e.g., transform skip, DCT2, DST7, DCT8), etc.
[0196] Sub-blocks of the current block may be predicted / reconstructed sequentially based on a predetermined priority. In this case, a first sub-block of the current block may be predicted / reconstructed, and a second sub-block may be predicted / reconstructed with reference to a first sub-block that has already been decoded. In relation to the priority, prediction / reconstruction may be performed in a top-to-bottom order, with the top and bottom sub-blocks predicted / reconstructed left-to-right. Alternatively, prediction / reconstruction may be performed in a top-to-bottom order, with the top and bottom sub-blocks predicted / reconstructed right-to-left. Alternatively, prediction / reconstruction may be performed in a bottom-to-top order, with the bottom and top sub-blocks predicted / reconstructed left-to-right. Alternatively, prediction / reconstruction may be performed in a bottom-to-top order, with the bottom and top sub-blocks predicted / reconstructed right-to-top. Alternatively, prediction / reconstruction may be performed in a left-to-right order, with the left and right sub-blocks predicted / reconstructed top-to-bottom. Alternatively, prediction / reconstruction may be performed in a left-to-right order, with the left and right sub-blocks predicted / reconstructed bottom-to-top. Alternatively, prediction / reconstruction is performed in the order from right to left, but each sub-block on the right and left can be predicted / reconstructed in the order from top to bottom. Alternatively, prediction / reconstruction is performed in the order from right to left, but each sub-block on the right and left can be predicted / reconstructed in the order from bottom to top.
[0197] The encoding / decoding device may define one of the above orders and use it fixedly. Alternatively, the encoding / decoding device may define at least two of the above orders and selectively use one of them. For this purpose, an index or flag specifying one of the predefined orders may be coded and signaled.
[0198] FIG. 6 shows a method for substituting unavailable pixels in a reference area as an embodiment to which the present invention is applied.
[0199] As described above, the reference area can be determined to be any one of the first to fourth pixel lines. However, for convenience of explanation, this embodiment assumes that the reference area is the first pixel line. This embodiment can also be applied to the second to fourth pixel lines in the same or similar manner.
[0200] If all pixels in the reference area are unavailable, the pixels can be replaced with one of the pixel value ranges expressed by the bit depth or the actual pixel value range of the image. For example, the maximum, minimum, median, or average value of the pixel value range can be substituted. If the bit depth is 8 and the median value of the bit depth is used for substitution, all pixels in the reference area can be filled with 128.
[0201] However, if this is not the case, i.e., if not all pixels in the reference area are unavailable but at least one pixel in the reference area is unavailable, a substitution process may be performed on at least one of the top reference area, left reference area, right reference area, or bottom reference area of the current block. For convenience of explanation, the following description will be focused on the left, top, and right reference areas of the current block.
[0202] (STEP 1) Determine whether the top left pixel (TL) adjacent to the current block (prediction block) is unavailable. If the top left pixel (TL) is unavailable, the pixel can be replaced with the median value of the bit depth.
[0203] (STEP 2) The top reference region may be sequentially searched for unavailable pixels. Here, the top reference region may include at least one of the pixel lines adjacent to the top or upper right corner of the current block. The length of the top reference region may be equal to the width (nW), (2*nW), or the sum of the width and height (nW+nH) of the current block.
[0204] Here, the search can be performed from left to right. In this case, if it is determined that pixel p[x][-1] is unavailable, pixel p[x][-1] can be replaced with neighboring pixel p[x-1][-1]. Alternatively, the search can be performed from right to left. In this case, if it is determined that pixel p[x][-1] is unavailable, pixel p[x][-1] can be replaced with neighboring pixel p[x+1][-1].
[0205] (STEP 3) The left reference area may be sequentially searched for unavailable pixels. Here, the left reference area may include at least one of the pixel lines adjacent to the left or bottom left edge of the current block. The length of the left reference area may be equal to the height (nH), (2*nH), or the sum of the width and height (nW+nH) of the current block.
[0206] Here, the search can be performed from the top to the bottom. If it is determined that pixel p[-1][y] is unavailable, pixel p[-1][y] can be replaced with the surrounding pixel p[-1][y-1]. Alternatively, the search can be performed from the bottom to the top. If it is determined that pixel p[-1][y] is unavailable, pixel p[-1][y] can be replaced with the surrounding pixel p[-1][y+1].
[0207] (STEP 4) The right reference area may be sequentially searched for unavailable pixels. Here, the right reference area may include a pixel line adjacent to the right of the current block. The length of the right reference area may be equal to the height (nH) of the current block.
[0208] Here, the search may be performed in a direction from the top to the bottom. In this case, if it is determined that pixel p[nW][y] is unavailable, pixel p[nW][y] may be replaced with a neighboring pixel p[nW][y-1]. Alternatively, the search may be performed in a direction from the bottom to the top. In this case, if it is determined that pixel p[nW][y] is unavailable, pixel p[nW][y] may be replaced with a neighboring pixel p[nW][y+1].
[0209] Alternatively, a separate search process for the right reference area can be omitted. Instead, the unavailable pixels in the right reference area can be filled with the median value of the bit depth. Alternatively, the unavailable pixels in the right reference area can be replaced with either the upper right pixel TR or the lower right pixel BR adjacent to the current block, or with a representative value thereof. Here, the representative value can be expressed as an average value, a maximum value, a minimum value, a mode value, a median value, or the like. Alternatively, the unavailable pixels in the right reference area can be derived by applying a predetermined weight to each of the upper right pixel TR and the lower right pixel BR. Here, the weight can be determined taking into account a first distance between the unavailable pixel in the right reference area and the upper right pixel TR and a second distance between the unavailable pixel in the right reference area and the lower right pixel BR. The lower right pixel BR can be filled with either one of the upper left pixel TL, the upper right pixel TR, or the lower left pixel BL adjacent to the current block, or can be replaced with representative values of at least two of the upper left pixel TL, the upper right pixel TR, or the lower left pixel BL. Here, the representative value is as described above. Alternatively, the bottom right pixel BR can be derived by applying a predetermined weight to at least two of the top left pixel TL, the top right pixel TR, and the bottom left pixel BL, where the weight can be determined taking into account the distance from the bottom right pixel BR.
[0210] Meanwhile, the substitution process is not limited to being performed in the order of top → left → right. For example, the substitution process may be performed in the order of left → top → right. Alternatively, the substitution process may be performed in parallel on the top and left reference areas, and then on the right reference area. Furthermore, STEP 1 may be omitted if the substitution process is performed in the order of left → top → right.
[0211] FIG. 7 is a diagram showing a method for changing / correcting an intra prediction mode as an embodiment to which the present invention is applied.
[0212] The decoded intra prediction mode (IntraPredMode) may be changed based on a predetermined offset. The offset may be selectively applied based on at least one of block attributes, i.e., size, shape, partition information, partition depth, intra prediction mode value, or component type. Here, the block may refer to the current block and / or neighboring blocks of the current block.
[0213] The partition information may include at least one of first information indicating whether the current block is divided into a plurality of sub-blocks, second information indicating a division direction (e.g., horizontal or vertical), or third information regarding the number of divided sub-blocks. The partition information may be coded and signaled by an encoding device. Alternatively, a part of the partition information may be variably determined by a decoding device based on the above-mentioned block attributes, or may be set to a fixed value predefined in the encoding / decoding device.
[0214] For example, if the first information is a first value, the current block is divided into a plurality of sub-blocks; otherwise, the current block is not divided into a plurality of sub-blocks (NO_SPLIT). If the current block is divided into a plurality of sub-blocks, the current block can be divided horizontally (HOR_SPLIT) or vertically (VER_SPLIT) based on the second information. In this case, the current block can be divided into k sub-blocks, where k can be an integer of 2, 3, 4, or more. Alternatively, k can be limited to an exponential power of 2, such as 1, 2, 4, etc. Alternatively, if the current block is a block in which at least one of the width or height is 4 (e.g., 4x8, 8x4), k can be set to 2; otherwise, k can be set to 4, 8, or 16. If the current block is not divided (NO_SPLIT), k can be set to 1.
[0215] The current block may be divided into sub-blocks having the same width and height, or into sub-blocks having different widths and heights. The current block may be divided into NxM block units (e.g., 2x2, 2x4, 4x4, 8x4, 8x8, etc.) that are already defined for the encoding / decoding device, regardless of the block attributes.
[0216] The offset can be applied only when the size of the current block is equal to or smaller than a predetermined threshold T1. Here, the threshold T1 may represent the maximum block size to which the offset is applied. Alternatively, the offset can be applied only when the size of the current block is equal to or larger than a predetermined threshold T2. In this case, the threshold T2 may represent the minimum block size to which the offset is applied. The threshold can be signaled via a bitstream. Alternatively, the threshold can be variably determined by the decoding device based on at least one of the above-mentioned block attributes, or can be a fixed value already assigned to the encoding / decoding device.
[0217] Alternatively, the offset can be applied only when the shape of the current block is non-square. For example, if the following condition is satisfied, a predetermined offset (for example, 65) can be added to the IntraPredMode of the current block.
[0218] -nW is greater than nH -IntraPredMode is greater than or equal to 2 -IntraPredMode is less than(whRatio>1)?(8+2*whRatio):8
[0219] Here, nW and nH mean the width and height of the current block, respectively, and whRatio can be set to Abs(Log2(nW / nH)).
[0220] Alternatively, if the following condition is met, a predetermined offset (for example, 67) can be subtracted from the IntraPredMode of the current block.
[0221] -nH is greater than nW -IntraPredMode is less than or equal to 66 -IntraPredMode is greater than(whRatio>1)?(60-2*whRatio):60
[0222] As described above, the final intra prediction mode may be determined by adding / subtracting an offset to / from the intra prediction mode (IntraPredMode) of the current block, taking into account the attributes of the current block. However, without being limited thereto, the application of the offset may be performed in the same / similar manner, taking into account the attributes (e.g., size, shape) of a sub-block instead of the current block.
[0223] FIG. 8 shows a prediction method based on inter-component reference as an embodiment to which the present invention is applied.
[0224] The current block can be classified into a luma block and a chroma block according to the type of component. The chroma block can be predicted using pixels of the already reconstructed luma block, which is called inter-component referencing. In this embodiment, it is assumed that the chroma block has a size of (nTbW×nTbH), and the luma block corresponding to the chroma block has a size of (2*nTbW×2*nTbH).
[0225] Referring to FIG. 8, an intra prediction mode for a chrominance block may be determined (S800).
[0226] As discussed in FIG. 5, the intra prediction mode of the chrominance block may be determined to be one of one or more inter-component reference-based prediction modes belonging to a first group based on the third flag. The first group may be composed of only inter-component reference-based prediction modes. The encoding / decoding apparatus may define at least one of INTRA_LT_CCLM, ITRA_L_CCLM, or INTRA_T_CCLM as the inter-component reference-based prediction mode. INTRA_LT_CCLM may be a mode that references both the left and top regions adjacent to the luma / chrominance block, INTRA_L_CCLM may be a mode that references the left region adjacent to the luma / chrominance block, and INTRA_T_CCLM may be a mode that references the top region adjacent to the luma / chrominance block.
[0227] A predetermined index can be used to select one of the inter-component reference-based prediction modes. The index may be information identifying one of INTRA_LT_CCLM, INTRA_L_CCLM, or INTRA_T_CCLM. The index may be signaled only when the third flag is a first value. Inter-component reference-based prediction modes belonging to the first group and indexes assigned to each prediction mode are shown in Table 6 below.
[0228] [Table 6]
[0229] Table 6 is merely an example of an index assigned to each prediction mode, and is not limited thereto. That is, as shown in Table 6, indexes may be assigned in the order of priority: INTRA_LT_CCLM, INTRA_L_CCLM, INTRA_T_CCLM, or INTRA_LT_CCLM, INTRA_T_CCLM, INTRA_L_CCLM. Alternatively, INTRA_LT_CCLM may have a lower priority than INTRA_T_CCLM or INTRA_L_CCLM. The third flag may be selectively signaled based on information indicating whether inter-component reference is permitted. For example, if the value of the information is 1, the third flag may be signaled, and if not, the third graph may not be signaled. Here, the information may be determined to be 0 or 1 based on a predetermined condition, which will be described later.
[0230] (Condition 1) If a fourth flag indicating whether prediction based on inter-component reference is allowed is 0, the information can be set to 0. The fourth flag can be signaled in at least one of a VPS, an SPS, a PPS, and a slice header.
[0231] (Condition 2) If at least one of the following sub-conditions is satisfied, the information can be set to 1:
[0232] If the value of -qtbtt_dual_tree_intra_flag is 0
[0233] -If the slice type is not I-slice
[0234] -If the size of the coding tree block is smaller than 64x64
[0235] In condition 2, qtbtt_dual_tree_intra_flag may indicate whether a coding tree block is implicitly divided into coding blocks of size 64x64 and whether a coding block of size 64x64 is divided into a dual tree. The dual tree may refer to a scheme in which luma and chroma components are divided using independent division structures. The size of the coding tree block (CtbLog2Size) may be a size (e.g., 64x64, 128x128, 256x256) predefined in the encoding / decoding device, or may be coded and signaled by the encoding device.
[0236] (Condition 3) If at least one of the following sub-conditions is satisfied, the information can be set to 1:
[0237] -If the width and height of the first upper block are 64
[0238] -When the depth of the first upper block is equal to (CtbLog2Size-6), the first upper block is divided by Horizontal BT division, and the second upper block is 64x32
[0239] -If the depth of the first upper block is greater than (CtbLog2Size-6)
[0240] - When the depth of the first upper block is the same as (CtbLog2Size-6), the first upper block is split by Horizontal BT splitting, and the second upper block is split by Vertical BT splitting.
[0241] In Condition 3, the first upper block may be a block that includes the current chrominance block as a subordinate block. For example, if the depth of the current chrominance block is k, the depth of the first upper block may be (kn), where n may be 1, 2, 3, 4, or greater. The depth of the first upper block may refer only to a depth based on quadtree-based division, or may refer to a depth based on at least one of quadtree, binary tree, and ternary tree division. The second upper block is a subordinate block belonging to the first upper block and may have a depth smaller than that of the current chrominance block but a depth larger than that of the first upper block. For example, if the depth of the current chrominance block is k, the depth of the second upper block may be (kn), where m may be a natural number smaller than n.
[0242] If none of the above conditions 1 to 3 is satisfied, the information can be set to 0.
[0243] However, even if at least one of the conditions 1 to 3 is satisfied, the information can be reset to 0 if at least one of the following sub-conditions is satisfied.
[0244] When the first upper block is 64x64 and the above-mentioned sub-block unit prediction is performed
[0245] - If at least one of the width or height of the first upper block is smaller than 64 and the depth of the first upper block is equal to (CtbLog2Size-6)
[0246] Referring to FIG. 8, a luma domain for inter-component reference of a chroma block can be identified (S810).
[0247] The luminance region may include at least one of a luminance block or an adjacent region adjacent to the luminance block. Here, the luminance block may be defined as a region including pixels pY[x][y] (x=0..nTbW*2-1, y=0..nTbH*2-1). The pixels may represent restored values before an in-loop filter is applied.
[0248] The neighboring region may include at least one of a left-side neighboring region, a top-edge neighboring region, or a top-left-edge neighboring region. The left-side neighboring region may be set to a region including pixels pY[x][y] (x=-1..-3, y=0..2*numSampL-1). This setting may be performed only if the value of numSampL is greater than 0. The top-edge neighboring region may be set to a region including pixels pY[x][y] (x=0..2*numSampT-1, y=-1..-3). This setting may be performed only if the value of numSampT is greater than 0. The top-left neighboring region may be set to a region including pixels pY[x][y] (x=-1, y=-1, -2). This setting may be performed only if the top-left region of the luminance block is available.
[0249] The above-mentioned numSampL and numSampT may be determined based on the intra prediction mode of the current block, where the current block may refer to a chrominance block.
[0250] For example, if the intra prediction mode of the current block is INTRA_LT_CCLM, it can be derived as in Equation 1. Here, INTRA_LT_CCLM may refer to a mode in which inter-component reference is performed based on the regions adjacent to the left and top of the current block.
[0251] [Formula 1] numSampT=availT?nTbW:0 numSampL=availL?nTbH:0
[0252] According to Equation 1, numSampT may be induced to nTbW if the top neighboring region of the current block is available, and may be induced to 0 if not. Similarly, numSampL may be induced to nTbH if the left neighboring region of the current block is available, and may be induced to 0 if not.
[0253] On the other hand, if the intra prediction mode of the current block is not INTRA_LT_CCLM, it can be derived as shown in Equation 2 below.
[0254] [Formula 2] numSampT=(availT&&predModeIntra==INTRA_T_CCLM)?(nTbW+numTopRight):0 numSampL=(availL&&predModeIntra==INTRA_L_CCLM)?(nTbH+numLeftBelow):0
[0255] In Equation 2, INTRA_T_CCLM may refer to a mode in which inter-component reference is performed based on the region adjacent to the top of the current block, and INTRA_L_CCLM may refer to a mode in which inter-component reference is performed based on the region adjacent to the left of the current block. numTopRight may refer to the number of all or a portion of pixels belonging to the region adjacent to the top right of the chrominance block. The "some pixels" may refer to usable pixels among the pixels belonging to the bottommost pixel row of the region. The usability determination is performed by sequentially determining whether pixels are usable from left to right, and this may be repeated until an unusable pixel is found. numLeftBelow may refer to the number of all or a portion of pixels belonging to the region adjacent to the bottom left of the chrominance block. The "some pixels" may refer to usable pixels among the pixels belonging to the rightmost pixel column of the region. The usability determination is performed by sequentially determining whether pixels are usable from top to bottom, and this may be repeated until an unusable pixel is found.
[0256] Referring to FIG. 8, downsampling may be performed on the luminance region identified in S810 (S820).
[0257] The downsampling may include at least one of: 1. downsampling for the luminance block; 2. downsampling for the left-neighboring region of the luminance block; or 3. downsampling for the top-neighboring region of the luminance block, which will be discussed in detail below.
[0258] 1. Downsampling of luminance blocks (Embodiment 1) The downsampled pixel pDsY[x][y] (x=0..nTbW-1, y=0..nTbH-1) of the luminance block can be derived based on the corresponding pixel pY[2*x][2*y] of the luminance block and neighboring pixels. The neighboring pixels may refer to pixels adjacent to the corresponding pixel in at least one of the left, right, top, and bottom directions. For example, the pixel pDsY[x][y] can be derived as shown in Equation 3 below.
[0259] [Formula 3] pDsY[x][y]=(pY[2*x][2*y-1]+pY[2*x-1][2*y]+4*pY[2*x][2*y]+pY[2*x+1][2*y]+pY[2*x][2*y+1]+4)>>3
[0260] However, there may be cases where the left / top neighboring region of the current block is unavailable. If the left neighboring region of the current block is unavailable, pixel pDsY[0][y] (y=1..nTbH-1) of the downsampled luminance block can be derived based on corresponding pixel pY[0][2*y] of the luminance block and neighboring pixels. The neighboring pixels may refer to pixels neighboring in at least one direction, either the top or bottom, of the corresponding pixel. For example, pixel pDsY[0][y] (y=1..nTbH-1) can be derived as shown in Equation 4 below.
[0261] [Formula 4] pDsY[0][y]=(pY[0][2*y-1]+2*pY[0][2*y]+pY[0][2*y+1]+2)>>2
[0262] If the top neighboring region of the current block is unavailable, pixel pDsY[x][0] (x=1..nTbW-1) of the downsampled luminance block can be derived based on corresponding pixel pY[2*x][0] of the luminance block and neighboring pixels. The neighboring pixels may refer to pixels adjacent to at least one of the left and right sides of the corresponding pixel. For example, pixel pDsY[x][0] (x=1..nTbW-1) can be derived as shown in Equation 5 below.
[0263] [Formula 5] pDsY[x][0]=(pY[2*x-1][0]+2*pY[2*x][0]+pY[2*x+1][0]+2)>>2
[0264] Meanwhile, the pixel pDsY[0][0] of the downsampled luminance block can be derived based on the corresponding pixel pY[0][0] of the luminance block and / or the surrounding pixels. The location of the surrounding pixels can be determined differently depending on the availability of the left / top neighboring regions of the current block.
[0265] For example, if the left adjacent region is available and the top adjacent region is not available, pDsY[0][0] can be derived as shown in Equation 6 below.
[0266] [Formula 6] pDsY[0][0]=(pY[-1][0]+2*pY[0][0]+pY[1][0]+2)>>2
[0267] On the other hand, if the left adjacent region is unavailable and the top adjacent region is available, pDsY[0][0] can be derived as shown in Equation 7 below.
[0268] [Formula 7] pDsY[0][0]=(pY[0][-1]+2*pY[0][0]+pY[0][1]+2)>>2
[0269] On the other hand, if neither the left nor the top neighboring regions are available, pDsY[0][0] can be set to the corresponding pixel pY[0][0] of the luminance block.
[0270] (Embodiment 2) A pixel pDsY[x][y] (x=0..nTbW-1, y=0..nTbH-1) of the downsampled luminance block can be derived based on a corresponding pixel pY[2*x][2*y] of the luminance block and neighboring pixels. The neighboring pixels may refer to pixels adjacent to the corresponding pixel in at least one direction of the bottom, left, right, bottom left, or bottom right. For example, the pixel pDsY[x][y] can be derived as shown in Equation 8 below.
[0271] [Formula 8] pDsY[x][y]=(pY[2*x-1][2*y]+pY[2*x-1][2*y+1]+2*pY[2*x][2*y]+2*pY[2*x][2*y+1]+pY[2*x+1][2*y]+pY[2*x+1][2*y+1]+4)>>3
[0272] However, if the left neighboring region of the current block is unavailable, the pixel pDsY[0][y] (y=0..nTbH-1) of the downsampled luminance block can be derived based on the corresponding pixel pY[0][2*y] of the luminance block and the neighboring pixels at the bottom. For example, the pixel pDsY[0][y] (y=0..nTbH-1) can be derived as shown in Equation 9 below.
[0273] [Formula 9] pDsY[0][y]=(pY[0][2*y]+pY[0][2*y+1]+1)>>1
[0274] The downsampling of the luminance block may be performed according to either the first or second embodiment described above. In this case, either the first or second embodiment may be selected based on a predetermined flag. The flag may indicate whether the downsampled luminance pixels have the same positions as the original luminance pixels. For example, if the flag is a first value, the downsampled luminance pixels have the same positions as the original luminance pixels. On the other hand, if the flag is a second value, the downsampled luminance pixels have the same positions as the original luminance pixels in the horizontal direction but are shifted by half a pel in the vertical direction.
[0275] 2. Downsampling of the left adjacent region of the luminance block (Embodiment 1) The downsampled pixel pLeftDsY[y] (y=0..numSampL-1) of the left-neighboring region can be derived based on the corresponding pixel pY[-2][2*y] of the left-neighboring region and surrounding pixels. The surrounding pixels may refer to pixels adjacent to the corresponding pixel in at least one direction of the left, right, top, or bottom. For example, the pixel pLeftDsY[y] can be derived as shown in Equation 10 below.
[0276] [Formula 10] pLeftDsY[y]=(pY[-2][2*y-1]+pY[-3][2*y]+4*pY[-2][2*y]+pY[-1][2*y]+pY[-2][2*y+1]+4)>>3
[0277] However, if the upper left neighboring region of the current block is unavailable, the pixel pLeftDsY[0] of the downsampled left neighboring region can be derived based on the corresponding pixel pY[-2][0] of the left neighboring region and surrounding pixels. The surrounding pixels may refer to pixels adjacent to at least one of the left and right sides of the corresponding pixel. For example, the pixel pLeftDsY[0] can be derived as shown in Equation 11 below.
[0278] [Formula 11] pLeftDsY[0]=(pY[-3][0]+2*pY[-2][0]+pY[-1][0]+2)>>2
[0279] (Embodiment 2) The downsampled pixel pLeftDsY[y] (y=0..numSampL-1) of the left-neighboring region can be derived based on the corresponding pixel pY[-2][2*y] of the left-neighboring region and surrounding pixels. The surrounding pixels may refer to pixels adjacent to the corresponding pixel in at least one direction below, to the left, to the right, or in the lower left or lower right direction. For example, the pixel pLeftDsY[y] can be derived as shown in Equation 12 below.
[0280] [Formula 12] pLeftDsY[y]=(pY[-1][2*y]+pY[-1][2*y+1]+2*pY[-2][2*y]+2*pY[-2][2*y+1]+pY[-3][2*y]+pY[-3][2*y+1]+4)>>3
[0281] Similarly, downsampling of the left-side adjacent region can be performed based on either of the above-described embodiments 1 and 2. In this case, either embodiment 1 or 2 can be selected based on a predetermined flag. The flag indicates whether the downsampled luminance pixel has the same position as the original luminance pixel, as described above.
[0282] Meanwhile, downsampling for the left neighboring region can be performed only when the numSampL value is greater than 0. A case where the numSampL value is greater than 0 may mean that the left neighboring region of the current block is available and the intra prediction mode of the current block is INTRA_LT_CCLM or INTRA_L_CCLM.
[0283] 3. Downsampling of the upper adjacent region of the luminance block (Embodiment 1) The pixels pTopDsY[x] (x=0..numSampT-1) of the downsampled top-neighboring region can be derived by considering whether the top-neighboring region belongs to a different CTU from the luminance block.
[0284] If the top-neighboring region belongs to the same CTU as the luminance block, the downsampled pixel pTopDsY[x] of the top-neighboring region can be derived based on the corresponding pixel pY[2*x][-2] of the top-neighboring region and surrounding pixels. The surrounding pixels may refer to pixels adjacent to the corresponding pixel in at least one direction of the left, right, top, or bottom. For example, the pixel pTopDsY[x] can be derived as shown in Equation 13 below.
[0285] [Formula 13] pTopDsY[x]=(pY[2*x][-3]+pY[2*x-1][-2]+4*pY[2*x][-2]+pY[2*x+1][-2]+pY[2*x][-1]+4)>>3
[0286] On the other hand, if the top-neighboring region belongs to a CTU different from the luminance block, the downsampled pixel pTopDsY[x] of the top-neighboring region can be derived based on the corresponding pixel pY[2*x][-1] of the top-neighboring region and surrounding pixels. The surrounding pixels may refer to pixels adjacent to at least one of the left and right sides of the corresponding pixel. For example, the pixel pTopDsY[x] can be derived as shown in Equation 14 below.
[0287] [Formula 14] pTopDsY[x]=(pY[2*x-1][-1]+2*pY[2*x][-1]+pY[2*x+1][-1]+2)>>2
[0288] Alternatively, if the top left neighboring region of the current block is unavailable, the neighboring pixels may refer to pixels adjacent to the corresponding pixel in at least one of the upper and lower directions. For example, pixel pTopDsY[0] can be derived as shown in Equation 15 below.
[0289] [Formula 15] pTopDsY[0]=(pY[0][-3]+2*pY[0][-2]+pY[0][-1]+2)>>2
[0290] Alternatively, if the top-left neighboring region of the current block is unavailable and the top neighboring region belongs to a different CTU than the luminance block, pixel pTopDsY[0] can be set to pixel pY[0][-1] of the top neighboring region.
[0291] (Embodiment 2) The pixels pTopDsY[x] (x=0..numSampT-1) of the downsampled top-neighboring region can be derived by considering whether the top-neighboring region belongs to a different CTU from the luminance block.
[0292] If the top-neighboring region belongs to the same CTU as the luminance block, the pixel pTopDsY[x] of the downsampled top-neighboring region can be derived based on the corresponding pixel pY[2*x][-2] of the top-neighboring region and surrounding pixels. The surrounding pixels may refer to pixels adjacent to the corresponding pixel in at least one direction among the bottom, left, right, bottom-left, and bottom-right. For example, the pixel pTopDsY[x] can be derived as shown in Equation 16.
[0293] [Formula 16] pTopDsY[x]=(pY[2*x-1][-2]+pY[2*x-1][-1]+2*pY[2*x][-2]+2*pY[2*x][-1]+pY[2*x+1][-2]+pY[2*x+1][-1]+4)>>3
[0294] On the other hand, if the top-neighboring region belongs to a CTU different from the luminance block, the downsampled pixel pTopDsY[x] of the top-neighboring region can be derived based on the corresponding pixel pY[2*x][-1] of the top-neighboring region and surrounding pixels. The surrounding pixels may refer to pixels adjacent to at least one of the left and right sides of the corresponding pixel. For example, the pixel pTopDsY[x] can be derived as shown in Equation 17 below.
[0295] [Formula 17] pTopDsY[x]=(pY[2*x-1][-1]+2*pY[2*x][-1]+pY[2*x+1][-1]+2)>>2
[0296] Alternatively, if the top left neighboring region of the current block is unavailable, the neighboring pixels may refer to pixels adjacent to the corresponding pixel in at least one direction of the top or bottom. For example, pixel pTopDsY[0] can be derived as shown in Equation 18 below.
[0297] [Formula 18] pTopDsY[0]=(pY[0][-2]+pY[0][-1]+1)>>1
[0298] Alternatively, if the top-left neighboring region of the current block is unavailable and the top neighboring region belongs to a different CTU than the luminance block, pixel pTopDsY[0] can be set to pixel pY[0][-1] of the top neighboring region.
[0299] Similarly, downsampling of the top-adjacent region may be performed based on either of the above-described embodiments 1 and 2. In this case, either embodiment 1 or 2 can be selected based on a predetermined flag. The flag indicates whether the downsampled luminance pixel has the same position as the original luminance pixel, as described above.
[0300] Meanwhile, downsampling for the top-neighboring region can be performed only when the numSampT value is greater than 0. A case where the numSampT value is greater than 0 may mean that the top-neighboring region of the current block is available and the intra prediction mode of the current block is INTRA_LT_CCLM or INTRA_T_CCLM.
[0301] The downsampling of at least one of the left and top neighboring regions of the luminance block (hereinafter referred to as the luminance reference region) may be performed using only the corresponding pixel pY[-2][2*y] and surrounding pixels at a specific position. Here, the specific position may be determined based on the position of a pixel selected from a plurality of pixels belonging to at least one of the left and top neighboring regions of the chrominance block (hereinafter referred to as the chrominance reference region).
[0302] The selected pixels may be odd-numbered pixels or even-numbered pixels in the chrominance reference region. Alternatively, the selected pixels may be a starting pixel and one or more pixels located at predetermined intervals from the starting pixel. Here, the starting pixel may be the first, second, or third pixel in the chrominance reference region. The interval may be one, two, three, four, or more sample intervals. For example, if the interval is one sample interval, the selected pixels may include the nth pixel, the (n+2)th pixel, etc. The number of selected pixels may be two, four, six, eight, or more.
[0303] The number of selected pixels, the starting pixel, and the interval may be variably determined based on at least one of the length of the chrominance reference region (i.e., numSampL and / or numSampT) or the intra-prediction mode of the chrominance block, or the number of selected pixels may be a fixed number (e.g., four) already assigned to the encoding / decoding apparatus, regardless of the length of the chrominance reference region and the intra-prediction mode of the chrominance block.
[0304] Referring to FIG. 8, parameters for inter-component reference of the chrominance block can be derived (S830).
[0305] The parameter may include at least one of a weight or an offset. The parameter may be determined taking into account an intra prediction mode of the current block. The parameter may be derived using selected pixels of a chrominance reference domain and pixels obtained through downsampling of a luma reference domain.
[0306] Specifically, the n pixels can be classified into two groups by comparing the magnitudes of the n pixels obtained through downsampling of the luminance reference region. For example, the first group can be a group of pixels having relatively large values among the n pixels, and the second group can be a group of pixels remaining among the n samples excluding the pixels in the first group. That is, the second group can be a group of pixels having relatively small values. Here, n can be 4, 8, 16, or more. The average value of the pixels belonging to the first group can be set to a maximum value (MaxL), and the average value of the pixels belonging to the second group can be set to a minimum value (MinL).
[0307] Selected pixels in the chrominance reference region may be grouped based on the grouping of the n pixels obtained through downsampling of the luminance reference region. Pixels in the chrominance reference region corresponding to pixels in the first group for the luminance reference region may be used to form a first group for the chrominance reference region, and pixels in the chrominance reference region corresponding to pixels in the second group for the luminance reference region may be used to form a second group for the chrominance reference region. Similarly, the average value of the pixels in the first group may be set to a maximum value (MaxC), and the average value of the pixels in the second group may be set to a minimum value (MinC).
[0308] Based on the calculated maximum (MaxL, MaxC) and minimum (MinL, MinC) values, weights and / or offsets for the parameters can be derived.
[0309] The chrominance block can be predicted based on the downsampled luma block and the parameters (S840).
[0310] The chrominance block can be predicted by applying at least one of the weights or offsets previously derived to the pixels of the downsampled luma block.
[0311] FIG. 9 is a diagram showing a method for constructing a reference region as an embodiment to which the present invention is applied.
[0312] The reference area according to the present invention may be an area adjacent to the current block. Next, a method for classifying reference areas by category and configuring each reference area with available pixels will be described. For convenience of explanation, the left, top, and right reference areas of the current block will be mainly described. For explanations not described in the following embodiments, please refer to or be guided by the embodiment described in FIG. 6.
[0313] Referring to FIG. 9, the pixels of the reference area can be divided into predetermined categories (SA00).
[0314] The reference area can be divided / classified into k categories, where k can be an integer of 1, 2, 3 or more. Alternatively, k can be limited to an integer of 2 or less. The reference area can be classified into one of the predetermined categories based on the image type (I / P / B), component type (Y / Cb / Cr, etc.), block attributes (size, shape, division information, division depth, etc.), position of the reference pixel, etc. Here, the block can refer to the current block and / or neighboring blocks of the current block.
[0315] For example, blocks can be classified into specific categories according to their size. In this case, the support range for block size can be determined according to the threshold size. Each threshold size can be expressed as W, H, W x H, or W * H, where W and H are width (W) and height (H). W and H can be natural numbers such as 4, 8, 16, and 32. Two or more threshold sizes are supported, and can be used to set the support range, such as the minimum and maximum values that a block can have.
[0316] Alternatively, the reference pixels may be classified into a predetermined category according to their positions, which may be defined in pixel units or according to the direction of the block to which the reference pixels belong (left, right, top, bottom, top-left, top-right, bottom-left, bottom-right).
[0317] Referring to Figure 6, the definition can be based on whether the pixel is included in the positions of the top left pixel (TL), top right pixel (TR), bottom left pixel (BL), and bottom right pixel (BR). It can also be based on whether it is included in pixel positions TL0, TL1, TR0, TR1, BL0, BL1, BR0, and BR1, which are located based on the width (2*nW), height (2*nH), or the sum of the width and height (nW+nH) of the current block. Referring to Figure 12, the definition can be based on whether it is included in pixel positions T0, T3, B0, B3, L0, L3, R0, and R3, which are pixels located at both ends of the top, bottom, left, and right blocks. It can also be based on whether it is included in pixels (T1, L2, B2, R1, etc.) located in the middle of each block.
[0318] Based on the various coding elements, the reference regions (reference pixels) can be classified into categories.
[0319] Referring to FIG. 9, the unavailable pixels belonging to the reference region can be searched (SA10).
[0320] A search can be performed sequentially to determine whether there are any unavailable pixels in the reference area. Referring to Figure 6, the start position of the search can be determined from TL, TR, BL, and BR, but is not limited thereto. In this case, if the search is performed sequentially, the start position of the search can be set to one, but if the search is performed in parallel, two or more start positions can be set.
[0321] The search area for unavailable pixels can be determined based on the starting position of the search. If one search position is specified (assuming TL), the top reference area or the left reference area can be searched before the right reference area, but this may not be the case in a coding setting (parallel processing).
[0322] Here, the search direction can be determined as either clockwise or counterclockwise. In this case, one of the clockwise or counterclockwise directions can be selected for the entire reference area. Alternatively, it can be adaptively selected based on the position of the reference area. That is, either one of the clockwise or counterclockwise search directions can be supported for the top / bottom / left / right reference areas. Here, it should be understood that the position of the reference area is not limited to the width (nW) and height (nH) of the current block (i.e., it also includes reference areas included in 2*nW, 2*nH, nW+nH, etc.).
[0323] Here, the clockwise direction can mean the direction from the bottom to the top in the left reference area, the direction from the left to the right in the top reference area, the direction from the top to the bottom in the right reference area, and the direction from the right to the left in the bottom reference area. The counterclockwise direction can be derived in the reverse direction from the clockwise direction.
[0324] For example, when starting a search from the top left pixel TL adjacent to the current block, the top reference area and the right reference area can be searched in a clockwise direction (left to right, top to bottom).The left reference area and the bottom reference area can be searched in a counterclockwise direction (top to bottom, left to right).However, the above explanation is only a partial example, and various modifications are possible.
[0325] Referring to FIG. 9, available pixels can be substituted using a method set for each category (SA20).
[0326] Unusable pixels can be replaced with a predetermined default value (e.g., the median value of a pixel value range). Alternatively, they can be replaced based on a predetermined available pixel, and the available pixel can be replaced with a value obtained by copying, linear extrapolating, interpolating, or the like from one or more adjacent available pixels. First, a process of classifying each pixel into categories based on its position is performed. Next, examples of how each of the above methods is applied according to a plurality of categories are provided, and the unavailable pixels are referred to as target pixels.
[0327] An example <1> If there are available pixels in the reference area, the target pixel can be replaced with a pixel value obtained based on the available pixels, and if there are no available pixels in the reference area, the target pixel can be replaced with a default value.
[0328] An example <2> If there are available pixels before the target pixel, including the start position of the search, the target pixel can be replaced with a pixel value obtained based on the available pixels, and if there are no available pixels before the target pixel, the target pixel can be replaced with a default value.
[0329] An example <3> As such, the target pixel can be replaced with a default value.
[0330] <1> In the case of (1), a method of substituting available pixels depending on whether or not there are available pixels in the reference area is described. <2> In the case of , a method is described in which an available pixel is substituted depending on whether or not an available pixel exists during the previous search process. <3> In this case, we explain how to substitute one fixed available pixel.
[0331] If one category is supported, <1> ~ <3> If two or more categories are supported, the pixels belonging to one category can be substituted with one available pixel. <1> ~ <3> Select one of the categories, and the reference pixels belonging to other categories are <1> ~ <3> You can select and use one of them.
[0332] 9, a reference region can be configured using available pixels (SA30), and then intra-frame prediction (intra-prediction) can be performed (SA40).
[0333] In the above-described embodiment, a method for dividing unavailable pixels in a reference area into categories and substituting available pixels according to the category has been described. In addition to unavailable pixels, available pixels may also be substituted with default values, other available pixels, or values obtained based on other available pixels according to the category.
[0334] FIG. 10 is a diagram illustrating an example of configuring an intra prediction mode set for each step according to an embodiment of the present invention.
[0335] Referring to (Step 1) of Figure 10, various prediction mode candidate groups can be supported, and any one of them can be selected implicitly or explicitly. The prediction mode candidate groups can be classified according to the number of modes, gradient information (dy / dx) of directional modes, support range of directional modes, etc. In this case, even if the number of modes is k (a directional modes and b non-directional modes), there may be prediction mode candidate groups where a or b is different.
[0336] The prediction mode candidate group selection information may be explicitly generated and signaled at at least one level of the VPS, SPS, PPS, PH, and slice header. Alternatively, the prediction mode candidate group may be implicitly selected based on coding settings. In this case, the coding settings may be defined based on the image type (I / P / B), component type, and block attributes (size, shape, partition information, partition depth, etc.). Here, the term "block" may refer to a current block and / or neighboring blocks of the current block, and may be a description that applies equally in this embodiment.
[0337] When one prediction mode candidate group is selected (B is selected in this example) through (Step 1), intra prediction or prediction mode encoding can be performed based on the selected group. Alternatively, work can be performed to configure an efficient candidate group, which will be considered through (Step 2).
[0338] Referring to (Step 2) of Figure 10, some prediction modes can be configured in various ways, and one of them can be selected implicitly or explicitly. (B0) to (B2) may be candidate configurations assuming that some prediction modes (dotted lines in the drawing) are not frequently used.
[0339] The prediction mode selection information can be generated explicitly and signaled in the CTU, coding block, prediction block, transform block, etc. Alternatively, the prediction mode selection can be implicitly based on the coding configuration, which can be defined by various previous coding elements.
[0340] Here, the shape of the block can be refined based on the width / height ratio (W:H) of the block, and the prediction mode candidate set can be configured to be different for all possible W:H ratios, or the prediction mode candidate set can be configured to be different only for a certain ratio of W:H.
[0341] When one prediction mode candidate group is selected (B1 in this example) through (Step 2), intra prediction or prediction mode encoding can be performed based on the selected group. Alternatively, work can be performed to configure an efficient candidate group, which will be considered through (Step 3).
[0342] Referring to (Step 3) of FIG. 10, since the number of prediction mode candidates is large, the candidate configuration may be based on the assumption that some prediction modes (dotted lines in the drawing) are not frequently used.
[0343] The prediction mode candidate set selection information may be generated explicitly and may be signaled by a CTU, a coding block, a prediction block, a transform block, etc. Alternatively, the prediction mode candidate set may be selected implicitly based on the coding configuration, where the coding configuration may be defined by various previous coding elements, and the prediction mode and position of the block may be further considered factors in defining the coding configuration.
[0344] In this case, the prediction mode and the block position may refer to information about neighboring blocks of the current block. That is, a prediction mode that is estimated to be less frequently used may be derived based on attribute information of the current block and the neighboring blocks, and the mode may be excluded from a prediction mode candidate group.
[0345] For example, when neighboring blocks include positions TL, T0, TR0, L0, and BL0 in Figure 12, it is assumed that the prediction mode of the corresponding block has some directionality (from the upper left to the lower right). In this case, it can be predicted with high probability that the prediction mode of the current block has some directionality, but the following processing possibilities may exist.
[0346] For example, intra prediction can be performed for any of the modes in the prediction mode candidate group, and prediction mode coding can be performed based on the prediction mode candidate group.
[0347] Alternatively, some modes may perform intra prediction for modes in the removed prediction mode candidate set, and some modes may perform prediction mode coding based on the removed prediction mode candidate set.
[0348] Comparing the above examples, it can be distinguished whether a prediction mode that is considered to have a low occurrence probability is included in the actual prediction and encoding process (which can be included as non-MPM, etc.) or whether it is removed.
[0349] In the case of (B21), it may be a configuration in which some prediction modes having a directionality different from that of the prediction mode of the adjacent block are removed, but it may also be an example in which some sparsely arranged modes are removed in case the directional mode actually occurs.
[0350] FIG. 11 shows a method for classifying intra prediction modes into a plurality of candidate groups as an embodiment to which the present invention is applied.
[0351] Next, a case where one or more candidate groups are classified for decoding of intra prediction modes will be described. This may be the same or similar concept as the MPM candidate group and non-MPM candidate group described above. Therefore, parts not mentioned in this embodiment may be the same or similar to those described in the previous embodiments. In this embodiment, the number of candidate groups may be varied and a method for constructing the candidate groups will be described later.
[0352] The intra prediction mode of the current block can be derived by selectively using one of a plurality of candidate groups, and for this purpose, the number of selection flags available is equal to or less than the number of candidate groups minus 1.
[0353] For example, when prediction modes are classified into three candidate groups (A, B, C), a flag (first flag) indicating whether the intra prediction mode of the current block is derived from candidate group A can be used.
[0354] In this case, if the first flag is a first value, the A candidate group is used, and if the first flag is a second value, a flag (second flag) indicating whether the intra prediction mode of the current block is derived from the B candidate group may be used.
[0355] In this case, if the second flag is the first value, the candidate group B is used, and if the second flag is the second value, the candidate group C is used. In the above example, three candidate groups are supported, and therefore, a total of two selection flags, the first flag and the second flag, can be used.
[0356] When one of the candidates is selected, the intra prediction mode of the current block may be determined based on the candidate group and a candidate group index. The candidate group index may be information identifying one of the candidates belonging to the candidate group. The candidate group index may be signaled only when multiple candidates belong to the candidate group.
[0357] The above example has been used to describe a configuration related to selection flags when three candidate groups are supported. As in the above configuration, flags indicating whether a prediction mode for a current block is induced can be supported sequentially from the candidate group with the highest priority (e.g., in the order of the first flag and then the second flag). That is, if a candidate group selection flag is generated and the candidate group is not selected, a selection flag for the next candidate group can be generated.
[0358] Alternatively, the selection flag EE may have a different meaning from the selection flag EE. For example, a flag (first flag) indicating whether the intra prediction mode of the current block is derived from candidate group A or B may be used.
[0359] In this case, if the first flag is a first value, the C candidate group is used, and if the first flag is a second value, a flag (second flag) indicating whether the intra prediction mode of the current block is derived from the A candidate group may be used.
[0360] At this time, when the second flag is the first value, the A candidate group may be used, and when the second flag is the second value, the B candidate group may be used.
[0361] Candidate groups A, B, and C may have m, n, and p candidates, respectively, where m may be an integer of 1, 2, 3, 4, 5, 6, or more. n may be an integer of 1, 2, 3, 4, 5, 6, or more. Alternatively, n may be an integer between 10 and 40. p may be (total number of prediction modes - mn). Here, m may be equal to or less than n, and n may be equal to or less than p.
[0362] As another example, if prediction modes are classified into four candidate groups (A, B, C, D), flags (first flag, second flag, third flag) can be used to indicate whether the intra prediction mode of the current block is derived from candidate group A, B, or C.
[0363] In this case, the second flag can be generated when the first flag is a second value, and the third flag can be generated when the second flag is a second value. That is, when the first flag is a first value, the A candidate group can be used, and when the second flag is a first value, the B candidate group can be used. When the third flag is a first value, the C candidate group can be used, and when the third flag is a second value, the D candidate group can be used. This example can also be configured as shown in the partial example (EE) of a three-candidate group configuration.
[0364] Alternatively, the second flag may be generated when the first flag is a first value, and the third flag may be generated when the first flag is a second value. When the second flag is a first value, candidate group A may be used, and when the second flag is a second value, candidate group B may be used. When the third flag is a first value, candidate group C may be used, and when the third flag is a second value, candidate group D may be used.
[0365] The candidate groups A, B, C, and D may each have m, n, p, and q candidates, where m may be an integer of 1, 2, 3, 4, 5, 6, or greater. n may be an integer of 1, 2, 3, 4, 5, 6, or greater. Alternatively, n may be an integer between 8 and 24. p may be an integer such as 6, 7, 8, 9, 10, 11, or 12. Alternatively, p may be an integer between 10 and 32. q may be (the total number of prediction modes - mnp). Here, m may be equal to or less than n, n may be equal to or less than p, and p may be equal to or less than q.
[0366] Next, a method for configuring each candidate group when multiple candidate groups are supported will be described. The reason for supporting multiple candidate groups is to achieve efficient intra-prediction mode decoding. That is, the prediction mode that is expected to be the same as the intra-prediction mode of the current block is configured as a high-priority candidate group, and the prediction mode that is not expected to be the same as the intra-prediction mode of the current block is configured as a low-priority candidate group.
[0367] For example, in FIG. 11, if category 2 (a), category 3 (b and c), and category 4 (d) are the candidate groups with the lowest priorities, the candidate groups can be configured with prediction modes that are not included in any of the candidate groups with the previous priorities. In this embodiment, the candidate groups are configured with the remaining prediction modes that are not included in the previous candidate groups, and are therefore assumed to be candidate groups that are unrelated to the priorities of the prediction modes that configure each candidate group described below (i.e., are configured only with the remaining modes that are not included in the previous candidate groups, without considering the priorities). It is also assumed that the priorities (importance) among the candidate groups are in ascending order (category 1 → category 2 → category 3 → category 4) as shown in FIG. 11, and that the priorities in the examples described below are terms used to list the prediction modes to configure each candidate group.
[0368] As with the above-described MPM candidate group, a candidate group can be configured using prediction modes of neighboring blocks, default modes, predetermined directional modes, etc. Among them, a candidate group can be configured by determining a predetermined priority for candidate group configuration in order to allocate a small amount of bits to the most likely prediction mode.
[0369] Referring to (a) of Figure 11, a priority order can be set for the first candidate group (Category 1). After the first candidate group is configured according to the number of first candidate groups based on the priority order, the remaining prediction modes (b, j, etc.) can be configured in the second candidate group (Category 2).
[0370] Referring to (b) of FIG. 11, a common priority can be supported for the first and second candidate groups. The first candidate group is configured according to the number of the first candidate group based on the priority. The second candidate group is configured according to the number of the second candidate group based on the priority (c and onward) of the prediction mode (e) finally included in the first candidate group. The remaining prediction modes (b, j, etc.) can be configured in the third candidate group (Category 3).
[0371] Referring to (c) of FIG. 11, individual priorities (first priority, second priority) for the first candidate group and the second candidate group can be supported. The first candidate group is configured according to the number of first candidate groups based on the first priority. Then, the second candidate group is configured according to the number of second candidate groups based on the second priority. The remaining prediction modes (b, w, x, etc.) can be configured with the third candidate group.
[0372] Here, the second priority is the prediction mode of the adjacent block, as in the conventional priority (first priority). <1> , default mode <2> , a given directional mode <3> However, the priority may be set based on a different importance than the first priority (for example, if the first priority is configured in the order of 1-2-3, the second priority may be configured in the order of 3-2-1). Furthermore, the second priority may be variably configured depending on the modes included in the previous candidate group, and may also be affected by the mode of the block adjacent to the current block. While the basic principle may be that the second priority is configured differently from the first priority, it should be understood that the configuration of the second priority may be partially affected by the configuration of the first candidate group. It should also be understood that in this paragraph, the first priority (previous order), the first candidate group (previous candidate group), and the second priority (current order), and the second candidate group (current candidate group) are not limited to being described in a fixed order such as numbers.
[0373] Referring to (d) of FIG. 11, a common priority (first priority) for the first and second candidate groups is supported, and an individual priority (second priority) for the third candidate group may be supported. The first candidate group is configured according to the number of the first candidate group based on the first priority. The second candidate group is configured according to the number of the second candidate group based on the priority (e and subsequent) after the prediction mode (d) finally included in the first candidate group. Then, the third candidate group is configured according to the number of the third candidate group based on the second priority. The remaining prediction modes (u, f, etc.) may be configured in a fourth candidate group.
[0374] According to an embodiment of the present invention, the intra prediction mode of the current block may be configured as one or more candidate groups, and prediction mode decoding may be performed based on the candidate groups. In this case, one or more priorities used for configuring the candidate groups may be supported. The priorities may be used in the process of configuring one or more candidate groups.
[0375] In the above example, when two or more priorities are supported, a different priority (second priority) from the priority (first priority) used for a previous candidate group (first candidate group) is used for another candidate group (second candidate group). This can be the case when all candidates in one candidate group are configured based on a single priority.
[0376] Furthermore, the first priority used for the first candidate group can be used to configure some of the candidates in the second candidate group. That is, some of the candidates in the second candidate group (cand_A) can be determined based on the first priority (starting from modes not included in the first candidate group), and some of the candidates in the second candidate group (or the remaining candidates, cand_B) can be determined based on the second priority. In this case, cand_A can be an integer of 0, 1, 2, 3, 4, 5, or more. That is, prediction modes not yet included when configuring the first candidate group can be configured to be included in the second candidate group.
[0377] For example, three candidate groups are supported, with the first candidate group consisting of two candidates, and the first priority can be determined based on the prediction mode of the neighboring block (e.g., left, top), a predetermined directional mode (e.g., Ver, Hor, etc.), and a predetermined non-directional mode (e.g., DC, Planar, etc.) (e.g., Pmode_L-Pmode_U-DC-Planar-Ver-Hor, etc.). Based on the first priority, the candidate group configuration is completed according to the number of candidates in the first candidate group (e.g., Pmode_L, Pmode_U).
[0378] In this case, the second candidate group is composed of six candidates, and the second priority can be determined by a directional mode having a predetermined interval (1, 2, 3, 4 or an integer greater than 1) between the prediction modes of the neighboring blocks, or a directional mode having a constant gradient (dy / dx 1:1, 1:2, 2:1, 4:1, 1:4, etc.) (for example, diagonal down left-diagonal down right-diagonal up right-di ... down right-diagonal up right-diagonal down left-diagonal down right-diagonal down right-diagonal up right-diagonal down left-diagonal down right-diagonal down right-diagonal up right-diagonal down left-diagonal down right-diagonal down right-diagonal up right-diagonal down<Pmode_L+2> ,<Pmode_U+2> etc.).
[0379] The second candidate group may be configured based on a second priority, according to the number of the second candidate group, or two candidates in the second candidate group may be configured based on a first priority (DC, Planar), and the remaining four candidates may be configured based on a second priority (DDL, DDR, DUR).
[0380] For ease of explanation, the same terms as in the previous explanation, such as first candidate group, second candidate group, and first priority, will be used, but it should be noted that the priority between candidate groups among multiple candidate groups is not fixed as first and second.
[0381] In summary, when multiple candidate sets are classified for predictive mode decoding and one or more priorities are supported, a given candidate set can be organized based on one or more priorities.
[0382] FIG. 12 is an exemplary diagram showing a current block and its neighboring pixels as an embodiment to which the present invention is applied.
[0383] 12 shows pixels a to p belonging to the current block and their adjacent pixels. Specifically, the figure shows referenceable pixels Ref_T, Ref_L, Ref_TL, Ref_TR, and Ref_BL adjacent to the current block, as well as non-referenceable pixels B0 to B3, R0 to R3, and BR adjacent to the current block. This figure is based on the assumption that some coding orders, scanning orders, etc. are fixed (blocks to the left, above, above-left, above-right, and below-left of the current block are referenceable), and it should be understood that the figure can be modified to a different configuration depending on changes in the coding settings.
[0384] FIG. 13 shows a method for performing intra prediction step by step as an embodiment to which the present invention is applied.
[0385] The current block may be subjected to intra prediction using pixels located in the left, right, above, and below directions. In this case, not only referable pixels but also non-referable pixels may exist, as shown in Figure 12. By accurately estimating and utilizing pixel values not only for referable pixels but also for non-referable pixel positions, coding efficiency may be improved.
[0386] Referring to FIG. 13, an arbitrary pixel value can be obtained (SB00).
[0387] Here, the arbitrary pixel may be a non-referenceable pixel around the current block or an internal pixel of the current block. Next, the arbitrary pixel position will be described with reference to the drawings.
[0388] FIG. 14 is an exemplary diagram of an arbitrary pixel for intra prediction as an embodiment to which the present invention is applied.
[0389] 14, not only pixel c, which is located outside the current block and cannot be referenced, but also pixels a and b, which are located inside the current block, can be the arbitrary pixel targets. In the following, it is assumed that the size of the current block is Width x Height and the coordinates of the upper left corner are (0, 0).
[0390] In FIG. 14, a and b can be located between (0, 0) and (Width-1, Height-1).
[0391] For example, it can be located on a boundary line (left, right, top, bottom) of the current block, e.g., (Width-1, 0)~(Width-1, Height-1) which is the right column of the current block, or (0, Height-1)~(Width-1, Height-1) which is the bottom row of the current block.
[0392] For example, the suffix k may be located in odd or even columns and rows of the current block. For example, the suffix k may be located in the even columns of the current block, or in the odd columns of the current block. Alternatively, the suffix k may be located in the even and odd rows and columns of the current block, or in the odd and odd rows and columns of the current block. Here, in addition to the odd and even numbers, the suffix k may be located in multiples or exponents of k. k and k can be an integer of 1, 2, 3, 4, 5 or more.
[0393] FIG. 14c shows a pixel that can be referenced and can be located outside the current block.
[0394] For example, it may be located on the boundary line (right, bottom in this example) of the current block, e.g., (Width, 0)~(Width, Height), which is the right boundary of the current block, or (0, Height)~(Width, Height), which is the bottom boundary of the current block.
[0395] For example, it can be located in odd or even columns and rows of the current block. For example, it can be located in an even row beyond the right boundary of the current block, or in an odd column beyond the bottom boundary of the current block. Here, in addition to the odd and even numbers, it can also be located in a multiple or exponent of k. k and k can be an integer of 1, 2, 3, 4, 5 or more.
[0396] The number of arbitrary pixels used / referenced for intra prediction of the current block may be m, where m may be 1, 2, 3, 4, or more. Alternatively, the number of arbitrary pixels may be set based on the size (width or height) of the current block. For example, Width / w_factor, Height / h_factor, or (Width*Height) / wh_factor arbitrary pixels may be used for intra prediction. Here, w_factor and h_factor may be predetermined values used as division values based on the width and height of the block, respectively, and may be integers such as 1, 2, 4, and 8. Here, wh_factor may be a predetermined value used as a division value based on the size of the block, and may be an integer such as 2, 4, 8, and 16.
[0397] The number of any number of pixels may also be determined based on all or some coding factors such as image type, component type, block attribute, intra-prediction mode, and the like.
[0398] For any pixel obtained by the above process, the pixel value can be obtained from a referenceable region, for example, based on referenceable pixels located horizontally or vertically relative to the pixel.
[0399] At this time, ( <1> Horizontal / <2> The pixel value of any pixel can be obtained by copying or averaging one or more pixels (k pixels, where k is an integer such as 1, 2, 3, 4, 5, 6, etc.) located in the vertical direction. <1> Horizontal / <2> Vertical direction) is the coordinate of any pixel, ( <1> y component / <2> x component) of the current block. <1> Left direction / <2> It is also possible to extend this to a form in which accessible pixels located in the upper direction (top direction) can be used to obtain the pixel value of any pixel.
[0400] Pixel values can be acquired based on either the horizontal or vertical direction, or both. In this case, pixel value acquisition settings can be determined based on various coding elements (such as image type, block attributes, etc., as described in the above examples).
[0401] For example, if the current block has a rectangular shape with a width greater than its height, the pixel value of an arbitrary pixel can be obtained based on referenceable pixels located in the vertical direction. Alternatively, if primary pixel values are obtained based on referenceable pixels located in the vertical and horizontal directions, a secondary pixel value (i.e., the pixel value of an arbitrary pixel) can be obtained by applying a weight to the primary pixel value obtained in the vertical direction more than to the primary pixel value obtained in the horizontal direction.
[0402] Alternatively, if the current block has a rectangular shape with its height greater than its width, the pixel value of an arbitrary pixel may be obtained based on referenceable pixels located in the horizontal direction. Alternatively, if first-order pixel values are obtained based on referenceable pixels located in the vertical and horizontal directions, a second-order pixel value (i.e., the pixel value of an arbitrary pixel) may be obtained by applying a weight to the first-order pixel value obtained in the horizontal direction more than to the first-order pixel value obtained in the vertical direction. Of course, the present invention is not limited to the above examples, and opposite configurations are also possible.
[0403] Alternatively, a predetermined candidate list may be formed, and at least one of the candidate lists may be selected to obtain the pixel value of a given pixel, and may be signaled in any one unit such as a CTU, a coding block, a prediction block, or a transform block. In this case, the candidate list may be formed of pre-determined values or may be formed based on referenceable pixels adjacent to the current block. In this case, the number of candidate lists may be an integer of 2, 3, 4, 5, or more. Alternatively, the number may be an integer between 10 and 20, or an integer between 20 and 40, or an integer between 10 and 40.
[0404] Here, the candidate list may be configured with pixel values as candidates, or with mathematical expressions or feature values that derive pixel values. In the latter case, various pixel values may be derived on an arbitrary pixel basis based on the position (x or y coordinate) of an arbitrary pixel and the mathematical expression or feature value that determines the pixel value.
[0405] As described above, one or more arbitrary pixels may be obtained, and intra prediction may be performed based on the obtained arbitrary pixels. For convenience of explanation, the following description will be made assuming that there is one arbitrary pixel. However, it is clear that the following example can be extended to a case where two or more arbitrary pixels are obtained in the same or similar manner.
[0406] Referring to FIG. 13, the image can be divided into a plurality of sub-regions based on any pixel (SB10).
[0407] Here, the sub-regions can be divided based on horizontal or vertical lines on the basis of any pixel, and the number of sub-regions can be 2, 3, 4 or an integer greater than 1. The configuration of the sub-regions will be described with reference to the following drawings.
[0408] FIG. 15 is an exemplary diagram showing an embodiment to which the present invention is applied, in which an image is divided into a plurality of sub-regions based on an arbitrary pixel.
[0409] Referring to FIG. 15, predetermined sub-regions b and c, which are vertical or horizontal lines of an arbitrary pixel d, can be obtained, and predetermined sub-region a can be obtained based on the vertical and horizontal lines.
[0410] In this case, sub-region b can be obtained between an arbitrary pixel and pixel T that can be referenced in the vertical direction, and sub-region c can be obtained between an arbitrary pixel and pixel L that can be referenced in the horizontal direction. In this case, sub-regions b and c are not always generated in a fixed manner, but it is possible for only one of the two to occur based on arbitrary pixel-related settings of the current block (for example, when b or c is also an arbitrary pixel). If only one of sub-regions b and c occurs, sub-region a may not occur.
[0411] As shown in Figure 15, T or L can refer to adjacent referenceable pixels of the current block. Alternatively, T or L can refer to any pixel e or f that is different from any pixel d and located vertically or horizontally from the pixel d. This means that any pixel d can also be T or L of any other pixel.
[0412] As in the above example, the size of the sub-region can be determined based on the number and arrangement of any pixels in the current block.
[0413] For example, if there are two or more arbitrary pixels located one square apart, sub-regions a, b, and c can each have a size of 1 x 1. Alternatively, if there is one arbitrary pixel located at c in Figure 14, sub-regions a, b, and c can each have a size of (Width x Height), (1 x Height), and (Width x 1), respectively.
[0414] Referring to FIG. 13, intra-frame prediction (intra-prediction) can be performed in a predetermined order (SB20).
[0415] Here, depending on the location of any pixel, some sub-regions can be used as predictors for intra prediction or as temporary references for intra prediction.
[0416] For example, if a given pixel (or sub-region d) is located at c in FIG. 14, sub-regions b and c are located outside the current block, and only sub-region a can be subject to intra prediction. Alternatively, if a given pixel is located at the top of c in FIG. 14, sub-region b is located outside the current block, and sub-regions a and c can be subject to intra prediction. Alternatively, if a given pixel is located at the left side of c in FIG. 14, sub-region c is located outside the current block, and sub-regions a and b can be subject to intra prediction. Alternatively, if a given pixel is located inside the current block, sub-regions a, b, and c can be subject to intra prediction.
[0417] As in the above example, depending on the location of any pixel, it can be subject to intra prediction or used as a temporary reference value.
[0418] Although the following description will be made assuming that a pixel is located inside the current block, the following examples can be understood to be equally or similarly applicable even if the pixel location is changed.
[0419] Since the arbitrary pixel position and pixel value have been obtained in the previous step, the sub-regions a, b, and c can be acquired according to a predetermined priority order. For example, the pixel values of each sub-region can be acquired in the order b → c or c → b, and then the pixel value of sub-region a can be acquired.
[0420] In the case of sub-region b, the pixel value can be obtained based on any pixel d or T. In the case of sub-region c, the pixel value can be obtained based on any pixel d or L.
[0421] Although not shown, let us assume that the pixel on the top left is TL (i.e., the intersection of the horizontal line T and the vertical line L). If TL and T are located at the top of the current block, the pixels between TL and T are referable. Also, if TL and L are located on the left side of the current block, the pixels between TL and L are referable because the referable pixels belong to the neighboring block of the current block.
[0422] On the other hand, if either TL or T is located inside the current block, the pixels between TL and T are referable. Also, if either TL or L is located inside the current block, the pixels between TL and L are referable. This is because the referable pixels may be sub-regions obtained based on any other pixels.
[0423] Therefore, for sub-region a, pixel values can be obtained based on sub-regions b and c, the accessible region between TL and T, and the accessible region between TL and L. Of course, TL, T, L, and d can also be used to obtain pixel values for sub-region a, which means that any pixel can be referenced.
[0424] Through this process, intra prediction of the current block can be performed based on any pixel.
[0425] Performing intra prediction based on any pixel can be configured as one of the intra prediction modes, or can be included as an alternative to the conventional modes.
[0426] Alternatively, it may be classified as one of the candidate prediction methods, and selection information for that may be generated. For example, it may be considered as an additional prediction method to the method of performing intra prediction based on a directional mode or a non-directional mode. The selection information may be signaled by a CTU, a coding block, a prediction block, a transform block, etc.
[0427] The following description will be made assuming that the arbitrary pixel position is c in Figure 14. However, this is not limited to this, and the contents described below can be applied in the same or similar manner even if the arbitrary pixel position is arranged at another position. Next, the description will be made with reference to Figure 12.
[0428] The current block and the blocks to the right and below the current block are not coded but can be estimated based on data in the referenceable region.
[0429] For example, data may be directly copied or derived from regions Ref_TR, Ref_BL, etc., adjacent to the right and bottom boundaries of the current block, and then filled into the right or bottom boundaries of the current block. For example, the right boundary may be filled by directly copying one of pixels T3, TR0, TR1, etc., or by applying filtering to T3, TR0, TR1, etc.
[0430] Alternatively, data may be directly copied or derived from neighboring regions of the current block, such as Ref_TL, Ref_T, Ref_L, Ref_TR, and Ref_BL, and then embedded in the lower right boundary of the current block. For example, values obtained based on one or more pixels in the neighboring regions may be embedded in the lower right boundary of the current block.
[0431] Here, the right boundary of the current block can be (d~p) or (R0~R3). The bottom boundary of the current block can be (m~p) or (B0~B3). The bottom right boundary of the current block can be one of p, BR, R3, and B3.
[0432] In the following example, assume that the right boundary is R0 to R3, the bottom boundary is B0 to B3, and the bottom right boundary is BR.
[0433] (Right and bottom boundary processing) For example, the right boundary can be filled by copying one of the vertically adjacent pixels T3, TR0, or TR1, and the bottom boundary can be filled by copying one of the horizontally adjacent pixels L3, BL0, or BL1.
[0434] Alternatively, the right boundary can be filled with a weighted average of the vertically adjacent values T3, TR0, and TR1, and the bottom boundary can be filled with a weighted average of the horizontally adjacent values L3, BL0, and BL1.
[0435] After obtaining the values of the right and bottom boundaries, intra prediction of the current block can be performed based on the values.
[0436] (bottom right border processing) For example, the lower right boundary can be filled by copying any one of T3, TR0, TR1, L3, BL0, and BL1. Or, it can be filled with a weighted average of any one of T3, TR0, and TR1 and any one of L3, BL0, and BL1. Or, it can be filled with any one of the weighted averages of T3, TR0, and TR1 and L3, BL0, and BL1. Or, it can be filled with a quadratic weighted average of the linear weighted average of T3, TR0, and TR1 and the linear weighted average of L3, BL0, and BL1.
[0437] After obtaining the value of the lower right boundary, the value of the right boundary or the lower boundary can be obtained based on the obtained value, and intra prediction of the current block can be performed based on the right boundary or the lower boundary.
[0438] Next, the processing of the lower right boundary will be described.
[0439] Assume a configuration that considers the positions of TL, TR0, BL0, and BR, where BR refers to the pixel at the bottom right boundary, TR0 refers to a referenceable pixel located in the vertical direction of BR, BL0 refers to a referenceable pixel located in the horizontal direction of BR, and TL can be a referenceable pixel located at the top left boundary of the current block or at the intersection of the horizontal direction of TR0 and the vertical direction of BL0.
[0440] Based on the pixel positions, the directionality and feature information (eg, edges) of the current block can be estimated.
[0441] An example <1> As such, when moving diagonally from TL to BR, the pixel values can gradually increase or decrease, where TL can be greater than or equal to TR0 and BL0, and BR can be less than or equal to TR0 and BL0, or vice versa.
[0442] An example <2> As such, when moving diagonally from BL0 to TR0, the pixel values can gradually increase or decrease, where BL0 can be greater than or equal to TL and BR, and TR0 can be less than or equal to TL and BR, or vice versa.
[0443] An example <3> Assuming that the pixel values gradually increase or decrease when moving horizontally from left (TL, BL0) to right (TR0, BR), TL can be greater than or equal to TR0, and BL0 can be greater than or equal to BR, or vice versa.
[0444] An example <4> As such, when moving vertically from the top (TL, TR0) to the bottom (BL0, BR), the pixel values may gradually increase or decrease. In this case, TL may be greater than or equal to BL0, and TR0 may be greater than or equal to BR, or the opposite may be true.
[0445] If the current block has image features like those in the example above, the lower right boundary can be predicted using these features. In this case, pixels located vertically or horizontally from the pixel to be estimated and pixels at the intersections of the pixels in the vertical or horizontal direction can be requested, and the pixel to be estimated can be predicted based on a comparison of these pixel values.
[0446] <1> For example, when the conditions are TL<=TR0 and TL<=BL0, it is possible to estimate that there is a tendency for the value to increase from TL to the BR side, and to induce (predict) the BR value based on the difference value between pixels.
[0447] For example, the BR pixel value can be derived by adding the (BL0-TL) value to TR0, or by adding the (TR0-TL) value to BL0, or by averaging or weighted averaging the previous two.
[0448] <2> For example, when the conditions are TL>=BL0 and TL<=TR0, it is possible to estimate that there is a tendency for the value to increase from BL0 toward TR0, and to derive the BR value based on the difference value between pixels.
[0449] For example, the BR pixel value is derived by subtracting the (TL-BL0) value from TR0, or by adding the (TR0-TL) value to BL0, or by averaging or weighted averaging the previous two.
[0450] The above example describes a case where a block characteristic is estimated based on predetermined pixels adjacent to the current block and a given pixel (BR) is predicted. However, it may be difficult to accurately grasp the block characteristic due to limited pixels. For example, if some of the pixels referenced to derive the BR have impulse components, it may be difficult to accurately grasp the characteristic.
[0451] For this reason, characteristic information (e.g., variance, standard deviation, etc.) of the top and left regions of the current block can be calculated. For example, if it is determined that characteristic information of pixels between TL and TR0 or characteristic information of pixels between TL and BL0 well reflects an increase or decrease in a block, a method of deriving a value of an arbitrary pixel such as BR based on predetermined pixels of the current block such as TL, TR0, and BL0 can be used.
[0452] The various embodiments of the present disclosure are not intended to enumerate all possible combinations but to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.
[0453] Additionally, various embodiments of the present disclosure may be implemented using hardware, firmware, software, or a combination thereof, etc. In the case of a hardware implementation, the implementation may be implemented using one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.
[0454] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of the various embodiments to be performed on a device or computer, as well as non-transitory computer-readable media on which such software or instructions are stored and executable on a device or computer. [Industrial Applicability]
[0455] The present invention can be used to encode / decode video signals.
Claims
1. determining a reference region for intra prediction of the current block; deriving an intra prediction mode of the current block; decoding the current block based on the reference region and the intra prediction mode; The intra prediction modes already defined in the decoding device are divided into an MPM candidate group and a non-MPM candidate group. The MPM candidate group includes at least one of a first candidate group or a second candidate group; An image decoding method, wherein the intra prediction mode of the current block is derived from either the first candidate group or the second candidate group.
2. The first candidate group is configured with a default mode already defined in the decoding device, The image decoding method according to claim 1 , wherein the second candidate group is made up of a plurality of MPM candidates.
3. The image decoding method according to claim 2 , wherein the default mode is at least one of a planar mode, a DC mode, a vertical mode, a horizontal mode, a vertical mode, and a diagonal mode.
4. The plurality of MPM candidates include at least one of an intra prediction mode of a surrounding block, a mode obtained by subtracting an n value from the intra prediction mode of the surrounding block, or a mode obtained by adding an n value to the intra prediction mode of the surrounding block, 3. The image decoding method according to claim 2, wherein n is a natural number of 1, 2 or more.
5. the plurality of MPM candidates include at least one of a DC mode, a vertical mode, a horizontal mode, a mode obtained by subtracting or adding an m value to the vertical mode, or a mode obtained by subtracting or adding an m value to the horizontal mode; 3. The image decoding method according to claim 2, wherein m is a natural number of 1, 2, 3, 4 or more.
6. The image decoding method according to claim 1 , wherein either the first candidate group or the second candidate group is selected based on a flag signaled by an encoding device.
7. The induced intra-prediction mode is changed by applying a predetermined offset to the induced intra-prediction mode; The image decoding method of claim 1 , wherein the offset is selectively applied based on at least one of a size, a shape, partition information, an intra-prediction mode value, or a component type of the current block.
8. The step of determining the reference region includes: searching for unavailable pixels belonging to the reference region; and substituting the unavailable pixels with available pixels.
9. The image decoding method according to claim 8 , wherein the available pixels are determined based on a bit depth value or are pixels adjacent to at least one of the left, right, top, and bottom edges of the unavailable pixels.
Citation Information
Patent Citations
Image encoding device and image decoding device
JP2021010046A
Intra prediction-based video encoding / decoding method and device
JP2025131835A
Intra prediction-based video encoding / decoding method and device
JP2025131836A
Intra prediction-based video encoding / decoding method and device
JP2025131840A
Intra prediction based image encoding / decoding method and apparatus
JP7467467B2
Cited By
Intra prediction-based video encoding / decoding method and device
JP2025131835A
Intra prediction-based video encoding / decoding method and device
JP2025131836A
Intra prediction-based video encoding / decoding method and device
JP2025131840A