Video coding method and apparatus using derivation of intra prediction mode

The video coding method and device address the challenges of increasing image sizes and resolutions by deriving intra prediction modes from restored reference samples, enhancing encoding efficiency and image quality.

WO2025116448A1PCT designated stage expired Publication Date: 2025-06-05HYUNDAI MOTOR CO LTD +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/018715
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-22
Filing Date
2024-11-25
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing video compression technologies, such as H.264/AVC, HEVC, and VVC, face challenges in efficiently encoding increasing image sizes, resolutions, and frame rates, leading to a need for improved encoding efficiency and image quality.

Method used

A video coding method and device that derive an intra prediction mode for a current block using restored peripheral reference sample values, generating a prediction block based on the derived mode, thereby enhancing encoding efficiency.

Benefits of technology

The proposed method improves encoding efficiency by effectively utilizing restored reference samples to determine intra prediction modes, leading to better image quality and reduced data requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024018715_05062025_PF_FP_ABST
    Figure KR2024018715_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a video coding method using a derivation of an intra prediction mode. The method comprises the steps of: deriving an intra prediction mode of a current block using reconstructed neighboring reference sample values; and generating a prediction block of the current block on the basis of the derived intra prediction mode.
Need to check novelty before this filing date? Find Prior Art

Description

Video coding method and device using intra prediction mode derivation

[0001] The present invention relates to encoding and decoding of images, and more particularly, to a video coding method and device using the derivation of an intra prediction mode.

[0002] The content described below merely provides background information related to the present embodiment and does not constitute prior art.

[0003] Since video data has a large amount of data compared to voice data or still image data, it requires a lot of hardware resources, including memory, to store or transmit it without processing for compression.

[0004] Therefore, when storing or transmitting video data, the encoder compresses the video data and stores or transmits it, and the decoder receives the compressed video data, decompresses it, and plays it back. These video compression technologies include H.264 / AVC, HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding), which improves encoding efficiency by about 30% compared to HEVC.

[0005] However, as the size, resolution, and frame rate of images are gradually increasing, and the amount of data that needs to be encoded is also increasing, a new compression technology that has better encoding efficiency and better image quality improvement than existing compression technologies is required.

[0006] The present disclosure is intended to solve these problems, and provides a video coding method and device for deriving an intra prediction mode of a current block using restored surrounding reference sample values, and then generating a prediction block of the current block based on the derived prediction mode.

[0007] The problems to be solved by the present invention are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below.

[0008] One aspect of the present disclosure provides an image decoding method for generating a prediction block of a current block. The image decoding method comprises: determining a size of a sampling unit for gradient operation based on a size or area of ​​the current block; calculating gradients for sampling units of the determined size within a gradient calculation region (template) including reconstructed reference samples around the current block, and generating a histogram for directional prediction modes corresponding to the gradients; deriving at least one dominant prediction mode based on the histogram; and generating the prediction block using the at least one dominant prediction mode.

[0009] Another aspect of the present disclosure provides an image encoding method for generating a prediction block of a current block. The image encoding method includes the steps of: determining a size of a sampling unit for a gradient operation based on a size or area of ​​the current block; calculating gradients for sampling units of the determined size within a gradient calculation region (template) including reconstructed reference samples around the current block, and generating a histogram for directional prediction modes corresponding to the gradients; deriving at least one dominant prediction mode based on the histogram; and generating the prediction block using the at least one dominant prediction mode.

[0010] Another aspect of the present disclosure provides a method for storing or transmitting a bitstream generated by the aforementioned image encoding method.

[0011] Another aspect of the present disclosure provides a computer-readable recording medium for storing a bitstream generated by the aforementioned image encoding method.

[0012] As described above, according to the present embodiment, there is provided a video coding method and device that derives an intra prediction mode of a current block using restored surrounding reference sample values, and then generates a prediction block of the current block based on the derived prediction mode, thereby making it possible to improve encoding efficiency.

[0013] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.

[0014] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the techniques of the present disclosure.

[0015] Figure 2 is a drawing for explaining a method of dividing a block using the QTBTTT structure.

[0016] FIGS. 3A and 3B are diagrams illustrating multiple intra prediction modes, including wide-angle intra prediction modes.

[0017] Figure 4 is an example diagram of the surrounding blocks of the current block.

[0018] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the techniques of the present disclosure.

[0019] FIG. 6 is a block diagram illustrating in detail a portion of an image decoding device according to one embodiment of the present disclosure.

[0020] FIG. 7 is an exemplary diagram showing a gradient calculation area used for calculating gradient values ​​according to one embodiment of the present disclosure.

[0021] FIGS. 8A to 8C are diagrams illustrating various embodiments of a method for configuring sampling units for applying a gradient operator.

[0022] Figure 9 is a diagram showing an example of a directional mode corresponding to a gradient for some sampling units.

[0023] Figures 10a to 10c are drawings showing various examples of dividing a gradient calculation area into multiple areas.

[0024] FIG. 11 is a flowchart illustrating an intra prediction method using prediction mode derivation according to one embodiment of the present disclosure.

[0025] Hereinafter, some embodiments of the present disclosure will be described in detail with reference to exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, in describing the present embodiments, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present embodiments.

[0026] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the techniques of the present disclosure. Hereinafter, the image encoding device and its subcomponents will be described with reference to the illustration in FIG. 1.

[0027] The video encoding device may be configured to include a picture segmentation unit (110), a prediction unit (120), a subtractor (130), a transformation unit (140), a quantization unit (145), a reordering unit (150), an entropy encoding unit (155), an inverse quantization unit (160), an inverse transformation unit (165), an adder (170), a loop filter unit (180), and a memory (190).

[0028] Each component of the video encoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.

[0029] A single image (video) is composed of one or more sequences containing multiple pictures. Each picture is divided into multiple regions, and encoding is performed for each region. For example, a single picture is divided into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile or / slice is divided into one or more Coding Tree Units (CTUs). Each CTU is then divided into one or more Coding Units (CUs) by a tree structure. Information applied to each CU is encoded as the syntax of the CU, and information commonly applied to CUs included in a CTU is encoded as the syntax of the CTU. In addition, information commonly applied to all blocks within a single slice is encoded as the syntax of the slice header, and information applied to all blocks constituting one or more pictures is encoded in the Picture Parameter Set (PPS) or the picture header. Furthermore, information commonly referenced by multiple pictures is encoded in a Sequence Parameter Set (SPS). And, information commonly referenced by one or more SPS is encoded in a Video Parameter Set (VPS). In addition, information commonly applied to one tile or tile group may be encoded as syntax of a tile or tile group header. Syntaxes included in an SPS, PPS, slice header, tile or tile group header may be referred to as high level syntax.

[0030] The picture segmentation unit (110) determines the size of the CTU (Coding Tree Unit). Information about the size of the CTU (CTU size) is encoded as the syntax of SPS or PPS and transmitted to the image decoding device.

[0031] The picture segmentation unit (110) divides each picture constituting an image into a plurality of Coding Tree Units (CTUs) having a predetermined size, and then recursively divides the CTUs using a tree structure. A leaf node in the tree structure becomes a coding unit (CU), which is the basic unit of encoding.

[0032] The tree structure may be a QuadTree (QT) in which an upper node (or parent node) is divided into four lower nodes (or child nodes) of the same size, a BinaryTree (BT) in which an upper node is divided into two lower nodes, or a TernaryTree (TT) in which an upper node is divided into three lower nodes in a 1:2:1 ratio, or a structure that mixes two or more of the QT structures, BT structures, and TT structures. For example, a QTBT (QuadTree plus BinaryTree) structure may be used, or a QTBTTT (QuadTree plus BinaryTree TernaryTree) structure may be used. Here, BTTT may be combined and referred to as a MTT (Multiple-Type Tree).

[0033] Figure 2 is a drawing for explaining a method of dividing a block using the QTBTTT structure.

[0034] As illustrated in FIG. 2, a CTU may first be split into a QT structure. The quadtree splitting may be repeated until the size of the splitting block reaches the minimum block size (MinQTSize) of the leaf node allowed in the QT. A first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of the lower layer is encoded by the entropy encoding unit (155) and signaled to the image decoding device. If the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, it may be further split into one or more of the BT structure or the TT structure. In the BT structure and / or the TT structure, multiple splitting directions may exist. For example, there may be two directions in which the block of the corresponding node is split horizontally and two directions in which the block is split vertically. As illustrated in FIG. 2, when MTT splitting begins, a second flag (mtt_split_flag) indicating whether nodes have been split, and if splitting has occurred, a flag indicating the splitting direction (vertical or horizontal) and / or a flag indicating the splitting type (Binary or Ternary) are encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0035] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node is split into four nodes of a lower layer, a CU split flag (split_cu_flag) indicating whether the node is split may be encoded. If the CU split flag (split_cu_flag) value indicates that the node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU (coding unit), which is a basic unit of encoding. If the CU split flag (split_cu_flag) value indicates that the node is split, the video encoding device starts encoding from the first flag in the above-described manner.

[0036] As another example of a tree structure, when QTBT is used, there may be two types: a type that horizontally splits the block of the corresponding node into two blocks of the same size (i.e., symmetric horizontal splitting) and a type that vertically splits it (i.e., symmetric vertical splitting). A split flag (split_flag) indicating whether each node of the BT structure is split into blocks of a lower layer and split type information indicating the type of split are encoded by the entropy encoding unit (155) and transmitted to the image decoding device. Meanwhile, there may additionally be a type that splits the block of the corresponding node into two blocks of an asymmetrical shape. The asymmetric shape may include a shape that splits the block of the corresponding node into two rectangular blocks with a size ratio of 1:3, or a shape that splits the block of the corresponding node in a diagonal direction.

[0037] A CU can have various sizes depending on the QTBT or QTBTTT partitioning from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block." Depending on the QTBTTT partitioning employed, the current block may be rectangular as well as square.

[0038] The prediction unit (120) predicts the current block and generates a prediction block. The prediction unit (120) includes an intra prediction unit (122) and an inter prediction unit (124).

[0039] In general, each current block within a picture can be predictively coded. Prediction of the current block can typically be performed using either intra-prediction (using data from the picture containing the current block) or inter-prediction (using data from a picture coded before the picture containing the current block). Inter-prediction encompasses both unidirectional and bidirectional prediction.

[0040] The intra prediction unit (122) predicts pixels within the current block using pixels (reference pixels) located around the current block within the current picture including the current block. There are multiple intra prediction modes depending on the prediction direction. For example, as shown in Fig. 3a, the multiple intra prediction modes may include two non-directional modes including a planar mode and a DC mode, and 65 directional modes. The surrounding pixels to be used and the calculation formula are defined differently depending on each prediction mode.

[0041] For efficient directional prediction for a rectangular current block, directional modes (intra prediction modes 67 to 80 and -1 to -14) indicated by dotted arrows in Fig. 3b may be additionally used. These may be referred to as "wide-angle intra-prediction modes." In Fig. 3b, the arrows point to corresponding reference samples used for prediction, and do not indicate the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. Wide-angle intra-prediction modes are modes that perform prediction in the opposite direction of a specific directional mode without additional bit transmission when the current block is rectangular. At this time, among the wide-angle intra-prediction modes, some wide-angle intra-prediction modes available for the current block may be determined based on the ratio of the width and height of the rectangular current block. For example, wide-angle intra prediction modes (intra prediction modes 67 to 80) having an angle less than 45 degrees are available when the current block is a rectangular shape whose height is smaller than its width, and wide-angle intra prediction modes (intra prediction modes -1 to -14) having an angle greater than -135 degrees are available when the current block is a rectangular shape whose width is larger than its height.

[0042] The intra prediction unit (122) can determine an intra prediction mode to be used to encode the current block. In some examples, the intra prediction unit (122) can encode the current block using multiple intra prediction modes and select an appropriate intra prediction mode to be used from the tested modes. For example, the intra prediction unit (122) can calculate bit-rate distortion values ​​using rate-distortion analysis for multiple tested intra prediction modes and select an intra prediction mode with the best bit-rate distortion characteristics among the tested modes.

[0043] The intra prediction unit (122) selects one intra prediction mode from among multiple intra prediction modes and predicts the current block using surrounding pixels (reference pixels) and an operation formula determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.

[0044] The inter prediction unit (124) generates a prediction block for the current block using a motion compensation process. The inter prediction unit (124) searches for a block most similar to the current block within reference pictures that were encoded and decoded before the current picture, and generates a prediction block for the current block using the searched block. Then, a motion vector (MV) corresponding to the displacement between the current block within the current picture and the prediction block within the reference picture is generated. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma component. The motion information including information on the reference picture used to predict the current block and information on the motion vector is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.

[0045] The inter prediction unit (124) may perform interpolation on a reference picture or a reference block to improve prediction accuracy. That is, subsamples between two consecutive integer samples are interpolated by applying filter coefficients to a plurality of consecutive integer samples including the two integer samples. When a process of searching for a block most similar to the current block is performed on the interpolated reference picture, the motion vector can be expressed up to a precision in decimal units rather than a precision in integer sample units. The precision or resolution of the motion vector can be set differently for each target region to be encoded, such as a slice, tile, CTU, CU, etc. When such adaptive motion vector resolution (AMVR) is applied, information on the motion vector resolution to be applied to each target region must be signaled for each target region. For example, when the target region is a CU, information on the motion vector resolution applied to each CU is signaled. Information on the motion vector resolution may be information indicating the precision of a differential motion vector, which will be described later.

[0046] Meanwhile, the inter prediction unit (124) can perform inter prediction using bi-prediction. In the case of bi-prediction, two reference pictures and two motion vectors indicating the block position most similar to the current block within each reference picture are used. The inter prediction unit (124) selects a first reference picture and a second reference picture from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively, and searches for a block similar to the current block within each reference picture to generate a first reference block and a second reference block. Then, the first reference block and the second reference block are averaged or weighted averaged to generate a prediction block for the current block. Then, motion information including information on two reference pictures used to predict the current block and information on two motion vectors is transmitted to the encoding unit (150). Here, reference picture list 0 may be composed of pictures that are before the current picture in display order among the restored pictures, and reference picture list 1 may be composed of pictures that are after the current picture in display order among the restored pictures. However, this is not necessarily limited to this, and restored pictures that are after the current picture in display order may be additionally included in reference picture list 0, and conversely, restored pictures that are before the current picture may be additionally included in reference picture list 1.

[0047] Various methods can be used to minimize the number of bits required to encode motion information.

[0048] For example, if the reference picture and motion vector of the current block are identical to those of a neighboring block, the motion information of the current block can be transmitted to the image decoding device by encoding information that can identify the neighboring block. This method is called 'merge mode.'

[0049] In merge mode, the inter prediction unit (124) selects a predetermined number of merge candidate blocks (hereinafter referred to as 'merge candidates') from the surrounding blocks of the current block.

[0050] As the surrounding blocks for deriving merge candidates, all or part of the left block (A0), the lower left block (A1), the upper block (B0), the upper right block (B1), and the upper left block (A2) adjacent to the current block within the current picture may be used, as illustrated in FIG. 4. In addition, a block located within a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the current block is located may be used as a merge candidate. For example, a block co-located with the current block within the reference picture or blocks adjacent to the block at the co-located block may be additionally used as a merge candidate. If the number of merge candidates selected by the method described above is less than a preset number, a 0 vector is added to the merge candidates.

[0051] The inter prediction unit (124) uses these surrounding blocks to construct a merge list containing a predetermined number of merge candidates. Among the merge candidates included in the merge list, a merge candidate to be used as motion information of the current block is selected and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the encoding unit (150) and transmitted to the video decoding device.

[0052] Merge Skip mode is a special case of merge mode. After quantization, when all transform coefficients for entropy encoding are close to zero, only neighboring block selection information is transmitted without residual signals. By utilizing merge skip mode, relatively high encoding efficiency can be achieved for low-motion images, still images, and screen content images.

[0053] Hereinafter, merge mode and merge skip mode are collectively referred to as merge / skip mode.

[0054] Another method for encoding motion information is Advanced Motion Vector Prediction (AMVP) mode.

[0055] In AMVP mode, the inter prediction unit (124) derives predicted motion vector candidates for the motion vector of the current block using neighboring blocks of the current block. As neighboring blocks used to derive predicted motion vector candidates, all or some of the left block (A0), the lower left block (A1), the upper block (B0), the upper right block (B1), and the upper left block (A2) adjacent to the current block in the current picture as shown in FIG. 4 may be used. In addition, a block located in a reference picture (which may or may not be the same as the reference picture used to predict the current block) other than the current picture in which the current block is located may be used as the neighboring block used to derive predicted motion vector candidates. For example, a block located in the same position as the current block (collocated block) in the reference picture or blocks adjacent to the block in the same position may be used. If the number of motion vector candidates is less than a preset number by the method described above, a 0 vector is added to the motion vector candidates.

[0056] The inter prediction unit (124) derives predicted motion vector candidates using the motion vectors of these surrounding blocks, and determines a predicted motion vector for the motion vector of the current block using the predicted motion vector candidates. Then, the predicted motion vector is subtracted from the motion vector of the current block to produce a differential motion vector.

[0057] The predicted motion vector can be obtained by applying a predefined function (e.g., median, mean, etc.) to the predicted motion vector candidates. In this case, the image decoding device also knows the predefined function. In addition, since the surrounding blocks used to derive the predicted motion vector candidates are blocks that have already been encoded and decoded, the image decoding device also already knows the motion vectors of the surrounding blocks. Therefore, the image encoding device does not need to encode information for identifying the predicted motion vector candidates. Therefore, in this case, information about the differential motion vector and information about the reference picture used to predict the current block are encoded.

[0058] Alternatively, the predicted motion vector can be determined by selecting one of the predicted motion vector candidates. In this case, information for identifying the selected predicted motion vector candidate is additionally encoded, along with information about the differential motion vector and the reference picture used to predict the current block.

[0059] The subtractor (130) subtracts the prediction block generated by the intra prediction unit (122) or inter prediction unit (124) from the current block to generate a residual block.

[0060] The transformation unit (140) transforms residual signals in a residual block having pixel values ​​in a spatial domain into transform coefficients in a frequency domain. The transformation unit (140) may transform residual signals in the residual block using the entire size of the residual block as a transformation unit, or may divide the residual block into a plurality of sub-blocks and use the sub-blocks as transformation units to perform transformation. Alternatively, the residual signals may be transformed using only the transformation domain sub-block as a transformation unit by dividing the sub-blocks into two sub-blocks, that is, a transformation domain and a non-transform domain. Here, the transformation domain sub-block may be one of two rectangular blocks having a size ratio of 1:1 with respect to the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicating that only a sub-block has been converted, directionality (vertical / horizontal) information (cu_sbt_horizontal_flag), and / or position information (cu_sbt_pos_flag) are encoded by the entropy encoding unit (155) and signaled to the image decoding device. In addition, the size of the conversion area sub-block may have a size ratio of 1:3 with respect to the horizontal axis (or vertical axis), and in this case, a flag (cu_sbt_quad_flag) distinguishing the corresponding division is additionally encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0061] Meanwhile, the transformation unit (140) can individually perform transformations on the residual block in the horizontal and vertical directions. For the transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a Multiple Transform Set (MTS). The transformation unit (140) can select one transformation function pair with the best transformation efficiency among the MTS and transform the residual block in the horizontal and vertical directions, respectively. Information (mts_idx) on the transformation function pair selected among the MTS is encoded by the entropy encoding unit (155) and signaled to the image decoding device.

[0062] The quantization unit (145) quantizes the transform coefficients output from the transform unit (140) using quantization parameters and outputs the quantized transform coefficients to the entropy encoding unit (155). The quantization unit (145) may directly quantize a related residual block without transformation for a certain block or frame. The quantization unit (145) may also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients within the transform block. The quantization matrix applied to the quantized transform coefficients arranged in two dimensions may be encoded and signaled to an image decoding device.

[0063] The rearrangement unit (150) can perform rearrangement of coefficient values ​​for quantized residual values.

[0064] The reordering unit (150) can change a two-dimensional coefficient array into a one-dimensional coefficient sequence by using coefficient scanning. For example, the reordering unit (150) can output a one-dimensional coefficient sequence by scanning from the DC coefficient to the coefficients of the high-frequency region by using a zig-zag scan or a diagonal scan. Depending on the size of the transformation unit and the intra prediction mode, a vertical scan that scans the two-dimensional coefficient array in the column direction or a horizontal scan that scans the two-dimensional block-shaped coefficients in the row direction may be used instead of the zig-zag scan. That is, depending on the size of the transformation unit and the intra prediction mode, the scanning method to be used may be determined among the zig-zag scan, the diagonal scan, the vertical scan, and the horizontal scan.

[0065] The entropy encoding unit (155) generates a bitstream by encoding a sequence of one-dimensional quantized transform coefficients output from the rearrangement unit (150) using various encoding methods such as CABAC (Context-based Adaptive Binary Arithmetic Code) and Exponential Golomb.

[0066] In addition, the entropy encoding unit (155) encodes information related to block division, such as CTU size, CU division flag, QT division flag, MTT division type, and MTT division direction, so that the image decoding device can divide the block in the same manner as the image encoding device. In addition, the entropy encoding unit (155) encodes information about a prediction type indicating whether the current block is encoded by intra prediction or inter prediction, and encodes intra prediction information (i.e., information about an intra prediction mode) or inter prediction information (information about an encoding mode of motion information (merge mode or AMVP mode), a merge index in the case of a merge mode, and a reference picture index and a differential motion vector in the case of an AMVP mode) according to the prediction type. In addition, the entropy encoding unit (155) encodes information related to quantization, that is, information about a quantization parameter and information about a quantization matrix.

[0067] The inverse quantization unit (160) inversely quantizes the quantized transform coefficients output from the quantization unit (145) to generate transform coefficients. The inverse transform unit (165) transforms the transform coefficients output from the inverse quantization unit (160) from the frequency domain to the spatial domain to restore the residual block.

[0068] The addition unit (170) adds the restored residual block and the prediction block generated by the prediction unit (120) to restore the current block. The pixels within the restored current block are used as reference pixels when intra-predicting the next block.

[0069] The loop filter unit (180) performs filtering on restored pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transformation / quantization. The filter unit (180) may include all or part of a deblocking filter (182), a sample adaptive offset (SAO) filter (184), and an adaptive loop filter (ALF, 186) as an in-loop filter.

[0070] The deblocking filter (182) filters the boundaries between restored blocks to remove blocking artifacts caused by block-based encoding / decoding, and the SAO filter (184) and alf (186) perform additional filtering on the deblocking-filtered image. The SAO filter (184) and alf (186) are filters used to compensate for differences between restored pixels and original pixels caused by lossy coding. The SAO filter (184) improves not only subjective image quality but also encoding efficiency by applying an offset in units of CTUs. In contrast, the ALF (186) performs block-based filtering, and compensates for distortion by applying different filters by distinguishing the edge and degree of variation of the corresponding block. Information on filter coefficients to be used in the ALF can be encoded and signaled to an image decoding device.

[0071] The restored blocks filtered through the deblocking filter (182), SAO filter (184), and ALF (186) are stored in the memory (190). When all blocks within a picture are restored, the restored picture can be used as a reference picture for inter-predicting blocks within a picture to be encoded later.

[0072] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the techniques of the present disclosure. Hereinafter, the image decoding device and its subcomponents will be described with reference to FIG. 5.

[0073] The video decoding device may be configured to include an entropy decoding unit (510), a rearrangement unit (515), an inverse quantization unit (520), an inverse transformation unit (530), a prediction unit (540), an adder (550), a loop filter unit (560), and a memory (570).

[0074] Similar to the video encoding device of FIG. 1, each component of the video decoding device may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor executing the software functions corresponding to each component.

[0075] The entropy decoding unit (510) decodes the bitstream generated by the image encoding device to extract information related to block division, thereby determining the current block to be decoded, and extracts prediction information and information on residual signals required to restore the current block.

[0076] The entropy decoding unit (510) extracts information about the CTU size from the Sequence Parameter Set (SPS) or the Picture Parameter Set (PPS), determines the size of the CTU, and divides the picture into CTUs of the determined size. Then, the CTU is determined as the top layer of the tree structure, i.e., the root node, and the CTU is divided using the tree structure by extracting division information about the CTU.

[0077] For example, when splitting a CTU using the QTBTTT structure, first, the first flag (QT_split_flag) related to the splitting of QT is extracted, and each node is split into four nodes of the lower layer. Then, for the nodes corresponding to the leaf nodes of QT, the second flag (MTT_split_flag) related to the splitting of MTT and the split direction (vertical / horizontal) and / or split type (binary / ternary) information are extracted, and the corresponding leaf nodes are split into the MTT structure. Accordingly, each node below the leaf nodes of QT are recursively split into the BT or TT structure.

[0078] As another example, when splitting a CTU using the QTBTTT structure, the CU split flag (split_cu_flag) indicating whether the CU is split is first extracted, and if the block is split, the first flag (QT_split_flag) may be extracted. During the splitting process, each node may undergo zero or more repeated QT splits followed by zero or more repeated MTT splits. For example, a CTU may undergo an MTT split right away, or conversely, may undergo only multiple QT splits.

[0079] As another example, when splitting a CTU using the QTBT structure, the first flag (QT_split_flag) related to the splitting of QT is extracted, and each node is split into four nodes of the lower layer. Furthermore, for nodes corresponding to leaf nodes of QT, a split flag (split_flag) indicating whether to further split into BTs and splitting direction information are extracted.

[0080] Meanwhile, when the entropy decoding unit (510) determines the current block to be decoded by using the division of the tree structure, it extracts information on the prediction type indicating whether the current block is intra-predicted or inter-predicted. If the prediction type information indicates intra-prediction, the entropy decoding unit (510) extracts syntax elements for intra-prediction information (intra-prediction mode) of the current block. If the prediction type information indicates inter-prediction, the entropy decoding unit (510) extracts syntax elements for inter-prediction information, i.e., information indicating a motion vector and a reference picture referenced by the motion vector.

[0081] Additionally, the entropy decoding unit (510) extracts information about the quantized transform coefficients of the current block as information related to quantization and information about the residual signal.

[0082] The rearrangement unit (515) can change the sequence of one-dimensional quantized transform coefficients entropy-decoded in the entropy decoding unit (510) back into a two-dimensional coefficient array (i.e., block) in the reverse order of the coefficient scanning performed by the image encoding device.

[0083] The inverse quantization unit (520) inversely quantizes the quantized transform coefficients and inversely quantizes the quantized transform coefficients using the quantization parameters. The inverse quantization unit (520) may also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in two dimensions. The inverse quantization unit (520) may perform inverse quantization by applying a matrix of quantized coefficients (scaling values) from an image encoding device to a two-dimensional array of quantized transform coefficients.

[0084] The inverse transform unit (530) inversely transforms the inverse quantized transform coefficients from the frequency domain to the spatial domain to restore residual signals, thereby generating a residual block for the current block.

[0085] In addition, when the inverse transform unit (530) inversely transforms only a portion of a transform block (sub-block), it extracts a flag (cu_sbt_flag) indicating that only a sub-block of the transform block has been transformed, directionality (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block, and inversely transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to restore residual signals, and fills “0” values ​​with residual signals for areas that have not been inversely transformed, thereby generating a final residual block for the current block.

[0086] In addition, when MTS is applied, the inverse transform unit (530) determines a transform function or a transform matrix to be applied in the horizontal and vertical directions using MTS information (mts_idx) signaled from the image encoding device, and performs inverse transform on the transform coefficients within the transform block in the horizontal and vertical directions using the determined transform function.

[0087] The prediction unit (540) may include an intra prediction unit (542) and an inter prediction unit (544). The intra prediction unit (542) is activated when the prediction type of the current block is intra prediction, and the inter prediction unit (544) is activated when the prediction type of the current block is inter prediction.

[0088] The intra prediction unit (542) determines the intra prediction mode of the current block among a plurality of intra prediction modes from the syntax elements for the intra prediction mode extracted from the entropy decoding unit (510), and predicts the current block using reference pixels around the current block according to the intra prediction mode.

[0089] The inter prediction unit (544) uses the syntax elements for the inter prediction mode extracted from the entropy decoding unit (510) to determine the motion vector of the current block and the reference picture referenced by the motion vector, and predicts the current block using the motion vector and the reference picture.

[0090] An adder (550) adds a residual block output from an inverse transform unit and a predicted block output from an inter-prediction unit or an intra-prediction unit to restore the current block. The pixels within the restored current block are used as reference pixels when intra-predicting a block to be decoded later.

[0091] The loop filter unit (560) may include a deblocking filter (562), an SAO filter (564), and an ALF (566) as in-loop filters. The deblocking filter (562) deblocks the boundaries between restored blocks to remove blocking artifacts caused by block-by-block decoding. The SAO filter (564) and the ALF (566) perform additional filtering on restored blocks after deblocking filtering to compensate for differences between restored pixels and original pixels caused by lossy coding. The filter coefficients of the ALF are determined using information about filter coefficients decoded from the non-stream.

[0092] The restored blocks filtered through the deblocking filter (562), SAO filter (564), and ALF (566) are stored in the memory (570). When all blocks within a picture are restored, the restored picture is used as a reference picture for inter-predicting blocks within a picture to be encoded later.

[0093]

[0094] The present embodiment relates to encoding and decoding of images (videos) as described above. More specifically, a video coding method and device are provided for generating a prediction block of a current block using a DIMD (Decoder-side Intra Mode Derivation) technique that derives an intra prediction mode based on a gradient derived from a restored area around the current block.

[0095] The following embodiments may be performed by a prediction unit (120) within a video encoding device. Additionally, they may be performed by a prediction unit (540) within a video decoding device.

[0096] The video encoding device can generate information related to the present embodiment in terms of rate distortion optimization in encoding the current block. The video encoding device can encode the information using an entropy encoding unit (155) and then transmit it to an video decoding device. The video decoding device can decode information related to decoding the current block from a bitstream using an entropy decoding unit (510).

[0097] In the following description, the term "target block" may be used interchangeably with the current block or coding unit (CU). Alternatively, the term "target block" may also refer to a portion of a coding unit.

[0098] In the following description, the aspect ratio of a block is defined as the value obtained by dividing the horizontal length (W: Width) of the block by the vertical length (H: Height), i.e., the ratio between the horizontal length and the vertical length.

[0099] Also, a value of a flag being true indicates that the flag is set to 1. Also, a value of a flag being false indicates that the flag is set to 0.

[0100] I. Device Configuration

[0101] FIG. 6 is a block diagram illustrating in detail a portion of an image decoding device according to one embodiment of the present disclosure.

[0102] The video decoding device according to the present embodiment determines prediction and transformation units, and performs prediction and inverse transformation on the current block corresponding to the determined unit using the determined prediction technique and prediction mode, thereby finally generating a restoration block of the current block. The example illustrated in FIG. 6 may be performed by the inverse transformation unit (530), the prediction unit (540), and the adder (550) of the video decoding device. Meanwhile, the same operations as the example illustrated in FIG. 6 may be performed by the inverse transformation unit (165), the picture division unit (110), the prediction unit (120), and the adder (170) of the video encoding device. At this time, the video decoding device uses encoding information parsed from the bitstream, but the video encoding device may use encoding information set from a higher level in terms of minimizing bit rate distortion. Hereinafter, for convenience, the present embodiment will be described with reference to the video decoding device.

[0103] Referring to FIG. 6, the image decoding device according to the present embodiment includes all or part of a prediction unit (540), a transform coefficient entropy decoding unit (605), a transform coefficient inverse quantization unit (606), an inverse transform unit (530), and an adder (550). Here, as in the example of FIG. 5, the prediction unit (540) includes an intra prediction unit (542) and an inter prediction unit (544) according to a prediction technique, and as illustrated in FIG. 6, it may include all or part of a prediction unit determination unit (601), a prediction technique determination unit (602), a prediction mode determination unit (603), and a prediction execution unit (604). In addition, the inverse transform unit (530) may include all or part of an inverse transform unit determination unit (607), an inverse transform kernel determination unit (608), and an inverse transform execution unit (609).

[0104] The prediction unit determination unit (601) determines a prediction unit (PU). Here, the prediction unit can be a current block or one of the sub-blocks into which the current block is divided. The prediction technique determination unit (602) determines a prediction technique (e.g., inter-prediction (inter-screen prediction), intra-prediction (within-screen prediction), IBC (Intra Block Copy), and palette) for each prediction unit determined by the prediction unit determination unit (601). The prediction mode determination unit (603) determines a detailed prediction mode for the prediction technique. The prediction execution unit (604) generates a prediction block of the current block according to the determined prediction mode.

[0105] Detailed prediction modes in inter prediction technology may include, for example, the following modes.

[0106] - A mode that generates a prediction signal of the current block as a weighted sum of at least one prediction signal generated through the motion information of the current block and motion compensation.

[0107] - A mode that divides the current block into one or more sub-regions through geometric division, and generates a prediction signal of the current block through a weighted sum of at least one prediction signal generated through motion information and motion compensation for each sub-region.

[0108] Detailed prediction modes in intra prediction technology may include, for example, the following modes.

[0109] - A mode that generates a prediction signal for the current block using one directional prediction mode, planar mode (e.g. Horizontal Planar, Vertical Planar, or Regular Planar), or DC mode.

[0110] - A mode that generates a prediction signal for the current block using a predefined matrix.

[0111] - A mode that defines the restored area around the current block as a template and generates a prediction signal for the current block by performing matching between the defined template and the restored area.

[0112] - A mode that defines a restoration area around the current block as a template, derives at least one directional prediction mode from the template, and then generates a prediction signal for the current block using the derived directional prediction mode.

[0113] Meanwhile, the prediction unit (540) can generate a prediction block of the current block using a mixed mode that simultaneously uses a prediction mode that can be included in the inter prediction technology and a prediction mode that can be included in the intra prediction technology, and the following modes can be included in such a mixed mode as examples.

[0114] - A mode that generates a final prediction signal for the current block by weighting the prediction signal for the current block generated using the intra prediction mode and the prediction signal for the current block generated using the inter prediction mode.

[0115] - A mode that divides the current block into one or more sub-regions through geometric division, and generates the final prediction signal of the current block through a weighted sum of the prediction signal generated using the inter prediction mode for some sub-regions and the prediction signal generated using the intra prediction mode for the remaining sub-regions.

[0116] The transform coefficient entropy decoding unit (605) restores the quantized transform coefficient from the bitstream. Here, the transform coefficient entropy decoding unit (605) can restore the quantized second transform coefficient if a second transform is applied to the transform coefficient, and can perform restoration of the quantized first transform coefficient if the second transform is not applied, i.e., only the first transform is applied.

[0117] The transform coefficient inverse quantization unit (606) parses information such as quantization method and quantization parameter information for the transform coefficient restored by the transform coefficient entropy decoding unit (605), and performs inverse quantization on the restored transform coefficient.

[0118] The inverse transform unit determination unit (607) determines the inverse transform unit for the transform coefficients. The inverse transform unit may be the same as the transform unit (TU), and may be set to have the same size as the current block or one of the sub-blocks into which the current block is divided. The inverse transform kernel determination unit (608) determines a separable vertical and horizontal first-order inverse transform kernel and / or a non-separable second-order inverse transform kernel, or a non-separable first-order inverse transform kernel, as a kernel for inverse transform of the transform coefficients. The inverse transform performing unit (609) performs inverse transform on the inverse quantized transform coefficients using the inverse transform kernel determined by the inverse transform kernel determination unit (608) to restore residual signals. Meanwhile, whether to perform the non-separable first-order inverse transform and whether to perform the non-separable second-order inverse transform may be explicitly determined by parsing signaled information, or may be implicitly determined based on the size of the current TU, etc.

[0119] The adder (550) adds the residual signals restored by the inverse transform unit (530) and the final prediction signal generated from the prediction unit (540) to generate a final restoration signal for the current block. The final restoration signal for the current block, i.e., the restoration block, is stored in memory and can be used to predict other blocks thereafter.

[0120] As described above, the prediction technology of the current block can be determined in the prediction technology determination unit (602). The prediction technology can be one of the following technologies: inter prediction, intra prediction, IBC, palette, etc.

[0121] According to one embodiment of determining the prediction technique or prediction mode of the current block, the prediction unit (540) determines whether intra prediction is applied to the current block. If intra prediction is not applied to the current block, the prediction unit (540) parses a flag indicating whether the mode of the current block is skip mode. If the flag indicates that the prediction mode of the current block corresponds to skip mode, the prediction unit (540) can determine the prediction mode of the current block as inter merge mode or IBC merge mode. If the prediction mode of the current block corresponds to skip mode, the inverse transformation process of the transform coefficients can be omitted, and the prediction block generated from the prediction unit (540) can be used as a restored block.

[0122] According to another embodiment of determining the prediction technique or prediction mode of the current block, the prediction unit (540) determines whether the prediction mode of the current block corresponds to the skip mode. If the prediction mode of the current block does not correspond to the skip mode, the prediction unit (540) can parse a flag indicating the prediction technique of the current block and predict the current block using one of the techniques such as inter prediction, intra prediction, IBC, and palette.

[0123] According to another embodiment of determining the prediction technique or prediction mode of the current block, if the prediction mode of the current block does not correspond to the skip mode and inter prediction or IBC technique is applied to the current block, the prediction unit (540) can determine whether to perform prediction on the current block in merge mode or AMVP mode by parsing a flag indicating merge or AMVP mode.

[0124] According to another embodiment of determining the prediction technique or prediction mode of the current block, the prediction unit (540) determines whether an intra prediction technique is applied to the current block. If the intra prediction technique is applied to the current block, the prediction unit (540) can predict the current block based on the prediction mode indicated by information parsed from a flag and / or index indicating one of various prediction modes. Herein, the indicated prediction mode may be a prediction mode derivation technique or mode (hereinafter referred to as prediction mode derivation technique or DIMD (Decoder side Intra Mode Derivation) technique). Hereinafter, a process of generating a prediction block or prediction signals of the current block using the prediction mode derivation technique according to the present disclosure will be described in detail using drawings.

[0125]

[0126] II. DIMD (Decoder side Intra Mode Derivation)

[0127] The DIMD technique is a technique for generating a final prediction signal of a current block based on a gradient calculated for at least a portion of a previously restored area around the current block. Specifically, the DIMD technique is a technique for calculating at least one gradient for an 'L'-shaped area on the left and top in a previously restored area around the current block having a size of W×H, and generating a final prediction signal using at least one directional and / or non-directional prediction mode selected based on the calculated gradient.

[0128] In general, since DIMD technology can require a lot of computational power from the processor, its application may be limited depending on the size of the current block or the slice type. For example, DIMD technology can be applied to the current block only when the size of the current block (width × height) is greater than or equal to 128 and less than or equal to 1024, or DIMD technology can be applied to the current block only when the current slice type is I slice.

[0129] A video decoding device can parse a flag and / or an index from a bitstream that indicates whether a prediction mode derivation technique is applied to a current block. Here, the flag and / or index can be parsed after considering the aforementioned DIMD application conditions. Thereafter, the video decoding device can determine whether a prediction mode indicated by the parsed flag and / or index is a prediction mode derivation technique. If the prediction mode indicated by the parsed flag and / or index is a prediction mode derivation technique, the video decoding device generates a prediction block of the current block using the prediction mode derivation technique. On the other hand, if the prediction mode indicated by the parsed flag and / or index is not a prediction mode derivation technique, the video decoding device generates a prediction block of the current block using another indicated prediction mode that is not a prediction mode derivation technique. Here, if the prediction mode indicated by the parsed flag and / or index is a prediction mode derivation technique, the decoding device may omit parsing of the flag and / or index indicating whether other modes such as BDPCM, MIP, and MRL modes are applied.

[0130] Hereinafter, various embodiments of processes performed by an image decoding device to generate a prediction block of a current block using a prediction mode derivation technique are described in detail.

[0131] 1. Setting the gradient calculation area

[0132] As described above, if the prediction mode indicated by the parsed flag and / or index is a prediction mode derivation technique, the image decoding device generates a prediction block of the current block using the prediction mode derivation technique. To this end, the prediction unit (540) can set a gradient calculation area for gradient calculation from surrounding restoration samples of the current block.

[0133] FIG. 7 is an exemplary diagram showing a gradient calculation area used for calculating gradient values ​​according to one embodiment of the present disclosure.

[0134] According to one embodiment, the restored reference samples or reference pixels located in the upper n rows and the left m columns relative to the current block among the restored regions may be included in the gradient calculation region. Here, n and m may be integers greater than or equal to 2. If some or all of the upper n rows or the left m columns are unavailable, only the available region may be included in the gradient calculation region.

[0135] In some embodiments, m and n may be determined based on the size or area of ​​the current block (the total number of samples in the current block). For example, if the size or area of ​​the current block is smaller than a predefined size or equal to a predefined small block size (e.g., 4x4, 8x4, 4x8, etc.), m and n may be set to a first value, and otherwise, they may be set to a second value. Here, the second value may be a value greater than the first value. For example, the first value may be 2 and the second value may be 3.

[0136] In some other embodiments, m and n may be determined differently based on the width (W) and height (H) of the current block. That is, m may be determined based on the width of the current block, and n may be determined based on the height of the current block.

[0137] For example, if the width of the current block is less than or equal to a predefined threshold, m may be set to a first value, otherwise m may be set to a second value greater than the first value. Additionally, if the height of the current block is less than or equal to a predefined threshold, n may be set to the first value, otherwise n may be set to a second value. For example, the predefined threshold may be 8, the first value may be 2, and the second value may be 3.

[0138] Assuming that the top n rows and the left m columns are available, the length M of the top reference samples and the length N of the left reference samples can be determined based on the width (W) and height (H) of the current block. For example, assuming that the top left pixel position of the current block is (0, 0), based on the top left sample position (-m, -n) of the gradient calculation region, M can be set to the value of W+m+b, and N can be determined to the value of H+n+a. Here, n, m, M, and N are all integers greater than or equal to 1, and a and b can be integer values ​​determined depending on whether the restored region is available. Specifically, if the m×a region (the bottom left region of the current block) is restored based on FIG. 7 and can be used as a reference sample, a can be an integer greater than or equal to 1 and less than or equal to an integer multiple of H. On the other hand, if the m×a region cannot be used as a reference sample, a can be determined to the value 0. Similarly, if the b×n (upper right area of ​​the current block) area based on Fig. 7 is restored and can be used as a reference sample, b can be an integer having a value greater than or equal to 1 and less than or equal to an integer multiple of W, and if the b×n area cannot be used as a reference sample, b can be determined as a value of 0.

[0139] According to another embodiment, the gradient calculation region can be determined to include at least a portion of the left and upper regions relative to the current block by parsing a flag or index indicating the region to be included in the gradient calculation region from the bitstream.

[0140] In another embodiment, the gradient calculation region may be implicitly determined based on an agreement between the image encoding device and the image decoding device. In this case, the gradient calculation region determination operation described above may be omitted.

[0141] 2. Gradient calculation

[0142] According to the above-described embodiments, if the gradient calculation region has been determined, the prediction unit (540) can calculate gradients for at least some regions within the gradient calculation region. That is, the gradient can be calculated in units of at least some regions. Here, at least some regions can be set in units of p×q blocks, and hereinafter, at least one region of the p×q block units is referred to as a sampling unit. The width p and the height q of the sampling unit can be adaptively determined according to the size of the current block, the resolution of the image, the aspect ratio of the current block, etc.

[0143] In one embodiment, the sizes p and q of the sampling units may be determined based on the size or area of ​​the current block. For example, if the size or area of ​​the current block is smaller than a predefined size or equal to a predefined small block size (e.g., 4x4, 8x4, 4x8, etc.), p and q may be set to a first value, and otherwise, they may be set to a second value. Here, the second value is larger than the first value. For example, the first value may be 2 and the second value may be 3.

[0144] In another embodiment, p and q may be determined differently based on the width (W) and height (H) of the current block. That is, p may be determined based on the width of the current block, and q may be determined based on the height of the current block.

[0145] For example, if the width of the current block is less than or equal to a predefined threshold, p may be set to a first value, otherwise p may be set to a second value greater than the first value. Additionally, if the height of the current block is less than or equal to the predefined threshold, q may be set to the first value, otherwise q may be set to a second value. For example, the predefined threshold may be 8, the first value may be 2, and the second value may be 3.

[0146] Below, for convenience of explanation, an example with a sampling unit of 3×3 is provided.

[0147] Prior to gradient estimation, if the value of m is less than p or the value of n is less than q, the prediction unit (540) may pad the outermost part of the gradient estimation region so that the value of m or n that is less than p or q has a value greater than or equal to p or q. That is, the prediction unit (540) may pad the outermost part of the gradient estimation region so that the sampling unit does not go outside the gradient estimation region. For example, if m and n have a value of 2, the prediction unit (540) performs padding by copying the sample values ​​corresponding to the positions (-m,-n) to (-m,H+a-1) with respect to the current block to the positions (-m-1,-n) to (-m-1,H+a-1), and copying the sample values ​​corresponding to the positions (-m,-n) to (W+b-1,-n) to the positions (-m,-n-1) to (W+b-1,-n-1), thereby finally making m and n have a value of 3. In other words, the gradient calculation region can be expanded by padding to a size that can accommodate the sampling unit.

[0148] Thereafter, the prediction unit (540) can calculate the gradient along with preprocessing for the gradient calculation region. Specifically, the prediction unit (540) can determine whether to perform preprocessing based on a flag and / or index indicating whether a filter has been applied to the gradient calculation region, and then calculate the gradient. The filter can be applied to each sampling unit, to each region defined separately from the sampling unit within the gradient calculation region, or can be applied to the entire gradient calculation region at once.

[0149] If the flag and / or index indicating whether a filter is applied to the gradient calculation region indicates that a filter is applied, the prediction unit (540) can calculate a gradient by applying a filter to the gradient calculation region and then applying a gradient operator to each sampling unit included in the filtered gradient calculation region. Here, the filter to be applied to the gradient calculation region may be a smoothing filter expressed as a box filter or an averaging filter, or a sharpening filter expressed as a Laplacian filter or an unsharp mask filter. In addition, an edge detection filter such as a Sobel filter or a Prewitt filter may be used as the gradient operator.

[0150] In contrast, if the flag and / or index indicating whether a filter is applied to the gradient calculation region indicates that the filter is not applied, the prediction unit (540) may not apply the filter to the gradient calculation region, but may calculate the gradient by applying a gradient operator to each sampling unit included in the gradient calculation region. Here, as the gradient operator, an edge detection filter such as a Sobel filter or a Prewitt filter may be used, as described above.

[0151] Meanwhile, the configuration of sampling units for applying the gradient operator can be performed in various ways.

[0152] FIGS. 8A to 8C are diagrams illustrating various embodiments of a method for configuring sampling units for applying a gradient operator.

[0153] FIG. 8a is a diagram illustrating one embodiment of a method for configuring sampling units for applying a gradient operator.

[0154] Referring to FIG. 8A, a prediction unit (540) according to one embodiment configures sampling units continuously and overlappingly for the entire gradient calculation region, and calculates gradients by applying a gradient operator to the configured sampling units. In other words, the sampling units may be configured so that the center of each sampling unit is continuously arranged with the center of an adjacent sampling unit. Describing with reference to FIG. 8A, each sampling unit may be configured so that the center is (-2,-2) to (-2,H-1) and (-2,-2) to (W-1,-2) with respect to the current block.

[0155] The application of the gradient operator can be performed continuously and overlappingly over the entire gradient calculation region, as illustrated in Fig. 8a, but can also be performed over a portion of the gradient calculation region, or can be performed discontinuously and / or non-overlappingly.

[0156] FIG. 8b is a diagram illustrating another embodiment of a method for configuring sampling units for applying a gradient operator.

[0157] Referring to FIG. 8B, a prediction unit (540) according to another embodiment configures sampling units for a portion of a gradient calculation region, and calculates gradients by applying a gradient operator to the configured sampling units. Specifically, the center of each sampling unit may be configured to have a distance of 2 pixels from the center of an adjacent sampling unit. Referring to FIG. 8B, each sampling unit may be configured to have centers at (-2,H-2), (-2,H-4), (-2,H-6), ... and (W-2,-2), (W-4,-2), (W-6,-2), ... with respect to the current block. In addition, the centers of each sampling unit may be configured to have a distance of x pixels as well as a distance of 2 pixels from each other. Here, x may be adaptively determined depending on the size of the current block, the size of the sampling unit, etc.

[0158] Meanwhile, the sampling unit may be configured based on the partition information of the restored blocks around the current block.

[0159] FIG. 8c is a diagram illustrating another embodiment of a method for configuring sampling units for applying a gradient operator.

[0160] The dotted lines illustrated in Fig. 8c correspond to the boundaries of adjacent blocks adjacent to the current block. Referring to Fig. 8c, a prediction unit (540) according to another embodiment configures a sampling unit based on the segmentation information of previously restored blocks around the current block, and calculates gradients by applying a gradient operator to the configured sampling units. Specifically, the prediction unit (540) may not calculate a gradient for a certain pixel area based on the boundary where block segmentation is performed, but may calculate a gradient only for an area that is not adjacent to the block segmentation boundary. Referring to Fig. 8c, a sampling unit may be configured to be centered on samples spaced at least 2 pixels from the boundary of an adjacent block of the current block, excluding samples that touch the boundary of an adjacent block of the current block. Unlike Fig. 8c, a sampling unit may also be configured to be centered on samples spaced at least x pixels from the boundary of an adjacent block, where x may be adaptively determined depending on the size of the current block, the size of the sampling unit, etc.

[0161] 3. Derivation of available modes

[0162] The prediction unit (540) may calculate a gradient for each sampling unit based on the above-described methods, and derive at least one usable mode for predicting the current block based on the calculated gradient. To derive at least one usable mode, the prediction unit (540) may calculate a gradient for each sampling unit, and derive the usable modes using information based on the calculated gradient. The information based on the calculated gradient may be configured to include information on a directional mode corresponding to or mapped to the calculated gradient and information on the strength of the gradient. Here, the calculated gradient may correspond to at least one of the intra directional modes based on its directionality, and in some embodiments, at least one or more of an aspect ratio of the current block, a range of a gradient calculation region, or a position of a sampling unit may be additionally considered. Additionally, the strength of the gradient can be determined, for example, as the absolute value of the x-direction gradient (i.e., horizontal gradient), the absolute value of the y-direction gradient (i.e., vertical gradient), or the sum of the absolute values ​​of the x-direction gradient and the y-direction gradient. Alternatively, the gradient strength can be based on the sum of the squares of the x-direction gradient and the squares of the y-direction gradient.

[0163] According to one embodiment of associating a gradient with a directional mode, the calculated gradient may be associated with one of the intra directional modes based on its directionality, and additionally, the directional modes that may be associated with the gradient may be restricted based on the aspect ratio of the current block. Specifically, if the calculated gradient for a specific sampling unit corresponds to a restricted directional mode, the mode corresponding to the calculated gradient may be replaced with another unrestricted directional mode.

[0164] Table 1 shows an example of directional modes that are restricted according to the aspect ratio of the current block.

[0165]

[0166] The directional modes listed in Table 1 correspond to the directional modes illustrated in Fig. 3b. Referring to Table 1, an example of matching the calculated gradient with the directional mode, taking into account the aspect ratio of the current block, is exemplarily described as follows.

[0167] When the aspect ratio of the current block is 16 / 1, intra prediction modes 2 to 15 are restricted, and accordingly, the gradient calculated for each sampling unit can correspond to one of intra prediction modes 16 to 80. For example, when the aspect ratio of the current block is 16 / 1, even if the gradient calculated for a specific sampling unit corresponds to intra prediction mode 15, the mode corresponding to the gradient for the sampling unit can be mapped to intra prediction mode 80. In addition, when the aspect ratio of the current block is 1 / 1, the gradient calculated for each sampling unit can be mapped to correspond to one of intra prediction modes 2 to 66. In the same way, the mode corresponding to the gradient can be mapped to one of the intra prediction modes 14 to 78 when the aspect ratio is 8 / 1, one of the intra prediction modes 12 to 76 when the aspect ratio is 4 / 1, one of the intra prediction modes 8 to 72 when the aspect ratio is 2 / 1, one of the intra prediction modes -6 to -1 and 2 to 60 when the aspect ratio is 1 / 2, one of the intra prediction modes -10 to -1 and 2 to 56 when the aspect ratio is 1 / 4, one of the intra prediction modes -12 to -1 and 2 to 54 when the aspect ratio is 1 / 8, and one of the intra prediction modes -14 to -1 and 2 to 52 when the aspect ratio is 1 / 16.

[0168] If the directional mode A derived through the grain calculation corresponds to a limited mode as illustrated in Table 1, the directional mode A can be replaced by another mode. For example, the directional mode A can be mapped to another mode pointing in the opposite direction through the following relationship.

[0169] (1) When the aspect ratio (W / H) is greater than 1, directional mode A is replaced with directional mode (A+65).

[0170] (2) When the aspect ratio (W / H) is less than 1, directional mode A is replaced by directional mode (A-67).

[0171]

[0172] According to another embodiment of associating gradients with directional modes, the calculated gradients may be associated with intra-directional modes based on their directionality, and additionally, the directional modes that may be associated with the gradients may be restricted depending on the range of the gradient calculation region. According to this embodiment, if the gradients calculated for a specific sampling unit correspond to a restricted directional mode, the gradients calculated for the sampling unit may be associated with other unrestricted modes or may be excluded from the process of deriving usable modes.

[0173] Referring again to Figure 7, another embodiment of matching the calculated gradient with the directional mode, a method that takes into account the range of the gradient calculation area, is exemplarily described as follows.

[0174] Depending on the length b of the gradient calculation region located at the upper right of the current block, the available intra prediction modes may be determined differently. For example, when b is greater than or equal to W, modes 67 to 80 may be determined as available modes. In other words, the use of modes 2 to 15 is restricted. When the intra prediction mode corresponding to the direction of the gradient is included in the range of 2 to 15, the mode is replaced with one of modes 67 to 80 according to the above-described relationship. When b is greater than 0 and less than W, the available modes among modes 67 to 80 may be determined in proportion to the available modes when b = 0 and the available modes when b = W. As another example, the available modes may be predetermined according to the value of b by a convention of the encoder / decoder.

[0175] Similarly, the available intra prediction modes may be determined differently depending on the length a of the gradient calculation region located at the lower left of the current block. For example, when a is greater than or equal to H, modes -1 to -14 may be determined as available modes. In other words, the use of modes 53 to 66 is restricted. When the intra prediction mode corresponding to the direction of the gradient is included in the range of modes 53 to 66, the mode is replaced with one of modes -1 to -14 according to the relationship described above. On the other hand, when a is greater than 0 and less than H, the available modes among modes -1 to -14 may be determined in proportion to the available modes when a = 0 and the available modes when a = H. As another example, the available modes may be predetermined by a convention of the encoder / decoder depending on the value of a.

[0176] According to another embodiment of matching gradients and directional modes, the generated gradients can be matched with intra directional modes based on their directionality, and additionally, the directional modes that can be matched with the gradients can be restricted by considering both the aspect ratio of the current block and the range of the gradient calculation region. For example, if the aspect ratio of the current block is 16 / 1 and a is equal to W / 2, modes other than intra prediction modes 16 to 72 can be restricted. The restricted modes here can be replaced with other modes that are not restricted according to the above-described relationship, for example, or can be excluded in the process of deriving usable modes.

[0177] According to another embodiment of associating gradients and directional modes, directional modes that can be mapped to a corresponding sampling unit depending on the location of the sampling unit may be pre-specified according to an agreement between an encoding device and a decoding device.

[0178] Figure 9 is a diagram showing an example of a directional mode corresponding to a gradient for some sampling units.

[0179] In some embodiments, if the gradient for some sampling units corresponds to a specific directional mode, the gradient for the sampling units may be excluded from the derivation of available modes. For example, as shown in FIG. 9, if the gradient for a sampling unit located in the left gradient calculation region corresponds to a vertical prediction mode, i.e., intra prediction mode 50 based on FIG. 3B, the sampling unit may not be considered in the process of predicting the current block. In other words, if the intra prediction mode derived from the sampling unit located in the left gradient calculation region is a vertical prediction mode, the vertical prediction mode may be ignored. Similarly, as shown in FIG. 9, if the gradient for a sampling unit located in the upper gradient calculation region corresponds to a horizontal prediction mode, i.e., intra prediction mode 18 based on FIG. 3B, the sampling unit may not be considered in the process of predicting the current block.

[0180] 4. Selection of representative mode

[0181] The prediction unit (540) may perform a process of selecting at least one representative mode to be applied to the current block simultaneously with the process of deriving the modes available in the sampling unit, or after the modes available in the sampling unit are deriving. The selection of at least one representative mode may be performed based on a histogram calculated by aggregating the available modes determined in each sampling unit, for example. Specifically, the selection of at least one representative mode may be performed as a process of constructing a table or a histogram (hereinafter, the terms 'table' and 'histogram' may be used interchangeably) by aggregating the available modes derived in each sampling unit and / or the gradient strength information for each sampling unit, deriving at least one dominant mode from the constructed table, and then selecting at least one dominant mode as at least one representative mode. Here, if the dominant mode is derived using the gradient strength information for each sampling unit, the histogram may be calculated by recording the number of occurrences of the available modes derived from each sampling unit with a weight applied. Specifically, for each mode derived from each sampling unit, a table can be constructed that weights the modes based on gradient strength, rather than simply recording the number of occurrences. For example, a histogram can be generated by accumulating the modes derived from each sampling unit by the gradient strength.

[0182] According to one embodiment, the table or histogram can be constructed by aggregating the available modes derived from each sampling unit and / or the gradient strength information for each sampling unit for all gradient estimation regions during the histogram construction process.

[0183] Alternatively, the histogram construction process may be performed only until the total accumulated gradient intensity value on the histogram reaches a predefined threshold while constructing the histogram. That is, if the accumulated gradient intensity value is greater than or equal to the predefined threshold, the table or histogram construction process of the current block may be terminated early. To this end, mode information and intensity information may be constructed for each sampling unit in the table or histogram of the current block, and the intensity information may be accumulated and managed for each sampling unit. Depending on the embodiment, this process may be performed separately for the upper region and the left region of the current block among the gradient calculation regions.

[0184] In some embodiments, the accumulated gradient strength value may have different values ​​depending on the size of the current block, and a table or histogram early termination process may be performed depending on the size of the current block.

[0185] In another embodiment, early termination may be determined based on the size or area of ​​the current block. For example, early termination may only be performed if the area of ​​the current block is greater than or equal to a predefined threshold. The predefined threshold may be 128.

[0186] Meanwhile, in order to aggregate the available modes, the gradient calculation region can be defined by dividing it into multiple regions according to various embodiments.

[0187] Figures 10a to 10c are drawings showing various examples of dividing a gradient calculation area into multiple areas.

[0188] Figure 10a is a drawing showing an example of dividing a gradient calculation area into multiple areas.

[0189] According to an example of dividing the gradient calculation area into multiple areas, if a and b have values ​​of 0 based on Fig. 7, the gradient calculation area can be divided into a total of three areas, and each area can be defined as a left, upper, and upper left area based on the current block. Referring to Fig. 10a, the left area corresponds to area B, the upper area corresponds to area C, and the upper left area corresponds to area A.

[0190] Figure 10b is a diagram showing another example of dividing the gradient calculation area into multiple areas.

[0191] According to another example of dividing the gradient calculation area into multiple areas, if a and b have values ​​of 0 based on FIG. 7, the gradient calculation area can be defined by dividing it into a total of three areas similar to the example described above. However, unlike the example described above, as can be seen in FIG. 10b, the three areas can be defined based on the extension of the line that divides the height and width of the current block in half.

[0192] Figure 10c is a diagram showing another example of dividing the gradient calculation area into multiple areas.

[0193] According to another example of dividing the gradient calculation area into multiple areas, if a and / or b have non-zero values ​​based on FIG. 7, areas corresponding to areas different from the examples described above can be additionally defined. For example, as can be seen in FIG. 10c, if both a and b have non-zero values ​​based on FIG. 7, areas in the gradient calculation area that do not touch the current block, i.e., the lower left gradient area and the upper right gradient area of ​​the current block, can be additionally defined as the X area and the Y area. Furthermore, if only one of a and b has a non-zero value based on FIG. 7, areas in the gradient calculation area that do not touch the current block can be additionally defined as the X area or the Y area. In the case of areas in the gradient calculation area that touch the current block, areas may be defined in a manner similar or identical to the examples described above, or may be defined in a manner dissimilar to the examples described above, such as defining the entire area touching the current block as one area.

[0194] The method of dividing the gradient calculation area into multiple areas is not limited to the examples described above. For example, in the examples illustrated in FIG. 10a or FIG. 10b, areas A and B or C may be defined as a single area, or the gradient calculation area may be divided into multiple areas based on the boundaries of restored blocks adjacent to the current block, and some adjacent areas among the multiple areas divided from the gradient calculation area may be defined as a single area.

[0195] In some embodiments, in the process of aggregating available modes, the prediction unit (540) may aggregate only a selected portion of directional modes, rather than aggregate all directional modes. Here, the selection of some directional modes may be performed differently for each of a plurality of regions divided from the gradient calculation region. In other words, the gradient calculation region may be divided into a plurality of regions, and the prediction unit (540) aggregates the available modes to derive a dominant directional mode for each of the plurality of regions. However, the available modes that can be aggregated may be limited depending on the location of each region with respect to the current block.

[0196] As an example of aggregating only for some directional modes, when multiple regions are partitioned as illustrated in FIG. 10a, only directional modes corresponding to an angular range of 90° to 180° may be recorded for region A, only directional modes corresponding to an angular range of 135° to 225° may be recorded for region B, and only directional modes corresponding to an angular range of 45° to 135° may be recorded for region C. In this example, when region X and / or region Y are additionally partitioned, that is, when multiple regions are partitioned in a form in which the examples illustrated in FIG. 10a and FIG. 10c are combined, only directional modes corresponding to an angular range of 225° to 270° may be recorded for region X, and only directional modes corresponding to an angular range of 0° to 45° may be recorded for region Y. Here, the illustrated angles may correspond to the intra prediction modes illustrated in FIG. 3b. Specifically, the illustrated angles may correspond to the angles formed counterclockwise by a line vertically connecting the center of the rectangle illustrated in FIG. 3b and the right side of the rectangle, and lines indicating intra prediction modes.

[0197] In another example of aggregating only for some directional modes, when multiple regions are partitioned as illustrated in FIG. 10b, only directional modes corresponding to an angular range of 45° to 225° may be recorded for region A, only directional modes corresponding to an angular range of 135° to 225° may be recorded for region B, and only directional modes corresponding to an angular range of 45° to 135° may be recorded for region C. In this example, when region X and / or region Y are additionally partitioned, i.e., when multiple regions are partitioned in a form in which the examples illustrated in FIG. 10b and FIG. 10c are combined, only directional modes corresponding to an angular range of 180° to 270° may be recorded for region X, and only directional modes corresponding to an angular range of 0° to 90° may be recorded for region Y. Here, the angles illustrated may correspond to the intra prediction modes illustrated in FIG. 3b as described above.

[0198] The method of aggregating only certain directional modes is not limited to the examples described above. For example, the selection of certain directional modes for aggregation may be adaptively performed based on at least one of the aspect ratio of the current block, the size or range of the gradient calculation region, and the position or size of each of the multiple regions segmented from the gradient calculation region. Alternatively, the selection of certain directional modes for aggregation may be performed based on information provided by the video encoding device, such as a flag or index.

[0199] In some embodiments, during the process of aggregating the available modes, the prediction unit (540) may additionally aggregate directional modes having a vertical orientation if the horizontal slope for some sampling units is 0 or below a specific threshold, i.e., if the vertical slope is dominant. Similarly, during the process of aggregating the available modes, the prediction unit (540) may additionally aggregate directional modes having a horizontal orientation if the vertical slope for some sampling units is 0 or below a specific threshold, i.e., if the horizontal slope for some sampling units is 0 or below a specific threshold, i.e., if the horizontal slope for some sampling units is dominant. In addition, when the horizontal slope for a sampling unit located on the left side of the current block is 0 or below a specific threshold, the directional mode that can correspond to the corresponding sampling unit is likely to be a vertical directional mode, and therefore, the available modes determined for the corresponding sampling unit may not be aggregated. Similarly, when the vertical slope for a sampling unit located on the top side of the current block is 0 or below a specific threshold, the available modes determined for the corresponding sampling unit may not be aggregated. Here, the specific threshold value can be set as a fixed value or as the average or median luminance value of the gradient calculation area.

[0200] According to some embodiments, the prediction unit (540) may individually aggregate directional modes for each of a plurality of regions divided from the gradient calculation region, perform normalization for each aggregate, and then sum the normalized data, and select at least one representative mode based on the summed data.

[0201] Alternatively, the representative mode may be selected from each table (histogram) for multiple regions. Specifically, one or more of the most dominant modes in the table for each region may be selected as the representative mode.

[0202] As another example, the prediction unit (540) may aggregate directional modes at once for the entire gradient calculation area and select at least one representative mode based on the aggregated directional modes.

[0203] 5. Generating the final prediction signal

[0204] If at least one representative mode is selected, the prediction unit (540) can generate a final prediction signal for the current block using the selected representative mode.

[0205] According to one embodiment of generating a final prediction signal for a current block using a representative mode, based on a table constructed according to the aforementioned processes, the prediction unit (540) may select n dominant directional modes (n is an integer, hereinafter the same) as representative modes. Thereafter, the prediction unit (540) may generate a prediction signal for the current block using each representative mode, and weight and sum them to generate a final prediction signal for the current block. Here, the number of n or the maximum value of n may be provided as a fixed value, or may be adaptively determined according to the generated table. Alternatively, only modes aggregated above a specific threshold value in the table may be selected as representative modes. Alternatively, the number of n may be determined differently based on the size or area of ​​the current block. As an example, when the area of ​​the current block is greater than or equal to a predefined threshold value (e.g., 128), the value of n may be determined to be greater than otherwise. For example, when the area of ​​the current block is greater than or equal to the predefined threshold value, n may be set to a first value. On the other hand, if the area of ​​the current block is less than a predefined threshold, n may be set to a second value that is less than the first value.

[0206] According to another embodiment of generating a final prediction signal for the current block using a representative mode, the prediction unit (540) selects n dominant directional modes as representative modes, and then weights and adds the prediction signal generated using each representative mode and the prediction signal generated using the planar mode to generate a final prediction signal for the current block.

[0207] According to another embodiment of generating a final prediction signal for the current block using the representative mode, if the vertical direction mode is the most dominant among n representative modes according to the generated table, the prediction unit (540) can generate a final prediction signal for the current block by weighting the prediction signal generated using each representative mode and the prediction signal generated using the vertical planar mode. Alternatively, in this case, the prediction unit (540) can generate a final prediction signal for the current block by weighting the prediction signal generated using each representative mode and the prediction signal generated using the planar mode and the vertical planar mode.

[0208] According to another embodiment of generating a final prediction signal for the current block using a representative mode, if the horizontal directional mode is the most dominant among n representative modes according to the generated table, the prediction unit (540) can generate a final prediction signal for the current block by weighting the prediction signal generated using each representative mode and the prediction signal generated using the horizontal planar mode. Alternatively, in this case, the prediction unit (540) can generate a final prediction signal for the current block by weighting the prediction signal generated using each representative mode and the prediction signal generated using the planar mode and the horizontal planar mode.

[0209] The vertical planar mode is a mode that predicts the target sample by weighting the values ​​of the upper and lower surrounding samples that are located in the same row as the target sample in the current block and adjacent to the current block. The horizontal planar mode is a mode that predicts the target sample by weighting the values ​​of the left and right surrounding samples that are located in the same row as the target sample in the current block and adjacent to the current block. The planar mode is a mode that predicts the target sample by weighting the two surrounding samples used in the vertical planar mode and the two surrounding samples used in the horizontal planar mode.

[0210] Here, the right and bottom peripheral samples of the current block may be unrestored samples. In this case, the right peripheral sample may be replaced with the peripheral sample located in the upper right corner of the current block, and the bottom peripheral sample may be replaced with the peripheral sample located in the lower left corner of the current block.

[0211] In one embodiment of weighting multiple prediction signals, the weights may be provided as fixed values ​​having the same weight for all modes.

[0212] In another embodiment, the weights for each representative mode may be determined based on the histogram size (cumulative value of gradient intensities) of that representative mode. For example, a mode with a larger histogram size may be assigned a greater weight. In other words, a prediction block predicted by a mode with a larger histogram size may be assigned a greater weight.

[0213] As an alternative or additional embodiment, the weights to be applied to samples within a prediction block predicted by the corresponding representative mode may be determined differently depending on the sample location. For example, the prediction unit (540) checks the histogram size for a specific representative mode among the histograms for each of a plurality of regions and compares the sizes. Then, samples closer to regions with larger histogram sizes within the prediction block predicted by the corresponding representative mode are assigned greater weights.

[0214] For example, if the histogram size of the corresponding representative mode in the upper gradient region is larger than the histogram size of the corresponding representative mode in the left gradient region by a certain percentage or more, a larger weight is assigned to samples closer to the upper boundary of the prediction block among the prediction samples within the prediction block predicted by the corresponding representative mode. Conversely, in the opposite case, a larger weight is assigned to samples closer to the left boundary among the prediction samples within the prediction block predicted by the corresponding representative mode.

[0215] In another embodiment of weighting multiple prediction signals, the weights can be implicitly and adaptively provided using at least one of the height and width of the current block, the table, the location of the sampling unit from which the corresponding representative mode is derived, and information about the calculated gradient.

[0216] According to another embodiment of weighting multiple prediction signals, the weights may be provided as fixed values ​​having the same weight for some modes, and may be provided implicitly and adaptively for the remaining modes using at least one of the height and width of the current block, the table, the location of the sampling unit determined for the corresponding mode, and information about the calculated gradient.

[0217] Meanwhile, the weights applied to each of several prediction signals can be integerized based on a look-up table (LUT).

[0218]

[0219] FIG. 11 is a flowchart illustrating an intra prediction method using prediction mode derivation according to one embodiment of the present disclosure.

[0220] Hereinafter, with reference to FIG. 11, a method for predicting a current block using a prediction mode derivation by an image decoding device or an image encoding device according to the present disclosure will be described. While the operation of an image decoding device will be described below, it will be apparent that the same operation can be performed by an image encoding device.

[0221] The image decoding device calculates a gradient for each sampling unit within a gradient calculation area composed of restored reference samples around the current block (S1110). Here, the range of the gradient calculation area and / or the size of the sampling unit may be determined based on the size or area of ​​the current block.

[0222] Thereafter, the image decoding device derives directional prediction modes based on the gradient for each of the sampling units and generates a histogram for the directional prediction modes (S1120). The directional prediction modes may be derived based on at least one of the position of the sampling unit, the aspect ratio of the current block, and the range of the gradient calculation region, and the gradient corresponding to the sampling unit.

[0223] The video decoding device derives at least one dominant prediction mode based on the histogram (S1130) and generates a prediction block using the at least one dominant prediction mode (S1140). When multiple dominant prediction modes are used, prediction blocks are generated using each prediction mode (S1142) and the prediction blocks are weighted and combined to generate a final prediction block for the current block (S1144).

[0224]

[0225] Each component of the device or method according to the present disclosure may be implemented in hardware, software, or a combination of hardware and software. Furthermore, the functions of each component may be implemented in software, with a microprocessor configured to execute the software functions corresponding to each component.

[0226] Various implementations of the systems and techniques described herein may be implemented as digital electronic circuits, integrated circuits, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations of one or more computer programs executable on a programmable system. The programmable system includes at least one programmable processor (which may be a special purpose processor or a general purpose processor) coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device. Computer programs (also known as programs, software, software applications, or code) include instructions for the programmable processor and are stored on a "computer-readable recording medium."

[0227] A computer-readable recording medium includes any type of recording device that stores data that can be read by a computer system. Such a computer-readable recording medium may be a non-volatile or non-transitory medium such as a ROM, CD-ROM, magnetic tape, floppy disk, memory card, hard disk, magneto-optical disk, storage device, and may further include a transitory medium such as a data transmission medium. Furthermore, the computer-readable recording medium may be distributed across network-connected computer systems, so that computer-readable code can be stored and executed in a distributed manner.

[0228] Although the flowchart / timing diagram of this specification describes each process as being executed sequentially, this is merely an illustrative description of the technical idea of ​​one embodiment of the present disclosure. In other words, a person of ordinary skill in the art to which one embodiment of the present disclosure belongs may modify and apply various modifications and variations by changing the order described in the flowchart / timing diagram without departing from the essential characteristics of one embodiment of the present disclosure, or by executing one or more of the processes in parallel. Therefore, the flowchart / timing diagram is not limited to a chronological order.

[0229] The above description is merely an example of the technical idea of ​​the present embodiment, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of ​​the present embodiment, but rather to explain it, and the scope of the technical idea of ​​the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment.

[0230] CROSS-REFERENCE TO RELATED APPLICATION

[0231] This patent application claims priority to Korean patent application No. 10-2023-0169361, filed in Korea on November 29, 2023, and Korean patent application No. 10-2024-0168696, filed in Korea on November 22, 2024, the entire contents of which are incorporated herein by reference.

Claims

1. In an image decoding method for generating a prediction block of a current block, A step of determining the size of a sampling unit for gradient operation based on the size or area of ​​the current block; A step of computing gradients for sampling units of the determined size within a gradient estimation area (template) including restored reference samples around the current block, and generating a histogram for directional prediction modes corresponding to the gradients; a step of deriving at least one dominant prediction mode based on the histogram; and A step of generating the prediction block using at least one dominant prediction mode. A method for decrypting an image, comprising:

2. In paragraph 1, An image decoding method, wherein each of the above sampling units is defined such that its center is spaced apart from other adjacent sampling units within the gradient calculation area by a distance of m (an integer greater than or equal to 2) pixels.

3. In paragraph 2, A method for decoding an image, wherein the above m pixel distance is determined based on at least one of the size of the current block, the resolution of the image, and the aspect ratio of the current block.

4. In paragraph 1, An image decoding method, further comprising a step of defining the sampling unit based on segmentation information of surrounding blocks adjacent to the current block.

5. In paragraph 1, The steps for generating the above histogram are: A step of filtering samples within a target sampling unit; and A step of computing a gradient for the target sampling unit using filtered samples within the target sampling unit. A method for decrypting an image, comprising:

6. In paragraph 1, The steps for generating the above histogram are: A step of deriving the directional prediction modes based on at least one of the positions of the sampling units, the aspect ratio of the current block, or the range of the gradient calculation region and the gradients; and A step of generating the histogram based on the above directional prediction modes. A method for decrypting an image, comprising:

7. In paragraph 6, The steps for generating the above histogram are: A step of determining a first directional prediction mode corresponding to the directionality of a gradient derived from a target sampling unit among a plurality of directional prediction modes; and A step of replacing the first directional prediction mode with the second directional prediction mode or setting it to an unavailable mode based on at least one of the aspect ratio of the current block or the range of the gradient calculation region. A method for decrypting an image, comprising:

8. In paragraph 6, The steps for generating the above histogram are: A step of determining a first directional prediction mode corresponding to the directionality of a gradient derived from a target sampling unit among a plurality of directional prediction modes; A step of determining whether the first directional prediction mode is available based on the position of the target sampling unit within the gradient calculation area; and If the first directional prediction mode is determined to be available, a step of using the first directional prediction mode to generate the histogram A method for decrypting an image, comprising:

9. In paragraph 1, The step of generating the above histogram includes an early termination process, The above early termination process is: A step of generating and updating the cumulative strength value of gradients by sequentially computing the gradients for the sampling units in a predefined order; and A step of comparing the above accumulated intensity value with a predefined threshold value to determine whether to end the generation of the histogram. A method for decrypting an image, comprising:

10. In paragraph 1, Further comprising a step of decoding a flag for indicating the gradient calculation area from the bitstream, An image decoding method, wherein the above gradient calculation area is determined as one or more of the left or upper areas of the current block based on the flag.

11. In paragraph 1, The size of the above sampling unit is, If the size of the current block is smaller than a predefined threshold or equal to a predefined block size, it is determined as a first value, Otherwise, the image decoding method is determined by a second value greater than the first value.

12. In an image encoding method for generating a prediction block of a current block, A step of determining the size of a sampling unit for gradient operation based on the size or area of ​​the current block; A step of computing gradients for sampling units of the determined size within a gradient estimation area (template) including restored reference samples around the current block, and generating a histogram for directional prediction modes corresponding to the gradients; a step of deriving at least one dominant prediction mode based on the histogram; and A step of generating the prediction block using at least one dominant prediction mode. A method of encoding an image, comprising:

13. A method for providing a bitstream including image data to an image decoding device, A step of encoding image data generated based on a prediction for the current block into the bitstream; and Including a step of transmitting the above bitstream to the image decoding device, The prediction for the current block above is: A step of determining the size of a sampling unit for gradient operation based on the size or area of ​​the current block; A step of computing gradients for sampling units of the determined size within a gradient estimation area (template) including restored reference samples around the current block, and generating a histogram for directional prediction modes corresponding to the gradients; a step of deriving at least one dominant prediction mode based on the histogram; and A method comprising the step of generating the prediction block using at least one dominant prediction mode.

Citation Information

Patent Citations

  • Video decoding method and device, encoding method and device thereof

    KR1020180075558A

  • Energy saving system and method for Cargo Hold within ventilation system

    KR1020240045630A

  • Method and device for exchanging secret keys based on reconfigurable and unclonable cryptographic component

    KR1020250052002A

  • Semiconductor device

    KR1020250061470A

  • Remote assistance apparatus using augmented reality

    KR102312015B1