Intra-predictive method for video coding using intra-predictive mode guidance
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HYUNDAI MOTOR CO LTD
- Filing Date
- 2022-03-03
- Publication Date
- 2026-08-05
AI Technical Summary
【0009】 本発明によると、予め復元された周辺参照サンプル値を用いて現在ブロックのイントラ予測モードを誘導した後、誘導された予測モードに基づいて現在ブロックの予測ブロックを生成するビデオコーディングのためのイントラ予測方法及び装置を提供することで、符号化効率を向上させる効果がある。
Smart Images

Figure 0007900787000007 
Figure 0007900787000008 
Figure 0007900787000009
Abstract
Description
[Technical Field]
[0001] This invention relates to video coding using intra-predictive mode induction. Intranet prediction method Regarding. [Background technology]
[0002] The following information is provided solely as background information related to the present invention and does not constitute prior art.
[0003] Video data has a much larger data volume than audio or still image data, and therefore requires a lot of hardware resources, including memory, to store or transmit it without compression.
[0004] Therefore, when storing or transmitting video data, an encoder is typically used to compress the video data before storage or transmission, and a decoder receives the compressed video data, decompresses it, and plays it back. Such video compression technologies include H.264 / AVC and HEVC (High Efficiency Video Coding), as well as VVC (Versatile Video Coding), which improves encoding efficiency by approximately 30% or more compared to HEVC.
[0005] However, as video size, resolution, and frame rate gradually increase, the amount of data to be encoded also increases. Therefore, there is a need for new compression technologies that offer better encoding efficiency and greater image quality improvement than conventional compression techniques. In particular, in terms of encoding efficiency, it is necessary to consider methods for inducing the intra-predictive mode of blocks instead of analyzing it. [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] The present invention has been made in view of the above-mentioned prior art, and the object of the present invention is for video coding that generates predicted blocks of current blocks. Intranet prediction method The objective is to provide. [Means for solving the problem]
[0007] An intra-prediction method according to one aspect of the present invention, made to achieve the above objective, is an intra-prediction method of a current block performed by a video decoding device, comprising the steps of: analyzing a prediction mode induction flag indicating whether or not the prediction mode of the current block can be derivated from a bitstream; confirming the prediction mode induction flag, and if the prediction mode induction flag is true, determining a calculation region used to calculate a gradient value from peripheral reconstructed samples of the current block; calculating a histogram of the directional mode gradient in the calculation region for the current block; inducing a prediction mode of the current block based on the gradient histogram; and generating a prediction block of the current block by performing intra-prediction using the derivated prediction mode.
[0008] An intra-prediction device according to one aspect of the present invention, made to achieve the above objective, is characterized by comprising: a prediction mode induction feasibility determination unit that analyzes a prediction mode induction flag from a bitstream and determines whether or not the prediction mode of the current block can be derived; a gradient calculation area determination unit that determines a calculation area used to calculate a gradient value from the surrounding reconstruction samples of the current block; a histogram calculation unit that calculates a histogram of the directional mode gradient in the calculation area for the current block; a prediction mode induction unit that induces the prediction mode of the current block based on the gradient histogram; and an intra-prediction execution unit that generates a prediction block of the current block by performing intra-prediction using the induced prediction mode. [Effects of the Invention]
[0009] According to the present invention, an intra-prediction method and apparatus for video coding is provided, which involves inducing an intra-prediction mode for the current block using pre-recovered peripheral reference sample values, and then generating a predicted block for the current block based on the induced prediction mode, thereby improving coding efficiency. [Brief explanation of the drawing]
[0010] [Figure 1] This is an exemplary block diagram of a video encoding device that embodies the technology of the present invention. [Figure 2] This diagram illustrates how to divide a block using the QTBTTT structure. [Figure 3a] This figure shows multiple intra-prediction modes, including a wide-angle intra-prediction mode. [Figure 3b] This figure shows multiple intra-prediction modes, including a wide-angle intra-prediction mode. [Figure 4] This is an illustrative diagram showing the surrounding blocks of the current block. [Figure 5] This is an exemplary block diagram of an image decoding device that embodies the technology of the present invention. [Figure 6] A block diagram of an intra-predictive device using predictive mode induction according to one embodiment of the present invention. [Figure 7] This is an illustrative diagram showing the calculation region used for calculating gradient values according to one embodiment of the present invention. [Figure 8] This is an illustrative diagram showing a histogram of the gradient of the directional mode according to one embodiment of the present invention. [Figure 9] This is an illustrative diagram showing subsampled pixels from which gradient values are calculated according to one embodiment of the present invention. [Figure 10] This is an illustrative diagram showing matrix-based weight values according to one embodiment of the present invention. [Figure 11] This diagram illustrates the induction of a prediction mode during subblock partitioning according to one embodiment of the present invention. [Figure 12]A flowchart showing an intra prediction method using prediction mode induction according to an embodiment of the present invention.
Embodiments for Carrying Out the Invention
[0011] Hereinafter, specific examples of embodiments for carrying out the present invention will be described in detail while referring to the drawings. It should be noted that when adding reference numerals to the components of each drawing, the same components are given the same reference numerals as much as possible even when shown in other drawings. In addition, in describing the present embodiment, when it is determined that a detailed description of related known configurations or functions will obscure the gist of the present invention, the detailed description thereof will be omitted.
[0012] FIG. 1 is an exemplary block diagram of a video encoding apparatus embodying the technology of the present invention. Hereinafter, the video encoding apparatus and the subordinate configurations of this apparatus will be described with reference to FIG. 1.
[0013] The video encoding apparatus is configured to include a picture partitioning unit 110, a prediction unit 120, a subtractor 130, a conversion unit 140, a quantization unit 145, a rearrangement unit 150, an entropy encoding unit 155, an inverse quantization unit 160, an inverse conversion unit 165, an adder 170, a loop filter unit 180, and a memory 190.
[0014] Each component of the video encoding apparatus may be implemented by hardware or software, or a combination of hardware and software. Further, the functions of each component may be implemented by software such that a microprocessor executes the functions of the software corresponding to each component.
[0015] A single video consists of one or more sequences containing multiple pictures. Each picture is divided into multiple regions, and encoding is performed for each region. For example, a single picture is divided into one or more tiles and / or slices. Here, one or more tiles are defined as a tile group. Each tile and / or slice is divided into one or more Coding Tree Units (CTUs). Each CTU is then divided into one or more Coding Units (CUs) by a tree structure. Information applicable to each CU is encoded as the CU syntax, and information applicable to all CUs contained in a single CTU is encoded as the CTU syntax. Furthermore, information applicable to all blocks within a slice is encoded as the slice header syntax, and information applicable to all blocks constituting one or more pictures is encoded in the Picture Parameter Set (PPS) or picture header. In addition, information commonly referenced by multiple pictures is encoded in the Sequence Parameter Set (SPS). Information that one or more SPSs commonly reference is encoded in a Video Parameter Set (VPS). Furthermore, information that applies commonly to a single tile or tile group is encoded as the syntax of the tile or tile group header. The syntax contained in SPSs, PPSs, slice headers, tiles, or tile group headers is referred to as high-level syntax.
[0016] The picture segmentation unit 110 determines the size of the Coding Tree Unit (CTU). Information regarding the size of the CTU (CTU size) is encoded as SPS or PPS syntax and transmitted to the video decoding device.
[0017] The picture division unit 110 divides each picture constituting the video into multiple Coding Tree Units (CTUs) of predetermined sizes, and then recursively divides the CTUs using a tree structure. In the tree structure, the leaf nodes become coding units (CUs), which are the basic units of encoding.
[0018] In tree structures, there are quad trees (QT) where the top node (or parent node) is divided into four lower nodes (or child nodes) of the same size, binary trees (BT) where the top node is divided into two lower nodes, or ternary trees (TT) where the top node is divided into three lower nodes in a 1:2:1 ratio, or structures that combine two or more of these QT, BT, and TT structures. For example, a QTBT (QuadTree plus BinaryTree) structure or a QTBTTT (QuadTree plus BinaryTree TernaryTree) structure is used. Here, BTTT is collectively called MTT (Multiple-Type Tree).
[0019] Figure 2 is a diagram illustrating how to divide a block using the QTBTTT structure.
[0020] As shown in Figure 2, the CTU is initially split into a QT structure. Quad-tree splitting is repeated until the size of the splitting block reaches the minimum block size of a leaf node allowed in the QT, MinQTSize. A first flag, QT_split_flag, indicating whether each node in the QT structure is split into four lower-layer nodes, is encoded by the entropy encoding unit 155 and signaled to the video decoder. If the leaf nodes of the QT are not larger than the maximum block size of a root node allowed in the BT, they are further split into one or more BT or TT structures. In the BT and / or TT structures, there are multiple splitting directions. For example, there are two directions in which the block of the node is split: horizontally and vertically. As shown in Figure 2, when MTT splitting is initiated, a second flag, mtt_split_flag, indicating whether or not a node has been split, and additionally, if split, a flag indicating the splitting direction (vertical or horizontal) and / or the splitting type (binary or terrary), are encoded by the entropy encoding unit 155 and signaled to the video decoding device.
[0021] Alternatively, before encoding the first flag QT_split_flag, which indicates whether each node will be split into four lower-layer nodes, the CU split flag split_cu_flag, which indicates whether the node will be split, is encoded. If the value of the CU split flag split_cu_flag indicates that the node will not be split, the block of that node becomes a leaf node in the split tree structure and becomes a coding unit (CU), which is the basic unit of encoding. If the value of the CU split flag split_cu_flag indicates that the node will be split, the video encoding device starts encoding from the first flag using the method described above.
[0022] When QTBT is used as another example of a tree structure, there are two types: one in which the block of the node is divided horizontally into two blocks of the same size (i.e., symmetric horizontal splitting) and another in which it is divided vertically (i.e., symmetric vertical splitting). A splitting flag, split_flag, indicating whether each node in the BT structure is divided into a block of a lower layer, and splitting type information indicating the type of division are encoded by the entropy encoding unit 155 and transmitted to the video decoding device. On the other hand, there is an additional type in which the block of the node is divided into two blocks in an asymmetrical manner. The asymmetrical form includes a form in which the block of the node is divided into two rectangular blocks with a size ratio of 1:3, or a form in which the block of the node is divided diagonally.
[0023] CUs can have various sizes depending on the QTBT or QTBTTT partitioning from the CTU. Hereafter, the block corresponding to the CU to be encoded or decoded (i.e., a leaf node in the QTBTTT) will be referred to as the "current block." Depending on the QTBTTT partitioning used, the shape of the current block may be a rectangle as well as a square.
[0024] The prediction unit 120 predicts the current block and generates a predicted block. The prediction unit 120 includes an intra-prediction unit 122 and an inter-prediction unit 124.
[0025] Generally, each current block in a picture is coded predictively. Generally, the prediction of the current block is performed using intra-prediction techniques (using data from the picture containing the current block) or inter-prediction techniques (using data from pictures coded before the picture containing the current block). Inter-prediction includes both one-way and two-way prediction.
[0026] The intra-prediction unit 122 predicts pixels within the current block using pixels (reference pixels) located around the current block within the current picture that contains the current block. Multiple intra-prediction modes exist depending on the prediction direction. For example, as seen in Figure 3a, the multiple intra-prediction modes include two non-directional modes, including the planar mode and DC mode, and 65 directional modes. The surrounding pixels used by each prediction mode are defined to have different calculation formulas.
[0027] For efficient directional prediction of rectangular current blocks, additional directional modes (67-80, -1-14 intra-prediction modes) are used, indicated by dashed arrows in Figure 3b. These are referred to as "wide-angle intra-prediction modes." In Figure 3b, the arrows point to the corresponding reference samples used for prediction, not to the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. Wide-angle intra-prediction modes are modes that perform prediction in the opposite direction of a specific directional mode without transmitting additional bits when the current block is rectangular. In this case, some of the wide-angle intra-prediction modes available to the current block are determined by the ratio of the width to the height of the rectangular current block. For example, wide-angle intra-prediction modes with angles smaller than 45 degrees (intra-prediction modes 67-80) are available when the current block is in the form of a rectangle where the height is smaller than the width, and wide-angle intra-prediction modes with angles larger than -135 degrees (intra-prediction modes -1-14) are available when the current block is in the form of a rectangle where the width is larger than the height.
[0028] The intra-prediction unit 122 determines the intra-prediction mode to be used to encode the current block. In some examples, the intra-prediction unit 122 encodes the current block using various intra-prediction modes and selects the appropriate intra-prediction mode to be used from the tested modes. For example, the intra-prediction unit 122 calculates bitrate distortion values using bitrate distortion analysis for various tested intra-prediction modes and selects the intra-prediction mode with the best bitrate distortion characteristics among the tested modes.
[0029] The intra-prediction unit 122 selects one intra-prediction mode from among several intra-prediction modes and predicts the current block using the surrounding pixels (reference pixels) and calculation formula determined by the selected intra-prediction mode. Information regarding the selected intra-prediction mode is encoded by the entropy coding unit 155 and transmitted to the video decoding device.
[0030] The interpretation unit 124 generates a predicted block for the current block using a motion compensation process. The interpretation unit 124 searches for the block most similar to the current block in the reference picture, which has been encoded and decoded before the current picture, and uses the found block to generate a predicted block for the current block. It then generates a motion vector (MV) corresponding to the displacement between the current block in the current picture and the predicted block in the reference picture. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma and chroma components. Motion information, including information about the reference picture used to predict the current block and information about the motion vector, is encoded by the entropy encoding unit 155 and transmitted to the video decoding device.
[0031] The interpretation unit 124 performs interpolation on a reference picture or reference block to improve the accuracy of the prediction. That is, a subsample between two consecutive integer samples is interpolated by applying a filter coefficient to a series of consecutive integer samples that include those two integer samples. After performing the step of searching for the block most similar to the current block on the interpolated reference picture, the motion vector is expressed with decimal precision rather than integer sample precision. The precision or resolution of the motion vector is set differently for each unit of the target area to be encoded, such as slice, tile, CTU, CU, etc. When such Adaptive Motion Vector Resolution (AMVR) is applied, information about the motion vector resolution applied to each target area must be signaled for each target area. For example, if the target area is a CU, information about the motion vector resolution applied to each CU is signaled. The information about the motion vector resolution is information that indicates the precision of the differential motion vector, which will be described later.
[0032] On the other hand, the inter-prediction unit 124 performs inter-prediction using bi-prediction. In the case of bi-prediction, two reference pictures and two motion vectors representing the block position most similar to the current block within each reference picture are used. The inter-prediction unit 124 selects a first reference picture and a second reference picture from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively, and searches for a block similar to the current block within each reference picture to generate a first reference block and a second reference block. Then, it generates a predicted block for the current block by averaging or weighting the first reference block and the second reference block. Finally, it transmits motion information, including information about the two reference pictures and the two motion vectors used to predict the current block, to the encoding unit 150. Here, reference picture list 0 consists of previously restored pictures that precede the current picture in display order, and reference picture list 1 consists of previously restored pictures that precede the current picture in display order. However, it is not necessarily limited to this; previously restored pictures that precede the current picture in display order may be added to reference picture list 0, or conversely, previously restored pictures that precede the current picture may be added to reference picture list 1.
[0033] Various methods are used to minimize the number of bits required to encode motion information.
[0034] For example, if the reference picture and motion vector of the current block are the same as those of a surrounding block, the motion information of the current block is transmitted to the video decoder by encoding information that identifies the surrounding block. This method is called "merge mode".
[0035] In merge mode, the interpretation unit 124 selects a predetermined number of merge candidate blocks (hereinafter referred to as "merge candidates") from the surrounding blocks of the current block.
[0036] As shown in Figure 4, the surrounding blocks used to guide merge candidates are the left block A0, the lower left block A1, the upper block B0, the upper right block B1, and the upper left block adjacent to the current block within the current picture. B All or part of option 2 is used. Furthermore, blocks located in a reference picture (which may or may not be the same as the reference picture used to predict the current block) rather than the current picture in which the current block is located are used as merge candidates. For example, a block in the reference picture that is in the same position as the current block (a co-located block) or a block adjacent to that block in the same position is additionally used as a merge candidate. If the number of merge candidates selected by the method described above is less than a predetermined number, a 0 vector is added to the merge candidates.
[0037] The interpretation unit 124 constructs a merge list containing a predetermined number of merge candidates using such surrounding blocks. From the merge candidates included in the merge list, it selects a merge candidate to be used as the motion information of the current block and generates merge index information to identify the selected candidate. The generated merge index information is encoded by the encoding unit 150 and transmitted to the video decoding device.
[0038] The merge skip mode is a special case of the merge mode in which, after quantization, if all transformation coefficients for entropy coding are close to zero, only the peripheral block selection information is transmitted without transmitting the residual signal. By using the merge skip mode, relatively high coding efficiency can be achieved for videos with little motion, still images, and screen content.
[0039] Hereafter, merge mode and merge skip mode will be collectively referred to as merge / skip mode.
[0040] Another method for encoding motion information is AMVP (Advanced Motion Vector Prediction) mode.
[0041] In AMVP mode, the interpretation unit 124 uses the surrounding blocks of the current block to induce predicted motion vector candidates for the motion vector of the current block. The surrounding blocks used to induce predicted motion vector candidates are the left block A0, the lower left block A1, the upper block B0, the upper right block B1, and the upper left block adjacent to the current block in the current picture shown in Figure 4. B All or part of option 2 is used. Furthermore, blocks located in a reference picture (which may or may not be the same as the reference picture used to predict the current block) rather than the current picture in which the current block is located are used as surrounding blocks to guide the predicted motion vector candidates. For example, a block in the same position as the current block in the reference picture (a collocated block), or a block adjacent to that block in the same position, is used. If the number of motion vector candidates obtained by the method described above is less than a predetermined number, a 0 vector is added to the motion vector candidates.
[0042] The interpretation unit 124 uses the motion vectors of the surrounding blocks to derive candidate predicted motion vectors, and uses these candidate predicted motion vectors to determine the predicted motion vector relative to the current block's motion vector. Then, it subtracts the predicted motion vector from the current block's motion vector to calculate the difference motion vector.
[0043] The predicted motion vector is obtained by applying a predefined function (e.g., median, mean calculation) to the predicted motion vector candidate. In this case, the video decoder also knows the predefined function. Furthermore, since the surrounding blocks used to guide the predicted motion vector candidate are already encoded and decoded blocks, the video decoder also already knows the motion vectors of those surrounding blocks. Therefore, the video encoder does not need to encode information to identify the predicted motion vector candidate. Consequently, in this case, information about the differential motion vector and information about the reference picture used to predict the current block are encoded.
[0044] On the other hand, the predicted motion vector is determined by selecting one of the candidate predicted motion vectors. In this case, information to identify the selected candidate predicted motion vector is additionally encoded, along with information about the differential motion vector and the reference picture used to predict the current block.
[0045] The subtractor 130 generates a residual block by subtracting the predicted block generated by the intra-prediction unit 122 or the inter-prediction unit 124 from the current block.
[0046] The conversion unit 140 converts the residual signals in a residual block having pixel values in the spatial domain into conversion coefficients in the frequency domain. The conversion unit 140 converts the residual signals in the residual block using the overall size of the residual block as the conversion unit, or divides the residual block into multiple subblocks and converts using those subblocks as the conversion unit. Alternatively, it divides the residual block into two subblocks, a conversion region and a non-conversion region, and converts the residual signals using only the conversion region subblock as the conversion unit. Here, the conversion region subblock is one of two rectangular blocks having a 1:1 size ratio based on the horizontal axis (or vertical axis). In this case, a flag cu_sbt_flag indicating that only the subblock was converted, directional (vertical / horizontal) information cu_sbt_horizontal_flag, and / or position information cu_sbt_pos_flag are encoded by the entropy encoding unit 155 and signaled to the video decoding device. Furthermore, the size of the subblocks in the conversion region has a size ratio of 1:3 based on the horizontal axis (or vertical axis). In such cases, a flag cu_sbt_quad_flag that distinguishes the relevant division is additionally encoded by the entropy encoding unit 155 and signaled to the video decoding device.
[0047] Meanwhile, the transformation unit 140 performs transformations on the residual blocks individually in the horizontal and vertical directions. Various types of transformation functions or transformation matrices are used for the transformations. For example, a pair of transformation functions for horizontal and vertical transformations is defined as an MTS (Multiple Transform Set). The transformation unit 140 selects the one transformation function pair with the best transformation efficiency from among the MTS and transforms the residual blocks in the horizontal and vertical directions, respectively. Information about the selected transformation function pair from among the MTS, mts_idx, is encoded by the entropy coding unit 155 and signaled to the video decoding device.
[0048] The quantization unit 145 quantizes the conversion coefficients output from the conversion unit 140 using quantization parameters and outputs the quantized conversion coefficients to the entropy coding unit 155. For any block or frame, the quantization unit 145 immediately quantizes the residual blocks associated with no conversion. The quantization unit 145 applies different quantization coefficients (scaling values) to each other depending on the position of the conversion coefficients within the conversion block. The quantization matrix applied to the two-dimensionally arranged quantized conversion coefficients is encoded and signaled to the video decoding device.
[0049] The sorting unit 150 performs sorting of coefficient values for the quantized residual values.
[0050] The sorting unit 150 converts a two-dimensional coefficient array into a one-dimensional coefficient sequence using coefficient scanning. For example, the sorting unit 150 scans from the DC coefficients to the high-frequency region coefficients using a zig-zag scan or diagonal scan to output a one-dimensional coefficient sequence. Depending on the size of the conversion unit and the intra-prediction mode, a vertical scan that scans the two-dimensional coefficient array in the column direction and a horizontal scan that scans the two-dimensional block-shaped coefficients in the row direction may be used instead of a zig-zag scan. In other words, the scanning method used from among zig-zag scan, diagonal scan, vertical scan, and horizontal scan is determined by the size of the conversion unit and the intra-prediction mode.
[0051] The entropy coding unit 155 generates a bitstream by coding the sequence of one-dimensional quantized transformation coefficients output from the sorting unit 150 using various coding schemes such as CABAC (Context-based Adaptive Binary Arithmetic Code) and Exponential Golomb.
[0052] Furthermore, the entropy coding unit 155 encodes information related to block partitioning, such as the CTU size, CU partitioning flag, QT partitioning flag, MTT partitioning type, and MTT partitioning direction, so that the video decoder can partition blocks in the same way as the video encoder. The entropy coding unit 155 also encodes information about the prediction type, indicating whether the block was currently encoded by intra-prediction or inter-prediction, and encodes intra-prediction information (i.e., information about the intra-prediction mode) or inter-prediction information (information about the motion information encoding mode (merge mode or AMVP mode), the merge index in the case of merge mode, and the reference picture index and differential motion vector in the case of AMVP mode) depending on the prediction type. The entropy coding unit 155 also encodes information related to quantization, i.e., information about quantization parameters and information about the quantization matrix.
[0053] The inverse quantization unit 160 inverse quantizes the quantized conversion coefficients output from the quantization unit 145 to generate conversion coefficients. The inverse conversion unit 165 converts the conversion coefficients output from the inverse quantization unit 160 from the frequency domain to the spatial domain to restore the residual block.
[0054] The addition unit 170 adds the restored residual block and the predicted block generated by the prediction unit 120 to restore the current block. The pixels in the restored current block are used as reference pixels when intra-predicting the next block in order.
[0055] The loop filter section 180 performs filtering on the restored pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc., that occur due to block-based prediction and transformation / quantization. The filter section 180 includes all or part of a deblocking filter 182, a Sample Adaptive Offset (SAO) filter 184, and an Adaptive Loop Filter (ALF) 186 as in-loop filters.
[0056] The deblocking filter 182 filters the boundaries between restored blocks to remove blocking artifacts that occur during block-level encoding / decoding, and the SAO filter 184 and ALF 186 performs additional filtering on the deblocked filtered video. SAO filter 184 and ALF Filter 186 is used to compensate for the difference between the restored pixels and the original pixels that occurs due to lossy coding. The SAO filter 184 improves not only subjective image quality but also coding efficiency by applying an offset in units of CTU. In contrast, ALF 186 performs block-level filtering, compensating for distortion by applying different filters to the edges and degree of change of the relevant block. Information regarding the filter coefficients used for ALF is encoded and signaled to the video decoder.
[0057] The recovered blocks filtered through the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in memory 190. Once all blocks in a picture have been recovered, the recovered picture is used as a reference picture to interpret the blocks in the picture to be encoded later.
[0058] Figure 5 is an illustrative block diagram of an image decoding device embodying the technology of the present invention. The image decoding device and its sub-configurations will be described below with reference to Figure 5.
[0059] The video decoding device is configured to include an entropy decoding unit 510, a sorting unit 515, an inverse quantization unit 520, an inverse transform unit 530, a prediction unit 540, an adder 550, a loop filter unit 560, and a memory 570.
[0060] Similar to the video encoding device in Figure 1, each component of the video decoding device is embodied in hardware or software, or in a combination of hardware and software. Furthermore, the function of each component may be embodied in software, and a microprocessor may be configured to execute the software function corresponding to each component.
[0061] The entropy decoding unit 510 decodes the bitstream generated by the video encoding device and extracts information related to block division to determine the current block to be decoded, and extracts prediction information and residual signal information necessary to restore the current block.
[0062] The entropy decoding unit 510 extracts information about the CTU size from the SPS (Sequence Parameter Set) or PPS (Picture Parameter Set) to determine the size of the CTU and divides the picture into CTUs of the determined size. Then, it determines the CTU as the top layer of the tree structure, i.e., the root node, and divides the CTU using the tree structure by extracting division information related to the CTU.
[0063] For example, when splitting a CTU using a QTBTTTT structure, first, the first flag QT_split_flag related to QT splitting is extracted, and each node is split into four lower layer nodes. Then, for nodes corresponding to QT leaf nodes, the second flag MTT_split_flag related to MTT splitting and split direction (vertical / horizontal) and / or split type (binary / ternary) information are extracted, and the corresponding leaf node is split into an MTT structure. In this way, each node below the QT leaf node is recursively split into a BT or TT structure.
[0064] As another example, when splitting a CTU using a QTBTTT structure, first, the CU splitting flag split_cu_flag, which indicates whether the CU can be split, is extracted. If the block is split, the first flag QT_split_flag is extracted. During the splitting process, each node undergoes zero or more QT splits followed by zero or more MTT splits. For example, a CTU may undergo an MTT split immediately, or conversely, only multiple QT splits.
[0065] As another example, when splitting a CTU using a QTBT structure, the first flag QT_split_flag related to the splitting of QT is extracted, and each node is split into four lower layer nodes. Then, for nodes corresponding to the leaf nodes of QT, the splitting flag split_flag indicating whether or not it will be further split by BT and splitting direction information are extracted.
[0066] Meanwhile, the entropy decoding unit 510, after determining the current block to be decoded using the tree structure division, extracts information about the prediction type indicating whether the current block is intra-predicted or inter-predicted. If the prediction type information indicates intra-prediction, the entropy decoding unit 510 extracts syntax elements related to the intra-prediction information (intra-prediction mode) of the current block. If the prediction type information indicates inter-prediction, the entropy decoding unit 510 extracts syntax elements related to inter-prediction information, i.e., information representing the motion vector and the reference picture that the motion vector refers to.
[0067] Furthermore, the entropy decoding unit 510 extracts information regarding the quantized transformation coefficients of the current block as information related to quantization and information regarding the residual signal.
[0068] The sorting unit 515 converts the sequence of one-dimensional quantized transformation coefficients, which have been entropically decoded by the entropy decoding unit 510, back into a two-dimensional coefficient array (i.e., blocks) in the reverse order of the coefficient scanning sequence performed by the video encoding device.
[0069] The inverse quantization unit 520 inversely quantizes the quantized transformation coefficients and inversely quantizes the quantized transformation coefficients using quantization parameters. The inverse quantization unit 520 applies different quantization coefficients (scaling values) to the quantized transformation coefficients arranged in two dimensions. The inverse quantization unit 520 performs inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video encoding device to the two-dimensional array of quantized transformation coefficients.
[0070] The inverse transform unit 530 generates a residual block for the current block by inversely transforming the inversely quantized transformation coefficients from the frequency domain to the spatial domain and restoring the residual signal.
[0071] Furthermore, when the inverse transformer 530 inversely transforms only a portion of the transformation block (subblock), it extracts a flag cu_sbt_flag indicating that only the subblock of the transformation block has been transformed, a directional (vertical / horizontal) information cu_sbt_horizontal_ flag of the subblock, and / or position information cu_sbt_pos_flag of the subblock. It then reconstructs the residual signal by inversely transforming the transformation coefficients of the corresponding subblock from the frequency domain to the spatial domain, and generates the final residual block for the current block by satisfying the value of "0" in the residual signal for the region that has not been inversely transformed.
[0072] Furthermore, when MTS is applied, the inverse transformation unit 530 uses the MTS information mts_idx signaled from the video encoding device to determine the transformation function or transformation matrix to be applied in the horizontal and vertical directions, respectively, and performs an inverse transformation on the transformation coefficients within the transformation block in the horizontal and vertical directions using the determined transformation function.
[0073] The prediction unit 540 includes an intra-prediction unit 542 and an inter-prediction unit 544. The intra-prediction unit 542 is activated when the prediction type of the current block is intra-prediction, and the inter-prediction unit 544 is activated when the prediction type of the current block is inter-prediction.
[0074] The intra-prediction unit 542 determines the intra-prediction mode for the current block from among multiple intra-prediction modes based on the syntax elements for the intra-prediction modes extracted from the entropy decoding unit 510, and predicts the current block using the reference pixels surrounding the current block according to the intra-prediction mode.
[0075] The interprediction unit 544 uses the syntax elements for the interprediction mode extracted from the entropy decoding unit 510 to determine the motion vector of the current block and the reference picture that the motion vector refers to, and then predicts the current block using the motion vector and the reference picture.
[0076] The adder 550 restores the current block by adding the residual block output from the inverse transform unit and the predicted block output from the inter-prediction unit or intra-prediction unit. The pixels in the restored current block are used as reference pixels when intra-predicting the block to be decoded later.
[0077] The loop filter section 560 includes a deblocking filter 562, an SAO filter 564, and an ALF 566 as in-loop filters. The deblocking filter 562 deblocks the boundaries between restored blocks to remove blocking artifacts that occur due to block-level decoding. The SAO filter 564 and ALF 566 perform additional filtering on the restored blocks after deblocking to compensate for the difference between the restored pixels and the original pixels that occurs due to lossy coding. The filter coefficients of the ALF are determined using information about the filter coefficients decoded from the bitstream.
[0078] The recovered blocks filtered through the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in memory 570. Once all blocks in a picture have been recovered, the recovered picture is used as a reference picture to interpret the blocks in the picture to be encoded later.
[0079] This embodiment relates to the encoding and decoding of video as described above. More specifically, it provides a video coding method and apparatus that first induces an intra-prediction mode of the current block using previously restored peripheral reference sample values, and then generates a predicted block of the current block based on the induced prediction mode.
[0080] The following embodiments apply to the entropy decoding unit 510 and the intra prediction unit 542 in the video decoding device. They also apply to the intra prediction unit 122 in the video encoding device.
[0081] In the following explanation, the aspect ratio of a block is defined as the ratio between the width (W) and height (H) of the block.
[0082] In the following, a specific flag being true indicates that the value of that flag is 1, and a specific flag being false indicates that the value of that flag is 0.
[0083] The following description focuses on the video decoding device and intra-prediction, with references to the video encoding device only when necessary for convenience. The information below also applies to the video encoding device.
[0084] The phrases "the video decoding device or the entropy decoding unit 510 within the video decoding device decodes data from the bitstream" and "the data analyzes" are interchangeable.
[0085] ≪I. Intra Prediction and ISP (Intra Sub-Partitions)≫ In VVC technology, the intra-prediction modes of the rumabloc include non-directional modes (i.e., Planar and DC) as well as subdivided directional modes (i.e., 2-66), as illustrated in Figure 3a. Furthermore, as added to the example in Figure 3b, the intra-prediction modes of the rumabloc include directional modes (-14--1 and 67-80) corresponding to wide-angle intra-prediction.
[0086] Based on the prediction mode of a rumor block, various techniques exist to improve the coding efficiency of intra-prediction. ISP techniques currently subdivide a block into smaller blocks of the same size, then share the intra-prediction mode across all subblocks, but apply a transformation to each subblock. In this case, the subdivision of the block can be horizontal or vertical.
[0087] In the following explanation, the large block before subdivision will be referred to as the current block, and each of the smaller subdivided blocks will be referred to as a subblock.
[0088] The operation of ISP technology is as follows:
[0089] The video encoder transmits intra_subpartitions_mode_flag and intra_subpartitions_split_flag if the ISP activation flag sps_isp_enabled_flag on the higher-level SPS is true. The video decoder first parses the ISP activation flag sps_isp_enabled_flag from the bitstream. If the ISP activation flag is true, the video decoder decodes intra_subpartitions_mode_flag and intra_subpartitions_split_flag from the bitstream.
[0090] The video encoding device signals the video decoding device with intra_subpartitions_mode_flag, which indicates whether ISP is applicable, and intra_subpartitions_split_flag, which indicates the subpartitioning method. The subpartitioning types IntraSubPartitionsSplitType, determined by intra_subpartitions_mode_flag and intra_subpartitions_split_flag, are shown in Table 1.
[0091] [Table 1]
[0092] ISP technology sets the partition type IntraSubPartitionsSplitType as follows:
[0093] If intra_subpartitions_mode_flag is 0, IntraSubPartitionsSplitType is set to 0, and subblock splitting is not performed (ISP_NO_SPLIT). In other words, ISP is not applied.
[0094] If intra_subpartitions_mode_flag is not 0, ISP is applied. In this case, IntraSubPartitionsSplitType is set to the value of 1 + intra_subpartitions_split_flag, and subblock splitting is performed according to the splitting type. When IntraSubPartitionsSplitType=1, subblock splitting is performed horizontally (ISP_HOR_SPLIT), and when IntraSubPartitionsSplitType=2, subblock splitting is performed vertically (ISP_VER_SPLIT). In other words, intra_subpartitions_split_flag indicates the direction of subblock splitting.
[0095] For example, if the ISP mode, which subdivides horizontally, is currently applied to a block, then IntraSubPartitionsSplitType is 1, intra_subpartitions_mode_flag is 1, and intra_subpartitions_split_flag is 0.
[0096] In the following explanation, intra_subpartitions_mode_flag is represented as the sub-block splitting application flag, intra_subpartitions_split_flag is represented as the sub-block splitting direction flag, and IntraSubPartitionsSplitType is represented as the sub-block splitting type.
[0097] ≪II. Induction of Intra Prediction Mode≫ The intra-prediction device using prediction mode guidance is described below with reference to Figure 6.
[0098] Figure 6 is a block diagram showing an intra-predictive device using predictive mode induction according to one embodiment of the present invention.
[0099] The intra-prediction device according to this embodiment calculates a histogram from the gradient of previously restored peripheral reference sample values, uses the calculated histogram to induce an intra-prediction mode for the current block, and then generates a predicted block for the current block based on the induced prediction mode. The intra-prediction device includes all or part of a prediction mode induction feasibility determination unit 602, a gradient calculation area determination unit 604, a histogram calculation unit 606, a prediction mode induction unit 608, and an intra-prediction execution unit 610.
[0100] The prediction mode induction feasibility determination unit 602 analyzes a flag indicating whether the prediction mode of the current block can be induced and determines whether the intra-prediction mode can be induced. Hereinafter, such a flag will be referred to as the prediction mode induction flag. The prediction mode induction flag is set by the video encoding device from the viewpoint of bitrate distortion optimization and then transmitted to the video decoding device. After decoding the prediction mode induction flag from the bitstream, the video decoding device performs the steps described below. Meanwhile, the video encoding device obtains the prediction mode induction flag from the higher level and performs the subsequent steps.
[0101] If the prediction mode guidance flag is true, the intra-predictor guides the intra-prediction mode based on the peripheral reconstructed samples of the current block, without analyzing the intra-prediction mode of the current block. In this case, if the current block includes picture boundaries, slice boundaries, tile boundaries, etc., and left and / or upper reference samples are unavailable, decoding of the prediction mode guidance flag is implicitly omitted.
[0102] In another embodiment, the intra-prediction device analyzes one or more flags from the bitstream to determine whether the prediction mode of the current block is a non-directional mode such as Planar, DC mode, or matrix-based mode, before deciding whether or not to induce a prediction mode. If all of the relevant flags are 0, the prediction mode induction feasibility determination unit 602 analyzes the prediction mode induction flags and then decides whether or not to induce a prediction mode.
[0103] Furthermore, after the prediction mode induction feasibility determination unit 602 decides to induce a prediction mode, the intra prediction device decodes the subblock division application flag from the bitstream and determines whether or not ISP technology can be applied, that is, whether or not the subblock of the current block can be divided.
[0104] Figure 7 is an illustrative diagram showing the calculation region used for calculating the gradient value according to one embodiment of the present invention.
[0105] The gradient calculation region determination unit 604 determines the calculation region used to calculate the gradient value from the surrounding reconstructed samples of the current block in order to induce the intra prediction mode. As illustrated in Figure 7, the three rows of reconstructed reference pixels located to the left and above the current block within the reconstructed region are used as the gradient calculation region. In the example in Figure 7, the length M of the upper reference sample and the length N of the left reference sample are set based on the width W and height H of the current block. For example, based on the (-3,-3) pixel located in the upper left of the current block, M is set to a value such as W+3, 2×W+3, or W+H+3, and N is set to a value such as H+3, 2×H+3, or W+H+3.
[0106] In another embodiment, the gradient calculation area determination unit 604 analyzes a flag indicating the calculation area from the bitstream and sets the area to be used as the calculation area from the left or upper row relative to the current block. Alternatively, the gradient calculation area determination unit 604 analyzes an index indicating the calculation area from the bitstream and determines the area to be used as the calculation area from the left, upper row, or left / upper row relative to the current block.
[0107] In another embodiment, the calculation domain is implicitly determined by an agreement between the video encoding device and the video decoding device. In such a case, the operation of the gradient calculation domain determination unit 604 is omitted.
[0108] The histogram calculation unit 606 calculates the histogram of the gradient of the directional mode in the gradient calculation domain (Histogram: (H) The gradient is calculated. First, vertical and horizontal gradient values are calculated in a 3x3, 3x1, or 1x3 region, based on the restored reference sample located on the second line relative to the current block. The histogram calculation unit 606 calculates the gradient using an edge detection filter such as a Sobel filter or a Prewitt filter. Table 2 shows embodiments of the filters used to calculate the gradient.
[0109] [Table 2]
[0110] The histogram calculation unit 606 determines the gradient direction θ and gradient size I at the relevant pixel based on the calculated vertical / horizontal gradient Gv / Gh, as shown in Equation 1.
[0111]
number
[0112] Figure 8 is an illustrative diagram showing a histogram of the gradient of the directional mode according to one embodiment of the present invention.
[0113] The histogram calculation unit 606 calculates the directional mode of the intra-prediction that is closest to the corresponding direction according to the gradient direction θ for each pixel in the calculation region, and then accumulates the gradient size I in the histogram of the corresponding directional mode to generate a histogram (H) of the gradients of the directional modes, as illustrated in Figure 8.
[0114] When the histogram calculation unit 606 calculates the gradient histogram, as shown in the example in Figure 9, it subsamples pixels based on a preset sampling interval and then calculates the gradient histogram for the subsampled pixels. At this time, the pixel sampling interval and subsampling position are defined according to the agreement between the video encoding device and the video decoding device. Alternatively, the pixel sampling interval and subsampling position are determined based on the size of the current block and / or the aspect ratio of the current block.
[0115] The prediction mode guidance unit 608 guides the prediction mode of the current block based on the gradient histogram, as illustrated in Figure 8. The guided intra-prediction mode is either a directional mode or a non-directional mode. Furthermore, the guided intra-prediction mode includes both directional and non-directional modes.
[0116] The prediction mode guidance unit 608 guides the directional mode as follows:
[0117] The predictive mode guidance unit 608 determines the mode M with the largest value among the calculated histograms. b This is determined to be the current block's intra-prediction mode.
[0118] Alternatively, the prediction mode guidance unit 608 selects the mode M having the largest value. b and the mode with the second largest value M a We determine this as the intra-prediction mode for the current block. These are two modes M. b and M a The intra-prediction device generates a final predicted signal by weighting the predicted signals P1 and P2, which were predicted using the method, with weight values corresponding to the histogram values of the two modes.
[0119] In another embodiment, the predictive mode guidance unit 608 determines the mode M with the largest gradient histogram value. b and the second largest mode M a An additional flag indicating one of the modes is analyzed, and the prediction mode of the current block is determined according to the analyzed flag value. Hereinafter, this flag will be referred to as the prediction mode indicator flag. The prediction mode guidance unit 608 analyzes the prediction mode indicator flag from the bitstream if the difference between the histogram values of the two modes is less than or equal to a preset threshold. On the other hand, if the difference between the histogram values of the two modes is greater than a preset threshold, the prediction mode guidance unit 608 omits the analysis of the prediction mode indicator flag. Alternatively, as shown in Equation 2, if the difference between the two histogram values is greater than or equal to a preset ratio compared to the sum of the total histogram values, the prediction mode guidance unit 608 omits the analysis of the prediction mode indicator flag.
[0120]
number
[0121] In another embodiment, the prediction mode guidance unit 608 determines the prediction mode by analyzing the delta mode index from the bitstream with respect to the basic mode, which is the mode with the largest histogram value. For example, the delta mode is an offset relative to the basic mode. Therefore, the prediction mode guidance unit 608 determines the prediction mode by adding the delta mode to the basic mode. At this time, the index indicating the delta mode is set according to the agreement between the video encoding device and the video decoding device. Also, the delta mode is changed according to the basic mode.
[0122] On the other hand, the prediction mode induction unit 608 induces a non-directional mode as follows.
[0123] If the histogram value of the directional mode with the largest value is smaller than a preset threshold, or if the sum of the total histogram values is smaller than a preset threshold, the prediction mode guidance unit 608 determines the non-directional mode as the prediction mode for the current block. In this case, the non-directional mode is either DC mode or planar mode, and the prediction mode guidance unit 608 always determines the prediction mode for the current block as DC mode (or planar mode). Alternatively, the prediction mode guidance unit 608 analyzes a flag from the bitstream indicating one of these modes, and then determines one of the two modes according to the analyzed flag. Meanwhile, the preset threshold is set according to an agreement between the video encoding device and the video decoding device and is transmitted from the video encoding device to the video decoding device for each higher-level unit such as picture or slice.
[0124] In another embodiment, as shown in Equation 3, if the histogram value of the directional mode having the largest value is smaller than a preset ratio compared to the sum of the total histogram values, the prediction mode guidance unit 608 determines the non-directional mode as the prediction mode for the current block.
[0125]
number
[0126] At this time, the pre-set threshold is determined based on the current block size.
[0127] In another embodiment, if the value obtained by dividing the sum of the total histogram values by the number of pixels used in gradient calculation is smaller than a preset threshold, the prediction mode guidance unit 608 determines the non-directional mode as the prediction mode for the current block. Also, if the value obtained by dividing the histogram value of the directional mode with the largest value by the number of pixels used in gradient calculation is smaller than a preset threshold, the prediction mode guidance unit 608 determines the non-directional mode as the prediction mode for the current block.
[0128] Also, when the value obtained by dividing the sum of the overall histogram values by the number of gradients used in histogram calculation is smaller than a preset threshold value, the prediction mode induction unit 608 determines the non-directional mode as the prediction mode of the current block. Also, when the value obtained by dividing the histogram value of the directional mode having the largest value by the number of gradients used in histogram calculation is smaller than a preset threshold value, the prediction mode induction unit 608 determines the non-directional mode as the prediction mode of the current block.
[0129] On the other hand, as illustrated in Table 2, the number of pixels used in gradient calculation is 9 times or 3 times the number of gradients used in histogram calculation.
[0130] As another embodiment, when the induced prediction mode of the current block is a directional mode, the prediction mode induction unit 608 adds the non-directional mode as the prediction mode of the current block. At this time, the added non-directional mode is set according to the agreement between the video encoding device and the video decoding device.
[0131] The intra prediction execution unit 610 generates a prediction block of the current block by performing intra prediction using the induced prediction mode. The intra prediction execution unit 610 performs intra prediction using the reference samples included in the leftmost and upper lines closest to the current block among the restored reference samples on the left side and the upper side without additional analysis. Alternatively, the intra prediction execution unit 610 determines the reference sample line to be used for intra prediction by analyzing an index that designates which reference sample line among the multi-line reference samples is to be used.
[0132] As one embodiment, the intra prediction execution unit 610 generates the predicted signal P b using the signal P1 predicted in the directional mode M having the largest value among the gradient histograms. Also, as described above, the intra prediction execution unit 610 uses the directional mode M having the largest value among the gradient histograms d b The predicted signal P1 and mode M, which has the second largest value. a The predicted signal P2 is weighted and summed with the predicted signal P d This generates the histograms. At this time, the weight values are determined in proportion to the histogram values of each mode.
[0133] In another embodiment, if the induced mode is a directional mode as described above, the intra-prediction execution unit 610 additionally utilizes a preset non-directional mode (e.g., planar). The intra-prediction execution unit 610 generates a prediction signal P according to the directional mode prediction, as shown in Equation 4. d And the predicted signal P predicted in non-directional mode. nd The final predicted signal P is generated by weighting and summing these two values.
[0134]
number
[0135] Here, b represents the shift value for integer operations. Furthermore, the weight values w1 and w2 used in the weighted sum are in the form of scalars or matrices. Such weight values are set according to the agreement between the video encoding device and the video decoding device.
[0136] Figure 10 is an illustrative diagram showing matrix-formatted weight values according to one embodiment of the present invention.
[0137] On the other hand, the weighting values of the matrix are determined based on the directional mode. For example, depending on the prediction mode, the matrix is realized in such a way that the weighting decreases as the reference sample used for prediction moves further away from the current block sample.
[0138] In another embodiment, the intra-prediction execution unit 610 analyzes the bitstream for an index that represents one of a predefined k weight values or matrix, and determines the weight value or weight matrix.
[0139] In another embodiment, when ISP technology is applied, the divided subblocks share the same intra-prediction mode, thus sharing the induced prediction mode of the current block. Alternatively, the intra-prediction device induces the prediction mode of the current subblock based on the restored sample of the previous subblock, and then performs intra-prediction on the current subblock using the induced prediction mode. In this case, the feasibility of inducing a subblock-specific prediction mode is implicitly determined based on the size of the subblock.
[0140] As illustrated in Figure 11, if the current block is vertically divided into two subblocks, the intra-prediction device induces a prediction mode for the current subblock based on the reconstructed sample of the previous subblock. On the other hand, while the example in Figure 11 shows a subblock that has been vertically divided, the same prediction mode for the current subblock is induced for a subblock that has been horizontally divided.
[0141] Furthermore, when ISP technology is applied, the intra-prediction device sequentially restores each subblock and performs intra-prediction of the current subblock using the restored reference sample restored in the previous subblock.
[0142] Figure 12 is a flowchart showing an intra-prediction method using predictive mode induction according to one embodiment of the present invention.
[0143] The video decoding device analyzes the bitstream for a prediction mode induction flag that indicates whether the prediction mode for the current block can be induced (S1200). Meanwhile, the prediction mode induction flag is set by the video encoding device from the perspective of optimizing bitrate distortion and then transmitted to the video decoding device.
[0144] If the block currently contains picture boundaries, slice boundaries, tile boundaries, etc., and the left and / or upper reference samples are unavailable, the analysis of the prediction mode guidance flag is implicitly omitted.
[0145] In another embodiment, the video decoder analyzes one or more flags from the bitstream indicating whether the predicted mode of the current block is a non-directional mode such as planar, DC mode, or matrix-based mode, before determining whether predictive mode induction is possible. If all of the relevant flags are false, the predictive mode induction flags are analyzed.
[0146] The video decoding device checks the prediction mode guidance flag (S1202).
[0147] If the prediction mode guidance flag is false, the video decoder analyzes the prediction mode of the current block from the bitstream and then generates a prediction block for the current block using the analyzed prediction mode (S1204).
[0148] On the other hand, if the prediction mode guidance flag is true, the video decoding device proceeds to the next steps (S1210~S1216) without analyzing the intra-prediction mode of the current block.
[0149] The video decoding device determines the calculation region used to calculate the gradient value from the surrounding reconstructed samples of the current block (S1210).
[0150] The video decoding device determines the three rows of restored reference pixels located to the left and above the current block as calculation areas. At this time, the lengths of the calculation areas located to the left and above are set based on the width and height of the current block.
[0151] In another embodiment, the video decoding device analyzes a flag indicating the calculation area from the bitstream and sets the area to be used as the calculation area from the left or upper row relative to the current block.
[0152] In another embodiment, the calculation domain is implicitly determined according to an agreement between the video encoding device and the video decoding device. In this case, the step of determining the calculation domain is omitted.
[0153] The video decoding device calculates a histogram of the directional mode gradient in the calculation domain for the current block (S1212).
[0154] The video decoder calculates vertical and horizontal gradient values for the restored reference sample located on the second line relative to the current block, using a pre-configured boundary detection filter. The video decoder uses the vertical and horizontal gradient values to calculate the gradient direction and gradient size for the restored reference sample located on the second line. After calculating the intra-predicted directional mode closest to the gradient direction, the video decoder calculates a histogram of directional mode gradients by accumulating the gradient size on the histogram corresponding to the calculated directional mode.
[0155] The video decoder subsamples the restored reference sample located on the second line based on a preset sampling interval, and then calculates a gradient histogram for the subsampled pixels. At this time, the preset sampling interval and the position of the subsampled pixels are defined according to the agreement between the video encoder and the video decoder. Alternatively, the pixel sampling interval and subsampling position are determined based on the current block size and / or the aspect ratio of the current block.
[0156] The video decoding device derives the prediction mode for the current block based on the gradient histogram (S1214). The deriveted intra-prediction mode is either a directional mode or a non-directional mode.
[0157] The video decoding device induces the directional mode as follows:
[0158] The video decoding device determines the first directional mode, which has the largest value in the gradient histogram, as the prediction mode for the current block.
[0159] Alternatively, the video decoding device determines the first directional mode with the largest value in the gradient histogram, and the second directional mode with the second largest value, as the prediction modes for the current block.
[0160] In another embodiment, the video decoder further analyzes a prediction mode indicator flag that indicates either a first or second directional mode, and determines the prediction mode of the current block according to the analyzed flag value. The video decoder analyzes the prediction mode indicator flag from the bitstream if the difference between the histogram values of the first and second directional modes is less than or equal to a preset threshold. On the other hand, if the difference between the histogram values of the two modes is greater than a preset threshold, the video decoder omits the analysis of the prediction mode indicator flag. Alternatively, if the ratio between the difference between the two histogram values and the sum of the total histogram values is greater than or equal to a preset ratio, the video decoder omits the analysis of the prediction mode indicator flag.
[0161] The video decoding device induces a non-directional mode as follows:
[0162] If the histogram value of the first directional mode is less than a preset first threshold, or if the sum of the gradient histogram values is less than a preset second threshold, the video decoder determines the non-directional prediction mode as the prediction mode for the current block. In this case, the non-directional mode is either DC mode or planar mode.
[0163] In another embodiment, if the ratio between the sum of the histogram values of the first directional mode and the gradient histogram values is smaller than a preset ratio, the video decoder determines the non-directional prediction mode as the prediction mode for the current block.
[0164] In another embodiment, if the ratio between the sum of the gradient histogram values and the number of pixels used to calculate the gradient value is smaller than a preset first ratio, or if the ratio between the number of pixels used to calculate the histogram value and gradient value of the first directional mode is smaller than a preset second ratio, the video decoder determines the non-directional prediction mode as the prediction mode for the current block.
[0165] The video decoding device generates a predicted block for the current block by performing intra-prediction using the induced prediction mode (S1216).
[0166] The video decoder performs intra-prediction without additional analysis using the reference samples contained in the leftmost and uppermost lines closest to the current block from among the restored reference samples on the left and uppermost lines. Alternatively, the video decoder determines the reference sample lines to be used for intra-prediction by analyzing an index from the bitstream that specifies which reference sample lines from the multi-line reference samples to use.
[0167] In another embodiment, the video decoding device generates a first predicted block of the current block using a first directional mode, generates a second predicted block of the current block using a second directional mode, and then generates a predicted block of the current block by weighting the first and second predicted blocks together. At this time, the weight values are determined in proportion to the histogram values of each mode.
[0168] Although the flowcharts / timing diagrams in this specification describe each step as being performed sequentially, this is merely an illustrative example of the technical concept of the embodiments of the present invention. In other words, the flowcharts / timing diagrams are not limited to a chronological order, as they can be modified and adapted in various ways by changing the order described in the flowcharts / timing diagrams, or by performing one or more of the steps in parallel, without departing from the essential characteristics of the embodiments of the present invention.
[0169] It should be understood that the exemplary embodiments described above can be embodied in many different ways. Functions or methods described in one or more examples can be embodied in hardware, software, firmware, or any combination thereof. It should be understood that the functional components described herein are labeled as “units” to particularly emphasize their independent implementation.
[0170] On the other hand, the various functions or methods described in this embodiment are embodied in instruction words stored on a non-temporary recording medium that are read and executed by one or more processors. The non-temporary recording medium includes, for example, any kind of recording device in which data is stored in a form that can be read by a computer system. For example, the non-temporary recording medium includes storage media such as EPROM (erasable programmable read-only memory), flash drives, optical drives, magnetic hard drives, and solid-state drives (SSDs).
[0171] The above description is merely illustrative of the technical concept of the present invention, and any person with ordinary skill in the art to which the present invention belongs will be able to make various modifications and variations without departing from the essential characteristics of the present invention. Therefore, this embodiment is for illustrative purposes only and not to limit the technical concept of the present invention, and the scope of the technical concept of the present invention is not limited by such embodiments. The scope of protection of the present invention should be interpreted by the claims, and all technical concepts within an equivalent scope should be interpreted as being included in the scope of rights of the present invention.
[0172] CROSS-REFERENCE TO RELATED APPLICATION This patent application claims priority over patent application No. 10-2021-0028795 filed in Korea on March 4, 2021, and patent application No. 10-2022-0026549 filed in Korea on March 2, 2022, and all contents of those applications are merged into this patent application as references. [Explanation of Symbols]
[0173] 110 Picture division section 120, 540 prediction section 122, 542 Intra Prediction Unit 124, 544 Interpretation Unit 130 Subtractor 140 Conversion Unit 145 Quantization section 150, 515 Sorting section 155 Entropy coding unit 160, 520 Inverse quantization section 165, 530 Inverse Transform Section 170, 550 Adder 180, 560 Loop Filter Section 182, 562 deblocking filters 184, 564 SAO (Sample Adaptive Offset) filters 186, 566 ALF (Adaptive Loop Filter) 190,570 memory 510 Entropy Decoding Unit 542 Intra Prediction Unit 602 Prediction mode induction feasibility determination unit 604 Gradient Calculation Domain Determination Unit 606 Histogram Calculation Department 608 Predictive Mode Guidance Unit 610 Intra Prediction Execution Unit
Claims
1. An intra-prediction method for the current block performed by a video decoding device, The steps include: analyzing a prediction mode induction flag from the bitstream that indicates whether or not the prediction mode induction of the current block is possible; The step includes confirming the prediction mode guidance flag, If the aforementioned prediction mode guidance flag is true, The steps include determining the calculation region used to calculate the gradient value, For the current block, the steps include: calculating a histogram of the gradient of the directional mode in the calculation domain; The steps include: inducing at least one prediction mode for the current block based on the gradient histogram; The step of generating a predicted block for the current block by performing an intra-prediction using the induced prediction mode, The step of calculating the histogram of the gradient is: A step of subsampling the restored reference sample in the calculation domain based on a predetermined sampling interval, An intra-prediction method characterized by comprising the step of calculating a gradient histogram for the subsampled restored reference sample.
2. The step further includes analyzing the bitstream for one or more flags indicating whether the prediction mode of the current block is one of the non-directional modes, The intra prediction method according to claim 1, characterized in that if all flags indicating one of the non-directional modes are false, the prediction mode guidance flag is analyzed.
3. The intra prediction method according to claim 1, characterized in that, if the prediction mode induction flag is true, the subblock division application flag is analyzed from the bitstream to determine whether or not the current block can be subblock divided.
4. The step of determining the calculation area involves determining the three rows of restored reference samples located to the left and above the current block as the calculation area, The intra prediction method according to claim 1, characterized in that the length of the calculation area located on the left and upper levels is set based on the width and height of the current block.
5. The process further includes the step of analyzing a flag indicating the calculation region from the bitstream, The intra prediction method according to claim 1, characterized in that, in accordance with the flag indicating the calculation area, the calculation area is set from the left or upper row based on the current block.
6. The intra-prediction method according to claim 1, characterized in that the step of calculating the histogram involves calculating vertical and horizontal gradient values for a subsampled restored reference sample located on the second line relative to the current block using a preset boundary detection filter.
7. The intra-prediction method according to claim 6, characterized in that the step of calculating the histogram is to use the vertical and horizontal gradient values to calculate the gradient direction and gradient size in the subsampled restored reference sample located on the second line, calculate the directional mode of the intra-prediction closest to the gradient direction, and accumulate the gradient size in the histogram corresponding to the directional mode.
8. The intra prediction method according to claim 1, characterized in that the preset sampling interval and the position of the subsampled restored reference sample are determined based on the size and / or aspect ratio of the current block.
9. The intra-prediction method according to claim 1, characterized in that the step of inducing the prediction mode is to determine the first directional mode having the largest value among the gradient histograms as the prediction mode for the current block.
10. The intra-prediction method according to claim 1, characterized in that the step of inducing the prediction mode is to determine a first directional mode having the largest value among the gradient histograms and a second directional mode having the second largest value as the prediction mode for the current block.
11. The intra-prediction method according to claim 9, characterized in that the step of inducing the prediction mode is to determine the non-directional prediction mode as the prediction mode for the current block if the histogram value of the first directional mode is smaller than a preset first threshold, or if the sum of the histogram values of the gradient is smaller than a preset second threshold.
12. The intra-prediction method according to claim 9, characterized in that the step of inducing the prediction mode is to determine the non-directional prediction mode as the prediction mode for the current block if the ratio between the sum of the histogram values of the first directional mode and the histogram values of the gradient is smaller than a preset ratio.
13. The intra-prediction method according to claim 3, characterized in that, if the sub-block division application flag is true, the current block is divided into sub-blocks and the induced prediction mode of the current block is shared as the same intra-prediction mode for the sub-blocks.
14. The intra prediction method according to claim 3, characterized in that, if the subblock division application flag is true, the current block is divided into subblocks, the prediction mode of the current subblock is induced based on the restored sample of the restored previous subblock, and whether or not the prediction mode of each subblock can be induced is implicitly determined based on the size of the subblock.
15. The intra-prediction method according to claim 10, characterized in that the step of generating the prediction block involves generating a first prediction block of the current block using the first directional mode, generating a second prediction block of the current block using the second directional mode, and then generating a prediction block of the current block by weighting the first prediction block and the second prediction block.
16. The intra-prediction method according to claim 15, characterized in that the step of generating the prediction block is, if the prediction mode of the current block is a directional mode, additionally using a non-directional mode, generating a third prediction block of the current block using the non-directional mode, and then generating a prediction block of the current block by weighting the first prediction block, the second prediction block, and the third prediction block.
17. An intra-prediction method for the current block performed by a video encoding device, The step includes encoding a prediction mode induction flag that indicates whether or not the prediction mode induction of the current block is possible. If the aforementioned prediction mode guidance flag is true, The steps include determining the calculation region used to calculate the gradient value, For the current block, the steps include: calculating a histogram of the gradient of the directional mode in the calculation domain; The steps include: inducing at least one prediction mode for the current block based on the gradient histogram; The steps include: performing intraprediction using the induced prediction mode to generate a predicted block of the current block; and predicting the current block by performing the following steps: The step of calculating the histogram of the gradient is: A step of subsampling the restored reference sample in the calculation domain based on a predetermined sampling interval, An intra-prediction method characterized by comprising the step of calculating a gradient histogram for the subsampled restored reference sample.
18. A method for providing video data to a video decoding device, The steps include encoding the aforementioned video data into a bitstream, The step of transmitting the bitstream to the video decoding device is included. The step of encoding the video data includes a step of encoding a prediction mode induction flag that indicates whether or not the prediction mode induction of the current block is possible. If the aforementioned prediction mode guidance flag is true, The steps include determining the calculation region used to calculate the gradient value, For the current block, the steps include: calculating a histogram of the gradient of the directional mode in the calculation domain; The steps include: inducing at least one prediction mode for the current block based on the gradient histogram; The steps include: performing intraprediction using the induced prediction mode to generate a predicted block of the current block; and predicting the current block by performing the following steps: The step of calculating the histogram of the gradient is: A step of subsampling the restored reference sample in the calculation domain based on a predetermined sampling interval, A method for providing video data, comprising the step of calculating a gradient histogram for the subsampled restored reference sample.