Method for extrapolation-based intra prediction
The use of an extrapolation filter for intra-prediction enhances video encoding efficiency and quality, addressing the challenges of increasing video sizes and resolutions, particularly in UHD, game broadcasting, and VR/AR applications.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2026-04-02
AI Technical Summary
Existing video compression technologies struggle to efficiently handle increasing video sizes, resolutions, and frame rates, requiring improved encoding efficiency and image quality for applications like UHD video, game broadcasting, and VR/AR content.
Implementing an extrapolation filter for intra-prediction that involves determining reference regions, deriving filter coefficients, and recursively applying them to predict samples within a prediction region, enhancing encoding and decoding processes.
Improves video encoding efficiency and quality, reducing network burden and energy consumption for devices handling high-definition and immersive video content.
Smart Images

Figure KR2025011823_02042026_PF_FP_ABST
Abstract
Description
Method for extrapolation-based intra-prediction
[0001] The present disclosure relates to an image encoding / decoding method, an apparatus, and a recording medium for storing a bitstream, and more specifically, to a method for constructing intra-prediction samples using an extrapolation filter in intra-prediction.
[0002] The following description merely provides background information related to the present invention and does not constitute prior art.
[0003] Because video data contains a large amount of data compared to audio or still image data, storing or transmitting it as is without compression processing requires significant hardware resources, including memory.
[0004] Therefore, typically, when storing or transmitting video data, the encoder compresses the video data for storage or transmission, and the decoder receives the compressed video data, decompresses it, and plays it. Such video compression technologies include H.264 / AVC, HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding), which improves coding efficiency by more than 30% compared to HEVC.
[0005] However, as video size, resolution, and frame rates are gradually increasing, and the amount of data that needs to be encoded is also growing accordingly, a new compression technology is required that offers better encoding efficiency and higher image quality improvement effects than existing compression technologies.
[0006] The present disclosure aims to provide an image encoding / decoding method and apparatus for configuring intra-predicted samples using an extrapolation filter in intra-predicted, and a recording medium for storing a bitstream generated by said image encoding method / apparatus.
[0007] According to an embodiment of the present disclosure, a method for restoring a current block performed by an image decoding device comprises: decoding a first reference region and a filter form of an extrapolation filter, wherein the first reference region includes restored reference samples of the current block and the filter form of the extrapolation filter includes an output pixel region and an input pixel region; determining a prediction region in which the current block is downsampled; determining a second reference region corresponding to the prediction region from the first reference region; deriving filter coefficients applied to the input pixel region based on the second reference region and the filter form; and generating predicted samples of the prediction region by recursively applying the filter coefficients to the surrounding restored reference samples of the prediction region and to the first predicted samples within the prediction region to sequentially predict samples within the prediction region.
[0008] According to another embodiment of the present disclosure, a method for encoding a current block performed by an image encoding device comprises: obtaining a first reference region and a filter shape of an extrapolation filter, wherein the first reference region includes restored reference samples of the current block and the filter shape of the extrapolation filter includes an output pixel region and an input pixel region; determining a prediction region in which the current block is downsampled; determining a second reference region corresponding to the prediction region from the first reference region; deriving filter coefficients applied to the input pixel region based on the second reference region and the filter shape; and generating predicted samples of the prediction region by recursively applying the filter coefficients to the surrounding restored reference samples of the prediction region and to the first predicted samples within the prediction region to sequentially predict samples within the prediction region.
[0009] According to another embodiment of the present disclosure, a method for providing video data to an image decoder comprises: encoding the video data into a bitstream; and transmitting the bitstream to the image decoder, wherein the step of encoding the video data comprises: obtaining a first reference region and a filter shape of an extrapolation filter, wherein the first reference region includes restored reference samples of a current block and the filter shape of the extrapolation filter includes an output pixel region and an input pixel region; determining a prediction region in which the current block is downsampled; determining a second reference region corresponding to the prediction region from the first reference region; deriving filter coefficients applied to the input pixel region based on the second reference region and the filter shape; and generating predicted samples of the prediction region by recursively applying the filter coefficients to the surrounding restored reference samples of the prediction region and to the first predicted samples within the prediction region to sequentially predict samples within the prediction region.
[0010] According to another embodiment of the present disclosure, a method for restoring a current block performed by an image decoding device comprises: decoding a first reference region and a filter form of an extrapolation filter, wherein the first reference region includes restoration reference samples of the current block and the filter form of the extrapolation filter includes an output pixel region and an input pixel region; deriving filter coefficients applied to the input pixel region based on the restoration samples of the first reference region or the residual samples of the first reference region using the filter form; generating predicted residual samples of the current block by recursively applying the filter coefficients to the surrounding residual samples of the current block and the first predicted residual samples within the current block to sequentially predict first residual samples within the current block; generating predicted samples by applying intra prediction or inter prediction to the current block; decoding second residual samples of the current block from a bitstream; and generating first residual samples by adding the second residual samples and the predicted residual samples. The present invention provides a method comprising the step of generating a restoration block of the current block by adding the first residual samples and the prediction samples.
[0011] According to another embodiment of the present disclosure, a method for encoding a current block performed by an image encoding device comprises: obtaining a first reference region and a filter form of an extrapolation filter, wherein the first reference region includes restored reference samples of the current block and the filter form of the extrapolation filter includes an output pixel region and an input pixel region; deriving filter coefficients applied to the input pixel region based on the restored samples of the first reference region or the residual samples of the first reference region using the filter form; generating predicted residual samples of the current block by recursively applying the filter coefficients to the surrounding restored residual samples of the current block and the first predicted residual samples within the current block to sequentially predict first residual samples within the current block; generating predicted samples by applying intra prediction or inter prediction to the current block; and generating first residual samples by subtracting the predicted samples from the samples of the current block. A method is provided comprising the steps of: generating second residual samples by subtracting the predicted residual samples from the first residual samples; and encoding the second residual samples.
[0012] According to another embodiment of the present disclosure, a method for restoring a current block performed by an image decoding device comprises: decoding a first reference region and a filter form of an extrapolation filter, wherein the first reference region includes restoration reference samples of the current block and the filter form of the extrapolation filter includes an output pixel region and an input pixel region; deriving filter coefficients applied to the input pixel region based on the filter form and the restoration samples of the first reference region; generating predicted samples of the current block by recursively applying the filter coefficients to surrounding restoration reference samples of the current block and to samples first predicted within the current block to sequentially predict samples within the current block; decoding residual samples of the current block without an inverse quantization process; and adding the residual samples and the predicted samples to generate a restored block of the current block.
[0013] According to another embodiment of the present disclosure, a method for encoding a current block performed by an image encoding device comprises: obtaining a first reference region and a filter shape of an extrapolation filter, wherein the first reference region includes reconstructed reference samples of the current block and the filter shape of the extrapolation filter includes an output pixel region and an input pixel region; deriving filter coefficients applied to the input pixel region based on the filter shape and the reconstructed samples of the first reference region; generating predicted samples of the current block by recursively applying the filter coefficients to surrounding reconstructed reference samples of the current block and to samples first predicted within the current block to sequentially predict samples within the current block; generating residual samples of the current block by subtracting the predicted samples from the samples of the current block; and encoding the residual samples without quantization.
[0014] As described above, by providing a video encoding / decoding method and device according to the present embodiment, and a recording medium storing a bitstream generated by the video encoding method / device, it is possible to improve video encoding efficiency and video quality.
[0015] In addition, by providing a video encoding / decoding method according to the present embodiment, it is possible to reduce the burden on the network based on bit rate reduction in various content such as UHD (Ultra High Definition) video, game broadcasting, 360-degree video streaming, VR / AR (Virtual Reality / Augmented Reality) video, online lectures, etc., and to reduce energy consumption for video playback devices.
[0016] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the technologies of the present disclosure.
[0017] Figure 2 is a diagram illustrating a method for dividing blocks using a QTBTTT (QuadTree plus BinaryTree TernaryTree) structure.
[0018] FIGS. 3a and 3b are diagrams showing a plurality of intra prediction modes including wide-angle intra prediction modes.
[0019] Figure 4 is an example diagram of the surrounding blocks of the current block.
[0020] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the technologies of the present disclosure.
[0021] Figure 6 is an example diagram showing the search area used in IntraTMP (Intra Template Matching Prediction) technology.
[0022] Figure 7 is an example diagram illustrating a method for deriving one or more intra-prediction modes in the Template-based intra-mode derivation (TIMD) technique.
[0023] Figures 8a and 8b are example diagrams showing the type of EIP (Extrapolation filter-based intra prediction) filter and the derivation of the EIP coefficients.
[0024] Figure 9 is an example diagram showing the generation order of EIP-based prediction samples.
[0025] Figure 10 is an example diagram illustrating RDPCM (Residual Differential Pulse Code Modulation) technology.
[0026] FIG. 11 is an exemplary diagram illustrating the application of EIP according to one embodiment of the present disclosure.
[0027] FIG. 12 is a flowchart illustrating a method for predicting a current block using an EIP according to one embodiment of the present disclosure.
[0028] Hereinafter, embodiments of the present invention will be described in detail with reference to the exemplary drawings. It should be noted that in assigning reference numerals to the components of each drawing, the same components are given the same reference numeral whenever possible, even if they are shown in different drawings. Furthermore, in describing these embodiments, if it is determined that a detailed description of related known components or functions could obscure the essence of these embodiments, such detailed description is omitted.
[0029] FIG. 1 is an exemplary block diagram of an image encoding device capable of implementing the technologies of the present disclosure. Hereinafter, the image encoding device and its sub-components will be described with reference to FIG. 1.
[0030] The video encoding device may be configured to include a picture splitting unit (110), a prediction unit (120), a subtractor (130), a conversion unit (140), a quantization unit (145), a reordering unit (150), an entropy encoding unit (155), an inverse quantization unit (160), an inverse conversion unit (165), an adder (170), a loop filter unit (180), and a memory (190).
[0031] Each component of the video encoding device may be implemented in hardware or software, or as a combination of hardware and software. Additionally, the function of each component may be implemented in software, and a microprocessor may be implemented to execute the software function corresponding to each component.
[0032] A single image (video) consists of one or more sequences containing multiple pictures. Each picture is divided into multiple regions, and encoding is performed for each region. For example, a single picture is divided into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile or slice is divided into one or more Coding Tree Units (CTUs). And each CTU is divided into one or more Coding Units (CUs) by a tree structure. Information applicable to each CU is encoded as the syntax of the CU, and information applicable to all CUs included in a single CTU is encoded as the syntax of the CTU. Additionally, information applicable to all blocks within a single slice is encoded as the syntax of the slice header, and information applicable to all blocks constituting one or more pictures is encoded in the Picture Parameter Set (PPS) or the picture header. Furthermore, information commonly referenced by multiple pictures is encoded in a Sequence Parameter Set (SPS). Also, information commonly referenced by one or more SPSs is encoded in a Video Parameter Set (VPS). Additionally, information commonly applicable to a single tile or tile group may be encoded as the syntax of a tile or tile group header. The syntax included in the SPS, PPS, slice header, and tile or tile group header may be referred to as high-level syntax.
[0033] The picture splitting unit (110) determines the size of the CTU. Information regarding the size of the CTU (CTU size) is encoded as a syntax of SPS or PPS and transmitted to an image decoding device.
[0034] The picture division unit (110) divides each picture constituting the image into multiple CTUs having a predetermined size, and then recursively divides the CTUs using a tree structure. The leaf nodes in the tree structure become the CUs, which are the basic units of encoding.
[0035] The tree structure may be a QuadTree (QT) in which an upper node (or parent node) is divided into four lower nodes (or child nodes) of equal size, a BinaryTree (BT) in which an upper node is divided into two lower nodes, a TernaryTree (TT) in which an upper node is divided into three lower nodes in a 1:2:1 ratio, or a structure that combines two or more of these QT, BT, and TT structures. For example, a QTBT (QuadTree plus BinaryTree) structure may be used, or a QTBTTT (QuadTree plus BinaryTree TernaryTree) structure may be used. Here, BTTT combined may be referred to as an MTT (Multiple-Type Tree).
[0036] Figure 2 is a diagram illustrating a method for dividing blocks using a QTBTTT structure.
[0037] As illustrated in FIG. 2, the CTU can first be split into a QT structure. Quadtree splitting can be repeated until the size of the splitting block reaches the minimum block size of the leaf node allowed in QT (MinQTSize). A first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of the lower layer is encoded by the entropy encoder (155) and signaled to the image decoder. If the leaf node of the QT is not larger than the maximum block size of the root node allowed in BT (MaxBTSize), it can be further split into one or more of the BT structure or TT structure. In the BT structure and / or TT structure, multiple splitting directions may exist. For example, there may be two directions in which the block of the corresponding node is split horizontally and vertically. As shown in Figure 2, when MTT splitting begins, a second flag (mtt_split_flag) indicating whether the nodes have been split, and if splitting has occurred, a flag indicating the splitting direction (vertical or horizontal) and / or the splitting type (binary or ternary) are encoded by the entropy encoding unit (155) and signaled to the image decoding device.
[0038] Alternatively, prior to encoding the first flag (QT_split_flag) indicating whether each node is split into four nodes of the lower layer, the CU split flag (split_cu_flag) indicating whether the node is split may be encoded. If the value of the CU split flag (split_cu_flag) indicates that it is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a coding unit (CU), which is the basic unit of encoding. If the value of the CU split flag (split_cu_flag) indicates that it is split, the video encoding device starts encoding from the first flag in the manner described above.
[0039] When QTBT is used as another example of a tree structure, there may be two types: a type that divides the block of the corresponding node horizontally into two blocks of the same size (i.e., symmetric horizontal splitting) and a type that divides it vertically (i.e., symmetric vertical splitting). A splitting flag (split_flag) indicating whether each node of the BT structure is split into a block of a lower layer and splitting type information indicating the type of splitting are encoded by the entropy encoding unit (155) and transmitted to the image decoding device. Meanwhile, there may also be an additional type that divides the block of the corresponding node into two blocks of an asymmetric shape. The asymmetric shape may include a shape that divides the block of the corresponding node into two rectangular blocks with a size ratio of 1:3, or a shape that divides the block of the corresponding node diagonally.
[0040] A CU can have various sizes depending on the QTBT or QTBTTT partitioning from a CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of QTBTTT) is referred to as the 'current block'. Depending on the adoption of QTBTTT partitioning, the shape of the current block may be not only square but also rectangular.
[0041] The prediction unit (120) predicts the current block and generates a prediction block. The prediction unit (120) includes an intra prediction unit (122) and an inter prediction unit (124).
[0042] Generally, current blocks within a picture can each be predictively coded. Typically, the prediction of a current block can be performed using an intra-prediction technique (using data from the picture containing the current block) or an inter-prediction technique (using data from a picture coded prior to the picture containing the current block). Inter-prediction includes both unidirectional and bidirectional prediction.
[0043] The intra prediction unit (122) predicts pixels within the current block using pixels (reference pixels) located around the current block within the current picture containing the current block. Multiple intra prediction modes exist depending on the prediction direction. For example, as shown in FIG. 3a, multiple intra prediction modes may include two non-directional modes, including Planar mode and DC mode, and 65 directional modes. The surrounding pixels to be used and the calculation formula are defined differently for each prediction mode.
[0044] For efficient directional prediction for a rectangular current block, directional modes (intra-prediction modes 67 through 80 and -1 through -14) illustrated by dashed arrows in FIG. 3b may be additionally used. These may be referred to as "wide angle intra-prediction modes." In FIG. 3b, the arrows indicate corresponding reference samples used for prediction and do not indicate the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. Wide angle intra-prediction modes are modes that perform prediction in the opposite direction of a specific directional mode without additional bit transmission when the current block is rectangular. Among the wide angle intra-prediction modes, some wide angle intra-prediction modes available for the current block may be determined by the ratio of the width to the height of the rectangular current block. For example, wide-angle intra-prediction modes with an angle less than 45 degrees (intra-prediction modes 67 to 80) are available when the current block is a rectangular shape with a height less than the width, and wide-angle intra-prediction modes with an angle greater than -135 degrees (intra-prediction modes -1 to -14) are available when the current block is a rectangular shape with a width greater than the height.
[0045] The intra prediction unit (122) can determine the intra prediction mode to use for encoding the current block. In some examples, the intra prediction unit (122) may encode the current block using several intra prediction modes and select an appropriate intra prediction mode to use from the tested modes. For example, the intra prediction unit (122) may calculate the rate-distortion values using a rate-distortion analysis of several tested intra prediction modes and select the intra prediction mode having the best rate-distortion features among the tested modes.
[0046] The intra prediction unit (122) selects one intra prediction mode among a plurality of intra prediction modes and predicts the current block using a calculation formula and surrounding pixels (reference pixels) determined according to the selected intra prediction mode. Information regarding the selected intra prediction mode is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.
[0047] The inter prediction unit (124) generates a prediction block for the current block using a motion compensation process. The inter prediction unit (124) searches for the block most similar to the current block within a reference picture that is encoded and decoded before the current picture, and generates a prediction block for the current block using the searched block. Then, it generates a motion vector (MV) corresponding to the displacement between the current block in the current picture and the prediction block in the reference picture. Generally, motion estimation is performed on the lumina component, and the motion vector calculated based on the lumina component is used for both the lumina component and the chroma component. Motion information including information about the reference picture used to predict the current block and information about the motion vector is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.
[0048] The inter prediction unit (124) may perform interpolation on a reference picture or reference block to increase the accuracy of the prediction. That is, subsamples between two consecutive integer samples are interpolated by applying filter coefficients to a plurality of consecutive integer samples including those two integer samples. When the process of searching for the block most similar to the current block is performed for the interpolated reference picture, the motion vector can be expressed with precision in fractional units rather than precision in integer sample units. The precision or resolution of the motion vector can be set differently for each unit of the target area to be encoded, such as slice, tile, CTU, CU, etc. When such Adaptive Motion Vector Resolution (AMVR) is applied, information regarding the motion vector resolution to be applied to each target area must be signaled for each target area. For example, if the target area is a CU, information regarding the motion vector resolution applied to each CU is signaled. The information regarding the motion vector resolution may be information indicating the precision of the difference motion vector described later.
[0049] Meanwhile, the inter prediction unit (124) can perform inter prediction using bi-prediction. In the case of bi-prediction, two reference pictures and two motion vectors representing the block location most similar to the current block within each reference picture are used. The inter prediction unit (124) selects a first reference picture and a second reference picture from the reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively, and generates a first reference block and a second reference block by searching for a block similar to the current block within each reference picture. Then, it generates a prediction block for the current block by averaging or weighting the first reference block and the second reference block. Then, it transmits motion information containing information about the two reference pictures used to predict the current block and information about the two motion vectors to the entropy encoding unit (155). Here, reference picture list 0 consists of restored pictures that are prior to the current picture in the display order, and reference picture list 1 may consist of restored pictures that are prior to the current picture in the display order. However, this is not necessarily limited to this, and restored pictures prior to the current picture in the display order may be additionally included in reference picture list 0, and conversely, restored pictures prior to the current picture may be additionally included in reference picture list 1.
[0050] Various methods can be used to minimize the amount of bits required to encode motion information.
[0051] For example, if the reference picture and motion vector of the current block are identical to the reference picture and motion vector of a neighboring block, the motion information of the current block can be transmitted to an image decoder by encoding information that can identify the neighboring block. This method is called 'merge mode'.
[0052] In merge mode, the inter prediction unit (124) selects a predetermined number of merge candidate blocks (hereinafter referred to as 'merge candidates') from the surrounding blocks of the current block.
[0053] As for the surrounding blocks for deriving merge candidates, as shown in FIG. 4, all or part of the left block (A0), bottom-left block (A1), top block (B0), top-right block (B1), and top-left block (B2) adjacent to the current block within the current picture may be used. Additionally, a block located within a reference picture (which may be the same as or different from the reference picture used to predict the current block) other than the current picture where the current block is located may be used as a merge candidate. For example, a block located at the same position as the current block within the reference picture (co-located block) or a block adjacent to that same position may be additionally used as a merge candidate. If the number of merge candidates selected by the method described above is less than a preset number, a 0 vector is added to the merge candidates.
[0054] The inter prediction unit (124) constructs a merge list containing a predetermined number of merge candidates using these surrounding blocks. Among the merge candidates included in the merge list, it selects a merge candidate to be used as movement information for the current block and generates merge index information to identify the selected candidate. The generated merge index information is encoded by the entropy encoding unit (155) and transmitted to the image decoding device.
[0055] Merge skip mode is a special case of merge mode; after quantization, when all transform coefficients for entropy coding are close to zero, only neighbor block selection information is transmitted without transmitting residual signals. By utilizing merge skip mode, relatively high coding efficiency can be achieved in images with minimal motion, still images, and screen content images.
[0056] Hereinafter, merge mode and merge skip mode will be collectively referred to as merge / skip mode.
[0057] Another method for encoding motion information is the AMVP (Advanced Motion Vector Prediction) mode.
[0058] In AMVP mode, the inter-prediction unit (124) derives predicted motion vector candidates for the motion vector of the current block using the surrounding blocks of the current block. As surrounding blocks used to derive predicted motion vector candidates, all or part of the left block (A0), bottom-left block (A1), top block (B0), top-right block (B1), and top-left block (B2) adjacent to the current block within the current picture shown in FIG. 4 may be used. Additionally, blocks located within a reference picture (which may be the same as or different from the reference picture used to predict the current block) other than the current picture where the current block is located may be used as surrounding blocks to derive predicted motion vector candidates. For example, blocks located at the same position as the current block within the reference picture (co-located blocks) or blocks adjacent to the blocks at the same position may be used. If the number of motion vector candidates is less than a preset number by the method described above, a 0 vector is added to the motion vector candidates.
[0059] The inter prediction unit (124) derives predicted motion vector candidates using the motion vectors of the surrounding blocks and determines a predicted motion vector for the current block's motion vector using the predicted motion vector candidates. Then, it calculates a difference motion vector by subtracting the predicted motion vector from the current block's motion vector.
[0060] Predicted motion vectors can be obtained by applying a predefined function (e.g., median, mean operation, etc.) to the predicted motion vector candidates. In this case, the image decoder is also aware of the predefined function. Furthermore, since the surrounding blocks used to derive the predicted motion vector candidates have already been encoded and decoded, the image decoder is also aware of the motion vectors of those surrounding blocks. Therefore, the image decoder does not need to encode information to identify the predicted motion vector candidates. Consequently, in this case, information regarding the difference motion vector and the reference picture used to predict the current block is encoded.
[0061] Meanwhile, the predicted motion vector may be determined by selecting one of the predicted motion vector candidates. In this case, information for identifying the selected predicted motion vector candidate is additionally encoded, along with information about the difference motion vector and information about the reference picture used to predict the current block.
[0062] The subtractor (130) generates a residual block by subtracting the prediction block generated by the intra prediction unit (122) or the inter prediction unit (124) from the current block.
[0063] The conversion unit (140) converts residual signals within a residual block having pixel values in a spatial domain into conversion coefficients in the frequency domain. The conversion unit (140) can convert the residual signals within the residual block using the entire size of the residual block as the conversion unit, or it can divide the residual block into multiple sub-blocks and use the sub-blocks as the conversion unit to perform the conversion. Alternatively, it can divide the residual signals into two sub-blocks, a conversion area and a non-conversion area, and use only the conversion area sub-block as the conversion unit to convert the residual signals. Here, the conversion area sub-block may be one of two rectangular blocks having a size ratio of 1:1 with respect to the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicating that only the sub-block has been converted, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or position information (cu_sbt_pos_flag) are encoded by the entropy encoding unit (155) and signaled to the image decoding device. Additionally, the size of the converted area sub-block may have a size ratio of 1:3 with respect to the horizontal axis (or vertical axis), and in this case, a flag (cu_sbt_quad_flag) distinguishing the corresponding division is additionally encoded by the entropy encoding unit (155) and signaled to the image decoding device.
[0064] Meanwhile, the transformation unit (140) can perform transformations on the residual block individually in the horizontal and vertical directions. For the transformation, various types of transformation functions or transformation matrices may be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a Multiple Transform Set (MTS). The transformation unit (140) can select one pair of transformation functions with the best transformation efficiency among the MTS and transform the residual block in the horizontal and vertical directions, respectively. Information (mts_idx) regarding the selected pair of transformation functions among the MTS is encoded by the entropy encoding unit (155) and signaled to the image decoder.
[0065] The quantization unit (145) quantizes the transformation coefficients output from the transformation unit (140) using quantization parameters and outputs the quantized transformation coefficients to the entropy encoding unit (155). The quantization unit (145) may quantize the associated residual block directly without transformation for any block or frame. The quantization unit (145) may apply different quantization coefficients (scaling values) depending on the position of the transformation coefficients within the transformation block. The quantization matrix applied to the quantized transformation coefficients arranged in two dimensions can be encoded and signaled to an image decoder.
[0066] The reordering unit (150) can perform reordering of coefficient values for quantized residual values.
[0067] The reordering unit (150) can convert a two-dimensional coefficient array into a one-dimensional coefficient sequence using coefficient scanning. For example, the reordering unit (150) can output a one-dimensional coefficient sequence by scanning from DC coefficients to coefficients in the high-frequency range using a zig-zag scan or a diagonal scan. Depending on the size of the conversion unit and the intra-prediction mode, a vertical scan that scans the two-dimensional coefficient array in the column direction and a horizontal scan that scans the two-dimensional block-shaped coefficients in the row direction may be used instead of a zig-zag scan. That is, depending on the size of the conversion unit and the intra-prediction mode, the scanning method to be used among a zig-zag scan, a diagonal scan, a vertical scan, and a horizontal scan may be determined.
[0068] The entropy encoding unit (155) generates a bitstream by encoding a sequence of one-dimensional quantized transformation coefficients output from the reordering unit (150) using various encoding methods such as CABAC (Context-based Adaptive Binary Arithmetic Code) and Exponential Golomb.
[0069] Additionally, the entropy encoding unit (155) encodes information related to block division, such as CTU size, CU division flag, QT division flag, MTT division type, and MTT division direction, so that the video decoder can divide the block in the same way as the video encoding unit. Additionally, the entropy encoding unit (155) encodes information regarding a prediction type indicating whether the current block is encoded by intra prediction or by inter prediction, and encodes intra prediction information (i.e., information regarding the intra prediction mode) or inter prediction information (information regarding the encoding mode of motion information (merge mode or AMVP mode), the merge index in the case of merge mode, and the reference picture index and difference motion vector in the case of AMVP mode) according to the prediction type. Additionally, the entropy encoding unit (155) encodes information related to quantization, i.e., information regarding quantization parameters and information regarding the quantization matrix.
[0070] The inverse quantization unit (160) inversely quantizes the quantized transformation coefficients output from the quantization unit (145) to generate transformation coefficients. The inverse transformation unit (165) converts the transformation coefficients output from the inverse quantization unit (160) from the frequency domain to the spatial domain to restore the residual block.
[0071] The adder (170) restores the current block by adding the restored residual block and the prediction block generated by the prediction unit (120). The pixels within the restored current block are used as reference pixels when intra-predicting the next block in sequence.
[0072] The loop filter section (180) performs filtering on the restored pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. caused by block-based prediction and transformation / quantization. The loop filter section (180) may include all or part of a deblocking filter (182), a SAO (Sample Adaptive Offset) filter (184), and an ALF (Adaptive Loop Filter, 186) as an in-loop filter.
[0073] The deblocking filter (182) filters the boundaries between restored blocks to remove blocking artifacts caused by block-unit encoding / decoding, and the SAO filter (184) and ALF (186) perform additional filtering on the deblocking filtered image. The SAO filter (184) and ALF (186) are filters used to compensate for the difference between restored pixels and original pixels caused by lossy coding. The SAO filter (184) improves not only subjective image quality but also encoding efficiency by applying an offset in CTU units. In contrast, the ALF (186) performs block-unit filtering, and compensates for distortion by applying different filters by distinguishing the degree of edge and change of the corresponding block. Information regarding the filter coefficients to be used in the ALF can be encoded and signaled to an image decoder.
[0074] The restored blocks filtered through the deblocking filter (182), SAO filter (184), and ALF (186) are stored in memory (190). Once all blocks within a picture are restored, the restored picture can be used as a reference picture for inter-predicting blocks within a picture to be encoded later.
[0075] The video encoding device can store the bitstream of encoded video data on a non-transient recording medium or transmit it to a video decoding device using a communication network.
[0076] FIG. 5 is an exemplary block diagram of an image decoding device capable of implementing the technologies of the present disclosure. Hereinafter, the image decoding device and its sub-components will be described with reference to FIG. 5.
[0077] The image decoding device may be configured to include an entropy decoding unit (510), a reordering unit (515), an inverse quantization unit (520), an inverse transformation unit (530), a prediction unit (540), an adder (550), a loop filter unit (560), and a memory (570).
[0078] Similar to the image encoding device of FIG. 1, each component of the image decoding device may be implemented in hardware or software, or in combination of hardware and software. Additionally, the function of each component may be implemented in software, and a microprocessor may be implemented to execute the function of the software corresponding to each component.
[0079] The entropy decoding unit (510) determines the current block to be decoded by decoding the bitstream generated by the video encoding device and extracting information related to block division, and extracts prediction information, information on residual signals, etc., necessary to restore the current block.
[0080] The entropy decoding unit (510) extracts information about the CTU size from the SPS (Sequence Parameter Set) or PPS (Picture Parameter Set) to determine the size of the CTU and divides the picture into CTUs of the determined size. Then, the CTU is determined as the top layer of the tree structure, i.e., the root node, and divides the CTU using the tree structure by extracting division information for the CTU.
[0081] For example, when splitting a CTU using a QTBTTT structure, first, a first flag (QT_split_flag) related to QT splitting is extracted to split each node into four nodes of the lower layer. Then, for the nodes corresponding to the leaf nodes of QT, a second flag (mtt_split_flag) related to MTT splitting and splitting direction (vertical / horizontal) and / or splitting type (binary / ternary) information are extracted to split the corresponding leaf nodes into an MTT structure. Accordingly, each node below the leaf nodes of QT is recursively split into a BT or TT structure.
[0082] As another example, when splitting a CTU using the QTBTTT structure, a CU splitting flag (split_cu_flag) indicating whether to split the CU is first extracted, and if the block is split, a first flag (QT_split_flag) is extracted. During the splitting process, each node may undergo zero or more iterative MTT splittings after zero or more iterative QT splittings. For example, the CTU may undergo MTT splitting immediately, or conversely, only multiple QT splittings may occur.
[0083] As another example, when splitting a CTU using a QTBT structure, a first flag (QT_split_flag) related to the splitting of QT is extracted to split each node into four nodes of the lower layer. Then, for the nodes corresponding to the leaf nodes of QT, a split flag (split_flag) indicating whether to further split into BTs and split direction information are extracted.
[0084] Meanwhile, when the entropy decoding unit (510) determines the current block to be decoded using the division of the tree structure, it extracts information regarding the prediction type indicating whether the current block is intra-predicted or inter-predicted. If the prediction type information indicates intra-predicted, the entropy decoding unit (510) extracts syntax elements for the intra-predicted information (intra-predicted mode) of the current block. If the prediction type information indicates inter-predicted, the entropy decoding unit (510) extracts syntax elements for the inter-predicted information, namely information indicating the motion vector and the reference picture that the motion vector refers to.
[0085] Additionally, the entropy decoder (510) extracts information regarding quantization-related information and information regarding residual signals, as well as information regarding the quantized transformation coefficients of the current block.
[0086] The reordering unit (515) can change the sequence of one-dimensional quantized transformation coefficients entropy-decoded in the entropy decoding unit (510) back into a two-dimensional coefficient array (i.e., block) in the reverse order of the coefficient scanning order performed by the image encoding device.
[0087] The inverse quantization unit (520) inversely quantizes the quantized transformation coefficients and inversely quantizes the quantized transformation coefficients using quantization parameters. The inverse quantization unit (520) may apply different quantization coefficients (scaling values) to the quantized transformation coefficients arranged in two dimensions. The inverse quantization unit (520) may perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from an image encoding device to a two-dimensional array of quantized transformation coefficients.
[0088] The inverse transformation unit (530) generates a residual block for the current block by inversely transforming the inversely quantized transformation coefficients from the frequency domain to the spatial domain and restoring the residual signals.
[0089] Additionally, when the inverse transformation unit (530) inversely transforms only a part of the transformation block (sub-block), it extracts a flag (cu_sbt_flag) indicating that only the sub-block of the transformation block has been transformed, information on the directionality (vertical / horizontal) of the sub-block (cu_sbt_horizontal_flag) and / or information on the position of the sub-block (cu_sbt_pos_flag), restores residual signals by inversely transforming the transformation coefficients of the corresponding sub-block from the frequency domain to the spatial domain, and creates a final residual block for the current block by filling the areas that have not been inversely transformed with "0" values of residual signals.
[0090] Additionally, when MTS is applied, the inverse transformation unit (530) determines a transformation function or transformation matrix to be applied in the horizontal and vertical directions, respectively, using MTS information (mts_idx) signaled from the video encoding device, and performs an inverse transformation on the transformation coefficients within the transformation block in the horizontal and vertical directions using the determined transformation function.
[0091] The prediction unit (540) may include an intra prediction unit (542) and an inter prediction unit (544). The intra prediction unit (542) is activated when the prediction type of the current block is an intra prediction, and the inter prediction unit (544) is activated when the prediction type of the current block is an inter prediction.
[0092] The intra prediction unit (542) determines the intra prediction mode of the current block among a plurality of intra prediction modes from the syntax elements for the intra prediction mode extracted from the entropy decoding unit (510), and predicts the current block using reference pixels around the current block according to the intra prediction mode.
[0093] The inter prediction unit (544) determines the motion vector of the current block and the reference picture that the motion vector refers to using the syntax elements for the inter prediction mode extracted from the entropy decoding unit (510), and predicts the current block using the motion vector and the reference picture.
[0094] The adder (550) restores the current block by adding the residual block output from the inverse transformation unit (530) and the prediction block output from the inter prediction unit (544) or the intra prediction unit (542). The pixels within the restored current block are used as reference pixels when intra-predicting the block to be decoded later.
[0095] The loop filter section (560) may include a deblocking filter (562), an SAO filter (564), and an ALF (566) as an in-loop filter. The deblocking filter (562) deblocks the boundaries between restored blocks to remove blocking artifacts caused by block-unit decoding. The SAO filter (564) and the ALF (566) perform additional filtering on the restored blocks after deblocking filtering to compensate for the difference between the restored pixels and the original pixels caused by lossy coding. The filter coefficients of the ALF are determined using information about the filter coefficients decoded from the bitstream.
[0096] The restored blocks filtered through the deblocking filter (562), SAO filter (564), and ALF (566) are stored in memory (570). When all blocks within a picture are restored, the restored picture is used as a reference picture to inter-predict blocks within the picture to be encoded later.
[0097] The present embodiment relates to the encoding and decoding of an image (video) as described above. More specifically, the present invention provides an image encoding / decoding method and apparatus for configuring intra-prediction samples using an extrapolation filter in intra-prediction, and a recording medium for storing a bitstream generated by the image encoding method / apparatus.
[0098] The following embodiments may be performed by an intra prediction unit (122) within a video encoding apparatus. Additionally, the following embodiments may be performed by an intra prediction unit (542) within a video decoding apparatus.
[0099] The video encoding device can generate signaling information related to the present embodiment in terms of rate distortion optimization during the encoding of the current block. The video encoding device can encode the signaling information using the entropy encoding unit (155) and then transmit it to the video decoder. The video decoder can decode the signaling information related to the decoding of the current block from the bitstream using the entropy decoder (510).
[0100] In the following description, the term 'target block' may be used interchangeably with 'current block' or 'Coding Unit (CU).' Alternatively, 'target block' may refer to a specific area of a Coding Unit.
[0101] Also, a value of one flag being true indicates that the flag is set to 1. Also, a value of one flag being false indicates that the flag is set to 0.
[0102] The decoder-side includes all or part of an inverse quantizer (160), an inverse transform (165), a prediction unit (120), an adder (170), a loop filter (180), and a memory (190) in the image encoding device illustrated in FIG. 1. Alternatively, the decoder-side includes all or part of an inverse quantizer (520), an inverse transform (530), a prediction unit (540), an adder (550), a loop filter (560), and a memory (570) in the image decoding device illustrated in FIG. 5. In relation to a series of decoding processes, the decoder-side of the image encoding device and the decoder-side of the image decoding device perform the same operation. The image encoding device determines information related to the operation of the decoder-side and signals the determined information to the image decoding device. The image decoding device can decode the signaled information and operate the decoder-side based on the decoded information.
[0103] I. VVC Intra-prediction Technology
[0104] In the intra prediction of VVC, the angular prediction directions are subdivided up to 65, as shown in the example of Fig. 3a. Depending on the prediction angle of the intra prediction mode, prediction modes 2 through 66 may be used. By introducing Wide-Angle Intra Prediction (WAIP), prediction modes -14 through -1 and 67 through 80, which are angular modes with larger angles depending on the aspect ratio of the block, may be used. In intra prediction, the prediction block can be generated based on 67 Intra Prediction Modes (IPM). The 67 IPMs refer to 67 intra prediction modes that can be signaled according to the aspect ratio of the block, among prediction modes -14 through 80, including non-directional prediction modes such as Planar and DC modes.
[0105] In addition, the intra prediction technology of VVC can generate intra prediction blocks using technologies such as MRLP (Multiple Reference Line intra Prediction), CCLM (Cross-Component Linear Model), PDPC (Position Dependent intra Prediction Combination), ISP (Intra Sub-Partitions), and MIP (Matrix-based Intra Prediction).
[0106] I-1. IBC (Intra Block Copy)
[0107] IBC performs intra-prediction of the current block by generating a prediction block of the current block by copying the reference block within the same frame using a block vector.
[0108] The video encoding device performs block matching to derive the optimal block vector. Here, the block vector represents the displacement from the current block to the reference block. To increase encoding efficiency, the video encoding device may not transmit the block vector as is, but instead divide it into a Block Vector Predictor (BVP) and a Block Vector Difference (BVD), encode the BVP and BVD, and then transmit them to the video decoder. An integer is used as the pixel precision of the block vector. That is, pixel precision of 1 / 2 or 1 / 4 is not used.
[0109] In terms of utilizing block vectors, IBC possesses the characteristics of inter-prediction. Therefore, IBC can be classified into IBC merge / skip mode and IBC AMVP mode.
[0110] In IBC merge / skip mode, the video encoder constructs an IBC merge list. For the purpose of optimizing encoding efficiency, the video encoder selects one block vector from the candidates included in the IBC merge list and can use the selected candidate as the Block Vector Predictor (BVP). The video encoder determines a merge index that indicates the selected block vector. However, the video encoder does not generate a BVD. The video encoder encodes the merge index and transmits it to the video decoder. The IBC merge list can be constructed in the same way by both the video encoder and the video decoder. After decoding the merge index, the video decoder can generate a block vector from the IBC merge list using the merge index.
[0111] In the case of IBC skip mode, the video encoding device uses the same block vector transmission method as in IBC merge mode, but does not transmit the residual block corresponding to the difference between the current block and the predicted block.
[0112] In IBC AMVP mode, the video encoder determines a block vector and constructs an IBC AMVP list to optimize encoding efficiency. The video encoder determines a candidate index that designates one of the candidate block vectors included in the IBC AMVP list as the BVP. The video encoder calculates the BVD, which is the difference between the BVP and the block vector. Subsequently, the video encoder encodes the candidate index and the BVD and transmits them to the video decoder.
[0113] The image decoder decodes the candidate index and the BVD. The image decoder can recover the block vector by obtaining the BVP indicated by the candidate index from the IBC AMVP list and adding the BVP and the BVD.
[0114] I-2. Template Matching Prediction
[0115] Template Matching Prediction (TMP) searches for a prediction block that minimizes the difference between templates in a predefined recovery region, i.e., a search region, within the current frame. The template of the current block (hereinafter referred to as the 'current template') consists of top and left neighbor samples. In this case, the difference between the template of the current block and the template of the prediction block found in the search region is defined as a cost function. Template Matching Prediction determines the prediction block with the minimum cost as the prediction block of the current block. Sum of Absolute Differences (SAD) is used as the cost function.
[0116] In the Enhanced Compression Model (ECM), a next-generation technology of VVC, the IntraTMP (Intra Template Matching Prediction) technology sets L-shaped / left / top templates around the current block, searches for the template most similar to the current template in the original region of the current frame, and then uses a block adjacent to the searched template and of the same size as the current block as the prediction block for the current block. The IntraTMP technology searches for similar templates based on a cost function, using the Sum of Absolute Differences (SAD) as the cost function. The video encoder transmits whether to use the IntraTMP mode, and the video decoder can perform the same template matching task if the IntraTMP mode is applied.
[0117] Figure 6 is an example diagram showing the search area used in IntraTMP technology.
[0118] To save memory, the current CTU (coding tree unit) where the current block is located, the top-left CTU, the top CTU, and the left CTU may be limited to the possible search area. In FIG. 6, the search area is represented as R1 to R4. Additionally, among the possible search areas, the area generated by multiplying the width and height (w, h) of the current block by a preset constant a may be adaptively set as the search range. For example, by setting a=5, the search range SearchRange_w and SearchRange_h may be determined as shown in Equation 1.
[0119]
[0120] In mathematical formula 1, BlkW and BlkH represent the width and height of the current block.
[0121] The template search area in IntraTMP mode can be predefined according to an agreement between the video encoding device and the video decoder.
[0122] IntraTMP technology may include sub-modes such as a mode using a single template (TMP single), a technology that fuses multiple templates (TMP fusion), a sub-pixel precision mode, and a linear filter model mode.
[0123] In single template mode, the template search process includes two steps. In the first step, a search is performed at intervals of 3 pixels, and block vectors (BVs) specifying 30 candidate templates (i.e., templates for candidate reference blocks) are included in the candidate list. In the second step, 19 BVs are finally determined by performing an additional search at intervals of 1 pixel in a surrounding 3×3 area for the 30 BVs in the candidate list. Subsequently, the optimal reference block is selected in terms of rate distortion optimization, and a candidate list index indicating the finally selected BV is signaled from the video encoding device to the video decoder.
[0124] TMP fusion combines multiple reference templates and reference blocks using weights.
[0125] The sub-pixel precision mode searches for similar templates based on 1 / 2-Pel, 1 / 4-Pel, and 3 / 4-Pel precision.
[0126] The linear filter model mode filters the prediction block using a 6-tap filter. The 6-tap filter is a cross-shaped 5-tap filter with a bias added. The filter coefficients can be calculated based on the relationship between the current template and the reference template. The current block is predicted using the reference block to which the generated filter has been applied.
[0127] I-3. DIMD (Decoder-side Intra Mode Derivation) technology
[0128] The DIMD method sets reconstructed pixels adjacent to the current block as a template and derives directional intra prediction modes by performing gradient analysis on the set template. The DIMD method extracts orientation and magnitude information by analyzing the gradients in the vertical and horizontal directions within the template. The DIMD method can calculate the horizontal gradient Gx and the vertical gradient Gy, respectively, by applying a horizontal Sobel filter and a vertical Sobel filter to the pixel location at the center of the template. The orientation of the corresponding pixel location is calculated via atan(Gy / Gx), and the sum of the absolute values of Gx and Gy can be calculated as the magnitude of the orientation. The size of the template may be equal to or larger than the size of the Sobel filter. The template may include adjacent samples from the bottom-left and top-right of the current block.
[0129] The angle and magnitude calculated for each pixel location in the center of the template area are used to generate a Histogram of Gradient (HoG), which is constructed by accumulating the angle on the x-axis and the magnitude on the y-axis. The DIMD method can derive or determine corresponding intra-prediction modes and corresponding weights by mapping a number of angles (e.g., up to 5) with relatively high accumulated magnitude values on the HoG to directional modes. The intra-prediction modes determined by the DIMD method may be referred to as DIMD modes and can subsequently be utilized to generate the Most Probable Mode (MPM).
[0130] To form the final prediction block, the predictors of the intra prediction modes determined according to DIMD can be weighted and summed with the undirected predictor (based on the Planar or block vector).
[0131] Meanwhile, regarding intra prediction, the MPM technology utilizes the intra prediction mode of surrounding blocks when predicting the intra of the current block. The video encoding device generates an MPM list to include intra prediction modes derived from predefined locations spatially adjacent to the current block. The video encoding device can improve the encoding efficiency of the intra prediction mode by transmitting the index of the MPM list instead of the index of the prediction mode.
[0132] I-4. TIMD (Template-based Intra Mode Derivation) technology
[0133] The TIMD technology performs template prediction by applying the intra prediction mode of the MPM list to the template area around the current block. The TIMD technology calculates the Sum of Absolute Transformed Differences (SATD) cost between the predicted value in the template area and the restored template value. Based on the calculated cost, the TIMD technology selects the mode representing the minimum cost (costMode1) (hereinafter referred to as the first mode) and the mode representing the second lowest cost (costMode2) (hereinafter referred to as the second mode). The undirected mode having the lower SATD cost (costMode3) among DC or Planar is selected as the third mode. An undirected mode may also be used if the following conditions are satisfied.
[0134] - When the third non-directional mode is different from the first and second modes
[0135] - When costMode3 < 1.5 × costMode1
[0136] When both of the above two conditions are true, the prediction blocks according to the three intra prediction modes are combined using weights as in Equation 2.
[0137]
[0138] If either of the two aforementioned conditions is false, the first mode and the second mode are used. The TIMD technique determines whether to combine the weights of the two TIMD modes according to Equation 3.
[0139]
[0140] If Equation 3 is satisfied, the prediction blocks according to the two modes are weightedly combined. On the other hand, if Equation 3 is not satisfied, the prediction mode with the smaller cost is used. When the two modes are combined, a larger weight is assigned to the mode with the smaller SATD cost according to Equation 4.
[0141]
[0142] Figure 7 is an example diagram illustrating a method for deriving one or more intra-prediction modes in TIMD technology.
[0143] The size of the template area can be determined based on the size of the current block as shown in FIG. 7. If the width (W) or height (H) of the current block is greater than 8, the template size L1 or L2 is set to 4. If the width or height of the current block is 8 or less, L1 or L2 is set to 2.
[0144] I-5. EMRL (Extended Multiple Reference Line) technology
[0145] In VVC, the 0th, 1st, and 2nd lines are used as reference lines for intra-prediction. On the other hand, in ECM, the 0th, 1st, 3rd, 5th, 7th, and 12th lines can be used as extended reference lines.
[0146] I-6. EIP (Extrapolation filter-based Intra Prediction) technology
[0147] In ECM, EIP technology is utilized in the intra prediction process. An image decoder derives EIP filter coefficients (hereinafter used interchangeably with EIP coefficients) using three types of filters as shown in FIG. 8a and a reference sample area around the current block as shown in FIG. 8b, and predicts the current block using the derived EIP filter. In FIG. 8a, the EIP filter includes an input pixel area corresponding to the input (hereinafter used interchangeably with the input area) and an output pixel area corresponding to the output (hereinafter used interchangeably with the output area). The output area corresponds to a single pixel. To generate output samples, EIP coefficients are applied to samples in the input pixel area. In FIG. 8b, fWhidth and fHeight represent the height and width of the EIP filter, respectively, and leftSize and aboveSize define the reference sample area used for EIP.
[0148] The EIP technique includes 1) a process of deriving EIP coefficients, 2) a process of recursively generating prediction samples within the current block, and 3) a step of deriving a mode of the prediction block by analyzing the gradient of the predicted block, and selecting a kernel such as MTS (Multiple Transform Set) or LFNST (Low-frequency Non-separable Transform) using the derived mode.
[0149] The video decoder derives the EIP coefficients as follows.
[0150] For a filter selected from the three types in Fig. 8a, EIP coefficients can be calculated by moving pixels one step at a time in the reference sample area. EIP coefficients can be calculated using an auto-correlation matrix and a cross-correlation vector to predict output samples in the output area of the EIP. The auto-correlation matrix can be calculated based on the samples in the input area shown in Fig. 8a, and the cross-correlation vector can be calculated based on the input samples and the reconstructed samples in the output area.
[0151] The image decoder can select a filter type based on the syntax transmitted from the image encoding device.
[0152] The image decoder can generate predicted samples of the current block as follows. The image decoder predicts samples at position (x,y) using 15 filter taps as in Equation 5.
[0153]
[0154] In mathematical formula 5, pred(x,y) is the predicted value at position (x, y) within the current block, and c i is the EIP coefficient. t(x-offsetXi, y-offsetYi) represents the reconstructed sample or the predicted sample. The prediction process uses a recursive structure that uses the value output by EIP as input again.
[0155] Figure 9 is an example diagram showing the generation order of EIP-based prediction samples.
[0156] The video decoder generates predicted samples of EIP in a diagonal order from the upper right to the lower left, as shown in Fig. 9.
[0157] If the current block is coded in EIP mode, the image encoder may signal an EIP Merge Flag indicating whether information inherited from a previously coded block in EIP mode can be used. If the EIP Merge Flag is true, the image decoder may inherit the filter type and filter coefficients from the previously coded block in EIP mode. On the other hand, if the EIP Merge Flag is false, the image decoder parses an index indicating the type of filter and the type of reference sample region, and derives EIP coefficients based on the parsed index. Table 1 shows the syntax associated with EIP mode.
[0158]
[0159] After generating prediction samples of the current block using an EIP filter, the image decoder applies the DIMD process to the prediction samples to derive the prediction mode of the current block. For each prediction sample, the image decoder calculates the horizontal gradient and the vertical gradient and constructs a histogram between the gradient angles and magnitudes. Subsequently, a set of transformations for LFNST, NSPT (Non-separable primary transform), or MTS can be determined based on the prediction mode corresponding to the largest histogram frequency. That is, LFNST, NSPT, and MTS can utilize the selected directional mode when determining the transformation kernel of the primary transform or secondary transform.
[0160] I-7. RDPCM (Residual Differential Pulse Code Modulation) technology
[0161] RDPCM technology is a lossy compression method designed to improve the performance of lossless compression, performing additional predictions in the vertical or horizontal direction from the nearest pixel in relation to residual samples after intra-prediction.
[0162] In the HEVC extension standard, for lossy compression, RDPCM is applied to residual samples after intra and inter prediction when the transform unit (TU) is encoded in transform skipping (TS) mode. Transform skipping mode is a technique that directly entropies the residual signal without transforming it, and generally, its encoding performance is not superior compared to the Discrete Cosine Transform (DCT). However, since screen content video contains many residual components in the high-frequency band at the boundaries of graphic elements with high color contrast, transform skipping mode can be effectively utilized for compression. RDPCM can provide excellent compression performance by reducing the total energy of the residual components for entropy encoding during the transform skipping process.
[0163] RDPCM in HEVC / SCC (HEVC Screen Content Extension Standard) utilizes an implicit RDPCM method that performs RDPCM in the same direction after intra prediction in horizontal and vertical directions, and an explicit RDPCM method that selectively performs horizontal RDPCM (H-RDPCM) or vertical RDPCM (V-RDPCM) after inter prediction and transmits RDPCM prediction direction information to the video decoder side using a bitstream. In the case of lossy compression, RDPCM can be performed on a 4×4 residual block, as shown in the example of FIG. 10. As shown in FIG. 10, RDPCM first performs prediction using the residual gample of the nearest left column or top row among the encoded samples according to the direction of RDPCM. Residual sample r of an N×N size block i,j Regarding, second-order residual samples resulting from the application of RDPCM in the vertical direction It can be expressed as in mathematical equation 6.
[0164]
[0165] In Equation 6, Q(r) are the recovered residual samples containing quantization noise. In the case of vertical RDPCM, the image decoder uses secondary residual samples It encodes the second residual samples and transmits the encoded second residual samples to the image decoder. The image decoder restores the second residual samples to predict the residual samples of the next row. This prediction process proceeds sequentially across all rows of the block. Meanwhile, the image decoder restores the residual samples of the i-th row by sequentially adding the restored second residual samples as shown in Equation 7.
[0166]
[0167] In the case of implicit RDPCM, information regarding the prediction direction of RDPCM can be derived from a pre-decoded HEVC intra prediction mode. In the case of explicit RDPCM, information regarding the prediction direction of RDPCM can be decoded from a bitstream.
[0168] The following embodiments are described with reference to an image decoding device, but may be implemented identically or similarly in an image encoding device. Alternatively, the following embodiments are described with reference to the decoder side of the image decoding device, but may be implemented identically or similarly in the decoder side of an image encoding device.
[0169] II. Embodiments according to the present disclosure
[0170] The embodiments are methods for improving the accuracy of a prediction block and reducing complexity by applying EIP, and include: 1) a method of interpolating prediction samples by applying EIP when the block size is large; 2) a method of performing EIP in a residual pixel domain; and 3) a method of applying EIP to lossless compression.
[0171] For example, in the case of a large block, applying an EIP filter to all samples within the block is inefficient in terms of block processing efficiency. For instance, the current block can be downsampled to reduce the block size, and after generating a prediction block based on the EIP application, the prediction block can be upsampled again. Based on the aforementioned method, the complexity associated with EIP application can be reduced and processing efficiency improved.
[0172] As another example, EIP may be applied to residual samples generated by applying intra-prediction or inter-prediction to the current block. Hereinafter, the EIP applied to residual samples is referred to as REIP (residual EIP).
[0173] As another example, existing EIP technology has been applied to lossy compression, but it can also be applied to lossless compression.
[0174] FIG. 11 is an exemplary diagram illustrating the application of EIP according to one embodiment of the present disclosure.
[0175] The application of EIP when the current block size is large is described below.
[0176] In the case of a conventional EIP, as shown in FIG. 11 (a), the image decoder configures an EIP filter and uses the EIP filter to generate prediction samples of the current block in sequence along a diagonal direction from the upper right to the lower left. The output samples are used as input samples for the EIP to output the next adjacent prediction sample, as shown in Equation 5. Since the samples are predicted in sequence by applying the aforementioned method to all samples constituting the current block, the efficiency regarding block processing may be significantly reduced. In this embodiment, as shown in FIG. 11 (b), the image decoder downsamples the current block and applies the EIP to the downsampled current block to generate a prediction block.
[0177] In FIG. 11(a), region p and region q represent the current sample region (i.e., the current block) and the reference sample region, respectively. The height of region p is denoted by H and the width by W. If the size of the sample region p is large, the image decoder downsamples region p to generate region p'. For example, if H and W are greater than 32, the image decoder downsamples the current block and performs EIP on the downsampled current block. Region p' can be generated according to one of the following methods.
[0178] As an example, the size of region p' is determined by reducing the height H and width W of region p by a certain ratio. The size of region p' can be determined as H / 2 and W / 2. Alternatively, the size of region p' can be determined as H / 4 and W / 4. Meanwhile, the height H and width W may be reduced by different ratios.
[0179] As another example, the width and height of the p' region can be determined as one of the preset values, such as 4, 8, or 16.
[0180] With respect to the p' region, the reference region q' for calculating the EIP filter coefficients can be determined according to one of the following methods.
[0181] The q' region can be determined to be the same as the q region.
[0182] The q' region can be determined by downsampling the q region.
[0183] The q' region can be determined as a part of the q region. For example, if the q region is determined by the sizes of aboveSize and leftSize defined in FIG. 8b, the q' region is determined as an area close to the p' region and can be determined by the sizes of aboveSize / 2 and leftSize / 2.
[0184] The p' region can be predicted at the (x,y) position using m filter coefficients as in Equation 7.
[0185]
[0186] p' (x,y) is the predicted value at the (x,y) position within the p' region, and c i is the EIP coefficient calculated in the q' region. t(x-offsetXi, y-offsetYi) represents the reconstructed sample or the predicted sample. The prediction process uses a recursive structure that uses the value output by EIP as input again.
[0187] m can be 15, as in the existing EIP. Alternatively, it can be determined differently depending on the downsampling rate.
[0188] Below, the restoration process is explained assuming the size of the p' region is W / 2 and H / 2.
[0189] p' (x,y) In this equation, x is an integer value between 1 and W / 2, and y is an integer value between 1 and H / 2. p' (x,y) (x,y) of is p (2x,2y) It is a position corresponding to . Or, (x,y) may be a position corresponding to p(2x-1,2y-1).
[0190] p' (x,y) (x,y) of p (2x,2y) In the case of a position corresponding to , the image decoder can recover the value of the remaining sample p(2x-1,2y-1) during the upsampling process. The value of p(2x-1,2y-1) is p' (x,y) , p' (x+1,y) , p' (x,y+1) , p' (x+1,y+1) It can be restored using some or all of it. As an example, the value of p(2x-1,2y-1) is p' (x,y) , p' (x+1,y) , p' (x,y+1) , p' (x+1,y+1) It can be calculated using the average or weighted average of. As another example, p'(x,y) Filter horizontally using all samples from , x=1...W / 2, and p' (x,y) The value of p(2x-1,2y-1) can be restored by filtering in the vertical direction using all samples of y=1...H / 2. DCT interpolation filter coefficients can be used as filter coefficients.
[0191] When upsampling samples existing at the boundary of region p', the image decoder can use samples from region q. p (2x-1,1) , x=1...W / 2 are samples located at the upper boundary of the block. When upsampling the aforementioned samples, the image decoder may use the upper reference region samples. When using the reference region samples and prediction samples together, the image decoder may weight the samples by assigning equal or different weights to the reference region samples and prediction samples. The upper reference samples q (2x-1,0) When represented as, p (2x-1,1) It can be expressed as in mathematical formula 9.
[0192]
[0193] In mathematical formula 9, w represents the weight.
[0194] In addition, in the process of using samples from the reference range, q (2x-2,0) , q (2x-1,0) . q (2x,0) Pixel changes between them may also be reflected.
[0195] p' (x,y) If (x,y) is the position corresponding to p(2x-1,2y-1), the image decoder [decodes] the remaining samples p (2x,2y) The value of is restored during the upsampling process. p (2x,2y) The value of is p' (x,y) , p' (x+1,y) , p' (x,y+1) , p' (x+1,y+1) It can be restored using some or all of it. As an example, p (2x,2y) The value of is p' (x,y), p' (x+1,y) , p' (x,y+1) , p' (x+1,y+1) It can be calculated using the average or weighted average of. As another example, p' (x,y) Filter horizontally using all samples from , x=1...W / 2, and p' (x,y) By filtering vertically using all samples of , y=1...H / 2, p (2x,2y) The value of can be restored. DCT interpolation filter coefficients can be used as filter coefficients.
[0196] The following explains REIP.
[0197] In REIP technology, the image decoder generates residual samples by intra-predicting or inter-predicting the current block, and generates predicted residual samples by predicting the residual samples once more using the EIP prediction technique.
[0198] In FIG. 11(a), the p region and the q region represent the current sample region (i.e., the current block) and the reference sample region (i.e., the reference region), respectively. The height of the p region is denoted as H, and the width as W. The image decoder generates a residual region r by applying intra-prediction or inter-prediction to the sample region p. The r region can be generated according to one of the following methods.
[0199] As an example, the r region is generated based on intra prediction. The intra prediction mode can be determined as one of the following.
[0200] The intra prediction mode of VVC is used. In this case, the intra prediction mode can be signaled from the video encoder to the video decoder. As another example, if eip_merge_flag is true, it can be derived based on the merge index.
[0201] An intra prediction mode can be derived in reference region q by using a method that derives gradients in a reference region, such as DIMD.
[0202] As another example, the r region can be generated using residual samples generated according to the inter-prediction of VVC.
[0203] As shown in Equation 10, the predicted residual sample r_pred predicts the residual region r (x,y) It can be calculated.
[0204]
[0205] In Equation 10, s(x-offsetXi, y-offsetYi) is the residual value recovered or predicted in the reference region or residual region. d i is the EIP filter coefficient derived from the reference region. i It can be derived according to one of the following methods.
[0206] As an example, d i is c used in mathematical formula 5. i It is identical to. That is, filter coefficients are derived in the same way as the existing EIP in the same region as the q region. For example, reconstructed samples within the q region can be used to derive filter coefficients. The filter type can be determined by the syntax element "eip_mode_idx" included in Table 1.
[0207] As another example, d i is c in mathematical formula 5. i Although a method for deriving is used, it can be derived after changing the reference region q to the residual region during the derivation process. That is, the filter coefficient d i It can be derived using residual samples associated with restored samples within the reference region q.
[0208] Meanwhile, residual samples of the reference region q can be generated according to the intra prediction mode.
[0209] As an example, the aforementioned intra prediction mode may be used. Any one of the previously defined intra prediction modes may be used. Alternatively, the intra prediction mode used to generate residual samples of the reference region q may be determined based on the intra prediction mode of the current block or the surrounding blocks of the current block (or the blocks containing the reference region).
[0210] As another example, a DIMD mode can be applied to a reference region to derive a gradient of the reference region, and an intra prediction can be performed based on the derived result to generate residual samples of the reference region.
[0211] As another example, residual samples of the reference region q can be generated according to the inter prediction.
[0212] Based on the directionality of the intra prediction mode used in the process of generating residual samples, d i The method of inducing it may be applied differently.
[0213] In each of the embodiments described above, the filter type of the EIP can be determined based on the directionality of the induced intra-prediction mode. For example, in the EIP, a horizontally elongated, vertically elongated, or square filter is used as shown in FIG. 8a. If the directionality of the intra-prediction mode is close to horizontal, a horizontally elongated filter is used. If the directionality of the intra-prediction mode is close to vertical, a vertically elongated filter is used. If the directionality of the intra-prediction mode is not close to horizontal or vertical, a square filter may be used.
[0214] As another example, if the current block is predicted using EIP mode, a filter of the same type as the EIP filter type used can be used to predict the residual samples. For instance, the filter type can be determined by the aforementioned syntax element "eip_mode_idx".
[0215] If a residual region is generated based on inter-prediction, a square filter can be used.
[0216] The video encoding device can generate second residual samples of the current block by subtracting predicted residual samples of the residual region from the residual samples of the current block (i.e., first residual samples). The video encoding device can generate quantized transform coefficients by transforming / quantizing the second residual samples, and generate a bitstream by encoding the quantized transform coefficients.
[0217] The image decoder can decode quantized transform coefficients from a bitstream and generate second residual samples by applying inverse quantization / inverse transform to the quantized transform coefficients. The image decoder can generate residual samples of the current block (i.e., first residual samples) by adding the second residual samples and the predicted residual samples of the residual region.
[0218] The application of EIP for lossless compression is described below. In lossless compression, the quantization / inverse quantization process is omitted. At this time, Equation 11 can be used, as in the conventional EIP process.
[0219]
[0220] Since it is lossless compression, the pixels of the restored image at the corresponding location can be used as t in Equation 11.
[0221] Hereinafter, a method for predicting a current block by applying an EIP to a downsampled current block using the illustration of FIG. 12 is described. The illustration of FIG. 12 can be performed by an image encoding device and an image decoder. For convenience, FIG. 12 is described based on the image encoding device, but if necessary, the operation by the image decoder is additionally described.
[0222] FIG. 12 is a flowchart illustrating a method for predicting a current block using an EIP according to one embodiment of the present disclosure.
[0223] The image encoding device determines the filter shape of the first reference region and the extrapolation filter (S1200).
[0224] An image encoding device can obtain a first reference region for extrapolation intra prediction and a filter shape of an extrapolation filter from a high level. The first reference region includes reconstructed reference samples of the current block, and the filter shape of the extrapolation filter includes an output pixel region and an input pixel region. The image encoding device can encode the information of the first reference region and the filter shape of the extrapolation filter. On the other hand, an image decoder can decode the information of the first reference region and the filter shape of the extrapolation filter from a bitstream.
[0225] The video encoding device determines the predicted region where the current block is downsampled (S1202).
[0226] For example, if the size of the current block is larger than a preset value, the video encoding device can determine the prediction region by downsampling the current block.
[0227] The video encoding device can determine the prediction area by reducing the height and width of the current block by a certain ratio.
[0228] The image encoding device determines a second reference region corresponding to the prediction region from the first reference region (S1204). Here, the second reference region is used to derive the filter coefficients of the extrapolation filter.
[0229] The video encoding device can determine a second reference area by downsampling a first reference area.
[0230] The image encoding device derives filter coefficients applied to the input pixel area based on the second reference area and the filter shape (S1206).
[0231] The video encoding device sequentially predicts samples within the prediction region by recursively applying filter coefficients to the surrounding restored reference samples and the first predicted samples within the prediction region, based on the prediction order of samples within the prediction region (S1208).
[0232] The video encoding device interpolates the remaining samples not included in the prediction area using the prediction samples within the prediction area (S1210).
[0233] The video encoding device can interpolate the remaining sample by averaging the predicted samples existing around the remaining sample.
[0234] The image encoding device can generate horizontal interpolated samples by interpolating and filtering the predicted samples in the horizontal direction. The image encoding device can interpolate the remaining samples by interpolating and filtering the predicted samples and the horizontal interpolated samples in the vertical direction.
[0235] If the remaining sample exists at the boundary of the current block, the image encoding device can interpolate the remaining sample existing at the boundary using the samples of the first reference region and the predicted samples.
[0236] The video encoding device can generate a prediction block of the current block by interpolating the remaining samples.
[0237] Subsequently, the video encoding device can generate a residual block by subtracting the prediction block from the current block. The video encoding device can generate transformation coefficients by applying transformation / quantization to the residual block and encode the generated transformation coefficients.
[0238] The image decoder can decode the quantized transform coefficients of the current block from the bitstream and generate a residual block by applying inverse quantization / inverse transform to the quantized transform coefficients. The image decoder can restore the current block by adding the prediction block and the residual block.
[0239] Although the flowcharts and timing diagrams in this specification describe each process as being executed sequentially, this is merely an illustrative explanation of the technical concept of one embodiment of the present disclosure. In other words, a person skilled in the art to which one embodiment of the present disclosure belongs may modify and adapt the flowcharts and timing diagrams in various ways, such as changing the order described in the flowcharts and timing diagrams or executing one or more of the processes in parallel, without departing from the essential characteristics of one embodiment of the present disclosure; therefore, the flowcharts and timing diagrams are not limited to a chronological order.
[0240] It should be understood that the exemplary embodiments described above may be implemented in many different ways. The functions or methods described in one or more examples may be implemented in hardware, software, firmware, or any combination thereof. It should be understood that the functional components described herein are labeled as "...unit" to particularly emphasize their implementation independence.
[0241] Meanwhile, the various functions or methods described in the present embodiment may be implemented as instructions stored in a non-transient recording medium that can be read and executed by one or more processors. A non-transient recording medium includes, for example, any type of recording device in which data is stored in a form readable by a computer system. For example, a non-transient recording medium includes storage media such as an EPROM (erasable programmable read-only memory), a flash drive, an optical drive, a magnetic hard drive, and a solid-state drive (SSD).
[0242] The above description is merely an illustrative explanation of the technical concept of the present embodiment, and a person skilled in the art to which the present embodiment belongs would be able to make various modifications and variations within the scope of the essential characteristics of the present embodiment. Accordingly, the present embodiments are intended to explain, not limit, the technical concept of the present embodiment, and the scope of the technical concept of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of the present embodiment.
[0243]
[0244] CROSS-REFERENCE TO RELATED APPLICATION
[0245] This patent application claims priority to Korean patent application No. 10-2024-0132645 filed on September 30, 2024, the entire contents of which are incorporated into this patent application by reference.
Claims
1. A method for restoring a current block performed by an image decoding device, A step of decoding a first reference region and a filter form of an extrapolation filter, wherein the first reference region includes restored reference samples of the current block, and the filter form of the extrapolation filter includes an output pixel region and an input pixel region; and A step of deriving filter coefficients applied to the input pixel area based on the filter shape above. A method including 2. In Paragraph 1, If the size of the current block mentioned above is larger than the preset value, A step of determining a prediction region downsampled from the above current block; and Step of determining a second reference area corresponding to the prediction area from the first reference area Includes more, The step of deriving the above filter coefficients is, A method for deriving filter coefficients applied to the input pixel area based on the second reference area and the filter shape.
3. In Paragraph 2, A step of generating predicted samples of the prediction region by recursively applying the filter coefficients to the surrounding restored reference samples of the prediction region and to the first predicted samples within the prediction region to sequentially predict the samples within the prediction region; and A step of interpolating the remaining samples not included in the prediction region using the above prediction samples. A method that further includes.
4. In Paragraph 2, The step of determining the above prediction area is, A method for determining the predicted area by reducing the height and width of the current block by a certain ratio.
5. In Paragraph 2, The step of determining the second reference area above is, A method for determining the second reference region by downsampling the first reference region.
6. In Paragraph 3, The step of interpolating the remaining samples mentioned above is, A method for interpolating the remaining sample by averaging the prediction samples existing around the remaining sample.
7. In Paragraph 1, The method further includes the step of generating predicted residual samples of the current block by recursively applying the filter coefficients to the surrounding residual samples of the current block and to the first predicted residual samples within the current block to sequentially predict the first residual samples within the current block. The step of deriving the above filter coefficients is, A method for deriving filter coefficients applied to the input pixel region based on the restoration samples of the first reference region or the residual samples of the first reference region using the filter shape above.
8. In Paragraph 7, A step of generating prediction samples by applying intra prediction or inter prediction to the current block above; A step of decoding second residual samples of the current block from the bitstream; A step of generating the first residual samples by adding the second residual samples and the predicted residual samples; and A step of generating a restoration block of the current block by adding the first residual samples and the prediction samples. A method that further includes.
9. In Paragraph 7, A method further comprising the step of generating residual samples of the first reference region based on intra prediction or inter prediction.
10. In Paragraph 1, The step of deriving the above filter coefficients is, A method for deriving filter coefficients applied to the input pixel region based on the filter shape and the restoration samples of the first reference region.
11. In Paragraph 10, A step of generating predicted samples of the current block by recursively applying the filter coefficients to surrounding restored reference samples of the current block and to samples first predicted within the current block to sequentially predict samples within the current block; A step of decoding residual samples of the current block without omitting the inverse quantization process; and A step of generating a restoration block of the current block by adding the above residual samples and the above prediction samples. A method that further includes.
12. A method for encoding a current block performed by an image encoding device, A step of obtaining a first reference region and a filter shape of an extrapolation filter, wherein the first reference region includes restored reference samples of the current block, and the filter shape of the extrapolation filter includes an output pixel region and an input pixel region; and A step of deriving filter coefficients applied to the input pixel area based on the filter shape above. A method including 13. In Paragraph 12, If the size of the current block mentioned above is larger than the preset value, A step of determining a prediction region downsampled from the above current block; and Step of determining a second reference area corresponding to the prediction area from the first reference area Includes more, The step of deriving the above filter coefficients is, A method for deriving filter coefficients applied to the input pixel area based on the second reference area and the filter shape.
14. In Paragraph 13, A step of generating predicted samples of the prediction region by recursively applying the filter coefficients to the surrounding restored reference samples of the prediction region and to the first predicted samples within the prediction region to sequentially predict the samples within the prediction region; A step of interpolating the remaining samples not included in the prediction region using the above prediction samples; and A step of encoding the information of the first reference region and the filter shape of the extrapolation filter. A method that further includes.
15. In Paragraph 12, A step of generating predicted residual samples of the current block by recursively applying the filter coefficients to the surrounding restored residual samples of the current block and to the first predicted residual samples within the current block to sequentially predict the first residual samples within the current block. Includes more, The step of deriving the above filter coefficients is, A method for deriving filter coefficients applied to the input pixel region based on the restoration samples of the first reference region or the residual samples of the first reference region using the filter shape above.
16. In Paragraph 15, A step of generating prediction samples by applying intra prediction or inter prediction to the current block above; A step of generating first residual samples by subtracting the predicted samples from the samples of the current block. A step of generating second residual samples by subtracting the predicted residual samples from the first residual samples; and Step of encoding the above second residual samples A method that further includes.
17. In Paragraph 12, The step of deriving the above filter coefficients is, A method for deriving filter coefficients applied to the input pixel region based on the filter shape and the restoration samples of the first reference region.
18. In Paragraph 17, A step of generating predicted samples of the current block by recursively applying the filter coefficients to surrounding restored reference samples of the current block and to samples first predicted within the current block to sequentially predict samples within the current block; A step of generating residual samples of the current block by subtracting the predicted samples from the samples of the current block; and A step of encoding the above residual samples without omitting the quantization process A method that further includes.
19. A method for providing video data to a video decoder, A step of encoding the above video data into a bitstream; and Step of transmitting the above bitstream to the above video decoder Includes, The step of encoding the above video data is, A step of obtaining a first reference region and a filter shape of an extrapolation filter, wherein the first reference region includes restored reference samples of the current block, and the filter shape of the extrapolation filter includes an output pixel region and an input pixel region; and A step of deriving filter coefficients applied to the input pixel area based on the filter shape above. A method including