Video coding method and device, and recording medium storing bitstream
The intra prediction method and device improve video compression efficiency by using an intraMerge list to generate predictive blocks, addressing the challenges of compressing high-resolution images and enhancing prediction accuracy and compression performance.
Patent Information
- Application Number
- PCT/KR2024/016899
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-09
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-08
AI Technical Summary
Current video compression technologies face challenges in efficiently compressing high-resolution images, particularly in intra prediction methods, which affect the overall compression performance and efficiency.
The proposed solution involves an intra prediction method and device that utilize an intraMerge list to generate predictive blocks for current blocks, incorporating candidates based on surrounding blocks, historical candidates, and various prediction modes, such as DIMD, TIMD, and MPM lists, to improve prediction accuracy and efficiency.
This approach enhances the compression performance of video encoding and decoding by improving the prediction accuracy and reducing the amount of residual signals, thereby increasing the overall efficiency of the video coding process.
Smart Images

Figure KR2024016899_08052025_PF_FP_ABST
Abstract
Description
Video coding method and device, and recording medium storing bitstream
[0001] The present invention relates to a video signal processing method and device.
[0002] The market demand for high-resolution video is growing, necessitating technologies capable of efficiently compressing high-resolution images. To address this market need, the ISO / IEC's Moving Picture Expert Group (MPEG) and the ITU-T's Video Coding Expert Group (VCEG) jointly formed the Joint Collaborative Team on Video Coding (JCT-VC). They completed development of the HEVC (High Efficiency Video Coding) video compression standard in January 2013 and have been actively conducting research and development on next-generation compression standards.
[0003] Video compression largely consists of intraprediction, interprediction, transform, quantization, entropy coding, and in-loop filtering. Among these, intraprediction refers to a technique that generates a prediction block for the current block using reconstructed pixels surrounding the current block. The encoder encodes the intraprediction mode used for intraprediction, and the decoder performs intraprediction by reconstructing the encoded intraprediction mode.
[0004] The present disclosure seeks to provide an intra prediction method and device.
[0005] The present disclosure provides a prediction method and device based on an intra merge list.
[0006] The present disclosure provides a block vector-based intra prediction method and device.
[0007] The video decoding method and device according to the present disclosure can construct an intra merge list including a plurality of intra merge candidates for a current block, derive an intra prediction mode of the current block based on the intra merge list, and generate a prediction block of the current block based on the intra prediction mode.
[0008] In the video decoding method and device according to the present disclosure, at least one of the plurality of intra merge candidates may be derived based on a neighboring block of the current block. Here, the neighboring block may include at least one of an adjacent neighboring block or a non-adjacent neighboring block.
[0009] In the video decoding method and device according to the present disclosure, the plurality of intra merge candidates may further include one or more history-based candidates.
[0010] In the video decoding method and device according to the present disclosure, when the surrounding block is encoded based on a DIMD (decoder side intra mode derivation) mode, the DIMD mode may be added as an intra merge candidate of the current block, or an intra prediction mode derived for the surrounding block based on the DIMD mode may be stored in an intra merge candidate of the current block.
[0011] In the image decoding method and device according to the present disclosure, when the surrounding block is encoded based on a TIMD (template-based intra mode derivation) mode, the TIMD mode may be added as an intra merge candidate of the current block, or an intra prediction mode derived for the surrounding block based on the TIMD mode may be stored in an intra merge candidate of the current block.
[0012] In the video decoding method and device according to the present disclosure, when the intra prediction mode of the surrounding block is derived based on the MPM list, at least one of the MPM candidates of the MPM list for the surrounding block can be added as an intra merge candidate of the current block.
[0013] In the image decoding method and device according to the present disclosure, index information used by the surrounding block for intra prediction can be stored in the intra merge candidate of the current block.
[0014] In the image decoding method and device according to the present disclosure, when the prediction block of the surrounding block is generated through a weighted sum of a plurality of prediction blocks, the intra prediction modes used to generate the plurality of prediction blocks and / or the weights for the weighted sum may be stored in the intra merge candidate of the current block.
[0015] In the image decoding method and device according to the present disclosure, when the surrounding block is encoded in a geometric segmentation mode, the geometric segmentation mode may be added as an intra merge candidate of the current block, or information regarding geometric segmentation of the surrounding block may be stored in the intra merge candidate of the current block.
[0016] In the image decoding method and device according to the present disclosure, when the surrounding block is encoded in a planar mode, the planar mode may be added as an intra merge candidate of the current block, or information identifying the type of the planar mode may be stored in the intra merge candidate of the current block.
[0017] In the video decoding method and device according to the present disclosure, one or more intra prediction modes can be derived from the intra merge list.
[0018] In the video decoding method and device according to the present disclosure, the one or more intra prediction modes may be selected from the top N intra merge candidates among a plurality of intra merge candidates belonging to the intra merge list.
[0019] In the video decoding method and device according to the present disclosure, the intra merge list can be divided into a plurality of sets, and one or more intra prediction modes can be derived from at least one of the plurality of sets.
[0020] In the video decoding method and device according to the present disclosure, costs can be calculated for intra prediction modes of the plurality of intra merge candidates, and the intra prediction mode of the current block can be derived based on the top L intra prediction modes with the minimum costs.
[0021] The video encoding method and device according to the present disclosure can configure an intra merge list including a plurality of intra merge candidates for a current block, derive an intra prediction mode of the current block based on the intra merge list, and generate a prediction block of the current block based on the intra prediction mode.
[0022] A computer-readable recording medium according to the present disclosure can store a bitstream encoded by the image encoding method.
[0023] According to the present disclosure, the compression performance of a decoder / encoder can be improved through prediction based on an intra merge list.
[0024] According to the present disclosure, the performance of a decoder / encoder can be improved through efficient signaling of intra merge candidates.
[0025] According to the present disclosure, more precise intra prediction can be performed through a combination of intra prediction and affine prediction, and the compression efficiency of the encoder can be improved by reducing the bit amount of the banquet signal.
[0026] FIG. 1 is a block diagram showing an image encoding device according to the present disclosure.
[0027] FIG. 2 is a block diagram showing an image decoding device according to the present disclosure.
[0028] FIG. 3 illustrates an embodiment according to the present disclosure, which illustrates a method for generating a prediction block of a current block based on an intra merge list.
[0029] FIG. 4 illustrates an embodiment according to the present disclosure, which illustrates a method of utilizing block vector-based affine prediction in intra prediction.
[0030] The video decoding method and device according to the present disclosure can induce an intra prediction mode for the current block by applying a filter to a template of the current block, induce a weight for the intra prediction mode, and generate a prediction block generated based on the intra prediction mode and a final prediction block of the current block based on the weight.
[0031] In the image decoding method and device according to the present disclosure, the template is a peripheral area adjacent to the current block, and the peripheral area may include at least one of a left area, an upper area, or an upper left area.
[0032] In the image decoding method and device according to the present disclosure, the range of the template to which the filter is applied can be variably determined based on the height and width of the current block.
[0033] In the image decoding method and device according to the present disclosure, the center sample among the reference samples input to the filter may belong to at least one of the first reference sample line that is 1 sample away from the boundary of the current block or the second reference sample line that is 2 samples away from the boundary of the current block.
[0034] In the video decoding method and device according to the present disclosure, if there is an unavailable sample among the reference samples input to the filter, the unavailable sample is replaced with an available sample, and the available sample can be generated based on a predetermined interpolation method or a predetermined intra prediction mode.
[0035] In the video decoding method and device according to the present disclosure, the step of deriving the intra prediction mode may include the step of generating a HoG table for a template of the current block, and the HoG table may further include a frequency for each intra prediction mode.
[0036] In the image decoding method and device according to the present disclosure, the step of deriving the intra prediction mode may include a step of generating a HoG of the current block based on a DIMD HoG of a surrounding block.
[0037] In the video decoding method and device according to the present disclosure, the step of deriving the intra prediction mode may include a step of generating an HoG for a reference block specified by a predetermined block vector.
[0038] In the image decoding method and device according to the present disclosure, the final prediction block can be generated based on a weighted sum of a prediction block generated based on the intra prediction mode and a prediction block generated based on the planar mode.
[0039] In the image decoding method and device according to the present disclosure, the Planar mode can be adaptively induced into any one of a general Planar mode, a vertical Planar mode, or a horizontal Planar mode.
[0040] The video encoding method and device according to the present disclosure can derive an intra prediction mode for the current block by applying a filter to a template of the current block, derive a weight for the intra prediction mode, and generate a prediction block generated based on the intra prediction mode and a final prediction block of the current block based on the weight.
[0041] A computer-readable recording medium according to the present invention can store a bitstream encoded by the image encoding method.
[0042] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings attached to this specification so that those skilled in the art can easily implement the present invention. However, the present invention may be implemented in various different forms and is not limited to the embodiments described herein. In addition, in the drawings, parts irrelevant to the description have been omitted to clearly explain the present invention, and similar parts have been designated with similar reference numerals throughout the specification.
[0043] Throughout this specification, when a part is said to be 'connected' to another part, this includes not only cases where they are directly connected, but also cases where they are electrically connected with another element in between.
[0044] Additionally, whenever a part throughout this specification is said to "include" a component, this does not mean that other components are excluded, but rather that other components may be included, unless specifically stated otherwise.
[0045] Additionally, while terms such as "first," "second," etc. may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another.
[0046] Additionally, in the embodiments of the devices and methods described herein, some components of the devices or some steps of the methods may be omitted. Furthermore, the order of some components of the devices or some steps of the methods may be changed. Furthermore, other components or other steps may be inserted into some components of the devices or some steps of the methods.
[0047] Additionally, some components or some steps of the first embodiment of the present invention may be added to the second embodiment of the present invention, or some components or some steps of the second embodiment may be replaced.
[0048] In addition, the components shown in the embodiments of the present invention are independently depicted to represent different characteristic functions, and this does not mean that each component is composed of separate hardware or a single software component. That is, each component is described by listing each component for convenience of explanation, and at least two components among each component may be combined to form a single component, or a single component may be divided into multiple components to perform a function. Such integrated and separate embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0049] In this specification, a block can be variously expressed as a unit, an area, a unit, a partition, etc., and a sample can be variously expressed as a pixel, a pel, a pixel, etc.
[0050] Hereinafter, embodiments of the present invention will be described in more detail with reference to the attached drawings. In describing the present invention, duplicate descriptions of identical components will be omitted.
[0051] FIG. 1 is a block diagram showing an image encoding device according to the present disclosure.
[0052] Referring to FIG. 1, a video encoding device (100) may include a picture segmentation unit (110), a prediction unit (120, 125), a transformation unit (130), a quantization unit (135), a reordering unit (160), an entropy encoding unit (165), an inverse quantization unit (140), an inverse transformation unit (145), a filter unit (150), and a memory (155).
[0053] The picture segmentation unit (110) can segment the input picture into at least one processing unit. At this time, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). Hereinafter, in the embodiments of the present disclosure, the coding unit may be used to mean a unit that performs encoding or a unit that performs decoding.
[0054] A prediction unit may be divided into at least one square or rectangular shape of the same size within a single coding unit, or may be divided such that one prediction unit among the divided prediction units within a single coding unit has a different shape and / or size from another prediction unit. When a prediction unit that performs intra prediction based on a coding unit is generated and is not the minimum coding unit, intra prediction can be performed without being divided into a plurality of NxN prediction units.
[0055] The prediction unit (120, 125) may include an inter prediction unit (120) that performs inter prediction or inter prediction, and an intra prediction unit (125) that performs intra prediction or intra prediction. It may determine whether to use inter prediction or intra prediction for a prediction unit, and determine specific information (e.g., intra prediction mode, motion vector, reference picture, etc.) according to each prediction method. A residual value (residual block) between the generated prediction block and the original block may be input to the transformation unit (130). In addition, prediction mode information, motion vector information, etc. used for prediction may be encoded together with the residual value by the entropy encoding unit (165) and transmitted to the decoder.
[0056] The inter prediction unit (120) may predict a prediction unit based on information of at least one picture among the previous or subsequent pictures of the current picture, and in some cases, may predict a prediction unit based on information of a portion of an encoded region within the current picture. The inter prediction unit (120) may include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.
[0057] The reference picture interpolation unit can receive reference picture information from the memory (155) and generate pixel information less than an integer pixel from the reference picture. In the case of luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 4 pixels. In the case of a chrominance signal, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 8 pixels.
[0058] The motion prediction unit can perform motion prediction based on a reference picture interpolated by the reference picture interpolation unit. Various methods such as FBMA (Full search-based Block Matching Algorithm), TSS (Three Step Search), and NTS (New Three-Step Search Algorithm) can be used to derive a motion vector. The motion vector can have a motion vector value in units of 1 / 2 or 1 / 4 pixels based on the interpolated pixel. The motion prediction unit can predict the current prediction unit by using different motion prediction methods. Various methods such as Skip Mode, Merge Mode, AMVP Mode, Intra Block Copy Mode, and Affine Mode can be used as motion prediction methods.
[0059] The intra prediction unit (125) can generate a prediction unit based on reference pixel information surrounding the current block, which is pixel information within the current picture. If the surrounding block of the current prediction unit is a block on which inter prediction has been performed and the reference pixel is a pixel on which inter prediction has been performed, the reference pixel included in the block on which inter prediction has been performed can be replaced and used with reference pixel information of the surrounding block on which intra prediction has been performed. That is, if the reference pixel is not available, the unavailable reference pixel information can be replaced and used with at least one reference pixel among the available reference pixels.
[0060] Additionally, a residual block containing residual value information, which is the difference between the prediction unit that performed the prediction based on the prediction unit generated in the prediction unit (120, 125) and the original block of the prediction unit, can be generated. The generated residual block can be input to the transformation unit (130).
[0061] In the transformation unit (130), the residual block including the residual value information of the prediction unit generated through the original block and the prediction unit (120, 125) can be transformed using a transformation method such as DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), or KLT. Whether to apply DCT, DST, or KLT to transform the residual block can be determined based on the intra prediction mode information of the prediction unit used to generate the residual block.
[0062] The quantization unit (135) can quantize values converted to the frequency domain by the transformation unit (130). The quantization coefficients can vary depending on the block or the importance of the image. The values produced by the quantization unit (135) can be provided to the dequantization unit (140) and the reordering unit (160).
[0063] The rearrangement unit (160) can perform rearrangement of coefficient values for quantized residual values.
[0064] The rearrangement unit (160) can change a two-dimensional block-shaped coefficient into a one-dimensional vector form through a coefficient scanning method. For example, the rearrangement unit (160) can change the two-dimensional block-shaped coefficient into a one-dimensional vector form by scanning from the DC coefficient to the coefficient of the high-frequency region using a zig-zag scan method. Depending on the size of the transformation unit and the intra prediction mode, a vertical scan that scans the two-dimensional block-shaped coefficient in the column direction or a horizontal scan that scans the two-dimensional block-shaped coefficient in the row direction may be used instead of the zig-zag scan. That is, depending on the size of the transformation unit and the intra prediction mode, it is possible to determine which scan method among the zig-zag scan, the vertical scan, and the horizontal scan is to be used.
[0065] The entropy encoding unit (165) can perform entropy encoding based on the values produced by the rearrangement unit (160). Entropy encoding can use various encoding methods such as, for example, Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC). In this regard, the entropy encoding unit (165) can encode residual value coefficient information of the encoding unit from the rearrangement unit (160) and the prediction units (120, 125). In addition, according to the present disclosure, it is possible to signal and transmit information indicating that motion information is derived and used on the decoder side and information on a technique used to derive motion information.
[0066] The inverse quantization unit (140) and the inverse transformation unit (145) inversely quantize the values quantized in the quantization unit (135) and inversely transform the values transformed in the transformation unit (130). The residual values generated in the inverse quantization unit (140) and the inverse transformation unit (145) can be combined with the predicted prediction units predicted through the motion estimation unit, motion compensation unit, and intra prediction unit included in the prediction unit (120, 125) to generate a reconstructed block.
[0067] The filter unit (150) may include at least one of a deblocking filter, an offset correction unit, and an ALF (Adaptive Loop Filter). The deblocking filter may remove block distortion caused by boundaries between blocks in a restored picture. The offset correction unit may correct the offset from the original image on a pixel-by-pixel basis for the image on which deblocking has been performed. In order to perform offset correction for a specific picture, a method may be used in which the pixels included in the image are divided into a certain number of regions, the regions to be offset are determined, and the offset is applied to the regions, or the offset is applied by considering edge information of each pixel. The ALF (Adaptive Loop Filtering) may be performed based on a value obtained by comparing the filtered restored image with the original image. After dividing the pixels included in the image into a predetermined group, one filter to be applied to the group is determined, and filtering may be performed differentially for each group.
[0068] The memory (155) can store a restored block or picture produced through the filter unit (150), and the stored restored block or picture can be provided to the prediction unit (120, 125) when performing inter prediction.
[0069] FIG. 2 is a block diagram showing an image decoding device according to the present disclosure.
[0070] Referring to FIG. 2, the image decoding device (200) may include an entropy decoding unit (210), a rearrangement unit (215), an inverse quantization unit (220), an inverse transformation unit (225), a prediction unit (230, 235), a filter unit (240), and a memory (245).
[0071] When a video bitstream is input to a video encoding device, the input bitstream can be decoded in the opposite procedure to that of the video encoding device.
[0072] The entropy decoding unit (210) can perform entropy decoding in a procedure opposite to that of the entropy encoding unit of the video encoder. For example, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied in response to the method performed in the video encoder.
[0073] The entropy decoding unit (210) can decode information related to intra prediction and inter prediction performed in the encoder.
[0074] The reordering unit (215) can perform reordering based on the method by which the bitstream entropy-decoded by the entropy decoding unit (210) is reordered by the encoding unit. The coefficients expressed in the form of a one-dimensional vector can be reordered by restoring them back to coefficients in the form of a two-dimensional block.
[0075] The inverse quantization unit (220) can perform inverse quantization based on the quantization parameters provided by the encoder and the coefficient values of the rearranged block.
[0076] The inverse transform unit (225) can perform inverse transform, i.e., inverse DCT, inverse DST, and inverse KLT, on the transforms performed by the transform unit, i.e., DCT, DST, and KLT, on the quantization result performed by the image encoder. The inverse transform can be performed based on the transmission unit determined by the image encoder. In the inverse transform unit (225) of the image decoder, a transform technique (e.g., DCT, DST, KLT) can be selectively performed according to a plurality of pieces of information, such as a prediction method, the size of the current block, and the prediction direction.
[0077] The prediction unit (230, 235) can generate a prediction block based on the prediction block generation related information provided by the entropy decoding unit (210) and the previously decoded block or picture information provided by the memory (245).
[0078] As described above, when performing intra prediction or intra prediction in the same manner as the operation in the image encoder, if the size of the prediction unit and the size of the transformation unit are the same, intra prediction for the prediction unit is performed based on the pixels on the left side of the prediction unit, the pixels on the upper left side, and the pixels on the upper side. However, when performing intra prediction, if the size of the prediction unit and the size of the transformation unit are different, intra prediction can be performed using reference pixels based on the transformation unit. In addition, intra prediction using NxN division only for the minimum coding unit can be used.
[0079] The prediction unit (230, 235) may include a prediction unit determination unit, an inter prediction unit, and an intra prediction unit. The prediction unit determination unit may receive various information such as prediction unit information input from the entropy decoding unit (210), prediction mode information of an intra prediction method, and motion prediction-related information of an inter prediction method, and may distinguish a prediction unit from a current encoding unit and determine whether the prediction unit performs inter prediction or intra prediction. On the other hand, if the encoder (100) does not transmit motion prediction-related information for the inter prediction, but instead transmits information indicating that motion information is to be derived and used on the decoder side and information on a technique used to derive motion information, the prediction unit determination unit determines whether the inter prediction unit (230) performs prediction based on the information transmitted from the encoder (100).
[0080] The inter prediction unit (230) can perform inter prediction on the current prediction unit based on information included in at least one picture among the previous picture or the subsequent picture of the current picture including the current prediction unit, using information required for inter prediction of the current prediction unit provided by the image encoder. In order to perform inter prediction, it can be determined based on the encoding unit whether the motion prediction method of the prediction unit included in the corresponding encoding unit is one of Skip Mode, Merge Mode, AMVP Mode, Intra Block Copy Mode, and Affine Mode.
[0081] The intra prediction unit (235) can generate a prediction block based on pixel information within the current picture. If the prediction unit is a prediction unit that has performed intra prediction, intra prediction can be performed based on intra prediction mode information of the prediction unit provided by the image encoder.
[0082] The intra prediction unit (235) may include an Adaptive Intra Smoothing (AIS) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a unit that performs filtering on the reference pixels of the current block and can determine whether to apply the filter based on the prediction mode of the current prediction unit and apply it. AIS filtering can be performed on the reference pixels of the current block using the prediction mode and AIS filter information of the prediction unit provided by the image encoder. If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.
[0083] The reference pixel interpolation unit can interpolate the reference pixel to generate a reference pixel of a pixel unit less than an integer value when the prediction mode of the prediction unit is a prediction unit that performs intra prediction based on the pixel value interpolated from the reference pixel. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter can generate a prediction block through filtering when the prediction mode of the current block is the DC mode.
[0084] The restored block or picture may be provided to a filter unit (240). The filter unit (240) may include a deblocking filter, an offset correction unit, and an ALF.
[0085] Information about whether a deblocking filter has been applied to a corresponding block or picture can be received from a video encoding device, and if a deblocking filter has been applied, information about whether a strong or weak filter has been applied. The deblocking filter of the video decoder can receive information related to the deblocking filter provided by the video encoder, and the video decoder can perform deblocking filtering on the corresponding block.
[0086] The offset correction unit can perform offset correction on the restored image based on the type of offset correction applied to the image during encoding and information on the offset value. ALF can be applied to the encoding unit based on information on whether ALF is applied and ALF coefficient information provided from the encoder. This ALF information can be provided by being included in a specific parameter set.
[0087] The memory (245) can store a restored picture or block so that it can be used as a reference picture or reference block, and can also provide the restored picture to an output unit.
[0088] The intra merge mode according to the present disclosure may refer to a method for deriving an intra prediction mode of a current block based on an intra merge list. Hereinafter, with reference to FIG. 3, a method for generating a prediction block of a current block based on an intra merge list will be described in detail.
[0089] Referring to FIG. 3, an intra-merge list for the current block can be constructed (S300). The intra-merge list according to the present disclosure can include multiple intra-merge candidates.
[0090] At least one of the plurality of intra merge candidates may be derived based on a neighboring block of the current block. That is, the at least one intra merge candidate may have intra prediction information of a neighboring block. Here, the neighboring block may include at least one of an adjacent neighboring block or a non-adjacent neighboring block. The intra prediction information may be information signaled or derived for intra prediction. For example, the intra prediction information may be understood as an intra prediction mode, information about an MPM index, information about an intra merge index, information about a reference line for intra prediction, information about an intra subpartition (ISP), weighted prediction information, information about a geometric partition (GPM), information about a motion vector (or block vector), etc.
[0091] The above adjacent neighboring blocks may include at least one of a neighboring block spatially adjacent to the current block (hereinafter referred to as a spatial neighboring block) or a neighboring block temporally adjacent to the current block (hereinafter referred to as a temporal neighboring block).
[0092] The spatial surrounding block may include at least one of a left surrounding block, a lower left surrounding block, an upper surrounding block, an upper right surrounding block, or an upper left surrounding block. For example, it is assumed that the coordinates of the upper left sample of the current block are (0, 0), and the width and height of the current block are W and H. In this case, the left surrounding block may include at least one of a block containing a sample of (-1, 0) or a block containing a sample of (-1, H-1). The lower left surrounding block may be a block containing a sample of (-1, H). The upper surrounding block may include at least one of a block containing a sample of (0, -1) or a block containing a sample of (W-1, -1). The upper right surrounding block may be a block containing a sample of (W, -1). The upper left surrounding block may be a block containing a sample of (-1, -1).
[0093] The temporal neighboring block may belong to a picture (hereinafter referred to as a temporal picture) having a different temporal order from the current picture to which the current block belongs. The temporal order may refer to an output order (POC) or a encoding / decoding order. The temporal neighboring block may be a block at the same position as the current block within the temporal picture. Alternatively, the temporal neighboring block may be a block at a position shifted by a predetermined offset vector from the same position as the current block within the temporal picture. Alternatively, the temporal neighboring block may be a block including the central position of a block at a position shifted by the predetermined offset vector.
[0094] The non-adjacent neighboring block may belong to the same picture as the current block. The non-adjacent neighboring block may be a block located N samples away from the boundary of the current block. Here, N may be an integer greater than or equal to 1.
[0095] The above multiple intra merge candidates may further include one or more history-based candidates. Here, the history-based candidates may be blocks that were sub-encoded / decoded prior to the current block, and may have intra prediction information of the sub-encoded / decoded blocks based on intra prediction. To this end, the intra prediction information of the blocks sub-encoded / decoded based on intra prediction may be stored in a history-based table (or buffer). The intra prediction information of the blocks stored in the history-based table may be added to the intra merge list as intra merge candidates.
[0096] However, in the process of constructing the intra merge list, information that overlaps with intra prediction information already added to the intra merge list through pruning may not be added.
[0097] For example, if a neighboring block is encoded based on a decoder-side intra mode derivation (DIMD) mode, the DIMD mode may be added as an intra merge candidate. An intra prediction mode derived for the neighboring block based on the DIMD mode may be stored in the intra merge candidate. The DIMD mode may be a method of calculating a gradient based on at least two samples belonging to a neighboring area of a current block, and deriving an intra prediction mode based on at least one of the calculated gradient or the magnitude of the gradient.
[0098] If a neighboring block is encoded based on a TIMD (template-based intra mode derivation) mode, the TIMD mode may be added as an intra merge candidate. An intra prediction mode derived for the neighboring block based on the TIMD mode may be stored in the intra merge candidate. The TIMD mode may be a method of calculating, for each of predetermined intra prediction modes, a difference between a prediction sample and a reconstructed sample of a template region adjacent to a current block, and deriving an intra prediction mode based on the calculated differences.
[0099] If the intra prediction mode of a surrounding block is derived based on the MPM list, at least one of the MPM candidates in the MPM list for the surrounding block can be added as an intra merge candidate.
[0100] Index information used by surrounding blocks for intra prediction may be stored in the intra merge candidate. Here, the index information may mean an index indicating one of directional modes and / or non-directional modes, an MPM index, an intra merge index, a reference line index, or a geometric segmentation index.
[0101] When a prediction block of a surrounding block is generated through a weighted sum of multiple prediction blocks, the intra prediction modes used to generate the multiple prediction blocks and / or the weights for the weighted sum may be stored in the intra merge candidate.
[0102] If the surrounding blocks have one or more motion vectors (or block vectors), each motion vector or a weighted sum of them can be stored in the intra merge candidate.
[0103] If the surrounding blocks are encoded in a geometric partitioning mode, the geometric partitioning mode may be added as an intra merge candidate for the current block. Information regarding the geometric partitioning of the surrounding blocks may be stored in the intra merge candidate. Here, the information regarding the geometric partitioning may include at least one of information specifying the position / angle of the partitioning line for the geometric partitioning or the intra prediction mode of each partition.
[0104] Information indicating the prediction mode of the surrounding blocks may be stored in the intra-merge candidate. Here, the information indicating the prediction mode may be an MPM flag, a planar mode flag, an MPM index, an intra-merge flag, an intra-merge index, a DIMD flag, a TIMD flag, a MIP flag, or a GPM flag.
[0105] If the surrounding blocks are encoded in planar mode, the planar mode may be added as an intra-merge candidate. The planar mode may be categorized into types of general planar mode, horizontal planar mode, and vertical planar mode. In this case, information identifying the type of the planar mode may be stored in the intra-merge candidate. For example, at least one of a first flag indicating whether the general planar mode is used or a second flag indicating either the horizontal or vertical planar modes may be stored in the intra-merge candidate. Alternatively, an index indicating either the general, horizontal, or vertical planar modes may be stored in the intra-merge candidate.
[0106] The encoder and decoder can generate an intra merge list using the same method described above.
[0107] The intra-merge mode according to the present disclosure can be adaptively applied based on a predetermined flag. The flag may be a flag indicating whether the intra-merge mode is applied (hereinafter referred to as an intra-merge flag).
[0108] The encoder can determine whether the current block is subject to intra-merge mode and signal this to the decoder by encoding an intra-merge flag to indicate this. The decoder can then determine whether the current block is subject to intra-merge mode based on the signaled intra-merge flag.
[0109] The intra-merge flag may be signaled before the information indicating the aforementioned prediction mode. Alternatively, the intra-merge flag may be signaled after the information indicating the aforementioned prediction mode. Alternatively, it may be signaled between the information indicating the aforementioned prediction mode.
[0110] If the intra-merge flag indicates that the intra-merge mode applies to the current block, an intra-merge list may be constructed for the current block. If the intra-merge flag indicates that the intra-merge mode does not apply to the current block, an intra-merge list may not be constructed for the current block.
[0111] Referring to FIG. 3, the intra prediction mode of the current block can be derived based on the intra merge list (S310).
[0112] One intra prediction mode can be derived from an intra merge list. Alternatively, two or more intra prediction modes can be derived from an intra merge list.
[0113] The intra prediction mode of the current block can be derived explicitly or implicitly.
[0114] For example, the encoder can encode an intra merge index that specifies one of multiple intra merge candidates in an intra merge list and signal it to the decoder. The decoder can then derive the intra prediction mode of the current block based on the intra merge candidate specified by the intra merge index.
[0115] Alternatively, the encoder may encode intra merge indices that specify at least two intra merge candidates among a plurality of intra merge candidates belonging to the intra merge list and signal the encoded intra merge indices to the decoder. The decoder may derive intra prediction modes of the current block based on the intra merge candidates specified by the at least two intra merge indices.
[0116] Alternatively, N intra merge candidate(s) may be selected from among multiple intra merge candidates belonging to the intra merge list, and the intra prediction mode of the current block may be derived based on the selected intra merge candidate(s). Here, N may be an integer of 1, 2, or a higher number. The N intra merge candidate(s) may be the top N intra merge candidate(s) in ascending order of the index of the intra merge list.
[0117] Alternatively, the intra merge list may be divided into M sets, and one or more intra prediction modes may be derived from each set. Alternatively, the intra merge list may be divided into M sets, and one or more intra prediction modes may be derived from any one of the M sets. Here, M may be an integer greater than or equal to 2. Information for specifying the value of M may be separately signaled.
[0118] Alternatively, the template region of the current block and / or surrounding blocks may be predicted based on the intra prediction modes of all or some of the multiple intra merge candidates belonging to the intra merge list. Through this prediction, a predetermined cost may be calculated for each intra prediction mode of the intra merge list. Here, the cost may be defined as the difference (e.g., SAD, MSE) between the predicted sample of the template region and the reconstructed sample. The intra prediction mode(s) of the current block may be derived based on the top L intra prediction mode(s) with the minimum cost. Here, L may be an integer greater than or equal to 1.
[0119] Referring to FIG. 3, a prediction block of the current block can be generated based on the intra prediction mode (S320).
[0120] If one intra prediction mode is derived for the current block, intra prediction can be performed based on the intra prediction mode to generate a prediction block of the current block.
[0121] Alternatively, if at least two intra prediction modes are derived for the current block, prediction blocks can be generated based on the at least two intra prediction modes, and a prediction block of the current block can be generated through a weighted sum of the generated prediction blocks.
[0122] Additionally, a linear or nonlinear filter having a 1D or 2D filter shape may be applied to the prediction block of the current block to correct the prediction block.
[0123] Below, with reference to Fig. 4, we will examine a method of utilizing block vector-based affine prediction in intra prediction.
[0124] Referring to FIG. 4, a candidate list for intra-affine prediction (hereinafter referred to as an intra-affine candidate list) can be generated (S400).
[0125] The intra-affine candidate list according to the present disclosure may include multiple intra-affine candidates. Each intra-affine candidate may include one or more control point block vector predictors (CPBVPs). Each CPBVP may be used as a control point vector at the top-left, bottom-left, top-right, or bottom-right positions of the current block. Here, each CPBVP may have an integer pixel or sub-pixel precision. Scaling of the block vector precision may be performed during the process of deriving the CPBVP.
[0126] For example, if a neighboring block adjacent to the current block has a control point block vector (CPBV), an intra-affine candidate of the current block can be derived based on the CPBV of the neighboring block.
[0127] For example, if a block located a predetermined distance away from the current block has a block vector, an intra-affine candidate of the current block can be derived based on the block vector of the block. Here, the predetermined distance can be expressed as a distance moved by N in the x-axis direction and / or M in the y-axis direction. The absolute values of N and M can be less than or equal to the width and height of the current picture.
[0128] For example, if multiple blocks located a predetermined distance from the current block have block vectors, an intra-affine candidate for the current block can be derived based on a combination of the block vectors. Here, the predetermined distance is as described above.
[0129] For example, a block vector list according to an intra block copy (IBC) mode may be constructed. The block vector list according to the IBC mode may be generated based on block vectors of neighboring blocks of the current block. The neighboring blocks include at least one of adjacent neighboring blocks or non-adjacent neighboring blocks, as described above. An intra-affine candidate of the current block may be derived based on a combination of one or more block vectors included in the block vector list.
[0130] For example, a block vector list may be constructed according to an intra-template matching prediction (intra TMP) mode. The block vector list according to the intra TMP mode may be generated by performing template matching based on the template of the current block. An intra-affine candidate for the current block may be derived based on a combination of one or more block vectors included in the block vector list.
[0131] For example, a history-based block vector list can be constructed. The history-based block vector list can be generated based on the block vectors of blocks previously encoded / decoded prior to the current block. Intra-affinity candidates for the current block can be derived based on a combination of one or more block vectors in the history-based block vector list.
[0132] For example, the encoder can determine one or more block vectors and explicitly signal the determined block vector(s) to the decoder. The decoder can derive intra-affine candidates for the current block based on a combination of the signaled one or more block vectors.
[0133] For example, one or more block vectors can be obtained by performing block matching based on templates of the current block and / or surrounding blocks, and an intra-affine candidate of the current block can be derived based on a combination of the obtained one or more block vectors.
[0134] An intra-affine candidate of the current block can be derived based on a combination of at least two of the various embodiments described above.
[0135] However, in the process of forming the intra-affine candidate list, a duplication check with all or part of the CPBVPs already added to the intra-affine candidate list may be performed, and intra-affine candidates with overlapping CPBVPs may not be added to the intra-affine candidate list.
[0136] The intra-affine candidates in the intra-affine candidate list generated using the aforementioned method can be reordered based on a predetermined cost. Here, the cost can be calculated based on the template difference (e.g., SAD, MAE, MSE, etc.) between the current block and the reference block.
[0137] The above reference block can be identified based on the CPBVP of the intra-affine candidate. If the intra-affine candidate has multiple CPBVPs, reference blocks corresponding to the multiple CPBVPs can be identified, and template differences can be calculated for each of the reference blocks. The cost can be calculated based on a weighted sum of the calculated template differences. Alternatively, the cost can be calculated based on an average of the calculated template differences.
[0138] Alternatively, a reference block may be identified based on only a specific CPBVP among multiple CPBVPs possessed by the intra-affine candidate, and a template difference may be calculated for the specific reference block. Here, the specific CPBVP may be a CPBVP corresponding to the upper left position.
[0139] Alternatively, among the multiple CPBVPs possessed by the intra-affine candidate, reference blocks corresponding to the top N CPBVPs can be individually identified based on importance. In this case, template differences can be calculated for each of the identified reference blocks. The cost can be calculated based on a weighted sum of the calculated template differences. Alternatively, the cost can be calculated based on an average of the calculated template differences.
[0140] Alternatively, intra-affine prediction can be performed based on the CPBVP of each intra-affine candidate to generate intra-affine prediction blocks. Costs can be calculated based on template differences between the intra-affine prediction blocks.
[0141] Referring to FIG. 4, the CPBV of the current block can be derived based on the intra-affine candidate list (S410).
[0142] The CPBV of the current block can be derived based on any one of multiple intra-affine candidates in the intra-affine candidate list.
[0143] The encoder can determine one intra-affine candidate among multiple intra-affine candidates and signal a candidate index indicating the determined intra-affine candidate to the decoder. The decoder can then specify one of the multiple intra-affine candidates based on the signaled candidate index. Alternatively, the encoder and decoder can implicitly derive the candidate index of the current block using the same method.
[0144] Meanwhile, a control point block vector difference (CPBVD) may be additionally signaled for the current block. In this case, the CPBV of the current block may be derived based on the CPBVP of the intra-affine candidate specified by the candidate index and the CPBVD. The CPBVD may have integer or sub-pixel precision, and block vector scaling may be performed to compensate for the difference in precision.
[0145] CPBVD can be signaled based on information representing the sign and distance of the x-coordinate and y-coordinate, respectively. Alternatively, CPBVD can be signaled based on information representing direction (up, down, left, right, top left, bottom left, top right, bottom right, etc.) and distance by integrating the x-coordinate and y-coordinate. CPBVD can be signaled together with information representing pixel precision.
[0146] A single CPBVD can be signaled for the current block. In this case, the single CPBVD can be applied equally to multiple CPBVPs of the intra-affine candidate. Alternatively, the signaled CPBVD can be corrected based on the control point positions corresponding to the multiple CPBVPs, and the corrected CPBVD can be applied to the corresponding CPBVP.
[0147] A CPBVD may also be signaled for each control point location corresponding to multiple CPBVPs. In this case, a CPBV can be derived for each control point location in the current block based on the corresponding CPBVP and CPBVD.
[0148] Meanwhile, the signaling of CPBVD for the current block may be omitted, in which case the CPBVP of the intra-affine candidate may be derived as the CPBV of the current block.
[0149] Referring to FIG. 4, an intra-affine prediction block of the current block can be generated based on the CPBV of the current block (S420).
[0150] If one CPBV is derived for the current block, an intra-affine prediction block of the current block can be generated based on the CPBV.
[0151] Alternatively, if at least two CPBVs are derived for the current block, a prediction block can be generated based on each CPBV. An intra-affine prediction block of the current block can be generated through a weighted sum of the at least two generated prediction blocks. The weights for the weighted sum can be derived based on the template difference corresponding to each CPBV.
[0152] Alternatively, if at least two CPBVs are derived for the current block, a prediction block can be generated based on each CPBV. A predetermined filter (e.g., a Wiener filter) can be applied to the at least two generated prediction blocks to generate an intra-affine prediction block of the current block.
[0153] Alternatively, the current block can be divided into subblocks of a predetermined size. A block vector can be derived for each subblock based on the CPBV of the current block. Each subblock can be predicted based on the corresponding block vector to generate an intra-affine prediction block of the current block. The predetermined size can be expressed as MxK, where M and K can be integers greater than or equal to 1. The current block can be divided into subblocks of the same size or subblocks of different sizes.
[0154] For example, when three CPBVs are derived for the current block, the block vector in subblock units can be derived as in the following mathematical expressions 1 and 2. In mathematical expressions 1 and 2, CPBV, BV, W, H, x, and y can represent the CPBV derived for the current block, the block vector in subblock units, the width and height of the current block, and the coordinates of the subblock, respectively.
[0155]
[0156]
[0157] Based on at least two of the above-described embodiments, intra-affine prediction blocks can be generated for the current block, and a linear or non-linear filter can be applied to the generated intra-affine prediction blocks to generate a final intra-affine prediction block of the current block. Here, the filter coefficients of the linear or non-linear filter can be derived based on a template region or a pre-defined table.
[0158] The CPBV of the current block can be corrected based on a predetermined block vector difference (BVD), and an intra-affine prediction block of the current block can be generated based on the corrected CPBV. Here, the BVD can be explicitly signaled or implicitly derived.
[0159] For example, the CPBV of the current block can be adjusted to the CPBV with the lowest cost within a given search area. Here, the cost can be calculated based on the aforementioned template difference. Alternatively, the CPBV of the current block can be adjusted based on bilateral matching or optical flow estimation.
[0160] When the current block is divided into sub-blocks, the block vector of the sub-block can be corrected based on a predetermined block vector quantization (BVD), and an intra-affine prediction block of the current block can be generated based on the corrected block vector. Here, the BVD can be explicitly signaled or implicitly derived.
[0161] For example, the block vector of a subblock can be corrected to the block vector with the lowest cost within a given search area. Here, the cost can be calculated based on the template difference described above. Alternatively, the block vector of a subblock can be corrected based on bilateral matching or optical flow estimation.
[0162] The encoder can encode and signal to the decoder information indicating whether the current block uses the aforementioned block vector-based affine prediction. Based on this information, the decoder can determine whether the current block uses the block vector-based affine prediction.
[0163] The various embodiments of the present disclosure are not intended to list all possible combinations but rather to illustrate representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combinations of two or more.
[0164] Additionally, various embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, the embodiments may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.
[0165] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium having such software or instructions stored thereon and executable on the device or computer.
Claims
1. A step of constructing an intra merge list for the current block; wherein the intra merge list includes a plurality of intra merge candidates. A step of deriving an intra prediction mode of the current block based on the intra merge list; and A method for decoding an image, comprising the step of generating a prediction block of the current block based on the intra prediction mode.
2. In paragraph 1, At least one of the plurality of intra merge candidates is derived based on a neighboring block of the current block, A method for decoding an image, wherein the above-mentioned peripheral block includes at least one of an adjacent peripheral block or a non-adjacent peripheral block.
3. In paragraph 1, A method for decoding an image, wherein the plurality of intra merge candidates further include one or more history-based candidates.
4. In paragraph 2, A video decoding method, wherein, when the surrounding block is encoded based on a DIMD (decoder side intra mode derivation) mode, the DIMD mode is added as an intra merge candidate of the current block, or an intra prediction mode derived for the surrounding block based on the DIMD mode is stored in an intra merge candidate of the current block.
5. In paragraph 2, A video decoding method, wherein, when the surrounding block is encoded based on a TIMD (template-based intra mode derivation) mode, the TIMD mode is added as an intra merge candidate of the current block, or an intra prediction mode derived for the surrounding block based on the TIMD mode is stored in an intra merge candidate of the current block.
6. In paragraph 2, A video decoding method, wherein, when the intra prediction mode of the above-mentioned surrounding block is derived based on the MPM list, at least one of the MPM candidates of the MPM list for the above-mentioned surrounding block is added as an intra merge candidate of the current block.
7. In paragraph 2, An image decoding method, wherein index information used by the above-mentioned surrounding blocks for intra prediction is stored in the intra merge candidate of the above-mentioned current block.
8. In paragraph 2, An image decoding method, wherein, when the prediction block of the above-mentioned surrounding block is generated through a weighted sum of a plurality of prediction blocks, the intra prediction modes used to generate the plurality of prediction blocks and / or the weights for the weighted sum are stored in the intra merge candidate of the current block.
9. In paragraph 2, An image decoding method, wherein when the surrounding block is encoded in a geometric division mode, the geometric division mode is added as an intra merge candidate of the current block, or information about the geometric division of the surrounding block is stored in the intra merge candidate of the current block.
10. In paragraph 1, A method for decoding an image, wherein when the above-mentioned surrounding block is encoded in a planar mode, the planar mode is added as an intra merge candidate of the current block, or information identifying the type of the planar mode is stored in the intra merge candidate of the current block.
11. In paragraph 1, A method for decoding an image, wherein one or more intra prediction modes are derived from the intra merge list.
12. In paragraph 11, A method for decoding an image, wherein the one or more intra prediction modes are selected from the top N intra merge candidates among a plurality of intra merge candidates belonging to the intra merge list.
13. In paragraph 11, The above intra merge list is divided into multiple sets, A method for decoding an image, wherein one or more intra prediction modes are derived from at least one of the plurality of sets.
14. In paragraph 11, Costs are calculated for the intra prediction modes of the above multiple intra merge candidates, A video decoding method, wherein the intra prediction mode of the current block is derived based on the top L intra prediction modes with the minimum cost.
15. A step of constructing an intra merge list for the current block; wherein the intra merge list includes a plurality of intra merge candidates. A step of deriving an intra prediction mode of the current block based on the intra merge list; and A video encoding method, comprising the step of generating a prediction block of the current block based on the intra prediction mode.
16. A computer-readable storage medium storing a bitstream generated based on the image encoding method according to Article 15.
Citation Information
Patent Citations
Image coding device
KR101538248B1
Semiconductor devices including substrate structure
KR1020250042997A
Intra merge prediction
US11172203B2
KR20230012098A