Image encoding / decoding method and apparatus, and recording medium on which bit stream is stored
By deriveing reference samples of directional plane mode and generating prediction blocks, applying PDPC and determining transformation cores, the problem of inefficiency of directional plane mode in intra prediction in the prior art is solved, and more efficient image compression is achieved.
Patent Information
- Application Number
- CN202380071647.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-10
- Filing Date
- 2023-10-10
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively use the directional plane mode for intra prediction, resulting in low image compression efficiency.
By deriveing a reference sample for the orientation plane mode and generating a prediction block based on this, the PDPC is applied to eliminate the discontinuity between the prediction block and the adjacent region, and the transformation core is determined to improve inverse transformation performance.
The intra prediction performance is improved, the encoding efficiency of residual signals is improved, and the overall efficiency of image compression is improved.
Smart Images

Figure CN119948866A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and a recording medium storing a bit stream. Background Art
[0002] Recently, demands for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images have been increasing in various application fields, and therefore, efficient image compression technology is being discussed.
[0003] There are various technologies, such as inter-frame prediction technology that uses video compression technology to predict pixel values included in the current picture from pictures before or after the current picture, intra-frame prediction technology that predicts pixel values included in the current picture by using pixel information in the current picture, entropy coding technology that assigns short symbols to values with high occurrence frequency and assigns long symbols to values with low occurrence frequency, etc., and these image compression technologies can be used to effectively compress image data and transmit or store it. Summary of the invention
[0004] Technical issues
[0005] The present disclosure provides methods and apparatus for intra prediction based on directional plane modes.
[0006] The present disclosure provides methods and apparatus for signaling intra prediction mode information for directional planar mode.
[0007] The present disclosure provides methods and apparatus for deriving reference samples for directional planar patterns.
[0008] The present disclosure provides methods and apparatus for applying PDPC to a prediction block according to a directional plane mode.
[0009] The present disclosure provides a method and an apparatus for determining a transformation kernel of a residual signal according to a directional plane pattern.
[0010] Technical Solution
[0011] According to the image decoding method and device disclosed in the present invention, the intra-frame prediction mode of the current block can be derived from the predefined intra-frame prediction mode, the prediction block of the current block can be generated based on the intra-frame prediction mode, the transformation coefficient of the current block can be at least one of dequantization or inverse transformation to obtain the residual block of the current block, and the current block can be reconstructed based on the prediction block and the residual block of the current block.
[0012] In the image decoding method and apparatus according to the present disclosure, the predefined intra prediction modes may include a non-directional plane mode, a directional plane mode, a horizontal mode, and a vertical mode, and the directional plane mode may include at least one of a horizontal plane mode or a vertical plane mode.
[0013] In the image decoding method and apparatus according to the present disclosure, when the intra prediction mode of the current block belongs to the directional plane mode, a transform kernel for inverse transform may be determined based on a transform kernel for a predefined mode.
[0014] In the image decoding method and device according to the present disclosure, when the intra-frame prediction mode of the current block is a horizontal plane mode, the transform core for inverse transformation can be determined based on the transform core for the vertical mode, and when the intra-frame prediction mode of the current block is a vertical plane mode, the transform core for inverse transformation can be determined based on the transform core for the horizontal mode.
[0015] In the image decoding method and device according to the present disclosure, when the intra-frame prediction mode of the current block is a horizontal plane mode, the transform core for inverse transformation can be determined based on the transform core for the horizontal mode, and when the intra-frame prediction mode of the current block is a vertical plane mode, the transform core for inverse transformation can be determined based on the transform core for the vertical mode.
[0016] In the image decoding method and apparatus according to the present disclosure, when the intra prediction mode of the current block belongs to the directional plane mode, the transform kernel for inverse transform may be determined based on the transform kernel for the non-directional plane mode.
[0017] In the image decoding method and device according to the present disclosure, the intra-frame prediction mode of the current block is derived based on the intra-frame prediction mode information, and the intra-frame prediction mode information may include at least one of a plane flag indicating whether the intra-frame prediction mode of the current block is a non-directional plane mode or belongs to a directional plane mode, or a plane direction flag indicating whether the intra-frame prediction mode of the current block is a horizontal plane mode.
[0018] In the image decoding method and apparatus according to the present disclosure, based on the plane flag indicating that the intra prediction mode of the current block does not belong to the directional plane mode, an MPM flag indicating whether the intra prediction mode of the current block is derived from the MPM list can be signaled.
[0019] In the image decoding method and apparatus according to the present disclosure, at least one of a plane flag or a plane direction flag may be adaptively signaled based on an availability flag indicating whether decoder-side intra mode derivation (DIMD) is available.
[0020] In the image decoding method and apparatus according to the present disclosure, at least one of a plane flag or a plane direction flag may be adaptively signaled based on an availability flag indicating whether template-based intra mode derivation (TIMD) is available.
[0021] In the image decoding method and apparatus according to the present disclosure, when the intra prediction mode of the current block belongs to the directional plane mode, the reference sample for the current block may be derived based on the reference sample for the horizontal mode or the vertical mode.
[0022] In the image decoding method and apparatus according to the present disclosure, when the intra prediction mode of the current block belongs to the directional plane mode, the reference sample for the current block may be derived based on the reference sample for the non-directional plane mode.
[0023] According to the image encoding method and apparatus disclosed herein, a prediction block of a current block can be generated based on one of the predefined intra prediction modes, a residual block of the current block can be derived based on the prediction block of the current block, a transform coefficient of the current block can be derived by performing at least one of transformation or quantization on the residual block, and the transform coefficient of the current block can be encoded. Here, the predefined intra prediction mode may include a non-directional plane mode, a directional plane mode, a horizontal mode, and a vertical mode, and the directional plane mode may include at least one of a horizontal plane mode or a vertical plane mode.
[0024] A computer-readable digital storage medium is provided, which stores encoded video / image information, thereby causing a decoding device according to the present disclosure to perform an image decoding method.
[0025] A computer-readable digital storage medium storing video / image information generated according to an image encoding method according to the present disclosure is provided.
[0026] A method and apparatus for transmitting video / image information generated according to an image encoding method according to the present disclosure are provided.
[0027] Beneficial Effects
[0028] According to the present disclosure, the existing plane mode can be extended to the directional plane mode, thereby improving the intra prediction performance based on the plane with a high selection probability.
[0029] According to the present disclosure, encoding efficiency of intra prediction mode information indicating non-directional / directional planar mode can be improved.
[0030] According to the present disclosure, prediction characteristics of a directional plane pattern may be considered to derive reference samples, thereby improving intra prediction performance.
[0031] According to the present disclosure, PDPC may be applied to eliminate discontinuity between a prediction block and a neighboring region according to a directional plane mode, thereby improving encoding efficiency of a residual signal.
[0032] According to the present disclosure, by determining a transform or a kernel based on the correlation of residuals between intra prediction modes, the performance of the transform can be improved and better energy compression can be expected. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 A video / image compilation system according to the present disclosure is shown.
[0034] Figure 2 A schematic block diagram showing an encoding device to which an embodiment of the present disclosure is applicable and which performs encoding of a video / image signal.
[0035] Figure 3 A schematic block diagram showing a decoding device to which an embodiment of the present disclosure is applicable and which performs decoding of a video / image signal.
[0036] Figure 4 An image decoding method performed by the decoding device 300 according to an embodiment of the present disclosure is shown.
[0037] Figure 5 A schematic configuration of a decoding device 300 that performs an image decoding method according to the present disclosure is shown.
[0038] Figure 6 An image encoding method performed by the encoding device 200 as an embodiment according to the present disclosure is shown.
[0039] Figure 7 A schematic configuration of an encoding device 200 that performs an image encoding method according to the present disclosure is shown.
[0040] Figure 8 An example of a content streaming system to which an embodiment of the present disclosure can be applied is shown. DETAILED DESCRIPTION
[0041] Because the present disclosure can make various changes and has several embodiments, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. However, it is not intended to limit the present disclosure to specific embodiments, and it should be understood to include all changes, equivalents and substitutes included in the spirit and technical scope of the present disclosure. When describing each of the drawings, similar reference numerals are used for similar components.
[0042] Terms such as first, second, etc. can be used to describe various components, but components should not be limited by these terms. These terms are only used to distinguish one component from other components. For example, without departing from the scope of the present disclosure, a first component can be referred to as a second component, and similarly, a second component can also be referred to as a first component. Terms and / or combinations of any one or more related statement items in a plurality of related statement items are included.
[0043] When a component is referred to as being "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to another component, but another component may also exist in between. On the other hand, when a component is referred to as being "directly connected" or "directly linked" to another component, it should be understood that another component does not exist in between.
[0044] The terms used in this application are only used to describe specific embodiments and are not intended to limit the present disclosure. Unless the context clearly indicates otherwise, singular expressions include plural expressions. In this application, it should be understood that terms such as "including" or "having" are intended to designate the existence of features, numbers, steps, operations, components, parts or combinations thereof described in the specification, but do not exclude the possibility of the existence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof in advance.
[0045] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the Universal Video Coding (VVC) standard. In addition, the methods / embodiments disclosed herein may be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation audio video coding standard (AVS2), or the next generation video / image coding standard (e.g., H.267 or H.268, etc.).
[0046] This specification proposes various embodiments of video / image coding, and unless otherwise stated, these embodiments may be performed in combination with each other.
[0047] Here, video may refer to a collection of a series of images over time. A picture generally refers to a unit representing an image within a specific time period, and a slice / tile is a unit that forms a part of a picture in coding. A slice / tile may include at least one coding tree unit (CTU). A picture may be composed of at least one slice / tile. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of a CTU having the same height as the picture and a width assigned by the syntax requirements of the picture parameter set. A tile row is a rectangular area of a CTU having a height assigned by the picture parameter set and the same width as the width of the picture. The CTU within a tile may be arranged continuously according to a CTU raster scan, and the tiles within a picture may be arranged continuously according to a raster scan of the tile. A slice may include an integer number of complete tiles or an integer number of continuous complete CTU rows within a tile of a picture that may be exclusively included in a single NAL unit. At the same time, a picture may be divided into at least two sub-pictures. A sub-picture may be a rectangular area of at least one slice within a picture.
[0048] Pixel, pixel or picture element may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component, or only a pixel / pixel value of a chrominance component.
[0049] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with the corresponding region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, a unit may be used interchangeably with terms such as a block or region. In general, an MxN block may include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.
[0050] Here, "A or B" may refer to "only A", "only B", or "both A and B". In other words, herein, "A or B" may be interpreted as "A and / or B". For example, herein, "A, B or C" may refer to "only A", "only B", "only C", or "any combination of A, B, and C".
[0051] As used herein, a slash ( / ) or a comma may mean "and / or". For example, "A / B" may mean "A and / or B". Thus, "A / B" may mean "only A", "only B", or "both A and B". For example, "A, B, C" may mean "A, B, or C".
[0052] Here, "at least one of A and B" may refer to "only A", "only B", or "both A and B". In addition, herein, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same manner as "at least one of A and B".
[0053] In addition, herein, “at least one of A, B, and C” may refer to “only A”, “only B”, “only C”, or “any combination of A, B, and C”. In addition, “at least one of A, B, or C” or “at least one of A, B and / or C” may refer to “at least one of A, B, and C”.
[0054] In addition, the brackets used herein may refer to "for example". Specifically, when the indication is "prediction (intra-frame prediction)", "intra-frame prediction" may be proposed as an example of "prediction". In other words, the "prediction" here is not limited to "intra-frame prediction", and "intra-frame prediction" may be proposed as an example of "prediction". In addition, even when the indication is "prediction (ie, intra-frame prediction)", "intra-frame prediction" may be proposed as an example of "prediction".
[0055] Here, technical features described individually in one drawing may be implemented individually or simultaneously.
[0056] Figure 1 A video / image compilation system according to the present disclosure is shown.
[0057] refer to Figure 1 , a video / image coding system may include a first device (source device) and a second device (receiving device).
[0058] The source device may send the encoded video / image information or data to the receiving device in the form of a file or stream transmission through a digital storage medium or a network. The source device may include a video source, an encoding device, and a sending unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0059] The video source may obtain the video / image through the process of capturing, synthesizing or generating the video / image. The video source may include a device for capturing the video / image and a device for generating the video / image. The device for capturing the video / image may include at least one camera, a video / image archive including previously captured videos / images, etc. The device for generating the video / image may include a computer, a tablet computer, a smart phone, etc., and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc., and in this case, the process of capturing the video / image may be replaced by the process of generating the relevant data.
[0060] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bit stream.
[0061] The sending unit may send the encoded video / image information or data output in the form of a bit stream to the receiving unit of the receiving device in the form of a file or stream transmission through a digital storage medium or a network. The digital storage medium may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The sending unit may include an element for generating a media file in a predetermined file format and may include an element for transmission through a broadcast / communication network. The receiving unit may receive / extract a bit stream and send it to a decoding device.
[0062] The decoding device may decode the video / image by performing a series of processes such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operations of the encoding device.
[0063] The renderer may render the decoded video / image. The rendered video / image may be displayed through a display unit.
[0064] Figure 2 A rough block diagram of an encoding device to which an embodiment of the present disclosure can be applied and which performs encoding of a video / image signal is shown.
[0065] refer to Figure 2 , the encoding device 200 may be composed of an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the above-mentioned image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include a memory 270 as an internal / external component.
[0066] The image divider 210 may divide the input image (or picture, frame) input to the encoding device 200 into at least one processing unit. As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure.
[0067] For example, one coding unit may be segmented into a plurality of coding units having a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quadtree structure. The coding process according to this specification may be performed based on a final coding unit that is no longer segmented. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit may be directly used as the final coding unit, or if necessary, the coding unit may be recursively segmented into coding units of a deeper depth, and the coding unit with the best size may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction described later.
[0068] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the above-mentioned final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0069] In some cases, a unit may be used interchangeably with terms such as a block or region. In general, an MxN block may represent a set of transform coefficients or samples consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component, or only a pixel / pixel value of a chroma component. A sample may be used as a term to make a picture (or image) correspond to a pixel or a picture element.
[0070] The encoding device 200 may subtract the prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device 200 may be referred to as a subtractor 231.
[0071] The predictor 220 may perform prediction on a block to be processed (hereinafter referred to as a current block) and generate a predicted block including a prediction sample for the current block. The predictor 220 may determine whether intra prediction or inter prediction is applied in units of a current block or CU. The predictor 220 may generate various information about the prediction, such as prediction mode information, and send it to the entropy encoder 240, as described later in the description of each prediction mode. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0072] The intra-frame predictor 222 can predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the referenced sample can be located near the current block or can be located at a certain distance away from the current block. In intra-frame prediction, the prediction mode may include at least one non-directional mode and multiple directional modes. The non-directional mode may include at least one of the DC mode or the plane mode. Depending on the detail level of the prediction direction, the directional mode may include 33 directional modes or 65 directional modes. However, this is only an example, and more or less directional modes may be used depending on the configuration. The intra-frame predictor 222 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0073] The inter-frame predictor 221 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information sent in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter-frame predictor 221 may configure a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, for skip mode and merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information of the current block. For skip mode, unlike merge mode, a residual signal may not be transmitted. For motion vector prediction (MVP) mode, motion vectors of surrounding blocks are used as motion vector predictors, and motion vector differences are signaled to indicate the motion vector of the current block.
[0074] The predictor 220 may generate a prediction signal based on various prediction methods described later. For example, the predictor may not only apply intra prediction or inter prediction to predict a block, but may also apply intra prediction and inter prediction at the same time. It may be referred to as a combined inter and intra prediction (CIIP) mode. In addition, the predictor may be based on an intra block copy (IBC) prediction mode or may be based on a palette mode for prediction of a block. The IBC prediction mode or the palette mode may be used for content image / video coding of games, such as screen content coding (SCC), etc. IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction because it derives a reference block within the current picture. In other words, IBC may use at least one of the inter prediction techniques described herein. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture may be signaled based on information about the palette table and the palette index. The prediction signal generated by the predictor 220 may be used to generate a reconstructed signal or a residual signal.
[0075] The transformer 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously reconstructed pixels. In addition, the transform process may be applied to square pixel blocks of the same size or may be applied to non-square blocks of variable size.
[0076] The quantizer 233 may quantize the transform coefficients and send them to the entropy encoder 240, and the entropy encoder 240 may encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in the block form into a one-dimensional vector form based on the coefficient scanning order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0077] The entropy encoder 240 may perform various encoding methods such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc. The entropy encoder 240 may encode information necessary for video / video image reconstruction (e.g., values of syntax elements, etc.) in addition to transform coefficients quantized together or individually.
[0078] The encoded information (e.g., encoded video / image information) can be transmitted or stored in units of network abstraction layer (NAL) units in the form of a bitstream. The video / image information may further include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. Here, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded and included in the bitstream through the above-mentioned encoding process. The bitstream may be transmitted through a network or may be stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmission and / or a storage unit (not shown) for storing a signal output from the entropy encoder 240 may be configured as an internal / external element of the encoding device 200, or the transmission unit may also be included in the entropy encoder 240.
[0079] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients through the inverse quantizer 234 and the inverse transformer 235. The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual of the block to be processed, such as when the skip mode is applied, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can also be used for inter-frame prediction of the next picture through filtering described later. At the same time, luminance mapping with chroma scaling (LMCS) can be applied in the picture encoding and / or reconstruction process.
[0080] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 270, specifically in the DPB of the memory 270. Various filtering methods may include deblocking filtering, sample adaptation offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information about filtering and send it to the entropy encoder 240. The information about filtering can be encoded in the entropy encoder 240 and output in the form of a bit stream.
[0081] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter predictor 221. When inter prediction is applied therethrough, the encoding apparatus can avoid prediction mismatch in the encoding apparatus 200 and the decoding apparatus, and can also improve encoding efficiency.
[0082] The DPB of the memory 270 may store the modified reconstructed picture to be used as a reference picture in the inter-frame predictor 221. The memory 270 may store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the block in the pre-reconstructed picture. The stored motion information may be sent to the inter-frame predictor 221 to be used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 270 may store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 222.
[0083] Figure 3 A rough block diagram of a decoding device to which an embodiment of the present disclosure can be applied and which performs decoding of a video / image signal is shown.
[0084] refer to Figure 3 , the decoding apparatus 300 may be configured by including an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321.
[0085] According to an embodiment, the above-mentioned entropy decoder 310, residual processor 320, predictor 330, adder 340 and filter 350 may be configured by one hardware component (e.g., decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0086] When a bit stream including video / image information is input, the decoding apparatus 300 may generate a decoded signal in response to the bit stream received in the decoded signal. Figure 2 The image is reconstructed by the process of processing video / image information in the encoding device of the decoding device. For example, the decoding device 300 can derive the unit / block based on the relevant information of the block segmentation obtained from the bit stream. The decoding device 300 can perform decoding by using the processing unit applied in the encoding device. Therefore, the processing unit of decoding can be a coding unit, and the coding unit can be divided from the coding tree unit or the maximum coding unit according to the quadtree structure, the binary tree structure and / or the ternary tree structure. At least one transform unit can be derived from the coding unit. And, the reconstructed image signal decoded and output by the decoding device 300 can be played by a playback device.
[0087] The decoding device 300 may receive the bit stream from Figure 2The received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device may further decode the picture based on the information about the parameter set and / or the general constraint information. The signaled / received information and / or the syntax elements described later in this document may be decoded and obtained from the bitstream by a decoding process. For example, the entropy decoder 310 may decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, CABAC, etc., and output the values of the syntax elements necessary for image reconstruction and the quantized values of the transform coefficients of the residual. In more detail, the CABAC entropy decoding method can receive a bin corresponding to each syntax element from a bitstream, determine a context model by using information of syntax elements to be decoded, decoding information of surrounding blocks and blocks to be decoded, or information of symbols / bins decoded in the previous step, perform arithmetic decoding on bins by predicting the probability of occurrence of bins according to the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using information about decoded symbols / bins of the context model for the next symbol / bin. Among the information decoded in the entropy decoder 310, information about prediction is provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value, that is, the quantized transform coefficient and related parameter information, which is entropy decoded in the entropy decoder 310, can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300 or the receiving unit may be a component of the entropy decoder 310 .
[0088] Meanwhile, the decoding device according to this specification may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of an inverse quantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0089] The inverse quantizer 321 may inverse quantize the quantized transform coefficient and output the transform coefficient. The inverse quantizer 321 may rearrange the quantized transform coefficient into a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantizer 321 may perform inverse quantization on the quantized transform coefficient by using a quantization parameter (e.g., quantization step size information) and obtain the transform coefficient.
[0090] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0091] The predictor 320 may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor 320 may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310, and determine a specific intra / inter prediction mode.
[0092] The predictor 320 can generate a prediction signal based on various prediction methods described later. For example, the predictor 320 can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction at the same time. It can be called a combined inter and intra prediction (CIIP) mode. In addition, the predictor can be based on an intra block copy (IBC) prediction mode or can be based on a palette mode for prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding of games, such as screen content coding (SCC), etc. IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction because it derives a reference block within the current picture. In other words, IBC can use at least one of the inter prediction techniques described herein. The palette mode can be considered as an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and the palette index can be included in the video / image information and sent with a signal.
[0093] The intra-frame predictor 331 can predict the current block by referring to samples within the current picture. Depending on the prediction mode, the referenced sample can be located near the current block or can be located at a certain distance away from the current block. In intra-frame prediction, the prediction mode may include at least one non-directional mode and multiple directional modes. The intra-frame predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0094] The inter-frame predictor 332 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information sent in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter-frame predictor 332 may configure a motion information candidate list based on the neighboring blocks, and derive a motion vector and / or a reference picture index for the current block based on the received candidate selection information. Inter-frame prediction may be performed based on various prediction modes, and information about the prediction may include information indicating an inter-frame prediction mode for the current block.
[0095] The adder 340 may add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331) to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual of the block to be processed, such as when the skip mode is applied, the prediction block may be used as the reconstructed block.
[0096] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, may be output through filtering described later, or may be used for inter prediction of the next picture. At the same time, luminance mapping with chroma scaling (LMCS) may be applied during picture decoding.
[0097] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and send the modified reconstructed picture to the memory 360, specifically the DPB of the memory 360. The various filtering methods may include deblocking filtering, sampling adaptation offset, adaptive loop filter, bilateral filter, etc.
[0098] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can be derived from the motion information in its current picture (or decoded) The motion information of the block and / or the motion information of the block in the pre-reconstructed picture. The stored motion information can be sent to the inter-frame predictor 332 to be used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 331.
[0099] Here, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may also be equally or correspondingly applied to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300, respectively.
[0100] Figure 4 An image decoding method performed by the decoding device 300 as an embodiment according to the present disclosure is shown.
[0101] refer to Figure 4 , the intra prediction mode of the current block can be derived S400.
[0102] The intra prediction mode of the current block can be derived from one of the predefined intra prediction modes. The predefined intra prediction mode may include at least one of a non-directional mode or a directional mode. The non-directional mode may include at least one of a plane mode or a DC mode. The plane mode according to the present disclosure includes a non-directional plane mode and a directional plane mode, and the directional plane mode may include at least one of a horizontal plane mode or a vertical plane mode. Alternatively, the directional plane mode may be defined as a mode independent of the non-directional plane mode, in which case the plane mode according to the present disclosure may mean a non-directional plane mode. The directional mode may mean a mode with a predetermined angle, such as a horizontal mode, a vertical mode, a diagonal mode, etc.
[0103] The intra prediction mode of the current block may be derived based on information for specifying the intra prediction mode (hereinafter, referred to as intra prediction mode information). The intra prediction mode information according to the present disclosure may include at least one of an MPM flag, a plane flag, a plane direction flag, an MPM index, or residual mode information.
[0104] The intra-frame prediction mode information may be defined differently depending on whether the planar mode is defined as a mode of a candidate mode independent of the MPM list. As an example, based on a mode in which the planar mode is not defined as a mode of a candidate mode independent of the MPM list, the intra-frame prediction mode information may include at least one of an MPM flag, an MPM index, or residual mode information. On the other hand, based on a mode in which the planar mode is defined as a mode of a candidate mode independent of the MPM list (i.e., the planar mode may be used as a candidate mode of the MPM list), the intra-frame prediction mode may include at least one of an MPM flag, a plane flag, a plane direction flag, an MPM index, or residual mode information.
[0105] The MPM flag may indicate whether the intra prediction mode of the current block is derived from an MPM list including multiple candidate modes (most probable modes, MPMs). The plane flag may include at least one of a first plane flag indicating whether the intra prediction mode of the current block belongs to a plane mode or a second plane flag indicating whether the intra prediction mode of the current block is a non-directional plane mode. The second plane flag may be defined as a flag indicating whether the intra prediction mode of the current block belongs to a directional plane mode. The plane direction flag may indicate whether the intra prediction mode of the current block is a horizontal plane mode. The MPM index may specify any one of the multiple candidate modes in the MPM list. The residual mode information may specify any one of the remaining modes among the predefined intra prediction modes that do not include the plane mode and the candidate modes belonging to the MPM list.
[0106] Hereinafter, a method for signaling intra prediction mode information of a current block when a planar mode is a mode independent of a candidate mode from an MPM list will be described.
[0107] The horizontal / vertical plane mode may be signaled as one of the plane modes as shown in Table 1 below.
[0108] [Table 1]
[0109]
[0110] Referring to Table 1, the MPM flag (intra_luma_mpm_flag) can be obtained from the bitstream. Based on the MPM flag derived from the MPM list indicating the intra prediction mode of the current block, the first plane flag (intra_luma_not_planar_flag) can be obtained from the bitstream. Based on the first plane flag indicating that the intra prediction mode of the current block does not belong to the planar mode, the MPM index can be obtained from the bitstream. The intra prediction mode of the current block can be derived as a candidate mode specified by the MPM index.
[0111] Based on the first plane flag indicating that the intra prediction mode of the current block belongs to the plane mode, the second plane flag (planar_flag) can be obtained from the bitstream. Based on the second plane flag indicating that the intra prediction mode of the current block is a non-directional plane mode, the intra prediction mode of the current block can be derived as a non-directional plane mode. Based on the second plane flag indicating that the intra prediction mode of the current block is not a non-directional plane mode, the plane direction flag (planar_dir_flag) can be obtained from the bitstream. Based on the plane direction flag indicating that the intra prediction mode of the current block is a horizontal plane mode, the intra prediction mode of the current block can be derived as a horizontal plane mode. On the other hand, based on the plane direction flag indicating that the intra prediction mode of the current block is a vertical plane mode, the intra prediction mode of the current block can be derived as a vertical plane mode. When the ISP mode (intra-frame sub-partition mode) is not applied to the current block, the second plane flag according to the present disclosure can be sent by signal. The ISP mode may mean a mode in which the current block is divided into a plurality of sub-partitions and intra prediction is performed based on the sub-partitions.
[0112] Based on the MPM flag indicating that the intra prediction mode of the current block is not derived from the MPM list, residual mode information (intra_luma_mpm_remainder) may be obtained from the bitstream. The intra prediction mode of the current block may be derived as a mode specified by the residual mode information.
[0113] The second plane flag and the plane direction flag may be entropy decoded based on the CABAC method. Alternatively, because the plane direction flag is not consistent when selecting the prediction direction, it may be bypass coded so that the context is not updated with a probability of 0.5 each. Flags indicating the availability of higher-level horizontal / vertical plane modes may be signaled, such as video parameter sets (VPS), sequence parameter sets (SPS), picture parameter sets (PPS), picture headers (PH), and slice headers (SH).
[0114] Alternatively, the horizontal / vertical plane mode may be signaled as a mode independent of the non-directional plane mode, as shown in Table 2 below.
[0115] [Table 2]
[0116]
[0117] Referring to Table 2, the second plane flag (planar_horver_flag) can be obtained from the bitstream. Here, the second plane flag can indicate whether the intra prediction mode of the current block belongs to the directional plane mode. As an example, based on the value of the second plane flag being 1, this can indicate that the intra prediction mode of the current block is a horizontal plane mode or a vertical plane mode. Based on the value of the second plane flag being 0, this can indicate that the intra prediction mode of the current block is neither a horizontal plane mode nor a vertical plane mode.
[0118] Based on the second plane flag indicating that the intra prediction mode of the current block belongs to the directional plane mode, the plane direction flag can be obtained from the bitstream. Based on the plane direction flag indicating that the intra prediction mode of the current block is a horizontal plane mode, the intra prediction mode of the current block can be derived as a horizontal plane mode. On the other hand, based on the plane direction flag indicating that the intra prediction mode of the current block is a vertical plane mode, the intra prediction mode of the current block can be derived as a vertical plane mode. When the ISP mode (intra-frame sub-partitioning mode) is not applied to the current block, the second plane flag according to the present disclosure can be sent by signal.
[0119] Based on the second plane flag indicating that the intra prediction mode of the current block does not belong to the directional plane mode, at least one of the MPM flag, the first plane flag, the MPM index or the residual mode information mentioned above can be obtained from the bitstream, and the intra prediction mode of the current block can be derived based on it. This is the same as described by reference to Table 1, and repeated description will be omitted here.
[0120] The second plane flag and the plane direction flag may be entropy decoded based on the CABAC method. Alternatively, because the plane direction flag is not consistent in selecting the prediction direction, it may be bypass coded so that the context is not updated with a probability of 0.5 each. A signal indicating the availability of a higher level horizontal / vertical plane mode may be sent, such as a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), and a slice header (SH).
[0121] As described above, based on at least one of the second plane flag or the plane direction flag, the intra prediction mode of the current block may be derived as one of the non-directional plane mode, the horizontal plane mode, or the vertical plane mode.
[0122] The second plane flag and the plane direction flag according to the present disclosure may be signaled via a bitstream, which is the same as described in Table 1 and Table 2.
[0123] Alternatively, in Table 1 and Table 2, the availability flag indicating the availability of DIMD (decoder-side intra-mode derivation) can be further regarded as a signaling condition of the second plane flag. In this case, the horizontal or vertical plane mode can be derived based on the intra-prediction mode derived by DIMD (hereinafter referred to as DIMD mode) without explicitly signaling the plane direction flag. In other words, based on the second plane flag indicating that the intra-prediction mode of the current block is a non-directional plane mode, the intra-prediction mode of the current block can be derived as a non-directional plane mode. On the other hand, based on the second plane flag indicating that the intra-prediction mode of the current block is not a non-directional plane mode, the intra-prediction mode of the current block can be derived as a horizontal or vertical plane mode based on the DIMD mode.
[0124] Specifically, the second plane flag may be adaptively signaled based on an availability flag indicating whether the ISP mode is applied to the current block and whether the DIMD is available. For example, when the ISP mode is not applied to the current block and the availability flag indicates that the DIMD is available, the second plane flag may be signaled, and when this is not the case, the flag may not be signaled. Alternatively, when the availability flag indicates that the DIMD is available, the second plane flag may be signaled, and otherwise not signaled, regardless of whether the ISP mode is applied to the current block. The availability flag may be signaled in at least one of the VPS, PPS, PH, or SH.
[0125] Hereinafter, a method for deriving a DIMD mode and a method for deriving an intra prediction mode of a current block will be described.
[0126] A gradient may be derived based on at least two samples of a neighboring area belonging to the current block. Here, the gradient may include at least one of a horizontal gradient or a vertical gradient. An intra-frame prediction mode may be derived based on at least one of the derived gradient or the magnitude of the gradient. Here, the magnitude of the gradient may be determined based on the sum of the horizontal gradient and the vertical gradient. Through this derivation method, one intra-frame prediction mode may be derived, or two or more intra-frame prediction modes may be derived.
[0127] As an example, the gradient can be calculated in units of windows of a predetermined size. An angle indicating the directionality of samples within the window can be calculated based on the calculated gradient. The calculated angle can correspond to any of the above-mentioned predefined intra-frame prediction modes. The magnitude of the gradient can be stored / updated for the intra-frame prediction mode corresponding to the calculated angle. Through this process, the intra-frame prediction mode corresponding to the calculated gradient can be determined for each window, and the magnitude of the gradient can be stored / updated for the determined intra-frame prediction mode. Among the magnitudes of the stored gradients, the first T intra-frame prediction modes with the largest magnitude can be selected. Here, T can be an integer of 1, 2, 3 or more. The selected intra-frame prediction mode can be set to DIMD mode.
[0128] The adjacent area used to calculate the slope may include at least one of a left area, an upper area, an upper left area, a lower left area, or an upper right area adjacent to the current block, which are previously reconstructed areas of the current block. The adjacent area may include at least one of an adjacent sample line adjacent to the current block, a first non-adjacent sample line 1-sample away from the current block, and a second non-adjacent sample line 2-sample away from the current block. However, it is not limited thereto, and may further include a non-adjacent sample line N-samples away from the current block, and N may be an integer greater than or equal to 3.
[0129] The neighboring area may be an area predefined in the same manner in the encoding device and the decoding device to calculate the slope. Alternatively, the neighboring area may be variably determined based on information specifying the position of the neighboring area. In this case, the information specifying the position of the neighboring area may be signaled through the bitstream. Alternatively, the position of the neighboring area may be determined based on at least one of whether the current block is located at the boundary of the coding tree unit, the size of the current block (e.g., width, height, ratio of width and height, product of width and height), the partition type of the current block, the prediction mode of the neighboring area, or the availability of the neighboring area.
[0130] As an example, based on the current block being located at the upper boundary of the coding tree unit, at least one of the upper area, the upper left area, or the upper right area of the current block may not be referenced to calculate the gradient. When the width of the current block is greater than the height, the gradient may be calculated with reference to the upper area or the left area (e.g., the upper area), and the gradient may not be calculated with reference to the other (e.g., the left area). On the contrary, based on the width of the current block being less than the height, the gradient may be calculated with reference to the upper area or the left area (e.g., the left area), and the gradient may not be calculated with reference to the other area (e.g., the upper area). Based on generating the current block by block partitioning in the horizontal direction, the gradient may not be calculated with reference to the upper area. On the contrary, based on generating the current block by block partitioning in the vertical direction, the gradient may not be calculated with reference to the left area. Based on encoding the neighboring area of the current block in inter-frame mode, the gradient may not be calculated with reference to the neighboring area. However, it is not limited thereto, and the gradient may be calculated with reference to the neighboring area regardless of the prediction mode of the neighboring area.
[0131] Based on the fact that the value of the DIMD mode (or the mode with the maximum gradient) is less than the value of the upper left diagonal mode, it is determined that the probability of the intra-frame prediction mode being a horizontal prediction mode is high, and the intra-frame prediction mode of the current block can be inferred as a horizontal plane mode. On the other hand, based on the fact that the value of the DIMD mode is greater than or equal to the value of the upper left diagonal mode, it is determined that the probability of the intra-frame prediction mode being a vertical prediction mode is high, and the intra-frame prediction mode of the current block can be inferred as a vertical plane mode. For example, when the predefined directional mode is defined from the lower left diagonal mode of mode number 2 to the upper right diagonal mode of mode number 66, the upper left diagonal mode can correspond to mode number 34.
[0132] Based on the DIMD mode being a non-directional plane mode or a DC mode, the intra-frame prediction mode of the current block can be inferred to be a horizontal plane mode. In this case, because mode numbers 0 and 1 are assigned to the non-directional plane mode and the DC mode, respectively, the intra-frame prediction mode can be inferred without additional conditions. Alternatively, based on the DIMD mode being a non-directional plane mode or a DC mode, the intra-frame prediction mode of the current block can be inferred to be a vertical plane mode. Typically, because the edge of an image may be in a vertical direction, improvements in prediction performance can be expected by using a vertical plane mode.
[0133] Table 3 below is an example of a signaling method for the second plane flag.
[0134] [Table 3]
[0135]
[0136] Referring to Table 3, planar_flag indicates whether the intra prediction mode of the current block is a non-directional plane mode, which may correspond to a second plane flag according to the present disclosure. For example, based on planar_flag being 1, this may indicate that the intra prediction mode of the current block is a non-directional plane mode. Based on planar_flag being 0, this may indicate that the intra prediction mode of the current block is a horizontal or vertical plane mode. intra_subpartitions_mode_flag may indicate whether the ISP mode is applied to the current block, and sps_dimd_enabled_flag may indicate whether DIMD is available.
[0137] When the ISP mode is not applied to the current block (intra_subpartitions_mode_flag=0) and DIMD is available (sps_dimd_enabled_flag=1), planar_flag can be obtained from the bitstream. Here, it is assumed that sps_dimd_enabled_flag is signaled in the sequence parameter set, but is not limited to this. Based on the value of planar_flag being 1, the intra-frame prediction mode of the current block can be derived as a non-directional planar mode. On the other hand, based on the value of planar_flag being 0, the intra-frame prediction mode of the current block can be derived as a horizontal or vertical plane mode based on the DIMD mode as described above.
[0138] When the ISP mode is applied to the current block or DIMD is not available, planar_flag can be derived as 1 without being obtained from the bitstream. In other words, the intra prediction mode of the current block can be derived as a non-directional planar mode.
[0139] In Table 3, planar_flag is signaled depending on intra_subpartitions_mode_flag and sps_dimd_enabled_flag, but this is only an example. In other words, planar_flag may be signaled depending on sps_dimd_enabled_flag regardless of intra_subpartitions_mode_flag.
[0140] Alternatively, in Table 1 and Table 2, the availability flag indicating the availability of DIMD may be further regarded as a signaling condition of the plane direction flag. When the availability flag indicates that DIMD is available, the plane direction flag may not be signaled, and otherwise the plane direction flag may be signaled. Based on the plane direction flag being signaled, the intra-frame prediction mode of the current block may be derived as a horizontal or vertical plane mode depending on the value of the plane direction flag. On the other hand, based on the plane direction flag not being signaled, the intra-frame prediction mode of the current block may be derived as a horizontal or vertical plane mode based on the DIMD mode as described above.
[0141] Table 4 below is an example of a signaling method of a plane direction flag.
[0142] [Table 4]
[0143]
[0144] Referring to Table 4, as seen in Table 3, planar_flag may indicate whether the intra prediction mode of the current block is a non-directional plane mode, intra_subpartitions_mode_flag may indicate whether the ISP mode is applied to the current block, and sps_dimd_enabled_flag may indicate whether DIMD is available. planar_dir_flag indicates whether the intra prediction mode of the current block is a horizontal plane mode, which may correspond to a plane direction flag according to the present disclosure.
[0145] When ISP mode is not applied to the current block (intra_subpartitions_mode_flag = 0), planar_flag can be obtained from the bitstream. Regardless of sps_dimd_enabled_flag, planar_flag can be obtained from the bitstream.
[0146] Based on the value of planar_flag being 1, the intra-frame prediction mode of the current block can be derived as a non-directional plane mode. On the other hand, based on the value of planar_flag being 0, planar_dir_flag can be adaptively signaled based on sps_dimd_enabled_flag. Based on the value of planar_flag being 0 and sps_dimd_enabled_flag being 0, planar_dir_flag can be obtained from the bitstream. When the value of planar_dir_flag is 1, the intra-frame prediction mode of the current block can be derived as a horizontal plane mode, and when the value of planar_dir_flag is 0, the intra-frame prediction mode of the current block can be derived as a vertical plane mode. Based on the value of planar_flag being 0 and sps_dimd_enabled_flag being 1, planar_dir_flag cannot be obtained from the bitstream. In this case, based on the DIMD mode as described above, the intra-frame prediction mode of the current block can be derived as a horizontal or vertical plane mode.
[0147] In Table 4, planar_flag is signaled depending on intra_subpartitions_mode_flag, but this is only an example. In other words, planar_flag may be signaled regardless of intra_subpartitions_mode_flag.
[0148] Alternatively, without signaling the above-mentioned second plane flag and plane direction flag, the intra-frame prediction mode of the current block can be derived as one of the non-directional plane mode, the horizontal plane mode or the vertical plane mode based on the DIMD mode.
[0149] When the DIMD mode falls within a predetermined first range determined based on the value (modeH) of the horizontal mode as a directional mode, the intra prediction mode of the current block can be inferred to be a horizontal plane mode. Here, the predetermined range may mean a range from a value obtained by subtracting M (modeH-M) from the value of the horizontal mode to a value obtained by adding M (modeH+M) to the value of the horizontal mode. For example, when the value of the horizontal mode is 18, M is 5, and the value of the DIMD mode is 16, the value of the DIMD mode falls within a range from 13 to 23, and therefore, the intra prediction mode of the current block can be inferred to be a horizontal plane mode.
[0150] Similarly, when the DIMD mode falls within a predetermined second range determined based on the vertical mode as the directional mode, the intra prediction mode of the current block can be inferred to be the vertical plane mode. Here, the predetermined range can represent a range from the value of the vertical mode minus M (modeV-M) to the value of the vertical mode plus M (modeV+M). For example, when the value of the vertical mode is 50, M is 5, and the value of the DIMD mode is 46, the value of the DIMD mode falls within the range of 45 to 55, and therefore, the intra prediction mode of the current block can be inferred to be the vertical plane mode.
[0151] When the value of the DIMD mode does not fall within any one of the first range and the second range, the intra prediction mode of the current block may be inferred to be a non-directional planar mode.
[0152] When the number of predefined directional patterns is 65, M used to determine the above range may be an integer greater than or equal to 0 and less than or equal to 16. Alternatively, based on the number of predefined directional patterns being K, M may be an integer greater than or equal to 0 and less than or equal to K / 4.
[0153] Alternatively, in Table 1 and Table 2, the availability flag indicating whether TIMD (template-based intra-mode derivation) is available can be further regarded as a signaling condition of the second plane flag. In this case, the horizontal or vertical plane mode can be derived based on the intra-prediction mode derived from TIMD (hereinafter, referred to as TIMD mode) without explicit signaling of the plane direction flag. In other words, based on the second plane flag indicating that the intra-prediction mode of the current block is a non-directional plane mode, the intra-prediction mode of the current block can be derived as a non-directional plane mode. On the other hand, based on the second plane flag indicating that the intra-prediction mode of the current block is not a non-directional plane mode, the intra-prediction mode of the current block can be derived as a horizontal or vertical plane mode based on the TIMD mode.
[0154] Specifically, the second plane flag may be adaptively signaled based on an availability flag indicating whether the ISP mode is applied to the current block and whether the TIMD is available. For example, when the ISP mode is not applied to the current block and the availability flag indicates that the TIMD is available, the second plane flag may be signaled, otherwise the second plane flag may not be signaled. Alternatively, when the availability flag indicates that the TIMD is available, the second plane flag may be signaled, and otherwise it may not be signaled, regardless of whether the ISP mode is applied to the current block. The availability flag may be signaled in at least one of the VPS, PPS, PH, or SH.
[0155] Hereinafter, a method for deriving a TIMD mode and a method for deriving an intra prediction mode of a current block based on the method are described.
[0156] The cost may be calculated for each of the horizontal plane mode and the vertical plane mode. Here, the cost may be calculated as the sum of absolute differences (SAD) between the predicted samples of the template area generated based on the horizontal / vertical plane mode and the pre-reconstructed samples of the template area. Alternatively, the cost may be calculated as the sum of absolute transform differences (SATD) between the predicted samples and the reconstructed samples of the template area. Here, SATD may mean the SAD transformed to the frequency domain. As an example of a transform, a Hadamard transform may be used, but is not limited thereto. The mode with the smallest cost among the horizontal and vertical plane modes may be derived as the intra-prediction mode of the current block as a TIMD mode.
[0157] Alternatively, the horizontal plane mode and the vertical plane mode may be reordered in ascending order of the calculated cost, and the encoded plane direction flag may be signaled based on the reordered order. For example, when the cost of the vertical plane mode is greater than the cost of the horizontal plane mode among the two modes, an index of 1 may be assigned to the vertical plane mode, and an index of 0 may be assigned to the horizontal plane mode, respectively. When the current block is coded in the vertical plane mode, the second plane flag may be signaled as 0, and the plane direction flag may be signaled as 1. In addition, because when the two modes are reordered by cost based on the template area, the mode with an index of 0 is more likely to be selected, CABAC-based entropy coding may be applied to the plane direction flag.
[0158] Based on the value of the TIMD mode being less than the value of the upper left diagonal mode, it is determined that the probability of the intra-frame prediction mode being the horizontal prediction mode is high, and the intra-frame prediction mode of the current block can be inferred as the horizontal plane mode. On the other hand, based on the value of the TIMD mode being greater than or equal to the value of the upper left diagonal mode, it is determined that the probability of the intra-frame prediction mode being the vertical prediction mode is high, and the intra-frame prediction mode of the current block can be inferred as the vertical plane mode. As an example, when the predefined directional mode is defined from the lower left diagonal mode of mode number 2 to the upper right diagonal mode of mode number 66, the upper left diagonal mode can correspond to mode number 34.
[0159] Based on the TIMD mode being a non-directional plane mode or a DC mode, the intra prediction mode of the current block can be inferred as a horizontal plane mode. In this case, since mode numbers 0 and 1 are assigned to the non-directional plane mode and the DC mode, respectively, the intra prediction mode can be inferred without additional conditions. Alternatively, based on the TIMD mode being a non-directional plane mode or a DC mode, the intra prediction mode of the current block can be inferred as a vertical plane mode. Typically, because the edge of an image may be in a vertical direction, improvements in prediction performance can be expected by using a vertical plane mode.
[0160] Table 5 below is an example of a signaling method for the second plane flag.
[0161] [Table 5]
[0162]
[0163] Referring to Table 5, planar_flag indicates whether the intra prediction mode of the current block is a non-directional plane mode, which may correspond to a second plane flag according to the present disclosure. For example, based on planar_flag being 1, this may indicate that the intra prediction mode of the current block is a non-directional plane mode. Based on planar_flag being 0, this may indicate that the intra prediction mode of the current block is a horizontal or vertical plane mode. intra_subpartitions_mode_flag may indicate whether the ISP mode is applied to the current block, and sps_timd_enabled_flag may indicate whether TIMD is available.
[0164] When the ISP mode is not applied to the current block (intra_subpartitions_mode_flag = 0) and TIMD is available (sps_timd_enabled_flag = 1), planar_flag can be obtained from the bitstream. Here, it is assumed that sps_timd_enabled_flag is signaled in the sequence parameter set, but it is not limited to this. Based on the value of planar_flag being 1, the intra-frame prediction mode of the current block can be derived as a non-directional planar mode. On the other hand, based on the value of planar_flag being 0, the intra-frame prediction mode of the current block can be derived as a horizontal or vertical plane mode based on the TIMD mode as described above.
[0165] Based on the ISP mode being applied to the current block or TIMD being unavailable, planar_flag may be derived as 1 without being obtained from the bitstream. In other words, the intra prediction mode of the current block may be derived as a non-directional planar mode.
[0166] In Table 5, planar_flag is signaled depending on intra_subpartitions_mode_flag and sps_timd_enabled_flag, but this is only an example. In other words, planar_flag may be signaled depending on sps_timd_enabled_flag regardless of intra_subpartitions_mode_flag.
[0167] Alternatively, in Table 1 and Table 2, the availability flag indicating the availability of TIMD may be further regarded as a signaling condition of the plane direction flag. When the availability flag indicates that TIMD is available, the plane direction flag may not be signaled, otherwise the plane direction flag may be signaled. When the plane direction flag is signaled, the intra-frame prediction mode of the current block may be derived as a horizontal or vertical plane mode depending on the value of the plane direction flag. On the other hand, when the plane direction flag is not signaled, the intra-frame prediction mode of the current block may be derived as a horizontal or vertical plane mode based on the TIMD mode as described above.
[0168] Table 6 below is an example of how to signal the plane direction flag.
[0169] [Table 6]
[0170]
[0171] Referring to Table 6, as seen in Table 5, planar_flag indicates whether the intra prediction mode of the current block is a non-directional plane mode, intra_subpartitions_mode_flag indicates whether the ISP mode is applied to the current block, and sps_timd_enabled_flag may indicate whether DIMD is available. planar_dir_flag indicates whether the intra prediction mode of the current block is a horizontal plane mode, which may correspond to a plane direction flag according to the present disclosure.
[0172] planar_flag can be obtained from the bitstream based on the fact that ISP mode is not applied to the current block (intra_subpartitions_mode_flag = 0). planar_flag can be obtained from the bitstream regardless of sps_timd_enabled_flag.
[0173] When the value of planar_flag is 1, the intra prediction mode of the current block can be derived as a non-directional plane mode. On the other hand, when the value of planar_flag is 0, planar_dir_flag can be adaptively signaled based on sps_timd_enabled_flag. When the value of planar_flag is 0 and sps_timd_enabled_flag is 0, planar_dir_flag can be obtained from the bitstream. When the value of planar_dir_flag is 1, the intra prediction mode of the current block can be derived as a horizontal plane mode, and when the value of planar_dir_flag is 0, the intra prediction mode of the current block can be derived as a vertical plane mode. When the value of planar_flag is 0 and sps_timd_enabled_flag is 1, planar_dir_flag cannot be obtained from the bitstream. In this case, based on the above-mentioned TIMD mode, the intra prediction mode of the current block can be derived as a horizontal or vertical plane mode.
[0174] In Table 6, planar_flag is signaled depending on intra_subpartitions_mode_flag, but this is only an example. In other words, planar_flag may be signaled regardless of intra_subpartitions_mode_flag.
[0175] Alternatively, the horizontal / vertical plane mode may be used as part of the TIMD candidate mode without signaling a separate second plane flag or plane direction flag. The TIMD candidate mode may include at least one of an intra-frame prediction mode of a neighboring block, a candidate mode belonging to an MPM list, a vertical mode, a horizontal mode, or a DC mode. In addition, at least one of the horizontal plane mode or the vertical plane mode may be included in the TIMD candidate mode. In other words, when the horizontal or vertical plane mode among the TIMD candidate modes of the current block has the lowest cost, the mode may be set as the intra-frame prediction mode of the current block. However, this mode may be used when the flag (TIMD flag) indicating whether TIMD is applied to the current block is 1. In this way, when the horizontal / vertical plane mode is used as one of the TIMD candidate modes, coding efficiency can be improved because there is no other flag signaling except the TIMD flag.
[0176] Based on the horizontal / vertical plane mode being used as one of the TIMD candidate modes, the horizontal / vertical plane mode can be added as a TIMD candidate mode without a separate condition. Alternatively, when all intra-frame prediction modes of the neighboring blocks are not directional modes (i.e., plane mode or DC mode), the horizontal / vertical plane mode can be added as a TIMD candidate mode. When all intra-frame prediction modes of the neighboring blocks are non-directional modes, it is unlikely that texture or object boundaries, etc. are included in the current block. By adding the horizontal / vertical plane mode as a candidate mode to such a block, the prediction performance can be improved. Alternatively, when even one of the intra-frame prediction modes of the neighboring blocks is not a non-directional mode, the horizontal / vertical plane mode can also be added as a TIMD candidate mode. Because the horizontal or vertical plane mode mainly uses the reference samples on the left or top to perform prediction, it has the characteristics of each mode, and at the same time, it has the characteristics of the plane mode predicted by reflecting the distance of the reference sample, so that the prediction performance can be improved. Here, the neighboring block may include at least one of the left neighboring block, the upper neighboring block, the lower left neighboring block, the upper right neighboring block, or the upper left neighboring block.
[0177] Among the TIMD candidate modes, the first two modes (i.e., the mode with the lowest cost and the mode with the second lowest cost) can be selected in ascending order of cost. Based on the selected mode being a horizontal / vertical plane mode, the mode can be used as an intra-frame prediction mode for the current block. In other words, a prediction block can be generated for each of the two modes, and a final prediction block can be generated by mixing them. Alternatively, based on the mode with the lowest cost being a horizontal or vertical plane mode, the mode can be used alone without mixing with the above-mentioned other prediction blocks. Alternatively, based on the mode with the lowest cost being a horizontal or vertical plane mode, it can be determined whether to perform the above-mentioned mixing by comparing it with the cost of the mode with the second lowest cost. For example, based on the second low cost being greater than 1.5 times the lowest cost, mixing may not be performed, and the horizontal or vertical plane mode as the mode with the lowest cost may be used alone. Alternatively, based on the mode with the second low cost being a horizontal or vertical plane mode, the mode with the lowest cost may be used alone without performing mixing with the modes with the lowest cost and the second low cost. Alternatively, based on the mode with the second low cost being a horizontal or vertical plane mode, it can be determined whether to perform the above-mentioned mixing by comparing with the cost of the mode with the second low cost. For example, based on the second low cost being less than 1.2 times the lowest cost, mixing may not be performed. This method can improve the prediction performance by maintaining the characteristics of the horizontal / vertical plane mode as much as possible. For intra-frame slices, the above-mentioned TIMD may be limited by the block size. For example, based on the current slice being an intra-frame slice, TIMD is performed only when the number of samples in the block is 1024 or less. As described above, when utilizing the horizontal / vertical plane mode based on TIMD, the block size limitation of TIMD can be followed only for intra-frame slices. Alternatively, when utilizing the horizontal / vertical plane mode based on TIMD, the block size limitation of TIMD may not be followed for intra-frame slices. By using the horizontal / vertical plane mode based on the TIMD mode for all blocks, the performance of intra-frame prediction can be improved.
[0178] The plane mode may not be defined as a mode independent of the candidate mode of the MPM list. In this case, the plane mode may be added as a candidate mode of the MPM list, and a method of constructing the MPM list will be described in detail below.
[0179] The MPM list may include a plurality of candidate modes. One or more modes may be derived based on at least one of the following methods (1) to (6), and a plurality of candidate modes may be derived based on the mode.
[0180] (1) Intra-frame prediction mode of neighboring blocks adjacent to the current block
[0181] (2) Intra-frame prediction mode (IPM mode) derived from adjacent inter-frame modes
[0182] (3) DIMD mode
[0183] (4) TIMD mode
[0184] (5) Export mode
[0185] (6) Default mode
[0186] The neighboring blocks may include at least one of a left neighboring block, an upper neighboring block, a lower left neighboring block, an upper right neighboring block, or an upper left neighboring block. Depending on the size of the current block, the order in which the intra-frame prediction modes of the neighboring blocks are added to the MPM list may be different. For example, when the height of the current block is greater than or equal to the width, the intra-frame prediction mode of the upper neighboring block may be added before the intra-frame prediction mode of the left neighboring block.
[0187] Based on the neighboring block being coded in inter mode rather than in non-intra mode, the intra prediction mode can be derived from the IPM buffer. When the position pointed to by the motion vector of the neighboring block coded in inter mode is an intra mode, the intra prediction mode of the corresponding position can be stored in the IPM buffer. The intra prediction mode stored in the IPM buffer can be used as a candidate mode for the current block.
[0188] The intra prediction mode derived based on the above DIMD may be used as a candidate mode. Based on the current block not being in DIMD mode, the intra prediction mode derived based on DIMD may be used as a candidate mode.
[0189] The intra prediction mode derived based on the above TIMD may be used as a candidate mode. Based on the current block not being in TIMD mode, the intra prediction mode derived based on TIMD may be used as a candidate mode.
[0190] The derived mode may mean a mode derived by adding or subtracting a predetermined constant value from a mode derived based on at least one of the above methods (1) to (4). Here, the constant value may be an integer of 1, 2, 3, 4 or more.
[0191] The default mode may be a mode defined identically for the encoding device and the decoding device, and may include at least one of a horizontal mode, a vertical mode, or a mode derived by adding or subtracting a predetermined constant value from the horizontal / vertical mode. Here, the constant value may be defined as a multiple of 4, such as 4, 8, or 12.
[0192] When the mode derived based on at least one of the above methods (1) to (6) is a horizontal or vertical plane mode, the mode may be added to the MPM list. In this case, the diversity of intra prediction modes can be increased without additional signaling, and thus improvement in prediction performance can be expected. Alternatively, when the mode derived based on at least one of the above methods (1) to (6) is a horizontal or vertical plane mode, a non-directional plane mode may be added to the MPM list instead of the horizontal or vertical plane mode. In this case, because the conditions for inserting the MPM list due to the increase in candidate modes are not increased, complexity can be reduced. Alternatively, when the intra prediction mode derived from the adjacent inter-frame mode is a horizontal or vertical plane mode, the mode may be added to the MPM list. In this case, the diversity of intra prediction modes can be increased, and thus improvement in prediction performance can be expected. Alternatively, when the intra prediction mode derived from the adjacent inter-frame mode is a horizontal or vertical plane mode, a non-directional plane mode may be added to the MPM list instead of the horizontal or vertical plane mode. In this way, by using the non-directional planar mode with the highest selection frequency, an improvement in prediction performance can be expected.
[0193] The MPM list according to the present disclosure may be configured as a main MPM list and a secondary MPM list, respectively. The candidate patterns of the main / secondary MPM list are derived based on at least one of the above methods (1) to (6), but the secondary MPM list may be configured with a pattern that is different from the candidate pattern of the main MPM list. The secondary MPM list may be configured with M candidate patterns, where M may be an integer of 16 or greater. When the secondary MPM list is not filled, a pattern derived by adding or subtracting a value of N (N=1, 2, 3, 4) to a candidate pattern with a candidate index of 0 in the main MPM list may be inserted first. In addition, a pattern derived by adding or subtracting a value of N to a candidate pattern with a candidate index of 1 in the main MPM list may be inserted, and a pattern derived by adding or subtracting a value of K (K=1, 2, 3) to a candidate pattern with a candidate index of 2 in the main MPM list may be inserted. However, when the secondary MPM list is not filled, a pattern that is not inserted in the main / secondary MPM list may be inserted in the default pattern.
[0194] refer to Figure 4 , a prediction block of the current block may be generated based on the intra prediction mode of the current block S410.
[0195] Reference samples may be derived based on the intra prediction mode of the current block, and a prediction block of the current block may be generated based on the derived reference samples.
[0196] The filtered neighboring samples may be derived as reference samples, or the unfiltered neighboring samples may be derived as reference samples. As an example, when the reference samples are derived by performing filtering on the neighboring samples, the filtered neighboring samples may be derived as follows.
[0197] p[ -1 ][ -1 ] = ( refUnfilt[ -1 ][ 0 ] + 2 * refUnfilt[ -1 ][ -1 ] +refUnfilt[ 0 ][ -1 ] + 2 ) >> 2
[0198] For y = 0..refH – 2, p[ -1 ][ y ] = ( refUnfilt[ -1 ][ y + 1 ] + 2 *refUnfilt[ -1 ][ y ] + refUnfilt[ -1 ][ y - 1 ] + 2 ) >> 2
[0199] p[ -1 ][ refH - 1 ] = refUnfilt[ -1 ][ refH - 1 ]
[0200] For x = 0..refW – 2, p[ x ][ -1 ] = ( refUnfilt[ x - 1 ][ -1 ] + 2 *refUnfilt[ x ][ -1 ] + refUnfilt[ x + 1 ][ -1 ] + 2 ) >> 2
[0201] p[ refW - 1 ][ -1 ] = refUnfilt[ refW - 1 ][ -1 ]
[0202] In the above formula, refUnfilt represents the unfiltered neighborhood sample, and [x][y] represents the x, y coordinates of the sample. This may represent the coordinates when the coordinates of the upper left sample in the current block are (0, 0). refH and refW may represent the height and width of the reference region for intra prediction, respectively.
[0203] When some or all of the following specific conditions are met, filtering of neighboring samples may be performed, otherwise it may not be performed.
[0204] – nTbW * nTbH is greater than 32 (the product of the width and height of the current block is greater than 32)
[0205] – cIdx is equal to 0 (the component type of the current block is the luminance component)
[0206] – IntraSubPartitionsSplitType is equal to ISP_NO_SPLIT (ISP mode is not applied to the current block)
[0207] – One or more of the following conditions are true:
[0208] – predModeIntra is equal to INTRA_PLANAR (the intra prediction mode of the current block is non-directional planar mode)
[0209] – predModeIntra is equal to INTRA_ANGULAR34 (the intra prediction mode of the current block is the upper left diagonal mode)
[0210] – predModeIntra is equal to INTRA_ANGULAR2 and nTbH is greater than or equal to nTbW (the intra prediction mode of the current block is the bottom-left diagonal mode, and the width of the current block is greater than or equal to the height.)
[0211] – predModeIntra is equal to INTRA_ANGULAR66 and nTbW is greater than or equal to nTbH (the intra prediction mode of the current block is the top-right diagonal mode, and the width of the current block is greater than or equal to the height.)
[0212] In the horizontal plane mode, a prediction block may be generated based on a left neighboring sample row of the current block and a right upper neighboring sample row of the current block. As an example, a prediction block according to the horizontal plane mode may be generated as described in Equation 1 or Equation 2 below.
[0213] [Equation 1]
[0214]
[0215] [Equation 2]
[0216]
[0217] In Equation 1, predH(x, y) may mean an intermediate prediction sample of the (x, y) coordinate. W and H may mean the width and height of the current block, respectively. rec(-1, y) may mean a left neighboring sample of the current block, and rec(W, -1) may mean an upper right neighboring sample of the current block. Planar Hor (x, y) may mean the final predicted sample of the (x, y) coordinates. This may also apply to Equation 2.
[0218] In the vertical plane mode, a prediction block may be generated based on sample rows around the top of the current block and sample rows around the bottom left of the current block. As an example, a prediction block according to the vertical plane mode may be generated as described in Equation 3 or Equation 4 below.
[0219] [Equation 3]
[0220]
[0221] [Equation 4]
[0222]
[0223] In Equation 3, predV(x, y) may mean the middle prediction sample of the (x, y) coordinate. W and H may mean the width and height of the current block, respectively. rec(x, -1) may mean the upper neighboring sample of the current block, and rec(-1, H) may mean the lower left neighboring sample of the current block. Planar Ver (x, y) may mean the final predicted sample of the (x, y) coordinates. This may also apply to Equation 4.
[0224] When the intra prediction mode of the current block belongs to the directional plane mode, the reference sample for the current block may be derived based on the reference sample for the directional mode which is the horizontal mode or the vertical mode.
[0225] As described in the above equation, in the case of the horizontal plane mode, similar to the horizontal mode which is a directional mode, the left neighboring sample can be mainly used as a reference sample, and in the case of the vertical plane mode similar to the vertical mode which is a directional mode, the upper neighboring sample can be mainly used as a reference sample.
[0226] The horizontal or vertical plane mode according to the present disclosure may have similar characteristics to the horizontal mode or the vertical mode due to the characteristics of mainly using the left or upper neighboring samples. Therefore, the block encoded in the horizontal plane mode can use the same reference sample derivation method used in the block encoded in the horizontal mode. Similarly, the block encoded in the vertical plane mode can use the same reference sample derivation method used in the block encoded in the vertical mode.
[0227] Alternatively, when the intra prediction mode of the current block belongs to the directional plane mode, the reference samples for the current block may be derived based on the reference samples for the non-directional plane mode.
[0228] Because the horizontal / vertical plane mode performs prediction by applying a weight based on the distance between the reference sample and the prediction sample, it can have similar characteristics to the non-directional plane mode. Therefore, a block encoded in the horizontal or vertical plane mode can use the same reference sample derivation method used in the block encoded in the non-directional plane mode.
[0229] Additionally, based on whether the intra prediction mode of the current block is a horizontal or vertical plane mode, position-dependent intra prediction (PDPC) may be applied to the prediction block. This may be a filtering process used to mitigate discontinuities between the prediction block and reconstructed neighboring samples.
[0230] As an example, since the horizontal / vertical plane mode has similar characteristics to the non-directional plane mode, the PDPC method applied to the non-directional plane mode can be applied in the same / similar manner. Based on the PDPC method applied to the non-directional plane mode being applied, the prediction block or prediction sample generated based on the horizontal or vertical plane mode can be modified as described in the following equation 5.
[0231] [Equation 5]
[0232]
[0233] In Equation 5, floorLog2(a) and min(a, b) may be defined as described in Equation 6 below.
[0234] [Equation 6]
[0235]
[0236] In Equation 5, W and H represent the width and height of the current block, respectively, and pred(x, y) may represent a predicted sample of the (x, y) coordinate in the predicted block. Here, the (x, y) coordinate means the coordinate when the coordinate of the upper left sample of the current block is (0, 0). rec(-1, y) and rec(x, -1) represent left neighboring samples and upper neighboring samples, respectively. The left / upper neighboring samples may be the filtered neighboring samples described above, or may be unfiltered neighboring samples. In Equation 6, n may be an integer less than or equal to log2a and greater than (log2a-1).
[0237] Alternatively, in the case of the horizontal plane mode, since prediction is mainly performed using the left neighboring samples, the PDPC method for the horizontal mode having similar characteristics can be applied in the same / similar manner. Similarly, in the case of the vertical plane mode, since prediction is mainly performed using the upper neighboring samples, the PDPC method for the vertical mode having similar characteristics can be applied in the same / similar manner. When the PDPC method applied to the vertical mode is applied, the prediction block or prediction sample generated based on the vertical plane mode can be modified as described in Equation 7 below.
[0238] [Equation 7]
[0239]
[0240] In Equation 7, Clip1(a) may be defined as described in Equation 8 below.
[0241] [Equation 8]
[0242]
[0243] In Equation 8, BitDepth represents the bit depth of the current sample, and rec(-1, -1) may represent the upper left neighboring sample.
[0244] In addition, when the PDPC method applied to the horizontal mode is applied, the prediction block or prediction sample generated based on the horizontal plane mode may be modified as described in Equation 9 below.
[0245] [Equation 9]
[0246]
[0247] Clip1(a) in Equation 9 is defined as in Equation 8.
[0248] refer to Figure 4 , a residual block of the current block may be obtained by performing at least one of dequantization or inverse transformation on the transformation coefficients of the current block S420.
[0249] The transform coefficients of the current block may be derived by decoding the residual information signaled from the bitstream.
[0250] A transform kernel for inverse transform may be determined based on at least one of a size of a current block or an intra prediction mode.
[0251] As described above, the horizontal plane mode may use the left adjacent sample column and the upper right adjacent sample, and the vertical plane mode may use the upper adjacent sample row and the lower left adjacent sample. In other words, the residual characteristics of the block encoded in the horizontal plane mode may be similar to the residual of the block encoded in the horizontal mode or the residual of the block encoded in the vertical mode. Similarly, the residual characteristics of the block encoded in the vertical plane mode may be similar to the residual of the block encoded in the vertical mode or the residual of the block encoded in the horizontal mode. Depending on the similarity of the residual characteristics, the transform kernel of the block encoded in the horizontal / vertical plane mode may be determined to be the same as the transform kernel of the block encoded in the horizontal / vertical mode, which is a directional mode.
[0252] As an example, the residual signal of a block encoded in a horizontal plane mode can be regarded as the residual signal of a block encoded in a horizontal mode, and the transform kernel used in the block encoded in the horizontal mode can be used in the same manner. The residual signal of a block encoded in a vertical plane mode can be regarded as the residual signal of a block encoded in a vertical mode, and the transform kernel used in the block encoded in the vertical mode can be used in the same manner. Alternatively, the residual signal of a block encoded in a horizontal plane mode can be regarded as the residual signal of a block encoded in a vertical mode, and the transform kernel used in the block encoded in the vertical mode can be used in the same manner. The residual signal of a block encoded in a vertical plane mode can be regarded as the residual signal of a block encoded in a horizontal mode, and the transform kernel used in the block encoded in the horizontal mode can be used in the same manner. By determining the transform kernel based on the correlation of these residual characteristics, the performance of the transform is improved, and thus better energy compression can be expected.
[0253] Alternatively, since the horizontal / vertical plane mode performs prediction by considering the distance between the reference sample and the prediction sample, it may have similar characteristics to the residual of the non-directional plane mode. Therefore, the residual signal of the block encoded in the horizontal or vertical plane mode can be regarded as the residual signal of the block encoded in the non-directional plane mode, and the transform kernel used in the block encoded in the non-directional plane mode can be used in the same way. By determining the transform kernel based on the correlation of these residual characteristics, the performance of the transform is improved, and thus better energy compression can be expected.
[0254] refer to Figure 4 , the current block may be reconstructed based on the prediction block and the residual block of the current block (S430).
[0255] The reconstructed block may be generated by adding the prediction block and the residual block of the current block. Here, the prediction block may be a prediction block to which PDPC is not applied, or a prediction block modified by PDPC.
[0256] Figure 5A schematic configuration of a decoding device 300 that performs an image decoding method according to the present disclosure is shown.
[0257] refer to Figure 5 According to the present disclosure, the decoding device 300 may include an intra-frame prediction mode deriver 510, a prediction block generator 520, a residual block generator 530, and a reconstructed block generator 540. The intra-frame prediction mode deriver 510 and the prediction block generator 520 may be configured in Figure 3 In the intra-frame predictor 331, the residual block generator 530 may be configured in Figure 3 The residual processor 320, and the reconstructed block generator 540 can be configured in Figure 3 In the adder 340.
[0258] The intra prediction mode deriver 510 may perform the same method as that of deriving the intra prediction mode of the current block according to step S400. In other words, the intra prediction mode deriver 510 may derive the intra prediction mode of the current block based on the intra prediction mode information, and a detailed description will be omitted here.
[0259] The prediction block generator 520 can perform the prediction block generation method according to step S410 in the same manner. In other words, the prediction block of the current block can be generated based on the intra prediction mode of the current block. In this case, the prediction block generator 520 can derive reference samples for intra prediction based on neighboring samples of the current block, and can also derive reference samples by applying filtering to neighboring samples under certain conditions. In addition, the prediction block generator 520 can also generate a modified prediction block by applying PDPC to the prediction block.
[0260] The residual block generator 530 can perform the residual block generation method according to step S420 in the same manner. In other words, the residual block of the current block can be obtained by performing at least one of inverse quantization or inverse transformation on the transformation coefficients of the current block. In this case, as in Figure 4 As seen in , the transform kernel used for inverse transform can be determined based on at least one of the size of the current block or the intra prediction mode.
[0261] The reconstructed block generator 540 may reconstruct the current block based on the prediction block and the residual block of the current block.
[0262] Figure 6 An image encoding method performed by the encoding device 200 as an embodiment according to the present disclosure is shown.
[0263] refer to Figure 6 , a prediction block of the current block may be generated S600.
[0264] A prediction block of a current block may be generated based on a predetermined intra prediction mode. The predetermined intra prediction mode may be one of the predefined intra prediction modes. The predefined intra prediction mode may include at least one of a non-directional mode or a directional mode. The non-directional mode may include at least one of a plane mode or a DC mode. The plane mode according to the present disclosure may include a non-directional plane mode and a directional plane mode, and the directional plane mode may include at least one of a horizontal plane mode or a vertical plane mode. Alternatively, the directional plane mode may be defined as a mode independent of the non-directional plane mode, in which case the plane mode according to the present disclosure may mean a non-directional plane mode. The directional mode may mean a mode with a predetermined angle, such as a horizontal mode, a vertical mode, a diagonal mode, etc.
[0265] Based on the intra prediction mode used to generate the prediction block of the current block, the intra prediction mode information may be encoded. The encoded intra prediction mode information may be inserted into the bitstream and sent using a signal. The intra prediction mode information may include at least one of an MPM flag, a first plane flag, a second plane flag, a plane direction flag, an MPM index, or residual mode information.
[0266] The planar mode according to the present disclosure may be signaled as a mode independent of the candidate mode of the MPM list. The method of signaling or deriving intra prediction mode information for this purpose is similar to that of referring to Figure 4 Same as described.
[0267] Alternatively, the mode according to the present disclosure may not be signaled as a mode independent of the candidate mode of the MPM list. In other words, the planar mode may be added as a candidate mode of the MPM list, which is consistent with the reference Figure 4 Same as described.
[0268] The prediction block of the current block can be generated based on the predetermined reference sample. Here, the predetermined reference sample can be a pre-reconstructed neighboring sample or a filtered neighboring sample adjacent to the current block, such as by referring to Figure 4 The method for obtaining the neighboring samples for filtering and whether to perform filtering by referring to Figure 4 Same as described.
[0269] In addition, PDPC may be applied to the prediction block of the current block. In particular, when the intra prediction mode of the current block is a horizontal or vertical plane mode, PDPC may be applied to the prediction block. In this case, with respect to the horizontal or vertical plane mode, considering the characteristics similar to the non-directional plane mode, the PDPC method applied to the non-directional plane mode may be applied. Alternatively, with respect to the horizontal plane mode (or vertical plane mode), considering the characteristics similar to the horizontal mode (or vertical mode), the PDPC method applied to the horizontal mode (or vertical mode) may be applied.
[0270] refer to Figure 6 , a residual block of the current block may be derived based on the prediction block of the current block (S610). Here, the residual block of the current block may be derived by subtracting the prediction block from the original block of the current block (S620).
[0271] refer to Figure 6 , a transform coefficient of the current block may be derived by performing at least one of transform or quantization on the residual block of the current block (S620).
[0272] A transform kernel used for transform may be determined based on at least one of a size of a current block or an intra prediction mode.
[0273] As an example, the residual signal of a block encoded in a horizontal plane mode can be regarded as the residual signal of a block encoded in a horizontal mode, and the transform kernel used in the block encoded in the horizontal mode can be used in the same manner. The residual signal of a block encoded in a vertical plane mode can be regarded as the residual signal of a block encoded in a vertical mode, and the transform kernel used in the block encoded in the vertical mode can be used in the same manner. Alternatively, the residual signal of a block encoded in a horizontal plane mode can be regarded as the residual signal of a block encoded in a vertical mode, and the transform kernel used in the block encoded in the vertical mode can be used in the same manner. The residual signal of a block encoded in a vertical plane mode can be regarded as the residual signal of a block encoded in a horizontal mode, and the transform kernel used in the block encoded in the horizontal mode can be used in the same manner. By determining the transform kernel based on the correlation of these residual characteristics, the performance of the transform can be improved, and better energy compression can be expected.
[0274] Alternatively, since the horizontal / vertical plane mode performs prediction by considering the distance between the reference sample and the prediction sample, it may have similar characteristics to the residual of the non-directional plane mode. Therefore, the residual signal of the block encoded in the horizontal or vertical plane mode can be regarded as the residual signal of the block encoded in the non-directional plane mode, and the transform kernel used in the block encoded in the non-directional plane mode can be used in the same way. By determining the transform kernel based on the correlation of these residual characteristics, the performance of the transform can be improved, and better energy compression can be expected.
[0275] refer to Figure 6 , the transform coefficients of the current block may be encoded S630.
[0276] Figure 7 A schematic configuration of an encoding device 200 that performs an image encoding method according to the present disclosure is shown.
[0277] refer to Figure 7 According to the present disclosure, the encoding device 200 may include a prediction block generator 710, a residual block generator 720, a transform coefficient deriver 730, and a transform coefficient encoder 740. The prediction block generator 710 may be configured in Figure 2 The inter-frame predictor 221, and the residual block generator 720 and the transform coefficient deriver 730 may be configured in Figure 2 The transform coefficient encoder 740 may be configured in Figure 2 in the entropy encoder 240.
[0278] The prediction block generator 710 may generate a prediction block of the current block based on a predetermined intra prediction mode, and may generate intra prediction mode information based on the intra prediction mode used to generate the prediction block of the current block. Figure 6 The generated intra prediction mode information can be sent to Figure 2 The entropy encoder 240 is encoded.
[0279] The residual block generator 720 may generate a residual block through a difference between an original block and a predicted block of a current block.
[0280] The transform coefficient deriver 730 may derive the transform coefficient of the current block by performing at least one of transform or quantization on the residual block of the current block. The transform kernel for transform may be determined based on at least one of the size of the current block or the intra prediction mode, such as by referring to Figure 6 described.
[0281] The transform coefficient encoder 740 may encode the transform coefficient of the current block.
[0282] In the above embodiments, the method is described as a series of steps or boxes based on the flowchart, but the corresponding embodiments are not limited to the order of the steps, and some steps may occur simultaneously or in a different order than other steps described above. In addition, those skilled in the art will appreciate that the steps shown in the flowchart are not exclusive, and other steps may be included or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of the present disclosure.
[0283] The above-mentioned method according to the embodiment of the present disclosure can be implemented in the form of software, and the encoding device and / or decoding device according to the present disclosure can be included in a device that performs image processing, such as a TV, a computer, a smart phone, a set-top box, a display device, etc.
[0284] In the present disclosure, when the embodiment is implemented as software, the above method can be implemented as a module (process, function, etc.) that performs the above functions. The module can be stored in a memory and can be executed by a processor. The memory can be located inside or outside the processor and can be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), another chipset, a logic circuit, and / or a data processing device. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. In other words, the embodiments described herein can be implemented on a processor, a microprocessor, a controller, or a chip. For example, the functional unit shown in each of the figures can be implemented on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information (e.g., information about instructions) or an algorithm for implementation can be stored in a digital storage medium.
[0285] In addition, the decoding device and the encoding device of the embodiment of the present disclosure can be included in multimedia broadcast sending and receiving devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, devices for providing video on demand (VoD) services, over-the-top video (OTT) devices, devices for providing Internet streaming services, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, videophone video devices, transportation terminals (e.g., vehicle (including autonomous driving vehicle) terminals, aircraft terminals, ship terminals, etc.) and medical video devices, etc., and can be used to process video signals or data signals. For example, over-the-top video (OTT) devices may include game consoles, Blu-ray players, networked TVs, home theater systems, smart phones, tablet computers, digital video recorders (DVRs), etc.
[0286] In addition, the processing method of the embodiment of the present disclosure can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to an embodiment of the present disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media may include, for example, Blu-ray discs (BD), universal serial buses (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, tapes, floppy disks, and optical media storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmitted via the Internet). In addition, the bit stream generated by the encoding method may be stored in a computer-readable recording medium or may be sent via a wired or wireless communication network.
[0287] In addition, the embodiments of the present disclosure may be implemented by a computer program product through a program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0288] Figure 8 An example of a content streaming system to which an embodiment of the present disclosure can be applied is shown.
[0289] refer to Figure 8 The content streaming transmission system to which the embodiments of the present disclosure are applied may mainly include an encoding server, a streaming transmission server, a web server, a media storage, a user device, and a multimedia input device.
[0290] The encoding server generates a bitstream by compressing content input from a multimedia input device such as a smartphone, a camera, a camcorder, etc. into digital data and transmits it to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, a camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0291] A bitstream may be generated by applying the encoding method or the bitstream generating method of the embodiment of the present disclosure, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0292] The streaming server sends multimedia data to the user device through the web server based on the user's request, and the web server serves as a medium to inform the user what services are available. When the user requests the required service from the web server, the web server delivers it to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system may include a separate control server, and in this case, the control server controls the command / response between each device in the content streaming system.
[0293] The streaming server may receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content may be received in real time. In this case, in order to provide a smooth streaming service, the streaming server may store the bitstream for a certain period of time.
[0294] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays (HMDs), digital televisions, desktops, digital signage, etc.).
[0295] Each server in the content streaming system may be operated as a distributed server, and in this case, data received from each server may be distributed and processed.
[0296] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined and implemented as a device, and the technical features of the device claims of the present disclosure can be combined and implemented as a method. In addition, the technical features of the method claims of the present disclosure and the technical features of the device claims can be combined and implemented as a device, and the technical features of the method claims of the present disclosure and the technical features of the device claims can be combined and implemented as a method.
Claims
1. An image decoding method, comprising: deriving an intra prediction mode of the current block from predefined intra prediction modes, wherein the predefined intra prediction modes include a non-directional plane mode, a directional plane mode, a horizontal mode, and a vertical mode, and the directional plane mode includes at least one of a horizontal plane mode or a vertical plane mode; generating a prediction block of the current block based on the intra prediction mode; Obtaining a residual block of the current block by performing at least one of dequantization or inverse transformation on a transformation coefficient of the current block; and The current block is reconstructed based on the prediction block and the residual block of the current block.
2. The image decoding method according to claim 1, wherein: When the intra prediction mode of the current block belongs to the directional plane mode, a transform kernel for the inverse transform is determined based on a transform kernel for a predefined mode.
3. The image decoding method according to claim 2, wherein: When the intra prediction mode of the current block is the horizontal plane mode, determining the transform kernel for the inverse transform based on the transform kernel for the vertical mode, and Wherein, when the intra prediction mode of the current block is the vertical plane mode, the transform kernel used for the inverse transform is determined based on the transform kernel used for the horizontal mode.
4. The image decoding method according to claim 2, wherein: When the intra prediction mode of the current block is the horizontal plane mode, determining the transform kernel for the inverse transform based on the transform kernel for the horizontal mode, and Wherein, when the intra prediction mode of the current block is the vertical plane mode, the transform kernel used for the inverse transform is determined based on the transform kernel used for the vertical mode.
5. The image decoding method according to claim 2, wherein: When the intra prediction mode of the current block belongs to the directional plane mode, the transform kernel used for the inverse transform is determined based on a transform kernel used for the non-directional plane mode.
6. The image decoding method according to claim 1, wherein: deriving the intra prediction mode of the current block based on intra prediction mode information, The intra-frame prediction mode information includes at least one of a plane flag or a plane direction flag, and The plane flag indicates whether the intra-frame prediction mode of the current block is the non-directional plane mode or belongs to the directional plane mode, and the plane direction flag indicates whether the intra-frame prediction mode of the current block is the horizontal plane mode.
7. The image decoding method according to claim 6, wherein: Based on the plane flag indicating that the intra prediction mode of the current block does not belong to the directional plane mode, an MPM flag indicating whether the intra prediction mode of the current block is derived from an MPM list is signaled.
8. The image decoding method according to claim 6, wherein: At least one of the plane flag or the plane direction flag is adaptively signaled based on an availability flag indicating whether decoder-side intra mode derivation (DIMD) is available.
9. The image decoding method according to claim 6, wherein: At least one of the plane flag or the plane direction flag is adaptively signaled based on an availability flag indicating whether template based intra mode derivation (TIMD) is available.
10. The image decoding method according to claim 1, wherein: When the intra prediction mode of the current block belongs to the directional plane mode, a reference sample for the current block is derived based on a reference sample for the horizontal mode or the vertical mode.
11. The image decoding method according to claim 1, wherein: When the intra prediction mode of the current block belongs to the directional plane mode, a reference sample for the current block is derived based on a reference sample for the non-directional plane mode.
12. An image encoding method, comprising: generating a prediction block of the current block based on one of the predefined intra prediction modes, wherein the predefined intra prediction modes include a non-directional plane mode, a directional plane mode, a horizontal mode, and a vertical mode, and the directional plane mode includes at least one of a horizontal plane mode or a vertical plane mode; deriving a residual block of the current block based on the prediction block of the current block; deriving a transform coefficient of the current block by performing at least one of transform or quantization on the residual block; and The transform coefficients of the current block are encoded.
13. A computer-readable storage medium storing a bit stream generated by the image encoding method according to claim 12.
14. A method for sending data, comprising: Obtaining a bitstream for image information, wherein the bitstream is generated by: generating a prediction block of a current block based on one of predefined intra prediction modes, deriving a residual block of the current block based on the prediction block of the current block, deriving transform coefficients of the current block by performing at least one of transform or quantization on the residual block, and encoding the transform coefficients of the current block; and sending said data comprising said bit stream, The predefined intra prediction modes include a non-directional plane mode, a directional plane mode, a horizontal mode and a vertical mode, and the directional plane mode includes at least one of a horizontal plane mode and a vertical plane mode.