Image encoding / decoding method and apparatus, and recording medium storing a bit stream
By inducing intra prediction modes using DIMD and template matching, the method optimizes video compression by selecting between normal, horizontal, and vertical planar modes, enhancing prediction accuracy and reducing overhead in video encoding/decoding processes.
Patent Information
- Application Number
- JP2025504241
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-08
- Filing Date
- 2023-07-27
- Publication Date
- 2025-08-01
Smart Images

Figure 2025524954000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video encoding / decoding method and apparatus, and a recording medium storing a bitstream.
Background Art
[0002] Recently, the demand for high-resolution and high-quality videos such as HD (High Definition) videos and UHD (Ultra High Definition) videos has been increasing in various application fields. Along with this, highly efficient video compression technologies have been discussed.
[0003] As video compression technologies, there are various technologies such as an inter prediction technology that predicts pixel values included in a current picture from pictures before or after the current picture, an intra prediction technology that predicts pixel values included in the current picture using pixel information within the current picture, and an entropy coding technology that assigns short codes to values with high occurrence frequencies and long codes to values with low occurrence frequencies. By using such video compression technologies, video data can be effectively compressed for transmission or storage.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present disclosure aims to provide a method and apparatus for defining various planar modes and performing prediction based on them.
[0005] The present disclosure aims to provide a method and apparatus for inducing a planar mode using DIMD (decoder side intra mode derivation).
[0006] The present disclosure aims to provide a method and apparatus for inducing a planar mode based on template matching.
Means for Solving the Problems
[0007] The video decoding method and apparatus according to the present disclosure can induce an intra prediction mode of a current block, configure reference samples of the current block, and generate a prediction sample of the current block based on the intra prediction mode and the reference samples.
[0008] In the video decoding method and apparatus according to the present disclosure, the intra prediction mode can be induced from a plurality of predefined intra prediction modes.
[0009] In the video decoding method and apparatus according to the present disclosure, the plurality of predefined intra prediction modes can include a normal planar mode, a horizontal planar mode, and a vertical planar mode.
[0010] In the video decoding method and apparatus according to the present disclosure, the horizontal planar mode indicates a mode of generating a prediction sample of a current sample by using a horizontal sample of the current sample and a sample at the upper right end within the current block, the horizontal sample indicates a sample that is horizontally adjacent to the current sample among the reference samples, and the sample at the upper right end can indicate a sample that is adjacent to the corner at the upper right end of the current block among the reference samples.
[0011] In the video decoding method and apparatus according to the present disclosure, the vertical planar mode indicates a mode of generating a prediction sample of a current sample by using a vertical sample of the current sample and a sample at the lower left end within the current block, the vertical sample indicates a sample that is vertically adjacent to the current sample among the reference samples, and the sample at the lower left end can indicate a sample that is adjacent to the corner at the lower left end of the current block among the reference samples.
[0012] The video decoding method and apparatus according to the present disclosure can obtain a first flag indicating whether the normal planar mode is used for the current block.
[0013] In the video decoding method and apparatus according to the present disclosure, when the first flag indicates that the normal planner mode is not used for the current block, a second flag indicating an intra prediction mode used for the current block can be obtained from among the horizontal planner mode or the vertical planner mode.
[0014] In the video decoding method and apparatus according to the present disclosure, whether to obtain the first flag may be determined based on whether ISP (Intra sub-partitions) is applied to the current block.
[0015] The video decoding method and apparatus according to the present disclosure can derive an intra induction mode based on DIMD (Decoder-side intra mode derivation), and can select one from among the horizontal planner mode and the vertical planner mode based on the intra induction mode.
[0016] In the video decoding method and apparatus according to the present disclosure, when the intra induction mode is a horizontal direction mode, the intra prediction mode of the current block is induced to the horizontal planner mode, and when the intra induction mode is a vertical direction mode, the intra prediction mode of the current block can be induced to the vertical planner mode.
[0017] The video decoding method and apparatus according to the present disclosure can sort the order of the plurality of predefined intra prediction modes based on a cost calculated by a template matching method.
[0018] The video encoding method and apparatus according to the present disclosure can derive an intra prediction mode of a current block, configure reference samples of the current block, and generate prediction samples of the current block based on the intra prediction mode and the reference samples.
[0019] In the video encoding method and apparatus according to the present disclosure, the intra prediction mode can be derived from a plurality of predefined intra prediction modes.
[0020] In the video encoding method and apparatus according to the present disclosure, the plurality of predefined intra prediction modes can include a normal planar mode, a horizontal planar mode, and a vertical planar mode.
[0021] In the video encoding method and apparatus according to the present disclosure, the horizontal planar mode indicates a mode of generating a predicted sample of the current sample by using a horizontal sample of the current sample and a sample at the upper right end within the current block, the horizontal sample indicates a sample that is horizontally adjacent to the current sample among the reference samples, and the sample at the upper right end can indicate a sample that is adjacent to the corner at the upper right end of the current block among the reference samples.
[0022] In the video encoding method and apparatus according to the present disclosure, the vertical planar mode indicates a mode of generating a predicted sample of the current sample by using a vertical sample of the current sample and a sample at the lower left end within the current block, the vertical sample indicates a sample that is vertically adjacent to the current sample among the reference samples, and the sample at the lower left end can indicate a sample that is adjacent to the corner at the lower left end of the current block among the reference samples.
[0023] The video encoding method and apparatus according to the present disclosure can obtain a first flag indicating whether the normal planar mode is used for the current block.
[0024] When the first flag indicates that the normal planar mode is not used for the current block, the video encoding method and apparatus according to the present disclosure can obtain a second flag indicating an intra prediction mode used for the current block from among the horizontal planar mode or the vertical planar mode.
[0025] In the video encoding method and apparatus according to the present disclosure, whether to obtain the first flag may be determined based on whether ISP (Intra sub-partitions) is applied to the current block.
[0026] The video encoding method and apparatus according to the present disclosure can derive an intra prediction mode based on DIMD (Decoder-side intra mode derivation), and select one of the horizontal planar mode and the vertical planar mode based on the intra prediction mode.
[0027] In the video encoding method and apparatus according to the present disclosure, when the intra prediction mode is a horizontal directional mode, the intra prediction mode of the current block is derived to the horizontal planar mode, and when the intra prediction mode is a vertical directional mode, the intra prediction mode of the current block can be derived to the vertical planar mode.
[0028] The video encoding method and apparatus according to the present disclosure can sort the order of the plurality of predefined intra prediction modes based on the cost calculated by the template matching method.
[0029] A computer-readable digital storage medium storing encoded video / video information that causes a decoding device according to the present disclosure to perform a video decoding method is provided.
[0030] A computer-readable digital storage medium storing video / video information generated by the video encoding method according to the present disclosure is provided.
[0031] A method and apparatus for transmitting video / video information generated by the video encoding method according to the present disclosure are provided.
Advantages of the Invention
[0032] The present disclosure can consider more diverse prediction methods in prediction by defining various planar modes and performing predictions based on them, improving the accuracy of predictions and enhancing the compression performance.
[0033] The present disclosure can reduce the signaling overhead and improve the compression efficiency by effectively inducing the planar mode using DIMD (decoder side intra mode derivation).
[0034] The present disclosure can reduce the signaling overhead and improve the compression efficiency by effectively inducing the planar mode based on template matching.
Brief Description of the Drawings
[0035]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Mode for Carrying Out the Invention
[0036] While the present disclosure can be subjected to various changes and can have various embodiments, specific embodiments are illustrated in the drawings and will be described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood to include all changes, equivalents, and alternatives included in the spirit and technical scope of the present disclosure. Similar reference numerals are used for similar components when describing each drawing.
[0037] Terms such as first, second, etc. can be used to describe various components, but the components should not be limited by the terms. The terms are used only for the purpose of distinguishing one component from another. For example, without departing from the scope of the rights of the present disclosure, the first component may be named the second component, and similarly, the second component may also be named the first component. The term "and / or" includes a combination of a plurality of related described items or any one of the plurality of related described items.
[0038] When it is mentioned that a certain component is "connected to" or "coupled to" another component, it should be understood that it may be directly connected or coupled to the other component, or there may be other components in between. On the contrary, when it is mentioned that a certain component is "directly connected to" or "directly coupled to" another component, it should be understood that there are no other components in between.
[0039] The terms used in this application are merely used to describe specific embodiments and are not intended to limit the present disclosure. Singular expressions include plural expressions unless the context clearly dictates otherwise. In this application, terms such as "including" or "having" are intended to specify the presence of the features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0040] The present disclosure relates to video / video coding. For example, the methods / embodiments disclosed herein may be applicable to the methods disclosed in the VVC (versatile video coding) standard. Also, the methods / embodiments disclosed in this specification may be applicable to the methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268, etc.).
[0041] This specification presents various embodiments related to video / video coding, and the embodiments may be implemented in combination with each other unless otherwise stated.
[0042] In this specification, "video" may mean a set of a series of images over time. "Picture" generally means a unit indicating one image in a specific time period, and "slice" / "tile" is a unit constituting a part of a picture in coding. A slice / tile can include one or more CTUs (Coding Tree Units). One picture can be composed of one or more slices / tiles. One tile is a rectangular area composed of a plurality of CTUs within a specific tile column and a specific tile row of one picture. A tile column is a rectangular area of CTUs having the same height as the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs having a height specified by the picture parameter set and the same width as the width of the picture. The CTUs within one tile are continuously arranged by a CTU raster scan, while the tiles within one picture can be continuously arranged by a tile raster scan. One slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows within the tiles of a picture that can be exclusively included in a single NAL unit. On the other hand, one picture can be divided into two or more sub-pictures. A sub-picture can be a rectangular area of one or more slices within a picture.
[0043] "Pixel", "pixel" or "pel" may mean the smallest unit constituting one picture (or video). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and may indicate only the pixel / pixel value of the luma component, or may indicate only the pixel / pixel value of the chroma component.
[0044] A "unit" can indicate the basic unit of video processing. A unit can include at least one of a specific area of a picture and information related to the corresponding area. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit may be used interchangeably with terms such as "block" or "area" in some cases. In general, an MxN block can include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0045] As used herein, "A or B" can mean "only A", "only B", or "all of A and B". In other words, "A or B" as used herein can be interpreted as "A and / or B". For example, "A, B or C" as used herein can mean "only A", "only B", "only C", or "any combination of A, B and C".
[0046] The slashes ( / ) and commas used herein can mean "and / or". For example, "A / B" can mean "A and / or B". Accordingly, "A / B" can mean "only A", "only B", or "all of A and B". For example, "A, B, C" can mean "A, B or C".
[0047] As used herein, "at least one of A and B" can mean "only A", "only B", or "all of A and B". Also, expressions such as "at least one of A or B" and "at least one of A and / or B" as used herein can be interpreted identically to "at least one of A and B".
[0048] Also, in this specification, "at least one of A, B and C" may mean "only A", "only B", "only C", or "any combination of A, B and C". Also, "at least one of A, B or C" and "at least one of A, B and / or C" may mean "at least one of A, B and C".
[0049] Also, the parentheses used in this specification may mean "for example". Specifically, when it is displayed as "prediction (intra prediction)", "intra prediction" may be proposed as an example of "prediction". In other words, "prediction" in this specification is not limited to "intra prediction", and "intra prediction" may be proposed as an example of "prediction". Also, when it is displayed as "prediction (i.e., intra prediction)", "intra prediction" may be proposed as an example of "prediction".
[0050] The technical features separately described within one drawing in this specification may be embodied separately or simultaneously.
[0051] FIG. 1 illustrates a video / video coding system according to the present disclosure.
[0052] Referring to FIG. 1, the video / video coding system can include a first device (source device) and a second device (receiving device).
[0053] The source device can transmit encoded video / image information or data in file or streaming form to the receiving device through a digital storage medium or a network. The source device can include a video source, an encoding device, and a transmission unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0054] The video source can obtain video / images through processes such as video / image capture, synthesis, or generation. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated through a computer or the like, and in this case, the video / image capture process can be replaced as the process of generating related data.
[0055] The encoding device can encode the input video / image. The encoding device can perform a series of procedures such as prediction, conversion, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in bitstream form.
[0056] The transmission unit can transmit the encoded video / video information or data output in bitstream form to the receiving unit of the receiving device through a digital storage medium or network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include elements for generating media files through a predetermined file format and can include elements for transmission through a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.
[0057] The decoding device can perform a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device to decode the video / video.
[0058] The renderer can render the decoded video / video. The rendered video / video can be displayed through the display unit.
[0059] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of video / video signals is performed.
[0060] Referring to FIG. 2, the encoding device 200 may include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor 220 may include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 may include a transformer (232), a quantizer 233, a dequantizer 234, and an inverse transformer (235). The residual processor 230 may further include a subtractor (231). The adder 250 may be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by one or more hardware components (e.g., an encoding device chipset or a processor) according to an embodiment. Also, the memory 270 may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.
[0061] The video segmentation unit 210 can divide the input video (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit may be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.
[0062] For example, one coding unit can be divided into multiple coding units with deeper depths based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. In this case, for example, the quad-tree structure may be applied first, and the binary-tree structure and / or the ternary-tree structure may be applied later. Or the binary-tree structure may be applied before the quad-tree structure. The coding procedure according to this specification can be performed based on the final coding unit that cannot be further divided. In this case, based on the coding efficiency according to the video characteristics, etc., the largest coding unit can be immediately used as the final coding unit, or if necessary, the coding unit can be recursively divided into coding units with a lower depth, and a coding unit with an optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later.
[0063] As another example, the processing unit may further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit may be respectively divided or partitioned from the aforementioned final coding unit. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0064] The unit may be used interchangeably with terms such as a block or an area in some cases. In a general case, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luma component, or may represent only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel in one picture (or video).
[0065] The encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from the input video signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input video signal (original block, original sample array) in the encoding device 200 may be called a subtraction unit 231.
[0066] The prediction unit 220 can perform a prediction on a processing target block (hereinafter referred to as the current block) and generate a predicted block including a prediction sample for the current block. The prediction unit 220 can determine whether intra prediction or inter prediction is applied in units of the current block or CU. As will be described later in the description of each prediction mode, the prediction unit 220 can generate various pieces of information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0067] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the vicinity (neighbor) of the current block according to the prediction mode, or can be located at a certain distance from the current block. In intra prediction, the prediction mode can include one or more non-directional modes and a plurality of directional modes. The non-directional mode can include at least one of the DC mode or the Planar mode. The directional mode can include 33 directional modes or 65 directional modes depending on the degree of detail of the prediction direction. However, this is an example, and a greater or smaller number of directional modes can be used depending on the setting. The intra prediction unit 222 may determine the prediction mode to be applied to the current block by using the prediction mode applied to the surrounding blocks.
[0068] The inter prediction unit 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted from the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (such as L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can be called by names such as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of neighboring blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of a neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0069] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply not only intra prediction or inter prediction for the prediction of one block, but also can apply intra prediction and inter prediction simultaneously. This can be called the combined inter and intra prediction (CIIP) mode. Also, the prediction unit may be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of the block. The IBC prediction mode or the palette mode can be used for content video / motion video coding such as games like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in a manner similar to inter prediction in terms of guiding a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this specification. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information regarding the palette table and the palette index. The prediction signal generated through the prediction unit 220 can be used to generate a restored signal or can be used to generate a residual signal.
[0070] The conversion unit 232 can generate transform coefficients by applying a conversion technique to the residual signal. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the transform obtained from the graph when representing the relationship information between pixels as a graph. CNT means the transform obtained based on generating a prediction signal using all previously restored pixels. Also, the conversion process may be applied to pixel blocks having the same size of a square, and may also be applied to blocks of variable size instead of a square.
[0071] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240, and the entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bit stream. The information regarding the quantized transform coefficients may be called residual information. The quantization unit 233 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and may generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0072] The entropy encoding unit 240 can perform various encoding methods such as exponential Golomb, CAVLC (context - adaptive variable length coding), CABAC (context - adaptive binary arithmetic coding), etc. The entropy encoding unit 240 may encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients.
[0073] The encoded information (e.g., encoded video / video information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / video information may further include information regarding various parameter sets such as an adaptation parameter set APS, a picture parameter set PPS, a sequence parameter set SPS, or a video parameter set VPS. Also, the video / video information may further include general constraint information. Information and / or syntax elements transmitted / signaled from the encoding device to the decoding device in this specification may be included in the video / video information. The video / video information can be encoded through the encoding procedure described above and included in the bitstream. The bitstream can be transmitted through a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu - ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 with a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit may be included in the entropy encoding unit 240.
[0074] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be restored by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit 234 and the inverse transformation unit 235. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 may be referred to as a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and may be used for inter prediction of the next picture after passing through filtering as described later. On the other hand, LMCS (luma mapping with chroma scaling) may be applied in the picture encoding and / or restoration process.
[0075] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.
[0076] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When the encoding device applies inter prediction through this, the encoding device 200 and the decoding device can avoid prediction mismatches, and the encoding efficiency can also be improved.
[0077] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks in which the motion information within the current picture is induced (or encoded) and / or the motion information of the blocks within the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 270 can store the restored samples of the restored blocks within the current picture and transmit them to the intra prediction unit 222.
[0078] FIG. 3 shows a schematic block diagram of a decoding apparatus to which an embodiment of the present disclosure can be applied and in which decoding of a video / video signal is performed.
[0079] Referring to FIG. 3, the decoding apparatus 300 may include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filtering unit (filter, 350), and a memory (memoery, 360). The predictor 330 may include an inter-prediction unit 332 and an intra-prediction unit 331. The residual processor 320 may include an inverse quantization unit (dequantizer, 321) and an inverse transform unit (inverse transformer, 321).
[0080] The entropy decoding unit 310, the residual processing unit 320, the prediction unit 330, the addition unit 340, and the filtering unit 350 described above may be configured by one hardware component (for example, a decoding apparatus chipset or a processor) according to an embodiment. Also, the memory 360 may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0081] When a bitstream including video / video information is input, the decoding device 300 can restore the video corresponding to the process in which the video / video information is processed by the encoding device of FIG. 2. For example, the decoding device 300 can derive units / blocks based on the block division related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding can be a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. Then, the restored video signal decoded and output through the decoding device 300 can be played back through a playback device.
[0082] The decoding device 300 can receive the signal output from the encoding device of FIG. 2 in the form of a bit stream, and the received signal can be decoded through the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bit stream to derive information (e.g., video / image information) necessary for video restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an adaptation parameter set APS, a picture parameter set PPS, a sequence parameter set SPS, or a video parameter set VPS. Also, the video / image information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this specification can be decoded through the decoding procedure and obtained from the bit stream. For example, the entropy decoding unit 310 can decode the information in the bit stream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bit stream, determines a context model using the syntax element information to be decoded, the surrounding and decoding information of the decoding target block, or the information of the symbol / bin decoded in the previous stage, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin after determining the context model.Of the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction units (inter prediction unit 332 and intra prediction unit 331), and the residual value for which entropy decoding has been performed by the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, the information related to filtering among the information decoded by the entropy decoding unit 310 can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives a signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit may be a component of the entropy decoding unit 310.
[0083] On the other hand, the decoding device according to the present specification may be referred to as a video / video / picture decoding device, and the decoding device may be classified into an information decoding device (video / video / picture information decoding device) and a sample decoding device (video / video / picture sample decoding device). The information decoding device can include the entropy decoding unit 310, and the sample decoding device can include at least one of the inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.
[0084] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output the transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) to obtain the transform coefficients.
[0085] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).
[0086] The prediction unit 320 can perform prediction on the current block and generate a predicted block including the predicted samples for the current block. The prediction unit 320 can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.
[0087] The prediction unit 320 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit 320 can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called the combined inter and intra prediction (CIIP) mode. Also, the prediction unit may be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of the block. The IBC prediction mode or the palette mode can be used for content video / moving image coding such as games like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in a way similar to inter prediction in that it induces a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this specification. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be included in and signaled in the video / video information.
[0088] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block according to the prediction mode, or may be located at a certain distance from the current block. In intra prediction, the prediction mode can include one or more non-directional modes and a plurality of directional modes. The intra prediction unit 331 may determine the prediction mode to be applied to the current block by using the prediction mode applied to the neighboring blocks.
[0089] The inter prediction unit 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted from the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (such as L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the peripheral blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on the peripheral blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the inter prediction mode for the current block.
[0090] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the block to be processed, such as when the skip mode is applied, the prediction block can be used as the restored block.
[0091] The addition unit 340 may be referred to as a restoration unit or a restoration block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described later, or may be used for inter prediction of the next picture. On the other hand, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0092] The filtering unit 350 can apply filtering to the restoration signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0093] The (modified) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the blocks in which the motion information within the current picture is induced (or decoded) and / or the motion information of the blocks within the already restored picture. The stored motion information can be transmitted to the inter prediction unit 332 for utilization as the motion information of spatially neighboring blocks or temporally neighboring blocks. The memory 360 can store the restored samples of the blocks restored within the current picture and can transmit them to the intra prediction unit 331.
[0094] In this specification, the examples described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding device 200 can be applied to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding device 300 in the same or corresponding manner, respectively.
[0095] FIG. 4 is an example according to the present disclosure, illustrating an inter prediction method performed by the decoding device 300.
[0096] Referring to FIG. 4, the decoding device can induce the intra prediction mode of the current block (S400). The intra prediction mode of the current block can be induced from a plurality of predefined intra prediction modes.
[0097] According to an embodiment of the present disclosure, the plurality of predefined intra prediction modes can include a horizontal planar mode and / or a vertical planar mode in addition to the existing planar mode. In one embodiment, the horizontal planar mode and / or the vertical planar mode can be prediction modes that utilize only a part of the process of generating (or inducing) prediction samples in the existing planar mode. In the present disclosure, the horizontal planar mode may be referred to as a horizontal planar prediction mode, and the vertical planar mode may be referred to as a vertical planar prediction mode. The existing planar mode is a planar mode used in general video compression technology, and in the present disclosure, the existing planar mode may be referred to as a general planar mode or a normal planar mode.
[0098] As an example, the normal planar mode can generate prediction samples using the following Mathematical Formula 1.
[0099] [Mathematical Formula 1]
Number
[0100] Referring to Equation 1, the normal planar mode can configure the prediction mode to perform blending of the vertical planar prediction sample (predV) and the horizontal planar prediction sample (predH). In the present disclosure, blending may be referred to as combining, weighted summation, and weighted prediction.
[0101] According to an embodiment of the present disclosure, predV and predH in Equation 1 can be added as separate prediction modes respectively. In other words, predV and predH in Equation 1 can be used for intra prediction as separate intra prediction modes respectively.
[0102] As an example, the vertical planar mode can be defined as in Equation 2 below.
[0103] [Equation 2] [Number]
[0104] Referring to Equation 2, the vertical planar mode can generate a prediction sample by using the vertical sample of the current sample in the current block and the sample at the lower left end. The vertical sample indicates the sample that is vertically adjacent to the current sample among the reference samples of the current block, and the sample at the lower left end indicates the sample that is adjacent to the lower left corner of the current block among the reference samples of the current block. A weighted summation for the vertical sample and the sample at the lower left end can be performed based on the position (or vertical coordinate) of the current sample and the height of the current block. The weighted-summed value can be shifted by a value obtained by taking the logarithm with base 2 of the width of the current block. Also, the horizontal planar mode can be defined as in Equation 3 below.
[0105] [Equation 3] [Number]
[0106] Referring to Equation 3, the horizontal planar mode can generate a predicted sample using the horizontal sample of the current sample and the sample at the upper right corner within the current block. The horizontal sample indicates a sample that is horizontally adjacent to the current sample among the reference samples of the current block, and the sample at the upper right corner indicates a sample that is adjacent to the upper right corner of the current block among the reference samples of the current block. Based on the position (or horizontal coordinate) of the current sample and the width of the current block, a weighted sum for the horizontal sample and the sample at the upper right corner can be performed. The weighted sum value can be shifted by a value obtained by taking the logarithm with base 2 of the height of the current block.
[0107] According to an embodiment of the present disclosure, the vertical planar mode and the horizontal planar mode can be signaled as one of the planar modes. As an example, a syntax structure as shown in Table 1 below can be defined. The syntax (or syntax structure) defined in the present disclosure is described based on the current video compression standard, the VVC (versatile video coding) standard, but the invention is not limited thereto. It can be deformed within substantially the same range in consideration of the essence of the invention.
[0108]
Table 1
[0109] Referring to Table 1, intra_luma_mpm_flag is a syntax element that indicates whether MPM is applied to the current block. If MPM is applied to the current block, it can be indicated by intra_luma_not_planar_flag whether the intra prediction mode of the current block is the planar mode.
[0110] According to an embodiment of the present disclosure, as shown in Table 1, when intra_luma_not_planar_flag indicates that the planar mode is used, a syntax element planar_flag indicating whether the normal planar mode is used as the intra prediction mode may be signaled. When planar_flag indicates that the normal planar mode is not used, a syntax element planar_dir_flag indicating whether it is a horizontal planar mode or a vertical planar mode may be signaled.
[0111] As an example, the method proposed in the present disclosure may be applied only when the current block is not in the ISP (intra-subpartitions) mode. In other words, the directional planar mode that is not the normal planar mode may be used only when the ISP mode is not applied to the current block. That is, whether to obtain planar_flag may be determined based on whether ISP (Intra sub-partitions) is applied to the current block. When planar_flag is not signaled, that is, when planar_flag does not exist, planar_flag may be inferred as 1.
[0112] As an example, mode information may be signaled independently of the normal planar mode transmission method. As an example, a syntax structure as shown in Table 2 below may be defined.
[0113]
Table 2
[0114] Referring to Table 2, prior to the MPM-based intra prediction mode induction process, a syntax element planar_horver_flag indicating whether the directional planner mode is applied to the current block can be signaled. Specifically, if the planar_horver_flag value is 1, it can indicate that the intra prediction mode of the current block (i.e., the current coding unit, the current coding block) is the horizontal planner mode or the vertical planner mode. If the planar_horver_flag value is 0, it can indicate that the intra prediction mode of the current block is not the horizontal planner mode and the vertical planner mode. If planar_horver_flag does not exist, its value can be inferred as 0.
[0115] When the planar_horver_flag value is 1, a syntax element planar_dir_flag indicating the horizontal or vertical planner mode can be signaled. As an example, if the planar_dir_flag value is 1, it can indicate that the intra prediction mode of the current block is the horizontal planner mode, and if the planar_dir_flag value is 0, it can indicate that the intra prediction mode of the current block is the vertical planner mode.
[0116] Also, the horizontal / vertical planner mode flag (planar_horver_flag) and the direction indication flag (planar_dir_flag) can precede the ISP mode flag. As an example, a syntax structure as shown in Table 3 below can be defined.
[0117]
Table 3
[0118] Referring to Table 3, the parsing of the ISP mode can be determined by whether the directional planner mode is applied. The content described in Table 1 and Table 2 above can be equally applied in Table 3. Related duplicate explanations are omitted.
[0119] Also, the horizontal / vertical planar mode flag (planar_horver_flag) and the direction indication flag (planar_dir_flag) can precede the MRL (multi reference line) mode flag. As an example, a syntax structure as shown in Table 4 below can be defined.
[0120]
Table 4
[0121] The content described in Table 1 and Table 2 above can be equally applied to Table 4. Related duplicate explanations are omitted.
[0122] Also, the horizontal / vertical planar mode flag (planar_horver_flag) and the direction indication flag (planar_dir_flag) can precede the MIP (matrix baesed intra prediction) mode flag. As an example, a syntax structure as shown in Table 4 below can be defined.
[0123]
Table 5
[0124] The content described in Table 1 and Table 2 above can be equally applied to Table 5. Related duplicate explanations are omitted.
[0125] Also, when the planar mode is applied, the flag can be parsed without considering the presence or absence of ISP application. A syntax as shown in Table 6 below can be defined.
[0126]
Table 6
[0127] The content described in Table 1 and Table 2 above can be equally applied to Table 6. Related duplicate explanations are omitted.
[0128] Alternatively, mode information proposed independently of the planar mode signaling method can be signaled. A syntax such as the following Table 7 can be defined.
[0129] [Table 7]
[0130] The content described in Table 1 and Table 2 above can be equally applied to Table 7. Related duplicate explanations are omitted.
[0131] Also, the horizontal / vertical planar mode flag (planar_horver_flag) and the direction indication flag (planar_dir_flag) can precede the MRL (multi reference line) mode flag. As an example, a syntax structure such as the following Table 8 can be defined.
[0132] [Table 8]
[0133] The content described in Table 1 and Table 2 above can be equally applied to Table 8. Related duplicate explanations are omitted.
[0134] Also, the horizontal / vertical planar mode flag (planar_horver_flag) and the direction indication flag (planar_dir_flag) can precede the MIP (matrix baesed intra prediction) mode flag. As an example, a syntax structure such as the following Table 9 can be defined.
[0135] [Table 9]
[0136] The content described in Table 1 and Table 2 above can be equally applied to Table 9. Related duplicate explanations are omitted.
[0137] Also, the horizontal / vertical planar mode flag (planar_horver_flag) can be context-model coded. Also, the direction indication flag (planar_dir_flag) can be context-model coded. Or the horizontal / vertical planar mode flag can be context-model coded and the direction indication flag can be bypass coded.
[0138] In the above examples of the syntax, although not separately described for the upper-level switch (i.e., on / off) to represent the relationship with other intra prediction modes, the prediction modes proposed in the embodiments of the present disclosure can have the presence or absence of activation signaled by upper-level syntax such as VPS (video parameter set), SPS (sequence parameter set), PPS (picture parameter set), picture header, slice header, etc. For example, after first signaling the presence or absence of application of the direction indication flag in the SPS, it is possible to first check the presence or absence of activation at the upper level in the lower-level syntax (e.g., coding unit syntax) to determine whether flag parsing is required.
[0139] In the above-described embodiments, examples for the case of application to the luminance component were given, but this prediction method can be equally applied to the chrominance components. For example, when coding the chrominance components, if the prediction mode of the corresponding luminance component is the horizontal planar mode or the vertical planar mode, the same prediction mode can be selected and coded through the DM (direct mode). Or additional information indicating which of the normal planar mode, vertical planar mode, and horizontal planar mode is applied to the chrominance components can be signaled.
[0140] In addition, in this embodiment, a signal line method for the plurality of planar modes described above is proposed. As an example, based on the mode induced by DIMD (decoder-side intra mode derivation), the planar mode to be used for the current block among the plurality of planar modes can be specified.
[0141] DIMD calculates HoG (Histogram of Gradient) based on the restored samples in the periphery, and then utilizes the N prediction modes with high amplitudes of the histogram for the prediction of the current block. When DIMD is applied, the inclination can be calculated based on at least two samples belonging to the peripheral region of the current block. Here, the inclination can include at least one of the horizontal inclination or the vertical inclination. Based on at least one of the calculated inclination or the amplitude of the inclination, the intra prediction mode of the current block can be induced. Here, the amplitude of the inclination can be determined based on the sum of the horizontal inclination and the vertical inclination. Through this induction method, one intra prediction mode may be induced for the current block, or two or more intra prediction modes may be induced.
[0142] As an example, when the inclination-based induction method is applied, one or more intra prediction modes can be induced from the restored peripheral samples. At this time, one or more prediction modes can be induced based on the inclination calculated from the peripheral samples restored according to the embodiments of the present disclosure. For example, one or two or more intra prediction modes can be induced from the restored peripheral samples. As an example, the maximum number of intra prediction modes induced through the inclination-based induction method can be predefined in the encoder / decoder. As an example, the maximum number of intra prediction modes induced through the inclination-based induction method can be N. For example, N can be defined as 2, 3, 4, 5, 6, 7.
[0143] As an example, the predictors (which may be referred to as prediction samples or prediction blocks) obtained by the induced intra prediction mode can be combined with the predictors obtained by the planar mode (which may be abbreviated as planar mode predictors). For example, the predictors obtained by the induced intra prediction mode can be weighted and summed with the planar mode predictors. At this time, as an example, the weighting value used for the weighted sum can be determined based on the slope calculated from the peripheral samples restored according to the embodiments of the present disclosure.
[0144] Also, as an example, when the number of intra prediction modes induced through the slope-based induction method is two or more, the predictors obtained by the induced intra prediction mode can be combined with the planar mode predictors. As an example, when the number of intra prediction modes induced through the slope-based induction method is one, it is not combined (or weighted and summed) with the planar mode predictors, and the predictors obtained by the induced single intra prediction mode can be output as prediction samples or prediction blocks by the slope-based induction method.
[0145] Also, as an example, the inclination can be calculated in units of a window having a predetermined size. Based on the calculated inclination, an angle indicating the directionality of the samples within the corresponding window can be calculated. The calculated angle can correspond to any one of the plurality of predefined intra prediction modes described above. The magnitude of the inclination can be stored / updated for the intra prediction mode corresponding to the calculated angle. In the present disclosure, the magnitude of the inclination can be referred to as the amplitude of the inclination, the histogram amplitude, the size of the histogram, etc. Through such a process, for each window, the intra prediction mode corresponding to the calculated inclination can be determined. The magnitude of the inclination can be stored / updated for the determined intra prediction mode. The top T intra prediction modes having the largest magnitude among the stored magnitudes of the inclination are selected, and the selected intra prediction mode can be set as the intra prediction mode of the current block. Here, T can be an integer of 1, 2, 3, or more.
[0146] As an example, the peripheral region of the current block can be the left side, upper left side, and upper side regions of the current block. For example, the horizontal inclination and the vertical inclination can be derived from the rows and columns of the second peripheral samples. The rows and / or columns of the second peripheral samples can indicate the rows and / or columns located next to the rows and columns of the peripheral samples immediately adjacent to the current block. As an example, a histogram of gradients (HoG) can be derived from the derived horizontal and vertical inclinations. The inclination or the histogram of the inclination can be derived by applying a window using the L-shaped rows and columns of 3 pixels (or pixel lines) around the current block. At this time, the window can be defined as having a size of 3×3. However, as an example and not limited thereto, the window may be defined as having a size of 2×2, 4×4, 5×5, etc. Also, as an example, the window can be a Sobel filter.
[0147] As an example, two or more intra prediction modes having the largest slope magnitude (or amplitude) may be selected. A final prediction block may be generated by blending (or combining, weighted summing) a prediction block predicted using the selected intra prediction mode and a prediction block predicted using the planar mode. At this time, the weighting values applied to the respective prediction blocks may be derived based on the histogram amplitude.
[0148] In one embodiment of the present disclosure, when the planar mode used for the current block is not the normal planar mode, in order to determine which of the horizontal / vertical planar modes is the prediction mode of the current block, the prediction mode with the highest HoG amplitude obtained in the DIMD process (hereinafter referred to as the DIMD mode) can be used. That is, in the above-described embodiment, although planar_flag is signaled, planar_dir_flag is not signaled, and the planar mode used for the current block among the horizontal / vertical planar modes can be determined based on the DIMD prediction mode.
[0149] As an example, when the DIMD mode (i.e., the mode with the highest HoG amplitude) is a horizontal direction mode, the intra prediction mode of the current block may be induced to be the horizontal planar mode. When the DIMD mode is a vertical direction mode, the intra prediction mode of the current block may be induced to be the vertical planar mode. When the DIMD mode is smaller than the diagonal mode (mode 34 of the VVC standard), it may be determined as the horizontal direction mode. When the DIMD mode is larger than or the same as the large angle mode, it may be determined as the vertical direction mode.
[0150] Or as an example, when the current block is in the planar mode, the prediction mode with the highest HoG amplitude obtained in the DIMD process can be used to determine which of the normal planar mode, vertical planar mode, and horizontal planar mode is the prediction mode of the current block.
[0151] That is, in the foregoing embodiments, without signaling planar_flag and planar_dir_flag, the predicted mode of the current block can be determined as the planar mode, vertical planar mode, or horizontal planar mode to be used for the current block based on the mode induced by DIMD.
[0152] For example, when the DIMD mode belongs to a mode within a certain range from the horizontal mode, the planar mode to be used for the current block can be determined as the horizontal planar mode. When the DIMD mode belongs to a mode within a certain range from the vertical mode, the planar mode to be used for the current block can be determined as the vertical planar mode. Otherwise, the planar mode to be used for the current block can be determined as the normal planar mode.
[0153] Specifically, when the DIMD mode is ±M modes from the horizontal mode, the current predicted mode can be induced to the horizontal planar mode. For example, when M is 5 and the DIMD mode is 16, since 13 ≦ DIMD mode ≦ 23, the intra prediction mode of the current block can be induced to the horizontal planar mode.
[0154] Similarly, when the DIMD mode is ±M modes from the vertical mode, the current predicted mode can be induced to the vertical planar mode. Otherwise, it can be induced to the normal planar mode. At this time, when there are 65 intra prediction directions, M can be an integer from 0 to 16 or less. Or when there are K intra prediction directions, M can be an integer from 0 to K / 4 or less.
[0155] The decoding device can configure the reference samples of the current block (S410). The decoding device can generate the predicted samples of the current block based on the determined intra prediction mode of the current block and the configured reference samples (S420).
[0156] The following describes a method of selecting a prediction mode based on the prediction accuracy in a predefined template area.
[0157] FIG. 5 is a drawing for explaining a template matching-based prediction mode determination method according to an embodiment of the present disclosure.
[0158] According to an embodiment of the present disclosure, the planar mode applied to the current block among a plurality of planar modes can be determined using template matching. The planar mode used for the current block among the normal planar mode, the vertical planar mode, and the horizontal planar mode may be determined using template matching, or the planar mode used for the current block among the vertical planar mode and the horizontal planar mode can be determined using template matching.
[0159] Examining the latter case in detail, when a planar mode is used for the current block, a flag (planar_flag in FIG. 4 above) indicating whether it is a normal planar mode or a directional planar mode can be signaled. When the normal planar mode is not applied, that is, when the directional planar mode is applied, the planar mode used for the current block among the vertical planar mode and the horizontal planar mode can be determined using the cost calculated by the template area.
[0160] As an example, a prediction sample (or prediction block) for the template region illustrated in FIG. 5 can be generated using the vertical / horizontal planar mode. At this time, a template reference region can be used. A cost can be calculated based on the difference between the generated prediction sample and the restored sample. At this time, SAD (sum of absolute difference), SATD (sum of absolute transformed difference), etc. can be used for cost calculation. For example, prediction can be performed in the template region illustrated in FIG. 5 in the horizontal or vertical planar mode, and the mode having the minimum cost among the two modes can be selected as the prediction mode of the current block.
[0161] As another example, as a template-based mode selection method, although signaling the planar_dir_flag described in the embodiment of FIG. 4 above, the order of the vertical / horizontal planar mode candidates is sorted in ascending order of cost through the template-based SAD and SATD, and the intra prediction mode of the current block can be determined based on the sorted candidate list.
[0162] At this time, a direction indication flag indicating the vertical / horizontal planar mode can be context-coded. In this case, the accuracy of the direction indication flag can be increased, and an improvement in coding performance can be expected. Also, as an example, the template region can be selectively utilized by the vertical / horizontal planar mode. As in the formula described in FIG. 4 above, the vertical planar mode mainly uses the reference samples at the upper end, and the horizontal planar mode mainly uses the reference samples on the left side.
[0163] Therefore, in the case of the horizontal planar mode, it may be a more accurate prediction to induce a prediction sample in the left template region. Similarly, in the case of the vertical planar mode, it may be more accurate to induce a prediction sample in the upper template region. In the case of the horizontal planar mode, a prediction sample can be generated in the left template region, and in the case of the vertical planar mode, a prediction sample can be generated in the upper template region.
[0164] Also, as an example, when the sizes of the left template and the upper template are different from each other, the prediction mode may be determined by comparing the results derived by normalizing the cost by the number of pixels in each template region. Looking closely, although only the information indicating whether the current block is in the planar mode is signaled, the planar mode used for the current block among the normal planar mode, the vertical planar mode, and the horizontal planar mode can be determined using the cost calculated by the template region.
[0165] As an example, a prediction sample (or prediction block) for the template region illustrated in FIG. 5 can be generated using the normal / vertical / horizontal planar mode. The cost can be calculated based on the difference between the generated prediction sample and the restored sample. At this time, SAD (sum of absolute difference), SATD (sum of absolute transformed difference), etc. can be used for cost calculation. For example, prediction can be performed in the template region illustrated in FIG. 5 in the normal, horizontal, or vertical planar mode, and the mode having the minimum cost among the three modes can be selected as the intra prediction mode of the current block.
[0166] As another example, as a template-based mode selection method, the order of the normal / vertical / horizontal planar mode candidates can be sorted in ascending order of cost through the SAD and SATD of the template base, and the intra prediction mode of the current block can be determined based on the sorted candidate list. Information indicating the candidate used for the current block within the sorted candidate list can be signaled.
[0167] As another example, as a template base mode selection method, the order of the normal / vertical / horizontal planar mode candidates can be sorted in ascending order of cost through the SAD and SATD of the template base, and the intra prediction mode of the current block can be determined based on the sorted candidate list. Information indicating the candidate used for the current block within the sorted candidate list can be signaled.
[0168] Also, as an example, the template area can be selectively utilized in the vertical / horizontal planar mode. As in the formula described in FIG. 4 above, the vertical planar mode mainly uses the upper reference samples, and the horizontal planar mode mainly uses the left reference samples.
[0169] Therefore, in the case of the horizontal planar mode, it may be a more accurate prediction to induce prediction samples in the left template area. Similarly, in the case of the vertical planar mode, it may be more accurate to induce prediction samples in the upper template area. Prediction samples can be generated in the left template area in the case of the horizontal planar mode, and prediction samples can be generated in the upper template area in the case of the vertical planar mode.
[0170] Also, as an example, when the sizes of the left template and the upper template are different from each other, the prediction mode may be determined by comparing the results induced by normalizing the cost by the number of pixels in each template area.
[0171] Also, according to an embodiment of the present disclosure, PDPC (position dependent intra prediction) can be applied to the prediction block generated by the prediction mode proposed in the above embodiment. In the present disclosure, PDPC can be referred to as position-based intra prediction sample filtering. PDPC is a filter for mitigating the discontinuity between the prediction block and the restored samples in the vicinity.
[0172] Similar to other existing prediction blocks, PDPC can also be applied to the prediction blocks configured by the methods described in FIGS. 4 and 5 above. As an example, since the vertical / horizontal planar mode has performance similar to that of the normal planar mode, the PDPC method applied to existing planar prediction blocks can be applied to the vertical / horizontal planar mode.
[0173] Or, the horizontal planar mode relatively utilizes many left reference samples for prediction. On the contrary, except for the upper right reference sample, it does not utilize the upper reference samples. Therefore, considering this, a PDPC method for a horizontal prediction mode with similar characteristics can be applied. Similarly, the vertical planar mode relatively utilizes many upper reference samples for prediction. On the contrary, except for the lower left reference sample, it does not utilize the left reference samples. Therefore, considering this, a PDPC method for a vertical prediction mode with similar characteristics can be applied.
[0174] On the other hand, since PDPC is an additional process for prediction blocks and is applied in pixel units, it can affect the complexity of the codec. Considering this, PDPC application conditions for horizontal / vertical planar prediction blocks can be defined. For example, the presence or absence of PDPC application for the horizontal / vertical planar mode can be determined by the size of the block, the ratio of width to height. Or the presence or absence of PDPC application for the horizontal / vertical planar mode can be determined by the prediction mode of surrounding blocks. Or the horizontal / vertical planar mode may not apply PDPC.
[0175] FIG. 6 illustrates a schematic configuration of an intra prediction unit 331 that performs the intra prediction method according to the present disclosure.
[0176] Referring to FIG. 6, the intra prediction unit 331 can include an intra prediction mode induction unit 600, a reference sample configuration unit 610, and a prediction sample generation unit 620.
[0177] The intra prediction mode induction unit 600 can induce the intra prediction mode of the current block. The intra prediction mode of the current block can be induced from a plurality of predefined intra prediction modes.
[0178] As described above, in addition to the existing planner mode, the plurality of predefined intra prediction modes can include a horizontal planner mode and / or a vertical planner mode. In one embodiment, the horizontal planner mode and / or the vertical planner mode can be prediction modes that utilize only a part of the process of generating (or inducing) prediction samples in the existing planner mode.
[0179] As described above, when the horizontal planner mode is applied, a prediction sample of the current sample can be generated using the horizontal sample of the current sample and the sample at the upper right end within the current block. When the vertical planner mode is applied, a prediction sample of the current sample can be generated using the vertical sample of the current sample and the sample at the lower left end within the current block.
[0180] Also, as described in Tables 1 to 9 above, a syntax element indicating whether the planner mode is used for the current block, a syntax indicating the planner mode applied to the current block among the plurality of planner modes, etc. can be signaled.
[0181] Also, as described above, the mode proposed in the present disclosure can be determined based on whether ISP (Intra sub-partitions) is applied to the current block. As an example, the directional planner mode proposed in the present disclosure can be used for intra prediction of the current block only when ISP is not applied.
[0182] Also, as described above, based on the mode induced by DIMD (decoder-side intra mode derivation), the planner mode used for the current block among the plurality of planner modes can be specified.
[0183] Also, as described with reference to FIG. 5 above, the planner mode applied to the current block among a plurality of planner modes can be determined using template matching. The planner mode used for the current block may be determined among the normal planner mode, the vertical planner mode, and the horizontal planner mode using template matching, and the planner mode used for the current block can be determined among the vertical planner mode and the horizontal planner mode using template matching.
[0184] The reference sample component 610 can configure the reference sample of the current block. The predicted sample generation unit 620 can generate the predicted sample of the current block based on the determined intra prediction mode of the current block and the configured reference sample.
[0185] Also, as described above, position-based intra prediction sample filtering can be applied to the predicted block generated by the planner mode proposed in the present disclosure.
[0186] FIG. 7 is an example according to the present disclosure, illustrating an intra prediction method performed by the encoding device 200.
[0187] In an embodiment of the present disclosure, a planner mode-based intra prediction method performed by an encoding device will be described. The embodiments described with reference to FIGS. 4 to 6 above can be applied in substantially the same manner, and duplicate descriptions are omitted here.
[0188] Referring to FIG. 7, the encoding device can determine the intra prediction mode of the current block (S700). The intra prediction mode of the current block can be determined from a plurality of predefined intra prediction modes.
[0189] As described above, in addition to the plurality of predefined intra prediction modes, the horizontal planner mode and / or the vertical planner mode can be included. In one embodiment, the horizontal planner mode and / or the vertical planner mode can be prediction modes that utilize only a part of the process of generating (or deriving) prediction samples in the existing planner mode.
[0190] As described above, when the horizontal planner mode is applied, the prediction sample of the current sample can be generated using the horizontal sample of the current sample and the sample at the upper right end within the current block. When the vertical planner mode is applied, the prediction sample of the current sample can be generated using the vertical sample of the current sample and the sample at the lower left end within the current block.
[0191] Also, as described in Tables 1 to 9 above, syntax elements indicating whether the planner mode is used for the current block, syntax indicating the planner mode applied to the current block among the plurality of planner modes, etc. can be signaled.
[0192] Also, as described above, the mode proposed in the present disclosure can be determined based on whether ISP (Intra sub-partitions) is applied to the current block. As an example, the directional planner mode proposed in the present disclosure can be used for intra prediction of the current block only when ISP is not applied.
[0193] Also, as described above, the planner mode used for the current block among the plurality of planner modes can be specified based on the mode induced by DIMD (decoder-side intra mode derivation).
[0194] Also, as described with reference to FIG. 5 above, the planner mode applied to the current block among a plurality of planner modes can be determined using template matching. The planner mode used for the current block among the normal planner mode, the vertical planner mode, and the horizontal planner mode may be determined using template matching, and the planner mode used for the current block among the vertical planner mode and the horizontal planner mode can be determined using template matching.
[0195] The encoding device can configure the reference samples of the current block (S710). The encoding device can generate prediction samples for the current block based on the determined intra prediction mode of the current block and the configured reference samples (S720).
[0196] Also, as described above, position-based intra prediction sample filtering can be applied to the prediction block generated by the planner mode proposed in the present disclosure.
[0197] FIG. 8 illustrates a schematic configuration of an intra prediction unit 222 that performs an inter prediction method according to the present disclosure.
[0198] Referring to FIG. 8, the intra prediction unit 222 can include an intra prediction mode determination unit 800, a reference sample configuration unit 810, and a prediction sample generation unit 820.
[0199] Specifically, the intra prediction mode determination unit 800 can determine the intra prediction mode of the current block. The intra prediction mode of the current block can be determined from a plurality of predefined intra prediction modes.
[0200] As described above, in addition to the predefined plurality of intra prediction modes, the horizontal planner mode and / or the vertical planner mode can be included. In one embodiment, the horizontal planner mode and / or the vertical planner mode can be prediction modes that utilize only a part of the process of generating (or deriving) prediction samples in the existing planner mode.
[0201] As described above, when the horizontal planner mode is applied, the prediction sample of the current sample can be generated using the horizontal sample of the current sample and the sample at the upper right end within the current block. When the vertical planner mode is applied, the prediction sample of the current sample can be generated using the vertical sample of the current sample and the sample at the lower left end within the current block.
[0202] Also, as described in Tables 1 to 9 above, a syntax element indicating whether the planner mode is used for the current block, a syntax indicating the planner mode applied to the current block among the plurality of planner modes, etc. can be signaled.
[0203] Also, as described above, the mode proposed in the present disclosure can be determined based on whether ISP (Intra sub-partitions) is applied to the current block. As an example, the directional planner mode proposed in the present disclosure can be used for intra prediction of the current block only when ISP is not applied.
[0204] Also, as described above, based on the mode induced by DIMD (decoder-side intra mode derivation), the planner mode used for the current block among the plurality of planner modes can be specified.
[0205] Also, as described with reference to FIG. 5 above, the planner mode to be applied to the current block among a plurality of planner modes can be determined using template matching. The planner mode to be used for the current block among the normal planner mode, the vertical planner mode, and the horizontal planner mode may be determined using template matching, and the planner mode to be used for the current block among the vertical planner mode and the horizontal planner mode can be determined using template matching.
[0206] The reference sample component 810 can constitute a reference sample of the current block. The predicted sample generation unit 820 can generate a predicted sample of the current block based on the determined intra prediction mode of the current block and the configured reference sample.
[0207] Also, as described above, position-based intra prediction sample filtering can be applied to the predicted block generated by the planner mode proposed in the present disclosure.
[0208] In the above-described embodiments, the method is described as a series of steps or blocks based on a flowchart, but the corresponding embodiments are not limited to the order of the steps, and a certain step may occur in a different order from the steps described above or simultaneously. Also, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included or one or more steps of the flowchart may be deleted without affecting the scope of the embodiments of this document.
[0209] The method according to the embodiments of the present document described above can be embodied in software form, and the encoding device and / or decoding device according to this document can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0210] When an embodiment is implemented in software in this document, the above-described method may be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules may be stored in a memory and executed by a processor. The memory may be inside or outside the processor and may be connected to the processor by various means well known in the art. The processor may include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory may include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document may be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in each drawing may be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm may be stored in a digital storage medium.
[0211] In addition, the decoding device and the encoding device to which the embodiments of the present specification are applied can be included in a multimedia broadcast transmission / reception device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video dialogue device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a video camera, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (augmented reality) device, a picture phone video device, a transportation means terminal (e.g., a vehicle terminal including an autonomous driving vehicle, an airplane terminal, a ship terminal, etc.) and a medical video device, etc., and can be used to process video signals or data signals. For example, an OTT video (Over the top video) device can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0212] In addition, the processing method to which the embodiments of this specification are applicable can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of this specification can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Further, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (for example, transmission through the Internet). Also, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted through a wired / wireless communication network.
[0213] In addition, the embodiments of this specification can be embodied as a computer program product by program code, and the program code can be executed by a computer according to the embodiments of this specification. The program code can be stored on a carrier readable by a computer.
[0214] FIG. 9 shows an example of a content streaming system to which the embodiments of the present disclosure can be applied.
[0215] Referring to FIG. 9, the content streaming system to which the embodiments of this specification are applicable can generally include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.
[0216] The encoding server compresses the content input from multimedia input devices such as smartphones, cameras, video cameras, etc. into digital data to generate a bitstream, and transmits this to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, video cameras, etc. directly generate a bitstream, the encoding server may be omitted.
[0217] The bitstream may be generated by an encoding method or a bitstream generation method to which the embodiments of this specification are applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0218] The streaming server transmits multimedia data to the user device based on a user request through a web server, and the web server serves as a medium to inform the user of what services are available. If the user requests a service desired by the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system may include a separate control server. In this case, the control server serves to control commands / responses between each device within the content streaming system.
[0219] The streaming server can receive content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time in order to provide a smooth streaming service.
[0220] Examples of the user device may include mobile phones, smart phones, laptop computers, digital broadcast terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head mounted displays (HMDs)), digital TVs, desktop computers, digital signage, and the like.
[0221] Each server in the content streaming system may be operated as a distributed server, and in this case, the data received by each server may be distributedly processed.
[0222] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented in an apparatus, and the technical features of the apparatus claims in this specification can be combined and implemented in a method. Also, the technical features of the method claims in this specification and the technical features of the apparatus claims can be combined and implemented in an apparatus, and the technical features of the method claims in this specification and the technical features of the apparatus claims can be combined and implemented in a method.
Claims
1. inducing an intra prediction mode of a current block; constructing reference samples of the current block; generating a predicted sample of the current block based on the intra prediction mode and the reference samples, wherein the intra prediction mode is induced from a plurality of predefined intra prediction modes, and the plurality of predefined intra prediction modes include a normal planar mode, a horizontal planar mode, and a vertical planar mode, a video decoding method.
2. The horizontal planar mode indicates a mode of generating a predicted sample of the current sample by using a horizontal sample of the current sample in the current block and a sample at the upper right end, wherein the horizontal sample indicates a sample horizontally adjacent to the current sample among the reference samples, and the sample at the upper right end indicates a sample adjacent to a corner at the upper right end of the current block among the reference samples. The video decoding method according to claim 1.
3. The vertical planar mode indicates a mode of generating a predicted sample of the current sample by using a vertical sample of the current sample in the current block and a sample at the lower left end, wherein the vertical sample indicates a sample vertically adjacent to the current sample among the reference samples, and the sample at the lower left end indicates a sample adjacent to a corner at the lower left end of the current block among the reference samples. The video decoding method according to claim 1.
4. The step of inducing the intra prediction mode of the current block includes: obtaining a first flag indicating whether the normal planar mode is used for the current block. The video decoding method according to claim 1.
5. The step of inducing the intra prediction mode of the current block further includes: when the first flag indicates that the normal planar mode is not used for the current block, obtaining a second flag indicating an intra prediction mode used for the current block from among the horizontal planar mode and the vertical planar mode. The video decoding method according to claim 4.
6. Whether to obtain the first flag is determined based on whether ISP (Intra sub - partitions) is applied to the current block. The video decoding method according to claim 4.
7. The step of deriving the intra prediction mode of the current block includes: a step of deriving an intra induction mode based on DIMD (Decoder-side intra mode derivation); and a step of selecting one of the horizontal planar mode and the vertical planar mode based on the intra induction mode, the video decoding method according to claim 1.
8. When the intra induction mode is a horizontal direction mode, the intra prediction mode of the current block is derived to the horizontal planar mode; When the intra induction mode is a vertical direction mode, the intra prediction mode of the current block is derived to the vertical planar mode, the video decoding method according to claim 7.
9. The step of deriving the intra prediction mode of the current block includes: a step of arranging the order of the plurality of predefined intra prediction modes based on a cost calculated by a template matching method, the video decoding method according to claim 1.
10. A step of determining an intra prediction mode of a current block; a step of constructing a reference sample of the current block; a step of generating a prediction sample of the current block based on the intra prediction mode and the reference sample, but the intra prediction mode is determined from a plurality of predefined intra prediction modes; the plurality of predefined intra prediction modes include a normal planar mode, a horizontal planar mode, and a vertical planar mode, a video encoding method.
11. A computer-readable storage medium for storing a bitstream generated by the video encoding method according to claim 10.
12. A step of determining an intra prediction mode of a current block; a step of constructing a reference sample of the current block; a step of generating a prediction sample of the current block based on the intra prediction mode and the reference sample; a step of generating a bitstream by encoding the current block based on the prediction sample; a step of transmitting data including the bitstream, but the intra prediction mode is derived from a plurality of predefined intra prediction modes; A data transmission method for video information, wherein the plurality of predefined intra prediction modes include a normal planner mode, a horizontal planner mode, and a vertical planner mode.