Matrix-based Intra Prediction Apparatus and Method
The matrix-based intra prediction method improves video coding efficiency by deriving and generating intra prediction samples, addressing the need for efficient compression of high-resolution and immersive media.
Patent Information
- Application Number
- JP2024099927
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-03
- Filing Date
- 2024-06-20
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-06-03
AI Technical Summary
The increasing demand for high-resolution and high-quality video data, including immersive media like VR and AR content, has led to a need for more efficient video/image compression technologies to reduce transmission and storage costs.
A matrix-based intra prediction method and apparatus for video coding, which includes deriving intra prediction samples using matrix-based intra prediction (MIP) mode information and generating restored samples through downsampling and upsampling processes.
This approach enhances coding efficiency, reduces implementation complexity, and improves prediction performance by efficiently coding index information for intra prediction.
Smart Images

Figure 0007701520000051 
Figure 0007701520000052 
Figure 0007701520000053
Abstract
Description
Technical Field
[0001] This document relates to video coding technology, and more particularly to video coding technology for a matrix-based intra prediction apparatus and method.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality video / images such as 4K or 8K and above UHD (Ultra High Definition) video / images has been increasing in various fields. As the video / image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to existing video / image data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband network, or storing video / image data using an existing storage medium, the transmission cost and storage cost increase.
[0003] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of video / images having video characteristics different from those of real-world video / images, such as game video, has been increasing.
[0004] Accordingly, there is a need for a highly efficient video / image compression technology to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality video / images having various characteristics as described above.
Summary of the Invention
Problems to be Solved by the Invention
[0005] The technical problem of this document is to provide a method and apparatus for improving the coding efficiency of video.
[0006] Another technical problem of this document is to provide an efficient intra prediction method and apparatus.
[0007] Another technical problem of this document is to provide a video coding method and apparatus for intra prediction based on a matrix.
[0008] Another technical problem of this document is to provide a video coding method and apparatus for coding mode information for intra prediction based on a matrix.
Means for Solving the Problem
[0009] According to one embodiment of this document, a video decoding method executed by a decoding apparatus is provided. The method includes receiving, for the current block, second flag information indicating whether the matrix-based intra prediction (MIP) is used based on first flag information indicating whether the matrix-based intra prediction (MIP) is applicable to the current block; receiving matrix-based intra prediction (MIP) mode information based on the second flag information; generating an intra prediction sample for the current block based on the MIP mode information; and generating a restored sample for the current block based on the intra prediction sample.
[0010] The first flag information can be signaled via syntax information of a sequence parameter set (SPS), and the second flag information can be signaled via syntax information of a coding unit.
[0011] The MIP mode information can be index information indicating the MIP mode applied to the current block.
[0012] The step of generating the intra prediction sample may include: a step of downsampling a reference sample adjacent to the current block to derive a reduced boundary sample; a step of deriving a reduced prediction sample based on the multiplication of the reduced boundary sample and a MIP matrix; and a step of upsampling the reduced prediction sample to generate the intra prediction sample for the current block.
[0013] Here, the reduced boundary sample may be downsampled by averaging the reference samples, and the intra prediction sample may be upsampled by linear interpolation of the reduced prediction sample.
[0014] The MIP matrix can be derived based on the size of the current block and the index information.
[0015] The MIP matrix can be selected from any one of three matrix sets classified according to the size of the current block, and each of the three matrix sets can include a plurality of MIP matrices.
[0016] According to one embodiment of the present document, a video encoding method executed by an encoding device is provided. The method includes: deriving whether matrix-based intra prediction (MIP) is applicable to a current block; when the MIP is applicable to the current block, deriving an intra prediction sample of the current block based on the MIP; deriving a residual sample for the current block based on the intra prediction sample; and encoding information about the residual sample and information about the MIP. The information about the MIP can include first flag information indicating whether matrix-based intra prediction (MIP) is applicable to the current block, second flag information indicating whether the MIP is applied to the current block, and matrix-based intra prediction (MIP) mode information.
[0017] According to another embodiment of the present document, a digital storage medium storing video data including encoded video information and a bitstream generated by a video encoding method executed by an encoding device can be provided.
[0018] According to another embodiment of the present document, a digital storage medium storing video data including encoded video information that causes a decoding device to perform the video decoding method and a bitstream can be provided.
Effects of the Invention
[0019] This document can have various effects. For example, according to one embodiment of this document, the compression efficiency of general video can be increased. Or, according to one embodiment of this document, through efficient intra prediction, the complexity of implementation can be reduced and the prediction performance can be improved, thereby improving the efficiency of general coding. Or, according to one embodiment of this document, during matrix-based intra prediction, the index information indicating this can be efficiently coded to improve the coding efficiency.
[0020] The effects obtained through a specific example of this document are not limited to the effects listed above. For example, there may be various technical effects that can be understood or induced from this document by a person having ordinary skill in the related art. Thus, the specific effects of this document are not limited to those explicitly described in this document and may include various effects that can be understood or induced from the technical features of this document.
Brief Description of the Drawings
[0021]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
MODE FOR CARRYING OUT THE INVENTION
[0022] This document can be modified in various ways and can have various embodiments. However, specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this document are merely used to explain specific embodiments and are not intended to limit the technical idea in this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this document, terms such as "including" or "having" are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the document, and should be understood not to preclude the possibility of the existence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof.
[0023] On the other hand, each configuration on the drawings described in this document is shown independently for the convenience of explaining different characteristic functions, and does not mean that each configuration is implemented by separate hardware or separate software. For example, among each configuration, two or more configurations may be combined to form one configuration, and one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of rights of this document as long as they do not deviate from the essence of this document.
[0024] In this document, "A or B" can mean "only A", "only B", or "both A and B". In other words, in this document, "A or B" can be interpreted as "A and / or B". For example, in this document, "A, B or C" can mean "only A", "only B", "only C", or "any combination of A, B and C".
[0025] The slashes ( / ) and commas used in this document may mean "and / or". For example, "A / B" may mean "A and / or B". Thus, "A / B" may mean "only A", "only B", or "both A and B". For example, "A, B, C" may mean "A, B or C".
[0026] In this document, "at least one of A and B" may mean "only A", "only B", or "both A and B". Also, in this document, expressions such as "at least one of A or B" and "at least one of A and / or B" may be interpreted in the same way as "at least one of A and B".
[0027] Also, in this document, "at least one of A, B and C" may mean "only A", "only B", "only C", or "any combination of A, B and C". Furthermore, "at least one of A, B or C" and "at least one of A, B and / or C" may mean "at least one of A, B and C".
[0028] Also, the parentheses used in this document may mean "for example". Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction". In other words, "prediction" in this document is not limited to "intra prediction", and "intra prediction" may be proposed as an example of "prediction". Also, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be proposed as an example of "prediction".
[0029] In this document, the technical features separately described within one drawing may be implemented separately or simultaneously.
[0030] This document relates to video / video coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or the next-generation video / video coding standard (e.g., H.267 or H.268, etc.).
[0031] In this document, various embodiments related to video / video coding are presented, and unless otherwise mentioned, the embodiments may be executed in combination with each other.
[0032] In this document, "video" can mean a collection of a series of images over time. "Picture" generally means a unit representing one image at a specific time period, and "slice" / "tile" is a unit that constitutes a part of a picture in coding. A slice / tile can contain one or more CTUs (Coding Tree Units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can contain one or more tiles. A brick may represent a rectangular region of CTU rows within a tile in a picture. A tile may be partitioned into multiple bricks, each of which consisting of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may be also referred to as a brick.A brick scan may indicate a specific sequential ordering of CTUs partitioning a picture, where the CTUs are ordered consecutively in a CTU raster scan within a brick, bricks within a tile are ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set.The tile row is a rectangular region of CTUs having a width specified by syntax elements in the picture parameter set and a height equal to the height of the picture. A tile scan is a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice may include an integer number of bricks of a picture that may be exclusively contained in a single NAL unit. A slice may consists of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile.In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0033] A pixel or pel may mean the smallest unit that constitutes one picture (or video). Also, the term "sample" may be used as a term corresponding to a pixel. A sample may generally indicate a pixel or a pixel value, may indicate only the pixel / pixel value of the luma component, or may indicate only the pixel / pixel value of the chroma component. Alternatively, a sample may mean a pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it may mean a conversion coefficient in the frequency domain.
[0034] A unit can indicate the basic unit of video processing. A unit can include at least one of a specific region of a picture and information regarding the corresponding region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit may, in some cases, be used interchangeably with terms such as a block or an area. In a general case, an M×N block can include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0035] Hereinafter, with reference to the attached drawings, the preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals will be used for the same components in the drawings, and duplicate descriptions of the same components may be omitted.
[0036] FIG. 1 schematically shows an example of a video / video coding system that can be applied to the embodiments of this document.
[0037] Referring to FIG. 1, a video / video coding system can include a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data in the form of a file or a stream to the receiving device via a digital storage medium or a network.
[0038] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can also be composed of a separate device or an external component.
[0039] The video source can obtain video / video through processes such as capture, synthesis, or generation of video / video. The video source can include a video / video capture device and / or a video / video generation device. The video / video capture device can include, for example, one or more cameras, a video / video archive including previously captured video / video, etc. The video / video generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / video. For example, virtual video / video can be generated via a computer or the like, and in this case, the video / video capture process can be replaced by the process of generating related data.
[0040] The encoding device can encode the input video / image. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0041] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.
[0042] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0043] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0044] Figure 2 is a diagram schematically explaining the configuration of a video / image encoding device that can be applied to the embodiments of this document. Hereinafter, the video encoding device can include the image encoding device.
[0045] Referring to FIG. 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be called a reconstructor or a reconstructed block generator. The aforementioned image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.
[0046] The video segmentation unit 210 can divide the input video (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit may be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or, the binary-tree structure can also be applied first. Based on the final coding unit that cannot be further divided, the coding procedure according to this document can be performed. In this case, based on the coding efficiency according to the video characteristics, etc., the largest coding unit can be immediately used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, conversion, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can be divided or partitioned from the above-mentioned final coding unit respectively.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.
[0047] The unit can be used interchangeably with terms such as block or area, as the case may be. In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel in one picture (or video).
[0048] The encoding device 200 can subtract the prediction signal (predicted block, predicted sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from the input video signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input video signal (original block, original sample array) within the encoding device 200 can be called the subtraction unit 231. The prediction unit can perform prediction on the processing target block (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0049] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the neighborhood of the current block or at a distance from the current block depending on the prediction mode. The prediction modes in intra prediction can include a plurality of non - directional modes and a plurality of directional modes. The non - directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the fineness of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.
[0050] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called by names such as a collocated reference block and a collocated CU (col CU), and the reference picture including the temporal neighboring block may also be called a collocated picture (colPic). For example, the inter prediction unit 221 can configure a motion information candidate list based on the peripheral block, and generate information indicating which candidate is used to derive the motion vector and / or the index of the reference picture of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 221 can use the motion information of the peripheral block as the motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0051] The prediction unit 220 can generate a prediction signal based on various prediction methods described below. For example, for the prediction of one block, the prediction unit can apply not only intra prediction or inter prediction, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit may be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for coding content videos / movies such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information regarding the palette table and the palette index.
[0052] The prediction signal generated via the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when expressing the relationship information between pixels in a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process may be applied to a pixel block having the same size of a square or may be applied to a block of a variable size that is not square.
[0053] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients may be referred to as residual information. The quantization unit 233 can reorder the block-form quantized transform coefficients in the form of a one-dimensional vector based on the scan order of the coefficients, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the form of the one-dimensional vector. The entropy encoding unit 240 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS), etc. Also, the video / video information can further include general constraint information. In this document, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / video information. The video / video information can be encoded via the encoding procedure described above and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 may be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit may be included in the entropy encoding unit 240.
[0054] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 250 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 may be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and as will be described later, can also be used for inter prediction of the next picture after passing through filtering.
[0055] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture encoding and / or restoration process.
[0056] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and can store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240 as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0057] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve the encoding efficiency.
[0058] The DPB in memory 270 can be saved for using the corrected restored picture as a reference picture in the inter prediction unit 221. Memory 270 can save the motion information of the block for which the motion information in the current picture has been derived (or encoded) and / or the motion information of the block in the already restored picture. The saved motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatial neighboring blocks or the motion information of temporal neighboring blocks. Memory 270 can save the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 222.
[0059] FIG. 3 is a diagram schematically explaining the configuration of a video / video decoding apparatus applicable to the embodiments of this document.
[0060] Referring to FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filtering unit 350, and a memory 360. The predictor 330 can include an inter prediction unit 331 and an intra prediction unit 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The above-described entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering unit 350 can be configured by one hardware component (for example, a decoder chipset or a processor) according to the embodiment. Also, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.
[0061] When a bitstream including video / video information is input, the decoding device 300 can restore the video corresponding to the process in which the video / video information was processed by the encoding device in FIG. 2. For example, the decoding device 300 can derive units / blocks based on the information regarding block division obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding may be, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit into a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. Then, the restored video signal decoded and output via the decoding device 300 can be played back via a playback device.
[0062] The decoding device 300 can receive the signal output from the encoding device in FIG. 2 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding and the block to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin according to the determined context model, executes arithmetic decoding of the bin, and can generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the next symbol / bin context model after determining the context model. Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit (inter prediction unit 332 and intra prediction unit 331), and the residual value for which entropy decoding is executed by the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, among the information decoded by the entropy decoding 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device according to the present document can be called a video / video / picture decoding device, and the decoding device can also be classified into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.
[0063] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output the transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the scan order of the coefficients executed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) to obtain the transform coefficients.
[0064] In the inverse transform unit 322, the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0065] The prediction unit can perform prediction on the current block and generate a predicted block including the predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.
[0066] The prediction unit 320 can generate a prediction signal based on various prediction methods described later. For example, for the prediction of one block, the prediction unit can apply not only intra prediction or inter prediction, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit may be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of the block. The IBC prediction mode or the palette mode can be used for content video / moving picture coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be signaled and included in the video / video information.
[0067] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples referred to may be located adjacent to or away from the periphery of the current block depending on the prediction mode. The prediction modes in intra prediction can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.
[0068] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on the peripheral blocks, and derive the motion vector of the current block and / or the index of the reference picture based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.
[0069] The addition unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the processing target block, as in the case where the skip mode is applied, the predicted block can be used as the restored block.
[0070] The addition unit 340 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, and may be output after passing through filtering as described later, or may be used for inter prediction of the next picture.
[0071] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture decoding process.
[0072] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically, the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0073] The (modified) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the blocks for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 332 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 331.
[0074] In this document, the embodiments described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding device 200 can be applied in the same or corresponding manner to the filtering unit 250, the inter prediction unit 332, and the intra prediction unit 331 of the decoding device 300, respectively.
[0075] As described above, when performing video coding, prediction is performed to improve the compression efficiency. Through this, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is derived in the same way in the encoding device and the decoding device. The encoding device can improve the efficiency of video coding by signaling information (residual information) regarding the residual between the original block and the predicted block, rather than the original sample values of the original block, to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, combine the residual block and the predicted block, generate a restored block including restored samples, and generate a restored picture including the restored block.
[0076] The residual information can be generated through conversion and quantization procedures. For example, an encoding device can derive a residual block between an original block and a predicted block, perform a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, perform a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and signal the related residual information to a decoding device (via a bitstream). Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can perform an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in the inter prediction of subsequent pictures to derive a residual block and generate a restored picture based on this.
[0077] FIG. 4 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.
[0078] Referring to FIG. 4, a content streaming system to which the embodiments of this document are applied can generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0079] The encoding server serves to compress the content input from a multimedia input device such as a smartphone, a camera, a camcorder, etc. into digital data to generate a bitstream and transmit this to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, a camcorder, etc. directly generates a bitstream, the encoding server can be omitted.
[0080] The bitstream can be generated by an encoding method or a bitstream generation method applicable to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0081] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as a medium to inform the user of what services are available. If the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include another control server. In this case, the control server plays a role in controlling commands / responses between each device in the content streaming system.
[0082] The streaming server can receive content from a media storage and / or encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0083] Examples of the user device may include mobile phones, smart phones, laptop computers, digital broadcast terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head mounted displays (HMDs)), digital TVs, desktop computers, digital signage, and the like.
[0084] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be distributedly processed.
[0085] FIG. 5 exemplarily shows context-adaptive binary arithmetic coding (CABAC) for encoding a syntax element. For example, in the encoding process of CABAC, when the input signal is not a binary value but a syntax element, the encoding device can binarize the value of the input signal to convert the input signal into a binary value. Also, when the input signal is already a binary value (i.e., when the value of the input signal is a binary value), it can be bypassed without binarization. Here, each binary number 0 or 1 that constitutes a binary value can be called a bin. For example, when the binary string after binarization is 110, 1, 1, and 0 are each called one bin. The bin for one syntax element can indicate the value of the syntax element. Such binarization can be based on various binarization methods such as Truncated Rice binarization process, Fixed-length binarization process, etc., and the binarization method for the target syntax element can be predefined. The binarization procedure can be executed by a binarization unit within the entropy encoding unit.
[0086] Thereafter, the binarized bin of the syntax element can be input to a regular encoding engine or a bypass encoding engine. The regular encoding engine of the encoding device can assign a context model that reflects a probability value to the corresponding bin, and can encode the corresponding bin based on the assigned context model. The regular encoding engine of the encoding device can update the context model for the corresponding bin after performing encoding for each bin. As described above, the bin to be encoded can be referred to as a context-coded bin.
[0087] On the one hand, when the evolved bins of the syntax element are input to the bypass encoding engine, they can be coded as follows. For example, the bypass encoding engine of the encoding device omits the procedure of estimating the probability for the input bin and the procedure of updating the probability model applied to the bin after encoding. When bypass encoding is applied, the encoding device can apply a uniform probability distribution instead of assigning a context model and encode the input bin, thereby improving the encoding speed. As described above, the bin to be encoded can be referred to as a bypass bin.
[0088] Entropy decoding can be shown as a process of executing the same process as the aforementioned entropy encoding in reverse order.
[0089] The decoding device (entropy decoding unit) can decode the encoded video / video information. The video / video information can include information related to partitioning, information related to prediction (e.g., inter / intra prediction partition information, intra prediction mode information, inter prediction mode information, etc.), residual information, information related to in-loop filtering, etc., or can include various syntax elements related thereto. The entropy encoding can be executed in units of syntax elements.
[0090] The decoding device can perform binarization on the target syntax element. Here, the binarization can be based on various binarization methods such as Truncated Rice binarization process, Fixed-length binarization process, etc., and the binarization method for the target syntax element can be defined in advance. The decoding device can derive an available bit string (a candidate for the bit string) for the available value of the target syntax element through the binarization procedure. The binarization procedure can be executed by a binarization unit within the entropy decoding unit.
[0091] The decoding device sequentially decodes and parses each bin for the target syntax element from the input bits in the bitstream, and compares the derived bit string with the available bit strings for the corresponding syntax element. If the derived bit string is the same as one of the available bit strings, the value corresponding to the corresponding bit string is derived as the value of the corresponding syntax element. If not, after further parsing the next bit in the bitstream, the above-described procedure is executed again. Through such a process, without using start bits or end bits for specific information (specific syntax elements) in the bitstream, variable-length bits can be used to signal the corresponding information. Through this, relatively fewer bits can be allocated for low values, and the overall coding efficiency can be improved.
[0092] The decoding device can decode each bin in the bit string from the bitstream based on an entropy coding technique such as CABAC or CAVLC, based on a context model, or based on bypass.
[0093] When a syntax element is decoded based on a context model, a decoding device can receive a bin corresponding to the syntax element via a bitstream, and use the decoding information of the syntax element and the block to be decoded or surrounding blocks, or information of symbols / bins decoded in a previous step to determine a context model. The determined context model can be used to predict the occurrence probability of the received bin and perform arithmetic decoding of the bin to derive the value of the syntax element. Subsequently, the context model of the bin to be decoded next can be updated with the determined context model.
[0094] The context model can be allocated and updated for each bin to be context-coded (entropy-coded), and the context model can be indicated based on ctxIdx or ctxInc. ctxIdx can be derived based on ctxInc. Specifically, for example, a context index (ctxIdx) indicating a context model for each of the bins to be entropy-coded can be derived as the sum of a context index increment (ctxInc) and a context index offset (ctxIdxOffset). Here, the ctxInc can be derived differently for each bin. The ctxIdxOffset can be represented by the lowest value of the ctxIdx. The ctxIdxOffset is generally a value used for distinction from context models for other syntax elements, and a context model for one syntax element can be distinguished / derived based on ctxInc.
[0095] In the entropy encoding procedure, it is possible to determine whether to perform encoding via the normal encoding engine or via the bypass encoding engine, and switch the encoding path. Entropy decoding performs the same process as entropy encoding in reverse.
[0096] On the other hand, for example, when a syntax element is bypass decoded, the decoding device can receive a bin corresponding to the syntax element via the bitstream and decode the input bin by applying a uniform probability distribution. In this case, the procedure for deriving the context model of the syntax element and the procedure for updating the context model applied to the bin after decoding can be omitted.
[0097] As described above, the residual samples can be derived as quantized transform coefficients through the conversion and quantization processes. The quantized transform coefficients can also be called transform coefficients. In this case, the transform coefficients within the block can be signaled in the form of residual information. The residual information can include a residual coding syntax. That is, the encoding device can construct a residual coding syntax with the residual information, encode this, and output it in the form of a bitstream. The decoding device can decode the residual coding syntax from the bitstream and derive the residual (quantized) transform coefficients. The residual coding syntax can include syntax elements indicating whether a transform is applied to the corresponding block, where the position of the last valid transform coefficient within the block is, whether there are valid transform coefficients within the sub-block, what the magnitude / symbol of the valid transform coefficients is, etc., as will be described later.
[0098] On the one hand, when intra prediction is performed, the correlation relationship between samples can be used, and the difference between the original block and the predicted block, that is, the residual, can be obtained. The above-mentioned transformation and quantization can be applied to the residual, and through this, spatial redundancy can be removed. Hereinafter, the encoding method and decoding method using intra prediction will be specifically described.
[0099] Intra prediction refers to a prediction that generates prediction samples for the current block based on reference samples outside the current block within the picture including the current block (hereinafter, the current picture). Here, the reference samples outside the current block can be samples located around the current block. When intra prediction is applied to the current block, neighboring reference samples used for intra prediction of the current block can be derived.
[0100] For example, when the size (width × height) of the current block is nW × nH, the neighboring reference samples of the current block can include samples adjacent to the left boundary of the current block, and a total of 2 × nH samples adjacent to the bottom - left, samples adjacent to the top boundary of the current block, and a total of 2 × nW samples adjacent to the top - right, and 1 sample adjacent to the top - left of the current block. Alternatively, the neighboring reference samples of the current block can also include multiple rows of upper neighboring samples and multiple columns of left neighboring samples. Also, the neighboring reference samples of the current block can include a total of nH samples adjacent to the right boundary of the current block with a size of nW × nH, a total of nW samples adjacent to the bottom boundary of the current block, and 1 sample adjacent to the bottom - right of the current block.
[0101] However, some of the neighboring reference samples of the current block may not have been decoded yet or may not be available. In this case, the decoding device can substitute the unavailable samples with available samples to form the neighboring reference samples used for prediction. Alternatively, the neighboring reference samples used for prediction can be formed through interpolation of the available samples.
[0102] When the neighboring reference samples are derived, (i) a predicted sample can be derived based on the average or interpolation of the neighboring reference samples of the current block, and (ii) a predicted sample can also be derived based on the reference samples that exist in a specific (prediction) direction with respect to the predicted sample among the neighboring reference samples of the current block. In the case of (i), it can be applied when the intra prediction mode is a non-directional mode or a non-angle mode, and in the case of (ii), it can be applied when the intra prediction mode is a directional mode or an angular mode.
[0103] Also, among the neighboring reference samples, a predicted sample can be generated through interpolation between a first neighboring sample located in the prediction direction of the intra prediction mode of the current block and a second neighboring sample located in the opposite direction of the prediction direction, with the predicted sample of the current block as a reference. In the case described above, it can be called linear interpolation intra prediction (LIP). Also, a chroma predicted sample can be generated based on the luma samples using a linear model. In this case, it can be called the LM mode.
[0104] Also, a temporary prediction sample of the current block is derived based on the filtered peripheral reference samples, and at least one reference sample derived by an intra prediction mode among the existing peripheral reference samples, i.e., the unfiltered peripheral reference samples, and the temporary prediction sample are weighted-summed to derive a prediction sample of the current block. In the case described above, it can be called PDPC (Position dependent intra prediction).
[0105] Also, among the multiple reference sample lines around the current block, the reference sample line with the highest prediction accuracy is selected, and a prediction sample is derived using the reference sample located in the prediction direction on the corresponding line. At this time, the encoding of intra prediction can be performed by a method of instructing (signaling) the used reference sample line to the decoding device. In the case described above, it can be called multi-reference line (MRL) intra prediction or intra prediction based on MRL.
[0106] Also, the current block is divided into vertical or horizontal sub-partitions, and intra prediction is performed based on the same intra prediction mode. Peripheral reference samples can be derived and used in units of sub-partitions. That is, in this case, the intra prediction mode for the current block is applied to the sub-partitions in the same way, and by deriving and using peripheral reference samples in units of sub-partitions, the performance of intra prediction can be enhanced as appropriate. Such a prediction method can be called intra sub-partitions (ISP) or intra prediction based on ISP.
[0107] The intra prediction method described above can be classified into an intra prediction mode and can be called an intra prediction type. The intra prediction type can be called by various terms such as an intra prediction technique or an additional intra prediction mode. For example, the intra prediction type (or an additional intra prediction mode, etc.) can include at least one of the above-described LIP, PDPC, MRL, and ISP. A general intra prediction method excluding specific intra prediction types such as the above-described LIP, PDPC, MRL, and ISP can be called a normal intra prediction type. The normal intra prediction type can be generally applied when the above-described specific intra prediction types are not applicable, and prediction may be performed based on the above-described intra prediction mode. On the other hand, post-processing filtering for the derived prediction sample may be performed as necessary.
[0108] On the other hand, in addition to the above-described intra prediction types, matrix-based intra prediction (hereinafter, MIP) can be used as a method for intra prediction. MIP can be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP).
[0109] When MIP is currently applied to a block, i) using the surrounding reference samples for which an averaging procedure has been performed, ii) performing a matrix-vector-multiplication procedure, and iii) further performing a horizontal / vertical interpolation procedure as necessary, a prediction sample for the current block can be derived. The intra prediction mode used for the MIP can be configured differently from the intra prediction modes used in the above-described LIP, PDPC, MRL, and ISP intra predictions and the normal intra prediction.
[0110] The intra prediction mode for MIP can be called the "affine linear weighted intra prediction mode" or the matrix-based intra prediction mode. For example, with the intra prediction mode for MIP, the matrix and offset used in matrix-vector multiplication can be set differently. Here, the matrix can be called an (affine) weighted value matrix, and the offset can be called an (affine) offset vector or an (affine) bias vector. In this document, the intra prediction mode for MIP can be called the MIP intra prediction mode, the linear weighted intra prediction mode, the matrix weighted intra prediction mode, or the matrix based intra prediction mode. Specific MIP methods will be described later.
[0111] The following drawings are created to illustrate a specific example of this document. Since the names of specific devices, specific terms, and names (such as syntax names, etc.) described in the drawings are presented by way of example, the technical features of this document are not limited to the specific names used in the following drawings.
[0112] FIG. 6 shows an example of a video encoding method based on schematic intra prediction to which the embodiments of this document can be applied.
[0113] Referring to FIG. 6, S600 can be executed by the intra prediction unit 222 of the encoding device, and S610 can be executed by the residual processing unit 230 of the encoding device. Specifically, S610 can be executed by the subtraction unit 231 of the encoding device. In S620, the prediction information is derived by the intra prediction unit 222 and can be encoded by the entropy encoding unit 240. In S620, the residual information is derived by the residual processing unit 230 and can be encoded by the entropy encoding unit 240. The residual information is information regarding the residual samples. The residual information can include information regarding the quantized transform coefficients for the residual samples. As described above, the residual samples are derived as transform coefficients via the transform unit 232 of the encoding device, and the transform coefficients can be derived as quantized transform coefficients via the quantization unit 233. The information regarding the quantized transform coefficients can be encoded by the entropy encoding unit 240 via the residual coding procedure.
[0114] The encoding device performs intra prediction on the current block (S600). The encoding device can derive the intra prediction mode / type for the current block and derive the surrounding reference samples of the current block, and generate prediction samples within the current block based on the intra prediction mode / type and the surrounding reference samples. Here, the determination of the intra prediction mode / type, the derivation of the surrounding reference samples, and the generation procedure of the prediction samples may be performed simultaneously, or any one of the procedures may be performed prior to the other procedures.
[0115] On the other hand, although not shown, when a filtering procedure for prediction samples is performed, the intra prediction unit 222 may further include a prediction sample filter unit (not shown). The encoding device can determine the mode / type applied to the current block among a plurality of intra prediction modes / types. The encoding device can compare the RD cost for the intra prediction mode / type and determine the optimal intra prediction mode / type for the current block.
[0116] As described above, the encoding device can also perform a filtering procedure for prediction samples. The filtering of prediction samples may be referred to as post-filtering. By the filtering procedure for prediction samples, some or all of the prediction samples can be filtered. Depending on the case, the filtering procedure for prediction samples may be omitted.
[0117] The encoding device generates a residual sample for the current block based on the (filtered) prediction samples (S610). The encoding device can compare the prediction samples with the original samples of the current block on a phase basis and derive the residual samples.
[0118] The encoding device can encode video information including information related to intra prediction (prediction information) and residual information related to the residual samples (S620). The prediction information can include intra prediction mode information and intra prediction type information. The residual information can include the syntax of residual coding. The encoding device can transform / quantize the residual samples and derive the quantized transform coefficients. The residual information can include information about the quantized transform coefficients.
[0119] The encoding device can output the encoded video information in the form of a bitstream. The output bitstream can be transmitted to the decoding device via a storage medium or a network.
[0120] As described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks). Therefore, the encoding device can re-inverse quantize / inverse transform the quantized transform coefficients to derive (modified) residual samples. The reason for performing inverse quantization / inverse transformation again after converting / quantizing the residual samples in this way is to derive the same residual samples as those derived from the decoding device, as described above. The encoding device can generate a reconstructed block including reconstructed samples for the current block based on the predicted samples and the (modified) residual samples. Based on the reconstructed block, a reconstructed picture for the current picture can be generated. As described above, in-loop filtering procedures and the like can be further applied to the reconstructed picture.
[0121] FIG. 7 shows an example of a video decoding method based on intra prediction to which the embodiments of this document can be applied.
[0122] Referring to FIG. 7, the decoding device can execute operations corresponding to the operations executed by the aforementioned encoding device. S700 to S720 can be executed by the intra prediction unit 331 of the decoding device, and the prediction information of S700 and the residual information of S730 can be obtained from the bitstream by the entropy decoding unit 310 of the decoding device. The residual processing unit 320 of the decoding device can derive residual samples for the current block based on the residual information. Specifically, the inverse quantization unit 321 of the residual processing unit 320 performs inverse quantization based on the quantized transform coefficients derived based on the residual information to derive the transform coefficients, and the inverse transform unit 322 of the residual processing unit performs inverse transform on the transform coefficients to derive residual samples for the current block. S740 can be executed by the addition unit 340 or the restoration unit of the decoding device.
[0123] The decoding device can derive the intra prediction mode / type for the current block based on the received prediction information (intra prediction mode / type information) (S700). The decoding device can derive the surrounding reference samples of the current block (S710). The decoding device can generate prediction samples within the current block based on the intra prediction mode / type and the surrounding reference samples (S720). In this case, the decoding device can perform a filtering procedure for the prediction samples. The filtering of the prediction samples can be referred to as post-filtering. By the filtering procedure of the prediction samples, some or all of the prediction samples can be filtered. Depending on the case, the filtering procedure of the prediction samples can be omitted.
[0124] The decoding device generates residual samples for the current block based on the received residual information (S730). The decoding device can generate restored samples for the current block based on the prediction samples and the residual samples, and derive a restored block including the restored samples (S740). Based on the restored block, a restored picture for the current picture can be generated. As described above, in-loop filtering procedures and the like can be further applied to the restored picture.
[0125] On the other hand, although not shown, when the filtering procedure for the prediction samples described above is performed, the intra prediction unit 331 can further include a prediction sample filter unit (not shown).
[0126] On the other hand, the intra prediction mode can include a non-directional (or non-angular) intra prediction mode and a directional (or angular) intra prediction mode. For example, in the HEVC standard, an intra prediction mode including two non-directional prediction modes and 33 directional prediction modes is used. The non-directional prediction modes can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional prediction modes can include the intra prediction modes numbered 2 to 34. The planar intra prediction mode may be called the planar mode, and the DC intra prediction mode may be called the DC mode.
[0127] Alternatively, in order to capture any edge direction presented in a natural video, the directional intra prediction mode can be extended from the existing 33 to 65 as shown in FIG. 10 described later. In this case, the intra prediction mode can include two non-directional intra prediction modes and 65 directional intra prediction modes. The non-directional intra prediction modes can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional intra prediction modes can include the intra prediction modes numbered 2 to 66. The extended directional intra prediction modes can be applied to blocks of all sizes and to all of the luma and chroma components. However, this is an example, and the embodiments of this document can also be applied when the number of intra prediction modes is different. Optionally, an intra prediction mode numbered 67 may be further used, and the intra prediction mode numbered 67 may indicate an LM (linear model) mode.
[0128] FIG. 8 shows an example of an intra prediction mode to which the embodiments of this document can be applied.
[0129] Referring to FIG. 8, the intra prediction modes having horizontal directionality can be distinguished from the intra prediction modes having vertical directionality, centering on the 34th intra prediction mode having a predicted direction of the upper left diagonal. H and V in FIG. 10 respectively represent horizontal directionality and vertical directionality, and the numbers from -32 to 32 indicate displacements in units of 1 / 32 on the sample grid position. The 2nd to 33rd intra prediction modes have horizontal directionality, and the 34th to 66th intra prediction modes have vertical directionality. The 18th intra prediction mode and the 50th intra prediction mode respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. The 2nd intra prediction mode can be called a lower left diagonal intra prediction mode, the 34th intra prediction mode can be called an upper left diagonal intra prediction mode, and the 66th intra prediction mode can be called an upper right diagonal intra prediction mode.
[0130] The intra prediction mode information can include, for example, flag information (e.g., intra_luma_mpm_flag) indicating whether the MPM (most probable mode) is applied to the current block or the remaining mode is applied. At this time, when the MPM is applied to the current block, the prediction mode information can further include index information (e.g., intra_luma_mpm_idx) indicating one of the candidates of the intra prediction mode (MPM candidates). The candidates of the intra prediction mode (MPM candidates) can be composed of an MPM candidate list or an MPM list. Also, when the MPM is not applied to the current block, the intra prediction mode information can further include remaining mode information (e.g., intra_luma_mpm_remainder) indicating one of the remaining intra prediction modes excluding the candidates of the intra prediction mode (MPM candidates). The decoding device can determine the intra prediction mode of the current block based on the intra prediction mode information.
[0131] Also, the intra prediction type information can be embodied in various forms. As an example, the intra prediction type information can include index information of the intra prediction type indicating one of the intra prediction types. As another example, the intra prediction type information can include reference sample line information (e.g., intra_luma_ref_idx) indicating whether the MRL is applied to the current block and, if the MRL is applied, which reference sample line is used, ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block, ISP type information (e.g., intra_subpartitions_split_flag) indicating the split type of the subpartition when the ISP is applied, flag information indicating whether PDCP is applicable, or at least one of flag information indicating whether LIP is applicable. Also, the intra prediction type information can include an MIP flag indicating whether the MIP is applied to the current block.
[0132] The above-mentioned intra prediction mode information and / or intra prediction type information can be encoded / decoded through the coding method described in this document. For example, the above-mentioned intra prediction mode information and / or intra prediction type information can be encoded / decoded through entropy coding (e.g., CABAC, CAVLC) based on truncated (rice) binary code.
[0133] On the one hand, when intra prediction is applied, the intra prediction mode applied to the current block can be determined using the intra prediction modes of the surrounding blocks. For example, the decoding device can select one of the mpm (most probable mode) candidates in the mpm list derived based on the intra prediction modes of the surrounding blocks (e.g., the left and / or upper surrounding blocks) of the current block and additional candidate modes according to the received mpm index, or can select one of the remaining intra prediction modes not included in the mpm candidates (and the planar mode) based on the remaining intra prediction mode information. The mpm list can be configured to include the planar mode as a candidate or not. For example, when the mpm list includes the planar mode as a candidate, the mpm list can have 6 candidates, and when the mpm list does not include the planar mode as a candidate, the mpm list can have 5 candidates. When the mpm list does not include the planar mode as a candidate, a not planar flag (e.g., intra_luma_not_planar_flag) indicating whether the intra prediction mode of the current block is not the planar mode can be signaled. For example, if the mpm flag is signaled first, the mpm index and the not planar flag can be signaled when the value of the mpm flag is 1. Also, the mpm index can be signaled when the value of the not planar flag is 1. Here, the configuration that the mpm list does not include the planar mode as a candidate is for signaling the flag (not planar flag) first and checking whether it is the planar mode first because the planar mode is always considered as an mpm rather than not being an mpm.
[0134] For example, whether the intra prediction mode currently applied to a block is among the mpm candidates (and the planar mode) or among the remaining modes can be indicated based on the mpm flag (e.g., intra_luma_mpm_flag). A value of 1 for the mpm flag can indicate that the intra prediction mode for the current block is within the mpm candidates (and the planar mode), and a value of 0 for the mpm flag can indicate that the intra prediction mode for the current block is not within the mpm candidates (and the planar mode). A value of 0 for the not planar flag (e.g., intra_luma_not_planar_flag) can indicate that the intra prediction mode for the current block is the planar mode, and a value of 1 for the not planar flag can indicate that the intra prediction mode for the current block is not the planar mode. The mpm index can be signaled in the form of a syntax element of mpm_idx or intra_luma_mpm_idx, and the remaining intra prediction mode information can be signaled in the form of a syntax element of rem_intra_luma_pred_mode or intra_luma_mpm_remainder. For example, the remaining intra prediction mode information can index the remaining intra prediction modes not included in the mpm candidates (and the planar mode) among the overall intra prediction modes in ascending order of the prediction mode numbers and point to one of them. The intra prediction mode can be the intra prediction mode for the luma component (samples). Hereinafter, the intra prediction mode information can include at least one of the mpm flag (e.g., intra_luma_mpm_flag), the not planar flag (e.g., intra_luma_not_planar_flag), the mpm index (e.g., mpm_idx or intra_luma_mpm_idx), and the remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder).In this document, the mpm list can be referred to by various terms such as the mpm candidate list, candidate mode list (candModeList), candidate intra prediction mode list, etc.
[0135] Generally, when a video is block - divided, the current block to be coded and the surrounding blocks will have similar video characteristics. Therefore, the current block and the surrounding blocks are likely to be identical to each other or have similar intra - prediction modes. Thus, the encoder can use the intra - prediction mode of the surrounding blocks to encode the intra - prediction mode of the current block. For example, the encoder / decoder can construct an MPM (most probable modes) list for the current block. The MPM list can also be referred to as the MPM candidate list. Here, MPM can mean a mode that is used to improve coding efficiency by considering the similarity between the current block and the surrounding blocks during the coding of the intra - prediction mode.
[0136] FIG. 9 shows an example of an intra - prediction method based on the MPM mode in an encoding device to which the embodiments of this document can be applied.
[0137] Referring to FIG. 9, the encoding device constructs an MPM list for the current block (S900). The MPM list can include candidate intra - prediction modes (MPM candidates) that are highly likely to be applied to the current block. The MPM list can also include the intra - prediction modes of the surrounding blocks and can further include specific intra - prediction modes by a predetermined method. The specific method for constructing the MPM list will be described later.
[0138] The encoding device determines the intra prediction mode of the current block (S910). The encoding device can perform prediction based on various intra prediction modes, and can determine the optimal intra prediction mode based on rate-distortion optimization (RDO) based on this. In this case, the encoding device can also determine the optimal intra prediction mode using only the MPM candidates configured in the MPM list and the planar mode, or can also determine the optimal intra prediction mode using not only the MPM candidates configured in the MPM list and the planar mode but also the remaining intra prediction modes.
[0139] Specifically, for example, if the intra prediction type of the current block is not the normal intra prediction type but a specific type (e.g., LIP, MRL, or ISP), the encoding device can consider only the MPM candidates and the planar mode as candidates for the intra prediction mode for the current block and determine the optimal intra prediction mode. That is, in this case, the intra prediction mode for the current block can be determined only from the MPM candidates and the planar mode, and in this case, the mpm flag may not be encoded / signaled. The decoding device can, in this case, assume that the mpm flag is 1 without receiving separate signaling.
[0140] Generally, when the intra prediction mode of the current block is not the planar mode but one of the MPM candidates in the MPM list, the encoding device generates an mpm index (mpm idx) indicating one of the MPM candidates. If the intra prediction mode of the current block is not in the MPM list, remaining intra prediction mode information indicating the same mode as the intra prediction mode of the current block is generated from among the remaining intra prediction modes not included in the MPM list (and the planar mode).
[0141] The encoding device can encode the intra prediction mode information and output it in the form of a bitstream (S920). The intra prediction mode information can include the mpm flag, not planar flag, mpm index, and / or remaining intra prediction mode information described above. Generally, the mpm index and the remaining intra prediction mode information are in an alternative relationship and are not signaled simultaneously when indicating the intra prediction mode for one block. That is, either the value 1 of the mpm flag and the not planar flag or the mpm index are both signaled, or the value 0 of the mpm flag and the remaining intra prediction mode information are both signaled. However, as described above, when a specific intra prediction type is applied to the current block, the mpm flag may not be signaled, and only the not planar flag and / or mpm index may be signaled. That is, in this case, the intra prediction mode information may include only the not planar flag and / or mpm index.
[0142] FIG. 10 shows an example of an intra prediction method based on the MPM mode in a decoding device to which the embodiments of this document can be applied. The decoding device in FIG. 10 can determine the intra prediction mode corresponding to the intra prediction mode information determined and signaled by the encoding device in FIG. 9.
[0143] Referring to FIG. 10, the decoding device acquires the intra prediction mode information from the bitstream (S1000). The intra prediction mode information can include at least one of the mpm flag, not planar flag, mpm index, and remaining intra prediction mode as described above.
[0144] The decoding device constructs an MPM list (S1010). The MPM list is constructed in the same way as the MPM list constructed by the encoding device. That is, the MPM list may include the intra prediction modes of the surrounding blocks and may further include specific intra prediction modes by a predetermined method. The specific method for constructing the MPM list will be described later.
[0145] Although S1010 is shown to be performed after S1000, this is an illustration, and S1010 may be performed before S1000 or may be performed simultaneously.
[0146] The decoding device determines the intra prediction mode of the current block based on the MPM list and the intra prediction mode information (S1020).
[0147] As an example, when the value of the mpm flag is 1, the decoding device can derive the planar mode as the intra prediction mode of the current block or derive the candidate pointed to by the mpm index from among the MPM candidates in the MPM list (based on the not planar flag). Here, the MPM candidates may refer only to the candidates included in the MPM list, or may include not only the candidates included in the MPM list but also the planar mode that can be applied when the value of the mpm flag is 1.
[0148] As another example, when the value of the mpm flag is 0, the decoding device can derive the intra prediction mode pointed to by the remaining intra prediction mode information as the intra prediction mode of the current block from among the remaining intra prediction modes not included in the MPM list and the planar mode.
[0149] As another example, when the intra prediction type of the current block is a specific type (e.g., LIP, MRL, or ISP, etc.), the decoding device can also derive the candidate indicated by the mpm index in the planner mode or the MPM list as the intra prediction mode of the current block without checking the mpm flag.
[0150] In constructing the MPM list, as one embodiment, the encoding device / decoding device can derive the left mode, which is the candidate intra prediction mode for the left peripheral block of the current block, and the upper mode, which is the candidate intra prediction mode for the upper peripheral block of the current block. Here, the left peripheral block can indicate the lowermost peripheral block among the left peripheral blocks adjacent to the left side of the current block, and the upper peripheral block can indicate the rightmost peripheral block among the upper peripheral blocks adjacent to the upper side of the current block. For example, when the size of the current block is W×H and the x component of the top-left sample position of the current block is xN and the y component is yN, the left peripheral block can be a block including the sample at the (xN - 1, yN + H - 1) coordinates, and the upper peripheral block can be a block including the sample at the (xN + W - 1, yN - 1) coordinates.
[0151] For example, if the encoding device / decoding device has an available left peripheral block and intra prediction is applied to the left peripheral block, the intra prediction mode of the left peripheral block can be derived as the left candidate intra prediction mode (i.e., the left mode). If the decoding device has an available upper peripheral block, intra prediction is applied to the upper peripheral block, and the upper peripheral block is included in the current CTU, the intra prediction mode of the upper peripheral block can be derived as the upper candidate intra prediction mode (i.e., the upper mode). Alternatively, if the encoding device / decoding device does not have an available left peripheral block or intra prediction is not applied to the left peripheral block, the planar mode can be derived as the left mode. If the decoding device does not have an available upper peripheral block, intra prediction is not applied to the upper peripheral block, or the upper peripheral block is not included in the current CTU, the planar mode can be derived as the upper mode.
[0152] Based on the left mode derived from the left peripheral block and the upper mode derived from the upper peripheral block, the encoding device / decoding device can derive the candidate intra prediction mode for the current block and construct the MPM list. At this time, the MPM list can include the left mode and the upper mode, and can also further include a specific intra prediction mode by a predetermined method.
[0153] Also, when constructing the MPM list, a general intra prediction method (normal intra prediction type) can be applied, or a specific intra prediction type (e.g., MRL, ISP) can be applied. In this document, not only such a general intra prediction method (normal intra prediction type) but also when a specific intra prediction type (e.g., MRL, ISP) is applied, a method for constructing the MPM list is proposed and will be described later.
[0154] On one hand, the intra prediction mode used in the aforementioned MIP is not an existing directional mode, but can indicate the matrix and offset used for intra prediction. That is, through the intra mode for MIP, the matrix and offset for intra prediction can be derived. In this case, when deriving the intra mode for generating the aforementioned normal intra prediction or MPM list, the intra prediction mode of the block predicted as MIP can be set to a preset mode, for example, the planar mode or the DC mode. Alternatively, according to another example, based on the block size, the intra mode for MIP can also be mapped to the planar mode, the DC mode, or the directional intra mode.
[0155] Hereinafter, we will look at MIP (Matrix based intra prediction), which is one method of intra prediction.
[0156] As described above, matrix-based intra prediction (Matrix based intra prediction, hereinafter referred to as MIP) can be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). To predict the samples of a rectangular block with width (W) and height (H), MIP uses one H line of the left boundary samples of the restored periphery of the block and one W line of the upper boundary samples of the restored periphery of the block as input values. If the restored samples are not available, reference samples can be generated by the interpolation method applied in normal intra prediction.
[0157] FIG. 11 is a diagram for explaining the generation procedure of prediction samples based on MIP according to an example. Referring to FIG. 11, the MIP procedure is explained as follows.
[0158] 1. Averaging process
[0159] Among the boundary samples, when W = H = 4, 4 samples, and in all other cases, 8 samples are extracted by the averaging process.
[0160] 2. Matix vector multiplication process
[0161] The multiplication of the matrix vector is performed with the averaged samples as the input, and subsequently an offset is added. Through such operations, reduced predicted samples for the subsampled sample set within the original block can be derived.
[0162] 3. (linear)Interpolation process
[0163] The predicted samples at the remaining positions are generated from the predicted samples of the subsampled sample set by linear interpolation, which is a single step linear interpolation in each direction.
[0164] The matrix and the offset vector required to generate the prediction block or predicted samples can be selected from three sets S0, S1, S2 for the matrix.
[0165] Set S0 consists of 16 matrices A0 i , where i ∈ {0, …, 15}, and each matrix can consist of 16 rows, 4 columns, and 16 offset vectors b0 i , where i ∈ {0, …, 15}. The matrices and offset vectors of set S0 can be used for blocks of size 4×4. According to another example, set S0 can also include 18 matrices.
[0166] Set S1 consists of 8 matrices A1 i, which is composed of \(i\in\{0,\ldots,7\}\), and each matrix has 16 rows, 8 columns, and 8 offset vectors \(b1\) i , which can be composed of \(i\in\{0,\ldots,7\}\). According to another example, set \(S1\) can also include 6 matrices. The matrices and offset vectors of set \(S1\) can be used for blocks of sizes \(4\times8\), \(8\times4\), and \(8\times8\). Alternatively, the matrices and offset vectors of set \(S1\) can be used for blocks of sizes \(4\times H\) or \(W\times4\).
[0167] Finally, set \(S2\) is composed of 6 matrices \(A2\) i , which is composed of \(i\in\{0,\ldots,5\}\), and each matrix has 64 rows, 8 columns, and 6 offset vectors \(b2\) i , which can be composed of \(i\in\{0,\ldots,5\}\). The matrices and offset vectors of set \(S2\), or a part of them, can be used for all other block forms where sets \(S0\) and \(S1\) are not applicable. For example, the matrices and offset vectors of set \(S2\) can be used for operations on blocks with a height and width of 8 or more.
[0168] The number of multiplication operations required for matrix-vector multiplication is always less than or equal to \(4\times W\times H\). That is, in MIP mode, a maximum of 4 multiplications per sample are required.
[0169] Hereinafter, the general MIP procedure will be outlined. The remaining blocks not described below can be processed in any of the four cases described.
[0170] Figures 12 to 15 are diagrams showing the MIP procedure according to the block size. Figure 12 is a diagram showing the MIP procedure for a \(4\times4\) block, Figure 13 is a diagram showing the MIP procedure for an \(8\times8\) block, Figure 14 is a diagram showing the MIP procedure for an \(8\times4\) block, and Figure 15 is a diagram showing the MIP procedure for a \(16\times16\) block.
[0171] As shown in FIG. 12, when a 4×4 block is given, the MIP takes the average of two samples along each axis of the boundary. As a result, four input samples become the input values for the matrix-vector multiplication, and the matrix is taken from set S0. When an offset is added, 16 final predicted samples are generated. For a 4×4 block, linear interpolation is not necessary to generate the predicted samples. Therefore, for each sample, a multiplication of (4×16) / (4×4)=4 can be performed.
[0172] As shown in FIG. 13, when an 8×8 block is given, the MIP takes the average of four samples along each axis of the boundary. As a result, eight input samples become the input values for the matrix-vector multiplication, and the matrix is taken from set S1. Sixteen samples are generated at odd positions by the matrix-vector multiplication.
[0173] For an 8×8 block, in order to generate the predicted samples, a multiplication of (8×16) / (8×8)=2 is performed for each sample. After adding the offset, the samples are interpolated vertically using the reduced upper boundary samples and horizontally using the original left boundary samples. In this case, since no multiplication operation is required in the interpolation procedure, a total of two multiplication operations per sample are required for the MIP.
[0174] As shown in FIG. 14, when an 8×4 block is given, the MIP takes the average of four samples along the horizontal axis of the boundary and uses the four sample values of the left boundary for the vertical axis. As a result, eight input samples become the input values for the matrix-vector multiplication, and the matrix is taken from set S1. Sixteen samples are generated at the odd horizontal positions and the corresponding vertical positions by the matrix-vector multiplication.
[0175] In the case of an 8×4 block, in order to generate prediction samples, a multiplication of (8×16) / (8×4) = 4 is performed for each sample. After adding the offset, interpolation is performed horizontally using the original left boundary sample. In this case, since the interpolation procedure does not require a multiplication operation, a total of 4 multiplication operations per sample are required for MIP.
[0176] As shown in FIG. 15, when a 16×16 block is given, MIP takes the average of 4 samples along each axis. As a result, 8 input samples become the input values of the matrix-vector multiplication, and the matrix is taken from set S2. By the matrix-vector multiplication, 64 samples are generated at odd positions. In the case of a 16×16 block, a multiplication of (8×64) / (16×16) = 2 is performed for each sample to generate prediction samples. After adding the offset, the samples are interpolated vertically using 8 reduced upper boundary samples, and interpolated horizontally using the original left boundary sample. In this case, since the interpolation procedure does not require a multiplication operation, a total of 2 multiplication operations per sample are required for MIP.
[0177] For larger blocks, the MIP procedure is essentially the same as the procedure described in detail, and it can be easily confirmed that the number of multiplications per sample is less than 4.
[0178] In the case of a W×8 block where the width is greater than 8 (W>8), only horizontal interpolation is required because samples are generated at odd horizontal and each vertical position. In this case, a multiplication of (8×64) / (W×8) = 64 / W is performed for each sample for the prediction operation of the reduced samples. When W = 16, no additional multiplication is required for linear interpolation, and when W>16, the number of additional multiplications per sample required for linear interpolation is less than 2. That is, the total number of multiplications per sample is less than or equal to 4.
[0179] Also, in the case of a W×4 block where the width is greater than 4 (W>4), a matrix A is generated by omitting all rows corresponding to odd entries along the horizontal axis of the downsampled block. k Let it be so. Therefore, the size of the output is 32, and only horizontal interpolation is performed. For the prediction operation of the downsampled samples, a multiplication of (8×32) / (W×4)=64 / W per sample is performed. When W = 16, no additional multiplication is required, and when W>16, the number of additional multiplications per sample required for linear interpolation is less than 2. That is, the total number of multiplications per sample is less than or equal to 4.
[0180] When the matrix is prefixed, it can be processed thereby.
[0181] FIG. 16 is a diagram for explaining the boundary averaging procedure among the MIP procedures. With reference to FIG. 16, the averaging procedure will be specifically described.
[0182] According to the averaging procedure, averaging is applied to each boundary, that is, the left boundary or the upper boundary. Here, the boundary indicates the peripheral reference samples adjacent to the boundary of the current block as shown in FIG. 16. For example, the left boundary (bdry left ) indicates the left peripheral reference samples adjacent to the left boundary of the current block, and the upper boundary (bdry top ) indicates the upper peripheral reference samples adjacent to the upper side.
[0183] When the current block is a 4×4 block, the size of each boundary can be reduced to 2 samples through the averaging procedure. If the current block is not a 4×4 block, the size of each boundary can be reduced to 4 samples through the averaging procedure.
[0184] The first step of the averaging procedure is to reduce the input boundaries (bdry left and bdry top ) to smaller boundaries. JPEG0007701520000001.jpg1151 JPEG0007701520000002.jpg964 consists of 2 samples in the case of a 4×4 block and 4 samples in all other different cases.
[0185] For a 4×4 block, for 0 ≦ i < 2 JPEG0007701520000003.jpg1020 can be expressed by the following mathematical formula, JPEG0007701520000004.jpg828 can also be defined similarly.
[0186]
Number
[0187] On the other hand, for 0 ≦ i < 4, when the width of the block is W = 4×2 k is given, JPEG0007701520000006.jpg838 can be expressed by the following mathematical formula, JPEG0007701520000007.jpg940 can also be defined similarly.
[0188]
Number
[0189] Two reduced boundaries JPEG0007701520000009.jpg1281 is concatenated to the reduced boundary vector bdry red so bdry red has a size of 4 in the case of a 4×4 form block and a size of 8 in the case of all other blocks.
[0190] When referring to "mode" as the MIP mode, the reduced boundary vector bdry redThe range of the MIP mode value (mode) can be defined based on the block size and the intra_mip_transposed_flag value as in the following formula.
[0191]
Number
[0192] In the above formula, intra_mip_transposed_flag can be referred to as MIP transpose, and such flag information can indicate whether the downsampled prediction samples are transposed or not. The semantics for such syntactic elements can be expressed as "intra_mip_transposed_flag[x0][y0] specifies whether the input vector for matrix-based intra prediction mode for luma samples is transposed or not."
[0193] Finally, for interpolation of the subsampled prediction samples, in a large block, a second version of the averaged boundary is required. That is, if the smaller value of the width and height is greater than 8 JPEG0007701520000011.jpg11101, If the width is equal to or greater than the height (W≧H), W = 8 * 2 l For 0 ≦ i < 8, JPEG0007701520000012.jpg12115 can be defined as in the following formula. Also, if the smaller value of the width and height is greater than 8 JPEG0007701520000013.jpg10115, if the height is greater than the width (H>W) JPEG0007701520000014.jpg990 can also be defined in the same way.
[0194]
Number
[0195] Next, let's look at the procedure for generating reduced prediction samples by matrix-vector multiplication.
[0196] Reduced input vector bdry red One of them generates a reduced prediction sample JPEG0007701520000016.jpg1048. The prediction sample is a signal for a downsampled block of width W red and height H red Here, W red and H red are defined as follows.
[0197]
Number
[0198] Reduced prediction sample JPEG0007701520000018.jpg1269 can be calculated by adding an offset after the matrix-vector multiplication operation and can be derived through the following formula.
[0199]
Number
[0200] Here, A is a matrix with W red ×h red rows and 4 columns when W and H are 4 (W = H = 4) and 8 columns in all other cases, and b is a vector of size W red ×h red .
[0201] Matrix A and vector b are selected from sets S0, S1, S2 as follows, and the index idx = idx(W, H) can be defined as in Equation 7 or Equation 8.
[0202]
Number
[0203]
Number
[0204] If idx is 1 or less (idx ≦ 1), or if idx is 2 and the smaller value of W and H is greater than 4 JPEG0007701520000022.jpg1196, A is set to JPEG0007701520000023.jpg1089. JPEG0007701520000024.jpg1184, b is set to JPEG0007701520000025.jpg1155. If idx is 2 and the smaller value of W and H is 4 When JPEG0007701520000026.jpg1098 and W is 4, A becomes the matrix obtained by removing each row of JPEG0007701520000027.jpg1097 corresponding to odd x - coordinates within the down - sampled block. Or, if H is 4, A becomes the matrix obtained by removing each column of JPEG0007701520000028.jpg10133 corresponding to odd y - coordinates within the down - sampled block. When JPEG0007701520000026.jpg1098 and W is 4, A becomes the matrix obtained by removing each row of JPEG0007701520000027.jpg1097 corresponding to odd x - coordinates within the down - sampled block. Or, if H is 4, A becomes the matrix obtained by removing each column of JPEG0007701520000028.jpg10133 corresponding to odd y - coordinates within the down - sampled block. When JPEG0007701520000026.jpg1098 and W is 4, A becomes the matrix obtained by removing each row of JPEG0007701520000027.jpg1097 corresponding to odd x - coordinates within the down - sampled block. Or, if H is 4, A becomes the matrix obtained by removing each column of JPEG0007701520000028.jpg10133 corresponding to odd y - coordinates within the down - sampled block.
[0205] Finally, the down - sized prediction sample can be replaced by its transpose in Equation 9.
[0206]
Number
[0207] For the multiplication required for the calculation of JPEG0007701520000030.jpg1196, when W = H = 4, since A is composed of 4 columns and 16 rows, it is 4. In all other cases, since A is composed of 8 columns and W red ×h red rows, so For the calculation of JPEG0007701520000031.jpg10109, it can be confirmed that a maximum of 4 multiplications per sample are required.
[0208] Figure 17 is a diagram for explaining linear interpolation in the MIP procedure. Referring to Figure 17, the linear interpolation procedure is specifically described as follows.
[0209] The interpolation procedure may be referred to as a linear interpolation or a bilinear interpolation procedure. The interpolation procedure can include two steps: 1) vertical interpolation and 2) horizontal interpolation, as shown.
[0210] If W ≥ H, vertical linear interpolation can be applied first, followed by horizontal linear interpolation. If W < H, horizontal linear interpolation can be applied first, followed by vertical linear interpolation. In the case of a 4×4 block, the interpolation procedure may be omitted.
[0211] For a W×H block that is JPEG0007701520000032.jpg11119, the predicted samples are W red ×H red of the reduced predicted samples above derived from JPEG0007701520000033.jpg11113. Depending on the form of the block, linear interpolation is performed vertically, horizontally, or in both directions. When linear interpolation is applied in both directions, if W < H, it is applied first horizontally, otherwise, it is applied first vertically.
[0212] It is 1265 in JPEG0007701520000034.jpg. In the case of a W×H block where W≧H, it can be considered that there is no loss of generality. Then, one-dimensional linear interpolation is performed as follows. If there is no loss of generality, linear interpolation in the vertical direction is sufficiently explained.
[0213] First, the reduced prediction samples are extended upward by the boundary signal. The coefficient of vertical upsampling defines 1079 in JPEG0007701520000035.jpg, and sets 1186 in JPEG0007701520000036.jpg. The extended reduced prediction samples can be set as in the following formula.
[0214]
Equation
[0215] After that, vertical linear interpolation prediction samples can be generated from such extended reduced prediction samples by the following formula.
[0216]
Equation
[0217] Here, x can be 0≦x<W red and y can be 0≦y<H red and k can be 0≦k<U ver
[0218] Next, let's look at a method to reduce the complexity and maximize the performance for the MIP technique. The embodiments described later may be executed independently or in combination.
[0219] On the one hand, when MIP is applied to the current block, an MPM list for the current block to which the MIP is applied can be separately configured. The MPM list can be called by various names such as MIP MPM list (or LWIP MPM list, candLwipModeList, etc.) in order to distinguish it from the MPM list when ALWIP is not applied to the current block. Hereinafter, for the sake of distinction, it is expressed as MIP MPM list, but this can be called an MPM list.
[0220] The MIP MPM list can include n candidates. For example, n can be 3. The MIP MPM list can be configured based on the left surrounding block and the upper surrounding block of the current block. Here, the left surrounding block can indicate the uppermost block among the surrounding blocks adjacent to the left boundary of the current block. Also, the upper surrounding block can indicate the leftmost block among the surrounding blocks adjacent to the upper boundary of the current block.
[0221] For example, when MIP is applied to the left surrounding block, the first candidate intra prediction mode (or candLwipModeA) can be set in the same way as the MIP mode of the left surrounding block. Also, for example, when MIP is applied to the upper surrounding block, the second candidate intra prediction mode (or candLwipModeB) can be set in the same way as the prediction mode of the MIP mode of the upper surrounding block.
[0222] On the one hand, the left peripheral block and the upper peripheral block can be coded based on an intra prediction that is not MIP. That is, when coding the left peripheral block or the upper peripheral block, other intra prediction types that are not MIP can be applied. In this case, it is not appropriate to directly use the general intra prediction mode number of the peripheral block (left peripheral block / upper peripheral block) where MIP is not applied as the candidate intra mode for the current block where MIP is applied. Therefore, in this case, as an example, the MIP mode of the peripheral block (left peripheral block / upper peripheral block) where MIP is not applied can be regarded as a prediction mode of the MIP mode with a specific value (for example, 0, 1, or 2, etc.). Alternatively, as another example, the general intra prediction mode of the peripheral block (left peripheral block / upper peripheral block) where MIP is not applied can be mapped to the MIP mode based on a predetermined mapping table and used for the construction of the MIP MPM list. In this case, the mapping can be executed based on the block size type of the current block.
[0223] Also, even if MIP is applied, depending on whether the peripheral block (for example, the left peripheral block / upper peripheral block) is not available (for example, located outside the current picture, located outside the current tile / tile group, etc.), a MIP mode that is not available for the current block can also be used according to the block size type. In this case, for the first candidate and / or the second candidate, a specific MIP mode defined in advance can be used as the first candidate intra prediction mode or the second candidate intra prediction mode. Also, for the third candidate, a specific MIP prediction mode defined in advance can be used as the third candidate intra prediction mode.
[0224] On the other hand, the existing MIP mode is classified into a non-MPM mode and an MPM mode in the same way as the derivation method of the existing intra prediction mode, and an MPM flag is sent. Based on the MPM mode or the non-MPM mode, the MIP mode of the current block is coded.
[0225] According to one example, for a block to which the MIP technique is applied, a structure can be proposed that directly codes the MIP mode without distinguishing between the MPM mode and the non-MPM mode. In such a video coding structure, a complex syntax structure can be simplified. Also, since the actual occurrence frequency for the MIP mode is relatively uniformly distributed for each mode and is clearly different from the occurrence frequency shown in the existing intra modes, the efficiency of encoding and decoding MIP mode information can be maximized through the proposed coding structure.
[0226] The video information transmitted and received for MIP according to this embodiment is as follows. The syntax described later can be included in the video / video information transmitted from the encoding device to the decoding device, can be configured / encoded by the encoding device, and can be signaled to the decoding device in the form of a bitstream. The decoding device can parse / decode the information (syntax elements) included according to the conditions / order disclosed in the syntax.
[0227]
Table 1
[0228] As shown in Table 1, the syntax intra_mip_flag and intra_mip_mode_idx for the MIP mode for the current block can be included in the syntax information for the coding unit and signaled.
[0229] On the other hand, sps_mip_enabled_flag, which is flag information indicating whether matrix-based intra prediction (MIP) can be applied to the current block, can be further received, whereby intra_mip_flag can be signaled.
[0230] The flag information sps_mip_enabled_flag can be signaled via the syntax information of the Sequence Parameter Set (SPS).
[0231] If intra_mip_flag is 1, it indicates that the intra prediction type for luma samples is matrix-based intra prediction, and if its value is 0, it indicates that the intra prediction type for luma samples is not matrix-based intra prediction.
[0232] When intra_mip_flag is 1 and signaled, intra_mip_mode_idx indicates the matrix-based intra prediction mode for luma samples. Such a matrix-based intra prediction mode can indicate the matrix and offset or matrix for MIP as described above.
[0233] Also, by way of an example, flag information, such as intra_mip_transposed_flag, can be further signaled via the syntax of the coding unit to indicate whether the input vector for matrix-based intra prediction is transposed. When intra_mip_transposed_flag is 1, the input vector for matrix-based intra prediction is transposed, and such flag information can reduce the number of matrices for matrix-based intra prediction.
[0234] On the other hand, intra_mip_mode_idx can be encoded and decoded in a Truncated Binarization manner as shown in the following table.
[0235]
Table 2
[0236] As shown in Table 2, intra_mip_flag is binary-coded in a Fixed Length Code, while intra_mip_mode_idx is binary-coded in a Truncated Binary Coding method, and the maximum length (cMax) of the binary coding can be set according to the size of the coding block. The maximum length of the binary coding is set to 34 when the width and height of the coding block are 4 (cbWidth == 4 && cbHeight == 4), and otherwise, it can be set to 18 or 10 depending on whether the width and height of the coding block are 8 or less ((cbWidth <= 8 && cbHeight <= 8)?).
[0237] On the other hand, intra_mip_mode_idx can be coded in a bypass method instead of being based on a context model. By coding in a bypass method, the coding speed and efficiency can be increased.
[0238] In another example, when intra_mip_mode_idx is binary-coded in a Truncated Binary Coding method, the maximum length of the binary coding is as shown in the following table.
[0239]
Table 3
[0240] As shown in Table 3, the maximum length of the binary evolution for intra_mip_mode_idx is set to 15 when the width and height of the coding block are 4 ((cbWidth == 4 && cbHeight == 4)), and otherwise, it can be set to 7 when the width or height of the coding block is 4 ((cbWith == 4 || cbHeight == 4)) or the width and height of the coding block are 8 (cbWith == 8 && cbHeight == 8), and it can be set to 5 when the width or height of the coding block is 4 ((cbWith == 4 || cbHeight == 4)) or the width and height of the coding block are not 8 (cbWith == 8 && cbHeight == 8).
[0241] In another example, intra_mip_mode_idx can be encoded with a Fixed Length Code. At this time, the number of MIP modes available to improve the encoding efficiency is limited to an exponential number of 2 for each block size (e.g., A = 2 K1 - 1, B = 2 K2 - 1, 2 K3 - 1, where K1, K2, K3 are positive constants).
[0242] Shown in a table, it is as follows.
[0243]
Table 4
[0244] In Table 4, when K1 = 5, K2 = 4, and K3 = 3, intra_mip_mode can be evolved in binary as follows.
[0245]
Table 5
[0246] Alternatively, by way of example, if K1 is set to 4, the maximum length of the binary evolution for the intra_mip_mode of a block where the width and height of the coding block are 4 can be set to 15. Also, if K2 is set to 3, when the width or height of the coding block is 4, or when the width and height of the coding block are 8, the maximum length of the binary evolution can be set to 7.
[0247] On the other hand, by way of example, a method can be proposed to use MIP only for specific blocks where the MIP technique can be efficiently applied. When the method according to this embodiment is applied, the number of matrix vectors required for MIP decreases, and the memory required to store the matrix vectors can be significantly reduced (by 50%). Despite such an effect, the coding efficiency is maintained almost constant (less than 0.1%).
[0248] The syntax including the specific conditions for applying the MIP technique according to this embodiment is as follows in the following table.
[0249]
Table 6
[0250] As shown in Table 6, a condition (cbWidth>K1||cbHeight>K2) is added so that MIP is applied only to large blocks, and the size of the block can be determined by preset values (K1 and K2). The reason for applying MIP only to large blocks is that the coding efficiency of MIP is shown in relatively large blocks.
[0251] The following table shows an example where K1 and K2 are predefined as 8 in Table 6.
[0252]
Table 7
[0253] The semantics for intra_mip_flag and intra_mip_mode_idx in Tables 6 and 7 are the same as those in Table 1.
[0254] On the other hand, when intra_mip_mode_idx has 11 possible modes, it can be encoded in a truncated binary evolution (cMax = 10) manner as follows.
[0255] [Table 8]
[0256] Alternatively, when the available MIP modes are limited to 8, intra_mip_mode_idx[x0][y0] can be encoded with a fixed-length code as follows.
[0257] [Table 9]
[0258] Both Table 8 and Table 9 can code intra_mip_mode_idx in a bypass manner.
[0259] k ), and an MIP technique that can apply the offset vector (b k ) to small blocks as well as the weight matrix (A
[0260] FIG. 18 is a diagram for explaining an MIP technique according to an example of this document.
[0261] As shown, (a) of FIG. 18 shows the operation of the matrix and the offset vector for the index i of the large block, and (b) of FIG. 18 shows the operation of the sampled matrix and the offset vector applied to the small block.
[0262] As shown in FIG. 18, the subsampled weighted value matrix (Sub(A k )) obtained by subsampling the weighted value matrix used for the large block and the offset vector (Sub(b k )) obtained by subsampling the offset vector used for the large block can be regarded as the weighted value matrix and the offset vector for the small block respectively, and the existing MIP procedure can be applied.
[0263] Here, subsampling may be applied to only one of the horizontal and vertical directions, or may be applied to both directions. In particular, the subsampling factor (for example, 1 out of 2, or 1 out of 4) and the sampling direction for vertical or horizontal can be set based on the width (Width) and height (Height) of the corresponding block.
[0264] Also, the number of intra prediction modes for MIP to which this embodiment is applied can be set differently based on the size of the current block. For example, i) when the height and width of the current block (coding block or transform block) are both 4, 35 intra prediction modes (that is, intra prediction modes 0 to 34) may be available, ii) when the height and width of the current block are both 8 or less, 19 intra prediction modes (that is, intra prediction modes 0 to 18) may be available, and iii) in other cases, 11 intra prediction modes (that is, intra prediction modes 0 to 10) may be available.
[0265] For example, when the height and width of the current block are both 4, it is called block size type 0. When the height and width of the current block are both 8 or less, it is called block size type 1. When it is other cases, it is called block size type 2. When this is the case, the number of intra prediction modes for MIP can be organized as shown in the following table.
[0266]
Table 10
[0267] In order to apply the weighted value matrix and offset vector used for a large block (for example, block size type = 2) to a small block (for example, block size = 0 or block size = 1), the number of available intra prediction modes for each block size can be applied in the same way as shown in the following table.
[0268]
Table 11
[0269] Alternatively, as shown in Table 12 below, apply MIP only to block size type 1 and 2. For block size type 1, the weighted value matrix and offset vector defined for block size type 2 can be subsampled and used. Through this, memory can be efficiently saved (50%).
[0270]
Table 12
[0271] The following drawings are created to explain a specific example of this specification. Since the names of specific devices and the names of specific signals / messages / fields described in the drawings are presented exemplarily, the technical features of this specification are not limited to the specific names used in the following drawings.
[0272] FIG. 19 is a flowchart schematically showing a decoding method that can be executed by a decoding apparatus according to an embodiment of the present document.
[0273] The method disclosed in FIG. 19 can be executed by the decoding apparatus 300 disclosed in FIG. 3. Specifically, steps S1900 to S1940 in FIG. 19 can be executed by the entropy decoding unit 310 and / or the prediction unit 330 (specifically, the intra prediction unit 331) disclosed in FIG. 3, and step S1950 in FIG. 19 can be executed by the addition unit 340 disclosed in FIG. 3. Also, the method disclosed in FIG. 19 can include the embodiments described above in this document. Therefore, in FIG. 19, specific descriptions of the content overlapping with the above-described embodiments will be omitted or simplified.
[0274] Referring to FIG. 19, the decoding apparatus can receive second flag information indicating whether matrix-based intra prediction (MIP) is used for the current block based on first flag information indicating whether MIP is applicable to the current block from a bitstream (S1900).
[0275] The first flag information can be signaled via syntax information of a sequence parameter set (SPS) as sps_mip_enabled_flag.
[0276] If such first flag information is 1, second flag information indicating whether matrix-based intra prediction (MIP) is used for the current block can be obtained from the bitstream.
[0277] Such second flag information can be signaled included in the syntax information of a coding unit in a syntax such as intra_mip_flag.
[0278] The decoding device can receive matrix-based intra prediction (MIP) mode information based on the received second flag information (S1910).
[0279] The MIP mode information can be represented by intra_mip_mode_idx and can be signaled when intra_mip_flag is 1. Intra_mip_mode_idx can be index information indicating the MIP mode applied to the current block, and such index information can be used to derive a matrix when generating prediction samples.
[0280] Also, by way of an example, flag information indicating whether an input vector for matrix-based intra prediction is transposed or not, for example, intra_mip_transposed_flag, can be further signaled via the syntax of the coding unit.
[0281] The decoding device can generate intra prediction samples for the current block based on the MIP information. The decoding device can derive at least one of the surrounding reference samples of the current block to generate intra prediction samples, and can generate prediction samples based on the surrounding reference samples.
[0282] When MIP is applied, the decoding device can downsample the reference samples adjacent to the current block and derive the reduced boundary samples (S1920).
[0283] The reduced boundary samples can be derived by being downsampled by averaging the reference samples.
[0284] When the width and height of the current block are 4, 4 reduced boundary samples can be derived, and in other remaining cases, 8 samples can be derived.
[0285] The averaging procedure for downsampling can be applied to each boundary, the left boundary or the upper boundary of the current block, which can be applied to the peripheral reference samples adjacent to the boundary of the current block.
[0286] By way of example, if the current block is a 4×4 block, the size of each boundary can be reduced to 2 samples via the averaging procedure, and if the current block is not a 4×4 block, the size of each boundary can be reduced to 4 samples via the averaging procedure.
[0287] Thereafter, the decoding device can derive reduced prediction samples based on the multiplication operation of the MIP matrix derived based on the size of the current block and the index information and the reduced boundary samples (S1930).
[0288] The MIP matrix can be derived based on the size of the current block and the received index information.
[0289] The MIP matrix can be selected from any of three matrix sets classified according to the size of the current block, and each of the three matrix sets can include a plurality of MIP matrices.
[0290] That is, three matrix sets for MIP can be set, and each matrix set can be composed of a plurality of matrices and offset vectors. Such matrix sets can be classified and applied according to the size of the current block.
[0291] For example, for a 4×4 block, a matrix set including 18 or 16 matrices composed of 16 rows and 4 columns and 18 or 16 offset vectors can be applied. The index information can be information indicating any one of the plurality of matrices included in one matrix set.
[0292] For 4×8, 8×4, and 8×8 blocks, or 4×H or W×4 blocks, a matrix set including 10 or 8 matrices composed of 16 rows and 8 columns, and 10 or 8 offset vectors can be applied.
[0293] Alternatively, for blocks other than the aforementioned blocks or blocks with a height and width of 8 or more, a matrix set including 6 matrices composed of 64 rows and 8 columns, and 6 offset vectors can be applied.
[0294] After the operation of multiplying the MIP matrix by the downsampled boundary samples, a downsampled predicted sample, that is, a predicted sample to which the MIP matrix is applied, is derived based on the operation of adding an offset.
[0295] The decoding device can upsample the downsampled predicted sample to generate an intra prediction sample for the current block (S1940).
[0296] The intra prediction sample can be upsampled by linear interpolation of the downsampled predicted sample.
[0297] The interpolation procedure can be referred to as a linear interpolation or bilinear interpolation procedure and can include two steps: 1) vertical interpolation and 2) horizontal interpolation.
[0298] If W >= H, vertical linear interpolation can be applied first, followed by horizontal linear interpolation. If W < H, horizontal linear interpolation can be applied first, followed by vertical linear interpolation. In the case of a 4×4 block, the interpolation procedure can be omitted.
[0299] The decoding device can generate a restored sample for the current block based on the predicted sample (S1950).
[0300] As one embodiment, the decoding device may immediately use the predicted sample as a restored sample according to the prediction mode, or may add a residual sample to the predicted sample to generate a restored sample.
[0301] When there is a residual sample for the current block, the decoding device can receive information regarding the residual for the current block. The information regarding the residual can include transform coefficients regarding the residual sample. The decoding device can derive a residual sample (or a residual sample array) for the current block based on the residual information. The decoding device can generate a restored sample based on the predicted sample and the residual sample, and can derive a restored block or a restored picture based on the restored sample. Thereafter, as described above, the decoding device can apply an in-loop filtering procedure such as deblocking filtering and / or SAO procedure to the restored picture in order to improve subjective / objective image quality as necessary.
[0302] FIG. 20 is a flowchart schematically showing an encoding method that can be executed by an encoding device according to an embodiment of this document.
[0303] The method disclosed in FIG. 20 can be executed by the encoding device 200 disclosed in FIG. 2. Specifically, steps S2000 to S2030 in FIG. 20 can be executed by the prediction unit 220 (specifically, the intra prediction unit 222) disclosed in FIG. 2, step S2040 in FIG. 20 can be executed by the subtraction unit 231 disclosed in FIG. 2, and step S2050 in FIG. 20 can be executed by the entropy encoding unit 240 disclosed in FIG. 2. Also, the method disclosed in FIG. 20 can include the embodiments described above in this document. Therefore, in FIG. 20, specific descriptions of the content overlapping with the above-described embodiments will be omitted or simplified.
[0304] Referring to FIG. 20, the encoding device can derive whether matrix-based intra prediction (MIP) is applied to the current block (S2000).
[0305] The encoding device can apply various prediction techniques to find the optimal prediction mode for the current block, and can determine the optimal intra prediction mode based on rate-distortion optimization (RDO).
[0306] When it is determined that MIP is applied to the current block, the encoding device can downsample the reference samples adjacent to the current block to derive the reduced boundary samples (S2010).
[0307] The reduced boundary samples can be derived by being downsampled by averaging the reference samples.
[0308] When the width and height of the current block are 4, 4 reduced boundary samples can be derived, and in other cases, 8 samples can be derived.
[0309] The averaging procedure for downsampling can be applied to each boundary, the left boundary or the upper boundary of the current block, which can be applied to the surrounding reference samples adjacent to the boundary of the current block.
[0310] By way of example, when the current block is a 4×4 block, the size of each boundary can be reduced to 2 samples through the averaging procedure, and when the current block is not a 4×4 block, the size of each boundary can be reduced to 4 samples through the averaging procedure.
[0311] When a reduced boundary sample is derived, the encoding device can derive a reduced prediction sample based on an operation of multiplying the MIP matrix selected based on the size of the current block by the reduced boundary sample (S2020).
[0312] The MIP matrix can be selected from any of three matrix sets classified according to the size of the current block, and each of the three matrix sets can include a plurality of MIP matrices.
[0313] That is, three matrix sets for MIP can be set, and each matrix set can be composed of a plurality of matrices and offset vectors. Such a matrix set can be classified and applied according to the size of the current block.
[0314] For example, for a 4×4 block, a matrix set including 18 or 16 matrices composed of 16 rows and 4 columns and 18 or 16 offset vectors can be applied. The index information can be information indicating any one of the plurality of matrices included in one matrix set.
[0315] For 4×8, 8×4, and 8×8 blocks, or 4×H or W×4 blocks, a matrix set including 10 or 8 matrices composed of 16 rows and 8 columns and 10 or 8 offset vectors can be applied.
[0316] Alternatively, for blocks other than the aforementioned blocks or blocks with a height and width of 8 or more, a matrix set including 6 matrices composed of 64 rows and 8 columns and 6 offset vectors can be applied.
[0317] After the multiplication operation of the MIP matrix and the reduced boundary samples, a reduced predicted sample can be derived based on the operation of adding an offset, that is, a predicted sample to which the MIP matrix is applied.
[0318] Thereafter, the encoding device can upsample the reduced predicted sample to generate an intra prediction sample for the current block (S2030).
[0319] The intra prediction sample can be upsampled by linear interpolation of the reduced predicted sample.
[0320] The interpolation procedure can be referred to as a linear interpolation or a bilinear interpolation procedure and can include two steps: 1) vertical interpolation, and 2) horizontal interpolation.
[0321] If W >= H, vertical linear interpolation can be applied first, followed by horizontal linear interpolation. If W < H, horizontal linear interpolation can be applied first, followed by vertical linear interpolation. In the case of a 4×4 block, the interpolation procedure can be omitted.
[0322] Also, the encoding device can derive a residual sample for the current block based on the predicted sample of the current block and the original sample of the current block (S2040).
[0323] Then, the encoding device can generate residual information for the current block based on the residual sample, and encode video information including the residual information, first flag information indicating whether matrix-based intra prediction (MIP) is applicable to the current block, second flag information indicating whether MIP is used for the current block, and MIP mode information (S2050).
[0324] Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the derived quantized conversion coefficients obtained by performing conversion and quantization on the residual samples.
[0325] The first flag information indicating whether matrix-based intra prediction (MIP) is applicable to the current block can be encoded via the syntax information of the sequence parameter set (SPS) as sps_mip_enabled_flag.
[0326] The second flag information indicating whether MIP is applied can be encoded in the syntax information of the coding unit with a syntax such as intra_mip_flag.
[0327] Also, the MIP mode information can be represented by intra_mip_mode_idx and can be encoded when intra_mip_flag is 1. intra_mip_mode_idx can be index information indicating the MIP mode applied to the current block, and such index information can be used to derive a matrix when generating prediction samples. The index information can indicate any one of the plurality of MIP matrices included in one matrix set.
[0328] Also, by way of an example, flag information, for example, intra_mip_transposed_flag, indicating whether the input vector for matrix-based intra prediction is transposed via the syntax of the coding unit can be further signaled.
[0329] That is, the encoding device can encode video information including the above-described MIP mode information and / or residual information of the current block and output it to the bitstream.
[0330] The bitstream can be transmitted to a decoding device via a network or a (digital) storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0331] The process of generating prediction samples for the current block described above can be executed by the intra prediction unit 222 of the encoding device 200 disclosed in FIG. 2, the process of deriving residual samples can be executed by the subtraction unit 231 of the encoding device 200 disclosed in FIG. 2, and the process of generating and encoding residual information can be executed by the residual processing unit 230 and the entropy encoding unit 240 of the encoding device 200 disclosed in FIG. 2.
[0332] In the above-described embodiments, the method is described based on a flowchart as a series of steps or blocks, but the embodiments of this document are not limited to the order of the steps, and a certain step can occur in a different order from the steps described above or simultaneously. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, and different steps may be included, or one or more steps of the flowchart may be deleted without affecting the scope of this document.
[0333] The method according to the above-described document can be embodied in the form of software, and the encoding device and / or decoding device according to this document can be included in a device that executes video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0334] In this document, when an embodiment is implemented in software, the above-described method can be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored in a digital storage medium.
[0335] In addition, the decoding device and encoding device to which this document is applicable may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video dialogue device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an order-made video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, an image phone video device, a transportation means terminal (e.g., a vehicle terminal including an autonomous driving vehicle terminal, an airplane terminal, a ship terminal, etc.) and a medical video device, etc., and can be used to process video signals or data signals. For example, as an OTT video (Over the top video) device, it can include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recoder), etc.
[0336] In addition, the processing method to which this document is applicable can be produced in the form of a program executed by a computer and can be stored in a recording medium readable by a computer. Multimedia data having a data structure related to this document can also be stored in a recording medium readable by a computer. The recording medium readable by the computer includes all types of storage devices and distributed storage devices in which data readable by a computer is stored. The recording medium readable by the computer can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Further, the recording medium readable by the computer includes a medium embodied in the form of a carrier wave (e.g., transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a recording medium readable by a computer or can be transmitted via a wired or wireless communication network.
[0337] In addition, embodiments of this document can be embodied in a computer program product by program code, and the program code can be executed by a computer according to embodiments of this document. The program code can be stored on a carrier readable by a computer.
[0338] The claims described in this document can be combined in various ways. For example, the technical features of the method claims in this document can be combined and implemented as an apparatus, and the technical features of the apparatus claims can be combined and implemented as a method. Also, the technical features of the method claims and the technical features of the apparatus claims in this document can be combined and implemented as an apparatus, and the technical features of the method claims and the technical features of the apparatus claims in this document can be combined and implemented as a method. (Claims in the present description can be combined in a various way. For instance, technical features in method claims of the present description can be combined to be implemented or performed in an apparatus, and technical features in apparatus claims can be combined to be implemented or performed in a method. Further, technical features in method claim(s) and apparatus claim(s) can be combined to be implemented or performed in an apparatus. Further, technical features in method claim(s) and apparatus claim(s) can be combined to be implemented or performed in a method.)
Claims
1. 1. A video decoding method performed by a decoding device, comprising: receiving a bitstream; receiving residual information from the bitstream; receiving second flag information related to whether an intra prediction type for a current block is the matrix-based intra prediction (MIP) based on first flag information related to whether the MIP is applicable, the first flag information being equal to 1 indicating that the MIP is applicable; receiving MIP mode information based on the second flag information, the second flag information being equal to 1 indicating that the intra prediction type for the current block is the MIP; generating intra-prediction samples for the current block based on the MIP mode information; deriving transformation coefficients based on the residual information; generating residual samples based on the transform coefficients; generating a reconstructed sample for the current block based on the intra-predicted sample and the residual sample; the MIP mode information is MIP mode index information indicating a MIP mode to be applied to the current block; the MIP mode index information is used to derive a MIP matrix for the current block; the intra prediction samples are generated using the MIP matrix derived from the MIP mode index information; The syntax element bin string for the MIP mode information is binarized by a truncated binarization (TB) method, A method according to claim 1, wherein a maximum length of the syntax element bin string is determined by a size of the current block.
2. A video encoding method performed by an encoding device, comprising: deriving whether matrix-based intra prediction (MIP) is applied to the current block; deriving intra-prediction samples for the current block based on the application of the MIP to the current block; deriving a residual sample for the current block based on the intra prediction samples; deriving transformation coefficients based on a transformation process for the residual samples; generating residual information based on the transformation coefficients; encoding image information including the residual information and information related to the MIP; generating a bitstream including the encoded video information; The information regarding the MIP includes: first flag information relating to whether the MIP is applicable; second flag information relating to whether an intra prediction type for the current block is the MIP; and MIP mode information; the second flag information is signaled based on the first flag information; the first flag information is equal to 1 to indicate that the MIP is applicable; the MIP mode information is signaled based on the second flag information; the second flag information is equal to 1 indicating that the intra prediction type for the current block is the MIP; the MIP mode information is MIP mode index information indicating a MIP mode to be applied to the current block; the MIP mode index information is used to derive a MIP matrix for the current block; the intra prediction samples are derived using the MIP matrix indicated by the MIP mode index information; The syntax element bin string for the MIP mode information is binarized by a truncated binarization (TB) method, A method according to claim 1, wherein a maximum length of the syntax element bin string is determined by a size of the current block.
3. A method for transmitting data for a video, comprising the steps of: obtaining a bitstream for the video, the bitstream comprising: deriving whether matrix-based intra prediction (MIP) is applied to the current block; deriving intra-prediction samples for the current block based on the application of the MIP to the current block; deriving a residual sample for the current block based on the intra prediction samples; deriving transformation coefficients based on a transformation process for the residual samples; generating residual information based on the transformation coefficients; encoding image information including the residual information and information related to the MIP; generating a bitstream including the encoded video information; transmitting the data including the bitstream; The information regarding the MIP includes: first flag information relating to whether the MIP is applicable; second flag information relating to whether an intra prediction type for the current block is the MIP; and MIP mode information; the second flag information is signaled based on the first flag information; the first flag information is equal to 1 to indicate that the MIP is applicable; the MIP mode information is signaled based on the second flag information; the second flag information is equal to 1 indicating that the intra prediction type for the current block is the MIP; the MIP mode information is MIP mode index information indicating a MIP mode to be applied to the current block; the MIP mode index information is used to derive a MIP matrix for the current block; the intra prediction samples are derived using the MIP matrix indicated by the MIP mode index information; The syntax element bin string for the MIP mode information is binarized by a truncated binarization (TB) method, A method according to claim 1, wherein a maximum length of the syntax element bin string is determined by a size of the current block.
Citation Information
Patent Citations
JPP7379543B
JPP7509978B
Simplified signaling method for affine linear weighted intra prediction mode
US20200322620A1
Simplified signaling method for affine linear weighted intra prediction mode
WO2020205705A1