Matrix-based intra prediction device and method
The matrix-based intra prediction method enhances video coding efficiency by simplifying the encoding and decoding processes for high-resolution and immersive video, addressing the need for cost-effective compression and transmission of high-quality video data.
Patent Information
- Application Number
- JP2025067998
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-03
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2040-06-03
AI Technical Summary
The increasing demand for high-resolution and high-quality video, including immersive media like VR and AR, necessitates a highly efficient video coding technique to compress, transmit, and store video data effectively, reducing transmission and storage costs.
A video coding method and apparatus utilizing matrix-based intra prediction (MIP) with truncated binary coding for MIP mode information, including steps like downsampling reference samples, matrix multiplication, and upsampling to generate intra prediction samples, and encoding residual samples with efficient coding methods.
Improves video coding efficiency by reducing implementation complexity and enhancing prediction performance through efficient intra prediction, particularly for high-resolution and immersive video content.
Smart Images

Figure 2025105664000001_ABST
Abstract
Description
Technical Field
[0001] This document relates to video coding technology, and more particularly to video coding technology for a matrix-based intra prediction apparatus and method.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality video / video such as 4K or 8K and above UHD (Ultra High Definition) video / video has been increasing in various fields. As the video / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing video / video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line, or storing video / video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of video / video having video characteristics different from those of real-world video, such as game video, has been increasing.
[0004] Accordingly, there is a need for a highly efficient video / video compression technique to effectively compress, transmit, store, and reproduce information of high-resolution and high-quality video / video having various characteristics as described above.
Summary of the Invention
Problems to be Solved by the Invention
[0005] The technical problem of this document is to provide a method and apparatus for improving video coding efficiency.
[0006] Another technical problem of this document is to provide an efficient intra prediction method and apparatus.
[0007] Another technical problem of this document is to provide a video coding method and apparatus for intra prediction based on a matrix.
[0008] Another technical problem of this document is to provide a video coding method and apparatus for coding mode information for intra prediction based on a matrix.
Means for Solving the Problem
[0009] According to one embodiment of this document, a video decoding method executed by a decoding apparatus is provided. The method includes receiving flag information indicating whether matrix-based intra prediction (MIP) is used for a current block; receiving matrix-based intra prediction (MIP) mode information based on the flag information; generating an intra prediction sample for the current block based on the MIP mode information; and generating a restored sample for the current block based on the intra prediction sample, wherein the bin string of the syntax element for the MIP mode information can be binary-coded by a truncated binary coding method.
[0010] The maximum length of the bin string of the syntax element for the MIP mode information can be set to different values according to the size of the current block.
[0011] The maximum length is set to three values according to the size of the current block, and the maximum length may be the largest when the width and height of the current block are 4.
[0012] Such MIP mode information can be decoded in a bypass mode.
[0013] The MIP mode information may be index information indicating the MIP mode applied to the current block.
[0014] The step of generating the intra prediction sample may include: a step of downsampling a reference sample adjacent to the current block to derive a reduced boundary sample; a step of deriving a reduced prediction sample based on the multiplication of the reduced boundary sample and the MIP matrix; and a step of upsampling the reduced prediction sample to generate the intra prediction sample for the current block.
[0015] Here, the reduced boundary sample may be downsampled by averaging the reference samples, and the intra prediction sample may be upsampled by linear interpolation of the reduced prediction sample.
[0016] The MIP matrix can be derived based on the size of the current block and the index information.
[0017] The MIP matrix can be selected from any one of three matrix sets classified according to the size of the current block, and each of the three matrix sets can include a plurality of MIP matrices.
[0018] According to one embodiment of this document, a video encoding method executed by an encoding device is provided. The method includes: deriving whether matrix-based intra prediction (MIP) is applicable to a current block; when the MIP is applicable to the current block, deriving intra prediction samples of the current block based on the MIP; deriving residual samples for the current block based on the intra prediction samples; and encoding information on the residual samples and information on the MIP. The information on the MIP includes matrix-based intra prediction (MIP) mode information, and a bin string of a syntax element for the MIP mode information can be evolved by a truncated binary evolution method.
[0019] According to another embodiment of this document, a digital storage medium storing encoded video information generated by a video encoding method executed by an encoding device and video data including a bitstream can be provided.
[0020] According to another embodiment of this document, a digital storage medium storing encoded video information and video data including a bitstream, which causes a decoding device to perform the video decoding method, can be provided.
Advantages of the Invention
[0021] This document can have various advantages. For example, according to one embodiment of this document, the compression efficiency of general video can be improved. Alternatively, according to one embodiment of this document, the overall coding efficiency can be improved by reducing the implementation complexity and improving the prediction performance through efficient intra prediction. Or, according to one embodiment of this document, when performing matrix-based intra prediction, the index information indicating this can be efficiently coded to improve the coding efficiency.
[0022] The effects obtained through a specific example of this document are not limited to the effects listed above. For example, there may be various technical effects that can be understood or induced by a person having ordinary skill in the related art from this document. Accordingly, the specific effects of this document are not limited to those explicitly described in this document, and may include various effects that can be understood or induced from the technical features of this document.
Brief Description of the Drawings
[0023]
Figure 1
[0024]
Figure 2
[0025]
Figure 3
[0026]
Figure 4
[0027]
Figure 5
[0028]
Figure 6
[0029]
Figure 7
[0030]
Figure 8
[0031]
Figure 9
[0032]
Figure 10
[0033]
Figure 11
[0034]
Figure 12
[0035]
Figure 13
[0036]
Figure 14
[0037]
Figure 15
[0038]
Figure 16
[0039]
Figure 17
[0040]
Figure 18
[0041]
Figure 19
[0042]
Figure 20
DETAILED DESCRIPTION OF THE INVENTION
[0043] This document can be modified in various ways and can have various embodiments, but specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this document are used merely to explain specific embodiments and are not intended to limit the technical idea in this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this document, terms such as "including" or "having" are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the document, and should be understood not to preclude in advance the possibility of the existence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof.
[0044] On the one hand, each component in the drawings described in this document is shown independently for the convenience of explaining different characteristic functions, and it does not mean that each component is implemented by separate hardware or separate software. For example, among the components, two or more components may be combined to form one component, or one component may be divided into multiple components. Embodiments in which the components are integrated and / or separated are also included in the scope of rights of this document as long as they do not deviate from the essence of this document.
[0045] In this document, "A or B" may mean "only A", "only B", or "both A and B". In other words, in this document, "A or B" may be interpreted as "A and / or B". For example, in this document, "A, B or C" may mean "only A", "only B", "only C", or "any combination of A, B and C".
[0046] The slashes ( / ) and commas used in this document may mean "and / or". For example, "A / B" may mean "A and / or B". Thus, "A / B" may mean "only A", "only B", or "both A and B". For example, "A, B, C" may mean "A, B or C".
[0047] In this document, "at least one of A and B" may mean "only A", "only B", or "both A and B". Also, in this document, expressions such as "at least one of A or B" and "at least one of A and / or B" may be interpreted in the same way as "at least one of A and B".
[0048] Also, in this document, "at least one of A, B and C" may mean "only A", "only B", "only C", or "any combination of A, B and C". Further, "at least one of A, B or C" and "at least one of A, B and / or C" may mean "at least one of A, B and C".
[0049] Also, the parentheses used in this document may mean "for example". Specifically, when it is displayed as "prediction (intra prediction)", "intra prediction" may be proposed as an example of "prediction". In other words, the "prediction" in this document is not limited to "intra prediction", and "intra prediction" may be proposed as an example of "prediction". Also, when it is displayed as "prediction (i.e., intra prediction)", "intra prediction" may be proposed as an example of "prediction".
[0050] The technical features separately described within one drawing in this document may be implemented separately or simultaneously.
[0051] This document relates to video / video coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268, etc.).
[0052] This document presents various embodiments related to video / video coding, and unless otherwise stated, the embodiments may be implemented in combination with each other.
[0053] In this document, "video" can mean a collection of a series of images over time. "Picture" generally means a unit representing one image at a specific time period, and "slice" / "tile" is a unit that constitutes a part of a picture in coding. A slice / tile can contain one or more CTUs (Coding Tree Units). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can contain one or more tiles. A brick may represent a rectangular region of CTU rows within a tile in a picture. A tile may be partitioned into multiple bricks, each of which consisting of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may be also referred to as a brick.A brick scan may indicate a specific sequential ordering of CTUs partitioning a picture, where the CTUs may be ordered consecutively in a CTU raster scan within a brick, bricks within a tile may be ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture may be ordered consecutively in a raster scan of the tiles of the picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set.The tile row is a rectangular region of CTUs having a width specified by syntax elements in the picture parameter set and a height equal to the height of the picture. A tile scan is a specific sequential ordering of CTUs partitioning a picture in which the CTUs are ordered consecutively in CTU raster scan in a tile whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice may include an integer number of bricks of a picture that may be exclusively contained in a single NAL unit. A slice may consists of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile.In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0054] A pixel or pel may mean the smallest unit that makes up one picture (or video). Also, the term "sample" may be used as a term corresponding to a pixel. A sample may generally indicate a pixel or a pixel value, or may only indicate the pixel / pixel value of the luma component, or may only indicate the pixel / pixel value of the chroma component. Alternatively, a sample may mean a pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it may mean a conversion coefficient in the frequency domain.
[0055] A unit can indicate the basic unit of video processing. A unit can include at least one of a specific area of a picture and information related to the corresponding area. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit may, in some cases, be used interchangeably with terms such as block or area. In general, an M×N block can include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0056] Hereinafter, with reference to the attached drawings, the preferred embodiments of this document will be described in more detail. Hereinafter, for the same components in the drawings, the same reference numerals will be used, and duplicate descriptions for the same components may be omitted.
[0057] FIG. 1 is a diagram schematically explaining the configuration of a video / video encoding apparatus applicable to an embodiment of this document. Hereinafter, the video encoding apparatus can include a video encoding apparatus.
[0058] Referring to FIG. 1, the encoding apparatus 100 can be configured to include an image partitioner 110, a predictor 120, a residual processor 130, an entropy encoder 140, an adder 150, a filter 160, and a memory 170. The predictor 120 can include an inter-predictor 121 and an intra-predictor 122. The residual processor 130 can include a transformer 132, a quantizer 133, a dequantizer 134, and an inverse transformer 135. The residual processor 130 can further include a subtractor 131. The adder 150 can be referred to as a reconstructor or a reconstructed block generator. The aforementioned image partitioner 110, predictor 120, residual processor 130, entropy encoder 140, adder 150, and filter 160 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 170 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 170 as an internal / external component.
[0059] The video segmentation unit 110 can divide the input video (or picture, frame) input to the encoding device 100 into one or more processing units. As an example, the processing unit may be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into multiple coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or, the binary-tree structure can also be applied first. Based on the final coding unit that cannot be further divided, the coding procedure according to this document can be performed. In this case, based on the coding efficiency according to the video characteristics, etc., the largest coding unit can be immediately used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can be divided or partitioned from the final coding unit described above, respectively.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.
[0060] The unit can be used interchangeably with terms such as block or area depending on the context. In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel for one picture (or video).
[0061] The encoding device 100 can subtract the predicted signal (predicted block, predicted sample array) output from the inter prediction unit 121 or the intra prediction unit 122 from the input video signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 132. In this case, as shown, the unit that subtracts the predicted signal (predicted block, predicted sample array) from the input video signal (original block, original sample array) within the encoding device 100 can be called the subtraction unit 131. The prediction unit can perform prediction on the processing target block (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 140 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 140 and output in the form of a bitstream.
[0062] The intra prediction unit 122 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located in the neighborhood of the current block or at a distance depending on the prediction mode. The prediction modes in intra prediction can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of fineness of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 122 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.
[0063] The inter prediction unit 121 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (col CU), and the reference picture including the temporal neighboring block can also be called a collocated picture (colPic). For example, the inter prediction unit 121 can construct a motion information candidate list based on the peripheral block and generate information indicating which candidate is used to derive the motion vector and / or the index of the reference picture of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 121 can use the motion information of the peripheral block as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector difference is signaled, so that the motion vector of the current block can be indicated.
[0064] The prediction unit 120 can generate a prediction signal based on various prediction methods described later. For example, for the prediction of one block, the prediction unit can apply not only intra prediction or inter prediction, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit may be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for coding content videos / movies such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information regarding the palette table and the palette index.
[0065] The prediction signal generated via the prediction unit (including the inter prediction unit 121 and / or the intra prediction unit 122) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 132 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when expressing the relationship information between pixels in a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process may be applied to a pixel block having the same size of a square or may be applied to a block of a variable size that is not a square.
[0066] The quantization unit 133 quantizes the transform coefficients and transmits them to the entropy encoding unit 140. The entropy encoding unit 140 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients may be referred to as residual information. The quantization unit 133 can reorder the block-form quantized transform coefficients in the form of a one-dimensional vector based on the scan order of the coefficients, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the form of the one-dimensional vector. The entropy encoding unit 140 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 140 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in units of NAL (network abstraction layer) units in the form of a bitstream. The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. In this document, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information can be encoded via the encoding procedure described above and be included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 140 may be configured as an internal / external element of the encoding device 100 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit may be included in the entropy encoding unit 140.
[0067] The quantized transform coefficients output from the quantization unit 133 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 134 and the inverse transform unit 135, a residual signal (residual block or residual sample) can be restored. The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 121 or the intra prediction unit 122. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 may be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture and, as will be described later, can also be used for inter prediction of the next picture after passing through filtering.
[0068] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture encoding and / or restoration process.
[0069] The filtering unit 160 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 160 can generate various information related to filtering and transmit it to the entropy encoding unit 140 as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 140 and output in the form of a bitstream.
[0070] The modified restored picture transmitted to the memory 170 can be used as a reference picture by the inter prediction unit 121. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve the encoding efficiency.
[0071] The DPB in memory 170 can be saved for use as a reference picture in the inter prediction unit 121 for the modified reconstructed picture. Memory 170 can save the motion information of the blocks for which the motion information in the current picture has been derived (or encoded) and / or the motion information of the blocks in the already reconstructed picture. The saved motion information can be transmitted to the inter prediction unit 121 for utilization as the motion information of spatially neighboring blocks or temporally neighboring blocks. Memory 170 can save the reconstructed samples of the reconstructed blocks in the current picture and can transmit them to the intra prediction unit 122.
[0072] FIG. 2 is a diagram schematically illustrating the configuration of a video / video decoding apparatus applicable to the embodiments of this document.
[0073] Referring to FIG. 2, the decoding apparatus 200 can be configured to include an entropy decoder 210, a residual processor 220, a predictor 230, an adder 240, a filtering unit 250, and a memory 260. The predictor 230 can include an inter prediction unit 231 and an intra prediction unit 232. The residual processor 220 can include a dequantizer 221 and an inverse transformer 221. The above-described entropy decoding unit 210, residual processing unit 220, prediction unit 230, addition unit 240, and filtering unit 250 can be configured by one hardware component (for example, a decoder chipset or a processor) according to the embodiment. Further, the memory 260 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 260 as an internal / external component.
[0074] When a bitstream including video / video information is input, the decoding device 200 can restore the video corresponding to the process in which the video / video information was processed by the encoding device in FIG. 1. For example, the decoding device 200 can derive units / blocks based on the information regarding block division obtained from the bitstream. The decoding device 200 can execute decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding may be, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit by a quad-tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. Then, the restored video signal decoded and output via the decoding device 200 can be played back via a playback device.
[0075] The decoding device 200 can receive the signal output from the encoding device of FIG. 1 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding and the block to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin according to the determined context model, executes arithmetic decoding of the bin, and can generate the symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the next symbol / bin context model after determining the context model. Among the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction unit (inter prediction unit 232 and intra prediction unit 231), and the residual value for which entropy decoding is executed by the entropy decoding unit 210, that is, the quantized transform coefficient and related parameter information can be input to the residual processing unit 220. The residual processing unit 220 can derive a residual signal (residual block, residual sample, residual sample array). Also, among the information decoded by the entropy decoding 210, the information related to filtering can be provided to the filtering unit 250. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device may be further configured as an internal / external element of the decoding device 200, or the receiving unit may be a component of the entropy decoding unit 210. On the other hand, the decoding device according to the present document may be called a video / video / picture decoding device, and the decoding device may be classified into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder may include the entropy decoding unit 210, and the sample decoder may include at least one of the inverse quantization unit 221, inverse transform unit 222, addition unit 240, filtering unit 250, memory 260, inter prediction unit 232, and intra prediction unit 231.
[0076] In the inverse quantization unit 221, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 221 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the scan order of the coefficients executed by the encoding device. The inverse quantization unit 221 can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) to obtain transform coefficients.
[0077] In the inverse transform unit 222, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).
[0078] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode.
[0079] The prediction unit 220 can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for the prediction of one block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit may be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for content video / moving picture coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be executed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be signaled and included in the video / video information.
[0080] The intra prediction unit 231 can predict the current block by referring to samples within the current picture. The samples to be referred to may be located adjacent to or away from the periphery of the current block depending on the prediction mode. The prediction modes in intra prediction can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 231 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the surrounding blocks.
[0081] The inter prediction unit 232 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. For example, the inter prediction unit 232 can construct a motion information candidate list based on the peripheral blocks, and derive the motion vector of the current block and / or the index of the reference picture based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.
[0082] The addition unit 240 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit 232 and / or the intra prediction unit 231). When there is no residual for the processing target block, as in the case where the skip mode is applied, the predicted block can be used as the restored block.
[0083] The addition unit 240 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block in the current picture, and may be output after filtering as described later, or may be used for inter prediction of the next picture.
[0084] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture decoding process.
[0085] The filtering unit 250 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 250 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 260, specifically, to the DPB of the memory 260. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, and the like.
[0086] The (modified) restored picture stored in the DPB of the memory 260 can be used as a reference picture in the inter prediction unit 232. The memory 260 can store the motion information of the blocks for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored pictures. The stored motion information can be transmitted to the inter prediction unit 232 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 260 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 231.
[0087] In this document, the embodiments described in the filtering unit 160, the inter prediction unit 121, and the intra prediction unit 122 of the encoding device 100 can also be applied to the filtering unit 250, the inter prediction unit 232, and the intra prediction unit 231 of the decoding device 200 so as to be the same or corresponding.
[0088] As described above, in performing video coding, prediction is carried out to improve the compression efficiency. Through this, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is derived in the same way in both the encoding device and the decoding device. The encoding device can improve the efficiency of video coding by signaling to the decoding device information (residual information) regarding the residual between the original block and the predicted block, rather than the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, combine the residual block and the predicted block, and generate a restored block including restored samples, and can generate a restored picture including the restored block.
[0089] The residual information can be generated through the conversion and quantization procedures. For example, an encoding device can derive a residual block between an original block and a predicted block, perform a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, perform a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and signal the related residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can perform an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in the inter prediction of subsequent pictures to derive a residual block and generate a restored picture based on this.
[0090] On the other hand, as described above, the encoding device can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context - adaptive variable length coding), CABAC (context - adaptive binary arithmetic coding), etc. Also, the decoding device can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements necessary for video restoration and the quantized values of the conversion coefficients related to the residuals.
[0091] For example, the coding methods described above can be performed as described in the content to be described later.
[0092] FIG. 3 exemplarily shows context - adaptive binary arithmetic coding (CABAC) for encoding a syntax element. For example, in the encoding process of CABAC, when the input signal is not a binary value but a syntax element, the encoding device can binarize the value of the input signal to convert the input signal into a binary value. Also, when the input signal is already a binary value (i.e., when the value of the input signal is a binary value), it can be bypassed without binarization. Here, each binary number 0 or 1 that constitutes a binary value can be called a bin. For example, when the binary string after binarization is 110, 1, 1, and 0 are each called one bin. The bin for one syntax element can indicate the value of the syntax element. Such binarization can be based on various binarization methods such as Truncated Rice binarization process, Fixed - length binarization process, etc., and the binarization method for the target syntax element can be predefined. The binarization procedure can be executed by a binarization unit within the entropy encoding unit.
[0093] Thereafter, the binarized bins of the syntax element can be input into a regular encoding engine or a bypass encoding engine. The regular encoding engine of the encoding device can assign a context model that reflects a probability value to the corresponding bin, and encode the corresponding bin based on the assigned context model. After executing the encoding for each bin, the regular encoding engine of the encoding device can update the context model for the corresponding bin. As described above, the bins to be encoded can be referred to as context - coded bins.
[0094] On the other hand, when the evolved bins of the syntax element are input to the bypass encoding engine, they can be coded as follows. For example, the bypass encoding engine of the encoding device omits the procedure of estimating the probability for the input bin and the procedure of updating the probability model applied to the bin after encoding. When bypass encoding is applied, the encoding device can apply a uniform probability distribution instead of assigning a context model and encode the input bin, thereby improving the encoding speed. As described above, the bin to be encoded can be indicated as a bypass bin.
[0095] Entropy decoding can be shown as a process of executing the same process as the above-described entropy encoding in reverse order.
[0096] A decoding device (entropy decoding unit) can decode the encoded video / video information. The video / video information can include information regarding partitioning, information regarding prediction (e.g., inter / intra prediction division information, intra prediction mode information, inter prediction mode information, etc.), residual information, information regarding in-loop filtering, etc., or can include various syntax elements regarding the same. The entropy coding can be executed in units of syntax elements.
[0097] The decoding device can perform binarization on the target syntax element. Here, the binarization can be based on various binarization methods such as Truncated Rice binarization process, Fixed-length binarization process, etc., and the binarization method for the target syntax element can be predefined. The decoding device can derive an available bit string (a candidate for the bit string) for the available value of the target syntax element through the binarization procedure. The binarization procedure can be executed by a binarization unit within the entropy decoding unit.
[0098] The decoding device sequentially decodes and parses each bin for the target syntax element from the input bits in the bitstream, and compares the derived bit string with the available bit strings for the corresponding syntax element. If the derived bit string is the same as one of the available bit strings, the value corresponding to the corresponding bit string is derived as the value of the corresponding syntax element. If not, after further parsing the next bit in the bitstream, the above-described procedure is executed again. Through such a process, without using start bits or end bits for specific information (specific syntax elements) in the bitstream, variable-length bits can be used to signal the corresponding information. Through this, relatively fewer bits can be allocated for low values, and the overall coding efficiency can be improved.
[0099] The decoding device can decode each bin in the bit string from the bitstream based on an entropy coding technique such as CABAC or CAVLC, based on a context model, or based on bypass.
[0100] When a syntax element is decoded based on a context model, the decoding apparatus can receive a bin corresponding to the syntax element via a bitstream, and can determine a context model using the decoding information of the syntax element and the block to be decoded or the surrounding blocks, or the information of the symbol / bin decoded in the previous step. Based on the determined context model, the occurrence probability of the received bin can be predicted to perform arithmetic decoding of the bin, and the value of the syntax element can be derived. Thereafter, the context model of the bin to be decoded next can be updated with the determined context model.
[0101] The context model can be allocated and updated for each bin to be context-coded (entropy-coded), and the context model can be indicated based on ctxIdx or ctxInc. ctxIdx can be derived based on ctxInc. Specifically, for example, the context index (ctxIdx) indicating the context model for each of the bins to be entropy-coded can be derived as the sum of the context index increment (ctxInc) and the context index offset (ctxIdxOffset). Here, the ctxInc can be derived differently for each bin. The ctxIdxOffset can be represented by the lowest value of the ctxIdx. The ctxIdxOffset is generally a value used for distinguishing from the context models for other syntax elements, and the context model for one syntax element can be distinguished / derived based on ctxInc.
[0102] In the entropy encoding procedure, it is possible to determine whether to perform encoding via the normal encoding engine or via the bypass encoding engine, and switch the coding path. Entropy decoding performs the same process as entropy encoding in reverse.
[0103] On the other hand, for example, when a syntax element is bypass decoded, the decoding device can receive bins corresponding to the syntax element via the bitstream and decode the input bins by applying a uniform probability distribution. In this case, the procedure for deriving the context model of the syntax element and the procedure for updating the context model applied to the bins after decoding can be omitted.
[0104] As described above, the residual samples can be derived as quantized transform coefficients through the conversion and quantization processes. The quantized transform coefficients may also be referred to as transform coefficients. In this case, the transform coefficients within the block can be signaled in the form of residual information. The residual information can include residual coding syntax. That is, the encoding device can construct a residual coding syntax with the residual information, encode it, and output it in the form of a bitstream. The decoding device can decode the residual coding syntax from the bitstream and derive the residual (quantized) transform coefficients. The residual coding syntax can include syntax elements indicating whether a transform is applied to the corresponding block, where the position of the last valid transform coefficient within the block is, whether there are valid transform coefficients within the sub-block, what the magnitude / symbol of the valid transform coefficients is, etc., as will be described later.
[0105] On the one hand, when intra prediction is performed, the correlation relationship between samples can be used, and the difference between the original block and the predicted block, that is, the residual, can be obtained. The above-mentioned transformation and quantization can be applied to the residual, and through this, the spatial redundancy can be removed. Hereinafter, the encoding method and the decoding method using intra prediction will be specifically described.
[0106] Intra prediction refers to a prediction that generates prediction samples for a current block based on reference samples outside the current block within a picture (hereinafter referred to as the current picture) including the current block. Here, the reference samples outside the current block can be said to be samples located around the current block. When intra prediction is applied to the current block, the neighboring reference samples used for intra prediction of the current block can be derived.
[0107] For example, when the size (width × height) of the current block is nW × nH, the neighboring reference samples of the current block can include samples adjacent to the left boundary of the current block and a total of 2 × nH samples adjacent to the bottom-left, samples adjacent to the top boundary of the current block and a total of 2 × nW samples adjacent to the top-right, and 1 sample adjacent to the top-left of the current block. Alternatively, the neighboring reference samples of the current block can also include a plurality of rows of upper neighboring samples and a plurality of columns of left neighboring samples. Also, the neighboring reference samples of the current block can include a total of nH samples adjacent to the right boundary of the current block with a size of nW × nH, a total of nW samples adjacent to the bottom boundary of the current block, and 1 sample adjacent to the bottom-right of the current block.
[0108] However, some of the neighboring reference samples of the current block may not have been decoded yet or may not be available. In this case, the decoding device can substitute the unavailable samples with available samples to form the neighboring reference samples used for prediction. Alternatively, the neighboring reference samples used for prediction can be formed through interpolation of the available samples.
[0109] When the neighboring reference samples are derived, (i) the predicted sample can be derived based on the average or interpolation of the neighboring reference samples of the current block, and (ii) the predicted sample can also be derived based on the reference samples existing in a specific (prediction) direction with respect to the predicted sample among the neighboring reference samples of the current block. In the case of (i), it can be applied when the intra prediction mode is a non-directional mode or a non-angle mode, and in the case of (ii), it can be applied when the intra prediction mode is a directional mode or an angular mode.
[0110] Also, among the neighboring reference samples, the predicted sample can be generated through interpolation between a first neighboring sample located in the prediction direction of the intra prediction mode of the current block and a second neighboring sample located in the opposite direction of the prediction direction with respect to the predicted sample of the current block. In the above-described case, it can be called linear interpolation intra prediction (LIP). Also, the chroma predicted sample can be generated based on the luma samples using a linear model. In this case, it can be called the LM mode.
[0111] Also, a provisional prediction sample of the current block is derived based on the filtered peripheral reference samples, and at least one reference sample derived by an intra prediction mode among the existing peripheral reference samples, i.e., the unfiltered peripheral reference samples, and the provisional prediction sample are weighted-summed to derive a prediction sample of the current block. In the case described above, it can be called PDPC (Position dependent intra prediction).
[0112] Also, among the multiple reference sample lines around the current block, the reference sample line with the highest prediction accuracy is selected, and a prediction sample is derived using the reference sample located in the prediction direction on the corresponding line. At this time, the encoding of intra prediction can be performed by a method of indicating (signaling) the used reference sample line to the decoding device. In the case described above, it can be called multi-reference line (MRL) intra prediction or intra prediction based on MRL.
[0113] Also, the current block is divided into vertical or horizontal sub-partitions, and intra prediction is performed based on the same intra prediction mode, and peripheral reference samples can be derived and used in units of sub-partitions. That is, in this case, the intra prediction mode for the current block is applied to the sub-partitions in the same way, and by deriving and using peripheral reference samples in units of sub-partitions, the performance of intra prediction can be improved as appropriate. Such a prediction method can be called intra sub-partitions (ISP) or intra prediction based on ISP.
[0114] The intra prediction method described above can be classified into intra prediction modes and can be referred to as intra prediction types. The intra prediction types can be referred to by various terms such as intra prediction techniques or additional intra prediction modes. For example, the intra prediction type (or additional intra prediction mode, etc.) can include at least one of the LIP, PDPC, MRL, and ISP described above. A general intra prediction method excluding specific intra prediction types such as the LIP, PDPC, MRL, and ISP can be referred to as a normal intra prediction type. The normal intra prediction type can be generally applied when the specific intra prediction types as described above are not applicable, and prediction may be performed based on the intra prediction mode described above. On the other hand, post-processing filtering for the derived prediction samples may be performed as necessary.
[0115] On the other hand, in addition to the intra prediction types described above, matrix-based intra prediction (Matrix based intra prediction, hereinafter referred to as MIP) can be used as a method for intra prediction. MIP can be referred to as Affine linear weighted intra predictio (ALWIP) or Matrix weighted intra prediction (MWIP).
[0116] When MIP is currently applied to a block, i) using the surrounding reference samples for which an averaging procedure has been performed, ii) performing a matrix-vector-multiplication procedure, and iii) further performing a horizontal / vertical interpolation procedure as necessary, prediction samples for the current block can be derived. The intra prediction mode used for the MIP can be configured differently from the intra prediction modes used in the LIP, PDPC, MRL, and ISP intra predictions and the normal intra prediction described above.
[0117] The intra prediction mode for MIP can be called the "affine linear weighted intra prediction mode" or the matrix-based intra prediction mode. For example, by the intra prediction mode for MIP, the matrix and offset used in matrix-vector multiplication can be set differently. Here, the matrix can be called an (affine) weighted value matrix, and the offset can be called an (affine) offset vector or an (affine) bias vector. In this document, the intra prediction mode for MIP can be called the MIP intra prediction mode, the linear weighted intra prediction mode, the matrix weighted intra prediction mode, or the matrix based intra prediction mode. Specific MIP methods will be described later.
[0118] The following drawings are created to illustrate a specific example of this document. Since the names of specific devices described in the drawings and specific terms and names (such as the names of syntax, etc.) are presented exemplarily, the technical features of this document are not limited to the specific names used in the following drawings.
[0119] FIG. 4 shows an example of a video encoding method based on intra prediction to which the embodiment of this document can be applied, and FIG. 5 schematically shows the intra prediction unit in the encoding device. The intra prediction unit in the encoding device of FIG. 5 can be applied to be the same as or corresponding to the intra prediction unit 122 of the encoding device 100 in FIG. 1 described above.
[0120] Referring to FIGS. 4 and 5, S400 can be executed by the intra prediction unit 122 of the encoding device, and S410 can be executed by the residual processing unit 130 of the encoding device. Specifically, S410 can be executed by the subtraction unit 131 of the encoding device. In S420, the prediction information is derived by the intra prediction unit 122 and can be encoded by the entropy encoding unit 140. In S420, the residual information is derived by the residual processing unit 130 and can be encoded by the entropy encoding unit 140. The residual information is information regarding the residual samples. The residual information can include information regarding the quantized transform coefficients for the residual samples. As described above, the residual samples are derived as transform coefficients via the transform unit 132 of the encoding device, and the transform coefficients can be derived as quantized transform coefficients via the quantization unit 133. The information regarding the quantized transform coefficients can be encoded by the entropy encoding unit 140 via the residual coding procedure.
[0121] The encoding device performs intra prediction on the current block (S400). The encoding device can derive the intra prediction mode / type for the current block and derive the neighboring reference samples of the current block, and generate the predicted samples within the current block based on the intra prediction mode / type and the neighboring reference samples. Here, the determination of the intra prediction mode / type, the derivation of the neighboring reference samples, and the generation procedure of the predicted samples may be performed simultaneously, or any one of the procedures may be performed prior to the other procedures.
[0122] For example, the intra prediction unit 122 of the encoding device can include an intra prediction mode / type determination unit 122-1, a reference sample derivation unit 122-2, and a prediction sample derivation unit 122-3. The intra prediction mode / type determination unit 122-1 determines the intra prediction mode / type for the current block, the reference sample derivation unit 122-2 derives the peripheral reference samples of the current block, and the prediction sample derivation unit 122-3 can derive the prediction samples of the current block. On the other hand, although not shown, when a filtering procedure for prediction samples is performed, the intra prediction unit 122 may further include a prediction sample filter unit (not shown). The encoding device can determine the mode / type to be applied to the current block among a plurality of intra prediction modes / types. The encoding device can compare the RD cost for the intra prediction mode / type and determine the optimal intra prediction mode / type for the current block.
[0123] As described above, the encoding device can also perform a filtering procedure for prediction samples. The filtering of prediction samples can be referred to as post-filtering. By the filtering procedure of prediction samples, some or all of the prediction samples can be filtered. Depending on the case, the filtering procedure of prediction samples can be omitted.
[0124] The encoding device generates residual samples for the current block based on the (filtered) prediction samples (S410). The encoding device can compare the prediction samples with the original samples of the current block based on phase and derive the residual samples.
[0125] The encoding device can encode video information including information related to intra prediction (prediction information) and residual information related to residual samples (S420). The prediction information can include intra prediction mode information and intra prediction type information. The residual information can include the syntax of residual coding. The encoding device can convert / quantize the residual samples and derive quantized transform coefficients. The residual information can include information regarding the quantized transform coefficients.
[0126] The encoding device can output the encoded video information in the form of a bitstream. The output bitstream can be transmitted to the decoding device via a storage medium or a network.
[0127] As described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks). Therefore, the encoding device can perform inverse quantization / inverse transformation on the quantized transform coefficients again to derive (corrected) residual samples. The reason for performing inverse quantization / inverse transformation again after converting / quantizing the residual samples in this way is to derive the same residual samples as those derived from the decoding device, as described above. The encoding device can generate a reconstructed block including reconstructed samples for the current block based on the prediction samples and the (corrected) residual samples. Based on the reconstructed block, a reconstructed picture for the current picture can be generated. As described above, loop filtering procedures and the like can be further applied to the reconstructed picture.
[0128] FIG. 6 shows an example of a video decoding method based on intra prediction to which the embodiments of this document can be applied, and FIG. 7 schematically shows an intra prediction unit in a decoding apparatus. The intra prediction unit in the decoding apparatus of FIG. 7 can be applied so as to be identical or corresponding to the intra prediction unit 231 of the decoding apparatus 200 of FIG. 2 described above.
[0129] Referring to FIGS. 6 and 7, the decoding apparatus can execute operations corresponding to the operations executed by the encoding apparatus described above. S600 to S620 can be executed by the intra prediction unit 231 of the decoding apparatus, and the prediction information of S600 and the residual information of S630 can be obtained from the bitstream by the entropy decoding unit 210 of the decoding apparatus. The residual processing unit 220 of the decoding apparatus can derive residual samples for the current block based on the residual information. Specifically, the inverse quantization unit 221 of the residual processing unit 220 performs inverse quantization based on the quantized transform coefficients derived based on the residual information to derive the transform coefficients, and the inverse transform unit 222 of the residual processing unit performs inverse transform on the transform coefficients to derive residual samples for the current block. S640 can be executed by the addition unit 240 or the restoration unit of the decoding apparatus.
[0130] The decoding device can derive the intra prediction mode / type for the current block based on the received prediction information (intra prediction mode / type information) (S600). The decoding device can derive the surrounding reference samples of the current block (S610). The decoding device generates prediction samples within the current block based on the intra prediction mode / type and the surrounding reference samples (S620). In this case, the decoding device can perform a filtering procedure for the prediction samples. The filtering of the prediction samples may be referred to as post-filtering. By the filtering procedure of the prediction samples, some or all of the prediction samples can be filtered. Depending on the case, the filtering procedure of the prediction samples may be omitted.
[0131] The decoding device generates residual samples for the current block based on the received residual information (S630). The decoding device can generate restored samples for the current block based on the prediction samples and the residual samples, and derive a restored block including the restored samples (S640). Based on the restored block, a restored picture for the current picture can be generated. As described above, an in-loop filtering procedure or the like can be further applied to the restored picture.
[0132] Here, the intra prediction unit 231 of the decoding device can include an intra prediction mode / type determination unit 231-1, a reference sample derivation unit 231-2, and a prediction sample derivation unit 231-3. The intra prediction mode / type determination unit 231-1 determines the intra prediction mode / type for the current block based on the intra prediction mode / type information obtained from the entropy decoding unit 210. The reference sample derivation unit 231-2 derives the surrounding reference samples of the current block. The prediction sample derivation unit 231-3 can derive the prediction samples of the current block. On the other hand, although not shown, when the above-described filtering procedure for the prediction samples is performed, the intra prediction unit 231 can further include a prediction sample filtering unit (not shown).
[0133] The intra prediction mode information can include, for example, flag information (e.g., intra_luma_mpm_flag) indicating whether MPM (most probable mode) is applied to the current block or whether the remaining mode is applied. At this time, when MPM is applied to the current block, the prediction mode information can further include index information (e.g., intra_luma_mpm_idx) pointing to one of the candidates of the intra prediction mode (MPM candidates). The candidates of the intra prediction mode (MPM candidates) can be composed of an MPM candidate list or an MPM list. Also, when MPM is not applied to the current block, the intra prediction mode information can further include remaining mode information (e.g., intra_luma_mpm_remainder) pointing to one of the remaining intra prediction modes excluding the candidates of the intra prediction mode (MPM candidates). The decoding device can determine the intra prediction mode of the current block based on the intra prediction mode information.
[0134] In addition, the intra prediction type information can be embodied in various forms. As an example, the intra prediction type information can include index information of the intra prediction type that indicates one of the intra prediction types. As another example, the intra prediction type information includes reference sample line information (e.g., intra_luma_ref_idx) that indicates whether the MRL is applied to the current block and, if the MRL is applied, which reference sample line is used, ISP flag information (e.g., intra_subpartitions_mode_flag) that indicates whether the ISP is applied to the current block, ISP type information (e.g., intra_subpartitions_split_flag) that indicates the split type of the subpartition when the ISP is applied, flag information indicating whether PDCP can be applied, or flag information indicating whether LIP can be applied. In addition, the intra prediction type information can include an MIP flag that indicates whether MIP is applied to the current block.
[0135] The intra prediction mode information and / or the intra prediction type information described above can be encoded / decoded through the coding method described in this document. For example, the intra prediction mode information and / or the intra prediction type information described above can be encoded / decoded through entropy coding (e.g., CABAC, CAVLC) based on truncated (rice) binary code.
[0136] On the one hand, when intra prediction is applied, the intra prediction mode applied to the current block can be determined using the intra prediction modes of the surrounding blocks. For example, the decoding device can select one of the mpm (most probable mode) candidates in the mpm list derived based on the intra prediction modes of the surrounding blocks (e.g., the left and / or upper surrounding blocks) of the current block and additional candidate modes according to the received mpm index, or can select one of the remaining intra prediction modes not included in the mpm candidates (and the planar mode) based on the remaining intra prediction mode information. The mpm list can be configured to include or not include the planar mode as a candidate. For example, when the mpm list includes the planar mode as a candidate, the mpm list may have 6 candidates, and when the mpm list does not include the planar mode as a candidate, the mpm list may have 5 candidates. When the mpm list does not include the planar mode as a candidate, a not planar flag (e.g., intra_luma_not_planar_flag) indicating whether the intra prediction mode of the current block is not the planar mode can be signaled. For example, if the mpm flag is signaled first, the mpm index and the not planar flag can be signaled when the value of the mpm flag is 1. Also, the mpm index can be signaled when the value of the not planar flag is 1. Here, the reason for configuring the mpm list not to include the planar mode as a candidate is to signal the flag (not planar flag) first and check whether it is the planar mode first because the planar mode is always considered as an mpm rather than not being an mpm.
[0137] For example, whether the intra prediction mode currently applied to a block is among the mpm candidates (and the planar mode) or among the remaining modes can be indicated based on the mpm flag (e.g., intra_luma_mpm_flag). A value of 1 for the mpm flag can indicate that the intra prediction mode for the current block is within the mpm candidates (and the planar mode), and a value of 0 for the mpm flag can indicate that the intra prediction mode for the current block is not within the mpm candidates (and the planar mode). A value of 0 for the not planar flag (e.g., intra_luma_not_planar_flag) can indicate that the intra prediction mode for the current block is the planar mode, and a value of 1 for the not planar flag can indicate that the intra prediction mode for the current block is not the planar mode. The mpm index can be signaled in the form of a syntax element of mpm_idx or intra_luma_mpm_idx, and the remaining intra prediction mode information can be signaled in the form of a syntax element of rem_intra_luma_pred_mode or intra_luma_mpm_remainder. For example, the remaining intra prediction mode information can index, in the order of prediction mode numbers, the remaining intra prediction modes not included in the mpm candidates (and the planar mode) among all the intra prediction modes and point to one of them. The intra prediction mode can be the intra prediction mode for the luma component (samples). Hereinafter, the intra prediction mode information can include at least one of the mpm flag (e.g., intra_luma_mpm_flag), the not planar flag (e.g., intra_luma_not_planar_flag), the mpm index (e.g., mpm_idx or intra_luma_mpm_idx), and the remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder).In this document, the mpm list can be referred to by various terms such as an mpm candidate list, a candidate mode list (candModeList), a candidate intra prediction mode list, and the like.
[0138] Generally, when a video is block - divided, the current block to be coded and the surrounding blocks will have similar video characteristics. Therefore, the current block and the surrounding blocks are likely to be identical to each other or have similar intra - prediction modes. Thus, the encoder can use the intra - prediction mode of the surrounding blocks to encode the intra - prediction mode of the current block. For example, the encoder / decoder can construct an MPM (most probable modes) list for the current block. The MPM list can also be referred to as an MPM candidate list. Here, MPM can mean a mode that is used to improve coding efficiency by considering the similarity between the current block and the surrounding blocks during the coding of the intra - prediction mode.
[0139] FIG. 8 shows an example of an intra - prediction mode to which the embodiments of this document are applicable.
[0140] Referring to FIG. 8, it is possible to distinguish between an intra prediction mode having horizontal directionality centered on the 34th intra prediction mode with a prediction direction of the upper left diagonal, and an intra prediction mode having vertical directionality. H and V in FIG. 8 respectively mean horizontal directionality and vertical directionality, and the numbers from -32 to 32 indicate displacements in units of 1 / 32 on the sample grid position. The 2nd to 33rd intra prediction modes have horizontal directionality, and the 34th to 66th intra prediction modes have vertical directionality. The 18th intra prediction mode and the 50th intra prediction mode respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. The 2nd intra prediction mode can be called a lower left diagonal intra prediction mode, the 34th intra prediction mode can be called an upper left diagonal intra prediction mode, and the 66th intra prediction mode can be called an upper right diagonal intra prediction mode.
[0141] On the other hand, the intra prediction mode used for the aforementioned MIP is not an existing directional mode, but can indicate the matrix and offset used for intra prediction. That is, through the intra mode for MIP, the matrix and offset for intra prediction can be derived. In this case, when deriving the intra mode for generating the aforementioned normal intra prediction or MPM list, the MIP and the intra prediction mode of the predicted block can be set to a preset mode, for example, a planar mode or a DC mode. Alternatively, by another example, based on the block size, the intra mode for MIP can also be mapped to a planar mode, a DC mode, or a directional intra mode.
[0142] Hereinafter, MIP (Matrix based intra prediction), which is one method of intra prediction, will be described.
[0143] As described above, matrix-based intra prediction (MIP) can be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). To predict samples of a rectangular block having a width (W) and a height (H), MIP uses one H-line of the left boundary samples of the restored periphery of the block and one W-line of the upper boundary samples of the restored periphery of the block as input values. If restored samples are not available, reference samples can be generated by the interpolation method applied in normal intra prediction.
[0144] FIG. 9 is a diagram for explaining a procedure for generating MIP-based prediction samples according to an example. Referring to FIG. 9, the MIP procedure is described as follows.
[0145] 1. Averaging process
[0146] Among the boundary samples, when W = H = 4, 4 samples, and in all other cases, 8 samples are extracted by the averaging process.
[0147] 2. Matrix vector multiplication process
[0148] Matrix-vector multiplication is performed with the averaged samples as input, and subsequently an offset is added. Through such operations, reduced prediction samples for the subsampled sample set within the original block can be derived.
[0149] 3. (Linear) Interpolation process
[0150] Prediction samples at the remaining positions are generated from the prediction samples of the subsampled set of samples by linear interpolation, which is a linear interpolation of single steps in each direction.
[0151] The matrices and offset vectors necessary to generate the prediction block or prediction samples can be selected from three sets S0, S1, S2 for the matrix.
[0152] Set S0 consists of 16 matrices A0 i , where i ∈ {0, …, 15}, and each matrix can consist of 16 rows, 4 columns, and 16 offset vectors b0 i , where i ∈ {0, …, 15}. The matrices and offset vectors of set S0 can be used for blocks of size 4×4. According to another example, set S0 can also include 18 matrices.
[0153] Set S1 consists of 8 matrices A1 i , where i ∈ {0, …, 7}, and each matrix can consist of 16 rows, 8 columns, and 8 offset vectors b1 i , where i ∈ {0, …, 7}. According to another example, set S1 can also include 6 matrices. The matrices and offset vectors of set S1 can be used for blocks of size 4×8, 8×4, and 8×8. Alternatively, the matrices and offset vectors of set S1 can be used for blocks of size 4×H or W×4.
[0154] Finally, set S2 consists of 6 matrices A2 i , where i ∈ {0, …, 5}, and each matrix can consist of 64 rows, 8 columns, and 6 offset vectors b2 ican be composed of i ∈ {0, …, 5}. The matrix and offset vector of set S2, or a part thereof, can be used in block forms of all other sizes where sets S0 and S1 are not applicable. For example, the matrix and offset vector of set S2 can be used for operations on blocks with a height and width of 8 or more.
[0155] The number of multiplications required for matrix-vector multiplication is always less than or equal to 4×W×H. That is, in the MIP mode, a maximum of 4 multiplications per sample are required.
[0156] Hereinafter, a general MIP procedure will be outlined. The remaining blocks not described below can be processed in any of the four cases described.
[0157] FIGS. 10 to 13 are diagrams showing the MIP procedure according to the block size. FIG. 10 is a diagram showing the MIP procedure for a 4×4 block, FIG. 11 is a diagram showing the MIP procedure for an 8×8 block, FIG. 12 is a diagram showing the MIP procedure for an 8×4 block, and FIG. 13 is a diagram showing the MIP procedure for a 16×16 block.
[0158] As shown in FIG. 10, when a 4×4 block is given, MIP takes the average of two samples along each axis of the boundary. As a result, four input samples become the input values for matrix-vector multiplication, and the matrix is taken from set S0. When the offset is added, 16 final predicted samples are generated. In the case of a 4×4 block, linear interpolation is not required to generate the predicted samples. Therefore, any of them can perform (4×16) / (4×4) = 4 multiplications per sample.
[0159] As shown in FIG. 11, when an 8×8 block is given, the MIP takes the average of four samples along each axis of the boundary. As a result, eight input samples become the input values for the matrix-vector multiplication, and the matrix is taken from set S1. The matrix-vector multiplication generates 16 samples at odd positions.
[0160] In the case of an 8×8 block, to generate predicted samples, a multiplication of (8×16) / (8×8)=2 is performed for each sample. After adding the offset, the reduced upper boundary samples are used to interpolate the samples vertically, and the original left boundary samples are used to interpolate horizontally. In this case, since the interpolation procedure does not require a multiplication operation, a total of two multiplication operations per sample are required for the MIP.
[0161] As shown in FIG. 12, when an 8×4 block is given, the MIP takes the average of four samples along the horizontal axis of the boundary and uses the four sample values of the left boundary for the vertical axis. As a result, eight input samples become the input values for the matrix-vector multiplication, and the matrix is taken from set S1. The matrix-vector multiplication generates 16 samples at odd horizontal positions and the corresponding vertical positions.
[0162] In the case of an 8×4 block, to generate predicted samples, a multiplication of (8×16) / (8×4)=4 is performed for each sample. After adding the offset, the original left boundary samples are used to interpolate horizontally. In this case, since the interpolation procedure does not require a multiplication operation, a total of four multiplication operations per sample are required for the MIP.
[0163] As shown in FIG. 13, when a 16×16 block is given, the MIP takes the average of four samples along each axis. As a result, eight input samples become the input values for the matrix-vector multiplication, and the matrix is taken from set S2. Through the matrix-vector multiplication, 64 samples are generated at odd positions. For the case of a 16×16 block, for each sample, a multiplication of (8×64) / (16×16)=2 is performed to generate the predicted samples. After adding the offset, eight reduced upper boundary samples are used to interpolate the samples vertically, and the original left boundary samples are used to interpolate horizontally. In this case, since no multiplication operation is required in the interpolation procedure, a total of two multiplication operations are required per sample for MIP.
[0164] For larger blocks, the MIP procedure is essentially the same as the procedure described in detail, and it can be easily confirmed that the number of multiplications per sample is less than 4.
[0165] For a W×8 block where the width is greater than 8 (W>8), since samples are generated at odd horizontal and each vertical position, only horizontal interpolation is required. In this case, a multiplication of (8×64) / (W×8)=64 / W is performed per sample for the prediction operation of the reduced samples. When W = 16, no additional multiplication is required for linear interpolation, and when W>16, the number of additional multiplications per sample required for linear interpolation is less than 2. That is, the total number of multiplications per sample is less than or equal to 4.
[0166] Also, for a W×4 block where the width is greater than 4 (W>4), the matrix A generated by omitting all rows corresponding to odd entries along the horizontal axis of the downsampled block kLet it be so. Therefore, the output size is 32, and only horizontal interpolation is performed. For the prediction operation of the downsampled samples, a multiplication of (8×32) / (W×4)=64 / W per sample is performed. When W = 16, no additional multiplication is required, and when W>16, the number of additional multiplications per sample required for linear interpolation is less than 2. That is, the total number of multiplications per sample is less than or equal to 4.
[0167] When the matrix is prefixed, it can thereby be processed.
[0168] FIG. 14 is a diagram for explaining the boundary averaging procedure among the MIP procedures. Referring to FIG. 14, the averaging procedure will be specifically described.
[0169] According to the averaging procedure, the averaging is applied to each boundary, that is, the left boundary or the upper boundary. Here, the boundary indicates the peripheral reference samples adjacent to the boundary of the current block as shown in FIG. 16. For example, the left boundary (bdry left ) indicates the left peripheral reference samples adjacent to the left boundary of the current block, and the upper boundary (bdry top ) indicates the upper peripheral reference samples adjacent to the upper side.
[0170] When the current block is a 4×4 block, the size of each boundary can be reduced to 2 samples through the averaging procedure. If the current block is not a 4×4 block, the size of each boundary can be reduced to 4 samples through the averaging procedure.
[0171] The first step of the averaging procedure is to reduce the input boundaries (bdry left and bdry top ) to smaller boundaries JPEG2025105664000002.jpg10127. JPEG2025105664000003.jpg11125 consists of 2 samples in the case of a 4×4 block and 4 samples in all other different cases.
[0172] For a 4×4 block, for 0≦i<2 JPEG2025105664000004.jpg13150 can be expressed by the following formula, JPEG2025105664000005.jpg13143 can also be defined similarly.
[0173]
Number
[0174] On the other hand, for 0≦i<4, when the width of the block is given as W = 4×2 k and given as, JPEG2025105664000007.jpg14156 can be expressed by the following formula, JPEG2025105664000008.jpg17151 can also be defined similarly.
[0175]
Number
[0176] Two reduced boundaries JPEG2025105664000010.jpg12121 is concatenated to the reduced boundary vector bdry red so bdry red is of size 4 for a 4×4 block and of size 8 for all other blocks.
[0177] When referring to "mode" as the MIP mode, the range of the reduced boundary vector bdry red and the MIP mode value (mode) can be defined based on the block size and the intra_mip_transposed_flag value as in the following formula.
[0178]
Number
[0179] In the above formula, intra_mip_transposed_flag can be referred to as MIP transpose, and such flag information can indicate whether the downsampled prediction samples are transposed. The semantics for such a syntax element can be expressed as "intra_mip_transposed_flag[x0][y0] specifies whether the input vector for matrix-based intra prediction mode for luma samples is transposed or not."
[0180] Finally, for the interpolation of the subsampled prediction samples, in a large block, a second version of the averaged boundary is required. That is, if the smaller value of the width and height is greater than 8 JPEG2025105664000012.jpg14148 If the width is equal to or greater than the height (W≧H), W = 8 * 2 l and for 0 ≦ i < 8 JPEG2025105664000013.jpg14154 can be defined as in the following formula. Also, if the smaller value of the width and height is greater than 8 JPEG2025105664000014.jpg12142 and the height is greater than the width (H>W) JPEG2025105664000015.jpg14145 can also be defined in the same way.
[0181]
Number
[0182] Next, let's look at the generation procedure of the downsampled prediction samples by matrix-vector multiplication.
[0183] Reduced input vector bdry red One of which is the reduced prediction sample generates JPEG2025105664000017.jpg12156. The prediction sample is a signal for the downsampled block of width W red and height H red Here, W red and H red are defined as follows.
[0184]
Equation
[0185] Reduced prediction sample JPEG2025105664000019.jpg15163 can be calculated by adding an offset after the matrix-vector multiplication operation and can be derived through the following formula.
[0186]
Equation
[0187] Here, A is a matrix with W red ×h red rows and 4 columns when W and H are 4 (W = H = 4) and 8 columns in all other cases, and b is a vector of size W red ×h red .
[0188] Matrix A and vector b are selected from the following sets S0, S1, S2, and the index idx = idx(W, H) can be defined as in Equation 7 or Equation 8.
[0189]
Equation
[0190]
Number
[0191] Whether idx is 1 or less (idx ≦ 1), or idx is 2 and the smaller value of W and H is greater than 4 JPEG2025105664000023.jpg14144, A is JPEG2025105664000024.jpg13137, b is Set to JPEG2025105664000025.jpg12132. When idx is 2 and the smaller value of W and H is 4 JPEG2025105664000026.jpg11140, when W is 4, A corresponds to the odd x - coordinates in the down - sampled block JPEG2025105664000027.jpg12139 becomes a matrix with each row removed. Or, when H is 4, A corresponds to the odd y - coordinates in the down - sampled block JPEG2025105664000028.jpg12141 becomes a matrix with each column removed.
[0192] Finally, the down - sized predicted sample can be replaced by its transpose in Equation 9.
[0193]
Number
[0194] For the calculation of JPEG2025105664000030.jpg11137, when W = H = 4, since A is composed of 4 columns and 16 rows, the number of multiplications required is 4. In all other cases, since A is composed of 8 columns and W red ×h red rows, so For the calculation of JPEG2025105664000031.jpg10127, it can be confirmed that a maximum of 4 multiplications per sample are required.
[0195] FIG. 15 is a diagram for explaining linear interpolation in the MIP procedure. Referring to FIG. 15, the linear interpolation procedure will be specifically described as follows.
[0196] The interpolation procedure may be referred to as a linear interpolation or a bilinear interpolation procedure. The interpolation procedure can include two steps: 1) vertical interpolation and 2) horizontal interpolation, as shown.
[0197] If W >= H, vertical linear interpolation can be applied first, followed by horizontal linear interpolation. If W < H, horizontal linear interpolation can be applied first, followed by vertical linear interpolation. In the case of a 4×4 block, the interpolation procedure may be omitted.
[0198] For a W×H block where JPEG2025105664000032.jpg11133, the predicted sample is W red ×H red The reduced predicted sample above is derived from JPEG2025105664000033.jpg13149. Depending on the form of the block, linear interpolation is performed vertically, horizontally, or in both directions. When linear interpolation is applied in both directions, if W < H, it is applied first in the horizontal direction; otherwise, it is applied first in the vertical direction.
[0199] For a W×H block where JPEG2025105664000034.jpg13143 and W >= H, it can be considered that there is no loss of generality. Then, one-dimensional linear interpolation is performed as follows. If there is no loss of generality, the linear interpolation for the vertical direction is sufficiently explained.
[0200] First, the reduced predicted sample is extended upward by the boundary signal. The coefficient of vertical upsampling is defined as JPEG2025105664000035.jpg13134, When setting 14130 for JPEG2025105664000036.jpg, the extended downsampling prediction samples can be set as in the following formula.
[0201]
Equation
[0202] Subsequently, the vertical linear interpolation prediction samples can be generated from such extended downsampling prediction samples according to the following formula.
[0203]
Equation
[0204] Here, x satisfies 0 ≦ x < W red , y satisfies 0 ≦ y < H red , and k satisfies 0 ≦ k < U ver and can hold.
[0205] In the following, we will look at a method to reduce the complexity and maximize the performance for the MIP technique. The embodiments described later may be executed independently or in combination.
[0206] On the other hand, when MIP is applied to the current block, an MPM list for the current block to which the MIP is applied can be configured separately. The MPM list can be called by various names such as MIP MPM list (or LWIP MPM list, candLwipModeList) to distinguish it from the MPM list when ALWIP is not applied to the current block. Hereinafter, for the sake of distinction, it is expressed as the MIP MPM list, but it can also be called the MPM list.
[0207] The MIP MPM list can include n candidates. For example, n can be 3. The MIP MPM list can be configured based on the left and upper neighboring blocks of the current block. Here, the left neighboring block can indicate the uppermost block among the neighboring blocks adjacent to the left boundary of the current block. Also, the upper neighboring block can indicate the leftmost block among the neighboring blocks adjacent to the upper boundary of the current block.
[0208] For example, when MIP is applied to the left neighboring block, the first candidate intra prediction mode (or candLwipModeA) can be set to be the same as the MIP mode of the left neighboring block. Also, for example, when MIP is applied to the upper neighboring block, the second candidate intra prediction mode (or candLwipModeB) can be set to be the same as the prediction mode of the MIP mode of the upper neighboring block.
[0209] On the other hand, the left peripheral block and the upper peripheral block can be coded based on an intra prediction other than MIP. That is, when coding the left peripheral block or the upper peripheral block, another intra prediction type other than MIP can be applied. In this case, it is not appropriate to directly use the general intra prediction mode number of the peripheral block (left peripheral block / upper peripheral block) where MIP is not applied as the candidate intra mode for the current block where MIP is applied. Therefore, in this case, as an example, the MIP mode of the peripheral block (left peripheral block / upper peripheral block) where MIP is not applied can be regarded as a prediction mode of the MIP mode with a specific value (for example, 0, 1, or 2, etc.). Alternatively, as another example, the general intra prediction mode of the peripheral block (left peripheral block / upper peripheral block) where MIP is not applied can be mapped to the MIP mode based on a predetermined mapping table and used for the configuration of the MIP MPM list. In this case, the mapping can be executed based on the block size type of the current block.
[0210] Also, even if MIP is applied, if the peripheral block (for example, the left peripheral block / upper peripheral block) is not available (for example, located outside the current picture, located outside the current tile / tile group, etc.), a MIP mode that is not available for the current block can also be used according to the block size type. In this case, for the first candidate and / or the second candidate, a specific MIP mode defined in advance can be used as the first candidate intra prediction mode or the second candidate intra prediction mode. Also, for the third candidate, a specific MIP prediction mode defined in advance can be used as the third candidate intra prediction mode.
[0211] On the other hand, the existing MIP mode is classified into a non-MPM mode and an MPM mode in the same way as the derivation method of the existing intra prediction mode, and an MPM flag is sent, and the MIP mode of the current block is coded based on the MPM mode or the non-MPM mode.
[0212] According to one example, for a block to which the MIP technique is applied, a structure can be proposed that directly codes in the MIP mode without distinguishing between the MPM mode and the non-MPM mode. In such a video coding structure, a complex syntax structure can be simplified. Also, since the actual occurrence frequency for the MIP mode is relatively uniformly distributed for each mode and is clearly different from the occurrence frequency shown in the existing intra modes, the efficiency of encoding and decoding MIP mode information can be maximized through the proposed coding structure.
[0213] The video information transmitted and received for MIP according to this embodiment is as follows. The syntax described later may be included in the video / video information transmitted from the encoding device to the decoding device, can be configured / encoded by the encoding device, and can be signaled to the decoding device in the form of a bitstream. The decoding device can parse / decode the information (syntax elements) included according to the conditions / order disclosed in the syntax.
[0214] [Table 1]
[0215] As shown in Table 1, the syntax intra_mip_flag and intra_mip_mode_idx for the MIP mode for the current block can be included in the syntax information for the coding unit and signaled.
[0216] When intra_mip_flag is 1, it indicates that the intra prediction type for luma samples is matrix-based intra prediction, and when its value is 0, it indicates that the intra prediction type for luma samples is not matrix-based intra prediction.
[0217] When intra_mip_flag is 1 and intra_mip_mode_idx is signaled, it indicates an intra prediction mode based on a matrix for luma samples. Such an intra prediction mode based on a matrix can indicate a matrix and an offset or a matrix for MIP, as described above.
[0218] Also, by way of an example, flag information, such as intra_mip_transposed_flag, can be further signaled via the syntax of a coding unit to indicate whether an input vector for intra prediction based on a matrix is transposed. When intra_mip_transposed_flag is 1, the input vector for intra prediction based on a matrix is transposed, and such flag information can reduce the number of matrices for intra prediction based on a matrix.
[0219] On the other hand, intra_mip_mode_idx can be encoded and decoded in a Truncated Binarization method as shown in the following table.
[0220]
Table 2
[0221] As shown in Table 2, intra_mip_flag is binary-coded with a Fixed Length Code, while intra_mip_mode_idx is binary-coded in a Truncated Binary Coding manner, and the maximum length (cMax) of the binary coding can be set according to the size of the coding block. The maximum length of the binary coding is set to 34 when the width and height of the coding block are 4 (cbWidth == 4 && cbHeight == 4), and otherwise, it can be set to 18 or 10 depending on whether the width and height of the coding block are 8 or less ((cbWidth <= 8 && cbHeight <= 8)?).
[0222] On the other hand, intra_mip_mode_idx can be coded in a bypass manner instead of being based on a context model. By coding in a bypass manner, the coding speed and efficiency can be increased.
[0223] In another example, when intra_mip_mode_idx is binary-coded in a Truncated Binary Coding manner, the maximum length of the binary coding is as shown in the following table.
[0224]
Table 3
[0225] As shown in Table 3, the maximum length of the binary evolution for intra_mip_mode_idx is set to 15 when the width and height of the coding block are 4 ((cbWidth == 4 && cbHeight == 4)), and otherwise, when either the width or height of the coding block is 4 ((cbWith == 4 || cbHeight == 4)) or the width and height of the coding block are 8 (cbWith == 8 && cbHeight == 8), it is set to 7, and when either the width or height of the coding block is 4 ((cbWith == 4 || cbHeight == 4)) or the width and height of the coding block are not 8 (cbWith == 8 && cbHeight == 8), it can be set to 5.
[0226] In another example, intra_mip_mode_idx can be encoded with a Fixed Length Code. At this time, the number of MIP modes available to improve the encoding efficiency is limited to an exponential number of 2 for each block size (e.g., A = 2 K1 - 1, B = 2 K2 - 1, 2 K3 - 1, where K1, K2, K3 are positive constants) can also be the case.
[0227] Shown in a table, it is as follows.
[0228]
Table 4
[0229] In Table 4, when K1 = 5, K2 = 4, and K3 = 3, intra_mip_mode can be evolved as follows in binary.
[0230]
Table 5
[0231] Alternatively, by way of example, when K1 is set to 4, the maximum length of the binary evolution for the intra_mip_mode of a block with a coding block width and height of 4 can be set to 15. Also, when K2 is set to 3, when the width or height of the coding block is 4, or when the width and height of the coding block are 8, the maximum length of the binary evolution can be set to 7.
[0232] On the other hand, by way of example, a method can be proposed to use MIP only for specific blocks where the MIP technique can be efficiently applied. When the method according to this embodiment is applied, the number of matrix vectors required for MIP decreases, and the memory required to store the matrix vectors can be significantly reduced (by 50%). Despite such an effect, the coding efficiency is maintained almost constant (less than 0.1%).
[0233] The syntax including the specific conditions under which the MIP technique is applied according to this embodiment is as follows in the following table.
[0234]
Table 6
[0235] As shown in Table 6, a condition (cbWidth>K1||cbHeight>K2) is added so that MIP is applied only to large blocks, and the size of the block can be determined by preset values (K1 and K2). The reason for applying MIP only to large blocks is that the coding efficiency of MIP is shown in relatively large blocks.
[0236] The following table shows an example where K1 and K2 are predefined as 8 in Table 6.
[0237]
Table 7
[0238] The semantics for intra_mip_flag and intra_mip_mode_idx in Tables 6 and 7 are the same as those in Table 1.
[0239] On the other hand, when intra_mip_mode_idx has 11 possible modes, it can be encoded in a truncated binary evolution (cMax = 10) manner as follows.
[0240] [Table 8]
[0241] Alternatively, when the available MIP modes are limited to 8, intra_mip_mode_idx[x0][y0] can be encoded with a fixed-length code as follows.
[0242] [Table 9]
[0243] In both Tables 8 and 9, intra_mip_mode_idx can be coded in a bypass manner.
[0244] On the other hand, by way of an example, a weighted value matrix (A k ) and an offset vector (b k ) used for large blocks so that the MIP technique can be efficiently applied from the perspective of memory saving can be proposed for small blocks. When the method according to this embodiment is applied, the number of matrix vectors required for MIP can be reduced, and the memory required to store the matrix vectors can be greatly reduced (by 50%). While such an effect is accompanied, the coding efficiency is maintained almost constant (less than 0.05%).
[0245] FIG. 16 is a diagram for explaining the MIP technique according to an example of this document.
[0246] As shown, (a) of FIG. 16 shows the operations of the matrix and the offset vector for the index i of the large block, and (b) of FIG. 16 shows the operations of the sampled matrix and the offset vector applied to the small block.
[0247] As shown in FIG. 16, the subsampled weighted value matrix (Sub(A k )) obtained by subsampling the weighted value matrix used for the large block and the offset vector (Sub(b k )) obtained by subsampling the offset vector used for the large block can be regarded as the weighted value matrix and the offset vector for the small block respectively, and the existing MIP procedure can be applied.
[0248] Here, subsampling may be applied only to either the horizontal direction or the vertical direction, or may be applied to both directions. In particular, the subsampling factor (for example, 1 out of 2, or 1 out of 4) and the sampling direction for vertical or horizontal can be set based on the width (Width) and height (Height) of the corresponding block.
[0249] Also, the number of intra prediction modes for MIP to which this embodiment is applied can be set differently based on the size of the current block. For example, i) when the height and width of the current block (coding block or transform block) are both 4, 35 intra prediction modes (that is, intra prediction modes 0 to 34) may be available; ii) when the height and width of the current block are both 8 or less, 19 intra prediction modes (that is, intra prediction modes 0 to 18) may be available; iii) in other cases, 11 intra prediction modes (that is, intra prediction modes 0 to 10) may be available.
[0250] For example, when the height and width of the current block are both 4, it is called block size type 0; when the height and width of the current block are both 8 or less, it is called block size type 1; and in other cases, it is called block size type 2. When this is the case, the number of intra prediction modes for MIP can be organized as shown in the following table.
[0251]
Table 10
[0252] To apply the weighted value matrix and offset vector used for large blocks (for example, block size type = 2) to small blocks (for example, block size = 0 or block size = 1), the number of available intra prediction modes for each block size can be applied in the same way as shown in the following table.
[0253]
Table 11
[0254] Alternatively, as shown in Table 12 below, apply MIP only to block size types 1 and 2. For block size type 1, the weighted value matrix and offset vector defined for block size type 2 can be subsampled and used. Through this, memory can be efficiently saved (50%).
[0255]
Table 12
[0256] The following drawings were created to illustrate a specific example of this specification. Since the names of the specific devices and the names of the specific signals / messages / fields described in the drawings are presented exemplarily, the technical features of this specification are not limited to the specific names used in the following drawings.
[0257] The following table shows the experimental results when, as in the above-described embodiment, the derivation of the MPM is omitted when applying the MIP to a block currently, and signals related to the MPM are also not signaled.
[0258] The experiment was conducted based on the reference software of VTM 5.0 under the normal test conditions defined in JVET-N1010.
[0259] [Table 13]
[0260] [Table 14]
[0261] [Table 15]
[0262] FIG. 17 is a flowchart schematically showing a decoding method that can be executed by a decoding apparatus according to an embodiment of the present document.
[0263] The method disclosed in FIG. 17 can be executed by the decoding apparatus 200 disclosed in FIG. 2. Specifically, steps S1700 to S1750 in FIG. 17 can be executed by the entropy decoding unit 210 and / or the prediction unit 230 (specifically, the intra prediction unit 231) disclosed in FIG. 2, and step S1760 in FIG. 17 can be executed by the addition unit 240 disclosed in FIG. 2. Further, the method disclosed in FIG. 17 can include the embodiments described above in this document. Therefore, in FIG. 17, specific descriptions of the content overlapping with the above-described embodiments will be omitted or simplified.
[0264] Referring to FIG. 17, the decoding device can receive, i.e., obtain from the bitstream, flag information indicating whether matrix-based intra prediction (MIP) is used for the current block in the bitstream (S1700).
[0265] Such flag information can be signaled included in the syntax information of the coding unit in a syntax such as intra_mip_flag.
[0266] Based on the received flag information, the decoding device can receive matrix-based intra prediction (MIP) mode information (S1710).
[0267] The MIP mode information can be represented by intra_mip_mode_idx and can be signaled when intra_mip_flag is 1. Intra_mip_mode_idx can be index information indicating the MIP mode applied to the current block, and such index information can be used to derive the matrix when generating prediction samples.
[0268] Also, by way of example, flag information indicating whether an input vector for matrix-based intra prediction is transposed, e.g., intra_mip_transposed_flag, can be further signaled via the syntax of the coding unit.
[0269] The decoding device can generate intra prediction samples for the current block based on the MIP information. The decoding device can derive at least one of the surrounding reference samples of the current block to generate prediction samples based on the surrounding reference samples in order to generate the intra prediction samples.
[0270] The decoding device can evolve the bit string of the syntax element for the MIP mode information by the truncated binary evolution method and decode it by the bypass method (S1720).
[0271] As described above, the maximum length of the bit string of the syntax element for the MIP mode information can be set to different values according to the size of the current block. Such a maximum length can be set to three values according to the size of the current block as in Table 2 or Table 3, and when the width and height of the current block are 4, the maximum length may be the largest.
[0272] For example, in Table 2, when the width and height of the coding block are 4, the maximum length of the evolution is set to 34, and in Table 3, it can be set to 15.
[0273] Also, such MIP mode information can be decoded by the bypass method instead of the context modeling method.
[0274] On the other hand, intra_mip_flag can be evolved by a Fixed Length Code.
[0275] When MIP is applied, the decoding device can downsample the reference samples adjacent to the current block and derive the reduced boundary samples (S1730).
[0276] The reduced boundary samples can be derived by being downsampled by averaging the reference samples.
[0277] When the width and height of the current block are 4, four reduced boundary samples are derived, and in other cases, eight samples can be derived.
[0278] The averaging procedure for downsampling can be applied to each boundary of the current block, the left boundary or the upper boundary, which can be applied to the peripheral reference samples adjacent to the boundary of the current block.
[0279] By way of example, if the current block is a 4×4 block, the size of each boundary can be reduced to 2 samples via the averaging procedure, and if the current block is not a 4×4 block, the size of each boundary can be reduced to 4 samples via the averaging procedure.
[0280] Thereafter, the decoding device can derive the reduced prediction samples based on the multiplication operation of the MIP matrix derived based on the size of the current block and the index information and the reduced boundary samples (S1740).
[0281] The MIP matrix can be derived based on the size of the current block and the received index information.
[0282] The MIP matrix can be selected from any of three matrix sets classified according to the size of the current block, and each of the three matrix sets can include a plurality of MIP matrices.
[0283] That is, three matrix sets for MIP can be set, and each matrix set can be composed of a plurality of matrices and offset vectors. Such matrix sets can be classified and applied according to the size of the current block.
[0284] For example, for a 4×4 block, a matrix set including 18 or 16 matrices composed of 16 rows and 4 columns and 18 or 16 offset vectors can be applied. The index information can be information indicating any of the plurality of matrices included in one matrix set.
[0285] For 4×8, 8×4, and 8×8 blocks, or 4×H or W×4 blocks, a matrix set including 10 or 8 matrices composed of 16 rows and 8 columns and 10 or 8 offset vectors can be applied.
[0286] Alternatively, for blocks other than the aforementioned blocks or blocks with a height and width of 8 or more, a matrix set including 6 matrices composed of 64 rows and 8 columns and 6 offset vectors can be applied.
[0287] After the operation of multiplying the MIP matrix by the reduced boundary samples, a reduced predicted sample, that is, a predicted sample to which the MIP matrix is applied, is derived based on the operation of adding an offset.
[0288] The decoding device can upsample the reduced predicted sample to generate an intra prediction sample for the current block (S1750).
[0289] The intra prediction sample can be upsampled by linear interpolation of the reduced predicted sample.
[0290] The interpolation procedure can be referred to as a linear interpolation or bilinear interpolation procedure and can include two steps: 1) vertical interpolation and 2) horizontal interpolation.
[0291] If W>=H, vertical linear interpolation can be applied first, followed by horizontal linear interpolation. If W<H, horizontal linear interpolation can be applied first, followed by vertical linear interpolation. In the case of a 4×4 block, the interpolation procedure can be omitted.
[0292] The decoding device can generate a restored sample for the current block based on the predicted sample (S1760).
[0293] As one embodiment, the decoding device may immediately use a predicted sample as a restored sample according to a prediction mode, or may generate a restored sample by adding a residual sample to the predicted sample.
[0294] When there is a residual sample for the current block, the decoding device can receive information regarding the residual for the current block. The information regarding the residual can include a transform coefficient regarding the residual sample. The decoding device can derive a residual sample (or a residual sample array) for the current block based on the residual information. The decoding device can generate a restored sample based on the predicted sample and the residual sample, and can derive a restored block or a restored picture based on the restored sample. Hereinafter, as described above, the decoding device can apply an in-loop filtering procedure such as deblocking filtering and / or SAO procedure to the restored picture in order to improve subjective / objective image quality as necessary.
[0295] FIG. 18 is a flowchart schematically showing an encoding method that can be executed by an encoding device according to an embodiment of this document.
[0296] The method disclosed in FIG. 18 can be executed by the encoding device 100 disclosed in FIG. 1. Specifically, steps S1800 to S1830 in FIG. 18 can be executed by the prediction unit 120 (specifically, the intra prediction unit 122) disclosed in FIG. 1, step S1840 in FIG. 18 can be executed by the subtraction unit 131 disclosed in FIG. 1, and steps S1850 and S1860 in FIG. 18 can be executed by the entropy encoding unit 140 disclosed in FIG. 1. Also, the method disclosed in FIG. 18 can include the embodiments described above in this document. Therefore, in FIG. 18, specific descriptions of contents overlapping with the above-described embodiments will be omitted or simplified.
[0297] Referring to FIG. 18, the encoding device can derive whether matrix-based intra prediction (MIP) is applied to the current block (S1800).
[0298] The encoding device can apply various prediction techniques to find the optimal prediction mode for the current block and can determine the optimal intra prediction mode based on rate-distortion optimization (RDO).
[0299] When it is determined that MIP is applied to the current block, the encoding device can downsample the reference samples adjacent to the current block to derive the reduced boundary samples (S1810).
[0300] The reduced boundary samples can be derived by being downsampled by averaging the reference samples.
[0301] When the width and height of the current block are 4, 4 reduced boundary samples are derived, and in other cases, 8 samples can be derived.
[0302] The averaging procedure for downsampling can be applied to each boundary, the left boundary or the upper boundary of the current block, which can be applied to the peripheral reference samples adjacent to the boundary of the current block.
[0303] By way of example, when the current block is a 4×4 block, the size of each boundary can be reduced to 2 samples through the averaging procedure, and when the current block is not a 4×4 block, the size of each boundary can be reduced to 4 samples through the averaging procedure.
[0304] When a reduced boundary sample is derived, the encoding device can derive a reduced prediction sample based on the operation of multiplying the MIP matrix selected based on the size of the current block by the reduced boundary sample (S1820).
[0305] The MIP matrix can be selected from any of three matrix sets classified according to the size of the current block, and each of the three matrix sets can include a plurality of MIP matrices.
[0306] That is, three matrix sets for MIP can be set, and each matrix set can be composed of a plurality of matrices and offset vectors. Such matrix sets can be classified and applied according to the size of the current block.
[0307] For example, for a 4×4 block, a matrix set including 18 or 16 matrices composed of 16 rows and 4 columns and 18 or 16 offset vectors can be applied. The index information can be information indicating any one of the plurality of matrices included in one matrix set.
[0308] For 4×8, 8×4, and 8×8 blocks, or 4×H or W×4 blocks, a matrix set including 10 or 8 matrices composed of 16 rows and 8 columns and 10 or 8 offset vectors can be applied.
[0309] Alternatively, for blocks other than the aforementioned blocks or blocks with a height and width of 8 or more, a matrix set including 6 matrices composed of 64 rows and 8 columns and 6 offset vectors can be applied.
[0310] After the multiplication operation of the MIP matrix and the reduced boundary samples, a reduced predicted sample can be derived based on the operation of adding an offset, that is, a predicted sample to which the MIP matrix is applied.
[0311] Thereafter, the encoding device can upsample the reduced predicted sample to generate an intra prediction sample for the current block (S1830).
[0312] The intra prediction sample can be upsampled by linear interpolation of the reduced predicted sample.
[0313] The interpolation procedure can be referred to as a linear interpolation or a bilinear interpolation procedure and can include two steps: 1) vertical interpolation, and 2) horizontal interpolation.
[0314] If W >= H, vertical linear interpolation can be applied first, followed by horizontal linear interpolation. If W < H, horizontal linear interpolation can be applied first, followed by vertical linear interpolation. In the case of a 4×4 block, the interpolation procedure can be omitted.
[0315] Also, the encoding device can derive a residual sample for the current block based on the predicted sample of the current block and the original sample of the current block (S1840).
[0316] Then, the encoding device can evolve the bit string of the syntax element for the MIP mode information in a truncated binary evolution manner and encode it in a bypass mode (S1850).
[0317] As described above, the maximum length of the bit string of the syntax element for the MIP mode information can be set to different values according to the size of the current block. Such a maximum length can be set to three values according to the size of the current block as in Table 2 or Table 3, and the maximum length may be the largest when the width and height of the current block are 4.
[0318] For example, in Table 2, when the width and height of the coding block are 4, the maximum length of the binary evolution is set to 34, and in Table 3, it can be set to 15.
[0319] Also, such MIP mode information can be encoded in a bypass mode instead of a context modeling method.
[0320] On the other hand, intra_mip_flag can be evolved in a Fixed Length Code.
[0321] The encoding device can generate residual information for the current block based on the residual samples, and output video information including the generated residual information, flag information indicating whether MIP is applied, and MIP mode information in the form of a bit stream (S1860).
[0322] Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the derived quantized conversion coefficient by performing conversion and quantization on the residual samples.
[0323] The flag information indicating whether MIP is applied can be encoded in the syntax information of the coding unit in a syntax like intra_mip_flag.
[0324] Also, the MIP mode information can be represented by intra_mip_mode_idx and can be encoded when intra_mip_flag is 1. Intra_mip_mode_idx can be index information indicating the MIP mode applied to the current block, and such index information can be used to derive a matrix when generating prediction samples. The index information can indicate any one of the plurality of MIP matrices included in one matrix set.
[0325] Also, by way of an example, flag information, such as intra_mip_transposed_flag, can further be signaled to indicate whether an input vector for matrix-based intra prediction is transposed via the syntax of a coding unit.
[0326] That is, the encoding device can encode video information including the MIP mode information and / or residual information of the current block described above and output it to a bitstream.
[0327] The bitstream can be transmitted to a decoding device via a network or a (digital) storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0328] The process of generating prediction samples for the current block described above can be executed by the intra prediction unit 122 of the encoding device 100 disclosed in FIG. 1, the process of deriving residual samples can be executed by the subtraction unit 131 of the encoding device 100 disclosed in FIG. 1, and the process of generating and encoding residual information can be executed by the residual processing unit 130 and the entropy encoding unit 140 of the encoding device 100 disclosed in FIG. 1.
[0329] In the foregoing embodiments, the method has been described based on a flowchart as a series of steps or blocks, but the embodiments of this document are not limited to the order of the steps, and a certain step may occur in a different order from the steps described above or simultaneously with different steps. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, and different steps may be included, or one or more steps of the flowchart may be deleted without affecting the scope of this document.
[0330] The method according to the foregoing document can be embodied in the form of software, and the encoding device and / or decoding device according to this document can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0331] In this document, when an embodiment is implemented in software, the above-described method can be implemented by modules (processes, functions, etc.) that perform the above-described functions. The modules can be stored in a memory and executed by a processor. The memory may be inside or outside the processor and may be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored in a digital storage medium.
[0332] In addition, the decoding device and the encoding device to which this document is applicable may include a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video intercom device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an over-the-top video (OTT) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a picture phone video device, a transportation means terminal (e.g., a vehicle terminal including an autonomous driving vehicle terminal, an airplane terminal, a ship terminal, etc.) and a medical video device, etc., and can be used to process video signals or data signals. For example, as the over-the-top video (OTT) device, a game console, a Blu-ray player, an Internet access TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc. can be included.
[0333] In addition, the processing method to which this document is applicable can be produced in the form of a program executed by a computer and can be stored in a recording medium readable by a computer. Multimedia data having a data structure related to this document can also be stored in a recording medium readable by a computer. The recording medium readable by the computer includes all types of storage devices and distributed storage devices in which data readable by a computer is stored. The recording medium readable by the computer can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Further, the recording medium readable by the computer includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a recording medium readable by a computer or can be transmitted via a wired or wireless communication network.
[0334] In addition, embodiments of this document can be embodied in a computer program product by program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a carrier readable by a computer.
[0335] FIG. 19 schematically shows an example of a video / image coding system applicable to embodiments of this document.
[0336] Referring to FIG. 19, the video / image coding system can include a first device (source device) and a second device (receiver device). The source device can transmit encoded video / image information or data to the receiver device in the form of a file or a stream via a digital storage medium or a network.
[0337] The source device can include a video source, an encoding device, and a transmission unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / video encoding device, and the decoding device may be referred to as a video / video decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can be composed of a separate device or an external component.
[0338] The video source can obtain video / video through processes such as video / video capture, synthesis, or generation. The video source can include a video / video capture device and / or a video / video generation device. The video / video capture device can include, for example, one or more cameras, a video / video archive including previously captured video / video, etc. The video / video generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / video. For example, virtual video / video can be generated via a computer or the like, and in this case, the video / video capture process can be replaced by the process of generating related data.
[0339] The encoding device can encode the input video / video. The encoding device can perform a series of procedures such as prediction, conversion, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / video information) can be output in the form of a bitstream.
[0340] The transmitting unit can transmit the encoded video / video information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or a network in the form of a file or a stream. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.
[0341] The decoding device can decode the video / video by performing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0342] The renderer can render the decoded video / video. The rendered video / video can be displayed via the display unit.
[0343] FIG. 20 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.
[0344] Referring to FIG. 20, the content streaming system applied to the embodiments of this document can generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0345] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream, and serves to transmit this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0346] The bitstream can be generated by an encoding method applied to the embodiments of this document or a method for generating a bitstream, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0347] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as a medium to inform the user of what services are available. If the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include another control server, and in this case, the control server serves to control commands / responses between each device within the content streaming system.
[0348] The streaming server can receive content from a media storage and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0349] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, a digital signage, etc.
[0350] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be distributedly processed.
[0351] The claims described in this document can be combined in various ways. For example, the technical features of the method claims in this document can be combined and implemented as an apparatus, and the technical features of the apparatus claims can be combined and implemented as a method. Also, the technical features of the method claims and the technical features of the apparatus claims in this document can be combined and implemented as an apparatus, and the technical features of the method claims and the technical features of the apparatus claims in this specification can be combined and implemented as a method. (Claims in the present description can be combined in a various way. For instance, technical features in method claims of the present description can be combined to be implemented or performed in an apparatus, and technical features in apparatus claims can be combined to be implemented or performed in a method. Further, technical features in method claim(s) and apparatus claim(s) can be combined to be implemented or performed in an apparatus. Further, technical features in method claim(s) and apparatus claim(s) can be combined to be implemented or performed in a method.)
Claims
1. In a video decoding method executed by a decoding device, obtaining flag information related to whether matrix-based intra prediction (MIP) is used for a current block from a bitstream; obtaining MIP mode information based on the fact that the value of the flag information is equal to 1, where the fact that the value of the flag information is equal to 1 is related to the fact that the MIP is used for the current block; deriving an MIP matrix based on the MIP mode information and the size of the current block; generating intra prediction samples for the current block based on the MIP matrix; generating restored samples for the current block based on the intra prediction samples, wherein the MIP mode information is index information related to the MIP matrix applied to the current block; the bitstring of the syntax element for the MIP mode information is evolved by a truncated binary (TB) method; the maximum length of the bitstring of the syntax element is set to three different values based on the size of the current block; A video decoding method in which, based on the width and height of the current block being 4, the maximum length has the largest value.
2. In a video encoding method executed by an encoding device, deriving whether matrix-based intra prediction (MIP) is applicable to a current block; deriving an MIP matrix based on the fact that the MIP is applicable to the current block; generating intra prediction samples for the current block based on the MIP matrix; generating flag information related to whether the MIP is applicable to the current block; encoding video information including the flag information, wherein based on the fact that the value of the flag information is equal to 1, the video information includes MIP mode information, and the fact that the value of the flag information is equal to 1 is related to the fact that the MIP is applicable to the current block. The MIP mode information is index information related to the MIP matrix applied to the current block, The bit string of the syntax element for the MIP mode information is evolved by the Truncated Binary (TB) method, The maximum length of the bit string of the syntax element is set to three different values based on the size of the current block, Based on the width and height of the current block being 4, the maximum length has the largest value, A video encoding method in which an MIP matrix is derived based on the MIP mode information and the size of the current block.
3. In a method for transmitting data for a video, A step of obtaining a bitstream for the video, the bitstream deriving whether matrix-based intra prediction (MIP) is applied to a current block, deriving an MIP matrix based on the MIP being applied to the current block, generating an intra prediction sample for the current block based on the MIP matrix, generating flag information related to whether the MIP is applied to the current block, and generating the video information including the flag information. A step of transmitting the data including the bitstream, Based on the value of the flag information being equal to 1, the video information includes MIP mode information, and the value of the flag information being equal to 1 is related to indicating that the MIP is applied to the current block, The MIP mode information is index information related to the MIP matrix applied to the current block, The bit string of the syntax element for the MIP mode information is evolved by the Truncated Binary (TB) method, The maximum length of the bit string of the syntax element is set to three different values based on the size of the current block, Based on the width and height of the current block being 4, the maximum length has the largest value, A transmission method in which an MIP matrix is derived based on the MIP mode information and the size of the current block.