Matrix-based Intra Prediction Apparatus and Method
Through the matrix-based intra prediction method, the effective compression and transmission of high-resolution, high-quality image/video data is solved, encoding efficiency and prediction performance are improved, and transmission and storage costs are reduced.
Patent Information
- Application Number
- CN202080041446.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-03
- Filing Date
- 2020-06-03
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-06-03
AI Technical Summary
The prior art is difficult to effectively compress and transmit high-resolution, high-quality image/video data, especially in applications of virtual reality and immersive media, resulting in increased transmission and storage costs.
The matrix-based intra prediction method is adopted, by receiving flag information indicating whether intra prediction is applied to the current block, an intra prediction sample of the current block is generated, and image encoding is performed based on the intra prediction mode information of the matrix.
Improve image/video compression efficiency, reduce implementation complexity, enhance prediction performance, and improve encoding efficiency.
Smart Images

Figure CN113950832B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to image encoding techniques, and more particularly, to an in - frame prediction apparatus based on a matrix and an image encoding technique for in - frame prediction based on a matrix. Background Art
[0002] Nowadays, the demand for high - resolution and high - quality images / videos such as 4K, 8K, or higher ultra - high - definition (UHD) images / videos has been continuously increasing in various fields. As image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases compared to traditional image data. Therefore, when transmitting image data using a medium such as a traditional wired / wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0003] In addition, nowadays, the interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and the broadcasting of images / videos having image characteristics different from those of real images such as game images is increasing.
[0004] Therefore, there is a need for an efficient image / video compression technique that can effectively compress, transmit, store, and reproduce the information of high - resolution and high - quality images / videos having various characteristics as described above. Summary of the Invention
[0005] Technical Problem
[0006] One technical aspect of the present disclosure is to provide a method and an apparatus for increasing image encoding efficiency.
[0007] Another technical aspect of the present disclosure is to provide an efficient in - frame prediction method and an efficient in - frame prediction apparatus.
[0008] Still another technical aspect of the present disclosure is to provide an image encoding method and an image encoding apparatus for in - frame prediction based on a matrix.
[0009] Still another technical aspect of the present disclosure is to provide an image encoding method and an image encoding apparatus for encoding mode information regarding in - frame prediction based on a matrix.
[0010] Technical Solution
[0011] According to an embodiment of the present disclosure, there is provided an image decoding method performed by a decoding device. The method may include: receiving flag information indicating whether matrix - based in - frame prediction (MIP) is used for a current block; receiving matrix - based in - frame prediction (MIP) mode information based on the flag information; generating in - frame prediction samples for the current block based on the MIP mode information; and generating reconstructed samples for the current block based on the in - frame prediction samples.
[0012] The MIP mode information may be index information indicating the MIP mode applied to the current block.
[0013] The generation of intra prediction samples may include: deriving reduced boundary samples by downsampling reference samples adjacent to the current block; deriving reduced prediction samples based on a multiplication operation of the reduced boundary samples and a MIP matrix; and generating intra prediction samples of the current block by upsampling the reduced prediction samples.
[0014] The reduced boundary samples may be downsampled by averaging the reference samples, and the intra prediction samples may be upsampled by linearly interpolating the reduced prediction samples.
[0015] The MIP matrix may be derived based on the size and index information of the current block.
[0016] The MIP matrix may be selected from any one of three matrix sets classified according to the size of the current block, and each of the three matrix sets may include multiple MIP matrices.
[0017] According to another embodiment of the present disclosure, there is provided an image encoding method performed by an encoding device. The method may include: deriving whether matrix-based intra prediction MIP is to be applied to the current block; when MIP is applied to the current block, deriving intra prediction samples of the current block based on MIP; deriving residual samples of the current block based on the intra prediction samples; and encoding information about the residual samples and information about MIP, wherein the information about MIP may include flag information indicating whether MIP is applied to the current block and matrix-based intra prediction (MIP) mode information.
[0018] According to still another embodiment of the present disclosure, there may be provided a digital storage medium storing image data including encoded image information and a bitstream generated according to an image encoding method performed by an encoding device.
[0019] According to still another embodiment of the present disclosure, there may be provided a digital storage medium storing image data including encoded image information and a bitstream to cause a decoding device to perform an image decoding method.
[0020] Technical effects
[0021] The present disclosure can have various effects. For example, according to an embodiment of the present disclosure, the overall image / video compression efficiency can be increased. In addition, according to an embodiment of the present disclosure, the implementation complexity can be reduced and the prediction performance can be enhanced through efficient intra prediction, thereby improving the overall coding efficiency. Additionally, according to an embodiment of the present disclosure, when performing matrix-based intra prediction, the index information indicating the matrix-based intra prediction can be efficiently encoded, thereby improving the coding efficiency.
[0022] The effects obtainable through the detailed examples in the specification are not limited to the above effects. For example, there can be various technical effects that can be understood or derived by those of ordinary skill in the relevant art from the specification. Therefore, the detailed effects of the specification are not limited to those explicitly described in the specification and can include various effects that can be understood or derived from the technical features of the specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 An example of a video / image coding system to which an embodiment of the present disclosure can be applied is schematically illustrated.
[0024] Figure 2 The configuration of a video / image coding device to which an embodiment of the present disclosure can be applied is schematically illustrated.
[0025] Figure 3 The configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied is schematically illustrated.
[0026] Figure 4 An example of an intra prediction-based image coding method to which an embodiment of the present disclosure can be applied is schematically illustrated.
[0027] Figure 5 An intra predictor in an encoding device is schematically illustrated.
[0028] Figure 6 An example of an intra prediction-based image decoding method to which an embodiment of the present disclosure can be applied is schematically illustrated.
[0029] Figure 7 An intra predictor in a decoding device is schematically illustrated.
[0030] Figure 8 An example of an MPM-based intra prediction method in an encoding device to which an embodiment of the present disclosure can be applied is illustrated.
[0031] Figure 9 An example of an MPM-based intra prediction method in a decoding device to which an embodiment of the present disclosure can be applied is illustrated.
[0032] Figure 10Illustrates an example of an intra prediction mode to which embodiments of the present disclosure can be applied.
[0033] Figure 11 Illustrates an MIP-based prediction sample generation process according to the example.
[0034] Figure 12 Illustrates the MIP process for a 4×4 block.
[0035] Figure 13 Illustrates the MIP process for an 8×8 block.
[0036] Figure 14 Illustrates the MIP process for an 8×4 block.
[0037] Figure 15 Illustrates the MIP process for a 16×16 block.
[0038] Figure 16 Illustrates the boundary averaging process in the MIP process.
[0039] Figure 17 Illustrates the linear interpolation in the MIP process.
[0040] Figure 18 Illustrates the MIP technology according to embodiments of the present disclosure.
[0041] Figure 19 Is a flowchart schematically illustrating a decoding method that can be executed by a decoding device according to embodiments of the present disclosure.
[0042] Figure 20 Is a flowchart schematically illustrating an encoding method that can be executed by an encoding device according to embodiments of the present disclosure.
[0043] Figure 21 Illustrates an example of a content stream transmission system to which embodiments of the present disclosure can be applied. Detailed implementation manners
[0044] This document can be modified in various ways and can have various embodiments, and specific embodiments will be illustrated and described in detail in the accompanying drawings. However, this is not intended to limit this document to specific embodiments. Terms commonly used in this specification are used to describe specific embodiments and are not used to limit the technical spirit of this document. Unless explicitly stated otherwise in the context, singular expressions include plural expressions. Terms such as "including" or "having" in this specification should be understood to indicate the existence of the features, quantities, steps, operations, elements, components, or combinations thereof described in the specification, and do not exclude the existence or addition of one or more other features, quantities, steps, operations, elements, components, or combinations thereof.
[0045] The elements in the accompanying drawings described in this document are independently illustrated for the convenience of describing features and functions related to different characteristics. This does not mean that each element is implemented as a separate piece of hardware or separate software. For example, at least two of the elements can be combined to form a single element, or a single element can be divided into multiple elements. Embodiments in which elements are combined and / or separated are also included within the scope of the rights of this document, unless it deviates from the essence of this document.
[0046] In this document, the term "A or B" can mean "only A", "only B", or "both A and B". In other words, in this document, the term "A or B" can be interpreted to indicate "A and / or B". For example, in this document, the term "A, B, or C" can mean "only A", "only B", "only C", or "any combination of A, B, and C".
[0047] The slash " / " or comma used in this document can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can represent "A, B, or C".
[0048] In this document, "at least one of A and B" can mean "only A", "only B", or "both A and B". In addition, in this document, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".
[0049] In addition, in this document, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0050] In addition, the parentheses used in this document may mean "for example". Specifically, in the case of expressing "prediction (intra prediction)", this may indicate that "intra prediction" is presented as an example of "prediction". In other words, "prediction" in this document is not limited to "intra prediction", and this may indicate that "intra prediction" is presented as an example of "prediction". Additionally, even in the case of expressing "prediction (i.e., intra prediction)", this may also indicate that "intra prediction" is presented as an example of "prediction".
[0051] In this document, the technical features separately described in one drawing may be implemented separately or may be implemented simultaneously.
[0052] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the Versatile Video Coding (VVC), Essential Video Coding (EVC) standard, AOMedia Video 1 (AV1) standard, Second Generation Audio Video Coding Standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.).
[0053] This document presents various embodiments of video / image coding, and unless otherwise specified, the embodiments can be executed in combination with each other.
[0054] In this document, video may refer to a series of images over time. A picture generally refers to a unit representing an image in a specific time region, while a slice / tile is a unit that is part of an encoded picture. A slice / tile may include one or more coding tree units (CTUs). A picture may consist of one or more slices / titles. A picture may consist of one or more tile groups. A tile group may include one or more tiles. A brick may represent a rectangular region of CTU rows within a tile in a picture. A tile may be divided into multiple bricks, each brick consisting of one or more CTU rows within the tile. A tile that is not divided into multiple bricks may also be referred to as a brick. Brick scan is a specific sequential ordering of the CTUs of the divided tiles within a brick, sequentially ordered in a CTU grid scan in a picture, and the tiles in the picture are sequentially ordered in a grid scan of the tiles of the picture. A tile is a rectangular region of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular region of CTUs, whose height is equal to the height of the picture and whose width is specified by a syntax element in the picture parameter set. A tile row is a rectangular region of CTUs, whose height is specified by a syntax element in the picture parameter set and whose width is equal to the width of the picture. Tile scan is a specific sequential ordering of the CTUs of the divided pictures within a tile, sequentially ordered in a CTU grid scan in a picture, while the tiles in the picture are sequentially ordered in a grid scan of the tiles of the picture. A slice includes an integral number of bricks of a picture that can be exclusively included in a single NAL unit. A slice may be composed of multiple complete tiles or only of a consecutive sequence of complete bricks of a single tile. Tile group and slice may be used interchangeably in this document. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0055] A pixel or pel may refer to the smallest unit that constitutes a picture (or image). Additionally, "sample" may be used as a term corresponding to a pixel. A sample generally may represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0056] A unit may represent the basic unit of image processing. A unit may include at least one of a specific region and information related to that region. A unit may include one luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the situation, terms such as unit and block, region, etc. may be used interchangeably. Generally, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0057] Hereinafter, preferred embodiments of this document will be described more specifically with reference to the accompanying drawings. Hereinafter, in the drawings, the same reference numerals are used for the same elements, and redundant descriptions of the same elements may be omitted.
[0058] Figure 1 An example of a video / image coding system to which embodiments of the present disclosure can be applied is schematically illustrated.
[0059] Referring to Figure 1 , the video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transfer the encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.
[0060] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0061] The video source may obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generating device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capturing process may be replaced by a process of generating relevant data.
[0062] The encoding device may encode the input video / image. The encoding device may perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.
[0063] The transmitter may send the encoded video / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating a media file in a predetermined file format and may include elements for sending via a broadcast / communication network. The receiver may receive / extract the bitstream and send the received / extracted bitstream to the decoding device.
[0064] The decoding device can decode video / images by performing a series of processes such as dequantization, inverse transformation, prediction, etc., corresponding to the operations of the encoding device.
[0065] The renderer can render the decoded video / images. The rendered video / images can be displayed through a display.
[0066] Figure 2 FIG. is a diagram schematically illustrating the configuration of a video / image encoding device to which the present disclosure can be applied. Hereinafter, the so-called video encoding device may include an image encoding device.
[0067] Refer to Figure 2 , the encoding device 200 may include an image splitter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the above-described image splitter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be constituted by one or more hardware components (e.g., an encoder chipset or a processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0068] The image divider 210 may divide an input image (or picture or frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit may be referred to as an encoding unit (CU). In this case, starting from a coding tree unit (CTU) or a largest coding unit (LCU), the encoding unit may be recursively divided according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, based on a quadtree structure, a binary tree structure, and / or a ternary tree structure, one encoding unit may be divided into multiple encoding units with a deeper depth. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or the ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final encoding unit that is not further divided. In this case, based on the encoding efficiency according to the image characteristics, the largest coding unit may be directly used as the final encoding unit. Alternatively, the encoding unit may be recursively divided into encoding units with a deeper depth as needed, whereby the encoding unit of the optimal size may be used as the final encoding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be separated or divided from the above-mentioned final encoding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.
[0069] Depending on the situation, the terms unit and terms such as block, region, etc. may be used in place of each other. In general, an M×N block may represent a set of samples or transformation coefficients composed of M columns and N rows. Samples generally may represent pixels or pixel values, and may represent only the pixels / pixel values of the luminance component, or only the pixels / pixel values of the chrominance component. Samples may be used as a term corresponding to the pixels or pels of a picture (or image).
[0070] In the encoding device 200, a prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 is subtracted from an input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown in the figure, the unit in the encoding device 200 for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) may be referred to as the subtractor 231. The predictor may perform prediction on a block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or the CU. As described later in the description of each prediction mode, the predictor may generate various information related to prediction such as prediction mode information and send the generated information to the entropy encoder 240. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0071] The intra-frame predictor 222 may predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the reference samples may be located near the current block or separated from the current block. In intra-frame prediction, the prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, the DC mode and the planar mode. Depending on the level of detail of the prediction direction, the directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-frame predictor 222 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0072] The inter-frame predictor 221 may derive a prediction block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information may be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same as or different from each other. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter-frame predictor 221 may configure a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter-frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter-frame predictor 221 may use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, a residual signal cannot be transmitted. In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of a neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling a motion vector difference.
[0073] The predictor 220 may generate a prediction signal based on various prediction methods described below. For example, the predictor may not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction at the same time. This may be referred to as combined inter-frame and intra-frame prediction (CIIP). In addition, the predictor may perform prediction of a block based on the intra-block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode may be used for content image / video coding such as games, for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but may be performed similarly to inter-frame prediction because a reference block is derived in the current picture. That is, IBC may use at least one of the inter-frame prediction techniques described in this document. The palette mode may be regarded as an example of intra-frame coding or intra-frame prediction. When the palette mode is applied, sample values within the picture may be signaled based on information about the palette table and the palette index.
[0074] The prediction signal generated by a predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT means a transform obtained from a graph when relationship information between pixels is represented by a graph. CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform process can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size rather than square blocks.
[0075] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, and entropy encoder 240 can encode the quantized signal (information on the quantized transform coefficients) and output a bitstream. The information on the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-shaped quantized transform coefficients into a one-dimensional vector based on the coefficient scan order, and generate information on the quantized transform coefficients based on the quantized transform coefficients in the form of a one-dimensional vector. Information on the quantized transform coefficients can be generated. Entropy encoder 240 can perform various coding methods such as, for example, exponential Golomb, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. Entropy encoder 240 can encode together or separately the information required for video / image reconstruction other than the quantized transform coefficients (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) can be sent or stored in the form of a bitstream in units of network abstraction layer (NAL). The video / image information can also include information on various parameter sets such as adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). Additionally, the video / image information can also include general constraint information. In this document, the information and / or syntax elements sent from the encoding device to / signaled to the decoding device can be included in the video / picture information. The video / image information can be encoded through the above encoding process and included in the bitstream. The bitstream can be sent through a network or stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for sending the signal output from entropy encoder 240 or a storage unit (not shown) for storing the signal can be included as an internal / external element of encoding device 200, and alternatively, the transmitter can be included in entropy encoder 240.
[0076] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients via the dequantizer 234 and the inverse transformer 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If there is no residual for the block to be processed, such as in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture and can be used for inter-frame prediction of the next picture through filtering described below.
[0077] In addition, a luminance mapping with chroma scaling (LMCS) can be applied during picture encoding and / or reconstruction.
[0078] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). The various filtering methods can include (for example) deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0079] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-frame predictor 221. When inter-frame prediction is applied by the encoding device, a prediction mismatch between the encoding device 200 and the decoding device can be avoided, and the encoding efficiency can be improved.
[0080] The DPB of the memory 270 can store the modified reconstructed picture to be used as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the blocks with derived (or encoded) motion information in the current picture and / or the motion information of the already reconstructed blocks in the picture. The stored motion information can be sent to the inter-frame predictor 221 and used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and can transfer the reconstructed samples to the intra-frame predictor 222.
[0081] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments of this document can be applied.
[0082] Referring to Figure 3 , the video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra predictor 331 and an inter predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be constituted by one or more hardware components (e.g., a decoder chipset or a processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0083] When receiving a bitstream including video / image information, the decoding device 300 may reconstruct an image corresponding to the processing of the video / image information that has been performed in the Figure 2 encoding device accordingly. For example, the decoding device 300 may derive units / blocks based on information related to block segmentation obtained from the bitstream. The decoding device 300 may perform decoding by using the processing units applied in the encoding device. Thus, the decoding processing unit may be, for example, an encoding unit, which may be divided along a quadtree structure, a binary tree structure, and / or a ternary tree structure with coding tree units or maximum coding units. One or more transform units may be derived with the encoding unit. And, the reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproducer.
[0084] The decoding device 300 may receive, in the form of a bitstream, from Figure 2The signal output by the encoding device and can decode the received signal through the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). Additionally, the video / image information can also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The signaled / received information and / or syntax elements described later in this document can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and outputs the syntax elements required for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the information of the symbols / bins decoded in the previous stage, the decoding target syntax element information, or the decoding information of the decoding target block to determine the context model, and perform arithmetic decoding on the bins by predicting the occurrence probability of the bins according to the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins of the context model for the next symbol / bin after determining the context model. The information related to prediction among the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values that have undergone entropy decoding in the entropy decoder 310, that is, the quantized transform coefficients and related parameter information, can be input to the residual processor 320. The residual processor 320 can derive the residual signal (residual block, residual sample, residual sample array). Additionally, the information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. Furthermore, the receiver (not shown) for receiving the signal output by the encoding device can also be configured as an internal / external component of the decoding device 300, or the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0085] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into the form of a two-dimensional block. In this case, the rearrangement can be performed based on the order of coefficient scanning that has been performed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.
[0086] The inverse transformer 322 obtains a residual signal (residual block, residual sample array) by performing an inverse transform on the transform coefficients.
[0087] The predictor can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and specifically can determine the intra / inter prediction mode.
[0088] The predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). Additionally, the predictor can perform prediction of a block based on the intra block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, e.g., screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction because a reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information about the palette table and the palette index.
[0089] The intra predictor 331 can predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the samples referred to can be located among the neighbors of the current block or can be located separately. In intra prediction, the prediction mode can include a plurality of non - directional modes and a plurality of directional modes. The intra predictor 331 can determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent block.
[0090] The inter-frame predictor 332 may derive a prediction block of a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, adjacent blocks may include spatially adjacent blocks present in the current picture and temporally adjacent blocks present in the reference picture. For example, the inter-frame predictor 332 may construct a motion information candidate list based on adjacent blocks and derive a motion vector and / or a reference picture index of the current block based on the received candidate selection information. The inter-frame prediction may be performed based on various prediction modes, and the information regarding the prediction may include information indicating a mode for the inter-frame prediction of the current block.
[0091] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (prediction block, prediction sample array) output from a predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331). If there is no residual for the block to be processed, such as when the skip mode is applied, the prediction block may be used as the reconstructed block.
[0092] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, may be output through filtering described below, or may be used for inter-frame prediction of the next picture.
[0093] In addition, a luminance mapping with chroma scaling (LMCS) may be applied during the picture decoding process.
[0094] The filter 350 may improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 may generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 360 (specifically, the DPB of the memory 360). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0095] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter - predictor 332. The memory 360 can store the motion information of the blocks of the derived (or decoded) motion information in the current picture and / or the motion information of the reconstructed blocks in the picture. The stored motion information can be sent to the inter - predictor 332 and used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and send the reconstructed samples to the intra - predictor 331.
[0096] In this document, the embodiments described in the filter 260, inter - predictor 221, and intra - predictor 222 of the encoding device 200 can be the same as, or respectively applied to correspond to, the filter 350, inter - predictor 332, and intra - predictor 331 of the decoding device 300.
[0097] As described above, when performing video encoding, prediction is performed to enhance the compression efficiency. A prediction block including prediction samples of a current block (i.e., a target encoding block) can be generated through prediction. In this case, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is equivalently derived in both the encoding device and the decoding device. The encoding device can enhance the image encoding efficiency by signaling information about the residual between the original block and the prediction block (residual information) instead of the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed picture including reconstructed samples by adding the residual block and the prediction block, and generate a reconstructed picture including reconstructed blocks.
[0098] The residual information can be generated through a transform and quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, derive quantized transform coefficients by performing a quantization process on the transform coefficients, and signal the relevant residual information to the decoding device (through the bitstream). In this case, the residual information can include information such as value information, position information, transform scheme, transform kernel, and quantization parameters of the quantized transform coefficients. The decoding device can perform a de - quantization / inverse - transform process based on the residual information and derive the residual samples (or residual block). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. In addition, the encoding device can derive a residual block by performing de - quantization / inverse - transform on the quantized transform coefficients for reference in the inter - prediction of subsequent pictures and generate a reconstructed picture.
[0099] In addition, if intra prediction is performed, the correlation between samples can be used, and the difference (i.e., the residual) between the original block and the predicted block can be obtained. The foregoing transformation and quantization can be applied to the residual. Therefore, spatial redundancy can be reduced. Hereinafter, an encoding method and a decoding method using intra prediction are specifically described.
[0100] Intra prediction refers to prediction for generating a predicted sample of a current block based on reference samples outside the current block in a picture (hereinafter referred to as the current picture) including the current block. In this case, the reference samples outside the current block may refer to samples adjacent to the current block. If intra prediction is applied to the current block, adjacent reference samples to be used for intra prediction of the current block can be derived.
[0101] For example, when the size (width × height) of the current block is nW × nH, the adjacent reference samples of the current block may include a total of 2 × nH samples adjacent to the left boundary and adjacent to the lower left of the current block, a total of 2 × nW samples adjacent to the upper boundary of the current block and adjacent to the upper right of the current block, and samples adjacent to the upper left of the current block. Alternatively, the adjacent reference samples of the current block may further include multiple columns of adjacent samples above and multiple rows of adjacent samples to the left. In addition, the adjacent reference samples of the current block may further include a total of nH samples adjacent to the right boundary of the current block having a size of nW × nH, a total of nW samples adjacent to the lower boundary of the current block, and one sample adjacent to the lower right of the current block.
[0102] In this case, some of the adjacent reference samples of the current block have not been decoded or may be unavailable. In this case, the decoding device can configure the adjacent reference samples to be used for prediction by replacing the unavailable samples with available samples. Alternatively, the adjacent reference samples to be used for prediction can be constructed by interpolation of available samples.
[0103] If adjacent reference samples are derived, (i) a predicted sample can be derived based on the average or interpolation of the adjacent reference samples of the current block, and (ii) a predicted sample can be derived based on the reference samples existing in a specific (prediction) direction among the adjacent reference samples of the current block for the predicted sample. When the intra prediction mode is a non - directional mode or a non - angular mode, (i) can be applied. When the intra prediction mode is a directional mode or an angular mode, (ii) can be applied.
[0104] In addition, a predicted sample can also be generated by interpolation between a first adjacent sample located in the prediction direction of the intra prediction mode of the current block and a second adjacent sample located in the opposite direction of the prediction direction based on the predicted sample of the current block among the adjacent reference samples. The foregoing case can be referred to as linear interpolation intra prediction (LIP). In addition, a chrominance prediction sample can also be generated based on luminance samples using a linear model. This case can be referred to as the LM mode.
[0105] In addition, the temporary prediction samples of the current block can be derived based on filtered neighboring reference samples, and the prediction samples of the current block can also be derived by weighted summation of at least one reference sample and the temporary prediction samples derived according to the intra prediction mode among the conventional neighboring reference samples (i.e., unfiltered neighboring reference samples). The foregoing case can be referred to as position-dependent intra prediction (PDPC).
[0106] In addition, the prediction samples can be derived by using the reference samples located in the prediction direction of the corresponding line by selecting the reference sample line with the highest prediction accuracy among multiple neighboring reference sample lines of the current block, and the intra prediction coding can be performed by a method for indicating (signaling) the reference sample line used at this time to the decoding device. The foregoing case can be called multi-reference line (MRL) intra prediction or MRL-based intra prediction.
[0107] In addition, the intra prediction can be performed based on the same intra prediction mode by separating the current block into vertical or horizontal sub-partitions, and the neighboring reference samples can be derived and used in units of sub-partitions. That is, in this case, the intra prediction mode of the current block is equally applied to the sub-partitions, and the neighboring reference samples can be derived and used in units of sub-partitions, thereby enhancing the intra prediction performance in some cases. This prediction method can be called intra sub-partition (ISP) intra prediction or ISP-based intra prediction.
[0108] The foregoing intra prediction methods can be called intra prediction types separately from the intra prediction mode. The intra prediction types can be called various terms, such as intra prediction techniques or additional intra prediction modes. For example, the intra prediction types (or additional intra prediction modes, etc.) can include at least one of the foregoing LIP, PDPC, MRL, and ISP. The general intra prediction method other than the specific intra prediction types such as LIP, PDPC, MRL, and ISP can be called the normal intra prediction type. If the specific intra prediction type is not applied, then the normal intra prediction type can generally be applied, and the prediction can be performed based on the foregoing intra prediction mode. In addition, if necessary, post-processing filtering for the derived prediction samples can also be performed.
[0109] In addition to the above intra prediction types, matrix-based intra prediction (hereinafter referred to as MIP) can be used as a method for intra prediction. MIP can be called affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP).
[0110] When MIP is applied to a current block, the prediction samples for the current block can be derived as follows: i) using neighboring reference samples that have undergone an averaging process; ii) performing a matrix-vector multiplication process; and (iii) further performing a horizontal / vertical interpolation process if necessary. The intra prediction mode for MIP can be configured differently from the intra prediction modes used for the aforementioned LIP, PDPC, MRL, or ISP intra prediction or for normal intra prediction.
[0111] The intra prediction mode of MIP can be referred to as an affine linear weighted intra prediction mode or a matrix-based intra prediction mode. For example, the matrix and offset used in the matrix-vector multiplication can be configured differently according to the intra prediction mode of MIP. Here, the matrix can be referred to as an (affine) weighting matrix, and the offset can be referred to as an (affine) offset vector or an (affine) bias vector. In the present disclosure, the intra prediction mode of MIP can be referred to as the MIP intra prediction mode, a linear weighted intra prediction mode, a matrix weighted intra prediction mode, or a matrix-based intra prediction mode. Specific MIP methods will be described later.
[0112] The following drawings have been prepared to explain specific examples of this document. Since the names of specific devices or specific words or names (e.g., grammatical names, etc.) described in the drawings are presented exemplarily, the technical features of this document are not limited to the specific names used in the following drawings.
[0113] Figure 4 An example of an intra prediction-based image coding method to which an embodiment of the present disclosure can be applied is schematically illustrated, and Figure 5 an intra predictor in an encoding device is schematically illustrated. Figure 5 The intra predictor in the illustrated encoding device can also be equally or correspondingly applied to Figure 2 the intra predictor 222 of the illustrated encoding device 200.
[0114] Referring to Figure 4 and Figure 5, S400 can be executed by the intra predictor 222 of the encoding device, and S410 can be executed by the residual processor 230 of the encoding device. Specifically, S410 can be executed by the subtractor 231 of the encoding device. In S420, the prediction information can be derived by the intra predictor 222 and encoded by the entropy encoder 240. In S420, the residual information can be derived by the residual processor 230 and encoded by the entropy encoder 240. The residual information indicates information about the residual samples. The residual information can include information about the quantized transform coefficients of the residual samples. As described above, the residual samples can be derived based on the transform coefficients passed through the transformer 232 of the encoding device, and the transform coefficients can be derived based on the quantized transform coefficients passed through the quantizer 233. The information about the quantized transform coefficients can be encoded by the entropy encoder 240 through the residual encoding process.
[0115] The encoding device performs intra prediction (S400) on the current block. The encoding device can derive the intra prediction mode / type of the current block, derive the neighboring reference samples of the current block, and generate the prediction samples in the current block based on the intra prediction mode / type and the neighboring reference samples. Here, the processes of determining the intra prediction mode / type, deriving the neighboring reference samples, and generating the prediction samples can also be executed simultaneously, and any one of the processes can also be executed earlier than the other processes.
[0116] For example, the intra predictor 222 of the encoding device can include an intra prediction mode / type determiner 222-1, a reference sample deriver 222-2, and a prediction sample deriver 222-3, where the intra prediction mode / type determiner 222-1 can determine the intra prediction mode / type of the current block, the reference sample deriver 222-2 can derive the neighboring reference samples of the current block, and the prediction sample deriver 222-3 can derive the prediction samples of the current block. In addition, although not shown, if the prediction sample filtering process is executed, the intra predictor 222 can also include a prediction sample filter (not shown). The encoding device can determine the mode / type applied to the current block among multiple intra prediction modes / types. The encoding device can compare the RD costs of the intra prediction modes / types and determine the best intra prediction mode / type of the current block.
[0117] As described above, the encoding device can also execute the prediction sample filtering process. The prediction sample filtering can be referred to as post-filtering. Some or all of the prediction samples can be filtered through the prediction sample filtering process. In some cases, the prediction sample filtering process can be omitted.
[0118] The encoding device generates the residual samples of the current block based on the (filtered) prediction samples (S410). The encoding device can compare the prediction samples based on the phase in the original samples of the current block and derive the residual samples.
[0119] The encoding device may encode image information including information on intra prediction (prediction information) and residual information on residual samples (S420). The prediction information may include intra prediction mode information and intra prediction type information. The residual information may include residual coding syntax. The encoding device may derive quantized transform coefficients by transforming / quantizing the residual samples. The residual information may include information on the quantized transform coefficients.
[0120] The encoding device may output the encoded image information in the form of a bitstream. The output bitstream may be delivered to the decoding device via a storage medium or a network.
[0121] As described above, the encoding device may generate a reconstructed picture (including reconstructed samples and reconstructed blocks). To this end, the encoding device may derive (modified) residual samples by dequantizing / inverse-transforming the quantized transform coefficients again. As described above, the reason for transforming / quantizing the residual samples and then dequantizing / inverse-transforming them again is to derive the same residual samples as those derived by the decoding device, as described above. The encoding device may generate a reconstructed block including the reconstructed samples of the current block based on the prediction samples and the (modified) residual samples. The reconstructed picture of the current picture may be generated based on the reconstructed block. As described above, an in-loop filtering process or the like may also be applied to the reconstructed picture.
[0122] Figure 6 An example of an intra prediction-based image decoding method to which embodiments of the present disclosure may be applied is schematically illustrated, and Figure 7 an intra predictor in the decoding device is schematically illustrated. Figure 7 The intra predictor in the illustrated decoding device may also be equally or correspondingly applied to Figure 3 the intra predictor 331 of the illustrated decoding device 300.
[0123] Referring to Figure 6 and Figure 7 , the decoding device may perform operations corresponding to the foregoing operations performed by the encoding device. S600 to S620 may be performed by the intra predictor 331 of the decoding device, and the entropy decoder 310 of the decoding device may obtain the prediction information in S600 and the residual information in S630 from the bitstream. The residual processor 320 of the decoding device may derive the residual samples of the current block based on the residual information. Specifically, the dequantizer 321 of the residual processor 320 may derive transform coefficients by performing dequantization based on the quantized transform coefficients derived based on the residual information, and the inverse transformer 322 of the residual processor may derive the residual samples of the current block by inverse-transforming the transform coefficients. S640 may be performed by the adder 340 or the reconstructor of the decoding device.
[0124] The decoding device can derive the intra prediction mode / type of the current block based on the received prediction information (intra prediction mode / type information) (S600). The decoding device can derive the neighboring reference samples of the current block (S610). The decoding device generates the prediction samples in the current block based on the intra prediction mode / type and the neighboring reference samples (S620). In this case, the decoding device can perform a prediction sample filtering process. The prediction sample filtering can be referred to as post-filtering. Some or all of the prediction samples can be filtered by the prediction sample filtering process. In some cases, the prediction sample filtering process can be omitted.
[0125] The decoding device generates the residual samples of the current block based on the received residual information (S630). The decoding device can generate the reconstructed samples of the current block based on the prediction samples and the residual samples, and derive the reconstructed block including the reconstructed samples (S640). The reconstructed picture of the current picture can be generated based on the reconstructed block. As described above, in-loop filtering processes and the like can also be applied to the reconstructed picture.
[0126] Here, the intra predictor 331 of the decoding device can include an intra prediction mode / type determiner 331-1, a reference sample deriver 331-2, and a prediction sample deriver 331-3. The intra prediction mode / type determiner 331-1 can determine the intra prediction mode / type of the current block based on the intra prediction mode / type information obtained by the entropy decoder 310. The reference sample deriver 331-2 can derive the neighboring reference samples of the current block, and the prediction sample deriver 331-3 can derive the prediction samples of the current block. In addition, although not shown, if the foregoing prediction sample filtering process is performed, the intra predictor 331 can also include a prediction sample filter (not shown).
[0127] The intra prediction mode information can include (for example) flag information (such as intra_luma_mpm_flag) indicating whether the most probable mode (MPM) is applied to the current block or whether the remaining modes are applied to it. At this time, if the MPM is applied to the current block, the prediction mode information can also include index information (such as intra_luma_mpm_idx) indicating one of the intra prediction mode candidates (MPM candidates). The intra prediction mode candidates (MPM candidates) can be constituted by an MPM candidate list or an MPM list. In addition, if the MPM is not applied to the current block, the intra prediction mode information can also include remaining mode information (such as intra_luma_mpm_remainder) indicating one of the remaining intra prediction modes other than the intra prediction mode candidates (MPM candidates). The decoding device can determine the intra prediction mode of the current block based on the intra prediction mode information.
[0128] In addition, the intra prediction type information can be implemented in various forms. As an example, the intra prediction type information may include intra prediction type index information indicating one of the intra prediction types. As another example, the intra prediction type information may include at least one of the following: reference sample line information (e.g., intra_luma_ref_idx), which indicates whether MRL is applied to the current block and which reference sample line is used if MRL is applied; ISP flag information (e.g., intra_subpartitions_mode_flag), which indicates whether ISP is applied to the current block; ISP type information (e.g., intra_subpartitions_split_flag), which indicates the split type of the subpartitions if ISP is applied; flag information indicating whether PDCP is applied; or flag information indicating whether LIP is applied. In addition, the intra prediction type information may include a MIP flag indicating whether MIP is applied to the current block.
[0129] The above intra prediction mode information and / or intra prediction type information can be encoded / decoded by the encoding methods described in this document. For example, the foregoing intra prediction mode information and / or intra prediction type information can be encoded / decoded by entropy coding based on truncated (Rice) binary codes (e.g., CABAC, CAVLC).
[0130] In addition, in the case of applying intra prediction, the intra prediction mode of adjacent blocks can be used to determine the intra prediction mode applied to the current block. For example, the decoding device can select one mpm candidate from the mpm candidates in the most probable mode (mpm) list derived based on the intra prediction modes of adjacent blocks (e.g., the left adjacent block and / or the upper adjacent block) of the current block and additional candidate modes according to the received mpm index, or can select one of the remaining intra prediction modes not included in the mpm candidates (and the planar mode) based on the remaining intra prediction mode information. The mpm list can be constructed to include or not include the planar mode as a candidate. For example, if the mpm list includes the planar mode as a candidate, the mpm list can have 6 candidates, while if the mpm list does not include the planar mode as a candidate, the mpm list can have 5 candidates. If the mpm list does not include the planar mode as a candidate, a non-planar flag (e.g., intra_luma_not_planar_flag) indicating whether the intra prediction mode of the current block is the planar mode can be signaled. For example, the mpm flag can be signaled first, and when the value of the mpm flag is equal to 1, the mpm index and the non-planar flag can be signaled. In addition, when the value of the non-planar flag is equal to 1, the mpm index can be signaled. Here, constructing the mpm list to not include the planar mode as a candidate is to first identify whether the intra prediction mode is the planar mode by signaling the flag (non-planar flag) first, because the planar mode has always been regarded as an mpm, while the non-planar mode is not an mpm.
[0131] For example, based on an MPM flag (e.g., intra_luma_mpm_flag), it can be indicated whether the intra prediction mode applied to the current block is among the MPM candidates (and planar mode) or among the remaining modes. An MPM flag value of 1 can indicate that the intra prediction mode of the current block is among the MPM candidates (and planar mode), and an MPM flag value of 0 can indicate that the intra prediction mode of the current block is not among the MPM candidates (and planar mode). A non-planar flag (e.g., intra_luma_not_planar_flag) value of 0 can indicate that the intra prediction mode of the current block is the planar mode, and a non-planar flag value of 1 can indicate that the intra prediction mode of the current block is not the planar mode. The MPM index can be signaled in the form of an mpm_idx or intra_luma_mpm_idx syntax element, and the remaining intra prediction mode information can be signaled in the form of a rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, the remaining intra prediction mode information can index the remaining intra prediction modes not included in the MPM candidates (and planar mode) among all the intra prediction modes in the order of their number of prediction modes, and can indicate one of them. The intra prediction mode can be the intra prediction mode of the luminance component (samples). Hereinafter, the intra prediction mode information can include at least one of an MPM flag (e.g., intra_luma_mpm_flag), a non-planar flag (e.g., intra_luma_not_plane_flag), an MPM index (e.g., mpm_idx or intra_luma_mpm_idx), and the remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list can be referred to by various terms, such as the MPM candidate list, the candidate mode list (candModeList), and the candidate intra prediction mode list.
[0132] Generally, when dividing a block of an image, the current block to be encoded and adjacent blocks have similar image attributes. Therefore, the current block and adjacent blocks are more likely to have the same or similar intra prediction modes. Thus, the encoder can use the intra prediction mode of an adjacent block to encode the intra prediction mode of the current block. For example, the encoder / decoder can form a most probable mode (MPM) list of the current block. The MPM list can also be referred to as the MPM candidate list. Here, MPM can mean a mode used to improve the encoding efficiency by considering the similarity between the current block and adjacent blocks when encoding the intra prediction mode.
[0133] Figure 8An example of an MPM-based intra prediction method in an encoding device to which an exemplary embodiment of the present disclosure can be applied is illustrated.
[0134] Referring to Figure 8 , the encoding device constructs an MPM list for the current block (S800). The MPM list may include candidate intra prediction modes (MPM candidates) that are more likely to be applied to the current block. The MPM list may also include the intra prediction modes of adjacent blocks, and may also include specific intra prediction modes according to a predetermined method. A specific method for constructing the MPM list will be described later.
[0135] The encoding device determines the intra prediction mode of the current block (S810). The encoding device may perform prediction based on various intra prediction modes, and determine the optimal intra prediction mode based on rate distortion optimization (RDO) based on the above prediction. In this case, the encoding device may also use only the MPM candidates and the planar mode configured in the MPM list to determine the optimal intra prediction mode, or may also use the remaining intra prediction modes and the MPM candidates and the planar mode configured in the MPM list to determine the optimal intra prediction mode.
[0136] Specifically, for example, if the intra prediction type of the current block is a specific type different from the normal intra prediction type (e.g., LIP, MRL, or ISP), the encoding device may consider only the MPM candidates and the planar mode as candidates for the intra prediction mode of the current block to determine the optimal intra prediction mode. That is, in this case, the intra prediction mode of the current block may be determined only among the MPM candidates and the planar mode, and in this case, the mpm flag may not be encoded / signaled. In this case, even without a separate signaling of the mpm flag, the decoding device may estimate the mpm flag as 1.
[0137] Generally, if the intra prediction mode of the current block is not the planar mode and is one of the MPM candidates in the MPM list, the encoding device generates an mpm index (mpm idx) indicating one of the MPM candidates. If the intra prediction mode of the current block does not even exist in the MPM list, the encoding device generates remaining intra prediction mode information indicating a mode among the remaining intra prediction modes not included in the MPM list (and the planar mode) as the intra prediction mode of the current block.
[0138] The encoding device may encode the intra prediction mode information to output the intra prediction mode information in the form of a bitstream (S820). The intra prediction mode information may include the aforementioned mpm flag, non - planar flag, mpm index, and / or the remaining intra prediction mode information. Generally, the mpm index and the remaining intra prediction mode information have an alternative relationship and are not signaled simultaneously when indicating the intra prediction mode of a block. That is, the value 1 of the mpm flag and the non - planar flag or the mpm index are signaled together, or the value 0 of the mpm flag and the remaining intra prediction mode information are signaled together. However, as described above, if a specific intra prediction type is applied to the current block, the mpm flag is not signaled and only the non - planar flag and / or the mpm index may be signaled. That is, in this case, the intra prediction mode information may also include only the non - planar flag and / or the mpm index.
[0139] Figure 9 An example of an MPM - based intra prediction method in an encoding device to which exemplary embodiments of the present disclosure can be applied is illustrated. Figure 9 The decoding device shown may determine the intra prediction mode corresponding to the intra prediction mode information determined and signaled by the Figure 8 encoding device shown.
[0140] Referring to Figure 9 , the decoding device obtains the intra prediction mode information from the bitstream (S900). As described above, the intra prediction mode information may include at least one of the mpm flag, non - planar flag, mpm index, and the remaining intra prediction modes.
[0141] The decoding device constructs an MPM list (S910). The MPM list consists of the same MPM list constructed in the encoding device. That is, the MPM list may also include the intra prediction modes of adjacent blocks and may also include specific intra prediction modes according to a predetermined method. A specific method for constructing the MPM list will be described later.
[0142] Although S910 is shown to be executed later than S900, it is illustrative, and S910 may also be executed earlier than S900, and S900 and S910 may also be executed simultaneously.
[0143] The decoding device determines the intra prediction mode of the current block based on the MPM list and the intra prediction mode information (S920).
[0144] As an example, if the value of the mpm flag is 1, the decoding device may derive the planar mode as the intra prediction mode of the current block, or may derive the candidate indicated by the mpm index among the MPM candidates in the MPM list (based on the non-planar flag) as the intra prediction mode of the current block. Here, the MPM candidates may also indicate candidates only included in the MPM list, or may also include the planar mode applicable to the case where the value of the mpm flag is 1 and the candidates included in the MPM list.
[0145] As another example, if the value of the mpm flag is 0, the decoding device may derive the intra prediction mode indicated by the remaining intra prediction mode information among the remaining intra prediction modes as the intra prediction mode of the current block, where the remaining intra prediction modes are not included in the MPM list and the planar mode.
[0146] As yet another example, if the intra prediction type of the current block is a specific type (e.g., LIP, MRL, or ISP), the decoding device may derive the candidate indicated by the mpm index in the MPM list or the planar mode as the intra prediction mode of the current block even without confirmation of the mpm flag.
[0147] When constructing the MPM list, in an embodiment, the encoding device / decoding device may derive the left mode, which is the candidate intra prediction mode of the left adjacent block of the current block, and may derive the up mode, which is the candidate intra prediction mode of the upper adjacent block of the current block. Here, the left adjacent block may represent the lowermost adjacent block among the left adjacent blocks located adjacent to the left of the current block, and the upper adjacent block may represent the rightmost adjacent block among the upper adjacent blocks located adjacent to the upper part of the current block. For example, if the size of the current block is W×H and the x component and y component of the upper left sample position of the current block are xN and yN, respectively, the left adjacent block may be a block including samples at the coordinates (xN - 1, yN + H - 1), and the upper adjacent block may be a block including samples at the coordinates (xN + W - 1, yN - 1).
[0148] For example, when the left adjacent block is available and intra prediction is applied to the left adjacent block, the encoding device / decoding device may derive the intra prediction mode of the left adjacent block as the left candidate intra prediction mode (i.e., the left mode). When the upper adjacent block is available, intra prediction is applied to the upper adjacent block, and the upper adjacent block is included in the current CTU, the decoding device may derive the intra prediction mode of the upper adjacent block as the upper candidate intra prediction mode (i.e., the up mode). Alternatively, when the left adjacent block is not available or intra prediction is not applied to the left adjacent block, the encoding device / decoding device may derive the planar mode as the left mode. When the upper adjacent block is not available or intra prediction is not applied to the upper adjacent block or the upper adjacent block is not included in the current CTU, the decoding device may derive the planar mode as the up mode.
[0149] An encoding device / decoding device may construct an MPM list by deriving candidate intra prediction modes of a current block based on a left mode derived from a left adjacent block and an upper mode derived from an upper adjacent block. In this case, the MPM list may include the left mode and the upper mode, and may also include specific intra prediction modes according to a predetermined method.
[0150] In addition, when constructing the MPM list, a normal intra prediction mode (normal intra prediction type) may be applied, or a specific intra prediction type (e.g., MRL or ISP) may be applied. This document proposes a method for constructing an MPM list when not only applying a normal intra prediction method (normal intra prediction type) but also applying a specific intra prediction type (e.g., MRL or ISP), and this will be described later.
[0151] In addition, the intra prediction mode may include a non - directional (or non - angular) intra prediction mode and a directional (or angular) intra prediction mode. For example, in the HEVC standard, an intra prediction mode including 2 non - directional prediction modes and 33 directional prediction modes is used. The non - directional prediction modes may include a planar intra prediction mode (i.e., No. 0) and a DC intra prediction mode (i.e., No. 1). The directional prediction modes may include intra prediction modes from No. 2 to No. 34. The planar mode intra prediction mode may be referred to as the planar mode, and the DC intra prediction mode may be referred to as the DC mode.
[0152] Alternatively, in order to capture a given edge direction proposed in natural videos, the directional intra prediction mode may be extended from the existing 33 modes to 65 modes as in Figure 10 . In this case, the intra prediction mode may include 2 non - directional intra prediction modes and 65 directional intra prediction modes. The non - directional intra prediction modes may include a planar intra prediction mode (i.e., No. 0) and a DC intra prediction mode (i.e., No. 1). The directional intra prediction modes may include intra prediction modes from No. 2 to No. 66. The extended directional intra prediction mode may be applied to blocks of all sizes and may be applied to both the luminance component and the chrominance component. However, this is an example, and the embodiments of this document may be applied to cases where the number of intra prediction modes is different. 67 intra prediction modes may also be used according to the situation. The 67th intra prediction mode may indicate a linear model (LM) mode.
[0153] Figure 10 Examples of intra prediction modes to which the embodiments of this document may be applied are illustrated.
[0154] Referring to Figure 10 , the modes may be divided into intra prediction modes with horizontal directivity and intra prediction modes with vertical directivity based on the 34th intra prediction mode having an upper - left diagonal prediction direction. InFigure 10 In this case, H and V respectively denote the horizontal directionality and the vertical directionality. Each of the numbers from -32 to 32 indicates a displacement of 1 / 32 unit at the sample grid position. The intra prediction modes within the 2nd to 33rd frames have horizontal directionality, and the intra prediction modes within the 34th to 66th frames have vertical directionality. The 18th intra prediction mode and the 50th intra prediction mode respectively indicate the horizontal intra prediction mode and the vertical intra prediction mode. The 2nd intra prediction mode can be referred to as the lower left diagonal intra prediction mode, the 34th intra prediction mode can be referred to as the upper left diagonal intra prediction mode, and the 66th intra prediction mode can be referred to as the upper right diagonal intra prediction mode.
[0155] The intra prediction mode for MIP can indicate the matrix and offset for intra prediction, rather than the existing directional modes. That is, the matrix and offset for intra prediction can be derived from the intra mode of MIP. In this case, when deriving the intra mode for general intra prediction or for generating the MPM list described above, the intra prediction mode of the block predicted by MIP can be set to a preset mode, for example, the planar mode or the DC mode. According to another example, the intra mode of MIP can be mapped to the planar mode, the DC mode, or the directional intra mode based on the block size.
[0156] Hereinafter, matrix-based intra prediction (MIP) as an intra prediction method is described.
[0157] As described above, matrix-based intra prediction (hereinafter, MIP) can be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). To predict the samples of a rectangular block having a width (W) and a height (H), MIP uses one H line among the reconstructed samples adjacent to the left boundary of the block and one W line among the reconstructed samples adjacent to the upper boundary of the block as input values. When no reconstructed samples are available, reference samples can be generated by an interpolation method applied to general intra prediction.
[0158] Figure 11 An example of the MIP-based prediction sample generation process according to an example is illustrated. Refer to the following Figure 11 for the description of the MIP process.
[0159] 1. Averaging process
[0160] Among the boundary samples, four samples for the case of W = H = 4 and eight samples for any other case are extracted through the averaging process.
[0161] 2. Matrix-vector multiplication process
[0162] Perform matrix-vector multiplication using the average sample as input, and then add an offset. Through this operation, a reduced predicted sample set of the downsampled samples in the original block can be derived.
[0163] 3. (Linear) interpolation process
[0164] Generate predicted samples at the remaining positions from the predicted samples of the downsampled sample set by linear interpolation, where the linear interpolation is single-step linear interpolation in each direction.
[0165] The matrices and offset vectors required to generate the predicted block or predicted samples can be selected from three sets S0, S1, and S2 of matrices.
[0166] Set S0 may include 16 matrices A0 i , i ∈ {0,..., 15}, and 16 offset vectors b0 i , i ∈ {0,..., 15}, and each matrix may include 16 rows and 4 columns. The matrices and offset vectors of set S0 can be used for 4×4 blocks. In another example, set S0 may include 18 matrices.
[0167] Set S1 may include eight matrices A1 i , i ∈ {0,..., 7}, and eight offset vectors b1 i , i ∈ {0,..., 7}, and each matrix may include 16 rows and 8 columns. In another example, set S1 may include six matrices. The matrices and offset vectors of set S1 can be used for 4×8, 8×4, and 8×8 blocks. Alternatively, the matrices and offset vectors of set S1 can be used for 4×H or W×4 blocks.
[0168] Finally, set S2 may include six matrices A2 i , i ∈ {0,..., 5}, and six offset vectors b2 i , i ∈ {0,..., 5}, and each matrix may include 64 rows and 8 columns. The matrices and offset vectors of set S2 or some of them can be used for any block with a different size not applicable to sets S0 and S1. For example, the matrices and offset vectors of set S2 can be used for operations on blocks with a height and width of 8 or greater.
[0169] The total number of multiplications required for the calculation of matrix-vector multiplication is always less than or equal to 4×W×H. That is, in the MIP mode, each sample requires at most four multiplications.
[0170] Hereinafter, the overall MIP process is briefly described. The remaining blocks not described below can be processed in any of the four cases described.
[0171] Figures 12 to 15 Illustrated is the MIP process according to the size of the block, whereFigure 12 Illustrates the MIP process for a 4×4 block, Figure 13 Illustrates the MIP process for an 8×8 block, Figure 14 Illustrates the MIP process for an 8×4 block, and Figure 15 Illustrates the MIP process for a 16×16 block.
[0172] As Figure 12 shown, given a 4×4 block, MIP averages two samples for each axis of the boundary. As a result, four input samples are inputs to a matrix-vector multiplication, and the matrix is taken from set S0. An offset is added, thereby generating 16 final prediction samples. For a 4×4 block, no linear interpolation is required to generate prediction samples. Thus, a total of (4×16) / (4×4) = 4 multiplications can be performed per sample.
[0173] As Figure 13 shown, given an 8×8 block, MIP averages four samples for each axis of the boundary. As a result, eight input samples are inputs to a matrix-vector multiplication, and the matrix is taken from set S1. Sixteen samples are generated at odd positions by matrix-vector multiplication.
[0174] For an 8×8 block, each sample performs a total of (8×16) / (8×8) = 2 multiplications to generate prediction samples. After adding the offset, the reduced upper boundary samples are used to vertically interpolate the samples, and the original left boundary samples are used to horizontally interpolate the samples. In this case, since no multiplication operations are required during the interpolation process, a total of two multiplications per sample are required for MIP.
[0175] As Figure 14 shown, given an 8×4 block, MIP averages four samples for the horizontal axis of the boundary and uses four sample values on the left boundary for the vertical axis. As a result, eight input samples are inputs to a matrix-vector multiplication, and the matrix is taken from set S1. Sixteen samples are generated at odd horizontal positions and corresponding vertical positions by matrix-vector multiplication.
[0176] For an 8×4 block, each sample performs a total of (8×16) / (8×4) = 4 multiplications to generate prediction samples. After adding the offset, the original left boundary samples are used to horizontally interpolate the samples. In this case, since no multiplication operations are required during the interpolation process, a total of four multiplications per sample are required for MIP.
[0177] As Figure 15As shown, given a 16×16 block, MIP averages four samples along each axis. As a result, eight input samples are the input to the matrix-vector multiplication, and the matrix is taken from set S2. Through the matrix-vector multiplication, 64 samples are generated at odd positions. For a 16×16 block, each sample performs a total of (8×64) / (16×16) = 2 multiplications to generate the predicted samples. After adding the offset, the samples are vertically interpolated using eight reduced upper boundary samples, and horizontally interpolated using the original left boundary samples. In this case, since no multiplication operations are required during the interpolation process, a total of two multiplications per sample are required for MIP.
[0178] For larger blocks, the MIP process is basically the same as the above process, and it is easy to identify that the number of multiplications per sample is less than 4.
[0179] For a W×8 block with width greater than 8 (W>8), since samples are generated at odd horizontal positions and each vertical position, only horizontal interpolation is required. In this case, each sample performs (8×64) / (W×8) = 64 / W multiplications for the prediction operation of the reduced samples. In the case of W = 16, no additional multiplications are required for linear interpolation, and in the case of W>16, the number of additional multiplications per sample required for linear interpolation is less than 2. That is, the total number of multiplications per sample is less than or equal to 4.
[0180] For a W×4 block with width greater than 4 (W>4), the matrix obtained by omitting all rows corresponding to odd entries along the horizontal axis of the downsampled block is defined as A k . Therefore, the output size is 32, and only horizontal interpolation is performed. For the prediction operation of the reduced samples, each sample performs (8×32) / (W×4) = 64 / W multiplications. In the case of W = 16, no additional multiplications are required, and in the case of W>16, the number of additional multiplications per sample required for linear interpolation is less than 2. That is, the total number of multiplications per sample is less than or equal to 4.
[0181] When the matrix is transposed, the processing can be performed accordingly.
[0182] Figure 16 Illustrates the boundary averaging process in the MIP process. Refer to Figure 16 for a detailed description of the averaging process.
[0183] According to the averaging process, averaging is applied to each boundary (left boundary or upper boundary). The boundary indicates the adjacent reference samples adjacent to the boundary of the current block, as Figure 16 shown. For example, the left boundary (bdry left ) indicates the left adjacent reference samples adjacent to the left boundary of the current block, and the upper boundary (bdrytop )Indicates the upper adjacent reference sample adjacent to the upper part.
[0184] When the current block is a 4×4 block, the size of each boundary can be reduced to two samples through an averaging process. When the current block is not a 4×4 block, the size of each boundary can be reduced to four samples through an averaging process.
[0185] The first step of the averaging process is to reduce the input boundaries (bdry left and bdry top ) to smaller boundaries ( and ). and Include two samples for 4×4 blocks and four samples for any other case.
[0186] For 4×4 blocks, where 0 ≤ i < 2, Can be represented by the following formula, and can be defined similarly
[0187] [Equation 1]
[0188]
[0189] When the width of the block is W = 4×2 k where 0 ≤ i < 4, Can be represented by the following formula, and can be defined similarly
[0190] [Equation 2]
[0191]
[0192] Since the two reduced boundaries and are cascaded into the reduced boundary vector bdry red , so bdry red has a size of 4 in 4×4 blocks and a size of 8 in any other block.
[0193] When "mode" indicates the MIP mode, the reduced boundary vector bdry red and the range of the MIP mode can be defined based on the block size and the intra_mip_transposed_flag value according to the following formula.
[0194] [Equation 3]
[0195]
[0196] In the above formula, intra_mip_transposed_flag can be referred to as MIP transpose, and this flag information can indicate whether to transpose the downsampled prediction samples. The semantics of this syntax element can be expressed as "intra_mip_transposed_flag[x0][y0] specifies whether the input vector of the matrix-based intra prediction mode for the luma samples is transposed.
[0197] Finally, for the interpolation of the downsampled prediction samples, a second version for averaging the boundaries for large blocks is required. That is, if the smaller value of the width and height is greater than 8 (min(W, H) > 8) and the width is equal to or greater than the height (W ≥ H), then W = 8 * 2 l , and can be defined by the following formula, where 0 ≤ i < 8. If the smaller value of the width and height is greater than 8 (min(W, H) > 8) and the height is greater than the width (H > W), it can be defined similarly
[0198] [Equation 4]
[0199]
[0200] Next, the process for generating the downsampled prediction samples through matrix-vector multiplication is described.
[0201] One of the downsampled input vectors bdry red is used to generate the downsampled prediction sample pred red . The prediction sample is a signal for the downsampled block with width W red and height H red . Here, W red and H red are defined as follows.
[0202] [Equation 5]
[0203]
[0204] The downsampled prediction sample pred red can be obtained by performing matrix-vector multiplication and then adding an offset, and can be derived by the following formula.
[0205] [Equation 6]
[0206] pred red = A · bdry red + b
[0207] Here, A is a matrix including W red × H reda matrix of W rows and four columns (where W and H are 4 (W = H = 4)) or eight columns (in any other case), and b is a W red ×H red vector.
[0208] Matrix A and vector b can be selected from S0, S1, and S2 as follows, and the index idx = idx(W, H) can be defined by Equation 7 or Equation 8.
[0209] [Equation 7]
[0210]
[0211] [Equation 8]
[0212]
[0213] When idx is 1 or less (idx ≤ 1) or when idx is 2 and the smaller value of W and H is greater than 4 (min(W, H) > 4), A is set to and b is set to When idx is 2, the smaller value of W and H is 4 (min(W, H) = 4), and W is 4, A is the matrix generated by removing each row corresponding to the odd x - coordinates in the down - sampling block. Alternatively, when H is 4, A is the matrix generated by removing each column corresponding to the odd y - coordinates in the down - sampling block.
[0214] Finally, in Equation 9, the reduced prediction sample can be replaced by its transpose.
[0215] [Equation 9]
[0216] ·W = H = 4 and intra_mip_transposed_flag = 1
[0217] ·max(W, H) = 8 and intra_mip_transposed_flag = 1
[0218] ·max(W, H) > 8 and intra_mip_transposed_flag = 1
[0219] When W = H = 4, since A includes four columns and 16 rows, the number of multiplications required to calculate pred red is 4. In any other case, since A includes eight columns and W red ×H red rows, at most four multiplications per sample are required to calculate pred red .
[0220] Figure 17 Illustrates the linear interpolation in the MIP process. Refer to Figure 17 Describe the linear interpolation process as follows.
[0221] The interpolation process can be called a linear interpolation process or a bilinear interpolation process. The interpolation process can include two steps, which are 1) vertical interpolation and 2) horizontal interpolation, as shown.
[0222] If W >= H, vertical linear interpolation can be applied first, followed by horizontal linear interpolation. If W < H, horizontal linear interpolation can be applied first, followed by vertical linear interpolation. In a 4×4 block, the interpolation process can be omitted.
[0223] In a W×H block (where max(W,H)≥8), the predicted samples are derived from the red ×H red reduced predicted samples pred red Depending on the block type, linear interpolation can be performed vertically, horizontally, or in both directions. When applying linear interpolation in both directions, if W < H, linear interpolation is applied first in the horizontal direction, otherwise first in the vertical direction.
[0224] For a W×H block ((where max(W,H)≥8) and W >= H), no general loss is considered to exist. In this case, one-dimensional linear interpolation is performed as follows. When there is no general loss, the linear interpolation in the vertical direction is fully explained.
[0225] First, the reduced predicted samples are extended upward by the boundary signals. When the vertical upsampling coefficient U ver = H / H red is defined and set , the extended reduced predicted samples can be set by the following formula.
[0226] [Equation 10]
[0227]
[0228] Subsequently, the vertical linear interpolation predicted samples can be generated from the extended reduced predicted samples by the following formula.
[0229] [Equation 11]
[0230]
[0231] Here, x can be 0≤x<W red , y can be 0≤y<H red , and k can be 0≤k<U ver .
[0232] In the following, a method for maximizing the performance of the MIP technology while reducing its complexity is described. The embodiments to be described below can be performed independently or in combination.
[0233] In an existing MIP mode, similar to the existing intra prediction mode derivation method, MPM flags are sent separately for non-MPM and MPM, and a MIP mode for a current block is encoded based on MPM or non-MPM.
[0234] According to an embodiment, for a block to which the MIP technology is applied, a structure for directly encoding a MIP mode without separating MPM and non-MPM can be proposed. This image encoding structure enables simplification of a complex syntax structure. Further, since the frequencies of respective modes actually occurring in the MIP mode are relatively evenly distributed, which is significantly different from the frequencies in the existing intra modes, the proposed encoding structure enables maximization of efficiency in encoding and decoding MIP mode information.
[0235] Image information for MIP transmission and reception according to an embodiment is as follows. The following syntax can be included in video / image information transmitted from the foregoing encoding device to a decoding device, and can be configured / encoded in the encoding device to be signaled to the decoding device in the form of a bitstream, and the decoding device can parse / decode the included information (syntax elements) according to the conditions / orders disclosed in the syntax.
[0236] [Table 1]
[0237]
[0238] As shown in Table 1, the syntax intra_mip_flag and intra_mip_mode_idx for concurrently signaling the MIP mode of a current block can be included in syntax information regarding an encoding unit, and in the syntax intra_mip_flag and intra_mip_mode_idx for the current block.
[0239] A value of intra_mip_flag equal to 1 indicates that the intra prediction type of a luma sample is matrix-based intra prediction, and a value equal to 0 indicates that the intra prediction type of the luma sample is not matrix-based intra prediction.
[0240] An intra_mip_mode_idx equal to 1 indicates a matrix-based intra prediction mode of a luma sample. As described above, the matrix-based intra prediction mode can indicate a matrix for MIP or a matrix and an offset.
[0241] In addition, according to the example, flag information (e.g., intra_mip_transposed_flag) indicating whether the input vector for matrix-based intra prediction is transposed can also be signaled through the syntax of the coding unit. An intra_mip_transposed_flag equal to 1 indicates that the input vector is transposed, and the number of matrices for matrix-based intra prediction can be reduced by this flag information.
[0242] intra_mip_mode_idx can be coded and decoded by the truncated binary method as shown in the following table.
[0243] [Table 2]
[0244]
[0245] As shown in Table 2, intra_mip_flag is binarized with a fixed-length code, while intra_mip_mode_idx is binarized using the truncated binary method, and the maximum binarization length (cMax) can be set according to the size of the coding block. If the width and height of the coding block are 4 (cbWidth == 4 && cbHeight == 4), the maximum binarization length can be set to 34; otherwise, the maximum binarization length can be set to 18 or 10 according to whether the width and height of the coding block are 8 or less ((cbWidth <= 8 && cbHeight <= 8)?).
[0246] intra_mip_mode_idx can be coded by the bypass method instead of the context model-based method. By coding in the bypass method, the coding speed and efficiency can be increased.
[0247] According to another example, when intra_mip_mode_idx is binarized by the truncated binary method, the maximum binarization length can be as shown in the following table.
[0248] [Table 3]
[0249]
[0250] As shown in Table 3, if the width and height of the coding block are 4 ((cbWidth == 4 && cbHeight == 4)), the maximum binarization length of intra_mip_mode_idx can be set to 15; otherwise, if the width or height of the coding block is 4 ((cbWidth == 4 || cbHeight == 4)), or if the width and height of the coding block are 8 (cbWidth == 8 && cbHeight == 8), the maximum binarization length can be set to 7, and if the width or height of the coding block is 4 ((cbWidth == 4 || cbHeight == 4)) or if the width and height of the coding block are not 8 (cbWidth == 8 && cbHeight == 8), the maximum binarization length can be set to 5.
[0251] According to another example, intra_mip_mode_idx can be encoded with a fixed - length code. In this case, to increase the coding efficiency, for each block size, the number of available MIP modes can be limited to a power of 2 (e.g., A = 2 K1 - 1, B = 2 K2 - 1, and C = 2 K3 - 1, where K1, K2, and K3 are positive integers).
[0252] The above description is shown in the following table.
[0253] [Table 4]
[0254]
[0255] In Table 4, when K1 = 5, K2 = 4, and K3 = 3, intra_mip_mode can be binarized as follows.
[0256] [Table 5]
[0257]
[0258] Alternatively, according to the example, when K1 is set to 4 and thus the width and height of the coding block are 4, the maximum binarization length of the intra_mip_mode of the block can be set to 15. Additionally, when K2 is set to 3 and thus the width or height of the coding block is 4 or the width and height of the coding block are 8, the maximum binarization length can be set to 7.
[0259] Embodiments can propose a method of using MIP only for specific blocks to which the MIP technique can be efficiently applied. When applying the method according to the embodiments, the number of matrix vectors required for MIP can be reduced and the memory required to store the matrix vectors can be significantly reduced (by 50%). With these effects, the coding efficiency remains almost the same (less than 0.1%).
[0260] The following table illustrates the syntax including specific conditions for applying the MIP technique according to the embodiments.
[0261] [Table 6]
[0262]
[0263] As shown in Table 6, a condition of applying MIP only to large blocks (cbWidth > K1 || cbHeight > K2) can be added, and the size of the block can be determined based on preset values (K1 and K2). The reason for applying MIP only to large blocks is that the coding efficiency of MIP appears in relatively large blocks.
[0264] The following table illustrates an example in which K1 and K2 in Table 6 are predefined as 8.
[0265] [Table 7]
[0266]
[0267] The semantics of intra_mip_flag and intra_mip_mode_idx in Tables 6 and 7 are the same as those shown in Table 1.
[0268] When intra_mip_mode_idx has 11 possible modes, intra_mip_mode_idx can be encoded by truncated binarization (cMax = 10) as follows.
[0269] [Table 8]
[0270]
[0271] Alternatively, when the available MIP modes are limited to eight modes, intra_mip_mode_idx[x0][y0] can be encoded with a fixed-length code as follows.
[0272] [Table 9]
[0273]
[0274] In Tables 8 and 9, intra_mip_mode_idx can be encoded by a bypass method.
[0275] Embodiments may propose a MIP technique for applying a weighted matrix (A k ) and an offset vector (B k ) for a large block to small blocks, so as to efficiently apply the MIP technique in terms of memory saving. When applying the method according to the embodiments, the number of matrix vectors required for MIP can be reduced, and the memory required to store the matrix vectors can be significantly reduced (by 50%). With these effects, the coding efficiency remains almost the same (less than 0.1%).
[0276] Figure 18 Illustrates a MIP technique according to an embodiment of the present disclosure.
[0277] As shown in the figure,[[]] Figure 18 (a) of shows the operation of the matrix and the offset vector for the large block index i, and Figure 18 (b) of shows the operation of the matrix applied to the sampling of the small block and the offset vector operation.
[0278] Referring to Figure 18 , the existing MIP process can be applied by considering the downsampled weighted matrix (Sub(A k )) obtained by downsampling the weighted matrix for the large block and the offset vector (Sub(b k )) obtained by downsampling the offset vector for the large block as the weighted matrix and the offset vector for the small blocks, respectively.
[0279] Here, downsampling can be applied only in one of the horizontal and vertical directions, or can be applied in both directions. Specifically, the downsampling factor (e.g., 1 in 2 or 1 in 4) and the vertical sampling direction or the horizontal sampling direction can be set based on the width and height of the corresponding block.
[0280] In addition, the number of intra prediction modes of MIP applying this embodiment can be set differently based on the size of the current block. For example, i) when both the height and width of the current block (coding block or transform block) are 4, 35 intra prediction modes (i.e., intra prediction modes 0 to 34) can be available, ii) when both the height and width of the current block are 8 or less than 8, 19 intra prediction modes (i.e., intra prediction modes 0 to 18) can be available, and iii) in other cases, 11 intra prediction modes (i.e., intra prediction modes 0 to 10) can be available.
[0281] For example, when the case where both the height and width of the current block are 4 is defined as block size type 0, the case where both the height and width of the current block are 8 or less is defined as block size type 1, and other cases are defined as block size type 2, the number of intra prediction modes of MIP can be as shown in the following table.
[0282] [Table 10]
[0283]
[0284] In order to apply the weighted matrix and offset vector for a large block (e.g., block size type = 2) to a small block (e.g., block size = 0 or block size = 1), the number of intra prediction modes available for each block size can be equally applied as shown in the following table.
[0285] [Table 11]
[0286]
[0287] Alternatively, as shown in Table 12 below, MIP can be applied only to block size types 1 and 2, and the weighted matrix and offset vector for block size type 2 can be downsampled for block size type 1. Thus, memory can be efficiently saved (50%).
[0288] [Table 12]
[0289]
[0290] The following drawings are provided to describe specific examples of the present disclosure. Since the specific names of the devices shown in the drawings or the names of specific signals / messages / fields are provided for illustration, the technical features of the present disclosure are not limited to the specific names used in the following drawings.
[0291] Figure 19 is a flowchart schematically illustrating a decoding method that can be performed by a decoding device according to an embodiment of the present disclosure.
[0292] Figure 19 The method shown can be performed by Figure 3 the decoding device 300 shown. Specifically, Figure 19 S1900 to S1940 of Figure 3 can be performed by the entropy decoder 310 and / or the predictor 330 (specifically, the intra predictor 331) shown, and Figure 19 S1950 of Figure 3 can be performed by the adder 340 shown. In addition, Figure 19 the method shown can be included in the foregoing embodiments of the present disclosure. Therefore, in Figure 19 , specific descriptions of details overlapping with the foregoing embodiments will be omitted or briefly described.
[0293] Referring to Figure 19 , the decoding device can receive (i.e., obtain) from the bitstream flag information indicating whether matrix-based intra prediction (MIP) is used for the current block (S1900).
[0294] The flag information is a syntax such as intra_mip_flag, and can be included and signaled in the coding unit syntax information.
[0295] The decoding device can receive matrix-based intra prediction (MIP) mode information (S1910) based on the received flag information.
[0296] The MIP mode information can be represented as intra_mip_mode_idx, and can be signaled when intra_mip_flag is equal to 1. intra_mip_mode_idx can be index information indicating the MIP mode applied to the current block, and this index information can be used to derive the matrix when generating prediction samples.
[0297] According to an example, flag information indicating whether the input vector of matrix-based intra prediction is transposed can also be signaled through the coding unit syntax, for example, intra_mip_transposed_flag.
[0298] The decoding device can generate intra prediction samples for the current block based on the MIP information. The decoding device can derive at least one neighboring reference sample from the neighboring reference samples of the current block to generate intra prediction samples, and can generate prediction samples based on the neighboring reference samples.
[0299] When applying MIP, the decoding device can derive reduced boundary samples (S1920) by downsampling the reference samples adjacent to the current block.
[0300] The reduced boundary samples can be derived by downsampling the reference samples using averaging.
[0301] When the width and height of the current block are 4, four reduced boundary samples can be derived, and in other cases, eight reduced boundary samples can be derived.
[0302] The averaging process for downsampling can be applied to each boundary of the current block (e.g., the left boundary or the upper boundary), and can be applied to the neighboring reference samples adjacent to the boundary of the current block.
[0303] According to an example, when the current block is a 4×4 block, the size of each boundary can be reduced to two samples through the averaging process, and when the current block is not a 4×4 block, the size of each boundary can be reduced to four samples through the averaging process.
[0304] Subsequently, the decoding device can derive reduced prediction samples (S1930) based on the multiplication operation of the MIP matrix derived based on the size and index information of the current block and the reduced boundary samples.
[0305] The MIP matrix can be derived based on the size of the current block and the received index information.
[0306] The MIP matrix can be selected from any one of three matrix sets classified according to the size of the current block, and each of the three matrix sets can include multiple MIP matrices.
[0307] That is to say, three matrix sets of MIP can be set, and each matrix set can include multiple matrices and multiple offset vectors. These matrix sets can be applied separately according to the size of the current block.
[0308] For example, a matrix set including 18 or 16 matrices with 16 rows and four columns and 18 or 16 offset vectors can be applied to a 4×4 block. The index information can be information indicating any one of the multiple matrices included in a matrix set.
[0309] A matrix set including 10 or 8 matrices with 16 rows and eight columns and 10 or 8 offset vectors can be applied to 4×8, 8×4, and 8×8 blocks or 4×H or W×4 blocks.
[0310] In addition, a matrix set including six matrices with 64 rows and eight columns and six offset vectors can be applied to blocks other than the foregoing blocks or blocks with a height and width of 8 or greater.
[0311] The reduced prediction samples (i.e., the prediction samples to which the MIP matrix has been applied) are derived based on the operation of adding an offset after the multiplication operation of the MIP matrix and the reduced boundary samples.
[0312] The decoding device can generate the intra prediction samples of the current block by upsampling the reduced prediction samples (S1940).
[0313] The intra prediction samples can be upsampled by linear interpolation of the reduced prediction samples.
[0314] The interpolation process can be referred to as a linear interpolation or bilinear interpolation process, and can include two steps, which are 1) vertical interpolation and 2) horizontal interpolation.
[0315] If W >= H, vertical linear interpolation can be applied first, followed by horizontal linear interpolation. If W < H, horizontal linear interpolation can be applied first, followed by vertical linear interpolation. In a 4×4 block, the interpolation process can be omitted.
[0316] The decoding device can generate the reconstructed samples of the current block based on the prediction samples (S1950).
[0317] In an embodiment, the decoding device may directly use the prediction sample as the reconstruction sample according to the prediction mode, or may generate the reconstruction sample by adding the residual sample to the prediction sample.
[0318] If there are residual samples of the current block, the decoding device may receive information about the residual of the current block. The information about the residual may include the transform coefficients of the residual samples. The decoding device may derive the residual samples (or an array of residual samples) of the current block based on the residual information. The decoding device may generate the reconstruction sample based on the prediction sample and the residual sample, and may derive the reconstructed block or the reconstructed picture based on the reconstruction sample. Thereafter, as needed, in order to enhance the subjective / objective picture quality, the decoding device may apply a deblocking filter and / or an in-loop filtering process (such as the SAO process) to the reconstructed picture, as described above.
[0319] Figure 20 is a flowchart schematically illustrating an encoding method that may be performed by an encoding device according to an embodiment of the present disclosure.
[0320] Figure 20 The method shown may be performed by Figure 2 the encoding device 200 shown. Specifically, Figure 20 S2000 to S2030 of Figure 2 may be performed by Figure 20 the predictor 220 (specifically, the intra predictor 222) shown, Figure 2 S2040 of Figure 20 may be performed by Figure 2 the subtractor 231 shown, and Figure 20 S2050 of Figure 20 may be performed by
[0321] the entropy encoder 240 shown. In addition, Figure 20 the method shown may be included in the foregoing embodiments of the present disclosure. Therefore, in
[0322] the details overlapping with the foregoing embodiments will be omitted or briefly described.
[0323] When it is determined that MIP is applied to the current block, the encoding device may derive the reduced boundary samples by downsampling the reference samples adjacent to the current block (S2010).
[0324] The reduced boundary samples may be derived by downsampling the reference samples using an average.
[0325] When the width and height of the current block are 4, four reduced boundary samples can be derived, and in other cases, eight reduced boundary samples can be derived.
[0326] The averaging process for downsampling can be applied to each boundary of the current block (e.g., the left boundary or the upper boundary), and can be applied to adjacent reference samples adjacent to the boundary of the current block.
[0327] According to an example, when the current block is a 4×4 block, the size of each boundary can be reduced to two samples through the averaging process, and when the current block is not a 4×4 block, the size of each boundary can be reduced to four samples through the averaging process.
[0328] When deriving the reduced boundary samples, the encoding device can derive the reduced prediction samples (S2020) based on the multiplication operation of the MIP matrix selected based on the size of the current block and the reduced boundary samples.
[0329] The MIP matrix can be selected from any one of three matrix sets classified according to the size of the current block, and each of the three matrix sets can include multiple MIP matrices.
[0330] That is, three matrix sets of MIP can be set, and each matrix set can include multiple matrices and multiple offset vectors. These matrix sets can be applied separately according to the size of the current block.
[0331] For example, a matrix set including 18 or 16 matrices with 16 rows and four columns and 18 or 16 offset vectors can be applied to a 4×4 block. The index information can be information indicating any one of the multiple matrices included in one matrix set.
[0332] A matrix set including 10 or eight matrices with 16 rows and eight columns and 10 or eight offset vectors can be applied to 4×8, 8×4, and 8×8 blocks or 4×H or W×4 blocks.
[0333] In addition, a matrix set including six matrices with 64 rows and eight columns and six offset vectors can be applied to blocks other than the foregoing blocks or blocks with a height and width of 8 or greater.
[0334] The reduced prediction samples (i.e., the prediction samples to which the MIP matrix has been applied) are derived based on the operation of adding an offset after the multiplication operation of the MIP matrix and the reduced boundary samples.
[0335] The encoding device can generate the intra prediction samples of the current block by upsampling the reduced prediction samples (S2030).
[0336] The intra prediction samples can be upsampled by linear interpolation of the reduced prediction samples.
[0337] The interpolation process may be referred to as a linear interpolation or bilinear interpolation process and may include two steps, which are 1) vertical interpolation and 2) horizontal interpolation.
[0338] If W >= H, vertical linear interpolation may be applied first, followed by horizontal linear interpolation. If W < H, horizontal linear interpolation may be applied first, followed by vertical linear interpolation. In a 4×4 block, the interpolation process may be omitted.
[0339] The encoding device may derive the residual samples of the current block based on the prediction samples of the current block and the original samples of the current block (S2040).
[0340] The encoding device may generate residual information of the current block based on the residual samples, and may encode the residual information and the picture information including flag information indicating whether MIP is applied and MIP mode information (S2050).
[0341] Here, the residual information may include quantization parameters, value information, position information, transform scheme, and transform kernel related to the quantized transform coefficients derived by transforming and quantizing the residual samples.
[0342] The flag information indicating whether MIP is applied is a syntax such as intra_mip_flag and may be included and encoded in the coding unit syntax information.
[0343] The MIP mode information may be represented as intra_mip_mode_idx and may be encoded when intra_mip_flag is equal to 1. intra_mip_mode_idx may be index information indicating the MIP mode applied to the current block, and this index information may be used to derive the matrix in generating the prediction samples. The index information may indicate any one of multiple matrices included in a matrix set.
[0344] According to an example, flag information indicating whether the input vector of the matrix-based intra prediction is transposed may also be signaled through the coding unit syntax, for example, intra_mip_transposed_flag.
[0345] That is, the encoding device may encode the picture information including the MIP mode information and / or the residual information of the current block, and may output the picture information as a bitstream.
[0346] The bitstream can be sent to a decoding device via a network or a (digital) storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0347] The foregoing process of generating prediction samples of the current block can be performed by Figure 2 the intra predictor 222 of the encoding device 200 shown, the process of deriving residual samples can be performed by Figure 2 the subtractor 231 of the encoding device 200 shown, and the process of generating and encoding residual information can be performed by Figure 2 the residual processor 230 and the entropy encoder 240 of the encoding device 200 shown.
[0348] In the above embodiments, the method is explained based on a flowchart by means of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step can be performed in an order or steps different from the above order or steps, or a certain step can be performed concurrently with other steps. In addition, those of ordinary skill in the art can understand that the steps shown in the flowchart are not exclusive, and one or more steps in the flowchart can be incorporated or deleted without affecting the scope of the present disclosure.
[0349] The above method according to the present disclosure can be implemented in software form, and the encoding device and / or decoding device according to the present disclosure can be included in devices for image processing such as a television, a computer, a smart phone, a set-top box, and a display device.
[0350] When the embodiments in the present disclosure are implemented by software, the above method can be implemented as modules (steps, functions, etc.) for performing the above functions. These modules can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor in various well-known ways. The processor can include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory can include a read only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present disclosure can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, the information for implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.
[0351] In addition, the decoding device and the encoding device to which this document is applied may be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home movie video device, a digital movie video device, a camera for surveillance, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VOD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a video phone device, a transportation device terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an aircraft terminal, and a ship terminal), and a medical video device, and may be used to process video signals or data signals. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, and a digital video recorder (DVR).
[0352] In addition, the processing method applied in this document may be generated in the form of a program executed by a computer and may be stored in a computer-readable recording medium. Multimedia data having a data structure according to this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated using an encoding method may be stored in a computer-readable recording medium or may be transmitted through a wired and wireless communication network.
[0353] In addition, an embodiment of this document may be implemented as a computer program product using program code. The program code may be executed by a computer according to an embodiment of this document. The program code may be stored on a computer-readable carrier wave.
[0354] Figure 21 Examples of a content streaming system to which the embodiments disclosed in this document may be applied are illustrated.
[0355] Referring to Figure 21 , a content streaming system to which the embodiments of this document are applied may basically include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0356] The encoding server compresses the content input from a multimedia input device such as a smart phone, a camera, a portable video camera, etc. into digital data to generate a bitstream, and sends the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, a camera, a portable video camera, etc. directly generates a bitstream, the encoding server can be omitted.
[0357] The bitstream can be generated by applying the encoding method or the bitstream generation method of the embodiments of this document, and the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0358] The streaming server sends the multimedia data to the user device through the web server based on the user's request, and the web server serves as a medium for notifying the user of the service. When the user requests a desired service from the web server, the web server passes it to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between the devices in the content streaming system.
[0359] The streaming server can receive the content from the media storage device and / or the encoding server. For example, when receiving the content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined time.
[0360] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc.
[0361] Each server in the content streaming system can be operated as a distributed server, in which case, the data received from each server can be distributed.
[0362] The claims in this specification can be combined in various ways. For example, the technical features in the method claims of this specification can be combined to be implemented or executed in a device, and the technical features in the device claims can be combined to be implemented or executed in a method. In addition, the technical features in the method claims and the device claims can be combined to be implemented or executed in a device. In addition, the technical features in the method claims and the device claims can be combined to be implemented or executed in a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Receiving flag information related to whether matrix-based intra prediction (MIP) is used for a current block; Receiving matrix-based intra prediction (MIP) mode information based on the flag information; Generating intra prediction samples for the current block based on the MIP mode information; And Generating reconstructed samples for the current block based on the intra prediction samples, wherein the MIP mode information is index information related to an MIP matrix applied to the current block, and wherein the MIP matrix is derived based on the size of the current block and the index information.
2. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: Deriving whether matrix-based intra prediction (MIP) is applied to a current block; Based on the MIP being applied to the current block, deriving intra prediction samples for the current block based on the MIP; Deriving residual samples for the current block based on the intra prediction samples; And Encoding information about the residual samples and information about the MIP, wherein the information about the MIP includes flag information related to whether the MIP is applied to the current block and matrix-based intra prediction (MIP) mode information, wherein the MIP mode information is index information related to an MIP matrix applied to the current block, and wherein the index information indicates any one of a plurality of MIP matrices included in a matrix set, wherein the MIP matrix is determined based on the size of the current block and the index information.
3. A sending method, the sending method comprising the following steps: Obtaining a bitstream, wherein the bitstream is generated by: deriving whether matrix-based intra prediction (MIP) is applied to a current block; based on the MIP being applied to the current block, deriving intra prediction samples for the current block based on the MIP; deriving residual samples for the current block based on the intra prediction samples; and encoding information about the residual samples and information about the MIP to generate the bitstream; and Sending the bitstream, wherein the information about the MIP includes flag information related to whether the MIP is applied to the current block and matrix-based intra prediction (MIP) mode information, wherein the MIP mode information is index information related to an MIP matrix applied to the current block, and wherein the index information indicates any one of a plurality of MIP matrices included in a matrix set, wherein the MIP matrix is determined based on the size of the current block and the index information.