Decoding device, encoding device, and data transmission device
Through the matrix-based intra prediction method, the efficient compression and transmission of high-resolution, high-quality image/video data is solved, the encoding efficiency is improved and complexity is reduced, and it is suitable for the field of image encoding technology.
Patent Information
- Application Number
- CN202510833980.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-03
- Filing Date
- 2020-06-03
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art is difficult to efficiently compress and transmit high-resolution, high-quality image/video data, especially in the case of increased transmission and storage costs, and the demand for immersive media such as VR and AR content is growing.
Using a matrix-based intra prediction method, an intra prediction sample of an image is generated by receiving and deriving matrix-based intra prediction mode information, and encoding efficiency is improved through downsampling and upsampling operations.
Improves image/video compression efficiency, reduces implementation complexity, and enhances prediction performance, thereby improving overall coding efficiency.
Smart Images

Figure CN120416474A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the original application number 202080041446.2 (International Application No.: PCT / KR2020 / 007215, filing date: June 3, 2020, invention title: Matrix-based Intra Prediction Apparatus and Method). Technical Field
[0002] The present disclosure relates to image coding technology, and more particularly, to image coding technology for a matrix-based intra prediction device and matrix-based intra prediction. Background Art
[0003] These days, the demand for high-resolution and high-quality images / videos such as 4K, 8K, or higher ultra-high definition (UHD) images / videos has been continuously increasing in various fields. As image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases compared to traditional image data. Therefore, when transmitting image data using a medium such as a traditional wired / wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0004] In addition, these days, the interest in and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and the broadcasting of images / videos having image characteristics different from those of real images such as game images is increasing.
[0005] Therefore, there is a need for an efficient image / video compression technology that can effectively compress, transmit, store, and reproduce information of high-resolution and high-quality images / videos having various characteristics as described above. Summary of the Invention
[0006] Technical Problem
[0007] One technical aspect of the present disclosure is to provide a method and apparatus for increasing image coding efficiency.
[0008] Another technical aspect of the present disclosure is to provide an efficient intra prediction method and an efficient intra prediction apparatus.
[0009] Still another technical aspect of the present disclosure is to provide an image coding method and an image coding apparatus for matrix-based intra prediction.
[0010] Still another technical aspect of the present disclosure is to provide an image coding method and an image coding apparatus for coding mode information regarding matrix-based intra prediction.
[0011] Technical Solution
[0012] According to an embodiment of the present disclosure, an image decoding method executed by a decoding device is provided. The method may include: receiving flag information indicating whether matrix-based intra prediction (MIP) is used for a current block; receiving matrix-based intra prediction (MIP) mode information based on the flag information; generating intra prediction samples for the current block based on the MIP mode information; and generating reconstructed samples for the current block based on the intra prediction samples.
[0013] The MIP mode information may be index information indicating the MIP mode applied to the current block.
[0014] The generation of the intra prediction samples may include: deriving reduced boundary samples by downsampling reference samples adjacent to the current block; deriving reduced prediction samples based on a multiplication operation of the reduced boundary samples and an MIP matrix; and generating intra prediction samples for the current block by upsampling the reduced prediction samples.
[0015] The reduced boundary samples may be downsampled by averaging the reference samples, and the intra prediction samples may be upsampled by linearly interpolating the reduced prediction samples.
[0016] The MIP matrix may be derived based on the size of the current block and the index information.
[0017] The MIP matrix may be selected from any one of three matrix sets classified according to the size of the current block, and each of the three matrix sets may include a plurality of MIP matrices.
[0018] According to another embodiment of the present disclosure, an image encoding method executed by an encoding device is provided. The method may include: deriving whether matrix-based intra prediction MIP is applied to a current block; when MIP is applied to the current block, deriving intra prediction samples for the current block based on MIP; deriving residual samples for the current block based on the intra prediction samples; and encoding information about the residual samples and information about MIP, wherein the information about MIP may include flag information indicating whether MIP is applied to the current block and matrix-based intra prediction (MIP) mode information.
[0019] According to still another embodiment of the present disclosure, a digital storage medium may be provided, which stores image data including encoded image information and a bitstream generated according to an image encoding method executed by an encoding device.
[0020] According to still another embodiment of the present disclosure, a digital storage medium may be provided, which stores image data including encoded image information and a bitstream so that a decoding device executes an image decoding method.
[0021] Technical effects
[0022] The present disclosure can have various effects. For example, according to embodiments of the present disclosure, the overall image / video compression efficiency can be increased. In addition, according to embodiments of the present disclosure, the implementation complexity can be reduced and the prediction performance can be enhanced through efficient intra prediction, thereby improving the overall coding efficiency. Additionally, according to embodiments of the present disclosure, when performing matrix-based intra prediction, the index information indicating matrix-based intra prediction can be efficiently encoded, thereby improving the coding efficiency.
[0023] The effects that can be obtained through the detailed examples in the specification are not limited to the above effects. For example, there can be various technical effects that those of ordinary skill in the relevant art can understand or derive from the specification. Therefore, the detailed effects of the specification are not limited to those explicitly described in the specification, and can include various effects that can be understood or derived from the technical features of the specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 An example of a video / image coding system to which embodiments of the present disclosure can be applied is schematically illustrated.
[0025] Figure 2 The configuration of a video / image coding device to which embodiments of the present disclosure can be applied is schematically illustrated.
[0026] Figure 3 The configuration of a video / image decoding device to which embodiments of the present disclosure can be applied is schematically illustrated.
[0027] Figure 4 An example of an intra prediction-based image coding method to which embodiments of the present disclosure can be applied is schematically illustrated.
[0028] Figure 5 An intra predictor in an encoding device is schematically illustrated.
[0029] Figure 6 An example of an intra prediction-based image decoding method to which embodiments of the present disclosure can be applied is schematically illustrated.
[0030] Figure 7 An intra predictor in a decoding device is schematically illustrated.
[0031] Figure 8 An example of an MPM-based intra prediction method in an encoding device to which embodiments of the present disclosure can be applied is illustrated.
[0032] Figure 9 An example of an MPM-based intra prediction method in a decoding device to which embodiments of the present disclosure can be applied is illustrated.
[0033] Figure 10 An example of an intra prediction mode to which the embodiments of the present disclosure are applicable is illustrated.
[0034] Figure 11 A MIP-based prediction sample generation process according to an example is illustrated.
[0035] Figure 12 The MIP process of a 4×4 block is illustrated.
[0036] Figure 13 The MIP process of an 8×8 block is illustrated.
[0037] Figure 14 The MIP process of an 8×4 block is illustrated.
[0038] Figure 15 The MIP process of a 16×16 block is illustrated.
[0039] Figure 16 The boundary averaging process in the MIP process is illustrated.
[0040] Figure 17 Illustrate linear interpolation in the MIP process.
[0041] Figure 18 The MIP technology according to the embodiment of the present disclosure is illustrated.
[0042] Figure 19 is a flowchart schematically illustrating a decoding method that can be performed by a decoding device according to an embodiment of the present disclosure.
[0043] Figure 20 is a flowchart schematically illustrating an encoding method that can be performed by an encoding device according to an embodiment of the present disclosure.
[0044] Figure 21 An example of a content streaming system to which embodiments of the present disclosure are applicable is illustrated. DETAILED DESCRIPTION
[0045] This document can be modified in various ways and can have various embodiments, and specific embodiments will be illustrated and described in detail in the accompanying drawings. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are used to describe specific embodiments and are not used to limit the technical spirit of this document. Unless otherwise clearly stated in the context, expressions in the singular include plural expressions. Terms such as "including" or "having" in this specification should be understood to indicate the presence of characteristics, quantities, steps, operations, elements, parts or combinations thereof described in the specification, without excluding the presence of one or more other characteristics, quantities, steps, operations, elements, parts or combinations thereof or the possibility of adding one or more other characteristics, quantities, steps, operations, elements, parts or combinations thereof.
[0046] The elements in the drawings described in this document are illustrated independently to facilitate descriptions related to different feature functions. This does not mean that each element is implemented as separate hardware or separate software. For example, at least two of the elements can be combined to form a single element, or a single element can be divided into multiple elements. Implementations in which elements are combined and / or separated are also included in the scope of the present document unless it deviates from the essence of this document.
[0047] In this document, the term "A or B" may mean "only A," "only B," or "both A and B." In other words, in this document, the term "A or B" may be interpreted to mean "A and / or B." For example, in this document, the term "A, B, or C" may mean "only A," "only B," "only C," or "any combination of A, B, and C."
[0048] As used in this document, a slash " / " or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Thus, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0049] In this document, “at least one of A and B” may mean “only A”, “only B”, or “both A and B”. In addition, in this document, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted as “at least one of A and B”.
[0050] In addition, in this document, "at least one of A, B, and C" may mean "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0051] Additionally, the parentheses used in this document can mean "for example". Specifically, in the expression "prediction (intra prediction)", this can indicate that "intra prediction" is presented as an example of "prediction". In other words, "prediction" in this document is not limited to "intra prediction", and this can indicate that "intra prediction" is presented as an example of "prediction". Additionally, even in the expression "prediction (i.e., intra prediction)", this can also indicate that "intra prediction" is presented as an example of "prediction".
[0052] In this document, the technical features separately described in one drawing can be implemented separately or can be implemented simultaneously.
[0053] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the Versatile Video Coding (VVC), Essential Video Coding (EVC) standard, AOMedia Video 1 (AV1) standard, Second Generation Audio Video Coding Standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.).
[0054] This document presents various embodiments of video / image coding, and unless otherwise stated, the embodiments can be executed in combination with each other.
[0055] In this document, video may refer to a series of images over time. A picture generally refers to a unit representing an image in a specific time region, while a slice / tile is a unit that is part of an encoded picture. A slice / tile may include one or more coding tree units (CTUs). A picture may consist of one or more slices / titles. A picture may consist of one or more tile groups. A tile group may include one or more tiles. A brick may represent a rectangular region of CTU rows within a tile in a picture. A tile may be divided into multiple bricks, each brick consisting of one or more CTU rows within the tile. A tile that is not divided into multiple bricks may also be referred to as a brick. Brick scan is a specific sequential ordering of the CTUs of the divided tiles within a brick, sequentially ordered in a CTU grid scan in a picture, and the tiles in a picture are sequentially ordered in a grid scan of the tiles of the picture. A tile is a rectangular region of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular region of CTUs, whose height is equal to the height of the picture and whose width is specified by a syntax element in the picture parameter set. A tile row is a rectangular region of CTUs, whose height is specified by a syntax element in the picture parameter set and whose width is equal to the width of the picture. Tile scan is a specific sequential ordering of the CTUs of the divided pictures within a tile, sequentially ordered in a CTU grid scan in a picture, while the tiles in a picture are sequentially ordered in a grid scan of the tiles of the picture. A slice includes an integral number of bricks of a picture that can be exclusively included in a single NAL unit. A slice may be composed of multiple complete tiles or only of a continuous sequence of complete bricks of a single tile. Tile group and slice may be used interchangeably in this document. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0056] A pixel or pel may refer to the smallest unit that constitutes a picture (or image). Additionally, "sample" may be used as a term corresponding to a pixel. A sample generally may represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0057] A unit may represent the basic unit of image processing. A unit may include at least one of a specific region and information related to that region. A unit may include one luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the situation, terms such as unit and block, region, etc. may be used interchangeably. Generally, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0058] In the following, preferred embodiments of this document are described more specifically with reference to the accompanying drawings. In the following, in the drawings, the same reference numerals are used for the same elements, and redundant descriptions of the same elements may be omitted.
[0059] Figure 1 An example of a video / image encoding system to which embodiments of the present disclosure can be applied is schematically illustrated.
[0060] Referring Figure 1 , the video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transfer the encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.
[0061] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0062] The video source may obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generating device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capturing process may be replaced by a process of generating relevant data.
[0063] The encoding device may encode the input video / image. The encoding device may perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.
[0064] The transmitter may send the encoded video / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating a media file in a predetermined file format and may include elements for transmitting via a broadcast / communication network. The receiver may receive / extract the bitstream and send the received / extracted bitstream to the decoding device.
[0065] The decoding device can decode video / images by performing a series of processes such as dequantization, inverse transformation, prediction, etc., corresponding to the operations of the encoding device.
[0066] The renderer can render the decoded video / images. The rendered video / images can be displayed through a display.
[0067] Figure 2 FIG. is a diagram schematically illustrating the configuration of a video / image encoding device to which the present disclosure can be applied. Hereinafter, the so-called video encoding device may include an image encoding device.
[0068] Refer to Figure 2 , the encoding device 200 may include an image splitter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the above-described image splitter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be constituted by one or more hardware components (e.g., an encoder chipset or a processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0069] The image divider 210 may divide an input image (or picture or frame) input to the encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a largest coding unit (LCU), the coding unit may be recursively divided according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, based on a quadtree structure, a binary tree structure, and / or a ternary tree structure, a coding unit may be divided into multiple coding units with a deeper depth. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or the ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final coding unit that is not further divided. In this case, based on the encoding efficiency according to the image characteristics, the largest coding unit may be directly used as the final coding unit. Alternatively, the coding unit may be recursively divided into coding units with a deeper depth as needed, whereby the coding unit with the optimal size may be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, a processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be separate from or divided from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.
[0070] Depending on the situation, the terms unit and terms such as block, region, etc. may be used in place of each other. In general, an M×N block may represent a set of samples or transformation coefficients composed of M columns and N rows. Samples generally may represent pixels or pixel values, and may represent only the pixels / pixel values of the luminance component, or only the pixels / pixel values of the chrominance component. Samples may be used as a term corresponding to the pixels or pels of a picture (or image).
[0071] In the encoding device 200, a prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 is subtracted from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown in the figure, the unit in the encoding device 200 for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) may be referred to as the subtractor 231. The predictor may perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As described later in the description of each prediction mode, the predictor may generate various information related to prediction such as prediction mode information and send the generated information to the entropy encoder 240. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0072] The intra-frame predictor 222 may predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the reference samples may be located near the current block or separated from the current block. In intra-frame prediction, the prediction mode may include a variety of non-directional modes and a variety of directional modes. The non-directional modes may include, for example, the DC mode and the planar mode. Depending on the level of detail of the prediction direction, the directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-frame predictor 222 may determine the prediction mode to be applied to the current block by using the prediction mode applied to the neighboring blocks.
[0073] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information can be predicted on a block, sub-block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, the neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks can be the same as or different from each other. The temporal neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including the temporal neighboring blocks can be referred to as a collocated picture (colPic). For example, the inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. The inter-frame prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter-frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, the residual signal cannot be transmitted. In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of a neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling a motion vector difference.
[0074] The predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction at the same time. This can be referred to as combined inter-frame and intra-frame prediction (CIIP). In addition, the predictor can perform prediction of a block based on the intra-block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter-frame prediction because a reference block is derived in the current picture. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be regarded as an example of intra-frame coding or intra-frame prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information about the palette table and the palette index.
[0075] The prediction signal generated by a predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) may be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT means a transform obtained from a graph when relationship information between pixels is represented by a graph. CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform process may be applied to square pixel blocks of the same size, or may be applied to blocks of variable size rather than square blocks.
[0076] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, and entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-shaped quantized transform coefficients into a one-dimensional vector based on the coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the form of a one-dimensional vector. Information about the quantized transform coefficients can be generated. Entropy encoder 240 can perform various coding methods such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. Entropy encoder 240 can encode together or separately the information required for video / image reconstruction other than the quantized transform coefficients (e.g., the values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) can be sent or stored in the form of a bitstream in units of network abstraction layer (NAL). The video / image information can also include information about various parameter sets such as adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). Additionally, the video / image information can also include general constraint information. In this document, the information and / or syntax elements sent from the encoding device to / signaled to the decoding device can be included in the video / picture information. The video / image information can be encoded through the above encoding process and included in the bitstream. The bitstream can be sent through a network or stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for sending the signal output from entropy encoder 240 or a storage unit (not shown) for storing the signal can be included as an internal / external element of encoding device 200, and alternatively, the transmitter can be included in entropy encoder 240.
[0077] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients via the dequantizer 234 and the inverse transformer 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If there is no residual for the block to be processed, such as in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture and can be used for inter-prediction of the next picture through filtering described below.
[0078] In addition, a luminance mapping with chroma scaling (LMCS) can be applied during picture encoding and / or reconstruction.
[0079] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0080] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-prediction unit 221. When inter-prediction is applied by the encoding device, prediction mismatch between the encoding device 200 and the decoding device can be avoided, and encoding efficiency can be improved.
[0081] The DPB of the memory 270 can store the modified reconstructed picture to be used as a reference picture in the inter-prediction unit 221. The memory 270 can store the motion information of the blocks with derived (or encoded) motion information in the current picture and / or the motion information of the already reconstructed blocks in the picture. The stored motion information can be sent to the inter-prediction unit 221 and used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and can transfer the reconstructed samples to the intra-prediction unit 222.
[0082] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which embodiments of this document can be applied.
[0083] Referring to Figure 3 , the video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra predictor 331 and an inter predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be constituted by one or more hardware components (e.g., a decoder chipset or a processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0084] When receiving a bitstream including video / image information, the decoding device 300 may reconstruct an image corresponding to the processing of the video / image information that has been performed in the Figure 2 encoding device accordingly. For example, the decoding device 300 may derive units / blocks based on information related to block segmentation obtained from the bitstream. The decoding device 300 may perform decoding by using the processing units applied in the encoding device. Thus, the decoding processing unit may be, for example, an encoding unit, which may be divided along a quadtree structure, a binary tree structure, and / or a ternary tree structure with a coding tree unit or a largest coding unit. One or more transform units may be derived with the coding unit. And, the reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproducer.
[0085] The decoding device 3 can receive in the form of a bitstream from Figure 2The signal output by the encoding device, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information can also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or the general constraint information. The signaled / received information and / or syntax elements described later in this document can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 decodes the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, or CABAC, and outputs the syntax elements required for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, determine the context model using the information of the symbols / bins decoded in the previous stage, the information of the syntax element to be decoded, or the decoded information of the block to be decoded, and perform arithmetic decoding on the bins by predicting the occurrence probability of the bins according to the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins of the context model for the next symbol / bin after determining the context model. The information related to prediction among the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values that have undergone entropy decoding in the entropy decoder 310, i.e., the quantized transform coefficients and the related parameter information, can be input to the residual processor 320. The residual processor 320 can derive the residual signal (residual block, residual sample, residual sample array). Additionally, the information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. Furthermore, the receiver (not shown) for receiving the signal output by the encoding device can also be configured as an internal / external component of the decoding device 300, or the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of the dequantizer 321, inverse transformer 322, adder 340, filter 350, memory 360, inter-frame predictor 332, and intra-frame predictor 331.
[0086] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in the form of a two-dimensional block. In this case, the rearrangement can be performed based on the order of coefficient scanning that has been performed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.
[0087] The inverse transformer 322 obtains a residual signal (residual block, residual sample array) by performing an inverse transform on the transform coefficients.
[0088] The predictor can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and specifically can determine the intra / inter prediction mode.
[0089] The predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). Additionally, the predictor can perform prediction of a block based on the intra block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, e.g., screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction because a reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information about the palette table and the palette index.
[0090] The intra predictor 331 can predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the samples referred to can be located in the neighbors of the current block or can be located separately. In intra prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to the adjacent block.
[0091] The inter-frame predictor 332 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, the adjacent blocks may include spatially adjacent blocks present in the current picture and temporally adjacent blocks present in the reference picture. For example, the inter-frame predictor 332 may construct a motion information candidate list based on the adjacent blocks and derive the motion vector and / or the reference picture index of the current block based on the received candidate selection information. The inter-frame prediction may be performed based on various prediction modes, and the information regarding the prediction may include information indicating the mode for the inter-frame prediction of the current block.
[0092] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (prediction block, prediction sample array) output from a predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331). If there is no residual for the block to be processed, such as when the skip mode is applied, the prediction block may be used as the reconstructed block.
[0093] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for the intra-frame prediction of the next block to be processed in the current picture, may be output through filtering described below, or may be used for the inter-frame prediction of the next picture.
[0094] In addition, a luminance mapping with chroma scaling (LMCS) may be applied during the picture decoding process.
[0095] The filter 350 may improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 may generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 360 (specifically, the DPB of the memory 360). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0096] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter - predictor 332. The memory 360 can store the motion information of the blocks of the derived (or decoded) motion information in the current picture and / or the motion information of the reconstructed blocks in the picture. The stored motion information can be sent to the inter - predictor 332 and thus be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and transmit the reconstructed samples to the intra - predictor 331.
[0097] In this document, the embodiments described in the filter 260, the inter - predictor 221, and the intra - predictor 222 of the encoding device 200 can be the same as, or be applied respectively to correspond to, the filter 350, the inter - predictor 332, and the intra - predictor 331 of the decoding device 300.
[0098] As described above, when performing video encoding, prediction is performed to enhance the compression efficiency. A prediction block including prediction samples of a current block (i.e., a target encoding block) can be generated through prediction. In this case, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived equally in the encoding device and the decoding device. The encoding device can enhance the image encoding efficiency by signaling information about the residual between the original block and the prediction block (residual information) instead of the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, can generate a reconstructed picture including reconstructed samples by adding the residual block and the prediction block, and can generate a reconstructed picture including reconstructed blocks.
[0099] The residual information can be generated through a transformation and quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, can derive transform coefficients by performing a transformation process on the residual samples (residual sample array) included in the residual block, can derive quantized transform coefficients by performing a quantization process on the transform coefficients, and can signal the relevant residual information to the decoding device (through the bitstream). In this case, the residual information can include information such as value information, position information, transformation scheme, transformation kernel, and quantization parameters of the quantized transform coefficients. The decoding device can perform a de - quantization / inverse - transformation process based on the residual information and can derive the residual samples (or the residual block). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. In addition, the encoding device can derive a residual block by performing de - quantization / inverse - transformation on the quantized transform coefficients for reference in the inter - prediction of subsequent pictures and can generate a reconstructed picture.
[0100] In addition, if intra prediction is performed, the correlation between samples can be used, and the difference (i.e., the residual) between the original block and the predicted block can be obtained. The foregoing transformation and quantization can be applied to the residual. Thus, spatial redundancy can be reduced. Hereinafter, encoding methods and decoding methods using intra prediction are specifically described.
[0101] Intra prediction refers to prediction for generating a predicted sample of a current block based on reference samples outside the current block in a picture (hereinafter referred to as the current picture) including the current block. In this case, the reference samples outside the current block may refer to samples adjacent to the current block. If intra prediction is applied to the current block, adjacent reference samples to be used for intra prediction of the current block can be derived.
[0102] For example, when the size (width × height) of the current block is nW × nH, the adjacent reference samples of the current block may include a total of 2 × nH samples adjacent to the left boundary and adjacent to the lower left of the current block, a total of 2 × nW samples adjacent to the upper boundary and adjacent to the upper right of the current block, and samples adjacent to the upper left of the current block. Alternatively, the adjacent reference samples of the current block may further include multiple columns of adjacent samples above and multiple rows of adjacent samples to the left. In addition, the adjacent reference samples of the current block may further include a total of nH samples adjacent to the right boundary of the current block having a size of nW × nH, a total of nW samples adjacent to the lower boundary of the current block, and one sample adjacent to the lower right of the current block.
[0103] In this case, some of the adjacent reference samples of the current block have not been decoded or may be unavailable. In this case, the decoding device can configure the adjacent reference samples to be used for prediction by replacing the unavailable samples with available samples. Alternatively, the adjacent reference samples to be used for prediction can be constructed by interpolation of available samples.
[0104] If adjacent reference samples are derived, (i) a predicted sample can be derived based on the average or interpolation of the adjacent reference samples of the current block, and (ii) a predicted sample can be derived based on the reference samples existing in a specific (prediction) direction among the adjacent reference samples of the current block for the predicted sample. When the intra prediction mode is a non - directional mode or a non - angular mode, (i) can be applied. When the intra prediction mode is a directional mode or an angular mode, (ii) can be applied.
[0105] In addition, a predicted sample can also be generated by interpolation between a first adjacent sample located in the prediction direction of the intra prediction mode of the current block and a second adjacent sample located in the opposite direction of the prediction direction based on the predicted sample of the current block among the adjacent reference samples. The foregoing case can be referred to as linear interpolation intra prediction (LIP). In addition, a chrominance prediction sample can also be generated based on luminance samples using a linear model. This case can be referred to as the LM mode.
[0106] In addition, the temporary prediction samples of the current block can be derived based on filtered neighboring reference samples, and the prediction samples of the current block can also be derived by performing a weighted sum of at least one reference sample and the temporary prediction samples derived according to the intra prediction mode among the regular neighboring reference samples (i.e., unfiltered neighboring reference samples). The foregoing case can be referred to as position-dependent intra prediction (PDPC).
[0107] In addition, the prediction samples can be derived by using the reference samples positioned in the prediction direction of the corresponding line by selecting the reference sample line with the highest prediction accuracy among multiple neighboring reference sample lines of the current block, and the intra prediction coding can be performed by a method for indicating (signaling) the reference sample line used at this time to the decoding device. The foregoing case can be referred to as multi-reference line (MRL) intra prediction or MRL-based intra prediction.
[0108] In addition, the intra prediction can be performed based on the same intra prediction mode by separating the current block into vertical or horizontal sub-partitions, and the neighboring reference samples can be derived and used in units of sub-partitions. That is, in this case, the intra prediction mode of the current block is equally applied to the sub-partitions, and the neighboring reference samples can be derived and used in units of sub-partitions, thereby enhancing the intra prediction performance in some cases. This prediction method can be referred to as intra sub-partition (ISP) intra prediction or ISP-based intra prediction.
[0109] The foregoing intra prediction methods can be referred to as intra prediction types separately from the intra prediction modes. The intra prediction types can be referred to by various terms, such as intra prediction techniques or additional intra prediction modes. For example, the intra prediction types (or additional intra prediction modes, etc.) can include at least one of the foregoing LIP, PDPC, MRL, and ISP. The general intra prediction method other than the specific intra prediction types such as LIP, PDPC, MRL, and ISP can be referred to as the normal intra prediction type. If the specific intra prediction type is not applied, then the normal intra prediction type can generally be applied, and the prediction can be performed based on the foregoing intra prediction mode. In addition, if necessary, post-processing filtering for the derived prediction samples can also be performed.
[0110] In addition to the above intra prediction types, matrix-based intra prediction (hereinafter referred to as MIP) can be used as a method for intra prediction. MIP can be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP).
[0111] When MIP is applied to a current block, prediction samples for the current block can be derived as follows: i) using neighboring reference samples that have undergone an averaging process; ii) performing a matrix-vector multiplication process; and (iii) further performing a horizontal / vertical interpolation process if necessary. The intra prediction mode for MIP can be configured differently from the intra prediction modes used for the aforementioned LIP, PDPC, MRL, or ISP intra prediction or for normal intra prediction.
[0112] The intra prediction mode of MIP can be referred to as an affine linear weighted intra prediction mode or a matrix-based intra prediction mode. For example, the matrix and offset used in the matrix-vector multiplication can be configured differently according to the intra prediction mode of MIP. Here, the matrix can be referred to as an (affine) weighting matrix, and the offset can be referred to as an (affine) offset vector or an (affine) bias vector. In the present disclosure, the intra prediction mode of MIP can be referred to as the MIP intra prediction mode, the linear weighted intra prediction mode, the matrix weighted intra prediction mode, or the matrix-based intra prediction mode. Specific MIP methods will be described later.
[0113] The following drawings have been prepared to explain specific examples of this document. Since the names of specific devices or specific words or names (e.g., grammatical names, etc.) described in the drawings are presented exemplarily, the technical features of this document are not limited to the specific names used in the following drawings.
[0114] Figure 4 An example of an intra prediction-based image coding method to which embodiments of the present disclosure can be applied is schematically illustrated, and Figure 5 an intra predictor in an encoding device is schematically illustrated. Figure 5 The intra predictor in the illustrated encoding device can also be equally or correspondingly applied to Figure 2 the intra predictor 222 of the illustrated encoding device 200.
[0115] Refer to Figure 4 and Figure 5, S400 can be performed by the intra predictor 222 of the encoding device, and S410 can be performed by the residual processor 230 of the encoding device. Specifically, S410 can be performed by the subtractor 231 of the encoding device. In S420, the prediction information can be derived by the intra predictor 222 and encoded by the entropy encoder 240. In S420, the residual information can be derived by the residual processor 230 and encoded by the entropy encoder 240. The residual information indicates information about the residual samples. The residual information can include information about the quantized transform coefficients of the residual samples. As described above, the residual samples can be derived based on the transform coefficients passed through the transformer 232 of the encoding device, and the transform coefficients can be derived based on the quantized transform coefficients passed through the quantizer 233. The information about the quantized transform coefficients can be encoded by the entropy encoder 240 through the residual encoding process.
[0116] The encoding device performs intra prediction (S400) on the current block. The encoding device can derive the intra prediction mode / type of the current block, derive the neighboring reference samples of the current block, and generate the prediction samples in the current block based on the intra prediction mode / type and the neighboring reference samples. Here, the processes of determining the intra prediction mode / type, deriving the neighboring reference samples, and generating the prediction samples can also be performed simultaneously, and any one of the processes can also be performed earlier than the other processes.
[0117] For example, the intra predictor 222 of the encoding device can include an intra prediction mode / type determiner 222-1, a reference sample deriver 222-2, and a prediction sample deriver 222-3, where the intra prediction mode / type determiner 222-1 can determine the intra prediction mode / type of the current block, the reference sample deriver 222-2 can derive the neighboring reference samples of the current block, and the prediction sample deriver 222-3 can derive the prediction samples of the current block. In addition, although not shown, if the prediction sample filtering process is performed, the intra predictor 222 can also further include a prediction sample filter (not shown). The encoding device can determine the mode / type applied to the current block among multiple intra prediction modes / type. The encoding device can compare the RD costs of the intra prediction modes / type and determine the best intra prediction mode / type of the current block.
[0118] As described above, the encoding device can also perform the prediction sample filtering process. The prediction sample filtering can be referred to as post-filtering. Some or all of the prediction samples can be filtered through the prediction sample filtering process. In some cases, the prediction sample filtering process can be omitted.
[0119] The encoding device generates the residual samples of the current block based on the (filtered) prediction samples (S410). The encoding device can compare the prediction samples based on the phases in the original samples of the current block and derive the residual samples.
[0120] The encoding device may encode image information including information on intra prediction (prediction information) and residual information on residual samples (S420). The prediction information may include intra prediction mode information and intra prediction type information. The residual information may include residual coding syntax. The encoding device may derive quantized transform coefficients by transforming / quantizing the residual samples. The residual information may include information on the quantized transform coefficients.
[0121] The encoding device may output the encoded image information in the form of a bitstream. The output bitstream may be delivered to the decoding device via a storage medium or a network.
[0122] As described above, the encoding device may generate a reconstructed picture (including reconstructed samples and reconstructed blocks). To this end, the encoding device may derive (modified) residual samples by dequantizing / inverse-transforming the quantized transform coefficients again. As described above, the reason for transforming / quantizing the residual samples and then dequantizing / inverse-transforming them again is to derive the same residual samples as those derived by the decoding device, as described above. The encoding device may generate a reconstructed block including the reconstructed samples of the current block based on the prediction samples and the (modified) residual samples. A reconstructed picture of the current picture may be generated based on the reconstructed blocks. As described above, an in-loop filtering process or the like may also be applied to the reconstructed picture.
[0123] Figure 6 An example of an intra prediction-based image decoding method to which embodiments of the present disclosure may be applied is schematically illustrated, and Figure 7 an intra predictor in the decoding device is schematically illustrated. Figure 7 The intra predictor in the illustrated decoding device may also be equally or correspondingly applied to Figure 3 the intra predictor 331 of the illustrated decoding device 300.
[0124] Referring to Figure 6 and Figure 7 , the decoding device may perform operations corresponding to the foregoing operations performed by the encoding device. S600 to S620 may be performed by the intra predictor 331 of the decoding device, and the entropy decoder 310 of the decoding device may obtain the prediction information in S600 and the residual information in S630 from the bitstream. The residual processor 320 of the decoding device may derive the residual samples of the current block based on the residual information. Specifically, the dequantizer 321 of the residual processor 320 may derive transform coefficients by performing dequantization based on the quantized transform coefficients derived based on the residual information, and the inverse transformer 322 of the residual processor may derive the residual samples of the current block by inverse-transforming the transform coefficients. S640 may be performed by the adder 340 or the reconstructor of the decoding device.
[0125] The decoding device may derive the intra prediction mode / type of the current block based on the received prediction information (intra prediction mode / type information) (S600). The decoding device may derive the neighboring reference samples of the current block (S610). The decoding device generates prediction samples in the current block based on the intra prediction mode / type and the neighboring reference samples (S620). In this case, the decoding device may perform a prediction sample filtering process. The prediction sample filtering may be referred to as post-filtering. Some or all of the prediction samples may be filtered by the prediction sample filtering process. In some cases, the prediction sample filtering process may be omitted.
[0126] The decoding device generates residual samples of the current block based on the received residual information (S630). The decoding device may generate reconstructed samples of the current block based on the prediction samples and the residual samples, and derive a reconstructed block including the reconstructed samples (S640). A reconstructed picture of the current picture may be generated based on the reconstructed block. As described above, an in-loop filtering process or the like may also be applied to the reconstructed picture.
[0127] Here, the intra predictor 331 of the decoding device may include an intra prediction mode / type determiner 331-1, a reference sample deriver 331-2, and a prediction sample deriver 331-3. The intra prediction mode / type determiner 331-1 may determine the intra prediction mode / type of the current block based on the intra prediction mode / type information obtained by the entropy decoder 310. The reference sample deriver 331-2 may derive the neighboring reference samples of the current block, and the prediction sample deriver 331-3 may derive the prediction samples of the current block. In addition, although not shown, if the foregoing prediction sample filtering process is performed, the intra predictor 331 may also include a prediction sample filter (not shown).
[0128] The intra prediction mode information may include (for example) flag information (e.g., intra_luma_mpm_flag) indicating whether the most probable mode (MPM) is applied to the current block or whether the remaining modes are applied thereto. At this time, if the MPM is applied to the current block, the prediction mode information may further include index information (e.g., intra_luma_mpm_idx) indicating one of the intra prediction mode candidates (MPM candidates). The intra prediction mode candidates (MPM candidates) may be constituted by an MPM candidate list or an MPM list. In addition, if the MPM is not applied to the current block, the intra prediction mode information may further include remaining mode information (e.g., intra_luma_mpm_remainder) indicating one of the remaining intra prediction modes other than the intra prediction mode candidates (MPM candidates). The decoding device may determine the intra prediction mode of the current block based on the intra prediction mode information.
[0129] In addition, the intra prediction type information can be implemented in various forms. As an example, the intra prediction type information can include intra prediction type index information indicating one of the intra prediction types. As another example, the intra prediction type information can include at least one of the following: reference sample line information (e.g., intra_luma_ref_idx), which indicates whether the MRL is applied to the current block and which reference sample line is used if the MRL is applied; ISP flag information (e.g., intra_subpartitions_mode_flag), which indicates whether the ISP is applied to the current block; ISP type information (e.g., intra_subpartitions_split_flag), which indicates the split type of the subpartitions if the ISP is applied; flag information indicating whether PDCP is applied; or flag information indicating whether LIP is applied. In addition, the intra prediction type information can include an MIP flag indicating whether the MIP is applied to the current block.
[0130] The above intra prediction mode information and / or intra prediction type information can be encoded / decoded by the encoding methods described in this document. For example, the foregoing intra prediction mode information and / or intra prediction type information can be encoded / decoded by entropy coding based on truncated (Rice) binary codes (e.g., CABAC, CAVLC).
[0131] In addition, in the case of applying intra prediction, the intra prediction mode of adjacent blocks can be used to determine the intra prediction mode applied to the current block. For example, the decoding device can select one mpm candidate from the mpm candidates in the most probable mode (mpm) list derived from the intra prediction modes of adjacent blocks (e.g., the left adjacent block and / or the upper adjacent block) of the current block and additional candidate modes based on the received mpm index, or can select one of the remaining intra prediction modes not included in the mpm candidates (and the planar mode) based on the remaining intra prediction mode information. The mpm list can be constructed to include or not include the planar mode as a candidate. For example, if the mpm list includes the planar mode as a candidate, the mpm list can have 6 candidates, while if the mpm list does not include the planar mode as a candidate, the mpm list can have 5 candidates. If the mpm list does not include the planar mode as a candidate, a non-planar flag (e.g., intra_luma_not_planar_flag) indicating whether the intra prediction mode of the current block is the planar mode can be signaled. For example, the mpm flag can be signaled first, and when the value of the mpm flag is equal to 1, the mpm index and the non-planar flag can be signaled. In addition, when the value of the non-planar flag is equal to 1, the mpm index can be signaled. Here, constructing the mpm list to not include the planar mode as a candidate is to first identify whether the intra prediction mode is the planar mode by signaling the flag (non-planar flag) first, because the planar mode has always been regarded as an mpm, while the non-planar mode is not an mpm.
[0132] For example, based on an MPM flag (e.g., intra_luma_mpm_flag), it can be indicated whether the intra prediction mode applied to the current block is among the MPM candidates (and the planar mode) or among the remaining modes. An MPM flag value of 1 can indicate that the intra prediction mode of the current block is among the MPM candidates (and the planar mode), and an MPM flag value of 0 can indicate that the intra prediction mode of the current block is not among the MPM candidates (and the planar mode). A non-planar flag (e.g., intra_luma_not_planar_flag) value of 0 can indicate that the intra prediction mode of the current block is the planar mode, and a non-planar flag value of 1 can indicate that the intra prediction mode of the current block is not the planar mode. The MPM index can be signaled in the form of an mpm_idx or intra_luma_mpm_idx syntax element, and the remaining intra prediction mode information can be signaled in the form of a rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, the remaining intra prediction mode information can index the remaining intra prediction modes among all the intra prediction modes that are not included in the MPM candidates (and the planar mode) in the order of their number of prediction modes, and can indicate one of them. The intra prediction mode can be the intra prediction mode of the luminance component (samples). Hereinafter, the intra prediction mode information can include at least one of an MPM flag (e.g., intra_luma_mpm_flag), a non-planar flag (e.g., intra_luma_not_plane_flag), an MPM index (e.g., mpm_idx or intra_luma_mpm_idx), and the remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list can be referred to by various terms, such as an MPM candidate list, a candidate mode list (candModeList), and a candidate intra prediction mode list.
[0133] Generally, when dividing a block of an image, the current block to be encoded and adjacent blocks have similar image attributes. Therefore, the current block and adjacent blocks are more likely to have the same or similar intra prediction modes. Thus, the encoder can use the intra prediction mode of an adjacent block to encode the intra prediction mode of the current block. For example, the encoder / decoder can form a most probable mode (MPM) list of the current block. The MPM list can also be referred to as an MPM candidate list. Here, MPM can mean a mode used to improve the encoding efficiency by considering the similarity between the current block and adjacent blocks when encoding the intra prediction mode.
[0134] Figure 8An example of an MPM-based intra prediction method in an encoding device to which an exemplary embodiment of the present disclosure can be applied is illustrated.
[0135] Referring to Figure 8 , the encoding device constructs an MPM list for the current block (S800). The MPM list may include candidate intra prediction modes (MPM candidates) that are more likely to be applied to the current block. The MPM list may also include the intra prediction modes of adjacent blocks, and may also include specific intra prediction modes according to a predetermined method. A specific method for constructing the MPM list will be described later.
[0136] The encoding device determines the intra prediction mode of the current block (S810). The encoding device may perform prediction based on various intra prediction modes, and determine the optimal intra prediction mode based on rate distortion optimization (RDO) based on the above prediction. In this case, the encoding device may also use only the MPM candidates and the planar mode configured in the MPM list to determine the optimal intra prediction mode, or may also use the remaining intra prediction modes and the MPM candidates and the planar mode configured in the MPM list to determine the optimal intra prediction mode.
[0137] Specifically, for example, if the intra prediction type of the current block is a specific type different from the normal intra prediction type (e.g., LIP, MRL, or ISP), the encoding device may consider only the MPM candidates and the planar mode as candidates for the intra prediction mode of the current block to determine the optimal intra prediction mode. That is, in this case, the intra prediction mode of the current block may be determined only among the MPM candidates and the planar mode, and in this case, the mpm flag may not be encoded / signaled. In this case, even without separate signaling of the mpm flag, the decoding device may estimate the mpm flag as 1.
[0138] Generally, if the intra prediction mode of the current block is not the planar mode and is one of the MPM candidates in the MPM list, the encoding device generates an mpm index (mpm idx) indicating one of the MPM candidates. If the intra prediction mode of the current block does not even exist in the MPM list, the encoding device generates remaining intra prediction mode information indicating a mode among the remaining intra prediction modes not included in the MPM list (and the planar mode) as the intra prediction mode of the current block.
[0139] The encoding device may encode intra prediction mode information to output the intra prediction mode information in the form of a bitstream (S820). The intra prediction mode information may include the aforementioned mpm flag, non - planar flag, mpm index, and / or the remaining intra prediction mode information. Generally, the mpm index and the remaining intra prediction mode information have an alternative relationship and are not signaled simultaneously when indicating the intra prediction mode of a block. That is, the value 1 of the mpm flag and the non - planar flag or the mpm index are signaled together, or the value 0 of the mpm flag and the remaining intra prediction mode information are signaled together. However, as described above, if a specific intra prediction type is applied to the current block, the mpm flag is not signaled and only the non - planar flag and / or the mpm index may be signaled. That is, in this case, the intra prediction mode information may also include only the non - planar flag and / or the mpm index.
[0140] Figure 9 An example of an MPM - based intra prediction method in an encoding device to which exemplary embodiments of the present disclosure can be applied is illustrated. Figure 9 The illustrated decoding device may determine an intra prediction mode corresponding to the intra prediction mode information determined and signaled by the Figure 8 illustrated encoding device.
[0141] Referring to Figure 9 , the decoding device obtains the intra prediction mode information from the bitstream (S900). As described above, the intra prediction mode information may include at least one of an mpm flag, a non - planar flag, an mpm index, and the remaining intra prediction mode.
[0142] The decoding device constructs an MPM list (S910). The MPM list consists of the same MPM list constructed in the encoding device. That is, the MPM list may also include the intra prediction modes of adjacent blocks and may also include specific intra prediction modes according to a predetermined method. A specific method for constructing the MPM list will be described later.
[0143] Although S910 is shown as being executed later than S900, it is illustrative, and S910 may also be executed earlier than S900, and S900 and S910 may also be executed simultaneously.
[0144] The decoding device determines the intra prediction mode of the current block based on the MPM list and the intra prediction mode information (S920).
[0145] As an example, if the value of the mpm flag is 1, the decoding device may derive the planar mode as the intra prediction mode of the current block, or may derive the candidate indicated by the mpm index among the MPM candidates in the MPM list (based on the non-planar flag) as the intra prediction mode of the current block. Here, the MPM candidates may also indicate only the candidates included in the MPM list, or may also include the planar mode applicable to the case where the value of the mpm flag is 1 and the candidates included in the MPM list.
[0146] As another example, if the value of the mpm flag is 0, the decoding device may derive the intra prediction mode indicated by the remaining intra prediction mode information among the remaining intra prediction modes as the intra prediction mode of the current block, where the remaining intra prediction modes are not included in the MPM list and the planar mode.
[0147] As yet another example, if the intra prediction type of the current block is a specific type (e.g., LIP, MRL, or ISP), the decoding device may derive the candidate indicated by the mpm index in the MPM list or the planar mode as the intra prediction mode of the current block even without confirmation of the mpm flag.
[0148] When constructing the MPM list, in an embodiment, the encoding device / decoding device may derive the left mode as the candidate intra prediction mode of the left adjacent block of the current block, and may derive the up mode as the candidate intra prediction mode of the upper adjacent block of the current block. Here, the left adjacent block may represent the lowermost adjacent block among the left adjacent blocks located adjacent to the left of the current block, and the upper adjacent block may represent the rightmost adjacent block among the upper adjacent blocks located adjacent to the upper part of the current block. For example, if the size of the current block is W×H and the x component and y component of the upper left sample position of the current block are xN and yN, respectively, the left adjacent block may be a block including the samples at the coordinates (xN-1, yN+H-1), and the upper adjacent block may be a block including the samples at the coordinates (xN+W-1, yN-1).
[0149] For example, when the left adjacent block is available and intra prediction is applied to the left adjacent block, the encoding device / decoding device may derive the intra prediction mode of the left adjacent block as the left candidate intra prediction mode (i.e., the left mode). When the upper adjacent block is available, intra prediction is applied to the upper adjacent block, and the upper adjacent block is included in the current CTU, the decoding device may derive the intra prediction mode of the upper adjacent block as the upper candidate intra prediction mode (i.e., the up mode). Alternatively, when the left adjacent block is not available or intra prediction is not applied to the left adjacent block, the encoding device / decoding device may derive the planar mode as the left mode. When the upper adjacent block is not available or intra prediction is not applied to the upper adjacent block or the upper adjacent block is not included in the current CTU, the decoding device may derive the planar mode as the up mode.
[0150] The encoding device / decoding device may construct an MPM list by deriving candidate intra prediction modes of a current block based on a left mode derived from a left adjacent block and an upper mode derived from an upper adjacent block. In this case, the MPM list may include the left mode and the upper mode, and may further include specific intra prediction modes according to a predetermined method.
[0151] In addition, when constructing the MPM list, a normal intra prediction mode (normal intra prediction type) may be applied, or a specific intra prediction type (e.g., MRL or ISP) may be applied. This document proposes a method for constructing an MPM list when not only a normal intra prediction method (normal intra prediction type) but also a specific intra prediction type (e.g., MRL or ISP) is applied, and this will be described later.
[0152] In addition, the intra prediction modes may include non - directional (or non - angular) intra prediction modes and directional (or angular) intra prediction modes. For example, in the HEVC standard, intra prediction modes including 2 non - directional prediction modes and 33 directional prediction modes are used. The non - directional prediction modes may include a planar intra prediction mode (i.e., No. 0) and a DC intra prediction mode (i.e., No. 1). The directional prediction modes may include intra prediction modes from No. 2 to No. 34. The planar mode intra prediction mode may be referred to as the planar mode, and the DC intra prediction mode may be referred to as the DC mode.
[0153] Alternatively, in order to capture a given edge direction proposed in natural videos, the directional intra prediction modes may be extended from the existing 33 modes to 65 modes as in Figure 10 . In this case, the intra prediction modes may include 2 non - directional intra prediction modes and 65 directional intra prediction modes. The non - directional intra prediction modes may include a planar intra prediction mode (i.e., No. 0) and a DC intra prediction mode (i.e., No. 1). The directional intra prediction modes may include intra prediction modes from No. 2 to No. 66. The extended directional intra prediction modes may be applied to blocks of all sizes and may be applied to both the luminance component and the chrominance component. However, this is an example, and the embodiments of this document may be applied to cases where the number of intra prediction modes is different. 67 intra prediction modes may also be used according to the situation. The 67th intra prediction mode may indicate a linear model (LM) mode.
[0154] Figure 10 Examples of intra prediction modes to which the embodiments of this document may be applied are illustrated.
[0155] Referring to Figure 10 , the modes may be divided into intra prediction modes with horizontal directivity and intra prediction modes with vertical directivity based on the 34th intra prediction mode having a top - left diagonal prediction direction. InFigure 10 In this case, H and V respectively denote horizontal directionality and vertical directionality. Each of the numbers from -32 to 32 indicates a displacement of 1 / 32 unit at the sample grid position. The intra prediction modes within the 2nd to 33rd frames have horizontal directionality, and the intra prediction modes within the 34th to 66th frames have vertical directionality. The 18th intra prediction mode and the 50th intra prediction mode respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. The 2nd intra prediction mode can be referred to as a lower left diagonal intra prediction mode, the 34th intra prediction mode can be referred to as an upper left diagonal intra prediction mode, and the 66th intra prediction mode can be referred to as an upper right diagonal intra prediction mode.
[0156] The intra prediction mode for MIP can indicate a matrix and an offset for intra prediction, rather than the existing directional modes. That is, the matrix and the offset for intra prediction can be derived from the intra mode of MIP. In this case, when deriving the intra mode for general intra prediction or for generating the MPM list described above, the intra prediction mode of the block predicted by MIP can be set to a preset mode, for example, the planar mode or the DC mode. According to another example, the intra mode of MIP can be mapped to the planar mode, the DC mode, or the directional intra mode based on the block size.
[0157] Hereinafter, matrix-based intra prediction (MIP) as an intra prediction method is described.
[0158] As described above, matrix-based intra prediction (hereinafter, MIP) can be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). To predict the samples of a rectangular block having a width (W) and a height (H), MIP uses one H line among the reconstructed samples adjacent to the left boundary of the block and one W line among the reconstructed samples adjacent to the upper boundary of the block as input values. When no reconstructed samples are available, reference samples can be generated by an interpolation method applied to general intra prediction.
[0159] Figure 11 An MIP-based predicted sample generation process according to an example is illustrated. Refer to the following Figure 11 for a description of the MIP process.
[0160] 1. Averaging process
[0161] Among the boundary samples, four samples for the case of W = H = 4 and eight samples for any other case are extracted through the averaging process.
[0162] 2. Matrix-vector multiplication process
[0163] Perform matrix-vector multiplication using the average sample as input, and then add an offset. Through this operation, a reduced set of predicted samples for the downsampled set of samples in the original block can be derived.
[0164] 3. (Linear) interpolation process
[0165] Generate predicted samples at the remaining positions from the predicted samples of the downsampled set of samples by linear interpolation, where the linear interpolation is single-step linear interpolation in each direction.
[0166] The matrices and offset vectors required to generate the predicted block or predicted samples can be selected from three sets of matrices S0, S1, and S2.
[0167] Set S0 may include 16 matrices A0 i , i ∈ {0,..., 15}, and 16 offset vectors b0 i , i ∈ {0,..., 15}, and each matrix may include 16 rows and 4 columns. The matrices and offset vectors of set S0 can be used for 4×4 blocks. In another example, set S0 may include 18 matrices.
[0168] Set S1 may include eight matrices A1 i , i ∈ {0,..., 7}, and eight offset vectors b1 i , i ∈ {0,..., 7}, and each matrix may include 16 rows and 8 columns. In another example, set S1 may include six matrices. The matrices and offset vectors of set S1 can be used for 4×8, 8×4, and 8×8 blocks. Alternatively, the matrices and offset vectors of set S1 can be used for 4×H or W×4 blocks.
[0169] Finally, set S2 may include six matrices A2 i , i ∈ {0,..., 5}, and six offset vectors b2 i , i ∈ {0,..., 5}, and each matrix may include 64 rows and 8 columns. The matrices and offset vectors of set S2 or some of them can be used for any block with a different size that sets S0 and S1 do not apply to. For example, the matrices and offset vectors of set S2 can be used for operations on blocks with a height and width of 8 or greater.
[0170] The total number of multiplications required for the calculation of matrix-vector multiplication is always less than or equal to 4×W×H. That is, in the MIP mode, each sample requires at most four multiplications.
[0171] Below, the overall MIP process is briefly described. The remaining blocks not described below can be processed in any of the four cases described.
[0172] Figures 12 to 15 Illustrates the MIP process according to the size of the block, whereFigure 12 Illustrates the MIP process for a 4×4 block, Figure 13 Illustrates the MIP process for an 8×8 block, Figure 14 Illustrates the MIP process for an 8×4 block, and Figure 15 Illustrates the MIP process for a 16×16 block.
[0173] As Figure 12 shown, given a 4×4 block, MIP averages two samples for each axis of the boundary. As a result, four input samples are the input for matrix-vector multiplication, and the matrix is taken from set S0. An offset is added, thereby generating 16 final prediction samples. For a 4×4 block, no linear interpolation is required to generate prediction samples. Thus, a total of (4×16) / (4×4) = 4 multiplications can be performed per sample.
[0174] As Figure 13 shown, given an 8×8 block, MIP averages four samples for each axis of the boundary. As a result, eight input samples are the input for matrix-vector multiplication, and the matrix is taken from set S1. Sixteen samples are generated at odd positions through matrix-vector multiplication.
[0175] For an 8×8 block, each sample performs a total of (8×16) / (8×8) = 2 multiplications to generate prediction samples. After adding the offset, the reduced upper boundary samples are used to vertically interpolate the samples, and the original left boundary samples are used to horizontally interpolate the samples. In this case, since no multiplication operations are required during the interpolation process, a total of two multiplications per sample are required for MIP.
[0176] As Figure 14 shown, given an 8×4 block, MIP averages four samples for the horizontal axis of the boundary and uses four sample values on the left boundary for the vertical axis. As a result, eight input samples are the input for matrix-vector multiplication, and the matrix is taken from set S1. Sixteen samples are generated at odd horizontal positions and corresponding vertical positions through matrix-vector multiplication.
[0177] For an 8×4 block, each sample performs a total of (8×16) / (8×4) = 4 multiplications to generate prediction samples. After adding the offset, the original left boundary samples are used to horizontally interpolate the samples. In this case, since no multiplication operations are required during the interpolation process, a total of four multiplications per sample are required for MIP.
[0178] As Figure 15As shown, given a 16×16 block, MIP averages four samples for each axis. As a result, eight input samples are the input for the matrix-vector multiplication, and the matrix is taken from set S2. Through the matrix-vector multiplication, 64 samples are generated at odd positions. For a 16×16 block, a total of (8×64) / (16×16) = 2 multiplications are performed for each sample to generate the predicted samples. After adding the offset, the samples are vertically interpolated using eight reduced upper boundary samples, and horizontally interpolated using the original left boundary samples. In this case, since no multiplication operation is required during the interpolation process, a total of two multiplications per sample are required for MIP.
[0179] For larger blocks, the MIP process is basically the same as the above process, and it is easy to identify that the number of multiplications for each sample is less than 4.
[0180] For a W×8 block with width greater than 8 (W>8), since samples are generated at odd horizontal positions and each vertical position, only horizontal interpolation is required. In this case, (8×64) / (W×8) = 64 / W multiplications are performed for each sample for the prediction operation of the reduced samples. In the case of W = 16, no additional multiplication is required for linear interpolation, and in the case of W>16, the number of additional multiplications for each sample required for linear interpolation is less than 2. That is, the total number of multiplications for each sample is less than or equal to 4.
[0181] For a W×4 block with width greater than 4 (W>4), the matrix obtained by omitting all rows corresponding to odd entries along the horizontal axis of the downsampled block is defined as A k . Therefore, the output size is 32, and only horizontal interpolation is performed. For the prediction operation of the reduced samples, (8×32) / (W×4) = 64 / W multiplications are performed for each sample. In the case of W = 16, no additional multiplication is required, and in the case of W>16, the number of additional multiplications for each sample required for linear interpolation is less than 2. That is, the total number of multiplications for each sample is less than or equal to 4.
[0182] When the matrix is transposed, the processing can be performed accordingly.
[0183] Figure 16 Illustrates the boundary averaging process in the MIP process. Refer to Figure 16 Describe the averaging process in detail.
[0184] According to the averaging process, averaging is applied to each boundary (left boundary or upper boundary). The boundary indicates adjacent reference samples adjacent to the boundary of the current block, as Figure 16 shown. For example, the left boundary (bdry left ) indicates the left adjacent reference samples adjacent to the left boundary of the current block, and the upper boundary (bdrytop )Indicates the upper adjacent reference sample adjacent to the upper part.
[0185] When the current block is a 4×4 block, the size of each boundary can be reduced to two samples through an averaging process. When the current block is not a 4×4 block, the size of each boundary can be reduced to four samples through an averaging process.
[0186] The first step of the averaging process is to reduce the input boundaries (bdry left and bdry top ) to smaller boundaries ( and ). and Include two samples for 4×4 blocks and four samples for any other case.
[0187] For 4×4 blocks, where 0 ≤ i < 2, Can be represented by the following formula, and can be defined similarly
[0188] [Equation 1]
[0189]
[0190] When the width of the block is W = 4×2 k , where 0 ≤ i < 4, Can be represented by the following formula, and can be defined similarly
[0191] [Equation 2]
[0192]
[0193] Since the two reduced boundaries and Are concatenated into the reduced boundary vector bdry red , so bdry red Has a size of 4 in 4×4 blocks and a size of 8 in any other block.
[0194] When "mode" indicates the MIP mode, the reduced boundary vector bdry red And the range of the MIP mode can be defined based on the block size and the intra_mip_transposed_flag value according to the following formula.
[0195] [Equation 3]
[0196]
[0197] In the above formula, intra_mip_transposed_flag can be referred to as MIP transpose, and this flag information can indicate whether to transpose the reduced prediction samples. The semantics of this syntax element can be expressed as "intra_mip_transposed_flag[x0][y0] specifies whether the input vector of the matrix-based intra prediction mode for the luma samples is transposed."
[0198] Finally, for the interpolation of the downsampled prediction samples, a second version for averaging the boundaries for large blocks is required. That is, if the smaller value of the width and height is greater than 8 (min(W,H)>8) and the width is equal to or greater than the height (W≥H), then W = 8 * 2 l , and can be defined by the following formula, where 0 ≤ i < 8. If the smaller value of the width and height is greater than 8 (min(W,H)>8) and the height is greater than the width (H>W), it can be defined similarly
[0199] [Equation 4]
[0200]
[0201] Next, the process for generating the reduced prediction samples by matrix-vector multiplication is described.
[0202] Reduced input vector bdry red One of them generates the reduced prediction sample pred red . The prediction sample is for the signal of the downsampled block with width W red and height H red . Here, W red and H red are defined as follows.
[0203] [Equation 5]
[0204]
[0205] The reduced prediction sample pred red can be obtained by performing matrix-vector multiplication and then adding an offset, and can be derived by the following formula.
[0206] [Equation 6]
[0207] pred red = A · bdry red + b
[0208] Here, A is a matrix including W red × H redA matrix with a number of rows and four columns (where W and H are 4 (W = H = 4)) or eight columns (in any other case), and b is a W red ×H red vector.
[0209] The matrix A and the vector b can be selected from S0, S1, and S2 as follows, and the index idx = idx(W, H) can be defined by Equation 7 or Equation 8.
[0210] [Equation 7]
[0211]
[0212] [Equation 8]
[0213]
[0214] When idx is 1 or less (idx ≤ 1) or when idx is 2 and the smaller value of W and H is greater than 4 (min(W, H) > 4), A is set to and b is set to When idx is 2, the smaller value of W and H is 4 (min(W, H) = 4), and W is 4, A is the matrix generated by removing each row corresponding to the odd x - coordinates in the down - sampling block. Alternatively, when H is 4, A is the matrix generated by removing each column corresponding to the odd y - coordinates in the down - sampling block.
[0215] Finally, in Equation 9, the reduced prediction sample can be replaced by its transpose.
[0216] [Equation 9]
[0217] ·W = H = 4 and intra_mip_transposed_flag = 1
[0218] ·max(W, H) = 8 and intra_mip_transposed_flag = 1
[0219] ·max(W, H) > 8 and intra_mip_transposed_flag = 1
[0220] When W = H = 4, since A includes four columns and 16 rows, the number of multiplications required to calculate pred red is 4. In any other case, since A includes eight columns and W red ×H red rows, at most four multiplications per sample are required to calculate pred red .
[0221] Figure 17 This illustrates linear interpolation in the MIP process. Figure 17 The linear interpolation process is described as follows.
[0222] The interpolation process may be referred to as a linear interpolation process or a bilinear interpolation process. The interpolation process may include two steps, which are 1) vertical interpolation and 2) horizontal interpolation, as shown.
[0223] If W>=H, vertical linear interpolation may be applied first, followed by horizontal linear interpolation. <H,则可以首先应用水平线性插值,随后是垂直线性插值。在4×4块中,可以省略插值过程。
[0224] In a W×H block (where max(W,H)≥8), the prediction samples are from W red ×H red The reduced prediction sample pred red Derived. Depending on the block type, linear interpolation can be performed vertically, horizontally, or in both directions. When linear interpolation is applied in both directions, if W <H,则首先在水平方向上应用线性插值,否则首先在垂直方向上应用线性插值。
[0225] For W×H blocks (where max(W,H)≥8 and W>=H), it can be assumed that there is no generality loss. In this case, one-dimensional linear interpolation is performed as follows. When there is no generality loss, linear interpolation in the vertical direction is fully explained.
[0226] First, the downsampled prediction samples are extended upwards by the boundary signal. When the vertical upsampling coefficient U is defined ver =H / H red And set When , the extended reduced prediction samples can be set by the following formula.
[0227] [Equation 10]
[0228]
[0229] Subsequently, vertical linear interpolation prediction samples can be generated from the expanded downscaled prediction samples by the following equation.
[0230] [Equation 11]
[0231]
[0232] Here, x can be 0≤x<W red , y can be 0≤y<H red , and k can be 0≤k<U ver .
[0233] In the following, a method for maximizing the performance of the MIP technique while reducing its complexity is described. The embodiments to be described below can be performed independently or in combination.
[0234] In existing MIP modes, similar to existing intra prediction mode derivation methods, MPM flags are sent separately for non-MPM and MPM, and the MIP mode for the current block is encoded based on MPM or non-MPM.
[0235] According to an embodiment, for a block to which the MIP technique is applied, a structure for directly encoding the MIP mode without separating MPM and non-MPM can be proposed. This image coding structure enables the simplification of complex syntax structures. In addition, since the frequencies of the respective modes actually occurring in the MIP mode are relatively uniformly distributed, which is significantly different from the frequencies in existing intra modes, the proposed coding structure enables the maximization of efficiency when encoding and decoding MIP mode information.
[0236] The image information for MIP transmission and reception according to an embodiment is as follows. The following syntax can be included in the video / image information sent from the aforementioned encoding device to the decoding device, and can be configured / encoded in the encoding device to be signaled to the decoding device in the form of a bitstream, and the decoding device can parse / decode the included information (syntax elements) according to the conditions / orders disclosed in the syntax.
[0237] [Table 1]
[0238]
[0239] As shown in Table 1, the syntax intra_mip_flag and intra_mip_mode_idx for concurrently signaling the MIP mode of the current block can be included in the syntax information regarding the coding unit, and in the syntax intra_mip_flag and intra_mip_mode_idx for the current block.
[0240] A value of intra_mip_flag equal to 1 indicates that the intra prediction type of the luma samples is matrix-based intra prediction, and a value equal to 0 indicates that the intra prediction type of the luma samples is not matrix-based intra prediction.
[0241] An intra_mip_mode_idx equal to 1 indicates the matrix-based intra prediction mode of the luma samples. As described above, this matrix-based intra prediction mode can indicate a matrix for MIP or a matrix and an offset.
[0242] In addition, according to the example, flag information (e.g., intra_mip_transposed_flag) indicating whether the input vector for matrix-based intra prediction is transposed can also be signaled through the syntax of the coding unit. An intra_mip_transposed_flag equal to 1 indicates that the input vector is transposed, and the number of matrices for matrix-based intra prediction can be reduced by this flag information.
[0243] intra_mip_mode_idx can be coded and decoded by the truncated binary method as shown in the following table.
[0244] [Table 2]
[0245]
[0246] As shown in Table 2, intra_mip_flag is binarized with a fixed-length code, while intra_mip_mode_idx is binarized using the truncated binary method, and the maximum binarization length (cMax) can be set according to the size of the coding block. If the width and height of the coding block are 4 (cbWidth == 4 && cbHeight == 4), the maximum binarization length can be set to 34; otherwise, the maximum binarization length can be set to 18 or 10 according to whether the width and height of the coding block are 8 or less ((cbWidth <= 8 && cbHeight <= 8)?).
[0247] intra_mip_mode_idx can be coded by the bypass method instead of the context model-based method. By coding in the bypass method, the coding speed and efficiency can be increased.
[0248] According to another example, when intra_mip_mode_idx is binarized by the truncated binary method, the maximum binarization length can be as shown in the following table.
[0249] [Table 3]
[0250]
[0251] As shown in Table 3, if the width and height of the coding block are 4 ((cbWidth == 4 && cbHeight == 4)), the maximum binary length of intra_mip_mode_idx can be set to 15; otherwise, if the width or height of the coding block is 4 ((cbWidth == 4 || cbHeight == 4)), or if the width and height of the coding block are 8 (cbWidth == 8 && cbHeight == 8), the maximum binary length can be set to 7, and if the width or height of the coding block is 4 ((cbWidth == 4 || cbHeight == 4)) or if the width and height of the coding block are not 8 (cbWidth == 8 && cbHeight == 8), the maximum binary length can be set to 5.
[0252] According to another example, intra_mip_mode_idx can be encoded with a fixed - length code. In this case, to increase the coding efficiency, for each block size, the number of available MIP modes can be limited to a power of 2 (e.g., A = 2 K1 - 1, B = 2 K2 - 1, and C = 2 K3 - 1, where K1, K2, and K3 are positive integers).
[0253] The above description is shown in the following table.
[0254] [Table 4]
[0255]
[0256] In Table 4, when K1 = 5, K2 = 4, and K3 = 3, intra_mip_mode can be binary - valued as follows.
[0257] [Table 5]
[0258]
[0259] Alternatively, according to the example, when K1 is set to 4 and thus the width and height of the coding block are 4, the maximum binary length of the intra_mip_mode of the block can be set to 15. Additionally, when K2 is set to 3 and thus the width or height of the coding block is 4 or the width and height of the coding block are 8, the maximum binary length can be set to 7.
[0260] Embodiments can propose a method of using MIP only for specific blocks to which the MIP technology can be efficiently applied. When applying the method according to the embodiments, the number of matrix vectors required for MIP can be reduced, and the memory required to store the matrix vectors can be significantly reduced (by 50%). With these effects, the coding efficiency remains almost the same (less than 0.1%).
[0261] The following table illustrates the syntax including specific conditions for applying the MIP technology according to the embodiments.
[0262] [Table 6]
[0263]
[0264] As shown in Table 6, a condition of applying MIP only to large blocks (cbWidth > K1 || cbHeight > K2) can be added, and the size of the block can be determined based on preset values (K1 and K2). The reason for applying MIP only to large blocks is that the coding efficiency of MIP appears in relatively large blocks.
[0265] The following table illustrates an example in which K1 and K2 in Table 6 are predefined as 8.
[0266] [Table 7]
[0267]
[0268] The semantics of intra_mip_flag and intra_mip_mode_idx in Tables 6 and 7 are the same as those shown in Table 1.
[0269] When intra_mip_mode_idx has 11 possible modes, intra_mip_mode_idx can be encoded by truncated binarization (cMax = 10) as follows.
[0270] [Table 8]
[0271]
[0272] Alternatively, when the available MIP modes are limited to eight modes, intra_mip_mode_idx[x0][y0] can be encoded with a fixed-length code as follows.
[0273] [Table 9]
[0274]
[0275] In Tables 8 and 9, intra_mip_mode_idx can be encoded by a bypass method.
[0276] Embodiments may propose a MIP technique for applying a weighted matrix (A k ) and an offset vector (B k ) for a large block to small blocks, so as to efficiently apply the MIP technique in terms of memory saving. When applying the method according to the embodiments, the number of matrix vectors required for MIP can be reduced, and the memory required to store the matrix vectors can be significantly reduced (by 50%). With these effects, the coding efficiency remains almost the same (less than 0.1%).
[0277] Figure 18 Illustrated is a MIP technique according to an embodiment of the present disclosure.
[0278] As shown in the figure, Figure 18 (a) of shows the operation of the matrix and the offset vector for the large block index i, and Figure 18 (b) of shows the operation of the matrix applied to the sampling of the small block and the offset vector operation.
[0279] Referring to Figure 18 , the existing MIP process can be applied by considering the downsampled weighted matrix (Sub(A k )) obtained by downsampling the weighted matrix for the large block and the offset vector (Sub(b k )) obtained by downsampling the offset vector for the large block as the weighted matrix and the offset vector for the small blocks, respectively.
[0280] Here, downsampling may be applied only in one of the horizontal and vertical directions, or may be applied in both directions. Specifically, the downsampling factor (e.g., 1 in 2 or 1 in 4) and the vertical sampling direction or the horizontal sampling direction may be set based on the width and height of the corresponding block.
[0281] In addition, the number of intra prediction modes for MIP applying this embodiment may be set differently based on the size of the current block. For example, i) when both the height and width of the current block (coding block or transform block) are 4, 35 intra prediction modes (i.e., intra prediction modes 0 to 34) may be available, ii) when both the height and width of the current block are 8 or less than 8, 19 intra prediction modes (i.e., intra prediction modes 0 to 18) may be available, and iii) in other cases, 11 intra prediction modes (i.e., intra prediction modes 0 to 10) may be available.
[0282] For example, when the case where both the height and width of the current block are 4 is defined as block size type 0, the case where both the height and width of the current block are 8 or less is defined as block size type 1, and other cases are defined as block size type 2, the number of intra prediction modes of MIP may be as shown in the following table.
[0283] [Table 10]
[0284]
[0285] In order to apply the weighted matrix and offset vector for a large block (e.g., block size type = 2) to a small block (e.g., block size = 0 or block size = 1), the number of intra prediction modes available for each block size can be applied equally as shown in the following table.
[0286] [Table 11]
[0287]
[0288] Alternatively, as shown in Table 12 below, MIP can be applied only to block size types 1 and 2, and the weighted matrix and offset vector for block size type 2 can be downsampled for block size type 1. Thus, memory can be efficiently saved (50%).
[0289] [Table 12]
[0290]
[0291] The following drawings are provided to describe specific examples of the present disclosure. Since the specific names of the devices shown in the drawings or the names of specific signals / messages / fields are provided for illustration, the technical features of the present disclosure are not limited to the specific names used in the following drawings.
[0292] Figure 19 is a flowchart schematically illustrating a decoding method that can be executed by a decoding device according to an embodiment of the present disclosure.
[0293] Figure 19 The method shown can be performed by Figure 3 the decoding device 300 shown. Specifically, Figure 19 S1900 to S1940 of Figure 3 can be performed by the entropy decoder 310 and / or the predictor 330 (specifically, the intra predictor 331) shown, and Figure 19 S1950 of Figure 3 can be performed by the adder 340 shown. In addition, Figure 19 the method shown can be included in the foregoing embodiments of the present disclosure. Therefore, in Figure 19 the specific description of details overlapping with the foregoing embodiments will be omitted or briefly described.
[0294] Referring to Figure 19 , the decoding device can receive (i.e., obtain) from the bitstream flag information indicating whether matrix-based intra prediction (MIP) is used for the current block (S1900).
[0295] The flag information is a syntax such as intra_mip_flag, and can be included and signaled in the coding unit syntax information.
[0296] The decoding device can receive matrix-based intra prediction (MIP) mode information (S1910) based on the received flag information.
[0297] The MIP mode information can be represented as intra_mip_mode_idx, and can be signaled when intra_mip_flag is equal to 1. intra_mip_mode_idx can be index information indicating the MIP mode applied to the current block, and this index information can be used to derive the matrix when generating prediction samples.
[0298] According to an example, flag information indicating whether the input vector of matrix-based intra prediction is transposed can also be signaled through the coding unit syntax, for example, intra_mip_transposed_flag.
[0299] The decoding device can generate intra prediction samples for the current block based on the MIP information. The decoding device can derive at least one neighboring reference sample from the neighboring reference samples of the current block to generate intra prediction samples, and can generate prediction samples based on the neighboring reference samples.
[0300] When applying MIP, the decoding device can derive reduced boundary samples (S1920) by downsampling the reference samples adjacent to the current block.
[0301] The reduced boundary samples can be derived by downsampling the reference samples using averaging.
[0302] When the width and height of the current block are 4, four reduced boundary samples can be derived, and in other cases, eight reduced boundary samples can be derived.
[0303] The averaging process for downsampling can be applied to each boundary of the current block (e.g., the left boundary or the upper boundary), and can be applied to the neighboring reference samples adjacent to the boundary of the current block.
[0304] According to an example, when the current block is a 4×4 block, the size of each boundary can be reduced to two samples through the averaging process, and when the current block is not a 4×4 block, the size of each boundary can be reduced to four samples through the averaging process.
[0305] Subsequently, the decoding device can derive reduced prediction samples (S1930) based on the multiplication operation of the MIP matrix derived based on the size and index information of the current block and the reduced boundary samples.
[0306] The MIP matrix can be derived based on the size of the current block and the received index information.
[0307] The MIP matrix can be selected from any one of three matrix sets classified according to the size of the current block, and each of the three matrix sets can include multiple MIP matrices.
[0308] That is, three matrix sets of MIP can be set, and each matrix set can include multiple matrices and multiple offset vectors. These matrix sets can be applied separately according to the size of the current block.
[0309] For example, a matrix set including 18 or 16 matrices with 16 rows and four columns and 18 or 16 offset vectors can be applied to a 4×4 block. The index information can be information indicating any one of the multiple matrices included in one matrix set.
[0310] A matrix set including 10 or 8 matrices with 16 rows and eight columns and 10 or 8 offset vectors can be applied to 4×8, 8×4, and 8×8 blocks or 4×H or W×4 blocks.
[0311] In addition, a matrix set including six matrices with 64 rows and eight columns and six offset vectors can be applied to blocks other than the aforementioned blocks or blocks with a height and width of 8 or greater.
[0312] The reduced prediction samples (i.e., the prediction samples to which the MIP matrix has been applied) are derived based on the operation of adding an offset after the multiplication operation of the MIP matrix and the reduced boundary samples.
[0313] The decoding device can generate intra prediction samples of the current block by upsampling the reduced prediction samples (S1940).
[0314] The intra prediction samples can be upsampled by linear interpolation of the reduced prediction samples.
[0315] The interpolation process can be referred to as a linear interpolation or bilinear interpolation process, and can include two steps, which are 1) vertical interpolation and 2) horizontal interpolation.
[0316] If W >= H, vertical linear interpolation can be applied first, followed by horizontal linear interpolation. If W < H, horizontal linear interpolation can be applied first, followed by vertical linear interpolation. In a 4×4 block, the interpolation process can be omitted.
[0317] The decoding device can generate reconstructed samples of the current block based on the prediction samples (S1950).
[0318] In an embodiment, the decoding device may directly use the prediction sample as a reconstruction sample according to a prediction mode, or may generate a reconstruction sample by adding a residual sample to the prediction sample.
[0319] If there are residual samples of a current block, the decoding device may receive information about the residual of the current block. The information about the residual may include transform coefficients of the residual samples. The decoding device may derive the residual samples (or an array of residual samples) of the current block based on the residual information. The decoding device may generate a reconstruction sample based on the prediction sample and the residual sample, and may derive a reconstructed block or a reconstructed picture based on the reconstruction sample. Thereafter, as needed, in order to enhance the subjective / objective picture quality, the decoding device may apply a deblocking filter and / or an in-loop filtering process (such as the SAO process) to the reconstructed picture, as described above.
[0320] Figure 20 is a flowchart schematically illustrating an encoding method that may be performed by an encoding device according to an embodiment of the present disclosure.
[0321] Figure 20 The method shown may be performed by Figure 2 the encoding device 200 shown. Specifically, Figure 20 S2000 to S2030 of Figure 2 may be performed by the predictor 220 (specifically, the intra predictor 222) shown, Figure 20 S2040 of Figure 2 may be performed by the subtractor 231 shown, and Figure 20 S2050 of Figure 2 may be performed by the entropy encoder 240 shown. In addition, Figure 20 the method shown may be included in the foregoing embodiments of the present disclosure. Thus, in Figure 20 the details overlapping with the foregoing embodiments will be omitted or described briefly.
[0322] Referring to Figure 20 , the encoding device may derive whether to apply matrix-based intra prediction (MIP) to a current block (S2000).
[0323] The encoding device may apply various prediction techniques to find an optimal prediction mode for the current block, and may determine the optimal intra prediction mode based on rate distortion optimization (RDO).
[0324] When it is determined that MIP is applied to the current block, the encoding device may derive downsampled boundary samples by downsampling reference samples adjacent to the current block (S2010).
[0325] The downsampled boundary samples may be derived by downsampling the reference samples using an average.
[0326] When the width and height of the current block are 4, four reduced boundary samples can be derived, and in other cases, eight reduced boundary samples can be derived.
[0327] The averaging process for downsampling can be applied to each boundary of the current block (e.g., the left boundary or the upper boundary), and can be applied to adjacent reference samples adjacent to the boundary of the current block.
[0328] According to an example, when the current block is a 4×4 block, the size of each boundary can be reduced to two samples through the averaging process, and when the current block is not a 4×4 block, the size of each boundary can be reduced to four samples through the averaging process.
[0329] When deriving the reduced boundary samples, the encoding device can derive the reduced prediction samples (S2020) based on the multiplication operation of the MIP matrix selected based on the size of the current block and the reduced boundary samples.
[0330] The MIP matrix can be selected from any one of three matrix sets classified according to the size of the current block, and each of the three matrix sets can include multiple MIP matrices.
[0331] That is, three matrix sets of MIP can be set, and each matrix set can include multiple matrices and multiple offset vectors. These matrix sets can be applied separately according to the size of the current block.
[0332] For example, a matrix set including 18 or 16 matrices with 16 rows and four columns and 18 or 16 offset vectors can be applied to a 4×4 block. The index information can be information indicating any one of the multiple matrices included in one matrix set.
[0333] A matrix set including 10 or eight matrices with 16 rows and eight columns and 10 or eight offset vectors can be applied to 4×8, 8×4, and 8×8 blocks or 4×H or W×4 blocks.
[0334] In addition, a matrix set including six matrices with 64 rows and eight columns and six offset vectors can be applied to blocks other than the foregoing blocks or blocks with a height and width of 8 or greater.
[0335] The reduced prediction samples (i.e., the prediction samples to which the MIP matrix has been applied) are derived based on the operation of adding an offset after the multiplication operation of the MIP matrix and the reduced boundary samples.
[0336] The encoding device can generate the intra prediction samples of the current block by upsampling the reduced prediction samples (S2030).
[0337] The in - frame prediction samples can be upsampled by linear interpolation of the reduced prediction samples.
[0338] The interpolation process can be referred to as a linear interpolation or bilinear interpolation process, and can include two steps, which are 1) vertical interpolation and 2) horizontal interpolation.
[0339] If W >= H, vertical linear interpolation can be applied first, followed by horizontal linear interpolation. If W < H, horizontal linear interpolation can be applied first, followed by vertical linear interpolation. In a 4×4 block, the interpolation process can be omitted.
[0340] The encoding device can derive the residual samples of the current block based on the prediction samples of the current block and the original samples of the current block (S2040).
[0341] The encoding device can generate residual information for the current block based on the residual samples, and can encode the residual information and the picture information including flag information indicating whether MIP is applied and MIP mode information (S2050).
[0342] Here, the residual information can include quantization parameters, value information, position information, transform scheme, and transform kernel related to the quantized transform coefficients derived by transforming and quantizing the residual samples.
[0343] The flag information indicating whether MIP is applied is a syntax such as intra_mip_flag, and can be included and encoded in the coded unit syntax information.
[0344] The MIP mode information can be represented as intra_mip_mode_idx, and can be encoded when intra_mip_flag is equal to 1. intra_mip_mode_idx can be index information indicating the MIP mode applied to the current block, and this index information can be used to derive the matrix when generating prediction samples. The index information can indicate any one of multiple matrices included in a matrix set.
[0345] According to an example, flag information indicating whether the input vector of matrix - based in - frame prediction is transposed, such as intra_mip_transposed_flag, can also be signaled through the coded unit syntax.
[0346] That is, the encoding device can encode the picture information including the MIP mode information and / or the residual information of the current block, and can output the picture information as a bitstream.
[0347] The bitstream can be sent to a decoding device via a network or a (digital) storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0348] The foregoing process of generating prediction samples for the current block can be performed by Figure 2 the intra predictor 222 of the encoding device 200 shown, the process of deriving residual samples can be performed by Figure 2 the subtractor 231 of the encoding device 200 shown, and the process of generating and encoding residual information can be performed by Figure 2 the residual processor 230 and the entropy encoder 240 of the encoding device 200 shown.
[0349] In the above embodiments, the method is explained based on a flowchart by means of a series of steps or blocks. However, the present disclosure is not limited to the order of the steps, and a certain step can be performed in an order or steps different from the above order or steps, or a certain step can be performed concurrently with other steps. In addition, those of ordinary skill in the art can understand that the steps shown in the flowchart are not exclusive, and one or more steps in the flowchart can be incorporated or deleted without affecting the scope of the present disclosure.
[0350] The above method according to the present disclosure can be implemented in software form, and the encoding device and / or decoding device according to the present disclosure can be included in devices for image processing such as a television, a computer, a smart phone, a set-top box, and a display device.
[0351] When the embodiments in the present disclosure are implemented by software, the above method can be implemented as modules (steps, functions, etc.) for performing the above functions. These modules can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor in various well-known ways. The processor can include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory can include a read only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present disclosure can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, each functional unit shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, the information for implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.
[0352] In addition, the decoding device and the encoding device to which this document is applied may be included in a multimedia broadcast transmission and reception device, a mobile communication terminal, a home movie video device, a digital movie video device, a camera for surveillance, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VOD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a video phone device, a transportation device terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an aircraft terminal, and a ship terminal), and a medical video device, and may be used to process video signals or data signals. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smart phone, a tablet PC, and a digital video recorder (DVR).
[0353] In addition, the processing method applied in this document may be generated in the form of a program executed by a computer and may be stored in a computer-readable recording medium. Multimedia data having a data structure according to this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, the bitstream generated using the encoding method may be stored in a computer-readable recording medium or may be transmitted through wired and wireless communication networks.
[0354] In addition, an embodiment of this document may be implemented as a computer program product using program code. The program code may be executed by a computer according to an embodiment of this document. The program code may be stored on a computer-readable carrier wave.
[0355] Figure 21 Examples of a content streaming system to which the embodiments disclosed in this document may be applied are illustrated.
[0356] Refer to Figure 21 , a content streaming system to which an embodiment of this document is applied may basically include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0357] The encoding server compresses the content input from a multimedia input device such as a smart phone, a camera, a portable video camera, etc. into digital data to generate a bitstream, and sends the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, a camera, a portable video camera, etc. directly generates a bitstream, the encoding server can be omitted.
[0358] The bitstream can be generated by applying the encoding method or the bitstream generation method of the embodiments of this document, and the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0359] The streaming server sends the multimedia data to the user device through the web server based on the user's request, and the web server serves as a medium for notifying the user of the service. When the user requests a desired service from the web server, the web server passes it to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between the devices in the content streaming system.
[0360] The streaming server can receive the content from the media storage device and / or the encoding server. For example, when receiving the content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined time.
[0361] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a slate PC, a tablet PC, a ultrabook, a wearable device (e.g., a smart watch, smart glasses, a head-mounted display), a digital TV, a desktop computer, a digital signage, etc.
[0362] Each server in the content streaming system can be operated as a distributed server, in which case, the data received from each server can be distributed.
[0363] The claims in this specification can be combined in various ways. For example, the technical features in the method claims of this specification can be combined to be implemented or executed in a device, and the technical features in the device claims can be combined to be implemented or executed in a method. In addition, the technical features in the method claims and the device claims can be combined to be implemented or executed in a device. In addition, the technical features in the method claims and the device claims can be combined to be implemented or executed in a method.
Claims
1. A decoding device for image decoding, the decoding device comprising: A memory; And At least one processor, the at least one processor being connected to the memory and configured to: Receive flag information related to whether matrix-based intra prediction (MIP) is used for a current block; Receive matrix-based intra prediction (MIP) mode information based on the flag information; Generate intra prediction samples for the current block based on the MIP mode information; And Generate reconstructed samples for the current block based on the intra prediction samples, Wherein the MIP mode information is index information related to an MIP matrix applied to the current block, and Wherein the MIP matrix is derived based on the size of the current block and the index information.
2. An encoding device for image encoding, the encoding device comprising: A memory; And At least one processor, the at least one processor being connected to the memory and configured to: Derive whether matrix-based intra prediction (MIP) is to be applied to a current block; Based on the MIP being applied to the current block, derive intra prediction samples for the current block based on the MIP; Derive residual samples for the current block based on the intra prediction samples; And Encode information about the residual samples and information about the MIP, Wherein the information about the MIP includes flag information related to whether the MIP is applied to the current block and matrix-based intra prediction (MIP) mode information, Wherein the MIP mode information is index information related to an MIP matrix applied to the current block, and Wherein the index information indicates any one of a plurality of MIP matrices included in a matrix set, Wherein the MIP matrix is determined based on the size of the current block and the index information.
3. A device for transmitting data for an image, the device comprising: At least one processor configured to obtain a bitstream, wherein the bitstream is generated by: deriving whether matrix-based intra prediction (MIP) is to be applied to a current block; based on the MIP being applied to the current block, deriving intra prediction samples for the current block based on the MIP; deriving residual samples for the current block based on the intra prediction samples; and encoding information about the residual samples and information about the MIP to generate the bitstream; and A transmitter configured to transmit the bitstream, Wherein the information about the MIP includes flag information related to whether the MIP is applied to the current block and matrix-based intra prediction (MIP) mode information, Wherein the MIP mode information is index information related to an MIP matrix applied to the current block, and Wherein the index information indicates any one of a plurality of MIP matrices included in a matrix set, Among them, the MIP matrix is determined based on the size of the current block and the index information.