Matrix-based intra prediction apparatus and method

By adopting a matrix-based intra prediction (MIP) method in image/video encoding, adaptive selection and flexible configuration of MIP mode, the problem of low compression and transmission efficiency of high-resolution image/video data is solved, and more efficient coding and better visual quality is achieved.

CN120166221APending Publication Date: 2025-06-17LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510288138.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-08-22
Filing Date
2020-08-24
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress and transmit high-resolution, high-quality image/video data, especially in the case of increasing demand for immersive media such as virtual reality (VR), artificial reality (AR), and holograms.

Method used

The matrix-based intra prediction (MIP) method is adopted to adaptively select whether to apply MIP, and flexibly select the number of MIP modes according to the type of the current block to improve image/video encoding efficiency.

Benefits of technology

Improves image/video compression efficiency, enhances subjective/objective visual quality, and simplifies signaling, with the advantages of hardware implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120166221A_ABST
    Figure CN120166221A_ABST
Patent Text Reader

Abstract

The invention relates to a matrix-based intra prediction device and method. According to one embodiment of the present document, a method for performing efficient intra prediction can be presented. In one embodiment, the MIP process can be performed regardless of the requirements related to the form of the current block to be compiled, and thus signaling is simplified and hardware implementation can be advantageous.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with the application number 202080073207.5 (PCT / KR2020 / 011231), international application date of August 24, 2020, and invention title of "Matrix-based Intra Prediction Apparatus and Method", which was filed with the Chinese Patent Office on April 19, 2022. Technical Field

[0002] This document relates to a matrix-based intra prediction apparatus and method. Background Art

[0003] Recently, there has been an increasing demand for high-resolution and high-quality images / videos such as 4K, 8K, or higher ultra-high definition (UHD) images / videos in various fields. As image / video data has high resolution and quality, the amount of information or bits to be transmitted increases compared to existing image / video data. Therefore, transmitting image data using media such as existing wired / wireless broadband lines or existing storage media or storing image / video data using existing storage media increases transmission costs and storage costs.

[0004] In addition, there has been a growing interest in and demand for virtual reality (VR) and artificial reality (AR) content, as well as immersive media such as holograms, and the broadcasting of images / videos with characteristics different from real images (such as game images) has also increased.

[0005] Therefore, highly efficient image / video compression techniques are needed to effectively compress, transmit, store, or reproduce the information of high-resolution and high-quality images / videos with various characteristics as described above.

[0006] In addition, to improve compression efficiency and enhance subjective / objective visual quality, a matrix-based intra prediction (MIP) process is performed, and schemes for enhancing the data transmission efficiency for the MIP process are discussed. Summary of the Invention

[0007] Technical Solution

[0008] According to an embodiment of this document, methods and apparatuses for improving image / video compilation efficiency are provided.

[0009] According to an embodiment of this document, efficient intra prediction methods and apparatuses are provided.

[0010] According to an embodiment of this document, efficient MIP application methods and apparatuses are provided.

[0011] According to an embodiment of this document, it is possible to adaptively select whether to apply MIP according to the type of the current block.

[0012] According to an embodiment of this document, MIP processing can be performed without restricting conditions related to the type of the current block.

[0013] According to an embodiment of this document, the number of MIP modes can be adaptively selected according to the type of the current block.

[0014] According to an embodiment of this document, MIP processing may not be performed on blocks with relatively low MIP prediction efficiency.

[0015] According to an embodiment of this document, a smaller number of MIP modes can be used for blocks with relatively low MIP prediction efficiency.

[0016] According to an embodiment of this document, a video / image decoding method executed by a decoding device is provided.

[0017] According to an embodiment of this document, a decoding device for performing video / image decoding is provided.

[0018] According to an embodiment of this document, a video / image encoding method executed by an encoding device is provided.

[0019] According to an embodiment of this document, an encoding device for performing video / image encoding is provided.

[0020] According to an embodiment of this document, a computer-readable digital storage medium is provided, in which encoded video / image information generated according to the video / image encoding method disclosed in at least one embodiment of the embodiments of this document is stored.

[0021] According to an embodiment of this document, a computer-readable digital storage medium is provided, in which encoded information or encoded video / image information that causes a decoding device to execute the video / image decoding method disclosed in at least one embodiment of the embodiments of this document is stored.

[0022] Beneficial effects

[0023] According to an embodiment of this document, the overall image / video compression efficiency can be improved.

[0024] According to an embodiment of this document, subjective / objective visual quality can be enhanced through effective intra prediction.

[0025] According to an embodiment of this document, MIP processing for image / video compilation can be effectively performed.

[0026] According to an embodiment of this document, MIP processing may not be performed on blocks with relatively low MIP prediction efficiency, thereby improving MIP prediction efficiency.

[0027] According to an embodiment of this document, since a smaller number of MIP modes are used for blocks with relatively low MIP prediction efficiency, the amount of transmitted data for MIP modes can be reduced, and the MIP prediction efficiency can be improved.

[0028] According to an embodiment of this document, since MIP processing is performed without restricting conditions related to the type of the current block, signaling can be simplified, and there are advantages in terms of hardware implementation. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 Schematically illustrates an example of a video / image compilation system to which embodiments of the present disclosure can be applied.

[0030] Figure 2 Is a diagram schematically explaining the configuration of a video / image encoding device to which embodiments of the present disclosure can be applied.

[0031] Figure 3 Is a diagram schematically explaining the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied.

[0032] Figure 4 Schematically illustrates an example of an intra-frame prediction-based image encoding method to which embodiments of the present disclosure are applicable.

[0033] Figure 5 Schematically illustrates an intra-frame predictor in an encoding device.

[0034] Figure 6 Illustrates an example of an image decoding method based on schematic intra-frame prediction to which embodiments of this document are applicable.

[0035] Figure 7 Schematically illustrates an intra-frame predictor in a decoding device.

[0036] Figure 8 Illustrates an example of an intra-frame prediction mode to which embodiments of this document are applicable.

[0037] Figure 9 Illustrates a process for generating prediction samples based on MIP according to an embodiment.

[0038] Figure 10 、 11 And 12 are flowcharts illustrating MIP processing according to an embodiment of this document.

[0039] Figure 13 And Figure 14 Schematically illustrates an example of a video / image encoding method and related components according to an embodiment of this document.

[0040] Figure 15 And Figure 16Schematically illustrate an example of a video / image decoding method and related components according to an embodiment of this document.

[0041] Figure 17 Illustrate an example of a content streaming system to which the embodiments disclosed in this document are applicable. Detailed implementation

[0042] This document can be modified in various ways and can have various embodiments, and specific embodiments will be illustrated and described in detail in the accompanying drawings. However, this is not intended to limit this document to specific embodiments. Commonly used terms in this specification are used to describe specific embodiments and do not limit the technical spirit of this document. Singular expressions include plural expressions unless clearly stated otherwise in the context. Terms such as "including" or "having" in this specification should be understood to indicate the presence of the features, quantities, steps, operations, elements, parts, or combinations thereof described in the specification, rather than excluding the presence or possibility of adding one or more other features, quantities, steps, operations, elements, parts, or combinations thereof.

[0043] At the same time, for the convenience of describing descriptions related to different feature functions, elements in the accompanying drawings described in this document are illustrated independently. This does not mean that each element is implemented as a separate hardware or separate software. For example, at least two elements can be combined to form a single element, or a single element can be divided into multiple elements. Embodiments in which elements are combined and / or separated are also included within the scope of the rights of this document unless it deviates from the essence of this document.

[0044] Hereinafter, the preferred embodiments of this document will be described more specifically with reference to the accompanying drawings. Hereinafter, in the accompanying drawings, the same reference numerals are used for the same elements, and repeated descriptions of the same elements are omitted.

[0045] This document relates to video / image compilation. For example, the methods / embodiments disclosed in this document may be related to the common video compilation (VVC) standard (ITU-T Rec. H.266), the next-generation video / image compilation standard after VVC, or other video compilation-related standards (such as the high efficiency video coding (HEVC) standard (ITU-T Rec. H.265), the essential video coding (EVC) standard, and the AVS2 standard).

[0046] This document presents various embodiments of video / image compilation, and the embodiments can be executed in combination with each other unless otherwise stated.

[0047] In this document, video may refer to a collection of a series of images over time. A picture generally refers to a unit representing an image in a specific time region, and a slice / tile is a unit that forms part of a picture in compilation. A slice / tile may include one or more Compilation Tree Units (CTUs). A picture may consist of one or more slices / tiles. A picture may consist of one or more tile groups. A tile group may include one or more tiles.

[0048] A pixel or pel may refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" may be used as a term corresponding to a pixel. A sample generally may represent a pixel or the value of a pixel, and may represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Further, a sample may refer to a pixel value in the spatial domain, or in the case where such a pixel value is transformed to the frequency domain, may refer to a transform coefficient in the frequency domain.

[0049] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. A unit may include one luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the terms unit and terms such as block, region, etc. may be used interchangeably. Generally, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0050] In this document, the terms " / " and "," are interpreted as indicating "and / or". For example, the expression "A / B" is interpreted as indicating "A and / or B", and "A, B" is interpreted as indicating "A and / or B". Additionally, "A / B / C" may mean "at least one of A, B, and / or C". Further, "A, B, C" may mean "at least one of A, B, and / or C".

[0051] Furthermore, in this document, the term "or" shall be interpreted as indicating "and / or". For example, the expression "A or B" may mean 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document may mean "additionally or alternatively".

[0052] In this specification, "at least one of A and B" may mean "only A", "only B", or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted the same as "at least one of A and B".

[0053] In addition, in this specification, "at least one of A, B, and C" may mean "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".

[0054] In addition, parentheses used in this specification may mean "for example". Specifically, in the case of expressing "prediction (intra prediction)", it may indicate that "intra prediction" is presented as an example of "prediction". In other words, the term "prediction" in this specification is not limited to "intra prediction", and it may indicate that "intra prediction" is presented as an example of "prediction". In addition, even in the case of expressing "prediction (i.e., intra prediction)", it may indicate that "intra prediction" is presented as an example of "prediction".

[0055] In this specification, technical features separately explained in one drawing may be implemented separately or may be implemented simultaneously.

[0056] Figure 1 Schematically illustrate an example of a video / image compilation system to which the embodiments of this document can be applied.

[0057] Reference Figure 1 , the video / image compilation system may include a source device and a receiving device. The source device may deliver encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.

[0058] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0059] The video source may obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. For example, the video / image capture device may include one or more cameras, a video / image archive including previously captured video / images, etc. For example, the video / image generation device may include a computer, a tablet computer, and a smart phone, and may generate video / images (electronically). For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capture process may be replaced by a process of generating relevant data.

[0060] The encoding device can encode the input video / image. For compression and compilation efficiency, the encoding device can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0061] The transmitter can send the encoded image / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating a media file in a predetermined file format and can include components for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to the decoding device.

[0062] The decoding device can decode the video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.

[0063] The renderer can render the decoded video / image. The rendered video / image can be displayed on a display.

[0064] Figure 2 is a diagram schematically explaining the configuration of the video / image encoding device to which this document is applicable. Hereinafter, the video encoding device can include an image encoding device.

[0065] Reference Figure 2 , the encoding device 200 includes an image splitter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can also include a subtractor 231. The adder 250 can be referred to as a reconstructor or a reconstruction block generator. According to an embodiment, the image splitter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 can be configured by at least one hardware component (e.g., an encoder chipset or a processor). Additionally, the memory 270 can include a decoded picture buffer (DPB), or can be configured by a digital storage medium. The hardware component can also include the memory 270 as an internal / external component.

[0066] The image splitter 210 may split an input image (or picture or frame) input to the encoding device 200 into one or more processors. For example, a processor may be referred to as a coding unit (CU). In this case, the coding units may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, a coding unit may be split into multiple coding units at a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure may be applied first, and / or the binary tree structure and / or the ternary tree structure may be applied later. Alternatively, the binary tree structure may also be applied first. The coding process according to the present disclosure may be performed based on the final coding units that are no longer split. In this case, based on the coding efficiency according to the image features, the largest coding unit may be used as the final coding unit, or when necessary, the coding units may be recursively split into coding units at a deeper depth, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction described later. As another example, the processor may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be partitioned or split from the aforementioned final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.

[0067] In some cases, a unit may be used interchangeably with terms such as a block or a region. In general, an M×N block may represent a set of samples or transformation coefficients composed of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. A sample may be used as a term corresponding to a pixel or a cell of a picture (or image).

[0068] The subtractor 231 may generate a residual signal (residual block, residual sample, or residual sample array) by subtracting the prediction signal (prediction block, prediction sample, or prediction sample array) output from the predictor 220 from the input image signal (original block, original sample, or original sample array), and may send the generated residual signal to the transformer 232. The predictor 220 may perform prediction on a processing target block (hereinafter referred to as “current block”) and may generate a prediction block including prediction samples for the current block. The predictor 220 may determine whether to apply intra prediction or inter prediction in units of the current block or CU. The predictor may generate and transmit various information about the prediction, such as prediction mode information described later in the explanation of each prediction mode, to the entropy encoder 240. The information about the prediction may be encoded by the entropy encoder 240 and output in the form of a bitstream.

[0069] The intra predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the samples referred to may be located near the current block or may be separated. In intra prediction, the prediction mode may include a plurality of non - directional modes and a plurality of directional modes. For example, the non - directional modes may include the DC mode and the planar mode. For example, depending on the level of detail of the prediction direction, the directional modes may include 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes may be used according to the settings. The intra predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to neighboring blocks.

[0070] The inter predictor 221 can derive the prediction block of the current block based on the reference block (reference sample array) specified by the motion vector on the reference picture. Here, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the skip mode and the merge mode, the inter predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, the residual signal may not be sent. In the case of the motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0071] Predictor 220 may generate a prediction signal based on various prediction methods described below. For example, the predictor may not only apply intra prediction or inter prediction to predict a block, but may also apply intra prediction and inter prediction simultaneously. This may be referred to as combined inter and intra prediction (CIIP). In addition, the predictor may perform intra block copy (IBC) for prediction of a block. Intra block copy may be used for content image / video compilation such as games, for example, screen content compilation (SCC). IBC basically performs prediction in the current picture, but may be performed similarly to inter prediction in terms of deriving a reference block in the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document.

[0072] The prediction signal generated by the inter predictor 221 and / or the intra predictor 222 may be used to generate a reconstruction signal or generate a residual signal. The transformer 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT means a transform obtained from a graph when relationship information between pixels is represented by the graph. CNT means a transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to a square pixel block having the same size, or may be applied to a block having a variable size rather than a square.

[0073] Quantizer 233 may quantize the transform coefficients and send them to entropy encoder 240, and entropy encoder 240 may encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients may be referred to as residual information. Quantizer 233 may rearrange the block-based quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Entropy encoder 240 may perform various encoding methods, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. Entropy encoder 240 may encode, together or separately, information required for video / image reconstruction in addition to the quantized transform coefficients (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) can be sent or stored in the form of a bitstream in units of NAL (network abstraction layer). The video / image information may further include information about various parameter sets, such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). In addition, the video / image information may further include general constraint information. In this document, the information and / or syntax elements signaled / sent later in this document may be encoded through the above encoding process and may be included in the bitstream. The bitstream may be sent through a network or may be stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that sends the signal output from entropy encoder 240 and / or a storage unit (not shown) that stores the signal may be included as an internal / external element of encoding device 200, and alternatively, the transmitter may be included in entropy encoder 240.

[0074] The quantized transform coefficients output from quantizer 233 may be used to generate a prediction signal. For example, the quantized transform coefficients may be dequantized and inverse-transformed by dequantizer 234 and inverse-transformer 235 to reconstruct a residual signal (residual block or residual sample). Adder 250 adds the reconstructed residual signal to the prediction signal output from predictor 220 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array). If the block to be processed has no residual, such as in the case of applying the skip mode, the predicted block may be used as the reconstructed block. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, and may be used for inter prediction of the next picture through filtering as described below.

[0075] Meanwhile, luminance mapping and chrominance scaling (LMCS) can be applied during picture encoding and / or reconstruction.

[0076] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in memory 270 (specifically, the DPB of memory 270). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. Filter 260 can generate various information related to the filtering and send the generated information to entropy encoder 240, as described later in the description of each filtering method. The information related to the filtering can be encoded by entropy encoder 240 and output in the form of a bitstream.

[0077] The modified reconstructed picture sent to memory 270 can be used as a reference picture in inter-frame predictor 221. When inter-frame prediction is applied by the encoding device, the prediction mismatch between encoding device 200 and decoding device 300 can be avoided, and the encoding efficiency can be improved.

[0078] The DPB of memory 270 can store the modified reconstructed picture to be used as a reference picture in inter-frame predictor 221. Memory 270 can store the motion information of the blocks from which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks in the already reconstructed pictures. The stored motion information can be sent to inter-frame predictor 221 and used as the motion information of spatially adjacent blocks or temporally adjacent blocks. Memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and can pass the reconstructed samples to intra-frame predictor 222.

[0079] Figure 3 FIG. is a diagram schematically explaining the configuration of a video / image decoding device to which this document is applicable.

[0080] Reference Figure 3 FIG., decoding device 300 can include entropy decoder 310, residual processor 320, predictor 330, adder 340, filter 350, and memory 360. Predictor 330 can include inter-frame predictor 331 and intra-frame predictor 332. Residual processor 320 can include dequantizer 321 and inverse transformer 321. According to an embodiment, entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 can be configured by hardware components (e.g., a decoder chipset or a processor). Additionally, memory 360 can include a decoded picture buffer (DPB) or can be configured by a digital storage medium. The hardware components can also include memory 360 as an internal / external component.

[0081] When the input includes a bitstream of video / image information, the decoding device 300 may reconstruct an image corresponding to the processing of the video / image information in the Figure 2 encoding device. For example, the decoding device 300 may derive units / blocks based on the block partitioning-related information obtained from the bitstream. The decoding device 300 may use the processor applied in the encoding device to perform decoding. Thus, the decoding processor may be, for example, a compilation unit, and the compilation unit may be partitioned from the compilation tree unit or the largest compilation unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the compilation unit. The reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproduction device.

[0082] The decoding device 300 may receive, in the form of a bitstream, from Figure 2The signal output by the encoding device, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can further include information about various parameter sets, such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information can further include general constraint information. The decoding device can further decode the picture based on the information about the parameter set and / or the general constraint information. The information and / or syntax elements signaled / received described later in this document can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, as well as the output values of the syntax elements required for image reconstruction and the quantization values of the transform coefficients for the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, determine the context model using the decoding target syntax element information, the neighborhood, and the decoding information of the decoding target block, or the information of the symbols / bins decoded in the previous stage, and perform arithmetic decoding on the bins by predicting the probability of the bin occurrence according to the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin after determining the context model. The information about prediction among the information decoded by the entropy decoder 310 can be provided to the predictor 330, and the information about the residuals for which entropy decoding has been performed in the entropy decoder 310, that is, the quantized transform coefficients and the related parameter information, can be input to the dequantizer 321. In addition, the information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. Meanwhile, the receiver (not shown) for receiving the signal output by the encoding device can be further configured as an internal / external component of the decoding device 300, or the receiver can be a component of the entropy decoder 310. Meanwhile, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of the dequantizer 321, the inverse transformer 322, the predictor 330, the adder 340, the filter 350, and the memory 360.

[0083] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order executed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients by using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.

[0084] The inverse transformer 322 performs an inverse transformation on the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0085] The predictor can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310 and can determine a specific intra / inter prediction mode.

[0086] The predictor can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can perform intra block copy (IBC) for prediction of a block. Intra block copy can be used for content image / video compilation such as games, for example, screen content compilation (SCC). IBC basically performs prediction in the current picture, but can be performed similar to inter prediction because a reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document.

[0087] The intra predictor 331 can predict the current block by referring to samples in the current picture. The samples referred to can be located near the current block or can be located at separate positions according to the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to neighboring blocks.

[0088] The inter - frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks can include spatially neighboring blocks present in the current picture and temporally neighboring blocks present in the reference picture. For example, the inter - frame predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter - frame prediction can be performed based on various prediction modes, and information about the prediction can include information indicating the mode of inter - frame prediction for the current block.

[0089] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (330). If there is no residual for the block to be processed, such as when the skip mode is applied, the prediction block can be used as the reconstructed block.

[0090] The adder 340 can be referred to as a reconstructor or a reconstructed - block generator. The generated reconstructed signal can be used for intra - frame prediction of the next block to be processed in the current picture, can be output through filtering as described below, or can be used for inter - frame prediction of the next picture.

[0091] Meanwhile, luminance mapping and chrominance scaling (LMCS) can be applied during picture decoding.

[0092] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 360 (specifically, the DPB of the memory 360). Various filtering methods can include, for example, de - blocking filtering, sample - adaptive offset, adaptive loop filter, bilateral filter, etc.

[0093] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the blocks from which the motion information in the current picture is derived (or decoded) and / or the motion information of the blocks in the already reconstructed pictures. The stored motion information can be sent to the inter-frame predictor 260 and used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and pass the reconstructed samples to the intra-frame predictor 331.

[0094] In this specification, the embodiments explained in the predictor 330, dequantizer 321, inverse transformer 322, and filter 350 of the decoding device 300 can be applied to or correspond to the predictor 220, dequantizer 234, inverse transformer 235, and filter 260 of the encoding device 200 in the same way, respectively.

[0095] As described above, in video compilation, prediction is performed to improve the compression efficiency. Through this operation, a prediction block including prediction samples for the current block (i.e., the block to be compiled) can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived in the same way in both the encoding device and the decoding device, and the encoding device decodes the information (residual information) regarding the residual between the original block and the prediction block (rather than the original sample values of the original block itself). Signaling to the device can improve the image compilation efficiency. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and generate a reconstructed picture including the reconstructed block.

[0096] The residual information can be generated through transformation processing and quantization processing. For example, the encoding device can derive the residual block between the original block and the prediction block, perform transformation processing on the residual samples (residual sample array) included in the residual block to derive transform coefficients, and then derive the quantized transform coefficients by performing quantization processing on the transform coefficients to signal the residual-related information (via the bitstream) to the decoding device. Here, the residual information can include value information, position information, transformation technology, transformation core, and quantization parameters of the quantized transform coefficients, etc. The decoding device can perform dequantization / inverse transformation processing based on the residual information and derive the residual samples (or residual block). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. The encoding device can also dequantize / inverse transform the quantized transform coefficients for inter-frame prediction reference of subsequent pictures to derive the residual block and generate a reconstructed picture based on this.

[0097] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or, for the sake of consistency in expression, may still be referred to as transform coefficients.

[0098] In this document, the quantized transform coefficients and the transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled by the residual coding syntax. The transform coefficients may be derived based on the residual information (or the information about the transform coefficients), and the scaled transform coefficients may be derived by inverse-transforming (scaling) the transform coefficients. The residual samples may be derived based on the inverse-transform (transform) of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.

[0099] The predictor of an encoding device / decoding device may derive a prediction sample by performing inter-frame prediction on a block-by-block basis. Inter-frame prediction may be prediction derived in a manner that depends on data elements (e.g., sample values or motion information) of pictures other than the current picture. When applying inter-frame prediction to a current block, a prediction block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index. Here, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information of the current block may be predicted on a block, sub-block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include information on an inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, a motion information candidate list may be configured based on neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) may be signaled to derive the motion vector and / or reference picture index of the current block. Inter-frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the motion information of the current block may be the same as the motion information of neighboring blocks. In the skip mode, different from the merge mode, a residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be signaled. In this case, the sum of the motion vector predictor and the motion vector difference may be used to derive the motion vector of the current block.

[0100] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), the motion information may include L0 motion information and / or L1 motion information. The motion vector in the L0 direction may be referred to as the L0 motion vector or MVL0, and the motion vector in the L1 direction may be referred to as the L1 motion vector or MVL1. The prediction based on the L0 motion vector may be referred to as L0 prediction, the prediction based on the L1 motion vector may be referred to as L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bidirectional prediction. Here, the L0 motion vector may indicate a motion vector associated with the reference picture list L0 (L0), and the L1 motion vector may indicate a motion vector associated with the reference picture list L1 (L1). The reference picture list L0 may include pictures earlier than the current picture in the output order as reference pictures, and the reference picture list L1 may include pictures later than the current picture in the output order. The previous pictures may be referred to as forward (reference) pictures, and the subsequent pictures may be referred to as backward (reference) pictures. The reference picture list L0 may also include pictures later than the current picture in the output order as reference pictures. In this case, the previous pictures may be indexed first in the reference picture list L0, and the subsequent pictures may be indexed later. The reference picture list L1 may also include pictures earlier than the current picture in the output order as reference pictures. In this case, the subsequent pictures may be indexed first in the reference picture list 1, and the previous pictures may be indexed later. The output order may correspond to the picture order count (POC) order.

[0101] In addition, in the case of performing intra-frame prediction, the correlation between samples may be used, and the difference between the original block and the predicted block (i.e., the residual) may be obtained. The above-mentioned transformation and quantization may be applied to the residual, and thereby the spatial redundancy may be removed. Hereinafter, the encoding method and the decoding method using intra-frame prediction will be described in detail.

[0102] Intra-frame prediction means generating a predicted sample for a current block based on reference samples outside the current block in a picture including the current block (hereinafter referred to as the current picture). Here, the reference samples outside the current block may be referred to as samples located around the current block. In the case of applying intra-frame prediction to the current block, the neighboring reference samples for the intra-frame prediction of the current block may be derived.

[0103] For example, when the size (width × height) of the current block is nW × nH, the neighboring reference samples of the current block may include samples that are adjacent to the left boundary of the current block and are a total of 2 × nH lower-left neighboring samples, samples that are adjacent to the upper boundary of the current block and are a total of 2 × nW upper-right neighboring samples, and one upper-left neighboring sample of the current block. In addition, the neighboring reference samples of the current block may include a plurality of column-top neighboring samples and a plurality of row-left neighboring samples. In addition, the neighboring reference samples of the current block may include a total of nH samples of size nW × nH that are adjacent to the right boundary of the current block, a total of nW samples that are adjacent to the bottom boundary of the current block, and one lower-right neighboring sample of the current block.

[0104] However, some of the neighboring reference samples of the current block may not have been decoded yet, or may not be available yet. In such a case, the decoding device may configure the neighboring reference samples for prediction by replacing the unavailable samples with available samples. In addition, the decoding device may configure the neighboring reference samples for prediction by interpolating the available samples.

[0105] In the case of deriving neighboring reference samples, (i) a prediction sample may be derived based on the average value or interpolation of the neighboring reference samples of the current block, or (ii) a prediction sample may be derived based on the reference samples that exist in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. The case of (i) may be applied to the case where the intra prediction mode is a non-directional mode or a non-angle mode, and the case of (ii) may be applied to the case where the intra prediction mode is a directional mode or an angle mode.

[0106] In addition, a prediction sample may be generated by interpolating between a first neighboring sample in the prediction direction of the intra prediction mode of the current block and a second neighboring sample in the direction opposite to the prediction direction based on the prediction sample of the current block among the neighboring reference samples. The above case may be referred to as linear interpolation intra prediction (LIP). In addition, chrominance prediction samples may be generated based on luminance samples using a linear model. This case may be referred to as the LM mode.

[0107] In addition, a temporary prediction sample of the current block may be derived based on filtered neighboring reference samples, and a prediction sample of the current block may be derived by performing a weighted sum of the temporary prediction sample and at least one reference sample derived according to the intra prediction mode among the existing neighboring reference samples (i.e., unfiltered neighboring reference samples). The above case may be referred to as position-dependent intra prediction (PDPC).

[0108] In addition, intra prediction compilation can be performed in the following method, which derives a prediction sample by selecting a corresponding line among adjacent multi-reference sample lines of a current block and using a reference sample in the prediction direction in the reference sample line with the highest prediction accuracy, and indicates (signals) the reference sample line used at that time to a decoding device. The above situation can be referred to as multi-reference line (MRL) intra prediction or MRL-based intra prediction.

[0109] In addition, when performing intra prediction based on the same intra prediction mode by dividing a current block into vertical or horizontal sub-partitions, adjacent reference samples can be derived and used in units of sub-partitions. That is, the intra prediction mode for the current block can be equally applied to the sub-partitions, and in this case, since adjacent reference samples are derived and used in units of sub-partitions, the intra prediction performance can be enhanced in some cases. This prediction method can be referred to as intra sub-partition (ISP) or ISP-based intra prediction.

[0110] Different from the intra prediction mode, the above intra prediction methods can be referred to as intra prediction types. Intra prediction types can be referred to by various terms such as intra prediction techniques or additional intra prediction modes. For example, an intra prediction type (or additional intra prediction mode) can include at least one of LIP, PDPC, MRL, and ISP as described above. A general intra prediction method other than a specific intra prediction type such as LIP, PDPC, MRL, or ISP can be referred to as a normal intra prediction type. In the case where the above specific intra prediction types are not applied, the normal intra prediction type can generally be applied, and prediction can be performed based on the above intra prediction mode. In addition, post-filtering can be performed on the derived prediction samples as needed.

[0111] In addition, in addition to the above intra prediction types, matrix-based intra prediction (hereinafter referred to as MIP) can be used as a method for intra prediction. MIP can be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (NWIP).

[0112] In the case of applying MIP to a current block, i) adjacent reference samples that have been averaged are used, ii) matrix-vector multiplication processing can be performed, and iii) a prediction sample for the current block can be derived by further performing horizontal / vertical interpolation processing as needed. The intra prediction mode for MIP can be the above LIP, PDPC, MRL, or ISP intra prediction, but can also be configured differently from the intra prediction mode used in normal intra prediction.

[0113] The intra prediction mode for MIP can be referred to as "affine linear weighted intra prediction mode" or matrix-based intra prediction mode. For example, according to the intra prediction mode for MIP, the matrix and offset used in matrix-vector multiplication can be configured differently. Here, the matrix can be referred to as (affine) weight matrix, and the offset can be referred to as (affine) offset vector or (affine) bias vector. In this document, the intra prediction mode for MIP can be referred to as MIP intra prediction mode, linear weighted intra prediction mode, matrix weighted intra prediction mode or matrix-based intra prediction mode. The detailed MIP method will be described later.

[0114] The following drawings have been prepared to explain detailed examples of this document. Since the names of detailed devices, detailed terms or names (e.g., names of grammars) described in the drawings are presented exemplarily, the technical features of this document are not limited to the detailed names used in the following drawings.

[0115] Figure 4 Schematically illustrate an example of an intra prediction-based image coding method to which an exemplary embodiment of this document is applicable, and Figure 5 Schematically illustrate an intra predictor in an encoding device. Figure 5 The intra predictor in the encoding device illustrated in Figure 2 can also be applied to the intra predictor 222 of the encoding device 200 illustrated in

[0116] Refer to Figure 4 and Figure 5 , S400 can be executed by the intra predictor 222 of the encoding device, and S410 can be executed by the residual processor 230 of the encoding device. Specifically, S410 can be executed by the subtractor 231 of the encoding device. In S420, the prediction information can be derived by the intra predictor 222 and encoded by the entropy encoder 240. In S420, the residual information can be derived by the residual processor 230 and encoded by the entropy encoder 240. The residual information indicates information about residual samples. The residual information can include information about the quantized transform coefficients of the residual samples. As described above, the residual samples can be derived by the transformer of the encoding device through transform coefficients, and the transform coefficients can be derived by the quantizer through the quantized transform coefficients. The information about the quantized transform coefficients can be encoded by the entropy encoder 240 through the residual compilation process.

[0117] The encoding device performs intra prediction on the current block (S400). The encoding device may derive an intra prediction mode / type for the current block, derive neighboring reference samples for the current block, and generate prediction samples in the current block based on the intra prediction mode / type and the neighboring reference samples. Here, the processes of determining the intra prediction mode / type, deriving the neighboring reference samples, and generating the prediction samples may also be executed simultaneously, and any one of the processes may be executed earlier than the other processes.

[0118] For example, the intra predictor 222 of the encoding device may include an intra prediction mode / type determiner 222-1, a reference sample deriver 222-2, and a prediction sample deriver 222-3. The intra prediction mode / type determiner 222-1 may determine the intra prediction mode / type for the current block, the reference sample deriver 222-2 may derive the neighboring reference samples for the current block, and the prediction sample deriver 222-3 may derive the prediction samples for the current block. Meanwhile, although not illustrated, if a prediction sample filtering process is performed, the intra predictor 222 may further also include a prediction sample filter (not illustrated). The encoding device may determine the mode / type applied to the current block among multiple intra prediction modes / types. The encoding device may compare the RD costs for the intra prediction modes / types and determine the best intra prediction mode / type for the current block.

[0119] As described above, the encoding device may also perform a prediction sample filtering process. Prediction sample filtering may be referred to as post-filtering. Some or all of the prediction samples may be filtered through the prediction sample filtering process. In some cases, the prediction sample filtering process may be omitted.

[0120] The encoding device generates residual samples for the current block based on the (filtered) prediction samples (S410). The encoding device may derive the residual samples by comparing the prediction samples with the original samples of the current block based on the phase.

[0121] The encoding device may encode the image information including information about intra prediction (prediction information) and residual information about the residual samples (S420). The prediction information may include intra prediction mode information and intra prediction type information. The residual information may include residual coding syntax. The encoding device may derive quantized transform coefficients by performing transformation / quantization of the residual samples. The residual information may include information about the quantized transform coefficients.

[0122] The encoding device may output the encoded image information in the form of a bitstream. The output bitstream may be transmitted to the decoding device through a storage medium or a network.

[0123] As described above, the encoding device can generate reconstructed pictures (including reconstructed samples and reconstructed blocks). To this end, the encoding device can derive (modified) residual samples by performing dequantization / inverse transformation of the quantized transform coefficients again. The reason for dequantizing / inverse transforming the residual samples again after transformation / quantization is to derive the same residual samples as those derived by the decoding device as described above. The encoding device can generate a reconstructed block including the reconstructed samples for the current block based on the prediction samples and the (modified) residual samples. Based on the reconstructed blocks, a reconstructed picture for the current picture can be generated. As described above, loop filter processing can be further applied to the reconstructed picture.

[0124] Figure 6 Schematically illustrate an example of an intra prediction-based image decoding method to which the exemplary embodiments of this document are applicable, and Figure 7 Schematically illustrate the intra predictor in the decoding device. Figure 7 The intra predictor in the illustrated decoding device can also be applied to the intra predictor 331 of the decoding device 300 illustrated in Figure 2 in the same or corresponding manner.

[0125] Refer to Figure 6 and Figure 7 , the decoding device can perform operations corresponding to the foregoing operations performed by the encoding device. S600 to S620 can be performed by the intra predictor 331 of the decoding device, and the prediction information in S600 and the residual information in S630 can be obtained from the bitstream by the entropy decoder 310 of the decoding device. The residual processor 320 of the decoding device can derive the residual samples for the current block based on the residual information. Specifically, the dequantizer 321 of the residual processor 320 can derive the transform coefficients by performing dequantization based on the quantized transform coefficients derived based on the residual information, and the inverter 322 of the residual processor can derive the residual samples for the current block by performing inverse transformation on the transform coefficients. S640 can be performed by the adder 340 or the reconstructor of the decoding device.

[0126] The decoding device can derive the intra prediction mode / type for the current block based on the received prediction information (intra prediction mode / type information) (S600). The decoding device can derive the neighboring reference samples of the current block (S610). The decoding device generates the prediction samples in the current block based on the intra prediction mode / type and the neighboring reference samples (S620). In this case, the decoding device can perform a prediction sample filtering process. Prediction sample filtering can be referred to as post-filtering. Some or all of the prediction samples can be filtered through the prediction sample filtering process. In some cases, the prediction sample filtering process can be omitted.

[0127] The decoding device generates residual samples for a current block based on received residual information (S630). The decoding device may generate reconstructed samples for the current block based on the prediction samples and the residual samples, and derive a reconstructed block including the reconstructed samples (S640). A reconstructed picture for the current picture may be generated based on the reconstructed block. As described above, processes such as a loop filter process may be further applied to the reconstructed picture.

[0128] Here, the intra predictor 331 of the decoding device may include an intra prediction mode / type determiner 331-1, a reference sample derivator 331-2, and a prediction sample derivator 331-3. The intra prediction mode / type determiner 331-1 may determine an intra prediction mode / type for the current block based on the intra prediction mode / type information obtained by the entropy decoder 310. The reference sample derivator 331-2 may derive neighboring reference samples of the current block, and the prediction sample derivator 331-3 may derive prediction samples of the current block. Meanwhile, although not illustrated, if the foregoing prediction sample filtering process is performed, the intra predictor 331 may further also include a prediction sample filter (not illustrated).

[0129] For example, the intra prediction mode information may include, for example, flag information (e.g., intra_luma_mpm_flag) that indicates whether the most probable mode (MPM) is applied to the current block or whether the remaining mode is applied to the current block. At this time, if the MPM is applied to the current block, the prediction mode information may further include index information (e.g., intra_luma_mpm_idx) that indicates one of the intra prediction mode candidates (MPM candidates). The intra prediction mode candidates (MPM candidates) may be composed of an MPM candidate list or an MPM list. In addition, if the MPM is not applied to the current block, the intra prediction mode information may further include remaining mode information (e.g., intra_luma_mpm_remainder) that indicates one of the remaining intra prediction modes other than the intra prediction mode candidates (MPM candidates). The decoding device may determine the intra prediction mode of the current block based on the intra prediction mode information.

[0130] In addition, the intra prediction type information can be implemented in various forms. As an example, the intra prediction type information may include intra prediction type index information indicating one of the intra prediction types. As another example, the intra prediction type information may include at least one of the following items: reference sample line information (e.g., intra_luma_ref_idx) indicating whether MRL is applied to the current block and which reference sample line is used if MRL is applied, ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether ISP is applied to the current block, ISP type information (e.g., intra_subpartitions_split_flag) indicating the split type of the subpartition in the case where ISP is applied, flag information indicating whether PDCP is applied, or flag information indicating whether LIP is applied. In addition, the intra prediction type information may include an MIP flag indicating whether MIP is applied to the current block.

[0131] The above intra prediction mode information and / or intra prediction type information can be encoded / decoded by the compilation methods described in this document. For example, the above intra prediction mode information and / or intra prediction type information can be encoded / decoded by entropy compilation based on truncated (Rice) binary codes (e.g., CABAC, CAVLC).

[0132] In addition, in the case of applying intra prediction, the intra prediction mode of neighboring blocks can be used to determine the intra prediction mode applied to the current block. For example, the decoding device may select one of the mpm candidates in the mpm list derived from the intra prediction modes of neighboring blocks (e.g., left and / or top neighboring blocks) of the current block and additional candidate modes based on the received most probable mode (mpm) index, or may select one of the remaining intra prediction modes not included in the mpm candidates (and the planar mode) based on the remaining intra prediction mode information. The mpm list may be constructed to include or not include the planar mode as a candidate. For example, if the mpm list includes the planar mode as a candidate, the mpm list may have 6 candidates, while if the mpm list does not include the planar mode as a candidate, the mpm list may have 5 candidates. If the mpm list does not include the planar mode as a candidate, a non-planar flag (e.g., intra_luma_not_planar_flag) indicating whether the intra prediction mode of the current block is not the planar mode may be signaled. For example, the mpm flag may be signaled first, and when the value of the mpm flag is 1, the mpm index and the non-planar flag may be signaled. In addition, the mpm index may be signaled when the value of the non-planar flag is 1. Here, constructing the mpm list that does not include the planar mode as a candidate first identifies whether the intra prediction mode is the planar mode by signaling the flag (non-planar flag) first, because the planar mode is always considered an mpm while the non-planar mode is not an mpm.

[0133] For example, based on an MPM flag (e.g., intra_luma_mpm_flag), it can be indicated whether the intra prediction mode applied to the current block is in the MPM candidates (and planar mode) or in the residual mode. An MPM flag value of 1 can indicate that the intra prediction mode for the current block is in the MPM candidates (and planar mode), and an MPM flag value of 0 can indicate that the intra prediction mode for the current block is not in the MPM candidates (and planar mode). A non-planar flag (e.g., intra_luma_not_planar_flag) value of 0 can indicate that the intra prediction mode for the current block is the planar mode, and a non-planar flag value of 1 can indicate that the intra prediction mode for the current block is not the planar mode. The MPM index can be signaled in the form of an mpm_idx or intra_luma_mpm_idx syntax element, and the remaining intra prediction mode information can be signaled in the form of a rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, the remaining intra prediction mode information can index the remaining intra prediction modes among all the intra prediction modes that are not included in the MPM candidates (and planar mode) in the order of their prediction mode numbers, and can indicate one of them. The intra prediction mode can be the intra prediction mode for the luminance component (samples). Hereinafter, the intra prediction mode information can include at least one of an MPM flag (e.g., intra_luma_mpm_flag), a non-planar flag (e.g., intra_luma_not_planar_flag), an MPM index (e.g., mpm_idx or intra_luma_mpm_idx), and the remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list can be referred to by various terms such as an MPM candidate list, a candidate mode list (candModelList), and a candidate intra prediction mode list.

[0134] Generally, when a block of an image is segmented, the current block to be coded and neighboring blocks have similar image attributes. Therefore, the current block and neighboring blocks are likely to have the same or similar intra prediction modes. Thus, the encoder can use the intra prediction mode of a neighboring block to code the intra prediction mode of the current block. For example, the encoder / decoder can form a most probable mode (MPM) list for the current block. The MPM list can also be referred to as an MPM candidate list. Here, the MPM can mean a mode used to improve the coding efficiency by taking into account the similarity between the current block and neighboring blocks when coding the intra prediction mode.

[0135] Figure 8The figure shows an example of an intra prediction mode to which the embodiments of this document are applicable.

[0136] Reference Figure 8 , the modes can be divided into intra prediction modes with horizontal directivity and intra prediction modes with vertical directivity around the 34th intra prediction mode of the upper left diagonal prediction direction. In Figure 8 , H and V respectively mean horizontal directivity and vertical directivity. Each of the numbers from -32 to 32 indicates a displacement of 1 / 32 unit at the sample grid position. The 2nd to 33rd intra prediction modes have horizontal directivity, and the 34th to 66th intra prediction modes have vertical directivity. The 18th intra prediction mode and the 50th intra prediction mode respectively indicate a horizontal intra prediction mode and a vertical intra prediction mode. The 2nd intra prediction mode can be called a lower left diagonal intra prediction mode, the 34th intra prediction mode can be called an upper left diagonal intra prediction mode, and the 66th intra prediction mode can be called an upper right diagonal intra prediction mode.

[0137] In addition, the intra prediction modes used in the above MIP are not existing directional modes, but can indicate a matrix and an offset for intra prediction. That is, a matrix and an offset for intra prediction can be derived through the intra mode of the MIP. In this case, in the case of deriving an intra mode for generating the above typical intra prediction or MPM list, the intra prediction mode of the block predicted by the MIP can be configured as a pre-configured mode (e.g., a planar mode or a DC mode). In addition, according to another example, the intra mode of the MIP can be mapped to a planar mode, a DC mode, or a directional intra mode based on the block size.

[0138] Hereinafter, as a method for intra prediction, matrix-based intra prediction (hereinafter, MIP) will be described.

[0139] As described above, matrix-based intra prediction can be called affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). To predict samples of a rectangular block with width W and height H, the MIP uses one H line among the reconstructed neighboring left boundary samples of the block and one W line among the reconstructed neighboring top boundary samples of the block as input values. If the reconstructed samples are not available, reference samples can be generated by an interpolation method already applied in typical intra prediction.

[0140] Figure 9 is a figure explaining the process of generating MIP-based prediction samples according to an embodiment. Reference Figure 9 , the MIP process will be described as follows.

[0141] 1. Average calculation process

[0142] By averaging, 4 boundary samples can be extracted when W = H = 4, and 8 boundary samples can be extracted in other cases.

[0143] 2. Matrix-vector multiplication processing

[0144] Perform matrix-vector multiplication using the input of the average samples, and then add an offset. Through such operations, the reduced prediction samples of the subsampled sample set in the original block can be derived.

[0145] 3. (Linear) interpolation processing

[0146] The prediction samples at the remaining positions are generated from the prediction samples of the subsampled sample set by linear interpolation, which is single-step linear interpolation in each direction.

[0147] The matrices and offset vectors required to generate the prediction block or prediction samples can be selected from three sets S0, S1, and S2 for the matrix.

[0148] Set S0 can consist of 18 matrices A0 i , i ∈ {0, …, 17} and 18 offset vectors b0 i , i ∈ {0, …, 17}. For example, each matrix A0 i , i ∈ {0, …, 17} can have 16 rows and 4 columns. In the example, each offset vector b0 i , i ∈ {0, …, 17} can have a size of 16. The matrices and offset vectors of set S0 can be used for blocks of size 4×4.

[0149] Set S1 can consist of 10 matrices A1 i , i ∈ {0, …, 9} and 10 offset vectors b1 i , i ∈ {0, …, 9}. For example, each matrix A1 i , i ∈ {0, …, 9} can have 16 rows and 8 columns. In the example, each offset vector b1 i , i ∈ {0, …, 9} can have a size of 16. The matrices and offset vectors of set S1 can be used for blocks of sizes 4×8, 8×4, and 8×8.

[0150] Finally, set S2 can consist of 6 matrices A2 i , i ∈ {0, …, 5} and 6 offset vectors b2 i , i ∈ {0, …, 5}. For example, each matrix A2 i , i ∈ {0, …, 5} can have 64 rows and 8 columns. In the example, each offset vector b2 i, the size of i ∈ {0, …, 5} can be 64. The matrix and offset vectors of set S2 or some of them can be used for all other block types of sizes that do not apply set S0 and set S1.

[0151] The total number of multiplications required to compute the matrix-vector product may always be equal to or less than 4 × W × H. For example, in the MIP mode, up to four multiplications may be required per sample.

[0152] Regarding the embodiment of averaging neighboring samples, outside the boundary samples, in the case of W = H = 4, four samples can be extracted by averaging, and eight samples can be extracted by averaging. For example, the input boundaries bdry left and bdry top can be reduced to a smaller boundary by averaging neighboring boundary samples according to predefined rules depending on the block size and .

[0153] Since the two reduced boundaries and are connected to the reduced boundary vector bdry red , in the case of a 4×4 type block, the size of bdry red is 4, and in the case of all other blocks, its size is 8.

[0154] In the case where the "mode" is referred to as the MIP mode, the range of the reduced boundary vector bdry red and the MIP mode value (mode) can be defined in the following formula.

[0155] [Equation 1]

[0156]

[0157] Regarding the embodiment of matrix-vector multiplication, matrix-vector multiplication can be performed using the averaged samples as input. One of the reduced input vectors bdry red generates a reduced prediction signal pred red . The predicted samples are signals for the downsampled block with width W red and height H red . For example, W red and H red can be defined in the following formula.

[0158] [Equation 2]

[0159]

[0160] The reduced predicted sample pred redIt can be calculated by adding an offset after performing a matrix-vector multiplication and can be derived by the following formula.

[0161] [Equation 3]

[0162]

[0163] Here, A represents a matrix with W red ×H red rows and 4 columns in the case where W and H are 4 (W = H = 4), or 8 columns in all other cases, and b represents a vector of size W red ×H red .

[0164] The matrix A and the offset vector b can be selected from among the sets S0, S1, and S2. For example, the index idx = idx(W, H) can be defined by the following formula.

[0165] [Equation 4]

[0166]

[0167] If idx is equal to or less than 1 (idx ≤ 1) or idx is 2, and the smaller value between W and H is greater than 4 (min(W, H)>4), then A is set to , and b is set to . If idx is 2, the smaller value between W and H is 4 (min(W, H) = 4), and W is 4, then A becomes the matrix from which the rows corresponding to the odd x coordinates in the downsampling block are removed. In addition, if H is 4, then A becomes the matrix from which the columns corresponding to the odd y coordinates in the downsampling block are removed.

[0168] Since A consists of 4 columns and 16 rows in the case of W = H = 4, the number of multiplications required to calculate pred red is 4. In all other cases, since A consists of 8 columns and W red ×H red rows, it can be confirmed that at most four multiplications per sample are required to calculate pred red .

[0169] The interpolation process can be referred to as linear interpolation or bilinear interpolation processing. As shown in the figure, the interpolation process can include two steps: 1) vertical interpolation and 2) horizontal interpolation.

[0170] In the case where W>=H, vertical linear interpolation can be applied first, and then horizontal interpolation can be applied. <H的情况下,可以首先应用水平线性插值,然后可以应用垂直线性插值。在4×4块的情况下,可以省略插值处理。

[0171] In the case of a W×H block with max(W,H)≥8, red ×H red The reduced prediction sample pred on red Derived prediction samples. Depending on the block type, linear interpolation is performed in the vertical, horizontal, or both directions. In the case where linear interpolation is applied in both directions, <H的情况下,首先在水平方向上应用线性插值,否则首先在垂直方向上应用线性插值。

[0172] In the case of a W×H block with max(W,H)≥8 and W>=H, it can be assumed that there is no generality loss. In this case, one-dimensional linear interpolation is performed as follows. If there is no generality loss, the linear interpolation in the vertical direction can be fully explained.

[0173] First, the downsampled prediction samples are extended to the top by the boundary signal. The vertical upsampling factor can be defined , and if configured as , the extended and reduced prediction samples can be configured as follows.

[0174] For each coding unit (CU) in intra mode, a flag indicates whether the MIP mode is applied to the corresponding current block. If the MIP mode is applied, the MPM flag can be signaled and can indicate whether the prediction mode is one of the MPM modes. For example, three modes can be considered for MPM. In the example, the MPM mode can be context-coded by truncated binarization. The non-MPM mode can be coded as a fixed length code (FLC). The derivation of such an MPM can be combined with the normal intra prediction mode by performing mode mapping between the normal intra prediction mode and the MIP intra prediction mode based on a predefined mapping table that depends on the block size (i.e., idx(W,H)∈{0,1,2}). The following formula can represent a forward (from normal mode to MIP mode) and / or reverse (from MIP mode to normal mode) mode mapping table.

[0175] [Formula 5]

[0176]

[0177] [Formula 6]

[0178]

[0179] In an example, 35 patterns can be used for a block (e.g., max(W,H)<=8&&W*H<32). In another example, 19 patterns can be used for a block (e.g., max(W,H)=8). 11 patterns can be used for a block (e.g., max(W,H)>8). Additionally, to reduce memory consumption, two patterns can share the same matrix and / or offset vector, as shown in the following equation.

[0180] [Equation 7]

[0181]

[0182] In the following, a method for maximizing performance while reducing the complexity of the MIP technique will be described. The embodiments to be described later can be executed individually or can be executed in combination.

[0183] According to an embodiment of this document, it is possible to adaptively select (determine) whether to apply MIP according to the type of the current (luminance) block in MIP processing. In existing embodiments, if the difference between the width and height of a luminance block exceeds four times (e.g., 32×4, 64×8, 64×4, 4×32, 8×64, 4×64, 128×4, 128×8, 128×16, 4×128, 8×128, and 16×128), MIP is not performed, and / or a flag indicating whether MIP is applied is not sent. In this embodiment, a method for improving encoding and decoding efficiency by extending or removing MIP in the case of applying MIP is provided.

[0184] In one of the examples of this embodiment, in the case where the width of the current (luminance) block is 4 and its height is equal to or greater than 16, MIP may not be applied, or in the case where the height of the current (luminance) block is 4 and the width is equal to or greater than 16, MIP may not be applied. Some compilation unit syntax according to this example may be the same as the following table.

[0185] [Table 1]

[0186]

[0187] In the above example, MIP is not applied to 4×16 and 16×4 blocks with low MIP efficiency, so MIP efficiency can be improved.

[0188] In another example of this embodiment, in the case where the width of the current (luminance) block is 4 and its height is equal to or greater than 32, MIP may not be applied, and in the case where the height of the current (luminance) block is 4 and its width is equal to or greater than 32, MIP may not be applied. Some compilation unit syntax according to this example may be as shown in the following table.

[0189] [Table 2]

[0190]

[0191] Through the above examples, for 8×64 and 64×8 blocks to which MIP can be effectively and easily applied, MIP can be applied, so the MIP efficiency can be improved.

[0192] According to another example of this embodiment, the block type restriction for MIP can be removed. For example, in this example, MIP can always be applied to all (compiled) blocks. When the MIP flag is signaled (parsed), the condition determination of the width and / or height of the current (luminance) block can be removed. Therefore, the MIP parsing condition can be simplified, and in addition, the complexity of software and hardware implementation can be reduced. Some compilation unit syntax according to this example can be as shown in the following table.

[0193] [Table 3]

[0194]

[0195] The test results by applying this example can be as shown in the following table.

[0196] [Table 4]

[0197]

[0198] Referring to the above table, the results of all intra (AI, allintra) and random access (RA, randomaccess) are given in the above test. As shown in the above table, it is confirmed that there is no performance degradation in this embodiment. That is, in the case of the proposed method 3, since the signaling (parsing) burden can be simplified by relaxing the MIP flag signaling (parsing) condition, it has an advantage in hardware implementation, and no coding and decoding performance degradation occurs.

[0199] According to this embodiment, in Figure 4 S400 and / or S410, it can be determined whether to apply MIP to the current block. In this case, as described above, it can be adaptively determined whether to apply MIP according to the type of the current (luminance) block. In this case, in S420, the encoding device can adaptively encode intra_mip_flag.

[0200] In an embodiment of this document, the number of MIP modes can be adaptively determined (selected) according to the type of the current (luminance) block in MIP processing. In existing embodiments, a 4×4 luminance block (MipSizeId 0) can use 35 MIP modes, a 4×4, 8×4, or 8×8 luminance block (MipSizeId 1) can use 19 MIP modes, and other luminance blocks (MipSizeId 2) can use 11 MIP modes.

[0201] In this embodiment, a smaller number of MIP modes can be used for block types with low MIP efficiency. Through this embodiment, the amount of transmission data of information about MIP modes can be reduced.

[0202] Figure 10 , 11 and 12 are flowcharts illustrating MIP processing according to an embodiment of this document.

[0203] Refer to Figure 10 , in an embodiment related to Figure 10 , if the difference between the width and height of the current (luminance) block exceeds four times the width or height of the current (luminance) block (if the height of the current block is greater than four times the width of the current block, or if the width of the current block is greater than four times the height of the current block), then MIP modes with a number less than 11 basic modes (MipSizeId 2) can be used, and information related to the MIP modes can be sent.

[0204] Refer to Figure 11 , in an embodiment related to Figure 11 , if the width of the current (luminance) block is 4 and its height is equal to or greater than 16, then MIP modes with a number less than 11 basic modes (MipSizeId 2) can be used, and information related to the MIP modes can be sent. Further, if the height of the current (luminance) block is 4 and its width is equal to or greater than 16, then MIP modes with a number less than 11 basic modes (MipSizeId 2) can be used, and information related to the MIP modes can be sent.

[0205] Refer to Figure 12 , in an embodiment related to Figure 12In related embodiments, if the width of the current (luma) block is 4 and its height is equal to or greater than 32, MIP modes with a number less than 11 basic modes (MipSizeId 2) can be used, and information related to the MIP modes can be transmitted. Further, if the height of the current (luma) block is 4 and its width is equal to or greater than 32, MIP modes with a number less than 11 basic modes (MipSizeId 2) can be used, and information related to the MIP modes can be transmitted.

[0206] By according to Figures 10 to 12 the embodiments, by using a small number of MIP modes for blocks with low MIP prediction efficiency (block types with a large difference between width and height), the amount of transmitted data of the MIP modes can be reduced, and thereby, the MIP prediction efficiency can be enhanced.

[0207] In addition, in the embodiments explained with reference to Figures 10 to 12 only a part of the 11 MIP modes can be used, and this is for increasing MIP efficiency through data reduction. In an example, only one of the 11 MIP modes can be used, and in this case, MIP mode data may not be transmitted. In another example, only two of the 11 MIP modes can be used, and in this case, 1-bit MIP mode data can be transmitted. In yet another example, only four of the 11 MIP modes can be used, and in this case, 2-bit MIP mode data can be transmitted. One, two, or four of the 11 modes can be selected in the order of the MIP core in the existing embodiments, or can be selected in the order of the MIP selection mode probability in the form of blocks as in the Figures 10 to 12 method explained above.

[0208] Figure 13 and Figure 14 schematically illustrate examples of a video / image encoding method and related components according to the embodiments of this document.

[0209] Figure 13 The method disclosed in Figure 2 or Figure 14 can be executed by the Figure 13 encoding device disclosed in Figure 14 Specifically, for example, Figure 13 S1300 and S1310 of Figure 14 can be executed by the predictor 220 of the Figure 13 encoding device, Figure 14 S1320 of Figure 13Although not illustrated, the prediction sample or prediction-related information may be derived by the predictor 220 of the encoding device of Figure 14 The residual information may be derived by the residual processor 230 of the encoding device from the original sample or the prediction sample, and the bitstream may be generated by the entropy encoder 240 of the encoding device from the residual information or the prediction-related information. Figure 13 The method disclosed in

[0210] Referring to Figure 13 , the encoding device may derive whether matrix-based intra prediction (MIP) is applied to the current block (S1300). If MIP is applied to the current block, the encoding device may generate intra MIP flag information indicating that MIP is applied to the current block. If MIP is not applied to the current block, the encoding device may generate intra MIP flag information indicating that MIP is not applied to the current block (e.g., intra_mip_flag). The encoding device may output the image information including the intra MIP flag information as a bitstream.

[0211] The encoding device may generate an intra prediction sample for the current block (S1310). The encoding device may generate the intra prediction sample based on determining that MIP is applied to the current block.

[0212] The encoding device may generate a residual sample for the current block (S1320). The encoding device may generate the residual sample based on the intra prediction sample. The encoding device may generate the residual sample based on the difference between the original sample for the current block and the intra prediction sample.

[0213] The encoding device may encode the image information including the information about MIP and the information about the residual sample (S1330). The information about MIP may include the intra MIP flag information related to whether MIP is applied to the current block. The intra MIP flag information may be related to whether the intra prediction mode type for the current block is MIP. In addition, the information about MIP may include the intra MIP mode information related to the MIP applied to the current block.

[0214] The information about the residual sample may be referred to as residual information. The encoding device may derive the residual information based on the residual sample. The encoding device may derive the transform coefficients based on the transform process for the residual sample. For example, the transform process may include at least one of DCT, DST, GBT, or CNT. The encoding device may derive the quantized transform coefficients based on the quantization process for the transform coefficients. The quantized transform coefficients may have a one-dimensional vector form based on the coefficient scan order. The encoding device may generate the residual information representing the quantized transform coefficients. The residual information may be generated by various coding methods, such as exponential Golomb, CAVLC, and CABAC.

[0215] It is capable of outputting encoded video / image information in the form of a bitstream. The bitstream can be sent to a decoding device via a network or a storage medium.

[0216] According to an embodiment of this document, the image / video information may include various information. For example, the image / video information may include the information disclosed in at least one of Tables 1 to 3 as described above.

[0217] In an embodiment, a matrix (MIP weight matrix) for MIP may be derived based on the width and / or height of the current block.

[0218] In an embodiment, an intra prediction sample for the current block may be generated based on the matrix.

[0219] In an embodiment, the image information may include a sequence parameter set (SPS). The SPS may include MIP availability flag information (e.g., sps_mip_enabled_flag) related to whether MIP is available.

[0220] In an embodiment, in order to generate an intra prediction sample, an encoding device may derive reduced boundary samples by downsampling reference samples (boundary samples) adjacent to the current block, generate MIP samples based on the product between the reduced boundary samples and the matrix, and perform upsampling of the MIP samples. The intra prediction sample may be generated based on the upsampled MIP samples.

[0221] In an embodiment, the height of the current block may be greater than four times the width of the current block.

[0222] In an embodiment, the width of the current block may be greater than four times the height of the current block.

[0223] In an embodiment, the size of the current block is 32×4, 4×32, 64×8, 8×64, 64×4, 4×64, 128×4, 128×8, 128×16, 4×128, 8×128, or 16×128.

[0224] In an embodiment, the image information may include MIP mode information. The matrix may be further derived based on the MIP mode information, and / or the syntax element bin string for the MIP mode information may be binarized by a truncation binarization method.

[0225] In an embodiment, the matrix may be derived based on three matrix sets classified according to the size of the current block, and / or each of the three matrix sets may include multiple matrices.

[0226] Figure 15 and Figure 16Schematically illustrate an example of a video / image decoding method and related components according to an embodiment of this document.

[0227] Figure 15 The method disclosed in Figure 3 or Figure 16 can be executed by the decoding device disclosed in Figure 15 Specifically, for example, Figure 16 S1500 of Figure 15 can be executed by the entropy decoder 310 of the decoding device of Figure 16 S1510 and S1520 of Figure 15 can be executed by the predictor 330 of the decoding device of Figure 16 and S1530 of Figure 15 can be executed by the adder 340 of the decoding device of Figure 16 In addition, although not illustrated in Figure 15 the prediction-related information or residual information can be derived from the bitstream by the entropy decoder 310 of the decoding device of

[0228] Refer to Figure 15 The decoding device can receive image information (S1500) including intra MIP flag information (e.g., intra_mip_flag) related to whether IMP is applied to the current block through the bitstream. For example, the decoding device can obtain the intra MIP flag information by parsing and decoding the bitstream. Here, the bitstream can be referred to as encoded (image) information.

[0229] The image / video information can include various information according to the embodiments of this document. For example, the image / video information can include the information disclosed in at least one of Tables 1 to 3 as described above.

[0230] The decoding device can derive the intra prediction type of the current block through MIP (S1510). The decoding device can derive the intra prediction type through MIP based on the intra MIP flag information.

[0231] The decoding device can generate an intra prediction sample for the current block (S1520). The decoding device can generate an intra prediction sample based on MIP. The decoding device can generate an intra prediction sample based on the neighboring reference samples in the current picture including the current block.

[0232] The decoding device may generate reconstructed samples for the current block (S1530). The reconstructed samples may be generated based on intra prediction samples. The decoding device may directly use the prediction samples as the reconstructed samples according to the prediction mode, or may generate the reconstructed samples by adding the residual samples to the prediction samples.

[0233] In an embodiment, a matrix (MIP weight matrix) for MIP may be derived based on the width and / or height of the current block.

[0234] In an embodiment, intra prediction samples for the current block may be generated based on the matrix.

[0235] In an embodiment, the picture information may include a sequence parameter set (SPS). The SPS may include MIP availability flag information (e.g., sps_mip_enabled_flag) related to whether MIP is available.

[0236] In an embodiment, to generate intra prediction samples, the decoding device may derive reduced boundary samples by downsampling reference samples (boundary samples) adjacent to the current block, generate MIP samples based on the product between the reduced boundary samples and the matrix, and / or perform upsampling of the MIP samples. The intra prediction samples may be generated based on the upsampled MIP samples.

[0237] In an embodiment, the height of the current block may be greater than four times the width of the current block.

[0238] In an embodiment, the width of the current block may be greater than four times the height of the current block.

[0239] In an embodiment, the size of the current block is 32×4, 4×32, 64×8, 8×64, 64×4, 4×64, 128×4, 128×8, 128×16, 4×128, 8×128, or 16×128.

[0240] In an embodiment, the picture information may include MIP mode information. The matrix may be further derived based on the MIP mode information, and / or the syntax element bin string for the MIP mode information may be binarized by a truncation binarization method.

[0241] In an embodiment, the matrix may be derived based on three matrix sets classified according to the size of the current block, and / or each of the three matrix sets may include multiple matrices.

[0242] In the presence of residual samples for the current block, the decoding device may receive information about the residuals for the current block. The information about the residuals may include transform coefficients for the residual samples. The decoding device may derive the residual samples (or an array of residual samples) for the current block based on the residual information. Specifically, the decoding device may derive the quantized transform coefficients based on the residual information. The quantized transform coefficients may have a one-dimensional vector form based on the coefficient scan order. The decoding device may derive the transform coefficients based on an inverse quantization process for the quantized transform coefficients. The decoding device may derive the residual samples based on the transform coefficients.

[0243] The decoding device may generate reconstructed samples based on (intra) prediction samples and residual samples, and may derive a reconstructed block or a reconstructed picture based on the reconstructed samples. Specifically, the decoding device may generate the reconstructed samples based on the sum of the (intra) prediction samples and the residual samples. Thereafter, as described above, if necessary, the decoding device may apply loop filter processing (e.g., deblocking filter and / or SAO processing) to the reconstructed picture to improve the subjective / objective picture quality.

[0244] For example, the decoding device may obtain image information including all or some of the above information (or syntax elements) by decoding a bitstream or encoded information. In addition, the bitstream or encoded information may be stored in a computer-readable storage medium, or the above decoding method may be executed.

[0245] In the above embodiments, the method is described based on a flowchart having a series of steps or blocks. The present disclosure is not limited to the order of the above steps or blocks. Some steps or blocks may occur simultaneously with or in a different order from other steps or blocks as described above. In addition, those skilled in the art will understand that the steps shown in the above flowchart are not exclusive, and may include additional steps, or one or more steps in the flowchart may be deleted without affecting the scope of this document.

[0246] The method according to the above embodiments of this document may be implemented in software form, and the encoding device and / or decoding device according to this document may be included, for example, in devices that perform image processing such as TVs, computers, smart phones, set-top boxes, and display devices.

[0247] When the embodiments in this document are implemented in software, the above methods can be implemented as modules (processes, functions, etc.) that perform the above functions. The modules can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory can include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information about the instructions or algorithms for the implementation can be stored in a digital storage medium.

[0248] In addition, the decoding device and the encoding device applying this document can be included in a multimedia broadcast transmission / reception device, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camera, a VoD service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a videoconference video device, a transportation user device (i.e., an in-vehicle user device, an aircraft user device, a ship user device, etc.), and a medical video device, and can be used to process video signals and data signals. For example, an over-the-top (OTT) video device can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smart phone, a tablet computer, a digital video recorder (DVR), etc.

[0249] Furthermore, the processing method applying this document can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure can also be stored in the computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices that store data readable by a computer system. For example, the computer-readable recording medium can include a BD, a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (i.e., transmission via the Internet). Additionally, the bit stream generated by an encoding method can be stored in the computer-readable recording medium or can be transmitted via a wired or wireless communication network.

[0250] In addition, a computer program product based on program code can be used to implement the embodiments of this document, and the program code can be executed in a computer according to the embodiments of this document. The program code can be stored on a computer-readable carrier.

[0251] Figure 17 An example of a content streaming system to which the embodiments disclosed in this document can be applied is shown.

[0252] Referring to Figure 17 , a content streaming system applying the embodiments of this document may mainly include an encoding server, a streaming server, a web server, a media storage, user equipment, and a multimedia input device.

[0253] The encoding server compresses the content input from a multimedia input device (e.g., a smart phone, a camera, a video camera, etc.) into digital data to generate a bitstream, and sends the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server can be omitted.

[0254] A bitstream can be generated by applying the encoding method or bitstream generation method of the present disclosure, and the streaming server can temporarily store the bitstream in the process of sending or receiving the bitstream.

[0255] The streaming server sends multimedia data to the user equipment via the web server based on the user's request, and the web server serves as a medium for notifying the user of the service. When the user requests a desired service from the web server, the web server passes the request to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between the devices in the content streaming system.

[0256] The streaming server can receive content from the media storage and / or the encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined time.

[0257] Examples of user equipment may include mobile phones, smartphones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system may be operated as a distributed server, and in this case, the data received from each server may be distributed.

[0258] Each server in the content streaming system may be operated as a distributed server, and in this case, the data received from each server may be processed in a distributed manner.

[0259] The claims described herein may be combined in various ways. For example, the technical features of the method claims in this document may be combined and implemented as a device, or the technical features of the device claims in this document may be combined and implemented as a method. Additionally, the technical features of the method claims in this document and the technical features of the device claims in this document may be combined to be implemented as a device, and the technical features of the method claims in this document and the technical features of the device claims in this document may be combined and implemented as a method.

Claims

1. An apparatus for decoding image information, the apparatus comprising: Memory; and at least one processor connected to the memory, the at least one processor being configured to: obtain, via a bitstream, image information including intra MIP flag information related to whether matrix-based intra prediction (MIP) is applied to a current block; derive, based on the intra MIP flag information, an intra prediction type for the current block via the MIP; generate, based on the MIP, intra prediction samples for the current block; and generate, based on the intra prediction samples, reconstructed samples for the current block, wherein, for the current block having a size equal to 32×4, the MIP flag information for the current block is obtained from the bitstream; wherein, based on the width and height of the current block, a matrix for the MIP is derived; and wherein the intra prediction samples are generated based on the matrix.

2. An apparatus for decoding image information, the apparatus comprising: Memory; and at least one processor connected to the memory, the at least one processor being configured to: determine whether matrix-based intra prediction (MIP) is applied to a current block; generate the MIP flag information based on whether the MIP is applied to the current block; generate intra prediction samples for the current block based on whether the MIP is applied to the current block; generate residual samples for the current block based on the intra prediction samples; and encode image information including information about the MIP, information about the residual samples, and the MIP flag information, wherein, for the current block having a size equal to 32×4, the MIP flag information for the current block is generated and included in the image information; wherein, based on the width and height of the current block, a matrix for the MIP is derived; and wherein the intra prediction samples are derived based on the matrix.

3. An apparatus for transmitting a bitstream for image information, the apparatus comprising: At least one processor configured to obtain a bitstream for the image information, wherein the bitstream is generated based on: determining whether matrix-based intra prediction (MIP) is applied to a current block, generating the MIP flag information based on whether the MIP is applied to the current block, generating intra prediction samples for the current block based on whether the MIP is applied to the current block, generating residual samples for the current block based on the intra prediction samples, and encoding image information including information about the MIP, information about the residual samples, and the MIP flag information; and a transmitter configured to transmit the data including the bitstream, wherein, for the current block having a size equal to 32×4, the MIP flag information for the current block is generated and included in the image information; wherein, based on the width and height of the current block, a matrix for the MIP is derived; and wherein the intra prediction samples are generated based on the matrix.