Image compiling method and device

Through the image compilation method based on intra prediction, the intra MIP syntax elements and variable MIP flags are used to improve the compression efficiency of high resolution and high-quality images/videos, solve the problem of increased transmission and storage costs, and achieve more efficient intra prediction.

CN120475147APending Publication Date: 2025-08-12LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510609365.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-05-11
Filing Date
2021-05-11
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art increases in the amount of information when transmitting and storing high-resolution, high-quality images/videos, and lacks efficient intra-prediction methods.

Method used

Using an image compilation method based on intra prediction, a variable MIP flag is derived by receiving intra MIP syntax elements, setting a preset specific area, exporting an intra prediction mode, and generating a reconstruction block based on the prediction sample.

Benefits of technology

The image/video compression efficiency is improved, and the efficiency of intra prediction is enhanced, especially intra prediction efficiency based on matrix and ISP.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475147A_ABST
    Figure CN120475147A_ABST
Patent Text Reader

Abstract

The invention relates to an image compiling method and device. An image decoding method according to the document comprises the steps of: receiving an intra MIP syntax element for a first target block; deriving a value of the intra MIP syntax element; setting a variable MIP flag for a preset specific region which is the same as the region of the first target block based on the value of the intra-frame MIP syntax element; deriving an intra prediction mode of the second target block based on the variable MIP flag; and deriving a prediction sample of the second target block based on the intra prediction mode of the second target block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with application number 202180033890.4 (PCT / KR2021 / 005837) filed on November 8, 2022, with an international application date of May 11, 2021, and the invention name is “Image compilation method and device thereof”. Technical Field

[0002] This document relates to image coding technology, and more particularly, to an image coding method and apparatus based on intra-frame prediction in an image coding system. Background Art

[0003] Nowadays, the demand for high-resolution and high-quality images / videos such as 4K, 8K or higher ultra-high-definition (UHD) images / videos has been growing in various fields. As image / video data becomes higher resolution and higher quality, the amount of information or bit volume transmitted increases compared to traditional image data. Therefore, when using a medium such as a traditional wired / wireless broadband line to transmit image data or using an existing storage medium to store image / video data, its transmission cost and storage cost increase.

[0004] In addition, today, interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and broadcasting of images / videos having image characteristics different from real images such as game images is increasing.

[0005] Therefore, there is a need for an efficient image / video compression technology that effectively compresses and transmits or stores and reproduces information of high-resolution and high-quality images / videos having various characteristics as described above. Summary of the Invention

[0006] Technical issues

[0007] A technical aspect of the present disclosure is to provide a method and apparatus for improving image coding efficiency.

[0008] Another technical aspect of this document is to provide an efficient intra-frame prediction method and device thereof.

[0009] Yet another technical aspect of this document is to provide an image coding method and device for matrix-based intra-frame prediction.

[0010] Yet another technical aspect of this document is to provide an image coding method and device for IPS-based intra-frame prediction.

[0011] Technical Solution

[0012] According to an embodiment of the present document, an image decoding method performed by a decoding device is provided. The method may include: receiving image information including intra-frame prediction type information from a bitstream, the intra-frame prediction type information including an intra-frame MIP syntax element for a first target block; deriving a value of the intra-frame MIP syntax element; setting a variable MIP flag for a preset specific area that is the same as the area of the first target block based on the value of the intra-frame MIP syntax element; deriving an intra-frame prediction mode for a second target block; deriving prediction samples for the second target block based on the intra-frame prediction mode for the second target block; and generating a reconstructed block based on the prediction samples, wherein the intra-frame prediction mode for the second target block can be derived based on the variable MIP flag for the first target block.

[0013] The first target block may be a left neighboring block of the second target block, and deriving the intra-frame prediction mode for the second target block may include: deriving a candidate intra-frame prediction mode based on a variable MIP flag; and deriving the intra-frame prediction mode for the second target block based on the candidate intra-frame prediction mode, the specific area may include a sample position of (xCb-1, yCb+cbHeight-1), (xCb, yCb) may be the upper left sample position of the second target block, and cbHeight may indicate the height of the second target block.

[0014] The first target block may be an upper neighboring block of the second target block, and deriving the intra-frame prediction mode for the second target block may include: deriving a candidate intra-frame prediction mode based on a variable MIP flag; and deriving the intra-frame prediction mode for the second target block based on the candidate intra-frame prediction mode, the specific area may include a sample position of (xCb+cbWidth-1, yCb-1), (xCb, yCb) may be the position of the upper left sample of the second target block, and cbWidth may indicate the width of the second target block.

[0015] The second target block may include a chrominance block, the first target block may be a luminance block related to the chrominance block, and deriving the intra-frame prediction mode for the second target block may include: deriving the corresponding luminance intra-frame prediction mode based on the variable MIP flag; and deriving the intra-frame prediction mode for the second target block based on the corresponding luminance intra-frame prediction mode.

[0016] Here, based on the fact that the tree type of the second target block is not a single tree or its chroma array type is not 3, the specific area may include a sample position of (xCb+cbWidth / 2, yCb+cbHeight / 2), (xCb, yCb) may indicate the upper left position of the chroma block in the luminance sample unit, cbWidth may indicate the width of the corresponding luminance block corresponding to the chroma block, and cbHeight may indicate the height of the corresponding luminance block.

[0017] The variable MIP flag may be set for a specific region based on the tree type of the first target block not being dual-tree chroma.

[0018] According to another embodiment of the present document, an image encoding method performed by an encoding device is provided. The method may include: deriving a value of an intra MIP flag for a first target block when an intra MIP mode is applied to the first target block; setting a variable MIP flag for a preset specific area that is the same as the area of the first target block based on the value of the intra MIP flag; deriving an intra prediction mode for a second target block; deriving prediction samples for the second target block based on the intra prediction mode for the second target block; deriving residual samples for the second target block based on the prediction samples; and encoding and outputting transform coefficient information generated based on the intra MIP flag and the residual samples, wherein the intra prediction mode for the second target block may be derived based on the variable MIP flag for the first target block.

[0019] According to still another embodiment of this document, there may be provided a digital storage medium storing image data including encoded image information and / or a bitstream generated according to an image encoding method performed by an encoding device.

[0020] According to yet another embodiment of this document, a digital storage medium may be provided that stores image data and / or a bit stream including encoded image information to enable a decoding device to perform an image decoding method.

[0021] Beneficial effects

[0022] According to this document, it is possible to increase overall image / video compression efficiency.

[0023] According to this document, efficiency in intra prediction can be improved.

[0024] According to this document, the image coding efficiency of matrix-based intra prediction can be improved.

[0025] According to this document, the image coding efficiency of ISP-based intra-frame prediction can be improved.

[0026] The effects that can be achieved through the specific examples of the present disclosure are not limited to the effects listed above. For example, there may be various technical effects that can be understood or derived from the present disclosure by a person of ordinary skill in the relevant field. Therefore, the specific effects of the present disclosure are not limited to the effects explicitly described in the present disclosure, and can include various effects that can be understood or derived from the technical features of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 An example of a video / image coding system to which the present disclosure is applicable is schematically illustrated.

[0028] Figure 2 is a diagram schematically illustrating the configuration of a video / image encoding device to which the present disclosure is applicable.

[0029] Figure 3 is a diagram schematically illustrating a configuration of a video / image decoding device to which the present disclosure is applicable.

[0030] Figure 4 The structure of a content streaming system to which the present disclosure is applied is illustrated.

[0031] Figure 5 Intra directional modes for 65 prediction directions are shown.

[0032] Figure 6 FIGURE 1 illustrates a MIP-based prediction sample generation process according to an example.

[0033] Figure 7 FIG. 1 illustrates an example of sub-blocks into which a coding block is divided.

[0034] Figure 8 FIG. 2 illustrates another example of sub-blocks into which one coding block is partitioned.

[0035] Figure 9 Schematically illustrates a multi-transformation scheme according to an embodiment of this document.

[0036] Figure 10 Illustrated is an RST according to an embodiment of this document.

[0037] Figure 11 is a flowchart illustrating the operation of a video decoding device according to an embodiment of this document.

[0038] Figure 12a and Figure 12b It is shown that a variable MIP flag for a first target block according to an embodiment of this document is used to derive an intra prediction mode for a second target block.

[0039] Figure 13 This figure illustrates the arrangement of samples according to the chroma format.

[0040] Figures 14a to 14c It is shown that a variable MIP flag for a corresponding luma block as a first target block according to an embodiment of this document is used to derive an intra prediction mode for a chroma block as a second target block.

[0041] Figure 15 Illustrate the operation of the video encoding device according to the embodiment of this document. DETAILED DESCRIPTION

[0042] Although the present disclosure may be susceptible to various modifications and includes various embodiments, specific embodiments thereof have been shown by way of example in the accompanying drawings and will now be described in detail. However, this is not intended to limit the present disclosure to the specific embodiments disclosed herein. The terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the technical ideas of the present disclosure. Unless the context clearly indicates otherwise, the singular form may include the plural form. Terms such as "including" and "having" are intended to indicate the presence of features, numbers, steps, operations, elements, components, or combinations thereof used in the following description and should therefore not be understood as precluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.

[0043] At the same time, in order to facilitate the description of different characteristic functions, each component in the drawings described herein is illustrated independently, however, it is not intended that each component is implemented by separate hardware or software. For example, any two or more of these components can be combined to form a single component, and any single component can be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of the patent rights of the present disclosure as long as they do not depart from the essence of the present disclosure.

[0044] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. In addition, in the accompanying drawings, the same reference numerals are used for the same components, and repeated description of the same components will be omitted.

[0045] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (Essential Video Coding) standard, the AVS2 standard, etc.).

[0046] In this document, various embodiments related to video / image coding may be provided, and unless otherwise specified, these embodiments may be combined with each other and performed.

[0047] In this document, video can refer to a collection of a series of images over a period of time. Generally, a picture refers to a unit of image that represents a specific time region, and a slice / tile is a unit that constitutes a part of a picture. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles.

[0048] A pixel or picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component. Alternatively, a sample may refer to a pixel value in a spatial domain, or when the pixel value is transformed into a frequency domain, it may refer to a transform coefficient in the frequency domain.

[0049] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region and information related to the region. A unit may include a luminance block and two chrominance (e.g., CB, CR) blocks. Depending on the situation, terms such as unit and block, region, etc. may be used interchangeably. In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0050] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". In addition, "A, B" may mean "A and / or B". In addition, "A / B / C" may mean "at least one of A, B, and / or C". In addition, "A / B / C" may mean "at least one of A, B, and / or C".

[0051] Additionally, in this document, the term "or" should be interpreted as meaning "and / or." For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as meaning "additionally or alternatively."

[0052] In the present disclosure, “at least one of A and B” may mean “only A”, “only B”, or “both A and B”. In addition, in the present disclosure, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted as “at least one of A and B”.

[0053] Furthermore, in the present disclosure, “at least one of A, B, and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.” Furthermore, “at least one of A, B, or C” or “at least one of A, B, and / or C” may mean “at least one of A, B, and C.”

[0054] In addition, the brackets used in this disclosure may indicate "for example." Specifically, when "prediction (intra-frame prediction)" is indicated, it may mean that "intra-frame prediction" is proposed as an example of "prediction." In other words, "prediction" in this disclosure is not limited to "intra-frame prediction," and "intra-frame prediction" is proposed as an example of "prediction." In addition, when "prediction (i.e., intra-frame prediction)" is indicated, it may also mean that "intra-frame prediction" is proposed as an example of "prediction."

[0055] Technical features described separately in one drawing in the present disclosure may be implemented separately or may be implemented simultaneously.

[0056] Figure 1 An example of a video / image coding system to which the present disclosure is applicable is schematically illustrated.

[0057] refer to Figure 1 The video / image coding system may include a first device (source device) and a second device (receiver device). The source device may deliver coded video / image information or data to the receive device in the form of a file or stream via a digital storage medium or a network.

[0058] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0059] The video source can obtain the video / image by capturing, synthesizing, or generating a video / image. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process that generates relevant data.

[0060] An encoding device can encode input video / images. It can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0061] The transmitter can transmit the encoded video / image information or data, output as a bitstream, to a receiver of a receiving device in the form of a file or stream via a digital storage medium or network. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating a media file in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received / extracted bitstream to a decoding device.

[0062] The decoding device may decode a video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.

[0063] The renderer can render the decoded video / image, and the rendered video / image can be displayed on a display.

[0064] Figure 2 is a diagram schematically illustrating a configuration of a video / image encoding device to which the present disclosure can be applied. Hereinafter, a term referred to as a video encoding device may include an image encoding device.

[0065] refer to Figure 2 , the encoding device 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above may be configured by one or more hardware components (e.g., an encoder chipset or processor) according to an embodiment. In addition, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.

[0066] The image partitioner 210 may partition the input image (or picture, or frame) input to the encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a largest coding unit (LCU), the coding units may be recursively divided according to a quadtree, binary tree, ternary tree (QTBTTT) structure. For example, a coding unit may be divided into multiple coding units of increasing depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding process according to the present disclosure may be performed based on the final coding unit without further division. In this case, the largest coding unit may be directly used as the final coding unit based on coding efficiency according to image characteristics. Alternatively, the coding unit may be recursively divided into coding units of increasing depth as needed, so that the optimally sized coding unit may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be split or partitioned from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.

[0067] Depending on the situation, terms such as unit and block, region, etc. may be used instead of each other. In general, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or pixel value, and may represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. A sample may be used as a term corresponding to a pixel or picture element (pel) of a picture (or image).

[0068] The subtractor 231 subtracts the prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) in the encoder 200 can be referred to as the subtractor 231. The predictor can perform prediction on a processing target block (hereinafter referred to as the "current block") and can generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction on a current block or CU basis. As discussed later, in the description of each prediction mode, the predictor can generate various information related to the prediction, such as prediction mode information, and send the generated information to the entropy encoder 240. The information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0069] The intra-frame predictor 222 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference sample can be located in a neighboring area of the current block or in a distant area away from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the configuration. The intra-frame predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.

[0070] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in blocks, subblocks, or sample units based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, the residual signal may not be transmitted. In the case of motion information prediction (motion vector prediction, MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0071] The predictor 220 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction to predict a block and can also apply intra prediction and inter prediction at the same time. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can be based on an intra block copy (IBC) prediction mode or a palette mode to perform prediction on the block. The IBC prediction mode or the palette mode can be used for content image / video coding in games, etc., such as screen content coding (SCC). Although IBC basically performs prediction in the current block, it can be performed similarly to inter prediction in terms of deriving a reference block within the current block. That is, IBC can use at least one of the inter prediction techniques described in this disclosure.

[0072] The prediction signal generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate a reconstructed signal or a residual signal. The transformer 232 can generate a transform coefficient by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT means a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size or can be applied to blocks of variable size other than square blocks.

[0073] The quantizer 233 can quantize the transform coefficients and transmit the quantized transform coefficients to the entropy encoder 240. The entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scanning order and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 can perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 can encode information necessary for video / image reconstruction (e.g., syntax element values, etc.) in addition to the quantized transform coefficients, either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream on a unit basis of a network abstraction layer (NAL). The video / image information may also include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc. In addition, the video / image information may also include general constraint information. In the present disclosure, the video / image information may include information and / or syntax elements sent / signaled from the encoding device to the decoding device. The video / image information may be encoded by the above-mentioned encoding process and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcast network, a communication network, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 and / or a storage device (not shown) that stores the signal may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoder 240.

[0074] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transformation to the quantized transform coefficients via the dequantizer 234 and the inverse transformer 235, a residual signal (residual block or residual sample) can be reconstructed. The adder 255 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222, so that a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) can be generated. When the processing target block has no residual, as in the case of applying skip mode, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current block, and as described later, the generated reconstructed signal can be used for inter-frame prediction of the next picture through filtering.

[0075] Meanwhile, during picture encoding and / or reconstruction, luminance mapping and chroma scaling (LMCS) may be applied.

[0076] The filter 260 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the filter 260 can store the modified reconstructed picture in the memory 270, more specifically, in the DPB of the memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. As discussed later in the description of each filtering method, the filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240. The information about filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0077] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-frame predictor 221. In doing so, the encoding device can avoid prediction mismatch in the encoding device 200 and the decoding device when applying inter-frame prediction, and can also improve coding efficiency.

[0078] The memory 270 DPB can store the modified reconstructed picture so that it can be used as a reference picture in the inter-frame predictor 221. The memory 270 can store motion information of blocks in the current picture from which motion information is derived (or encoded) and / or motion information of blocks in already (or previously) reconstructed pictures. The stored motion information can be sent to the inter-frame predictor 221 to be used as motion information of neighboring blocks or motion information of temporally neighboring blocks. The memory 270 can store reconstructed samples of reconstructed blocks in the current picture and can send the reconstructed samples to the intra-frame predictor 222.

[0079] Figure 3is a diagram schematically illustrating a configuration of a video / image decoding device to which the present disclosure can be applied.

[0080] refer to Figure 3 , the video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. The entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be configured by one or more hardware components (e.g., a decoder chipset or processor) according to an embodiment. In addition, the memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0081] When a bit stream including video / image information is input, the decoding device 300 can be used with the bit stream already in the Figure 2 The image is reconstructed accordingly by the process of processing video / image information in the encoding device. For example, the decoding device 300 can derive a unit / block based on information related to block partitioning obtained from the bitstream. The decoding device 300 can perform decoding by using a processing unit applied in the encoding device. Therefore, the processing unit of decoding can be, for example, a coding unit, which can be partitioned from a coding tree unit or a maximum coding unit along a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. And, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproducer.

[0082] The decoding device 300 can receive the data from the Figure 2The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc. In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter set and / or the general constraint information. In the present disclosure, the information and / or syntax elements sent / received using the signal described later can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on a coding method such as Exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, use the decoded information of the target syntax element and the decoded information of the neighboring and target blocks or the symbol / bin decoded in the previous step to determine a context model, predict the bin generation probability based on the determined context model, and perform arithmetic decoding of the bin to generate the symbol corresponding to each syntax element value. Here, after determining the context model, the CABAC entropy decoding method can update the context model using the symbol / bin information decoded by the context model for the next symbol / bin. Information about prediction from the information decoded in the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual values, i.e., quantized transform coefficients, and related parameter information, for which entropy decoding has been performed in the entropy decoder 310, can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, information about filtering from the information decoded in the entropy decoder 310 can be provided to the filter 350. At the same time, a receiver (not shown) that receives a signal output from the encoding device can further configure the decoding device 300 as an internal / external element, and the receiver can be a component of the entropy decoder 310. At the same time, the decoding device according to the present disclosure can be referred to as a video / image / picture coding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of the dequantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the inter-frame predictor 332, and the intra-frame predictor 331.

[0083] The dequantizer 321 can output the transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement process can be performed based on the order of coefficient scanning performed in the encoding device. The dequantizer 321 can use quantization parameters (e.g., quantization step size information) to dequantize the quantized transform coefficients and obtain the transform coefficients.

[0084] The inverse transformer 322 obtains a residual signal (residual block, residual sample array) by performing inverse transformation on the transformation coefficients.

[0085] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and more specifically, the predictor may determine an intra / inter prediction mode.

[0086] The predictor can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction to predict a block, and can also apply intra prediction and inter prediction at the same time. This can be called combined inter and intra prediction (CIIP). In addition, the predictor can perform intra block copying (IBC) to predict the block. Intra block copying can be used for content image / video coding in games, etc., such as screen content coding (SCC). Although IBC basically performs prediction in the current block, it can be performed similarly to inter prediction in terms of deriving a reference block within the current block. That is, IBC can use at least one of the inter prediction techniques described in this disclosure.

[0087] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the reference samples may be located in an adjacent area of the current block or in an area far away from the current block. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 may determine the prediction mode to be applied to the current block by using the prediction modes applied to the neighboring blocks.

[0088] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in block, sub-block, or sample units based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. The motion information can also include information regarding the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information regarding the prediction can include information indicating the inter-frame prediction mode used for the current block.

[0089] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor 330. When the processing target block has no residual, as in the case of applying the skip mode, the prediction block can be used as the reconstructed block.

[0090] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next processing target block in the current block, and as described later, the generated reconstructed signal may be output through filtering or used for inter-frame prediction of the next picture.

[0091] At the same time, during the picture decoding process, luminance mapping and chroma scaling (LMCS) can be applied.

[0092] The filter 350 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can send the modified reconstructed picture to the memory 360, more specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0093] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store motion information of blocks in the current picture from which motion information has been derived (or decoded) and / or motion information of blocks in already (or previously) reconstructed pictures. The stored motion information can be sent to the inter-frame predictor 260 to be used as motion information of neighboring blocks or motion information of temporally neighboring blocks. The memory 360 can store reconstructed samples of the reconstructed blocks in the current picture and can send the reconstructed samples to the intra-frame predictor 331.

[0094] In this specification, the examples described in the predictor 330, dequantizer 321, inverse transformer 322, and filter 350 of the decoding device 300 may be similarly or correspondingly applied to the predictor 220, dequantizer 234, inverter 235, and filter 260 of the encoding device 200, respectively.

[0095] As described above, prediction is performed to improve compression efficiency when performing video coding. In doing so, a prediction block including prediction samples for the current block as the coding target block can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived identically in the encoding device and the decoding device, and the encoding device can improve image coding efficiency by signaling to the decoding device not the original sample values of the original block itself but information about the residual between the original block and the prediction block (residual information). The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed picture including the reconstructed block.

[0096] Residual information can be generated through a transformation and quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transformation process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients, so that it can send the relevant residual information to the decoding device (via a bitstream) with a signal. Here, the residual information may include value information, position information, transformation technology, transformation kernel, quantization parameter, etc. of the quantized transform coefficients. The decoding device can perform a quantization / dequantization process and derive residual samples (or residual sample blocks) based on the residual information. The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can derive a residual block for reference for inter-frame prediction of the next picture by dequantizing / inverse transforming the quantized transform coefficients, and can generate a reconstructed picture based on the derived residual block.

[0097] Figure 4 The structure of a content streaming system to which the present disclosure is applied is illustrated.

[0098] Furthermore, the content streaming system to which the present disclosure is applied may generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0099] The encoding server is used to compress content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream, and then transmit the bitstream to the streaming server. As another example, if the multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method disclosed herein. The streaming server can also temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0100] The streaming server transmits multimedia data to the user device via a web server based on the user's request. The web server serves as a tool for notifying the user of available services. When the user requests a desired service, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this context, the content streaming system may include a separate control server, which in this case is used to control commands and responses between the corresponding devices in the content streaming system.

[0101] The streaming server can receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to smoothly provide a streaming service, the streaming server can store the bitstream for a predetermined time.

[0102] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system may operate as a distributed server, and in this case, data received by each server may be processed in a distributed manner.

[0103] When performing intra-frame prediction, the correlation between samples can be used, and the difference between the original block and the predicted block, i.e., the residual, can be obtained. The above-mentioned transformation and quantization can be applied to the residual, thereby reducing spatial redundancy. In the following, the encoding method and decoding method using intra-frame prediction are specifically described.

[0104] Intra-frame prediction refers to a method for generating prediction samples for a current block based on reference samples outside the current block within a picture including the current block (hereinafter, the current picture). In this case, the reference samples outside the current block may refer to samples adjacent to the current block. When intra-frame prediction is applied to the current block, neighboring reference samples to be used for intra-frame prediction of the current block may be derived.

[0105] For example, when the size (width×height) of the current block is nW×nH, the neighboring reference samples of the current block may include: a total of 2×nH samples including samples adjacent to the left boundary of the current block and lower-left neighboring samples, a total of 2×nW samples including samples adjacent to the upper boundary of the current block and upper-right neighboring samples, and one upper-left neighboring sample of the current block. Alternatively, the neighboring reference samples of the current block may include multiple rows of upper neighboring samples and multiple columns of left neighboring samples. Furthermore, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block having a size of nW×nH, a total of nW samples adjacent to the lower boundary of the current block, and one lower-right neighboring sample of the current block.

[0106] Here, some of the neighboring reference samples of the current block may not have been decoded or may be unavailable. In this case, the decoding device can configure the neighboring reference samples to be used for prediction by replacing the unavailable samples with available samples. Alternatively, the decoding device can configure the neighboring reference samples to be used for prediction by interpolating available samples.

[0107] When the neighboring reference samples are derived, (i) the prediction sample may be derived based on an average or interpolation of the neighboring reference samples of the current block, or (ii) the prediction sample may be derived based on reference samples existing in a specific (prediction) direction for the prediction sample among the neighboring reference samples of the current block. (i) may be applied when the intra prediction mode is a non-directional mode or a non-angular mode, and (ii) may be applied when the intra prediction mode is a directional mode or an angular mode.

[0108] The intra prediction mode may include a non-directional (or non-angle) intra prediction mode and a directional (or angle) intra prediction mode. For example, in HEVC, an intra prediction mode including two non-directional intra prediction modes and 33 directional intra prediction modes is used. The non-directional intra prediction mode may include a plane intra prediction mode as intra prediction mode 0 and a DC intra prediction mode as intra prediction mode 1, while the directional intra prediction mode may include intra prediction modes 2 to 34. The plane intra prediction mode may be referred to as a plane mode, and the DC intra prediction mode may be referred to as a DC mode.

[0109] Alternatively, to capture arbitrary edge directions present in natural videos, one can use the following Figure 10, the number of directional intra prediction modes is extended from 33 to 65 as shown in . In this case, the intra prediction mode may include two non-directional intra prediction modes and 65 directional intra prediction modes. The non-directional intra prediction mode may include a planar intra prediction mode as intra prediction mode 0 and a DC intra prediction mode as intra prediction mode 1, and the directional intra prediction mode may include intra prediction modes 2 to 66. The extended directional intra prediction mode can be applied to blocks of any size and can be applied to both luminance components and chrominance components. However, the above examples are for illustration, and the embodiments of this document can also be applied when the number of intra prediction modes is different. Intra prediction mode 67 may be further used depending on the situation, and intra prediction mode 67 may refer to a linear model (LM) mode.

[0110] Figure 5 An example of an intra prediction mode to which embodiments of this document are applicable is illustrated.

[0111] refer to Figure 5 , the intra prediction mode can be divided into an intra prediction mode having a horizontal directionality and an intra prediction mode having a vertical directionality based on the intra prediction mode 34 having an upper left diagonal prediction direction. Figure 5 In [ ], H and V indicate horizontal and vertical directionality, respectively, and each of the numbers -32 to 32 indicates a displacement in 1 / 32 units on the sample grid position. Intra-prediction modes 2 to 33 have horizontal directionality, while intra-prediction modes 34 to 66 have vertical directionality. Intra-prediction mode 18 and intra-prediction mode 50 refer to horizontal and vertical intra-prediction modes, respectively. Intra-prediction mode 2 may be referred to as a lower-left diagonal intra-prediction mode, intra-prediction mode 34 may be referred to as an upper-left diagonal intra-prediction mode, and intra-prediction mode 66 may be referred to as an upper-right diagonal intra-prediction mode.

[0112] Matrix-based intra prediction (hereinafter, MIP) may be used as a method for intra prediction. MIP may be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP).

[0113] When MIP is applied to the current block, the prediction sample for the current block can be derived by: i) using the neighboring reference samples that have been subjected to the averaging process, ii) performing a matrix-vector multiplication process, and (iii) further performing a horizontal / vertical interpolation process as necessary. The intra-frame prediction mode used for MIP can be configured differently from the intra-frame prediction mode used in the aforementioned LIP, PDPC, MRL or ISP intra-frame prediction or used in normal intra-frame prediction.

[0114] The intra prediction mode for MIP may be referred to as an affine linear weighted intra prediction mode or a matrix-based intra prediction mode. For example, the matrix and offset used in the matrix-vector multiplication may be configured differently depending on the intra prediction mode for the MIP. Here, the matrix may be referred to as an (affine) weighting matrix, and the offset may be referred to as an (affine) offset vector or an (affine) bias vector. In the present disclosure, the intra prediction mode for MIP may be referred to as a MIP intra prediction mode, an affine linear weighted intra prediction (ALWIP) mode, a matrix weighted intra prediction (MWI) mode, or a matrix-based intra prediction mode.

[0115] To predict samples of a rectangular block with width (W) and height (H), MIP uses as input one H line of reconstructed samples adjacent to the left boundary of the block and one W line of reconstructed samples adjacent to the top boundary of the block. When no reconstructed samples are available, reference samples can be generated by the interpolation method applied to general intra prediction.

[0116] Figure 6 The figure shows the process of generating prediction samples based on MIP according to an example. Figure 6 The MIP process is described as follows.

[0117] 1. Averaging process

[0118] Among the boundary samples, four samples for the case of W=H=4 and eight samples for any other cases are extracted through an averaging process.

[0119] 2. Matrix-vector multiplication process

[0120] A matrix-vector multiplication is performed with the average sample as input followed by the addition of an offset. Through this operation, a reduced prediction sample for the subsampled set of samples in the original block can be derived.

[0121] 3. (Linear) interpolation process

[0122] The prediction samples at the remaining positions are generated from the prediction samples of the subsampled sample set by linear interpolation, which is a single-step linear interpolation in each direction.

[0123] The matrices and offset vectors necessary to generate a prediction block or prediction sample may be selected from three sets S0, S1 and S2 for matrices.

[0124] Set S0 can include: 16 matrices A0 i , i∈{0,…,15}, and each matrix can include 16 rows, four columns; and 16 offset vectors b0 i, i∈{0,…,15}. The matrices and offset vectors of set S0 may be used for a 4×4 block. In another example, set S0 may include 18 matrices.

[0125] Set S1 can include: eight matrices A1 i , i∈{0,…,7}, and each matrix may include 16 rows, eight columns; and eight offset vectors b1 i , i∈{0,…,7}. In another example, set S1 may include six matrices. The matrices and offset vectors of set S1 may be used for 4×8, 8×4, and 8×8 blocks. Alternatively, the matrices and offset vectors of set S1 may be used for 4×H or W×4 blocks.

[0126] Finally, set S2 may include: six matrices A2 i , i∈{0,…,5}, and each matrix can include 64 rows, eight columns; and six offset vectors b2 i , i∈{0,…,5}. The matrices and offset vectors of set S2 or a portion thereof can be used for any block of a different size that sets S0 and S1 are not applicable to. For example, the matrices and offset vectors of set S2 can be used for operations on blocks with a height and width of 8 or more.

[0127] The total number of multiplications required for the calculation of the matrix-vector product is always less than or equal to 4 × W × H. That is, in MIP mode, up to four multiplications are required per sample.

[0128] In addition, the current block can be divided into vertical or horizontal sub-partitions, and intra prediction can be performed based on the same intra prediction mode, where neighboring reference samples can be derived and used in units of sub-partitions. That is, in this case, the intra prediction mode for the current block can be applied to the sub-partitions in the same manner, and neighboring reference samples can be derived and used in units of sub-partitions, thereby improving the performance of intra prediction in some cases. This prediction method can be referred to as intra sub-partitioning (ISP) or ISP-based intra prediction.

[0129] Intra-frame sub-partition (ISP) coding refers to performing intra-frame prediction coding by dividing the block to be currently coded in the horizontal direction or the vertical direction. In this case, a reconstructed block can be generated by performing encoding / decoding in units of partitioned blocks, and the reconstructed block can be used as a reference block for the next partitioned block. According to an example, in ISP coding, one coding block can be divided into two or four sub-blocks and then coded, and in ISP, one sub-block is subjected to intra-frame prediction with reference to the value of the reconstructed pixel of the left adjacent sub-block or the upper adjacent sub-block. As used herein, "coding" can be used as a concept that includes both coding performed by an encoding device and decoding performed by a decoding device.

[0130] Table 1 shows the number of sub-blocks partitioned according to the size of a block when ISP is applied, and the sub-partitions partitioned according to the ISP may be referred to as transform blocks (TUs).

[0131] [Table 1]

[0132] Block size (CU) Number of partitions 4×4 Unavailable 4×8、8×4 2 Any other circumstances 4

[0133] ISP partitions a block predicted by luma intra into two or four sub-partitions vertically or horizontally depending on the size of the block. For example, the minimum size of a block to which ISP applies is 4×8 or 8×4. When the block size is larger than 4×8 or 8×4, the block is partitioned into four sub-partitions.

[0134] Figure 7 and Figure 8 The diagram shows an example of a coding block being divided into sub-blocks. Specifically, Figure 7 An example of dividing a coding block (width (W) × height (H)) into a 4×8 block or an 8×4 block is illustrated, and Figure 8 An example of partitioning a coding block that is not a 4×8 block, an 8×4 block, or a 4×4 block is illustrated.

[0135] When ISP is applied, subblocks can be coded sequentially according to the partition type, such as horizontally, vertically, from left to right, or from top to bottom, where one subblock can be subjected to inverse transform, intra prediction, and reconstruction, after which the next subblock can be coded. For the leftmost or topmost subblock, as in the conventional intra prediction method, the reconstructed pixels of the already coded coding block are used for reference. In addition, when each side of the subsequent internal subblock is not adjacent to the previous subblock, as in the conventional intra prediction method, the reconstructed pixels of the already coded adjacent coding block are used for reference to derive the reference pixels adjacent to the side.

[0136] In ISP coding mode, all sub-blocks can be coded with the same intra prediction mode, and a flag indicating whether ISP coding is used and a flag indicating the direction in which the partitioning is performed (horizontal or vertical) can be signaled. Figure 7 and Figure 8 As shown, the number of sub-blocks can be adjusted to 2 or 4 depending on the block shape, and when the size (width×height) of one sub-block is less than 16, splitting into the corresponding number of sub-blocks may not be allowed or a restriction may be set not to apply ISP coding.

[0137] In the ISP prediction mode, one coding unit is partitioned into two or four partition blocks, ie, sub-blocks, to be predicted, and the same intra prediction mode is applied to the two or four partition blocks.

[0138] Figure 9The multiple transformation technology according to an embodiment of the present disclosure is schematically illustrated.

[0139] refer to Figure 9 , the converter can correspond to the aforementioned Figure 2 The converter in the encoding device, and the inverse converter may correspond to the aforementioned Figure 2 The inverse transformer in the encoding device, or corresponding to Figure 3 An inverse transformer in a decoding device.

[0140] The transformer may derive (primary) transform coefficients by performing a primary transform based on the residual samples (residual sample array) in the residual block (S910). This primary transform may be referred to as a core transform. Herein, the primary transform may be based on multiple transform selection (MTS), and when multiple transforms are applied as the primary transform, it may be referred to as a multi-core transform.

[0141] Multi-core transform may refer to a method for performing transforms using discrete cosine transform (DCT) type 2 and discrete sine transform (DST) type 7, DCT type 8, and / or DST type 1 in addition. In other words, multi-core transform may refer to a transform method that transforms a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. Herein, the primary transform coefficients may be referred to as temporary transform coefficients from the perspective of the transformer.

[0142] In other words, when a conventional transform method is applied, transform coefficients can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2. In contrast, when a multi-core transform is applied, transform coefficients (or primary transform coefficients) can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. In this document, DCT type 2, DST type 7, DCT type 8, and DST type 1 may be referred to as transform types, transform kernels, or transform cores. These DCT / DST transform types may be defined based on basis functions.

[0143] When performing multi-core transformation, a vertical transform kernel and a horizontal transform kernel for the target block can be selected from the transform kernels, a vertical transform can be performed on the target block based on the vertical transform kernel, and a horizontal transform can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform can indicate the transform of the horizontal component of the target block, and the vertical transform can indicate the transform of the vertical component of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target (CU or sub-block) including the residual block.

[0144] In addition, according to an example, if the primary transform is performed by applying MTS, the mapping relationship of the transform kernel can be set by setting a specific basis function to a predetermined value and combining the basis functions to be applied in the vertical transform or the horizontal transform. For example, when the horizontal transform kernel is represented as trTypeHor and the vertical transform kernel is represented as trTypeVer, trTypeHor or trTypeVer with a value of 0 can be set to DCT2, trTypeHor or trTypeVer with a value of 1 can be set to DST-7, and trTypeHor or trTypeVer with a value of 2 can be set to DCT-8.

[0145] In this case, MTS index information may be encoded and signaled to a decoding device to indicate any one of a plurality of transform kernel sets. For example, an MTS index of 0 may indicate that both trTypeHor and trTypeVer values are 0, an MTS index of 1 may indicate that both trTypeHor and trTypeVer values are 1, an MTS index of 2 may indicate that both trTypeHor and trTypeVer values are 1, an MTS index of 3 may indicate that both trTypeHor and trTypeVer values are 1 and 2, and an MTS index of 4 may indicate that both trTypeHor and trTypeVer values are 2.

[0146] In one example, the transformation kernel set according to the MTS index information is shown in the following table.

[0147] [Table 2]

[0148] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2

[0149] The transformer may perform a secondary transform based on the (primary) transform coefficients to derive modified (secondary) transform coefficients (S920). The primary transform is a transform from the spatial domain to the frequency domain, and the secondary transform refers to transforming into a more compressed expression using the correlation between the (primary) transform coefficients. The secondary transform may include a non-separable transform. In this case, the secondary transform may be referred to as a non-separable secondary transform (NSST) or a pattern-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform may represent a transform that performs a secondary transform on the (primary) transform coefficients derived by the primary transform based on a non-separable transform matrix to generate modified transform coefficients (or secondary transform coefficients) for the residual signal. Here, the vertical transform and the horizontal transform may not be applied separately (or the vertical transform and the horizontal transform may not be applied independently to the) (primary) transform coefficients, but the transform may be applied at one time based on the non-separable transform matrix. In other words, the non-separable secondary transform may refer to a transform method that is not applied separately to (primary) transform coefficients in the vertical and horizontal directions, and that rearranges a two-dimensional signal (transform coefficient) into a one-dimensional signal in a specific predetermined direction (e.g., a row-first direction or a column-first direction) and then generates modified transform coefficients (or secondary transform coefficients) based on the non-separable transform matrix. For example, according to row priority, MxN blocks are arranged in a line in the order of row 1, row 2, ..., row N. According to column priority, MxN blocks are arranged in a line in the order of column 1, column 2, ..., column M. NSST may be applied to the upper left region of a block configured with (primary) transform coefficients (hereinafter referred to as a transform coefficient block). For example, when the width W and height H of the transform coefficient block are both 8 or greater, 8×8 NSST may be applied to the upper left 8×8 region of the transform coefficient block. In addition, when the width (W) and height (H) of the transform coefficient block are both 4 or greater, when the width (W) or height (H) of the transform coefficient block is less than 8, 4×4 NSST may be applied to the upper left min(8,W)×min(8,H) area of the transform coefficient block. However, embodiments are not limited thereto, and for example, even if a condition is satisfied that only the width W or the height H of the transform coefficient block is 4 or greater, 4×4 NSST may be applied to the upper left min(8,W)×min(8,H) area of the transform coefficient block.

[0150] Here, in order to select the transform kernel, two inseparable secondary transform kernels can be configured per transform set for inseparable secondary transforms for both 8×8 transform and 4×4 transform, and there can be four transform sets. That is, four transform sets can be configured for 8×8 transform, and four transform sets can be configured for 4×4 transform. In this case, each of the four transform sets for 8×8 transform can include two 8×8 transform kernels, and each of the four transform sets for 4×4 transform can include two 4×4 transform kernels.

[0151] However, as the size of the transform (ie, the size of the region to which the transform is applied) may be other than 8×8 or 4×4, for example, the number of sets may be n, and the number of transform kernels in each set may be k.

[0152] The transform set may be referred to as an NSST set or a LFNST set. A specific set among the transform sets may be selected, for example, based on the intra prediction mode of the current block (CU or subblock). A low-frequency non-separable transform (LFNST) may be an example of a reduced non-separable transform, which will be described later and represents a non-separable transform for low-frequency components.

[0153] According to an example, four transform sets according to intra prediction modes may be mapped, for example, as shown in the following table.

[0154] [Table 3]

[0155] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0<=predModeIntra<=1 0 2<=predModeIntra<=12 1 13<=predModeIntra<=23 2 24<=predModeIntra<=44 3 45<=predModeIntra<=55 2 56<=predModeIntra<=80 1

[0156] As shown in Table 3, any one of four transform sets, ie, lfnstTrSetIdx, may be mapped to any one of four indexes (ie, 0 to 3) according to the intra prediction mode.

[0157] When it is determined that a specific set is used for an inseparable transform, one of the k transform cores in the specific set can be selected by an inseparable secondary transform index. The encoding device can derive an inseparable secondary transform index indicating a specific transform core based on a rate-distortion (RD) check, and can send the inseparable secondary transform index to the decoding device with a signal. The decoding device can select one of the k transform cores in the specific set based on the inseparable secondary transform index. For example, an lfnst index value 0 can refer to a first inseparable secondary transform core, an lfnst index value 1 can refer to a second inseparable secondary transform core, and an lfnst index value 2 can refer to a third inseparable secondary transform core. Alternatively, an lfnst index value 0 can indicate that the first inseparable secondary transform is not applied to the target block, and lfnst index values 1 to 3 can indicate three transform cores.

[0158] The transformer can perform a non-separable secondary transform based on the selected transform kernel and can obtain modified (secondary) transform coefficients. As described above, the modified transform coefficients can be derived as transform coefficients quantized by a quantizer, and can be encoded and sent to a decoding device with a signal, and transmitted to a dequantizer / inverse transformer in the encoding device.

[0159] At the same time, as described above, if the secondary transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform can be derived as transform coefficients quantized by the quantizer as described above, and can be encoded and sent to the decoding device with a signal, and transmitted to the dequantizer / inverse transformer in the encoding device.

[0160] The inverse transformer can perform a series of processes in the reverse order of the order in which the series of processes are performed in the above-mentioned transformer. The inverse transformer can receive (dequantized) transformer coefficients and derive (primary) transform coefficients by performing a secondary (inverse) transform (S950), and can obtain a residual block (residual sample) by performing a primary (inverse) transform on the (primary) transform coefficients (S960). In this regard, from the perspective of the inverse transformer, the primary transform coefficients can be referred to as modified transform coefficients. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.

[0161] The decoding device may further include a secondary inverse transform application determiner (or an element for determining whether to apply a secondary inverse transform) and a secondary inverse transform determiner (or an element for determining a secondary inverse transform). The secondary inverse transform application determiner may determine whether to apply a secondary inverse transform. For example, the secondary inverse transform may be NSST, RST, or LFNST, and the secondary inverse transform application determiner may determine whether to apply a secondary inverse transform based on a secondary transform flag obtained by parsing the bitstream. In another example, the secondary inverse transform application determiner may determine whether to apply a secondary inverse transform based on a transform coefficient of a residual block.

[0162] The secondary inverse transform determiner may determine the secondary inverse transform. In this case, the secondary inverse transform determiner may determine the secondary inverse transform to be applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra-frame prediction mode. In an embodiment, the secondary transform determination method may be determined depending on the primary transform determination method. Various combinations of the primary transform and the secondary transform may be determined according to the intra-frame prediction mode. In addition, in an example, the secondary inverse transform determiner may determine the area to which the secondary inverse transform is applied based on the size of the current block.

[0163] Meanwhile, as described above, if the secondary (inverse) transform is omitted, the (dequantized) transform coefficients may be received, the primary (separable) inverse transform may be performed, and the residual block (residual samples) may be obtained. As described above, the encoding device and the decoding device may generate a reconstructed block based on the residual block and the prediction block, and may generate a reconstructed picture based on the reconstructed block.

[0164] In this document, a reduced secondary transform (RST) in which the size of the transform matrix (kernel) is reduced can be applied within the concept of NSST to reduce the amount of computation and memory required for the non-separable secondary transform. RST is typically performed in low-frequency regions containing non-zero coefficients in a transform block, and therefore can be referred to as a low-frequency non-separable transform (LFNST). The transform index can be referred to as an LFNST index.

[0165] In this specification, LFNST may refer to a transform performed on residual samples for a target block based on a transform matrix with a reduced size. When performing a downscaled transform, the amount of computation required for the transform can be reduced due to the reduced size of the transform matrix. In other words, LFNST can be used to address the computational complexity issues that arise with non-separable transforms or transforms of large blocks.

[0166] When performing a secondary inverse transform based on LFNST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include: an inverse reduced secondary transformer that derives modified transform coefficients based on an inverse RST of the transform coefficients; and an inverse primary transformer that derives residual samples for the target block based on an inverse primary transform of the modified transform coefficients. The inverse primary transform refers to an inverse transform of the primary transform applied to the residual. In this document, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying a transform.

[0167] Figure 10 is a diagram illustrating an RST according to an embodiment of the present disclosure.

[0168] In this disclosure, a “target block” may refer to a current block to be encoded, a residual block, or a transform block.

[0169] In the RST according to the example, an N-dimensional vector can be mapped to an R-dimensional vector located in another space, so that a reduced transform matrix can be determined, where R is less than N. N can refer to the square of the length of the side of the block to which the transform is applied, or the total number of transform coefficients corresponding to the block to which the transform is applied, and the reduction factor can refer to an R / N value. The reduction factor can be referred to as a reduction factor, a shrinkage factor, a simplification factor, a simple factor, or various other terms. In addition, R can be referred to as a reduction coefficient, but depending on the situation, the reduction factor can refer to R. In addition, depending on the situation, the reduction factor can refer to an N / R value.

[0170] In this example, the reduction factor or the reduction coefficient may be signaled through the bitstream, but the example is not limited thereto. For example, a predefined value for the reduction factor or the reduction coefficient may be stored in each of the encoding device 200 and the decoding device 300, and in this case, the reduction factor or the reduction coefficient may not be signaled separately.

[0171] The size of the reduced transform matrix according to an example may be R×N, which is smaller than N×N (the size of a conventional transform matrix), and may be defined as in Equation 1 below.

[0172] [Formula 1]

[0173]

[0174] Figure 10 The matrix T in the reduced transform block shown in (a) may mean the matrix T of Equation 1. RxN .like Figure 10 As shown in (a), when the reduced transformation matrix T RxN When multiplied by the residual samples of the target block, transform coefficients for the target block can be derived.

[0175] In an example, if the size of the block to which the transform is applied is 8x8 and R=16 (ie, R / N=16 / 64=1 / 4), then the transform according to Figure 10 The RST of (a) is expressed as a matrix operation as shown in the following Equation 2. In this case, it is possible to reduce the memory and multiplication calculation to about 1 / 4 by the reduction factor.

[0176] In the present disclosure, a matrix operation may be understood as an operation of obtaining a column vector by multiplying a column vector by a matrix provided on the left side of the column vector.

[0177] [Formula 2]

[0178]

[0179] In formula 2, r1 to r 64 ∠ ... i , and export c i The process can be shown as formula 3.

[0180] [Formula 3]

[0181]

[0182] As a result of the calculation of Equation 3, transform coefficients c1 to c2 for the target block can be derived. R That is, when R=16, the transform coefficients c1 to c 16 Although 64 (N) transform coefficients are derived for the target block, if a regular transform is applied instead of RST and a transform matrix of 64x64 (NxN) size is multiplied by a residual sample of 64x1 (Nx1) size, only 16 (R) transform coefficients are derived for the target block because RST is applied. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data transmitted from the encoding device 200 to the decoding device 300 is reduced, which can improve the efficiency of transmission between the encoding device 200 and the decoding device 300.

[0183] When considering the size of the transformation matrix, the size of the conventional transformation matrix is 64×64 (N×N), but the size of the reduced transformation matrix is reduced to 16×64 (R×N). Therefore, when performing LFNST, the memory usage can be reduced by the R / N ratio compared to the case of performing the conventional transformation. In addition, when compared with the number of multiplication calculations (N×N) when using the conventional transformation matrix, the number of multiplication calculations (R×N) can be reduced by the R / N ratio when using the reduced transformation matrix.

[0184] In an example, the transformer 232 of the encoding device 200 may derive transform coefficients for the target block by performing a primary transform and an RST-based secondary transform on the residual samples for the target block. These transform coefficients may be passed to the inverse transformer of the decoding device 300, and the inverse transformer 322 of the decoding device 300 may derive modified transform coefficients based on an inverse reduced secondary transform (RST) on the transform coefficients, and may derive residual samples for the target block based on the inverse primary transform on the modified transform coefficients.

[0185] According to the example, the inverse RST matrix T NxR The size of is NxR which is smaller than the size of the conventional inverse transform matrix NxN and is the same as the reduced transform matrix T shown in Equation 4. RxN In a transposed relationship.

[0186] Figure 10 The matrix T in the reduced inverse transform block shown in (b) t It can be referred to as the inverse RST matrix T RxN T (The superscript T means transpose). Figure 10 (b) shows the inverse RST matrix T RxN TWhen multiplied by the transform coefficients for the target block, modified transform coefficients for the target block or residual samples for the current block can be derived. RxN T Expressed as (T RxN ) T NxR .

[0187] More specifically, when the inverse RST is used as the secondary inverse transform, when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. In addition, the inverse RST can be used as the inverse primary transform, and in this case, when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, residual samples of the target block can be derived.

[0188] In an example, if the size of the block to which the inverse transform is applied is 8x8 and R=16 (ie, R / N=16 / 64=1 / 4), the Figure 10 The RST of (b) is expressed as a matrix operation as shown in the following Equation 4.

[0189] [Formula 4]

[0190]

[0191] In formula 4, c1 to c 16 As a result of the calculation of Equation 4, r representing the modified transform coefficient of the target block or the residual sample of the target block can be derived. i , and derive r i The process can be shown as formula 5.

[0192] [Formula 5]

[0193]

[0194] As a result of the calculation of Equation 5, r1 to r2 representing the modified transform coefficients of the target block or the residual samples of the target block can be derived. N . From the perspective of the size of the inverse transform matrix, the size of the conventional inverse transform matrix is 64×64 (N×N), but the size of the inverse reduced transform matrix is reduced to 64×16 (R×N), so the storage usage in the case of performing inverse RST can be reduced by the R / N ratio compared to the case of performing the conventional inverse transform. In addition, when compared with the number of multiplication calculations N×N in the case of using the conventional inverse transform matrix, the use of the inverse reduced transform matrix can reduce the number of multiplication calculations (N×R) by the R / N ratio.

[0195] According to an embodiment of the present disclosure, for the transformation in the encoding process, only 48 pieces of data can be selected, and a maximum 16×48 transform kernel matrix can be applied thereto, instead of applying a 16×64 transform kernel matrix to 64 pieces of data forming an 8×8 area. Here, "maximum" means that m has a maximum value of 16 in the m×48 transform kernel matrix for generating m coefficients. That is, when RST is performed by applying an m×48 transform kernel matrix (m≤16) to an 8×8 area, 48 pieces of data are input and m coefficients are generated. When m is 16, 48 pieces of data are input and 16 coefficients are generated. That is, assuming that 48 pieces of data form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied in sequence, thereby generating a 16×1 vector. Here, the 48 pieces of data forming an 8×8 area can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on the 48 pieces of data constituting the area other than the lower right 4×4 area in the 8×8 area. Here, when the matrix operation is performed by applying a maximum 16×48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients can be arranged in the upper left 4×4 region according to the scanning order, and the upper right 4×4 region and the lower left 4×4 region can be filled with zeros.

[0196] For the inverse transform in the decoding process, a transposed matrix of the aforementioned transform kernel matrix may be used. That is, when inverse RST or LFNST is performed in the inverse transform process performed by the decoding device, input coefficient data to which inverse RST is applied is arranged in a one-dimensional vector according to a predetermined arrangement order, and modified coefficient vectors obtained by multiplying the one-dimensional vector by the corresponding inverse RST matrix on the left side of the one-dimensional vector may be arranged in a two-dimensional block according to the predetermined arrangement order.

[0197] In summary, during the transform process, when RST or LFNST is applied to an 8×8 region, the 48 transform coefficients in the upper left, upper right, and lower left regions of the 8×8 region, excluding the lower right region, are subjected to a matrix operation with the 16×48 transform kernel matrix. For the matrix operation, the 48 transform coefficients are input as a one-dimensional array. When the matrix operation is performed, 16 modified transform coefficients are derived and arranged in the upper left region of the 8×8 region.

[0198] On the contrary, in the inverse transform process, when the inverse RST or LFNST is applied to the 8×8 area, the 16 transform coefficients corresponding to the upper left area of the 8×8 area among the transform coefficients in the 8×8 area can be input in a one-dimensional array according to the scanning order, and can undergo a matrix operation with a 48×16 transform kernel matrix. That is, the matrix operation can be expressed as (48×16 matrix)*(16×1 transform coefficient vector)=(48×1 modified transform coefficient vector). Here, the n×1 vector can be interpreted as having the same meaning as the n×1 matrix, and can therefore be expressed as an n×1 column vector. In addition, * represents matrix multiplication. When the matrix operation is performed, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper left area, the upper right area, and the lower left area of the 8×8 area except the lower right area.

[0199] According to an example, syntax elements for a MIP may be signaled in a coding unit syntax table as in Table 4.

[0200] [Table 4]

[0201]

[0202]

[0203]

[0204] intra_mip_flag[x0][y0] signaled in the coding unit syntax table of Table 4 is flag information indicating whether the MIP intra mode is applied to the coding unit.

[0205] In addition to Table 4, as in Table 5, intra_mip_flag[x0][y0] is referenced multiple times in the specification text, and in particular, values of intra_mip_flag for positions other than (x0, y0), such as intra_mip_flag[xCb+cbWidth / 2][yCb+cbHeight / 2], are also referenced.

[0206] Here, x0 and y0 indicate the x-coordinate and y-coordinate based on the luma picture, respectively, where the x-coordinate increases from left to right in luma sample units when the leftmost position of the luma picture is defined as 0, and the y-coordinate increases from top to bottom in luma sample units when the topmost position of the luma picture is defined as 0. The x-coordinate and the y-coordinate may be expressed in a two-dimensional coordinate form such as (x, y).

[0207] However, intra_mip_flag[x0][y0] may be considered valid only for the position (x0, y0), and (x0, y0) corresponds to the top-left position of the coding unit (CU) in which intra_mip_flag[x0][y0] is signaled. That is, for each coding unit, intra_mip_flag[x0][y0] may be considered signaled only for (x0, y0) as the top-left position of the representative position in the coding unit.

[0208] In Table 5, since intra_mip_flag values for positions other than (x0, y0) which is the upper left position in the coding unit are referenced (underlined), information of intra_mip_flag values for positions other than the upper left position needs to be filled.

[0209] [Table 5]

[0210]

[0211]

[0212]

[0213]

[0214]

[0215] As in Tables 4 and 5, intra_mip_flag is defined as a syntax element, and as in the semantics of intra_mip_flag illustrated in Table 5 (7.4.11.5 Coding unit semantics), intra_mip_flag is inferred to be 0 when not present.

[0216] The value of intra_mip_flag may be used when deriving intra prediction modes (8.4.1 and 8.4.2). For example, as described in 8.4.2, the value of intra_mip_flag for a neighboring block or a specific position in a neighboring block may be used (intra_mip_flag[xNbX][yNbX] equal to 1).

[0217] In addition, as described in 8.3.4, when deriving the intra prediction mode for the chroma block, the value of the intra_mip_flag corresponding to the position of the luma block can be used (- the chroma intra prediction mode IntraPredModeC[xCb][yCb] is set equal to IntraPredModeY[xCb+cbWidth / 2][yCb+cbHeight / 2]).

[0218] In addition, the value of intra_mip_flag can be used not only in intra prediction but also in the transform process (8.7.4.1 - when intra_mip_flag[xTbY][yTbY] is equal to 1 and cIdx is equal to 0, predModeIntra is set equal to INTRA_PLANAR), and the value of intra_mip_flag for neighboring blocks can be used in the process of deriving the context index of the target block to be coded (Table 132: ctxInc specification using left and above syntax elements).

[0219] Therefore, considering intra_mip_flag as a two-dimensional array variable, ambiguity may arise in the method of filling the values of the intra_mip_flag array variable in positions other than (x0, y0). That is, since the semantics stipulate that intra_mip_flag[x0][y0] is inferred to be 0 when it does not exist, filling any non-zero value in the unfilled positions in the intra_mip_flag array may be ambiguous, that is, undefined.

[0220] Therefore, according to an example, it may be proposed to define a separate two-dimensional array, such as IntraLumaMipFlag[x][y] as shown in Table 6, and utilize this two-dimensional array in the subsequent image coding process.

[0221] [Table 6]

[0222]

[0223] In Table 6, (x0, y0) indicates the upper left position of the currently coded coding unit (CU), and cbWidth and cbHeight indicate the width and height of the coding unit, respectively. In addition, "x = x0..x0+cbWidth-1" means that the x coordinate value changes from x0 to x0+cbWidth-1, and "y = y0..y0+cbHeight-1" means that the y coordinate value changes from y0 to y0+cbHeight-1. Therefore, in the IntraLumaMipFlag[x][y] array in Table 6, the area corresponding to the coding unit is filled with the intra_mip_flag[x0][y0] value.

[0224] According to an example, all parts (underlined) that reference intra_mip_flag information in the current VVC specification text may be replaced or modified with the IntraLumaMipFlag variable as shown in Table 7.

[0225] [Table 7]

[0226]

[0227]

[0228]

[0229]

[0230]

[0231] In Table 7, the case where the IntraLumaMipFlag information of the upper left position in the coding unit is definitely referenced can be described using the existing intra_mip_flag as it is. Table 8 shows only a portion where the existing intra_mip_flag as extracted from Table 7 can be used.

[0232] [Table 8]

[0233]

[0234] In Table 4, the same problem may occur for intra_subpartitions_mode_flag other than intra_mip_flag in the current VVC specification text. That is, intra_subpartitions_mode_flag information of positions other than the upper left position (x0, y0) in the coding unit is referenced, which is shown in Table 9 (indicated by underline).

[0235] Therefore, it can be configured to define a two-dimensional array variable of IntraSubPartitionsModeFlag[x][y] in a similar manner to the IntraLumaMipFlag variable to fill all positions in the coding unit with the necessary information, and then reference the appropriate value for any position. In other words, all positions in the coding unit can be filled with the corresponding intra_subpartitions_mode_flag[x0][y0] value.

[0236] Table 10 shows that the portion (underline) referencing intra_subpartitions_mode_flag is replaced with the IntraSubPartitionsModeFlag variable using the IntraSubPartitionsModeFlag variable. In the case of IntraSubPartitionsModeFlag, when referencing information about the upper left position in a coding unit, intra_subpartitions_mode_flag may be configured to be referenced as is.

[0237] According to another example, the IntraLumaMipFlag variable or the IntraSubPartitionModeFlag variable presented in this embodiment can have different variable names. In other words, these variables can be expressed or referred to as different variables. For example, the IntraLumaMipFlag variable can be referred to as the MipFlag variable.

[0238] [Table 9]

[0239]

[0240]

[0241] [Table 10]

[0242]

[0243] The following drawings are provided to describe specific examples of the present disclosure. Since specific terms used for devices or specific terms used for signals / messages / fields shown in the drawings are provided for illustration, the technical features of the present disclosure are not limited to the specific terms used in the following drawings.

[0244] Figure 11 is a flowchart illustrating the operation of a video decoding apparatus according to an embodiment of the present disclosure.

[0245] Figure 11 Each process disclosed in the reference is based on Figures 5 to 10 Therefore, with reference to Figure 3 and Figures 5 to 10 Descriptions of those overlapping specific details will be omitted or will be made schematically.

[0246] According to an embodiment, the decoding device 300 can receive information about the intra-frame prediction mode, residual information, etc. from the bitstream, and can receive, for example, image information including intra-frame prediction type information, which intra-frame prediction type information includes an intra-frame MIP syntax element (intra_mip_flag) for the first target block (S1110).

[0247] Specifically, the decoding device 300 can decode information about the quantized transform coefficients for the current block from the bitstream, and can derive the quantized transform coefficients for the target block based on the information about the quantized transform coefficients for the current block. The information about the quantized transform coefficients for the target block may be included in a sequence parameter set (SPS) or a slice header, and may include information about whether to apply RST, information about a reduction factor, information about a minimum transform size for applying RST, information about a maximum transform size for applying RST, an inverse RST size, and at least one of information about a transform index indicating any one transform kernel matrix included in the transform set.

[0248] The decoding device may further receive information about the intra prediction mode for the current block and information about whether ISP is applied to the current block. The decoding device may receive and parse flag information indicating whether ISP coding or ISP mode is applied, thereby deriving whether the current block is partitioned into a predetermined number of sub-partition transform blocks. Here, the current block may be a coding block. In addition, the decoding device may derive the size and number of the partitioned sub-partition blocks using flag information indicating the direction in which the current block is partitioned.

[0249] The decoding apparatus 300 may decode the intra MIP syntax element for the first target block, thereby deriving a value of the intra MIP syntax element ( S1120 ).

[0250] The decoding apparatus may set a variable MIP flag for a preset specific region that is the same as the region of the first target block based on a value of the intra MIP syntax element ( S1130 ).

[0251] As described with reference to Table 6, the variable MIP flag (IntraLumaMipFlag[x][y]) may be set in the semantics for the intra-MIP syntax element and may be utilized in subsequent processes for decoding the image.

[0252] The specific area may be set to be the same as the area where the sample is located in the first target block (x=x0..x0+cbWidth-1 and y=y0..y0+cbHeight-1). Here, cbWidth and cbHeight indicate the width and height of the first target block.

[0253] The variable MIP flag may be set for a specific region based on the tree type of the first target block not being dual-tree chroma. That is, the variable MIP flag may be set to the value of the received intra MIP syntax element for a specific region only when the tree type of the first target block is single-tree or dual-tree luma but not dual-tree chroma.

[0254] The intra-MIP syntax element intra_mip_flag is signaled only when the tree type of the corresponding coding block is single-tree or dual-tree luma. It is not signaled and is therefore inferred to be 0 when the tree type is dual-tree chroma. Therefore, when the tree type of the first target block is dual-tree chroma, the value of intra_mip_flag is 0. In intra prediction or transform, when the intra prediction mode for a chroma block is derived, the intra prediction mode for the luma block is adopted, in which case the newly set variable MIP flag can be used. When the conditions for setting the variable MIP flag do not include the condition for setting the variable MIP flag only in the "single-tree and dual-tree luma" cases, the variable MIP flag can also be set in dual-tree chroma. In this case, since the value of intra_mip_flag is 0, the variable MIP flag does not include information about the luma component, and therefore luma information cannot be used when deriving the intra prediction mode for the chroma block. To prevent this problem, the variable MIP flag can be set only when the corresponding coding block (i.e., the first target block) is not dual-tree chroma.

[0255] The decoding apparatus may derive an intra prediction mode for the second target block based on the variable MIP flag of the first target block ( S1140 ).

[0256] Figure 12a and Figure 12b It is shown that the variable MIP flag for the first target block is used to derive the intra prediction mode for the second target block.

[0257] like Figure 12a As shown, the first target block may be a left neighboring block of the second target block, and the specific area may include a sample position of (xCb-1, yCb+cbHeight-1). In this case, (xCb, yCb) is the position of the upper left sample of the second target block, and cbHeight indicates the height of the second target block.

[0258] That is, a candidate intra-frame prediction mode for the second target block can be derived based on the variable MIP flag of the sample position (i.e., a specific area) of (xCb-1, yCb+cbHeight-1) included in the first target block, and the intra-frame prediction mode for the second target block can be derived based on the candidate intra-frame prediction mode.

[0259] In another example, Figure 12b As shown, the first target block may be an upper neighboring block of the second target block, and the specific area may include a sample position of (xCb+cbWidth-1, yCb-1). In this case, (xCb, yCb) is the position of the upper left sample of the second target block, and cbWidth indicates the width of the second target block.

[0260] That is, a candidate intra-frame prediction mode for the second target block can be derived based on the variable MIP flag of the sample position (i.e., a specific area) of (xCb+cbWidth-1, yCb-1) included in the first target block, and the intra-frame prediction mode for the second target block can be derived based on the candidate intra-frame prediction mode.

[0261] In yet another example, the second target block may include a chroma block, and the first target block may be a luma block associated with the chroma block.As described above, the intra prediction mode for the chroma block may be derived based on the variable MIP flag value for the luma block.

[0262] In this document, a picture / image may include a luma component array and, in some cases, may further include two chroma component (cb, cr) arrays. That is, one pixel of a picture / image may include a luma sample and a chroma sample (cb, cr).

[0263] The color format may indicate the configuration format of the luma component and the chroma component (cb, cr), and may be referred to as a chroma format or a chroma array type. The color format (or chroma format) may be predetermined or may be adaptively signaled. For example, the chroma format may be signaled based on at least one of chroma_format_idc and separate_colour_plane_flag as shown in Table 11.

[0264] [Table 11]

[0265] chroma_format_idc separate_colour_plane_flag Chroma format SubWidthC SubHeightC 0 0 monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1

[0266] Figure 13 The configuration of samples of the chroma format according to Table 11 is shown.

[0267] The 4:2:0 sampling with a chroma format index of 1, i.e., chroma array type 1, indicates that the height and width of the two chroma arrays are half the height and width of the luma array, respectively, while the 4:2:2 sampling with a chroma format index of 2, i.e., chroma array type 2, indicates that the height of the two chroma arrays is equal to the height of the luma array and their width is half the width of the luma array.

[0268] A 4:4:4 sampling with a chroma format index of 3, ie, chroma array type 3, indicates that the height and width of the chroma array are the same as the height and width of the luma array.

[0269] According to an example, based on the tree type of the second target block not being a single tree or the chroma array type thereof not being 3, the specific area may include a sample position of (xCb+cbWidth / 2, yCb+cbHeight / 2). Here, (xCb, yCb) may indicate the upper left position of the chroma block in luma sample units, cbWidth may indicate the width of a corresponding luma block corresponding to the chroma block, and cbHeight may indicate the height of the corresponding luma block.

[0270] Figures 14a to 14c It is shown that the variable MIP flag of the corresponding luma block as the first target block is used to derive the intra prediction mode for the chroma block as the second target block. Figures 14a to 14c The sample positions of (xCb+cbWidth / 2, yCb+cbHeight / 2) included in a specific area are shown according to the tree types and color formats of the luma block and the chroma block.

[0271] Figure 14a The luma block and chroma block are shown with a single tree type and a color format of 4:2:0. That is, Figure 14a A case where the condition that the chroma array type of the second target block is not 3 is satisfied among the conditions that the tree type of the second target block is not a single tree or the conditions that the chroma array type of the second target block is not 3 is shown.

[0272] The first sample position (I) of the chroma block indicates the upper left position of the chroma block, and the second sample position (II (xCb, yCb)) of the luma block indicates the upper left position of the chroma block in the luma sample unit. cbWidth indicates the width of the corresponding luma block corresponding to the chroma block, and cbHeight indicates the height of the corresponding luma block. Since the color format is 4:2:0, the width and height (cbWidth and cbHeight) of the corresponding luma block corresponding to the chroma block are twice the width and height (cbWidth / 2 and cbHeight / 2) of the chroma block.

[0273] The variable MIP flag value of the third sample position (III) indicated by (xCb+cbWidth / 2, yCb+cbHeight / 2) in the luma block can be used to derive the intra prediction mode for the chroma block. That is, a candidate intra prediction mode for the chroma block as the second target block can be derived based on the variable MIP flag of the sample position of (xCb+cbWidth / 2, yCb+cbHeight / 2) included in the luma block as the first target block, and the intra prediction mode for the chroma block can be derived based on the candidate intra prediction mode.

[0274] Figure 14bThe luminance block and chrominance block of the dual tree type and the color format of 4:4:4 are shown. That is, Figure 14b A case where the condition that the tree type of the second target block is not a single tree is satisfied among the conditions that the tree type of the second target block is not a single tree or the conditions that the chroma array type of the second target block is not 3 is shown.

[0275] As shown, the chroma blocks for the luma blocks indicated by the dotted lines are indicated by the dotted lines, and since the color format is 4:4:4, the width and height of the luma array and the chroma array are the same. Although the luma block is not split, the chroma block is split and coded.

[0276] When the lower right block of the segmented chroma block is the second target block to be predicted, the first sample position (I) indicates the upper left position of the second target block as shown, and the second sample position (II(xCb,yCb)) indicates the upper left position of the chroma block in the luma sample unit.

[0277] In addition, since the color format is 4:4:4, the third sample position (III) indicated by (xCb+cbWidth / 2, yCb+cbHeight / 2) in the luma block is located in the corresponding luma block, which is located at the lower right of the luma block.

[0278] The decoding device may derive an intra-frame prediction mode for the chroma block based on the variable MIP flag value of the third sample position (III). That is, a candidate intra-frame prediction mode for the chroma block as the second target block may be derived based on the variable MIP flag of the sample position of (xCb+cbWidth / 2, yCb+cbHeight / 2) included in the luma block as the first target block, and the intra-frame prediction mode for the chroma block may be derived based on the candidate intra-frame prediction mode.

[0279] Figure 14c The luminance block and chrominance block of the dual tree type and the color format of 4:2:0 are shown. That is, Figure 14c The case where both the condition that the tree type of the second target block is not a single tree and the condition that the chroma array type of the second target block is not 3 are satisfied is shown. In other words, Figure 14c A case where the tree type of the second target block is not single tree and its chroma array type is not 3 is shown.

[0280] The first sample position (I) of the chroma block indicates the upper left position of the chroma block, and the second sample position (II(xCb, yCb)) of the luma block indicates the upper left position of the chroma block in luma sample units. cbWidth indicates the width of the corresponding luma block corresponding to the chroma block, and cbHeight indicates the height of the corresponding luma block.

[0281] Therefore, the variable MIP flag value of the third sample position (III) indicated by (xCb+cbWidth / 2, yCb+cbHeight / 2) in the luma block can be used to derive the intra prediction mode for the chroma block. That is, a candidate intra prediction mode for the chroma block as the second target block can be derived based on the variable MIP flag of the sample position of (xCb+cbWidth / 2, yCb+cbHeight / 2) included in the luma block as the first target block, and the intra prediction mode for the chroma block can be derived based on the candidate intra prediction mode.

[0282] The decoding apparatus may derive a prediction sample for the second target block based on the intra prediction mode for the second target block ( S1150 ), and may generate a reconstructed block based on the prediction sample ( S1160 ).

[0283] The decoding device may generate a reconstructed block based on the received residual information and the prediction samples. The decoding device may derive transform coefficients through a transform process based on the residual information, and may utilize a variable MIP flag value of the first target block when an intra prediction mode for the second target block is required during the transform process.

[0284] According to another example, when the intra prediction type information includes a flag syntax element (intra_subpartitions_mode_flag) of the intra subpartition (ISP) mode for the first target block, the decoding device may derive the intra prediction mode for the second target block using a variable value set based on the value of intra_subpartitions_mode_flag.

[0285] That is, the decoding device can derive the value of the flag syntax element (intra_subpartitions_mode_flag), can set the variable ISP flag (IntraSubPartitionsModeFlag) for a preset specific area that is the same as the area of the first target block based on the value of the flag syntax element, and can derive the intra-frame prediction mode for the second target block based on the value of the variable ISP flag.

[0286] The variable ISP flag may also be set for a specific region based on the tree type of the first target block not being dual-tree chroma. That is, the variable ISP flag may be set to the value of the received flag syntax element for a specific region only when the tree type of the first target block is single-tree or dual-tree luma but not dual-tree chroma.

[0287] The following figures are provided to describe specific examples of the present disclosure. Because the specific terms of the devices or specific terms of signals / messages / fields illustrated in the figures are provided for explanation, the technical features of the present disclosure are not limited to the specific terms used in the following figures.

[0288] Figure 15 is a flowchart illustrating the operation of a video encoding apparatus according to an embodiment of the present disclosure.

[0289] exist Figure 15 Each process disclosed in the reference is based on Figures 5 to 10 Therefore, with reference to Figure 2 and Figures 5 to 10 Descriptions of those overlapping specific details will be omitted or will be made schematically.

[0290] When the intra MIP mode is applied to the first target block, the encoding apparatus 200 according to an embodiment may derive a prediction sample for the first target block and may derive a value of an intra MIP flag for the first target block ( S1510 ).

[0291] When the ISP is applied to the current block, the encoding apparatus may perform prediction through each sub-partitioned transform block.

[0292] The encoding device can determine whether to apply ISP coding or ISP mode to the current block, that is, the coding block, and can determine the direction in which the current block is partitioned, and can derive the size and number of partitioned sub-blocks according to the determination result.

[0293] The same intra prediction mode can be applied to the sub-partitioned transform blocks into which the current block is partitioned, and the encoding device can derive prediction samples for each sub-partitioned transform block. That is, the encoding device performs intra prediction sequentially according to the partitioning form of the sub-partitioned transform blocks, for example, horizontally or vertically, or from left to right or from top to bottom. For the leftmost or topmost sub-block, as in the conventional intra prediction method, reference is made to the reconstructed pixels of the already coded coding block. Furthermore, for each side of a subsequent internal sub-partitioned transform block that is not adjacent to the previous sub-partitioned transform block, in order to derive reference pixels adjacent to the side, reference is made to the reconstructed pixels of the already coded adjacent coding block, as in the conventional intra prediction method.

[0294] The encoding apparatus may set a variable MIP flag for a preset specific region that is the same as the region of the first target block based on a value of the intra MIP flag ( S1520 ).

[0295] As described with reference to Table 6, the variable MIP flag (IntraLumaMipFlag[x][y]) may be set in the semantics of the intra-MIP syntax element and may be utilized in subsequent processes for decoding the image.

[0296] The specific area may be set to be the same as the area where the sample is located in the first target block (x=x0..x0+cbWidth-1 and y=y0..y0+cbHeight-1). Here, cbWidth and cbHeight indicate the width and height of the first target block.

[0297] The variable MIP flag may be set for a specific region based on the tree type of the first target block not being dual-tree chroma. That is, the variable MIP flag may be set to the value of the received intra MIP syntax element for a specific region only when the tree type of the first target block is single-tree or dual-tree luma but not dual-tree chroma.

[0298] The intra-MIP syntax element intra_mip_flag is signaled only when the tree type of the corresponding coding block is single-tree or dual-tree luma. It is not signaled and is therefore inferred to be 0 when the tree type is dual-tree chroma. Therefore, when the tree type of the first target block is dual-tree chroma, the value of intra_mip_flag is 0. In intra prediction or transform, when the intra prediction mode for a chroma block is derived, the intra prediction mode for the luma block is adopted, in which case the newly set variable MIP flag can be used. When the conditions for setting the variable MIP flag do not include the condition for setting the variable MIP flag only in the "single-tree and dual-tree luma" cases, the variable MIP flag can also be set in dual-tree chroma. In this case, since the value of intra_mip_flag is 0, the variable MIP flag does not include information about the luma component, and therefore luma information cannot be used when deriving the intra prediction mode for the chroma block. To prevent this problem, the variable MIP flag can be set only when the corresponding coding block (i.e., the first target block) is not dual-tree chroma.

[0299] The encoding apparatus may derive an intra prediction mode for the second target block based on the variable MIP flag of the first target block ( S1530 ).

[0300] refer to Figures 12a to 14c The described details may be applied to an intra prediction mode derivation process performed by an encoding device.

[0301] The encoding apparatus may derive prediction samples for the second target block based on the derived intra prediction mode for the second target block ( S1540 ), and may derive residual samples for the second target block based on the prediction samples ( S1550 ).

[0302] In addition, as described above, when the intra prediction type information about the first target block is the intra sub-partitioning (ISP) mode, the variable ISP flag value of the first target block may be set based on the flag value indicating whether the intra sub-partitioning (ISP) mode is performed. The intra prediction mode for the second target block may be derived based on the variable ISP flag value.

[0303] The encoding apparatus may encode and output transform coefficient information generated based on the intra MIP flag and the residual sample ( S1560 ).

[0304] The encoding apparatus may derive a quantized transform coefficient by performing quantization based on the modified transform coefficient with respect to the current block, and may generate and output image information including the intra MIP flag.

[0305] The encoding device may generate residual information including information about the quantized transform coefficients. The residual information may include the aforementioned transform-related information / syntax elements. The encoding device may encode the image / video information including the residual information and may output the encoded image / video information in the form of a bitstream.

[0306] Specifically, the encoding apparatus 200 may generate information about a quantized transform coefficient and may encode the quantized information about the generated transform coefficient.

[0307] In the present disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for consistency of expression.

[0308] In addition, in the present disclosure, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled using residual coding syntax. The transform coefficients may be derived based on the residual information (or information about the transform coefficients), and the scaled transform coefficients may be derived by inverse transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. These details may also be applied / expressed in other parts of the present disclosure.

[0309] In the above embodiments, the method is described based on a flowchart with the aid of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a step may be performed in an order or step different from the above order or step or simultaneously with another step. In addition, it will be understood by those skilled in the art that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps of the flowchart may be removed without affecting the scope of the present disclosure.

[0310] The above-mentioned method according to the present disclosure may be implemented in software form, and the encoding device and / or decoding device according to the present disclosure may be included in an apparatus for image processing such as a TV, a computer, a smart phone, a set-top box, a display device, etc.

[0311] When the embodiments of the present disclosure are specifically implemented by software, the above methods can be embodied as modules (processes, functions, etc.) for performing the above functions. The modules can be stored in a memory and can be executed by a processor. The memory can be inside or outside the processor and can be connected to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices. In other words, the embodiments described in the present disclosure can be specifically implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units shown in each of the accompanying drawings can be specifically implemented and executed on a computer, a processor, a microprocessor, a controller or a chip.

[0312] In addition, the decoding device and the encoding device to which the present disclosure is applied may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), and the like.

[0313] In addition, the processing method of the present invention can be generated in the form of a program executed by a computer and stored in a computer-readable recording medium. Multimedia data with a data structure according to the present invention can also be stored in a computer-readable recording medium. Computer-readable recording media include all kinds of storage devices and distributed storage devices in which computer-readable data is stored. Computer-readable recording media may include, for example, Blu-ray discs (BDs), universal serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media embodied in the form of carrier waves (e.g., transmission via the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or sent via a wired or wireless communication network. Additionally, the embodiments of the present invention can be embodied as a computer program product through program code, and the program code can be executed on a computer through the embodiments of the present invention. The program code can be stored on a computer-readable carrier.

[0314] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined to be implemented or performed in a device, and the technical features of the device claims can be combined to be implemented or performed in a method. Furthermore, the technical features of the method claims and the device claims can be combined to be implemented or performed in a device, and the technical features of the method claims and the device claims can be combined to be implemented or performed in a method.

Claims

1. An image decoding method performed by a decoding device, the method comprising: Obtaining image information including intra prediction type information and residual information from a bitstream, the intra prediction type information including an intra MIP syntax element for a first coding block, the intra MIP syntax element indicating whether matrix-based intra prediction is applied to the first coding block; setting MIP flag variables for the first coding block based on a value of the intra MIP syntax element, the MIP flag variables respectively corresponding to sample positions in the first coding block, wherein each of the MIP flag variables is set equal to the value of the intra MIP syntax element indicating whether the matrix-based intra prediction is applied; deriving an intra prediction mode for a second coding block based on one of the MIP flag variables; deriving prediction samples for the second coding block based on the intra prediction mode for the second coding block; deriving transform coefficient information based on the residual information, and A reconstructed block is generated based on the prediction sample and the transform coefficient information.

2. An image encoding method performed by an encoding device, the method comprising: determining a value of an intra MIP syntax element indicating whether matrix-based intra prediction is applied to a first coding block; setting MIP flag variables for the first coding block based on whether the matrix-based intra prediction is applied to the first coding block, the MIP flag variables respectively corresponding to sample positions in the first coding block, wherein each of the MIP flag variables is set equal to a value of the intra MIP syntax element indicating whether the matrix-based intra prediction is applied; determining an intra prediction mode for a second coding block based on one of the MIP flag variables; deriving residual samples for the second coding block based on prediction samples for the second coding block, the prediction samples being obtained based on the intra prediction mode for the second coding block; and Encode transform coefficient information generated based on the residual samples.

3. A method for transmitting data for image information, comprising: generating a bitstream of the image information based on determining a value of an intra MIP syntax element indicating whether matrix-based intra prediction is applied to a first coding block, setting MIP flag variables for the first coding block based on whether the matrix-based intra prediction is applied to the first coding block, the MIP flag variables respectively corresponding to sample positions in the first coding block, wherein each of the MIP flag variables is set equal to the value of the intra MIP syntax element indicating whether the matrix-based intra prediction is applied, determining an intra prediction mode for a second coding block based on one of the MIP flag variables, deriving residual samples for the second coding block based on prediction samples for the second coding block, the prediction samples being obtained based on the intra prediction mode for the second coding block, and generating the bitstream of the image information including transform coefficient information generated based on the residual samples; and The data including the bitstream is transmitted.