MIP-based video encoding / decoding method, bitstream transmission method, and recording medium storing bitstream.

The MIP-based video encoding/decoding method addresses the challenge of high-resolution video transmission costs by using partial reference samples for prediction, enhancing efficiency and reducing costs.

JP2026511728APending Publication Date: 2026-04-14LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2024-03-27
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality video has led to a surge in transmission and storage costs due to the higher amount of information or bits required, necessitating highly efficient video compression technologies.

Method used

A video encoding/decoding method utilizing matrix-based intraprediction (MIP) that generates predicted blocks using only a portion of reference samples, such as the left or top-side samples, to enhance encoding/decoding efficiency.

Benefits of technology

This approach improves encoding/decoding efficiency and intra-screen prediction efficiency, allowing for more effective storage and transmission of high-resolution video while reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026511728000001_ABST
    Figure 2026511728000001_ABST
Patent Text Reader

Abstract

A video encoding / decoding method and apparatus are provided. The video decoding method according to this disclosure includes the steps of determining an in-screen prediction mode for the current block, and generating a predicted block for the current block by performing an in-screen prediction based on the determined in-screen prediction mode, wherein the in-screen prediction may be limited to using only reference samples within a predetermined range from the available reference samples of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to a MIP-based video encoding / decoding method, a method for transmitting a bitstream, and a recording medium for storing a bitstream, and more particularly to a MIP (matrix-based intraprediction)-based video encoding / decoding method, a method for transmitting a bitstream, and a recording medium for storing a bitstream using only a portion of a reference sample. [Background technology]

[0002] In recent years, demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing in various fields. The higher the resolution and quality of video data, the greater the amount of information or bits transmitted compared to existing video data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.

[0003] Therefore, highly efficient video compression technology is desired to effectively transmit, store, and play back high-resolution, high-quality video information. [Overview of the project] [Problems that the invention aims to solve]

[0004] The purpose of this disclosure is to provide a video encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Furthermore, this disclosure aims to provide a video encoding / decoding method and apparatus based on in-screen prediction and selected reference samples.

[0006] Furthermore, this disclosure aims to provide a video encoding / decoding method and apparatus that improves the intra-screen prediction efficiency of the Enhanced Compression Model (ECM).

[0007] Furthermore, this disclosure aims to provide a non-temporary computer-readable recording medium for storing a bitstream generated by a video encoding method or apparatus relating to this disclosure.

[0008] Furthermore, this disclosure aims to provide a non-temporary computer-readable recording medium that stores a bitstream that is received and decoded by the video decoding device relating to this disclosure and used for restoring video.

[0009] Furthermore, this disclosure aims to provide a method for transmitting a bitstream generated by a video encoding method or apparatus relating to this disclosure.

[0010] The technical challenges addressed in this disclosure are not limited to those mentioned above, and other technical challenges not mentioned above will be clearly understood by those with ordinary skill in the art to which this disclosure pertains from the following description. [Means for solving the problem]

[0011] A video decoding method according to one aspect of the present disclosure includes the steps of determining whether or not a matrix-based intraprediction (MIP) mode is applied to the current block, and generating a predicted block of the current block based on the application of the MIP mode to the current block, wherein the predicted block may be generated using only a portion of the left-side reference sample or the top-side reference sample of the current block based on the MIP mode.

[0012] On the other hand, as an example, the prediction block may be generated based on the MIP mode using only the upper reference sample of the current block, or it may be generated based on one or more upper reference sample lines.

[0013] On the one hand, as an example, the prediction block may be generated based on the MIP mode using only the upper reference sample of the current block, and the upper reference sample may further include the upper right reference sample of the current block.

[0014] On the one hand, as an example, the prediction block may be generated based on the MIP mode using only the left reference sample of the current block, and may be generated based on one or more left reference sample lines.

[0015] On the one hand, as an example, the prediction block may be generated based on the MIP mode using only the left reference sample of the current block, and the left reference sample may include the lower left reference sample of the current block.

[0016] On the one hand, as an example, the prediction block may be generated using a part of the left reference sample and the upper reference sample of the current block, and at least one of the left reference sample or the upper reference sample may be partially adjacent to the current block.

[0017] On the one hand, as an example, the prediction block may be generated using a part of the left reference sample and the upper reference sample of the current block, the left reference sample may further include the lower left reference sample of the current block, and the upper reference sample may further include the upper right reference sample of the current block.

[0018] On the one hand, as an example, based on the application of the MIP mode, it may be determined whether the MIP mode is performed based on the left reference sample or the upper reference sample.

[0019] On the one hand, as an example, whether the MIP mode is performed based on the left reference sample or the upper reference sample may be determined based on the information obtained from the bitstream.

[0020] A video encoding method according to one aspect of the present disclosure includes the steps of determining whether or not to apply a matrix-based intraprediction (MIP) mode to a current block, and generating a predicted block of the current block based on the application of the MIP mode to the current block, wherein the predicted block may be generated using only a portion of the left-side reference sample or the top-side reference sample of the current block based on the MIP mode.

[0021] A computer-readable recording medium according to yet another aspect of the present disclosure can store a bitstream generated by the video encoding method or video encoding apparatus of the present disclosure.

[0022] A transmission method relating to yet another aspect of the present disclosure may transmit a bitstream generated by a video encoding device or video encoding method of the present disclosure.

[0023] The features briefly summarized above are merely illustrative examples of the detailed description of the disclosure described below and do not limit the scope of the disclosure. [Effects of the Invention]

[0024] According to this disclosure, it is possible to provide a video encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0025] Furthermore, this disclosure provides a video encoding / decoding method and apparatus based on in-screen prediction and selected reference samples.

[0026] Furthermore, this disclosure provides a video encoding / decoding method and apparatus that can generate in-screen prediction blocks that are adaptive to the reference sample region during in-screen prediction.

[0027] Furthermore, this disclosure provides a video encoding / decoding method and apparatus that improves the intra-ECM prediction efficiency.

[0028] Furthermore, this disclosure makes it possible to provide a non-temporary computer-readable recording medium for storing a bitstream generated by a video encoding method or apparatus relating to this disclosure.

[0029] Furthermore, this disclosure makes it possible to provide a non-temporary computer-readable recording medium that stores a bitstream that is received and decoded by the video decoding device relating to this disclosure and used for restoring video.

[0030] Furthermore, this disclosure provides a method for transmitting a bitstream generated by a video encoding method or apparatus relating to this disclosure.

[0031] The effects obtained from this disclosure are not limited to those mentioned above, and any other effects not mentioned above will be clearly understood by a person with ordinary skill in the art to which this disclosure pertains from the following description. [Brief explanation of the drawing]

[0032] [Figure 1] This is a schematic diagram showing a video coding system to which the embodiments of this disclosure can be applied. [Figure 2] This is a schematic diagram showing a video encoding device to which the embodiments of this disclosure can be applied. [Figure 3] This is a schematic diagram showing an image decoding device to which the embodiments of this disclosure can be applied. [Figure 4] This flowchart shows an example of an intra-predictive mode signaling method in an encoding device. [Figure 5] This flowchart shows an example of a method for determining the intra-prediction mode in a decoding device. [Figure 6] This figure shows an example of a surrounding block used for MPM list induction. [Figure 7] This figure shows an example of the overall processes of averaging, matrix-vector multiplication, and linear interpolation applicable to this disclosure. [Figure 8]This figure shows an example of the overall processes of averaging, matrix-vector multiplication, and linear interpolation applicable to this disclosure. [Figure 9] This figure shows an example of the overall processes of averaging, matrix-vector multiplication, and linear interpolation applicable to this disclosure. [Figure 10] This figure shows an example of the overall processes of averaging, matrix-vector multiplication, and linear interpolation applicable to this disclosure. [Figure 11] This figure shows an example of the overall processes of averaging, matrix-vector multiplication, and linear interpolation applicable to this disclosure. [Figure 12] This figure shows an example of a boundary averaging process applicable to this disclosure. [Figure 13] This figure shows an example of a linear interpolation process applicable to this disclosure. [Figure 14] This figure illustrates a MIP prediction method applicable to this disclosure. [Figure 15] This figure illustrates a method for performing MIP prediction according to one embodiment of the present disclosure. [Figure 16] This figure illustrates an MIP-based video encoding or decoding method according to one embodiment of the present disclosure. [Figure 17] This figure illustrates a MIP prediction method applicable to this disclosure. [Figure 18] This figure illustrates a method for performing MIP prediction according to one embodiment of the present disclosure. [Figure 19] This figure illustrates an MIP-based video encoding or decoding method according to one embodiment of the present disclosure. [Figure 20] This figure illustrates a MIP prediction method applicable to this disclosure. [Figure 21] This figure illustrates a method for performing MIP prediction according to one embodiment of the present disclosure. [Figure 22] This figure illustrates an MIP-based video encoding or decoding method according to one embodiment of the present disclosure. [Figure 23] This figure shows a video decoding method according to one embodiment of the present disclosure. [Figure 24]This figure shows a video encoding method according to one embodiment of the present disclosure. [Figure 25] This figure illustrates a content streaming system to which the embodiments of this disclosure can be applied. [Modes for carrying out the invention]

[0033] Hereafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings, so that they can be easily implemented by a person with ordinary skill in the art to which the present disclosure pertains. However, the present disclosure may be embodied in various other forms and is not limited to the embodiments described herein.

[0034] In describing embodiments of this disclosure, if a specific description of a known configuration or function is deemed to obscure the gist of this disclosure, such detailed description will be omitted. In the figures, parts unrelated to the description of this disclosure will be omitted, and similar parts will be denoted by similar reference numerals.

[0035] In this disclosure, when one component is described as being “linked,” “joined,” or “connected” to another component, this may include not only direct linkages but also indirect linkages where other components exist in between. Furthermore, when one component is described as “containing” or “having” another component, this means, unless otherwise specified, that it may contain further other components rather than excluding them.

[0036] In this disclosure, terms such as "first," "second," etc., are used solely to distinguish one component from another, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.

[0037] In this disclosure, components are distinguished from each other solely to clearly describe their respective characteristics, and this does not necessarily mean that these components are separate. That is, multiple components may be integrated to constitute a single hardware or software unit, or a single component may be distributed to constitute multiple hardware or software units. Therefore, such integrated or distributed embodiments are also included in the scope of this disclosure, even without specific mention.

[0038] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, embodiments consisting of a subset of the components described in one embodiment are also included in the scope of this disclosure. Furthermore, embodiments that further include other components in addition to the components described in various embodiments are also included in the scope of this disclosure.

[0039] This disclosure relates to the encoding and decoding of video, and unless otherwise defined herein, the terms used herein may have their ordinary meanings in the art to which this disclosure pertains.

[0040] In this disclosure, "picture" generally refers to a unit representing a single video image for a specific time period, and "slice / tile" is an encoding unit that constitutes a part of a picture. A single picture may consist of one or more slices / tiles. A slice / tile may also contain one or more CTUs (coding tree units).

[0041] In this disclosure, “pixel” or “pel” can mean the smallest unit that constitutes a picture (or video). The term “sample” may also be used as a counterpart to pixel. A sample may generally represent a pixel or a pixel value, or it may represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.

[0042] In this disclosure, “unit” may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with that region. A unit may, as it may be, be replaced by terms such as “sample array,” “block,” or “area.” In general, an MxN block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0043] In this disclosure, “current block” can mean one of the following: “current coding block,” “current coding unit,” “block to encode,” “block to decode,” or “block to process.” When prediction is performed, “current block” can mean “current prediction block” or “block to predict.” When transformation (inverse transformation) / quantization (inverse quantization) is performed, “current block” can mean “current transformation block” or “block to transform.” When filtering is performed, “current block” can mean “block to filter.”

[0044] In this disclosure, "current block" may mean a block containing both a luma component block and a chroma component block, or "the luma block of the current block," unless otherwise explicitly stated as a chroma block. The luma component block of the current block may be expressed with an explicit mention of a luma component block, such as "luma block" or "current luma block." Similarly, the chroma component block of the current block may be expressed with an explicit mention of a chroma component block, such as "chroma block" or "current chroma block."

[0045] In this disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."

[0046] In this disclosure, “or” may be interpreted as “and / or.” For example, “A or B” may mean 1) “A” only, 2) “B” only, or 3) “A and B.” Alternatively, in this disclosure, “or” may mean “additionally or alternatively.”

[0047] Overview of the video coding system

[0048] Figure 1 is a schematic diagram showing a video coding system to which the embodiments of this disclosure can be applied.

[0049] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 in file or streaming form via a digital storage medium or network.

[0050] An encoding device 10 according to one embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to one embodiment may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be called a video / image encoding unit, and the decoding unit 22 may be called a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, and the display unit may be composed of a separate device or external component.

[0051] The video source generation unit 11 can acquire video / images through video / image capture, synthesis, or generation processes. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, and can generate video / images (electronically). For example, virtual video / images may be generated by a computer, in which case the video / image capture process may be replaced by a process in which related data is generated.

[0052] The encoding unit 12 can encode the input video / image data. The encoding unit 12 can perform a series of procedures such as prediction, transformation, and quantization for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0053] The transmitting unit 13 can acquire encoded video / image information or data output in bitstream form and transmit it in file or streaming form to the receiving unit 21 of the decoding device 20 or other external object via a digital storage medium or network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit 13 may include elements for generating media files in a predetermined file format and may include elements for transmission via a broadcast / communication network. The transmitting unit 13 may be provided as a transmission device separate from the encoding unit 12, in which case the transmission device may include at least one processor that acquires encoded video / image information or data output in bitstream form and a transmitting unit that transmits it in file or streaming form. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.

[0054] The decoding unit 22 can decode the video / image by performing a series of procedures such as inverse quantization, inverse transform, and prediction, which correspond to the operation of the encoding unit 12.

[0055] The rendering unit 23 can render the decoded video / image. The rendered video / image may be displayed through the display unit.

[0056] Overview of video encoding equipment

[0057] Figure 2 is a schematic diagram showing a video encoding device to which the embodiments of this disclosure can be applied.

[0058] As shown in Figure 2, the video encoding device 100 may include a video splitting unit 110, a subtraction unit 115, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse conversion unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter-prediction unit 180, an intra-prediction unit 185, and an entropy encoding unit 190. The inter-prediction unit 180 and the intra-prediction unit 185 may be collectively called the "prediction unit". The conversion unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse conversion unit 150 may be included in the residual processing unit. The residual processing unit may further include a subtraction unit 115.

[0059] Depending on the embodiment, all or at least some of the multiple components constituting the video encoding device 100 may be embodied as a single hardware component (e.g., an encoder or a processor). Furthermore, the memory 170 may include a DPB (decoded picture buffer) and may be embodied by a digital storage medium.

[0060] The video splitting unit 110 can split the input video (or picture, frame) input to the video encoding device 100 into one or more processing units. For example, the processing units may be called coding units (CUs). Coding units can be obtained by recursively splitting a coding tree unit (CTU) or the largest coding unit (LCU) using a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit may be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. For the splitting of coding units, a quad-tree structure may be applied first, followed by a binary-tree structure and / or a ternary-tree structure. The coding procedure according to this disclosure may be performed based on the final coding unit that is not further split. The maximum coding unit may be used directly as the final coding unit, or a lower-depth coding unit obtained by dividing the maximum coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or restoration, which will be described later. As another example, the processing unit of the coding procedure may be a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.

[0061] The prediction unit (inter-prediction unit 180 or intra-prediction unit 185) can make predictions for the block to be processed (current block) and generate a predicted block that includes prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block or on a CU basis. The prediction unit can generate various information regarding the prediction of the current block and transmit it to the entropy encoding unit 190. The prediction information may be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0062] The intra-prediction unit 185 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity of the current block or at a distance from it, depending on the intra-prediction mode and / or intra-prediction method. The intra-prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the accuracy of the prediction direction. However, this is an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.

[0063] The interprediction unit 180 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, colCU, etc. The reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 180 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation may be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 180 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, the residual signal does not need to be transmitted.In motion vector prediction (MVP) mode, the motion vectors of surrounding blocks are used as motion vector predictors, and the motion vector of the current block can be signaled by encoding the motion vector difference and an indicator for the motion vector predictor. The motion vector difference represents the difference between the motion vector of the current block and the motion vector predictor.

[0064] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described later. For example, the prediction unit may apply intra-prediction or inter-prediction to predict the current block, or it may apply intra-prediction and inter-prediction simultaneously. A prediction method that applies intra-prediction and inter-prediction simultaneously to predict the current block may be called CIIP (combined inter and intra prediction). The prediction unit can also perform intra-block copy (IBC) to predict the current block. Intra-block copy may be used, for example, for coding content images / videos such as games, as in SCC (screen content coding). IBC is a method of predicting the current block using a reference block that has already been restored in the current picture at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current picture may be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but it may be performed similarly to inter-prediction in that it derives the reference block within the current picture. In other words, IBC can use at least one of the interpretation methods described in this disclosure.

[0065] The predicted signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal output from the prediction unit (predicted block, predicted sample array) from the input video signal (original block, original sample array). The generated residual signal may be transmitted to the conversion unit 120.

[0066] The transformation unit 120 can generate transformation coefficients by applying a transformation method to the residual signal. For example, the transformation method may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to the transformation obtained from a graph when the relationship information between pixels is represented by this graph. CNT refers to the transformation obtained by generating a prediction signal using all previously reconstructed pixels and obtaining a transformation based on it. The transformation process may be applied to pixel blocks of the same size and square shape, or to blocks of a variable size instead of square shape.

[0067] The quantization unit 130 can quantize the conversion coefficients and transmit them to the entropy encoding unit 190. The entropy encoding unit 190 can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients may be called residual information. The quantization unit 130 can rearrange the block-shaped quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients.

[0068] The entropy encoding unit 190 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 190 can also encode information necessary for video / image restoration (e.g., the values ​​of syntax elements) together or separately. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as an adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The signaling information, transmitted information, and / or syntax elements referred to in this disclosure may be encoded by the encoding procedure described above and included in the bitstream.

[0069] The bitstream may be transmitted over a network or stored on a digital storage medium. Here, the network may include broadcasting networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting the signal output from the entropy encoding unit 190 and / or a storage unit (not shown) for storing it may be provided as an internal / external element of the video encoding device 100, or the transmitting unit may be provided as a component of the entropy encoding unit 190.

[0070] The quantized conversion coefficients output from the quantization unit 130 may be used to generate a resistive signal. For example, by applying inverse quantization and inverse transformation to the quantized conversion coefficients in the inverse quantization unit 140 and the inverse transformation unit 150, a resistive signal (residual block or resistive sample) can be reconstructed.

[0071] The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the predicted signal output from the inter-prediction unit 180 or the intra-prediction unit 185. When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block may be used as the reconstructed block. The adder 155 may be called the reconstruction unit or the reconstructed block generation unit. The generated reconstructed signal may be used for intra-prediction of the next block to be processed in the current picture, or, as described later, may be used for inter-prediction of the next picture after filtering.

[0072] The filtering unit 160 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 170, specifically in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 160 can generate various filtering information and transmit it to the entropy encoding unit 190, as will be described later in the description of each filtering method. The filtering information may be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0073] The corrected restored picture transmitted to memory 170 may be used as a reference picture in the interpretation unit 180. This allows the video encoding device 100 to avoid prediction mismatches between the video encoding device 100 and the video decoding device when interpretation is applied, and also improves encoding efficiency.

[0074] The DPB in memory 170 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 180. Memory 170 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information may be transmitted to the inter-prediction unit 180 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 170 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 185.

[0075] Overview of the video decoding device

[0076] Figure 3 is a schematic diagram showing an image decoding device to which the embodiments of this disclosure can be applied.

[0077] As shown in Figure 3, the video decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transformation unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265. The inter-prediction unit 260 and the intra-prediction unit 265 can be collectively referred to as the "prediction unit". The inverse quantization unit 220 and the inverse transformation unit 230 may be included in the residual processing unit.

[0078] All or at least some of the multiple components constituting the video decoding device 200 may be embodied as a single hardware component (e.g., a decoder or processor) depending on the embodiment. Furthermore, the memory 170 may include a DPB and may be embodied by a digital storage medium.

[0079] A video decoding device 200 that receives a bitstream containing video / image information can restore the image by performing a process corresponding to the process performed by the video encoding device 100 in Figure 2. For example, the video decoding device 200 can perform decoding using the processing unit applied in the video encoding device. Therefore, the decoding processing unit may be, for example, a coding unit. The coding unit may be a coding tree unit, or it may be obtained by dividing the largest coding unit. The restored video signal decoded and output by the video decoding device 200 may then be played back by a playback device (not shown).

[0080] The video decoding device 200 can receive the signal output from the video encoding device shown in Figure 2 in the form of a bitstream. The received signal may be decoded by the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information necessary for video restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also further include general constraint information. The video decoding device may further utilize the parameter set information and / or the general constraint information to decode the video. The signaling information, received information, and / or syntax elements referred to in this disclosure may be obtained from the bitstream by decoding through the decoding procedure. For example, the entropy decoding unit 210 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements necessary for image restoration and the quantized values ​​of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoding information of the surrounding blocks and the blocks to be decoded, or symbol / bin information decoded in a previous stage, predicts the probability of bin occurrence based on the determined context model, performs arithmetic decoding of the bins, and generates symbols corresponding to the values ​​of each syntax element.In this case, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Information related to prediction from the information decoded by the entropy decoding unit 210 is provided to the prediction unit (inter-prediction unit 260 and intra-prediction unit 265), and residual values ​​that have been entropy decoded by the entropy decoding unit 210, i.e., quantized conversion coefficients and related parameter information, may be input to the inverse quantization unit 220. In addition, information related to filtering from the information decoded by the entropy decoding unit 210 may be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives signals output from the video encoding device may be further provided as an internal / external element of the video decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.

[0081] On the other hand, the video decoding device according to this disclosure may be called a video / image / picture decoding device. The video decoding device may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transformation unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265.

[0082] The inverse quantization unit 220 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 220 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scan order performed by the video encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.

[0083] The inverse conversion unit 230 can inversely convert the conversion coefficients to obtain residual signals (residual blocks, residual sample arrays).

[0084] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 210, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode (prediction method).

[0085] As mentioned in the description of the prediction unit of the video coding device 100, the prediction unit can generate prediction signals based on various prediction methods (techniques) described later.

[0086] The intra-prediction unit 265 can predict the current block by referring to the samples in the current picture. The description of the intra-prediction unit 185 may also apply to the intra-prediction unit 265.

[0087] The interprediction unit 260 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in block, subblock, or sample units based on the correlation of motion information between the surrounding block and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding block may include spatially neighboring blocks present in the current picture and temporally neighboring blocks present in the reference picture. For example, the interprediction unit 260 can construct a motion information candidate list based on the surrounding blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction may be performed based on various prediction modes (methods), and the prediction information may include information indicating the mode (method) of interprediction for the current block.

[0088] The adder 235 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 260 and / or intra-prediction unit 265). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block may be used as the restored block. The description of the adder 155 may also apply to the adder 235. The adder 235 may be called the restore unit or the restored block generation unit. The generated restored signal may be used for intra-prediction of the next block to be processed in the current picture, or, as described later, may be used for inter-prediction of the next picture after filtering.

[0089] The filtering unit 240 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 250, specifically in the DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.

[0090] The restored picture stored (modified) in the DPB of memory 250 may be used as a reference picture in the inter-prediction unit 260. Memory 250 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 250 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 265.

[0091] In this specification, the embodiments described for the filtering unit 160, inter-prediction unit 180, and intra-prediction unit 185 of the video encoding device 100 may be applied identically or in a corresponding manner to the filtering unit 240, inter-prediction unit 260, and intra-prediction unit 265 of the video decoding device 200, respectively.

[0092] Intra Prediction Mode / Type Determination

[0093] When intraprediction is applied, the intraprediction mode applied to the current block may be determined using the intraprediction modes of the surrounding blocks. For example, the decoding device may select one of the mpm candidates in the mpm (most probable mode) list derived based on the intraprediction modes of the surrounding blocks of the current block (e.g., the surrounding blocks to the left and / or above) and further candidate modes, based on the received mpm index, or it may select one of the remaining intraprediction modes not included in the mpm candidates (and planar modes), based on the remaining intraprediction mode information. The mpm list may or may not include planar modes as candidates. For example, if the mpm list includes planar modes as candidates, the mpm list may have six candidates, and if the mpm list does not include planar modes as candidates, the mpm list may have three candidates. If the aforementioned mpm list does not include planar mode as a candidate, a not-planar flag (e.g., intra_luma_not_planar_flag) may be signaled to indicate that the current block's intra-prediction mode is not planar mode. For example, the mpm flag may be signaled first, and the mpm index and not-planar flag may be signaled if the value of the mpm flag is 1. Also, the mpm index may be signaled if the value of the not-planar flag is 1. Here, the reason why the mpm list is configured not to include planar mode as a candidate is not because the planar mode is not an mpm, but because the planar mode is always considered as an mpm, so the flag (not-planar flag) is signaled first to check whether it is a planar mode or not.

[0094] For example, whether the intra-prediction mode applied to the current block is within the MPM candidate (and planar mode) or within the remaining mode may be indicated based on the MPM flag (e.g., intra_luma_mpm_flag). A value of 1 for the MPM flag indicates that the intra-prediction mode for the current block is within the MPM candidate (and planar mode), and a value of 0 for the MPM flag indicates that the intra-prediction mode for the current block is not within the MPM candidate (and planar mode). A value of 0 for the not planar flag (e.g., intra_luma_not_planar_flag) indicates that the intra-prediction mode for the current block is planar mode, and a value of 1 for the not planar flag indicates that the intra-prediction mode for the current block is not planar mode. The mpm index may be signaled in the form of an mpm_idx or intra_luma_mpm_idx syntax element, and the remaining intra prediction mode information may be signaled in the form of a rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, the remaining intra prediction mode information can indicate one of the remaining intra prediction modes from the overall intra prediction modes that are not included in the mpm candidates (and planar modes), indexed in order of prediction mode number. The intra prediction mode may be an intra prediction mode for a luma component (sample). Hereinafter, the intra prediction mode information may include at least one of the following: the mpm flag (e.g., intra_luma_mpm_flag), the not planar flag (e.g., intra_luma_not_planar_flag), the mpm index (e.g., mpm_idx or intra_luma_mpm_idx), or the remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list may be referred to by various terms such as MPM candidate list, candModeList, etc.When MIP is applied to the current block, a separate mpm flag for MIP (e.g., intra_mip_mpm_flag), an mpm index (e.g., intra_mip_mpm_idx), and remaining intra prediction mode information (e.g., intra_mip_mpm_remainder) may be signaled, while the not planar flag is not signaled.

[0095] The intra-predictive mode signaling procedure in the encoding device and the intra-predictive mode determination procedure in the decoding device may be carried out, for example, as follows.

[0096] Figure 4 is a flowchart showing an example of an intra-predictive mode signaling method in an encoding device.

[0097] Referring to Figure 4, the encoding device configures an MPM list for the current block (S400). The MPM list may include candidate intra-prediction modes (MPM candidates) that are likely to be applied to the current block. The MPM list may also include intra-prediction modes for surrounding blocks, and may further include specific intra-prediction modes by a predetermined method. A specific method for configuring the MPM list will be described later.

[0098] The encoding device determines the intra-prediction mode for the current block (S410). The encoding device can perform predictions based on various intra-prediction modes and determine the optimal intra-prediction mode based on RDO (rate-distortion optimization) thereon. In this case, the encoding device may determine the optimal intra-prediction mode using only the MPM candidates and planar modes configured in the MPM list, or it may determine the optimal intra-prediction mode using the remaining intra-prediction modes in addition to the MPM candidates and planar modes configured in the MPM list. To give a specific example, if the intra-prediction type of the current block is not a normal intra-prediction type but a specific type (e.g., LIP, MRL, or ISP), the encoding device can determine the optimal intra-prediction mode by considering only the MPM candidates and planar modes as intra-prediction mode candidates for the current block. That is, in this case, the intra-prediction mode for the current block may be determined from among the MPM candidates and planar modes, and in this case, the mpm flag does not need to be encoded / signaled. In this case, the decoding device can infer that the mpm flag is 1 even if the mpm flag is not separately signaled.

[0099] On the other hand, generally, if the intra-prediction mode of the current block is not the planar mode but one of the MPM candidates in the MPM list, the encoding device generates an mpm index (mpm idx) indicating one of the MPM candidates. If the intra-prediction mode of the current block is not found in the MPM list, the device generates remaining intra-prediction mode information that indicates the same mode as the intra-prediction mode of the current block from among the remaining intra-prediction modes not included in the MPM list (and planar mode).

[0100] The encoding device can encode intra-prediction mode information and output it in the form of a bitstream. The intra-prediction mode information may include the mpm flag, not planar flag, mpm index, and / or remaining intra-prediction mode information described above. In general, the mpm index and the remaining intra-prediction mode information have an alternative relationship and are not signaled simultaneously when indicating an intra-prediction mode for a block. That is, either an mpm flag value of 1 and the not planar flag or mpm index are signaled together, or an mpm flag value of 0 and the remaining intra-prediction mode information are signaled together. However, as described above, if a specific intra-prediction type is applied to the current block, the mpm flag may not be signaled, and only the not planar flag and / or mpm index may be signaled. That is, in this case, the intra-prediction mode information may include only the not planar flag and / or mpm index.

[0101] The decoding device can determine the intra-prediction mode in response to the intra-prediction mode information determined and signaled by the encoding device.

[0102] Figure 5 is a flowchart showing an example of an intra-prediction mode determination method in a decoding device.

[0103] Referring to Figure 5, the decoding device obtains intra-prediction mode information from the bitstream (S500). As described above, the intra-prediction mode information may include at least one of the mpm flag, not planar flag, mpm index, and remaining intra-prediction mode.

[0104] The decoding device configures an MPM list (S510). The MPM list is configured identically to the MPM list configured by the encoding device. That is, the MPM list may include intra-prediction modes for peripheral blocks, and may further include specific intra-prediction modes by a predetermined method. A specific method for configuring the MPM list will be described later.

[0105] Although it is stated that S510 is performed after S500, this is merely an example, and S510 may be performed before S500, or at the same time.

[0106] The decoding device determines the intra-prediction mode of the current block based on the MPM list and the intra-prediction mode information (S520). For example, if the value of the mpm flag is 1, the decoding device may derive the planar mode as the intra-prediction mode of the current block, or derive the candidate indicated by the mpm index from among the MPM candidates in the MPM list (not planar flag-based) as the intra-prediction mode of the current block. As another example, if the value of the mpm flag is 0, the decoding device may derive the intra-prediction mode indicated by the remaining intra-prediction mode information from among the remaining intra-prediction modes not included in the MPM list and planar mode as the intra-prediction mode of the current block. On the other hand, as yet another example, if the intra-prediction type of the current block is a specific type (e.g., LIP, MRL, or ISP), the decoding device may derive the planar mode or the candidate indicated by the mpm index in the MPM list as the intra-prediction mode of the current block without checking the mpm flag.

[0107] For example, the not planar flag may be signaled when the MRL is not currently applied to a block (i.e., intra_luma_ref_idx == 0), and may be omitted when the MRL is currently applied to a block (i.e., intra_luma_ref_idx != 0). If the not planar flag is omitted, the decoding device may estimate its value to be 1.

[0108] On the other hand, the intra-prediction mode may include two directional intra-prediction modes and 65 directional intra-prediction modes. The non-directional intra-prediction mode may include a planar intra-prediction mode and a DC intra-prediction mode, and the directional intra-prediction mode may include intra-prediction modes 2 through 66. The extended directional intra-prediction mode may be applied to blocks of all sizes and may be applied to either the luma component or the chroma component.

[0109] On the other hand, the intra-prediction mode may further include a CCLM (cross-component linear model) mode for chroma samples, in addition to the intra-prediction mode described above. The CCLM mode may be divided into LT_CCLM, L_CCLM, and T_CCLM depending on whether the left sample, the upper sample, or both are considered for the derivation of the LM parameters, and may be applied only to chroma components.

[0110] The intra-prediction mode may be indexed as shown in Table 1 below, for example.

[0111] [Table 1]

[0112] On the other hand, the intra prediction type (or additional intra prediction mode, etc.) may include at least one of the above-mentioned LIP, PDPC, MRL, ISP, and MIP. The intra prediction type may be indicated based on intra prediction type information, and the intra prediction type information may be embodied in various forms. For example, the intra prediction type information may include intra prediction type index information that indicates one of the intra prediction types. As another example, the intra prediction type information may include at least one of the following: reference sample line information (e.g., intra_luma_ref_idx) indicating whether the MRL is applied to the current block and, if so, which reference sample line is used; ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block; ISP type information (e.g., intra_subpartitions_split_flag) indicating the split type of the subpartition when the ISP is applied; flag information indicating whether PDCP is applied; or flag information indicating whether LIP is applied. Furthermore, the intra prediction type information may include an MIP flag (or may be called intra_mip_flag) indicating whether or not MIP is applied to the current block.

[0113] On the other hand, as described above, when an MIP is applied to the current block (for example, when the value of intra_mip_flag is 1), a separate MPM list may be configured for the MIP, and the MPM flag that may be included in the intra prediction mode information for the MIP may be called intra_mip_mpm_flag, the MPM index may be called intra_mip_mpm_idx, and the remaining intra prediction mode information may be called intra_mip_mpm_remainder.

[0114] Furthermore, various prediction modes may be used for MIP, and the intra-prediction modes for MIP can be used to derive the matrix and offset for MIP. As described above, the matrix may be called the (MIP)weighted matrix, and the offset may be called the (MIP)offset vector or (MIP)bias vector. The number of intra-prediction modes for MIP may be set differently based on the size of the current block. For example, i) if the height and width of the current block (e.g., CB or TB) are both 4, then 35 intra-prediction modes (i.e., intra-prediction modes 0-34) may be available; ii) if both the height and width of the current block are 8 or less, then 19 intra-prediction modes (i.e., intra-prediction modes 0-18) may be available; and iii) in other cases, then 11 intra-prediction modes (i.e., intra-prediction modes 0-10) may be available. For example, if the current block height and width are both 4, then block size type 0 is defined as block size type 1 when both the current block height and width are 8 or less, and block size type 2 for all other cases, then the number of intra-prediction modes for MIP may be organized as shown in the following table. However, this is an example, and the number of block size types and available intra-prediction modes may be changed. In this document, intra-prediction modes for MIP may be referred to as MIP intra-prediction modes, MIP prediction modes, or MIP modes.

[0115] [Table 2]

[0116] On the other hand, the enhanced compression model (ECM) introduced a secondary MPM list. The existing primary MPM (PMPM) list consists of 6 entries, and the secondary MPM (SMPM) list contains 16 entries. First, a general MPM list with 22 entries is constructed, and then the first 6 entries from the general MPM list are included in the PMPM list, and the remaining entries are included in the SMPM list. In the general MPM list, the first entry is the planar mode, and the remaining entries consist of intra-modes for the left (L), upper (A), lower left (BL), upper right (AR), and upper left (AL) peripheral blocks as shown in Figure 6, directional modes with an offset added from the first two available directional modes of the peripheral block, and the default mode.

[0117] When the CU block is vertical, the order of the surrounding blocks may be top (A), left (L), lower left (BL), upper right (AR), and upper left (AL). Alternatively, the order may be left (L), top (A), lower left (BL), upper right (AR), and upper left (AL).

[0118] The PMPM flag is parsed, and if its value is 1, the PMPM index is parsed to determine which entry in the PMPM list is selected. Alternatively, the SPMPM flag is parsed to determine whether or not to parse the SPMPM index or the remaining mode.

[0119] Derivation of peripheral reference samples

[0120] When intraprediction is applied to the current block, peripheral reference samples to be used for intraprediction of the current block may be derived. The peripheral reference samples of the current block may include a total of 2 x nH samples adjacent to the left boundary and bottom-left of the current block of nW x nH size, a total of 2 x nW samples adjacent to the top boundary and top-right of the current block, and one sample adjacent to the top-left of the current block. Alternatively, the peripheral reference samples of the current block may include upper peripheral samples for multiple columns and left peripheral samples for multiple rows. Furthermore, the peripheral reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block of nW x nH size, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right of the current block.

[0121] On the other hand, when MRL, which will be described later, is applied, the reference sample may be located on lines 1 to 3 on the left / above side, rather than on line 0 which is adjacent to the current block. In this case, the number of surrounding reference samples may increase further. The specific regions and numbers of surrounding reference samples will be described later.

[0122] On the other hand, when the ISP described later is applied, the surrounding reference samples can be derived on a subpartition basis.

[0123] Currently, some of the surrounding reference samples in a block may not yet be decoded or available. In this case, the decoder can construct the surrounding reference samples to be used for prediction by interpolating the available samples.

[0124] Currently, some of the surrounding reference samples in a block may not yet be decoded or available. In this case, the decoder can construct the surrounding reference samples to be used for prediction by extrapolating the available samples. Starting from the bottom left corner and reaching the top right corner reference sample, the decoder can construct the samples by updating the available samples with the latest available sample, while substituting or padding undecoded or unavailable pixels with the latest available sample.

[0125] Derivation of intra-prediction mode / type-based prediction samples

[0126] The prediction unit of the encoding / decoding device can derive a reference sample from the surrounding reference samples of the current block using the intra-prediction mode, and can generate a predicted sample of the current block based on the reference sample.

[0127] For example, (i) a predicted sample can be derived based on the average or interpolation of neighboring reference samples of the current block, and (ii) a predicted sample can be derived based on a reference sample among the neighboring reference samples of the current block that is located in a specific (predicted) direction relative to the predicted sample. Case (i) may be called a non-directional mode or non-angular mode, and case (ii) may be called a directional mode or angular mode. Alternatively, the predicted sample may be generated by interpolation between the first and second neighboring samples, which are located in the opposite direction to the prediction direction of the intra-prediction mode of the current block, with respect to the predicted sample of the current block. In this case, it may be called linear interpolation intra-prediction (LIP). Furthermore, temporary predicted samples for the current block can be derived based on filtered peripheral reference samples, and predicted samples for the current block can be derived by performing a weighted sum on the temporary predicted samples and at least one reference sample derived by the intra-prediction mode from the existing peripheral reference samples, i.e., unfiltered peripheral reference samples. In this case, it may be called PDPC (Position dependent intra prediction). Alternatively, intra-prediction coding can be performed by selecting the reference sample line with the highest prediction accuracy from the peripheral multi-reference sample lines of the current block, deriving predicted samples using the reference samples located in the prediction direction on that line, and then instructing (signaling) the decoding device to use the reference sample line. In this case, it may be called MRL (multi-reference line intra prediction) or MRL-based intra-prediction. Additionally, the current block can be divided into vertical or horizontal subpartitions, and intra-prediction can be performed based on the same intra-prediction mode, but peripheral reference samples can be derived and used on a subpartition-by-subpartition basis.In other words, in this case, the intra-prediction mode for the current block is applied identically to the subpartition, but by deriving and utilizing surrounding reference samples on a subpartition-by-subpartition basis, the intra-prediction performance can be improved in some cases. Such a prediction method may be called ISP (intra sub-partitions) or ISP-based intra-prediction. The specific details will be described later. Furthermore, when the prediction direction based on the prediction sample indicates the area between surrounding reference samples, that is, when the prediction direction indicates a fractional sample position, the value of the prediction sample can be derived by interpolating multiple reference samples located around the prediction direction (around the fractional sample position).

[0128] The intra-prediction method described above may be called an intra-prediction type, distinct from the intra-prediction mode. The intra-prediction type may be referred to by various terms, such as an intra-prediction method or an additional intra-prediction mode. For example, the intra-prediction type (or additional intra-prediction mode, etc.) may include at least one of the LIP, PDPC, MRL, and ISP described above. Information regarding the intra-prediction type may be encoded by an encoding device and included in a bitstream, and signaled to a decoding device. Information regarding the intra-prediction type may be embodied in various forms, such as flag information indicating whether or not to apply each intra-prediction type, or index information indicating one of a plurality of intra-prediction types.

[0129] The MPM list for deriving the intra-prediction mode described above may be configured differently depending on the intra-prediction type. Alternatively, the MPM list may be configured commonly regardless of the intra-prediction type.

[0130] Derivation of peripheral reference samples

[0131] When intraprediction is applied to the current block, peripheral reference samples to be used for intraprediction of the current block may be derived. The peripheral reference samples of the current block may include a total of 2 x nH samples adjacent to the left boundary and bottom-left of the current block of nW x nH size, a total of 2 x nW samples adjacent to the top boundary and top-right of the current block, and one sample adjacent to the top-left of the current block. Alternatively, the peripheral reference samples of the current block may include upper peripheral samples for multiple columns and left peripheral samples for multiple rows. Furthermore, the peripheral reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block of nW x nH size, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right of the current block.

[0132] On the other hand, when MRL (Multiple Reference Line) is applied, the reference sample may be located on lines 1 to 3 on the left / above side, rather than line 0 which is adjacent to the current block. In this case, the number of peripheral reference samples may increase further. The specific regions and numbers of peripheral reference samples will be described later.

[0133] On the other hand, when ISP (Intra Sub-Partitions) is applied, the aforementioned peripheral reference samples can be derived on a sub-partition basis.

[0134] DIMD(Decoder side intra mode derivation)

[0135] In DIMD, the intra-prediction may be derived as a weighted average of the planar and two induced directions. For this purpose, two angular modes are selected from the HoG (Histogram of Gradient) calculated from the surrounding pixels of the current block. Once the two modes are selected, their predictors (prediction blocks) and planar predictors are normally calculated, and then the weighted average may be used as the final predictor (final prediction block) of the current block. In this case, the corresponding amplitudes in the HoG are used for each of the two modes to determine the weights.

[0136] Since the induced intra-mode is included in the primary list of the intra-MPM, the DIMD process may be performed before the MPM list is constructed. The temporarily induced intra-mode of a DIMD block is saved with the block and may be used to construct the MPM list of surrounding blocks.

[0137] The DIMD chroma mode can use the DIMD induction method to derive the chroma intra-prediction mode for the current block based on the already reconstructed peripheral Y, Cb, and Cr samples in the second peripheral row and column shown in Figure 12. Specifically, in order to build the HoG, horizontal and vertical gradients may be calculated for each collocated previously reconstructed chroma sample of the current chroma block, in addition to the already reconstructed Cb and Cr samples. Subsequently, chroma intra-prediction of the current chroma block may be performed using the intra-prediction mode with the maximum histogram amplitude value.

[0138] If the intra-prediction mode derived from the DIMD chroma mode is the same as the intra-prediction mode derived from the DM mode, the intra-prediction mode with the second largest histogram amplitude value may be used as the DIMD chroma mode. A predetermined CU level flag may be signaled to indicate whether or not the above-mentioned DIMD chroma mode is applied.

[0139] Fusion for TIMD (Fusion for template-based intra-mode derivation)

[0140] For each intra-prediction mode within the MPM, the SATD between the template's predicted sample and the recovered sample may be calculated. The first two intra-prediction modes with the smallest SATD may then be selected as the TIMD modes. These two TIMD modes may be fused using weighted values, and such weighted intra-predictions may now be used to code the CU. The derivation of the TIMD modes may include the position-dependent intra-prediction combination (PDPC) described above.

[0141] While comparing the costs of the two selected modes described above with a predetermined threshold, a cost factor 2 may be applied as shown in Equation 1 below.

[0142] [Formula 1]

number

[0143] If the condition in equation 1 is true, the fusion described above may be applied. Conversely, if the condition in equation 1 is false, only mode 1 may be used.

[0144] On the other hand, the weighted values ​​for the above modes may be calculated from the respective SATD costs as shown in the following equation 2.

[0145] [Formula 2]

number

[0146] MIP(Matrix-based intra prediction)

[0147] Figures 7 to 11 illustrate the MIP (Matrix-based intracellular prediction) process applicable to this disclosure.

[0148] MIP (Matrix-based intra prediction) may also be called ALWIP (Affine linear weighted intra prediction) or MIP (or MWIP) (Matrix weighted intra prediction). To predict samples for a rectangular block with width W and height H, MIP takes as input one line consisting of H reconstructed samples adjacent to the left boundary of the block, and one line consisting of W reconstructed samples adjacent to the top boundary. If reconstructed samples are unavailable, they may be generated in the same way as they would be generated by existing intra-prediction methods.

[0149] The prediction signal can be generated in basically the following steps.

[0150] 1. Of the boundary samples, if W=H=4, 4 samples can be extracted by averaging; otherwise, 8 samples can be extracted (averaging process).

[0151] 2. The matrix-vector multiplication performed after the offset addition may be done using the averaged samples as input. As a result of this process, a reduced prediction signal for the subsampled sample set of the original block may be obtained (matrix-vector multiplication process).

[0152] 3. Predicted signals at the remaining positions may be generated by linear interpolation, which is single-step linear interpolation in each direction from the predicted signals of the subsampling set ((linear) interpolation process).

[0153] The matrices and offset vectors required to generate the prediction signal (prediction block or prediction sample) may be obtained from three sets of matrices, S0, S1, and S2. Set S0 is an 18-matrix set. The image consists of JPEG2026511728000006.jpg693, with each matrix comprising 16 rows, 4 columns, and 18 offset vectors, each with a size of 16. It may have JPEG2026511728000007.jpg693. The matrices and offset vectors of this set may be used for a 4x4 size block. On the other hand, set S1 is a matrix of 10. The image consists of JPEG2026511728000008.jpg692, with each matrix comprising 16 rows, 8 columns, and 10 offset vectors, each with a size of 16. It may have JPEG2026511728000009.jpg692. The matrices and offset vectors of this set may be used for 4x8, 8x4, and 8x8 sized blocks. Finally, set S2 is a set of six matrices. The image consists of JPEG2026511728000010.jpg692, with each matrix having 64 rows, 8 columns, and 6 offset vectors, each with a size of 64. It may have JPEG2026511728000011.jpg692. The matrix and offset vector of this set may be used for any other size of block in addition to the size of block described above.

[0154] In matrix-vector product operations, the number of multiplication operations required may always be 4·W·H or less. In other words, in MIP mode, a maximum of four multiplication operations may be required for each sample.

[0155] Overview of the entire MIP process

[0156] The entire process of averaging, matrix-vector multiplication, and linear interpolation will be explained with reference to Figures 7 to 11. On the other hand, block forms not shown in Figures 7 to 11 may also be processed as shown in Figures 7 to 11.

[0157] 1. In a 4x4 block, the MIP may be calculated by taking two average values ​​along the axis of each boundary. The resulting four input samples may be input to a matrix vector multiplication. The matrix may be obtained from set S0. After offset addition, 16 final predicted samples may be obtained. Linear interpolation is not required to generate the predicted signal; that is, (4·16) / (4·4)=4 multiplications may be performed for each sample.

[0158] 2. In a 4x4 block, the MIP can take four average values ​​along the axis of each boundary. The resulting eight input samples may be input to a matrix vector multiplication. The matrix may be obtained from set S1. Thus, 16 samples may be obtained at odd positions in the prediction block. That is, (8·16) / (8·8)=2 multiplications may be performed for each sample. After offset addition, these samples may be vertically interpolated using the reduced upper boundary. Horizontal interpolation may be performed using the existing (original) left boundary. In this case, multiplication is not required in the interpolation process. That is, only a total of two multiplications per sample may be required to derive the MIP prediction.

[0159] In a 3.8x4 block, the MIP can take the four existing boundary values ​​of the left boundary and the four average values ​​along the horizontal axis of the boundary. The resulting eight input samples may be input to a matrix vector multiplication. The matrix may be obtained from set S1. Thus, 16 samples may be obtained at each vertical position and odd-numbered horizontal positions of the prediction block. That is, (8·16) / (8·4)=4 multiplications may be performed for each sample. After offset addition, these samples may be horizontally interpolated using the existing (original) left boundary. In this case, no additional multiplication process is required in the interpolation process. That is, only a total of four multiplications per sample may be necessary to derive the MIP prediction. Transposed cases may also be handled in this way.

[0160] 4. In a 16x16 block, the MIP can take four mean values ​​along each axis of the boundary. The resulting eight input samples may be input to a matrix vector multiplication. The matrix may be obtained from set S2. Thus, 64 samples may be obtained at odd positions in the prediction block. That is, each sample may undergo (8·64) / (16·16)=2 multiplications. After offset addition, these samples may be vertically interpolated using the eight mean values ​​at the upper boundary. Horizontal interpolation may be performed using the existing (original) left boundary. In this case, no additional multiplication process is required in the interpolation process. That is, only a total of two multiplications per sample may be necessary to derive the MIP prediction.

[0161] The process can be applied essentially the same way to larger blocks. In this case, it can be seen that the number of multiplication operations for each sample is less than 4.

[0162] In a Wx8 block where W>8, since samples are provided at odd horizontal and vertical positions, only horizontal interpolation may be required. In this case, to calculate the reduced prediction, a multiplication of (8·64) / (W·8)=64 / W may be performed for each sample. When W=16, no additional multiplication is required for linear interpolation, but when W>16, the number of additional multiplications per sample required for linear interpolation may be less than two. Therefore, the total number of multiplications per sample may be 4 or less.

[0163] Finally, in the case of a W x 4 block where W > 8, A k Assume that this is the matrix obtained by excluding all rows corresponding to odd items along the horizontal axis of the downsampled block. In this case, the output size is 32, and only horizontal interpolation can be performed. For the reduced prediction calculation, a multiplication of (8·32) / (W·4)=64 / W may be performed for each sample. If W=16, no additional multiplication is required, but if W>16, fewer than two multiplications may be required for each sample for linear interpolation. Therefore, the total number of multiplications may be 4 or less. The transposed case may also be handled in this way.

[0164] Boundary averaging

[0165] Figure 12 shows an example of a boundary averaging process applicable to this disclosure. For example, boundary averaging may be included in the MIP process.

[0166] Averaging can be applied to each boundary surface (left boundary surface or upper boundary surface) through an averaging process. Here, the boundary surface can mean the adjacent reference sample adjacent to the current block boundary, as explained with reference to Figures 7 to 11. For example, the left boundary surface (bdry left ) represents the left neighboring reference sample adjacent to the left boundary of the current block, and the upper boundary surface (bdry top) may represent the upper-next reference sample adjacent to the current upper boundary of the block. If the current block size is 4x4, the averaging process may reduce each boundary size to 2 samples. If the current block size is not 4x4, the averaging process may reduce each boundary size to 4 samples.

[0167] In the first stage, the input boundary is bdry. top and bdry left is a smaller boundary. It is acceptable to reduce the file size to JPEG2026511728000012.jpg737. Here, Each JPEG2026511728000013.jpg737 consists of two samples in the case of a 4x4 block, and may consist of four samples in all other cases.

[0168] For a 4x4 block, the definition may be as follows for 0 ≤ i < 2.

[0169] [Formula 3]

number

[0170] and, JPEG2026511728000015.jpg786 can also be defined in the same way as mathematical formula 3.

[0171] In other cases, the block width W is W = 4.2 k If so, then for 0 ≤ i < 4, it may be defined as follows:

[0172] [Equation 4]

number

[0173] and, JPEG2026511728000017.jpg786 can also be defined in the same way as mathematical formula 4.

[0174] Such two reduced boundaries JPEG2026511728000018.jpg737 is connected to the reduced boundary vector bdry red In the case of a 4x4 block form, the size may be 4, and in the case of other block forms, the size may be 8

[0175] On the other hand, when the variable mode indicates the MIP mode, the following concatenation may be defined

[0176] [Equation 5] [Number]

[0177] Finally, in the interpolation of the subsampled prediction signal, in the case of a large block, a corrected version of the averaged boundary, that is, the second version, may be further required. That is, min(W, H)>8 and W≧H, W = 8*2 l For 0≦i<8, the following may be defined

[0178] [Equation 6] [Number]

[0179] If min(W, H)>8 and H>W JPEG2026511728000021.jpg787 may be defined similarly

[0180] Reduced prediction signal generation by matrix-vector multiplication

[0181] One reduced input vector bdry red can generate a reduced prediction signal pred red The reduced prediction signal has a width W red and a height H redThe signal may be a downsampled block. Here, W red and H red It can be defined as follows:

[0182] [Equation 7]

number

[0183] Reduced predictive signal red This can be derived by calculating the matrix-vector product and adding the offsets.

[0184] [Equation 8]

number

[0185] Here, A is W red ·H red It may be a matrix having n rows and 4 columns when W=H=4, and 8 columns otherwise. Here, b is W red ·H red It can be a vector of sizes.

[0186] Matrix A and vector b may be obtained from any one of the sets S0, S1, or S2 as follows. The index idx = idx(W, H) may be defined as follows.

[0187] [Formula 9]

number

[0188] Then, if idx <= 1 or idx = 2 and min(W, H) > 4, It may be 638 for JPEG2026511728000026.jpg. When idx = 2 and min(W, H) = 4, if W = 4, A may correspond to the odd x - coordinates of the downsampled block, or if H = 4, A may correspond to the odd y - coordinates of the downsampled block.

[0189] Finally, in the following cases, the reduced prediction signal may be replaced with the transposed signal. JPEG2026511728000027.jpg22110

[0190] pred red The total number of multiplications required for the calculation of may be 4 when W = H = 4, but in this case, because A has 16 rows and 4 columns. In other cases, A has 8 columns and W red ·H red rows, and in this case, it can be seen that 8·W red ·H red <=4·W·H multiplications are required. That is, in this case, a maximum of 4 multiplications may be required for each sample to calculate pred red

[0191] Linear interpolation

[0192] FIG. 13 is a diagram showing an example of a linear interpolation process applicable to the present disclosure. As an example, the linear interpolation process may be included in the MIP process. As an example, the interpolation process may be a linear interpolation or a bilinear interpolation process. The interpolation process may include two steps consisting of 1) vertical interpolation and 2) horizontal interpolation as shown in FIG. 13. If W >= H, vertical linear interpolation may be applied first, and then horizontal linear interpolation may be applied. If W < H, horizontal linear interpolation may be applied first, and then vertical linear interpolation may be applied. In a 4x4 block, the interpolation process may be omitted.

[0193] In a W×H block where max(W, H)≧8, the prediction signal is linearly interpolated by W red ×H red ​The reduced prediction signal pred red may be generated based on this. Based on the block form, linear interpolation may be applied in the vertical direction, the horizontal direction, or both directions. When linear interpolation is applied in both directions, it may first be applied in the horizontal direction, or if W < H, it may first be applied in the vertical direction.

[0194] Assume there is a WxH block where max(W,H) >= 8 and W >= H. Here, 1D linear interpolation may be performed as follows. Considering generality, linear interpolation in the vertical direction may be performed. First, the reduced prediction signal may be extended to the upper end by the boundary signal. Let the vertical upsampling factor be JPEG2026511728000028.jpg674. In this case, the extended reduced prediction signal may be defined as follows.

[0195] [Equation 10] [Number]

[0196] In this case, the vertical linear interpolation prediction signal may be generated from the extended reduced prediction signal as follows.

[0197] [Equation 11] [Number]

[0198] Examples

[0199] This disclosure relates to intra (in-screen) prediction, specifically to a technique for performing MIP (Matrix-based Intra Prediction) based intra-screen prediction using only a portion of the reference samples. Figure 14 is a diagram illustrating an MIP prediction method applicable to this disclosure. According to an example in Figure 14, for the prediction of the current block, a predicted value of the current block can be obtained as an output value through a neural network using surrounding reference samples as input values.

[0200] On the other hand, current MIP prediction can generate predicted blocks using all reference samples on the left and top edges surrounding a block. However, the samples of the current block may have biased correlations due to some reference samples, such as those on the top or left edge. Therefore, using partial reference samples to predict the current block can yield higher prediction accuracy.

[0201] Therefore, this disclosure proposes a method for performing MIP prediction adaptively to the reference sample region. Specifically, it proposes a method for performing MIP prediction using only a portion of the surrounding reference samples of the current prediction block.

[0202] The embodiments of this disclosure will be described in detail below with reference to the drawings.

[0203] Figure 15 is a diagram illustrating a method for performing MIP prediction according to one embodiment of the present disclosure, and Figure 16 is a diagram illustrating an MIP-based video coding or decoding method according to one embodiment of the present disclosure. In this embodiment, a method for performing MIP prediction using only a portion of the reference samples from the current block will also be described with reference to Figures 14, 15, and 16.

[0204] On the other hand, the video encoding method or video decoding method shown in Figure 16 may be performed by a video encoder or video decoder, respectively, and the video encoder and video decoder may include the devices described above with reference to other drawings.

[0205] First, an MIP-based video encoding / decoding method (S1610) according to one embodiment of the present disclosure may be performed. For example, the S1610 method may be performed based on a selected set of reference samples.

[0206] For example, to perform MIP-based video coding or decoding, some of the left-side reference samples and top-edge reference samples may be selected. For example, the top-edge reference sample of the current block may be selected to be used for MIP (S1620). As mentioned above, the pixels of the current block may have a biased correlation with some reference samples, such as the top-edge or left-side reference samples. Therefore, when predicting the current block using partial reference samples, higher prediction accuracy can be obtained. Here, according to one embodiment of the present disclosure shown in Figure 15, only the top-edge reference sample can be selected for MIP prediction and used as the MIP prediction input value. In this case, only the top-edge reference sample may be selected as the neural network input value (input) for MIP prediction. On the other hand, as shown in Figure 15(a), the top-edge reference sample may correspond to a sample directly adjacent to the current block. That is, according to the embodiment in Figure 15(a), only the reference sample directly above the current block can be used. However, as another example, as shown in Figure 15(b), the top reference sample may include a number of adjacent samples (e.g., a form further extended to the right) adjacent to the top edge of the current block that is greater than the width of the current block, or it may include reference samples of multiple reference sample lines. According to the embodiment in Figure 15(b), multiple lines can also be used for the top edge and top-right reference samples of the current block. As yet another example, as shown in Figure 15(c), the top reference sample may include a number of adjacent samples (e.g., a form further extended to the left and / or right) adjacent to the top edge of the current block that is greater than the width of the current block, or it may include reference samples of multiple reference sample lines. According to the embodiment in Figure 15(c), the top edge, top-right, or top-left reference samples of the current block can also be used. On the other hand, the number of reference samples further extended to the left or right, or the number of reference sample lines, is not limited to those disclosed in Figure 15.Furthermore, for the method proposed in this embodiment, some samples within one or more reference sample lines may be used for MIP prediction, and currently, only a portion of the upper reference samples of the block may be used for MIP prediction as proposed in this embodiment.

[0207] Subsequently, an MIP mode may be selected (S1630) based on the MIP mode information. The MIP mode may be applied to the current block and may be explicitly signaled in the bitstream.

[0208] Subsequently, a neural network-based MIP may be applied (S1640) based on the selected MIP mode. In other words, an MIP-based prediction may be performed. Alternatively, as an example, based on the selected reference sample input, the predicted block sample value, which is the output, may be obtained through a neural network, as shown in Figures 14 and / or 15. Here, the neural network may be adaptively selected based on the reference sample input. In other words, step S1640 may include a process in which an MIP-based prediction is applied based on the selected input sample value and the selected neural network. As an example, a neural network consisting of a single matrix may be used to perform the MIP. On the other hand, the predicted block sample value, which is the output, may be derived as follows.

[0209] [Formula 12]

number

[0210] In the above formula, r represents the reference sample input value selected by the method proposed in this embodiment, and pred represents the predicted block sample value, which is the output value. In other words, the predicted block sample value may be derived based on the selected reference sample value. For example, A may represent a trained neural network matrix, and the Clip function may be a clipping function that adjusts (clips) the output predicted sample (pred) value to a certain range of sample values, and limits the lower and / or upper limits. In this case, the clipping range may be predefined. That is, the output value pred may be obtained by an operation on a matrix A with respect to the input value r. The size of matrix A may be adaptively changed depending on the size of the input and output values.

[0211] On the other hand, the neural network configuration for obtaining the output value pred for the input value r proposed in this embodiment is not limited to the above example. That is, in addition to a neural network using one simple matrix, multiple matrix operations may be used by configuring multiple hidden layers, and bias terms and activation functions may be used for each neural network layer.

[0212] Furthermore, when outputting predicted samples, values ​​may be output in the frequency domain in addition to the spatial domain. In this case, an inverse transform may be performed on the predicted values ​​to obtain the final predicted sample values.

[0213] On the other hand, as illustrated in the above example, the size of the neural network matrix may be adaptively selected based on the input and output values. In this embodiment, the selected input values, i.e., the selected reference samples, can be sampled using an appropriate downsampling method and used as input to the neural network. That is, reference samples sampled by downsampling the selected reference samples at a ratio of 1:2, 1:4, or 1:8 may be used as input values ​​to the neural network. In this case, the size of the input values ​​can be reduced, thereby reducing the size of the neural network matrix and significantly reducing the overall MIP computation time.

[0214] On the other hand, the sampling method and ratio of the reference samples selected in this example are not limited to the above example, and the input value may be increased by upsampling in addition to downsampling. The sampling method and ratio of the reference samples can be adaptively selected depending on the shape and size of the block and the number of reference samples selected. Furthermore, the selected reference samples can be used as neural network input values ​​after preprocessing with various filtering methods such as low-pass filtering and high-pass filtering. It is also possible to apply a transform to the selected reference samples to change them to the frequency domain before inputting them into the neural network, and after calculating the mean of the selected reference samples, the residual sample value obtained by subtracting the mean from all reference samples can also be used as the neural network input value.

[0215] On the other hand, the filtering and sampling methods for input values ​​described above can also be applied to output values. That is, if the number of output values ​​in the neural network, i.e., the number of output samples, is less than or greater than the predicted number of samples required for the current block size, the output values ​​can be upsampled or downsampled to match the predicted number of samples required for the current block size. Furthermore, if a transformation is applied to the input values ​​and frequency domain input values ​​are used, or if residual samples obtained by subtracting the mean are used as input values, this can also be corrected in the output values.

[0216] As illustrated above, the number of input and output samples for the neural network matrix operation proposed in this disclosure is not limited to a specific number.

[0217] Furthermore, post-processing filtering can be applied to the prediction blocks predicted by the method proposed in this disclosure. A Position-Dependent Intra Prediction Combination (PDPC) method can also be applied to the prediction blocks predicted by the method proposed in this disclosure. As an example, a smoothing filter can also be applied to the prediction blocks predicted by the method proposed in this invention.

[0218] On the other hand, as an example, when constructing a neural network for the method proposed in this embodiment, the coefficients of the matrix used in neural network operations may have various bit precisions. For example, when performing 10-bit precision matrix operations, the matrix coefficients may have a range of coefficient values ​​within a specific range (e.g., 0 to 1023 or -512 to 511), and the input and output values ​​may be adaptively adjusted accordingly. Such bit precision values ​​are not limited to the above examples and may be adaptively changed based on the number of layers in the neural network, etc.

[0219] On the other hand, in the method proposed in this embodiment, the number of neural matrix modes (the number of modes in the proposed MIP method) can be configured in various ways based on the mode (e.g., intra-mode or inter-mode), the width of the input block, the height of the input block, the number of pixels in the input block, the position of subblocks within the block, the explicitly signaled syntax elements, the statistical characteristics of surrounding pixels, and whether or not a quadratic transformation is used. For example, the number of neural matrix modes can be adaptively selected based on the current block size. Alternatively, the same number of neural matrix modes can be selected for all block sizes.

[0220] For example, the coefficients of the neural network matrix may be defined adaptively for all block configurations. For example, neural network coefficients corresponding to each block configuration from 4x4 to 256x256 blocks may be defined. The number of input and output values ​​of the neural network matrix may be predetermined according to the block configuration. For example, as illustrated in Figure 15(a), MIP prediction is performed using only one upper reference sample line, and 49 types of neural network matrices may be required that match the input values ​​of the neural network matrix of 4, 8, 16, 32, 64, 128, 256 and the output values ​​of 4, 8, 16, 32, 64, 128, 256. Furthermore, each of the 49 types of neural network matrices may have matrix coefficients based on the number of MIP modes proposed in this embodiment. For example, if the number of MIP modes in a neural network matrix with an input value of 8 and an output value of 16 (which may correspond to an 8x16 block configuration) is 35, then the neural network matrix may have a total of 8x16x35 matrix coefficients. Alternatively, as another example, only neural network matrices that fit a specific block shape may be defined, and they may have limited input and output values. In that case, the neural network matrices for the specific block shape may be utilized by appropriate pre-processing or post-processing of the input and output values. For example, if only neural network matrices that fit a square block shape are defined, the input values ​​of the square block neural network matrix can be made to correspond to the reference samples for non-square blocks by pre-processing (e.g., down-sampling or up-sampling). The output values ​​can also be made to correspond to the predicted values ​​for non-square blocks by post-processing (e.g., down-sampling or up-sampling).

[0221] On the other hand, according to the method proposed in this embodiment, the neural matrix mode (the mode of the proposed MIP method) may be selected based on at least one of the following: mode (e.g., intermode or intramode), input block width, input block height, number of pixels in the input block, subblock position within the block, explicitly signaled syntax elements, statistical characteristics of surrounding pixels, and whether or not a quadratic transformation is used. In this case, the encoder can transmit the selected mode information to the decoder by selecting and applying an appropriate binary method. For example, the encoder may set a specific mode to an MPM mode, such as the transmission of an intralumer mode, and then perform binary conversion on the remaining modes after setting them separately. In this case, binary bits can also be saved by appropriate context modeling during each binary conversion. Alternatively, as another example, the mode information can be transmitted to the decoder using binary conversion such as truncated binary / truncated unary / fixed length, taking into account the total number of MIP modes.

[0222] On the other hand, in order to select the type of primary and secondary transformation in a predicted block predicted by the method proposed in this embodiment, the transformation mode may be selected based on at least one of the following: neural matrix mode, input block width, input block height, number of pixels in the input block, subblock position within the block, explicitly signaled syntax elements, and statistical characteristics of surrounding pixels. For example, when selecting the type of primary transformation for a block to which the prediction method proposed in this disclosure is applied, the MTS (multiple transform set) transformation method may be applied. For example, when selecting the type of secondary transformation for a block to which the prediction method proposed in this disclosure is applied, a secondary transformation kernel that matches the planar intra mode may be selected.

[0223] On the other hand, for the MIP mode selection proposed in this embodiment, candidate MIP prediction modes for the current block may be determined, such as in DIMD or TIMD, or the MIP prediction mode for the current block may be determined. That is, candidate MIP prediction modes or the MIP prediction mode for the block may be inferred by utilizing surrounding reference samples, such as in DIMD or TIMD.

[0224] As another example, the mode candidate analogy method used in DIMD may be applied to predict the MIP mode in the method proposed in this embodiment. That is, a pixel gradient can be calculated using only the region of the upper reference sample, and the first or second MIP prediction mode obtained from this gradient may be set as the MIP prediction mode candidate for the current block.

[0225] As another example, the mode analogy method used in TIMD for MIP mode prediction in the method proposed in this embodiment may be applied. That is, TIMD template matching may be performed using only the region of the upper reference sample. The MIP prediction mode candidates to which TIMD template matching is applied may be selected as MPM, or all of the multiple MIP prediction modes may be selected.

[0226] On the other hand, whether or not the method proposed in this embodiment is applied may be signaled by information in HLS (high-level syntax) such as VPS, SPS, PPS, Picture Header, Slice Header, DCI, etc. As an example, in order to determine whether or not the method proposed in this embodiment is applied on a PPS basis, information on whether or not the method proposed in this embodiment is applied may be included in the PPS and signaled.

[0227] Additional information regarding the application of the method proposed in this embodiment may be explicitly signaled, but as an alternative, the method proposed in this embodiment may be inferred or adaptively selected without additional information signaling. For example, the application of the method proposed in this embodiment may be signaled by a 1-bit flag on a CTU or CU basis.

[0228] For example, the method proposed in this embodiment may only be used if the MIP mode is defined to be used in HLS (high-level syntax). Furthermore, the application of the method proposed in this embodiment may be determined by signaling additional information within the MIP mode. For example, if the value of information regarding the application of MIP (e.g., MIP flag) is a specific value (e.g., 1, true), the information regarding the application of the method proposed in this embodiment may be signaled by a 1-bit flag.

[0229] Alternatively, if the method proposed in this embodiment can be applied depending on the size or shape of a particular block, or the presence or absence of certain conditions, then, only in that case, the applicability of the method proposed in this embodiment may be signaled by a 1-bit flag. For example, if the height of a block is four times or more the width of the block, the method proposed in this embodiment does not need to be applied, and therefore, additional information signaling regarding its applicability may be omitted.

[0230] On the other hand, as another example, under certain conditions, the applicability of the method proposed in this embodiment may be implicitly inferred. For example, when there is no left-side reference sample for the current block, such as a block at the left-side boundary of the image, or when a left-side reference sample cannot be used, signaling information regarding the applicability of the method proposed in this embodiment may be omitted. Signaling information regarding the applicability of MIP mode may also be omitted. In this case, the system may always guide the user to apply the MIP method proposed in this embodiment. When the left-side reference sample for the current block is a CTU boundary / tile boundary / slice boundary / sub-picture boundary, signaling information regarding the applicability of the method proposed in this embodiment may be omitted. Signaling information regarding the applicability of MIP mode may also be omitted. In this case, the system may always guide the user to apply the MIP method proposed in this embodiment.

[0231] On the other hand, as an example, the decision of whether or not to apply the proposed method to this embodiment at a higher level defined by HLS (high-level syntax) may adaptively determine whether or not information transmission regarding the application of the proposed method to this embodiment is performed at the coding unit level. For example, if the value of the information signaled by SPS indicating whether or not the proposed method is applied to this embodiment is a specific value (e.g., 0, false) (i.e., indicating that the proposed method is not used on an SPS basis for this embodiment), then the proposed method is not applied to this embodiment at the coding unit level, and signaling of information regarding its use or non-use may not be performed.

[0232] Figures 17 and 18 illustrate a method for performing MIP prediction according to one embodiment of the present disclosure, and Figure 19 illustrates a MIP-based video coding or decoding method according to one embodiment of the present disclosure. In this embodiment, a method for performing MIP prediction using only a portion of the reference samples from the current block will also be described with reference to Figures 17 to 19.

[0233] On the other hand, the video encoding method or video decoding method shown in Figure 19 may be performed by a video encoder or video decoder, and the video encoder and video decoder may include the devices described above with reference to other drawings.

[0234] First, an MIP-based video encoding / decoding method (S1910) according to one embodiment of the present disclosure may be performed. For example, the S1910 method may be performed based on a selected set of reference samples.

[0235] For example, to perform MIP-based video coding or decoding, some of the left-hand reference samples and top-edge reference samples may be selected. For example, the left-hand reference sample of the current block may be selected to be used for MIP (S1920). As mentioned above, the pixels of the current block may have a correlation that is relatively biased towards some reference samples, such as the top-edge or left-hand reference sample. Therefore, when predicting the current block using partial reference samples, higher prediction accuracy can be obtained. Here, according to one embodiment of the present disclosure shown in Figure 18, only the left-hand reference sample can be selected for MIP prediction and used as the MIP prediction input value. In this case, only the left-hand reference sample may be selected as the neural network input value (input) for MIP prediction. On the other hand, as shown in Figure 18(a), the left-hand reference sample may be a sample directly adjacent to the current block. That is, according to the embodiment in Figure 18(a), only the reference sample immediately to the left of the current block can be used. However, as another example, as shown in Figure 18(b), the left-side reference sample may include a number of adjacent samples adjacent to the left side of the current block that exceeds the height of the current block (e.g., a form further extended to the bottom), or the left-side reference sample may include reference samples of multiple reference sample lines. According to the embodiment in Figure 18(b), multiple reference sample lines at the top and left-bottom of the current block can also be utilized. As yet another example, as shown in Figure 18(c), the left-side reference sample may include a number of adjacent samples adjacent to the left side of the current block that exceeds the height of the current block (e.g., a form further extended to the top and / or bottom), or the reference sample of multiple reference sample lines can be utilized. According to the embodiment in Figure 18(c), reference samples at the left side, left-top, or left-bottom of the current block can be utilized. On the other hand, the number of reference samples or reference sample lines further extended to the top or bottom is not limited to those disclosed in Figure 18.Also, for the method proposed in this embodiment, some samples in one or more reference sample lines may be utilized for MIP prediction, or only a partial region of all the left reference samples of the current block may be utilized for the MIP prediction proposed in this embodiment.

[0236] After that, an MIP mode may be selected (S1930) based on the MIP mode information. The MIP mode may be one applied to the current block or may be explicitly signaled in the bitstream.

[0237] After that, based on the selected MIP mode, neural network-based MIP may be applied (S1940). In other words, MIP-based prediction may be performed. On the other hand, as an example, based on the selected reference sample input value (input), as shown in FIG. 14 and / or FIG. 17, a predicted block sample value, which is an output value (output), may be obtained through a neural network. Here, the neural network may be adaptively selected based on the reference sample input value. In other words, the process of applying MIP-based prediction based on the selected input sample value and the selected neural network may be included in step S1940.

[0238] As an example, regarding the configuration of the neural network, the neural network matrix, the input value of the neural network, the processing for the output value, the signaling of information related to the embodiment, etc., it is as described in the above explanations referring to FIGS. 14, 15, and 16, and the overlapping explanations are omitted.

[0239] FIGS. 20 and 21 are diagrams for explaining a method of performing MIP prediction according to an embodiment of the present disclosure, and FIG. 22 is a diagram for explaining a method of MIP-based video encoding or decoding according to an embodiment of the present disclosure. In this embodiment, referring to FIGS. 20, 21, and 22, a method of performing MIP prediction using only some of the reference samples of the current block will also be explained.

[0240] On the other hand, the video encoding method or video decoding method shown in Figure 22 may be performed by a video encoder or video decoder, and the video encoder and video decoder may include the devices described above with reference to other drawings.

[0241] First, an MIP-based video encoding / decoding method (S2210) according to one embodiment of the present disclosure may be performed. For example, the S2210 method may be performed based on a selected set of reference samples.

[0242] As an example, some of the left-side reference samples and top-edge reference samples may be selected for MIP-based video coding or decoding. As an example, only some of the left-side and top-edge reference samples of the current block may be selected to be used for MIP (S2220). That is, in the embodiments of Figures 20 and 21, some or extended left-side reference samples and some or extended top-edge reference samples may be selected and used. As mentioned above, pixels in the current block may have a more biased correlation than some reference samples such as top-edge or left-side reference samples, and therefore, when predicting the current block using some or extended reference samples, higher prediction accuracy can be obtained. Here, according to one embodiment of the present disclosure shown in Figure 21, some or extended reference samples from adjacent reference samples located on the left side or top edge can be selected and used as MIP prediction input values ​​for MIP prediction. In this case, some left-side reference samples and some top-edge reference samples may be selected as neural network input values ​​(input) for MIP prediction. On the other hand, as shown in Figure 21(a), the left-side reference sample and the top-end reference sample may correspond to samples directly adjacent to the current block. That is, according to the embodiment in Figure 21(a), only the reference sample immediately to the left of the current block can be used, or only the reference sample immediately above the current block can be used. In other words, only reference samples adjacent to the current block may be used. However, as seen in Figure 21(a), while reference samples directly adjacent to the current block are used, the block located at the top of the left-side reference sample corresponding to the height of the current block is excluded, and the left-side reference sample further extended to the bottom may be included in the reference sample. Also, as seen in Figure 21(a), while reference samples directly adjacent to the current block are used, the block located at the left of the top-end reference sample corresponding to the width of the current block is excluded, and the top-end reference sample further extended to the right may be included in the reference sample.However, as another example, as shown in Figure 21(b), the left-side reference sample may include a number of adjacent reference samples adjacent to the left side of the current block, but fewer than the number of samples corresponding to the height of the current block, or the left-side reference sample may include reference samples from multiple reference sample lines. Also, as shown in Figure 21(b), the top-end reference sample may include a number of adjacent reference samples adjacent to the top end of the current block, but fewer than the number of samples corresponding to the width of the current block, or the top-end reference sample may include reference samples from multiple reference sample lines. Furthermore, the reference samples may include samples located at the top left end of the current block. That is, reference samples in a form further extended to the left from the top-end reference sample or a form further extended to the top from the left-side reference sample may be selected. According to the embodiment in Figure 21(b), reference samples from multiple lines on the left side, top end, and / or left-top end of the current block can be utilized. On the other hand, as yet another example, as shown in Figure 21(c), the left-side reference samples of the current block include reference samples corresponding to the height of the current block, but the top-end reference samples may include only a number of adjacent samples that are fewer than the number of reference samples corresponding to the width of the current block. In other words, only a portion of the top-end adjacent reference samples may be included in the top-end reference samples.

[0243] On the other hand, the number of reference samples or reference sample lines further extended to the upper or lower end is not limited to those disclosed in Figure 21. Furthermore, for the method proposed in this embodiment, some samples within one or more reference sample lines may be used for MIP prediction, and only a portion of all left-hand reference samples or upper-end reference samples in the current block may be used for MIP prediction as proposed in this embodiment.

[0244] Subsequently, an MIP mode may be selected (S2230) based on the MIP mode information. The MIP mode may be applied to the current block, or it may be explicitly signaled in the bitstream.

[0245] Subsequently, a neural network-based MIP may be applied (S2240) based on the selected MIP mode. In other words, MIP-based prediction may be performed. On the other hand, as an example, based on the selected reference sample input values, predicted block sample values, which are output values, may be obtained through a neural network, as shown in Figure 14 and / or Figure 20. Here, the neural network may be adaptively selected based on the reference sample input values. In other words, step S2240 may include a process in which MIP-based prediction is applied based on the selected input sample values ​​and the selected neural network. As an example, the neural network configuration, neural network matrix, processing of the neural network input values ​​and output values, and signaling of information related to the embodiment have been explained above with reference to Figures 14, 15, 16, 17, 18, 19, 20, and 21, and a redundant explanation will be omitted.

[0246] On the other hand, information related to the MIP according to one embodiment of the present disclosure may be signaled. The following describes the signalable information with reference to an exemplary table, but the names and order of the information described may be changed, and the information may be signaled at various levels (e.g., SPS (Sequence parameter set), PPS (Picture Parameter set), Slice Header, Picture Header, CU (Coding Unit), PU (Prediction Unit), etc.).

[0247] Table 3 below is an example of signaling information related to MIP according to one embodiment of the present disclosure.

[0248] [Table 3]

[0249] For example, the first syntax, intra_mip_flag, may be information for determining whether or not MIP is applied. For example, if the value of the first syntax is a first value (e.g., 1), it may be determined that MIP is applied, and if the value of the first syntax is a second value (e.g., 0), it may be determined that MIP is not applied. The first syntax may be obtained from the bitstream based on another syntax (e.g., sps_mip_enabled_flag). For example, the other syntax sps_mip_enabled_flag may be information indicating whether or not MIP is allowed. For example, if the first syntax is not present in the bitstream, the value of the first syntax may be induced to be a specific value (e.g., 0). For example, the second syntax, intra_mip_above_flag, may be information for determining whether or not MIP is applied that uses only upper-end reference samples. This may be the information for determining whether or not MIP is applied as described above with reference to Figures 14, 15, and 16. For example, if the value of the second syntax is the first value (e.g., 1), it may be determined that the MIP that uses only upper-edge reference samples is applied, and if the value of the second syntax is the second value (e.g., 0), it may be determined that the MIP that uses only upper-edge reference samples is not applied. For example, the second syntax may be obtained from the bitstream based on another syntax (e.g., the first syntax).

[0250] For example, the third syntax, intra_mip_above_mode, may be information for determining the mode of the MIP that uses only upper-end reference samples. This may be information for determining the mode of the MIP as described above with reference to Figures 14, 15, and 16. For example, the third syntax may be obtained from the bitstream based on another syntax (e.g., the second syntax). For example, the third syntax may be obtained from the bitstream only when the value of the second syntax is a specific value (e.g., 1).

[0251] On the other hand, if the value of the second syntax is a specific value (e.g., 0), other syntaxes other than the third syntax may be obtained. In this case, it is determined that MIP using general left-hand and top-hand reference samples is applied, and general MIP-related information (e.g., intra_mip_transposed_flag and / or intra_mip_mode) may be obtained.

[0252] On the other hand, Table 4 below provides other examples of information signaling related to MIP according to one embodiment of the present disclosure.

[0253] [Table 4]

[0254] For example, the first syntax, intra_mip_flag, may be information for determining whether or not MIP is applied. This is explained with reference to Table 3, and a redundant explanation will be omitted. For example, the fourth syntax, intra_mip_left_flag, may be information for determining whether or not MIP is applied that uses only left-reference samples. This is the same information for determining whether or not MIP is applied as explained above with reference to Figures 17 to 19. For example, if the value of the fourth syntax is the first value (e.g., 1), it is determined that MIP using only left-reference samples is applied, and if the value of the fourth syntax is the second value (e.g., 0), it is determined that MIP using only left-reference samples is not applied. For example, the fourth syntax may be obtained from the bitstream based on other syntax (e.g., the first syntax).

[0255] As an example, the fifth syntax intra_mip_left_mode may be information for determining a mode of MIP that uses only left reference samples. This may be information for determining the mode of MIP described above with reference to FIGS. 17 to 19. As an example, the fifth syntax may be obtained from a bitstream based on other syntaxes (e.g., the fourth syntax). For example, the fifth syntax may be obtained from the bitstream only when the value of the fourth syntax is a specific value (e.g., 1).

[0256] On the other hand, if the value of the fourth syntax is a specific value (e.g., 0), other syntaxes other than the fifth syntax may be obtained. In this case, it is determined that MIP using general left reference samples and top reference samples is applied, and general MIP related information (e.g., intra_mip_transposed_flag and / or intra_mip_mode) etc. may be obtained.

[0257] On the other hand, Table 5 below is another exemplification of signaling information related to MIP according to an embodiment of the present disclosure.

[0258]

Table 5

[0259] For example, the first syntax, intra_mip_flag, may be information for determining whether or not MIP is applied. This has been explained with reference to Table 3, and a redundant explanation will be omitted. For example, the sixth syntax, intra_mip_extended_flag, may be information for determining whether or not MIP is applied that uses only the upper reference sample or the left reference sample. This may be the information for determining whether or not MIP is applied as explained above with reference to Figures 14 to 22. For example, if the value of the sixth syntax is a first value (e.g., 1), it is determined that MIP using only the upper or left reference sample is applied, and if the value of the sixth syntax is a second value (e.g., 0), it is determined that MIP using only the upper or left reference sample is not applied. For example, the sixth syntax may be obtained from the bitstream based on other syntax (e.g., the first syntax).

[0260] For example, the second syntax, intra_mip_above_flag, may be obtained based on the value of the sixth syntax. For example, if the value of the sixth syntax is a specific value (e.g., 1), the second syntax may be obtained. The second syntax may be information for determining whether or not to apply MIP that uses only upper-level reference samples, as described above. However, for example, if the value of the second syntax in Table 5 is the first value (e.g., 1), it may be determined that MIP that uses only upper-level reference samples is applied, but if it is the second value (e.g., 0), it may be determined that MIP that uses only left-level reference samples is applied.

[0261] Subsequently, the third or fifth syntax may be obtained based on the value of the second syntax. For example, if the value of the second syntax is the first value (e.g., 1), the MIP that uses only upper reference samples is applied, and the third syntax intra_mip_above_mode may be obtained. The explanation for the third syntax is the same as that explained above with reference to Table 3, and the redundant explanation is omitted. As another example, if the value of the second syntax is the second value (e.g., 0), the MIP that uses only left reference samples is applied, and the fifth syntax intra_mip_left_mode may be obtained. The explanation for the fifth syntax is the same as that explained above with reference to Table 4, and the redundant explanation is omitted.

[0262] On the other hand, if the value of the sixth syntax is a specific value (e.g., 0), then other syntaxes other than the second, third, or fifth syntax may be obtained. In this case, it is determined that a MIP using general left-hand and top-end reference samples is applied, and general MIP-related information (e.g., intra_mip_transposed_flag and / or intra_mip_mode) may be obtained.

[0263] On the other hand, Tables 3 to 5 above signal separately whether or not MIP is applied using only the upper reference sample and whether or not MIP is applied using only the left reference sample; this is for clarity of explanation. Therefore, according to other embodiments of this disclosure, it is also possible to indicate as a single piece of information whether or not MIP is applied using only the upper reference sample, whether or not MIP is applied using only the left reference sample, or whether or not MIP is applied using only a portion of the left and upper reference samples, and it is clear that other embodiments are also included in this disclosure.

[0264] Figure 23 shows a video decoding method performed by a video decoding device (decoder) according to one embodiment of the present disclosure.

[0265] As an example, the video decoding device may determine whether or not to apply the MIP (matrix-based intra prediction) mode to the current block (S2310). As an example, the MIP mode whose application is determined at this stage may include the MIP described with reference to Figures 14 to 22. As an example, whether or not to apply the MIP mode may be determined based on the MIP application information described above, and this information may be obtained from the bitstream. On the other hand, step S2310 may further include a step of determining whether the MIP mode is applied using only the left reference sample or the top reference sample, or using only the left reference sample and a portion of the top reference sample.

[0266] As an example, a predicted block of the current block may be generated (S2320) based on the determination that MIP mode is applied to the current block. The predicted block may be generated using only a portion of the left-side reference samples or top-side reference samples of the current block based on MIP mode. As an example, when the predicted block is generated using only the top-side reference samples of the current block based on MIP mode, it may also be generated based on one or more top-side reference sample lines. Also, as an example, when the predicted block is generated using only the top-side reference samples of the current block based on MIP mode, the top-side reference samples may further include the right-side top-side reference samples of the current block. On the other hand, as another example, when the predicted block is generated using only the left-side reference samples of the current block based on MIP mode, it may also be generated based on one or more left-side reference sample lines. Also, as an example, when the predicted block is generated using only the left-side reference samples of the current block based on MIP mode, the left-side reference samples may include the left-side bottom-side reference samples of the current block. On the other hand, as another example, when a prediction block is generated using a portion of the left-side reference sample and top-side reference sample of the current block, at least one of the left-side reference sample or top-side reference sample may be partially adjacent to the current block. On the other hand, as yet another example, when a prediction block is generated using a portion of the left-side reference sample and top-side reference sample of the current block, the left-side reference sample may further include the left-side bottom-side reference sample of the current block, and the top-side reference sample may further include the right-side top-side reference sample of the current block. On the other hand, as yet another example, based on whether the MIP mode is applied, it may be determined whether the MIP mode is performed based on the left-side reference sample or top-side reference sample. That is, once it is determined that the MIP mode is applied, it may be determined how the subsequent MIP mode will be performed. As yet another example, whether or not the MIP mode is applied may be determined by information obtained from the bitstream, and whether or not the MIP mode is performed based on the left-side reference sample or top-side reference sample may be determined by information obtained from the bitstream.

[0267] On the other hand, since Figure 23 corresponds to one embodiment of the present disclosure, some steps may be modified, some steps may be in a different order, some steps may be omitted, or some steps may be performed simultaneously, and it is clear that these are also included in one embodiment of the present disclosure.

[0268] Figure 24 shows a video encoding method performed by a video encoding device according to one embodiment of the present disclosure.

[0269] For example, the video encoding device may decide whether or not to apply the MIP (matrix-based intra prediction) mode to the current block (S2410). For example, the MIP mode whose application is decided at this stage may include the MIP described with reference to Figures 14 to 22. For example, whether or not to apply the MIP mode may be encoded into the bitstream as the MIP application information described above. On the other hand, step S2410 may further include a step of deciding whether the MIP mode is applied using only the left reference sample or the top reference sample, or using only the left reference sample and a portion of the top reference sample.

[0270] As an example, a predicted block of the current block may be generated (S2420) based on the determination that the MIP mode is applied to the current block. The predicted block may be generated using only a portion of the left-side reference samples or top-side reference samples of the current block based on the MIP mode. As an example, when the predicted block is generated using only the top-side reference samples of the current block based on the MIP mode, it may also be generated based on one or more top-side reference sample lines. Also, as an example, when the predicted block is generated using only the top-side reference samples of the current block based on the MIP mode, the top-side reference samples may further include the right-side top-side reference samples of the current block. On the other hand, as another example, when the predicted block is generated using only the left-side reference samples of the current block based on the MIP mode, it may also be generated based on one or more left-side reference sample lines. Also, as an example, when the predicted block is generated using only the left-side reference samples of the current block based on the MIP mode, the left-side reference samples may include the left-side bottom-side reference samples of the current block. On the other hand, as another example, when a prediction block is generated using a portion of the left-side reference sample and top-side reference sample of the current block, at least one of the left-side reference sample or top-side reference sample may be partially adjacent to the current block. On the other hand, as yet another example, when a prediction block is generated using a portion of the left-side reference sample and top-side reference sample of the current block, the left-side reference sample may further include the left-side bottom-side reference sample of the current block, and the top-side reference sample may further include the right-side top-side reference sample of the current block. On the other hand, as yet another example, based on whether the MIP mode is applied, it may be determined whether or not the MIP mode is performed based on the left-side reference sample or top-side reference sample. That is, once it is determined that the MIP mode is applied, it may be determined how the MIP mode is performed thereafter. As yet another example, whether or not the MIP mode is applied may be encoded as one piece of information in the bitstream, and whether or not the MIP mode is performed based on the left-side reference sample or top-side reference sample may be encoded as one piece of information in the bitstream.

[0271] On the other hand, as an example, the bitstream generated by the video encoding method may be stored on a non-temporary computer-readable medium.

[0272] As an example, a method for transmitting a bitstream generated by a video encoding method may also be proposed.

[0273] On the other hand, since Figure 24 corresponds to one embodiment of the present disclosure, some steps may be modified, some steps may be in a different order, some steps may be omitted, or some steps may be performed simultaneously, and it is clear that this is also included as one embodiment of the present disclosure.

[0274] The exemplary methods in this disclosure are expressed as a series of actions for clarity of explanation, but this is not intended to limit the order in which the steps are performed, and each step may be performed simultaneously or in a different order, if necessary. To embody the methods relating to this disclosure, the exemplary steps may further include other steps, or the remaining steps may be included with some steps excluded, or yet another set of steps may be included with some steps excluded.

[0275] In this disclosure, a video encoding device or video decoding device that performs a predetermined operation (stage) may perform an operation (stage) to confirm the conditions or status of the execution of said operation (stage). For example, if it is stated that a predetermined operation is performed when a predetermined condition is met, the video encoding device or video decoding device may perform an operation to confirm whether or not the predetermined condition is met, and then perform the predetermined operation.

[0276] The various embodiments of this disclosure are not intended to enumerate all possible combinations, but rather to illustrate representative aspects of this disclosure. The matters described in the various embodiments may be applied independently or in combination of two or more.

[0277] Furthermore, various embodiments of this disclosure may be embodied in hardware, firmware, software, or a combination thereof. In the case of hardware embodiment, they may be embodied in one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.

[0278] Furthermore, the video decoding and video encoding devices to which the embodiments of this disclosure are applied may be included in multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video conferencing equipment, real-time communication equipment such as video communication, mobile streaming equipment, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video equipment, internet streaming service providers, 3D video equipment, image-phone video equipment, and medical video equipment, and may be used to process video signals or data signals. For example, over-the-top (OTT) video equipment may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).

[0279] Figure 21 illustrates a content streaming system to which the embodiments of this disclosure can be applied.

[0280] As shown in Figure 21, a content streaming system to which an embodiment of the present disclosure is applied may broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0281] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmitting this bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted.

[0282] The bitstream may be generated by a video encoding method and / or video encoding apparatus to which an embodiment of the present disclosure is applied, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0283] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server can act as an intermediary to inform users of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server can transmit multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server can play a role in controlling commands and responses between the devices within the content streaming system.

[0284] The streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0285] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs), digital TVs, desktop computers, and digital signage.

[0286] Each server within the aforementioned content streaming system may be operated as a distributed server, in which case the data received by each server may be processed in a distributed manner.

[0287] The scope of this disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that enable the operation of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or instructions are stored and executable on a device or computer. [Industrial applicability]

[0288] The embodiments described herein can be used for encoding / decoding video.

Claims

1. A video decoding method performed by a video decoding device, Currently, the stage is to determine the on-screen prediction mode for the block, The step includes: performing an in-screen prediction based on the determined in-screen prediction mode to generate a predicted block for the current block, The in-screen prediction mode is determined based on range limitation information indicating whether the in-screen prediction mode is limited to a specific range. A video decoding method in which, based on the range limitation information indicating that the in-screen prediction mode is limited to a specific range, the in-screen prediction is performed mainly on reference samples within a predetermined range from among the available reference samples of the current block, based on the in-screen prediction mode determined to be one of the in-screen prediction modes within the specific range.

2. The video decoding method according to claim 1, wherein the in-screen prediction mode is determined to be one directional mode within the specific range, based on the fact that it is used primarily for reference samples within the predetermined range.

3. The video decoding method according to claim 2, wherein the specified range is determined based on whether or not the in-screen prediction is based on a wide-angle in-screen prediction.

4. The video decoding method according to claim 2, wherein the specified range is one of the directional mode ranges of 50 to 66, 35 to 66, or 19 to 34.

5. The video decoding method according to claim 2, wherein the specified range is one of the directional mode ranges 2 to 18, 2 to 33, or 18 to 49.

6. The video decoding method according to claim 1, wherein the prediction block is generated by weighting together a first prediction block obtained mainly using reference samples within a first range and a second prediction block obtained mainly using reference samples within a second range.

7. The video decoding method according to claim 6, wherein the first range and the second range do not overlap with each other.

8. The video decoding method according to claim 6, wherein the first weighting value applied to the first prediction block is determined to be different from the weighting value applied to the second prediction block.

9. The video decoding method according to claim 8, wherein the first weighting value is determined based on the size of the current block.

10. The video decoding method according to claim 8, wherein the first weighting value and the second weighting value are predetermined values.

11. A video encoding method performed by a video encoding device, Currently, the stage is to determine the on-screen prediction mode for the block, The step includes encoding the prediction mode information of the current block based on the determined in-screen prediction mode, The in-screen prediction mode is determined based on whether or not the in-screen prediction mode is limited to a specific range. A video encoding method in which, based on the fact that the in-screen prediction mode is limited to a specific range, the in-screen prediction of the current block is performed mainly on reference samples within a predetermined range from among the available reference samples of the current block, based on the in-screen prediction mode which has been determined to be one of the in-screen prediction modes within the specific range.

12. A computer-readable medium for storing a bitstream generated by the video encoding method described in claim 11.

13. A method for transmitting a bitstream generated by a video encoding method, The aforementioned video encoding method is Currently, the stage is to determine the on-screen prediction mode for the block, The step includes encoding the prediction mode information of the current block based on the determined in-screen prediction mode, The in-screen prediction mode is determined based on whether or not the in-screen prediction mode is limited to a specific range. A video encoding method in which, based on the fact that the in-screen prediction mode is limited to a specific range, the in-screen prediction of the current block is performed mainly on reference samples within a predetermined range from among the available reference samples of the current block, based on the in-screen prediction mode which has been determined to be one of the in-screen prediction modes within the specific range.