Encoding / decoding methods, bitstream transmission methods, and readable / write media by non-forwarding computers.
Patent Information
- Authority / Receiving Office
- VN · VN
- Patent Type
- Applications
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2024-10-04
- Publication Date
- 2026-06-15
AI Technical Summary
The increasing demand for high-resolution and high-quality images, such as HD and UHD, leads to a significant increase in image data, resulting in higher transmission and storage costs due to the increased amount of information or bits required.
The development of an image coding/decryption method and device that utilizes Matrix-based Intra Prediction (MIP) to efficiently encode and decode images. This method applies MIP based on the size of the current block or the number of pixels, and generates predictive blocks using only a part of the reference samples, such as the upper or left reference samples.
The proposed method improves encoding/decryption efficiency, enhances prediction efficiency, and adapts predictions based on the current block's size or reference sample, resulting in better coding efficiency and image quality while reducing transmission and storage costs.
Smart Images

Figure VN1202603263_0
Abstract
Description
MIP-based video encoding / decoding method, method for transmitting bitstream, and recording medium storing bitstream
[0001] The present disclosure relates to a MIP-based image encoding / decoding method, a method for transmitting a bitstream, and a recording medium storing the bitstream, and more particularly, to a MIP (matrix-based intra prediction)-based image encoding / decoding method using only a portion of reference samples, a method for transmitting a bitstream, and a recording medium storing the bitstream.
[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, has been increasing across various fields. As image data becomes higher resolution and higher quality, the amount of information transmitted, or bits, increases relative to conventional image data. This increase in information or bits transmitted leads to increased transmission and storage costs.
[0003] Accordingly, a highly efficient image compression technology is required to effectively transmit, store, and play high-resolution, high-quality image information.
[0004] The present disclosure aims to provide a video encoding / decoding method and device with improved encoding / decoding efficiency.
[0005] In addition, the present disclosure aims to provide a method and device for encoding / decoding images based on in-screen prediction and selected reference samples.
[0006] In addition, the present disclosure aims to provide a video encoding / decoding method and device with improved intra-screen prediction efficiency.
[0007] In addition, the present disclosure aims to provide a video encoding / decoding method and device that can adaptively perform intra-screen prediction according to the size of a current block or a reference sample when performing intra-screen prediction.
[0008] In addition, the present disclosure aims to provide a non-transitory computer-readable recording medium that stores a bitstream generated by an image encoding method or device according to the present disclosure.
[0009] In addition, the present disclosure aims to provide a non-transitory computer-readable recording medium that stores a bitstream received and decoded by an image decoding device according to the present disclosure and used for restoring an image.
[0010] In addition, the present disclosure aims to provide a method for transmitting a bitstream generated by an image encoding method or device according to the present disclosure.
[0011] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below.
[0012] A video decoding method according to one aspect of the present disclosure comprises:
[0013] A step of deriving information related to MIP (Matrix-based Intra Prediction) based on a step of determining and the prediction mode being an intra prediction mode, wherein the MIP can be applied based on the size of the current block or the number of pixels included in the current block.
[0014] Meanwhile, as an example, the matrix for the MIP may be determined based on the size of the current block or the number of pixels included in the current block.
[0015] Meanwhile, as an example, if it is determined that the MIP is to be applied to the current block based on information related to the MIP, the value of the prediction block of the current block may be limited to a specific value based on a specific condition.
[0016] Meanwhile, as an example, the specific condition may be based on a reference sample value of the current block for the MIP.
[0017] Meanwhile, as an example, if the above specific condition is satisfied, application of the MIP to the current block may be omitted regardless of information related to the MIP.
[0018] Meanwhile, as an example, if it is determined that the MIP is to be applied to the current block based on information related to the MIP, the prediction block of the current block may be generated using only the upper reference sample of the current block.
[0019] Meanwhile, as an example, the upper reference sample may include the upper right reference sample of the current block.
[0020] Meanwhile, as an example, if it is determined that the MIP is to be applied to the current block based on information related to the MIP, the prediction block of the current block can be generated using only the left reference sample of the current block.
[0021] Meanwhile, as an example, the left reference sample may include the lower left reference sample of the current block.
[0022] Meanwhile, as an example, the prediction block of the current block is generated using a portion of the left reference sample and the upper reference sample of the current block, wherein at least one of the left reference sample or the upper reference sample may be partially adjacent to the current block.
[0023] Meanwhile, as an example, when it is determined that the MIP is to be applied to the current block based on information related to the MIP, the location of the reference sample of the current block is determined, and the location of the reference sample of the current block can be derived based on information signaled from the bitstream.
[0024] An image encoding method according to one aspect of the present disclosure includes a step of determining whether to apply MIP (matrix-based intra prediction) to a current block and a step of determining information related to the MIP, wherein the MIP can be applied based on the size of the current block or the number of pixels included in the current block.
[0025] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by an image encoding method or an image encoding device of the present disclosure.
[0026] A transmission method according to another aspect of the present disclosure can transmit a bitstream generated by an image encoding device or an image encoding method of the present disclosure.
[0027] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.
[0028] According to the present disclosure, a video encoding / decoding method and device with improved encoding / decoding efficiency can be provided.
[0029] Additionally, according to the present disclosure, a method and device for encoding / decoding an image based on in-screen prediction and some selected reference samples can be provided.
[0030] In addition, according to the present disclosure, when performing intra-screen prediction, coding efficiency and quality can be improved by adaptively performing intra-screen prediction according to the size of the current block, reference samples, etc.
[0031] In addition, according to the present disclosure, a video encoding / decoding method and device can be provided that can generate an adaptive intra-screen prediction block for a reference sample area and / or the size / number of pixels of a current block during intra-screen prediction.
[0032] In addition, according to the present disclosure, a video encoding / decoding method and device with improved intra-screen prediction efficiency can be provided.
[0033] In addition, according to the present disclosure, a non-transitory computer-readable recording medium for storing a bitstream generated by an image encoding method or device according to the present disclosure can be provided.
[0034] In addition, according to the present disclosure, a non-transitory computer-readable recording medium can be provided that stores a bitstream received and decoded by an image decoding device according to the present disclosure and used for restoring an image.
[0035] Additionally, according to the present disclosure, a method for transmitting a bitstream generated by an image encoding method or device according to the present disclosure can be provided.
[0036] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.
[0037] FIG. 1 is a diagram schematically illustrating a video coding system to which an embodiment according to the present disclosure can be applied.
[0038] FIG. 2 is a schematic diagram of an image encoding device to which an embodiment according to the present disclosure can be applied.
[0039] FIG. 3 is a schematic diagram illustrating an image decoding device to which an embodiment according to the present disclosure can be applied.
[0040] Fig. 4 is a flowchart illustrating an example of an intra prediction mode signaling method in an encoding device.
[0041] Fig. 5 is a flowchart illustrating an example of a method for determining an intra prediction mode in a decoding device.
[0042] Figure 6 is a diagram showing examples of peripheral blocks used to derive the MPM list.
[0043] FIGS. 7 to 11 are diagrams showing examples of the overall process of averaging, matrix vector multiplication, and linear interpolation that can be applied to the present disclosure.
[0044] FIG. 12 is a diagram illustrating an example of a boundary averaging process that can be applied to the present disclosure.
[0045] Figure 13 is a diagram showing an example of a linear interpolation process that can be applied to the present disclosure.
[0046] FIG. 14 is a diagram illustrating a MIP prediction method that can be applied to the present disclosure.
[0047] FIG. 15 illustrates a method for performing MIP prediction according to one embodiment of the present disclosure.
[0048] This is a drawing for.
[0049] FIG. 16 is a diagram for explaining a MIP-based image encoding or decoding method according to one embodiment of the present disclosure.
[0050] Figure 17 is a diagram for explaining a MIP prediction method that can be applied to the present disclosure.
[0051] FIG. 18 is a diagram for explaining a method for performing MIP prediction according to one embodiment of the present disclosure.
[0052] FIGS. 19a and 19b are diagrams for explaining a MIP-based image encoding or decoding method according to one embodiment of the present disclosure.
[0053] FIG. 20 is a diagram for explaining a MIP prediction method that can be applied to the present disclosure.
[0054] FIG. 21 illustrates a method for performing MIP prediction according to one embodiment of the present disclosure.
[0055] This is a drawing for.
[0056] FIG. 22 is a diagram for explaining a MIP-based image encoding or decoding method according to one embodiment of the present disclosure.
[0057] FIG. 23 is a diagram for explaining a MIP-based image encoding or decoding method according to one embodiment of the present disclosure.
[0058] FIG. 24 is a diagram for explaining a MIP-based image encoding or decoding method according to one embodiment of the present disclosure.
[0059] FIG. 25 is a diagram illustrating an image decoding method according to one embodiment of the present disclosure.
[0060] FIG. 26 is a diagram illustrating an image encoding method according to one embodiment of the present disclosure.
[0061] FIG. 27 is a diagram illustrating an example of a content streaming system to which an embodiment according to the present disclosure can be applied.
[0062] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein.
[0063] In describing embodiments of the present disclosure, detailed descriptions of known configurations or functions will be omitted if they are deemed to obscure the gist of the present disclosure. Furthermore, portions unrelated to the description of the present disclosure in the drawings have been omitted, and similar portions have been designated with similar reference numerals.
[0064] In the present disclosure, when a component is said to be "connected," "coupled," or "connected" to another component, this may include not only a direct connection, but also an indirect connection in which another component exists in between. Furthermore, when a component is said to "include" or "have" another component, unless otherwise specifically stated, this does not exclude the other component, but rather implies that the other component may be included.
[0065] In this disclosure, terms such as first, second, etc. are used solely to distinguish one component from another, and do not limit the order or importance of components unless specifically stated otherwise. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0066] In this disclosure, distinct components are used to clearly illustrate their respective characteristics, and do not necessarily imply that the components are separated. That is, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not specifically mentioned, such integrated or distributed embodiments are also included within the scope of this disclosure.
[0067] In the present disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, embodiments comprising a subset of the components described in one embodiment are also within the scope of the present disclosure. Furthermore, embodiments including other components in addition to the components described in various embodiments are also within the scope of the present disclosure.
[0068] The present disclosure relates to encoding and decoding of images, and terms used in the present disclosure may have their usual meanings commonly used in the technical field to which the present disclosure belongs, unless newly defined in the present disclosure.
[0069] In the present disclosure, a "picture" generally refers to a unit representing one image of a specific time period, and a slice / tile is a coding unit that constitutes a part of a picture, and a single picture may be composed of one or more slices / tiles. In addition, a slice / tile may include one or more coding tree units (CTUs).
[0070] In the present disclosure, "pixel" or "pel" may refer to the smallest unit that constitutes a picture (or image). Additionally, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component.
[0071] In the present disclosure, a "unit" may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. In some cases, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0072] In the present disclosure, the "current block" may mean one of the following: a "current coding block," a "current coding unit," a "block to be encoded," a "block to be decoded," or a "block to be processed." When prediction is performed, the "current block" may mean a "current prediction block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, the "current block" may mean a "current transformation block" or a "block to be transformed." When filtering is performed, the "current block" may mean a "block to be filtered."
[0073] In the present disclosure, a "current block" may mean a block that includes both a luma component block and a chroma component block, or a "luma block of the current block," unless explicitly described as a chroma block. The luma component block of the current block may be explicitly expressed by including an explicit description of the luma component block, such as "luma block" or "current luma block." Additionally, the chroma component block of the current block may be explicitly expressed by including an explicit description of the chroma component block, such as "chroma block" or "current chroma block."
[0074] In this disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Additionally, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."
[0075] In this disclosure, "or" may be interpreted as "and / or." For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B." Alternatively, "or" in this disclosure may mean "additionally or alternatively."
[0076] Overview of Video Coding Systems
[0077] FIG. 1 is a diagram schematically illustrating a video coding system to which an embodiment according to the present disclosure can be applied.
[0078] A video coding system according to one embodiment may include an encoding device (10) and a decoding device (20). The encoding device (10) may transmit encoded video and / or image information or data to the decoding device (20) in the form of a file or streaming via a digital storage medium or a network.
[0079] An encoding device (10) according to one embodiment may include a video source generation unit (11), an encoding unit (12), and a transmission unit (13). A decoding device (20) according to one embodiment may include a reception unit (21), a decoding unit (22), and a rendering unit (23). The encoding unit (12) may be referred to as a video / image encoding unit, and the decoding unit (22) may be referred to as a video / image decoding unit. The transmission unit (13) may be included in the encoding unit (12). The reception unit (21) may be included in the decoding unit (22). The rendering unit (23) may include a display unit, and the display unit may be configured as a separate device or an external component.
[0080] The video source generation unit (11) can obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source generation unit (11) can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced with a process of generating related data.
[0081] The encoding unit (12) can encode input video / images. The encoding unit (12) can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and encoding efficiency. The encoding unit (12) can output encoded data (encoded video / image information) in the form of a bitstream.
[0082] The transmission unit (13) can obtain encoded video / image information or data output in the form of a bitstream, and transmit it to the reception unit (21) of the decoding device (20) or another external object through a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) may include an element for generating a media file through a predetermined file format, and may include an element for transmission through a broadcasting / communication network. The transmission unit (13) may be provided as a separate transmission device from the encoding device (12), and in this case, the transmission device may include at least one processor for obtaining encoded video / image information or data output in the form of a bitstream, and a transmission unit for transmitting it in the form of a file or streaming. The reception unit (21) can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit (22).
[0083] The decoding unit (22) can decode video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding unit (12).
[0084] The rendering unit (23) can render the decrypted video / image. The rendered video / image can be displayed through the display unit.
[0085] Overview of the video encoding device
[0086] FIG. 2 is a schematic diagram illustrating an image encoding device to which an embodiment according to the present disclosure can be applied.
[0087] As illustrated in FIG. 2, the image encoding device (100) may include an image segmentation unit (110), a subtraction unit (115), a transformation unit (120), a quantization unit (130), an inverse quantization unit (140), an inverse transformation unit (150), an addition unit (155), a filtering unit (160), a memory (170), an inter prediction unit (180), an intra prediction unit (185), and an entropy encoding unit (190). The inter prediction unit (180) and the intra prediction unit (185) may be collectively referred to as a “prediction unit.” The transformation unit (120), the quantization unit (130), the inverse quantization unit (140), and the inverse transformation unit (150) may be included in a residual processing unit. The residual processing unit may further include a subtraction unit (115).
[0088] All or at least some of the plurality of components constituting the video encoding device (100) may be implemented as a single hardware component (e.g., an encoder or a processor) according to an embodiment. In addition, the memory (170) may include a decoded picture buffer (DPB) and may be implemented by a digital storage medium.
[0089] The image segmentation unit (110) can segment an input image (or picture, frame) input to the image encoding device (100) into one or more processing units. For example, the processing units may be referred to as coding units (CUs). The coding units can be obtained by recursively segmenting a coding tree unit (CTU) or a largest coding unit (LCU) according to a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit can be segmented into a plurality of coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. For segmenting the coding unit, the quad-tree structure may be applied first, and the binary-tree structure and / or the ternary-tree structure may be applied later. The coding procedure according to the present disclosure can be performed based on the final coding unit that is no longer segmented. The maximum coding unit can be used directly as the final coding unit, and the coding unit of the lower depth obtained by dividing the maximum coding unit can be used as the final concatenated unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or restoration described below. As another example, the processing unit of the coding procedure may be a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may each be divided or partitioned from the final coding unit. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit that derives a transform coefficient and / or a unit that derives a residual signal from a transform coefficient.
[0090] The prediction unit (inter-prediction unit (180) or intra-prediction unit (185)) can perform prediction on a block to be processed (current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block or CU unit. The prediction unit can generate various information regarding the prediction of the current block and transmit the information to the entropy encoding unit (190). The information regarding the prediction can be encoded by the entropy encoding unit (190) and output in the form of a bitstream.
[0091] The intra prediction unit (185) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the intra prediction mode and / or intra prediction technique. The intra prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of detail in the prediction direction. However, this is merely an example, and a greater or lesser number of directional prediction modes may be used depending on the settings. The intra prediction unit (185) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0092] The inter prediction unit (180) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. A reference picture including the above temporal neighboring blocks may be called a collocated picture (colPic). For example, the inter prediction unit (180) may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (180) may use the motion information of neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the current block can be signaled by using the motion vector of the surrounding blocks as the motion vector predictor and encoding the motion vector difference and an indicator for the motion vector predictor. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.
[0093] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit can apply intra prediction or inter prediction to predict the current block, and can also apply intra prediction and inter prediction simultaneously. A prediction method that simultaneously applies intra prediction and inter prediction to predict the current block may be called combined inter and intra prediction (CIIP). In addition, the prediction unit may perform intra block copy (IBC) to predict the current block. Intra block copy can be used for video / image coding of content such as games, such as screen content coding (SCC). IBC is a method of predicting the current block using a previously restored reference block within the current picture located at a predetermined distance from the current block. When IBC is applied, the location of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in the present disclosure.
[0094] The prediction signal generated through the prediction unit can be used to generate a restoration signal or a residual signal. The subtraction unit (115) can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit (120).
[0095] The transform unit (120) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously reconstructed pixels. The transform process can be applied to a pixel block having a square equal size, or can be applied to a block of a non-square variable size.
[0096] The quantization unit (130) can quantize the transform coefficients and transmit them to the entropy encoding unit (190). The entropy encoding unit (190) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (130) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0097] The entropy encoding unit (190) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit (190) can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in the form of a network abstraction layer (NAL) unit. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The signaling information, transmitted information and / or syntax elements mentioned in the present disclosure may be encoded through the encoding procedure described above and included in the bitstream.
[0098] The above bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting the signal output from the entropy encoding unit (190) and / or a storage unit (not shown) for storing the signal may be provided as an internal / external element of the video encoding device (100), or the transmission unit may be provided as a component of the entropy encoding unit (190).
[0099] The quantized transform coefficients output from the quantization unit (130) can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (140) and inverse transformation unit (150), a residual signal (residual block or residual samples) can be restored.
[0100] The addition unit (155) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (180) or the intra prediction unit (185). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit (155) can be called a reconstructor or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.
[0101] The filtering unit (160) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (160) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (170), specifically, in the DPB of the memory (170). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (160) can generate various information regarding filtering and transmit the information to the entropy encoding unit (190), as described later in the description of each filtering method. The information regarding filtering may be encoded by the entropy encoding unit (190) and output in the form of a bitstream.
[0102] The modified restored picture transmitted to the memory (170) can be used as a reference picture in the inter prediction unit (180). Through this, when inter prediction is applied, the image encoding device (100) can avoid prediction mismatch between the image encoding device (100) and the image decoding device, and can also improve encoding efficiency.
[0103] The DPB in the memory (170) can store a modified reconstructed picture to be used as a reference picture in the inter prediction unit (180). The memory (170) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (180) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (170) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (185).
[0104] Video Decryption Device Overview
[0105] FIG. 3 is a schematic diagram illustrating an image decoding device to which an embodiment according to the present disclosure can be applied.
[0106] As illustrated in FIG. 3, the image decoding device (200) may be configured to include an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (230), an addition unit (235), a filtering unit (240), a memory (250), an inter prediction unit (260), and an intra prediction unit (265). The inter prediction unit (260) and the intra prediction unit (265) may be collectively referred to as a “prediction unit.” The inverse quantization unit (220) and the inverse transformation unit (230) may be included in a residual processing unit.
[0107] All or at least some of the plurality of components constituting the video decoding device (200) may be implemented as a single hardware component (e.g., a decoder or processor) depending on the embodiment. In addition, the memory (170) may include a DPB and may be implemented by a digital storage medium.
[0108] The video decoding device (200) that receives a bitstream including video / image information can restore the image by performing a process corresponding to the process performed in the video encoding device (100) of FIG. 2. For example, the video decoding device (200) can perform decoding using a processing unit applied in the video encoding device. Therefore, the processing unit for decoding may be, for example, a coding unit. The coding unit may be a coding tree unit or may be obtained by dividing a maximum coding unit. In addition, the restored image signal decoded and output by the video decoding device (200) can be reproduced through a reproduction device (not shown).
[0109] The video decoding device (200) can receive a signal output from the video encoding device of FIG. 2 in the form of a bitstream. The received signal can be decoded through the entropy decoding unit (210). For example, the entropy decoding unit (210) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The video decoding device may additionally use information on the parameter set and / or the general constraint information to decode the image. The signaling information, received information, and / or syntax elements mentioned in the present disclosure can be obtained from the bitstream by being decoded through the decoding procedure. For example, the entropy decoding unit (210) can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and the decoding information of the surrounding block and the decoding target block or the information of the symbol / bin decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (210) is provided to the prediction unit (inter prediction unit (260) and intra prediction unit (265)), and the residual value on which entropy decoding is performed by the entropy decoding unit (210), i.e., quantized transform coefficients and related parameter information, can be input to the inverse quantization unit (220). In addition, information regarding filtering among the information decoded by the entropy decoding unit (210) can be provided to the filtering unit (240). Meanwhile, a receiving unit (not shown) that receives a signal output from an image encoding device may be additionally provided as an internal / external element of the image decoding device (200), or the receiving unit may be provided as a component of an entropy decoding unit (210).
[0110] Meanwhile, the video decoding device according to the present disclosure may be referred to as a video / video / picture decoding device. The video decoding device may include an information decoder (video / video / picture information decoder) and / or a sample decoder (video / video / picture sample decoder). The information decoder may include an entropy decoding unit (210), and the sample decoder may include at least one of an inverse quantization unit (220), an inverse transformation unit (230), an addition unit (235), a filtering unit (240), a memory (250), an inter prediction unit (260), and an intra prediction unit (265).
[0111] The inverse quantization unit (220) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (220) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the image encoding device. The inverse quantization unit (220) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0112] In the inverse transform unit (230), the transform coefficients can be inversely transformed to obtain a residual signal (residual block, residual sample array).
[0113] The prediction unit can perform a prediction on the current block and generate a predicted block containing prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block based on the prediction information output from the entropy decoding unit (210), and can determine a specific intra / inter-prediction mode (prediction technique).
[0114] The fact that the prediction unit can generate a prediction signal based on various prediction methods (techniques) described later is the same as that mentioned in the description of the prediction unit of the image encoding device (100).
[0115] The intra prediction unit (265) can predict the current block by referring to samples within the current picture. The description of the intra prediction unit (185) can be equally applied to the intra prediction unit (265).
[0116] The inter prediction unit (260) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (260) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information about the prediction can include information indicating the mode (technique) of inter prediction for the current block.
[0117] The addition unit (235) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (including the inter prediction unit (260) and / or the intra prediction unit (265)). When there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block can be used as the restoration block. The description of the addition unit (155) can be equally applied to the addition unit (235). The addition unit (235) can be called a restoration unit or a restoration block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after going through filtering as described below.
[0118] The filtering unit (240) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (240) can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory (250), specifically, in the DPB of the memory (250). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0119] The (modified) reconstructed picture stored in the DPB of the memory (250) can be used as a reference picture in the inter prediction unit (260). The memory (250) can store motion information of a block from which motion information is derived (or decoded) in the current picture and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (260) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (250) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (265).
[0120] In this specification, the embodiments described in the filtering unit (160), the inter prediction unit (180), and the intra prediction unit (185) of the image encoding device (100) can be applied to the filtering unit (240), the inter prediction unit (260), and the intra prediction unit (265) of the image decoding device (200) in the same or corresponding manner, respectively.
[0121] Determining intra prediction mode / type
[0122] When intra prediction is applied, the intra prediction mode to be applied to the current block can be determined using the intra prediction mode of the surrounding blocks. For example, the decoding device can select one of the MPM candidates in the MPM list derived based on the intra prediction modes of the surrounding blocks of the current block (e.g., left and / or upper surrounding blocks) and additional candidate modes based on the received MPM index, or can select one of the remaining intra prediction modes not included in the MPM candidates (and planar modes) based on the remaining intra prediction mode information. The MPM list can be configured to include or not include a planar mode as a candidate. For example, when the MPM list includes a planar mode as a candidate, the MPM list can have six candidates, and when the MPM list does not include a planar mode as a candidate, the MPM list can have three candidates. If the above mpm list does not include a planar mode as a candidate, a not planar flag (e.g., intra_luma_not_planar_flag) may be signaled to indicate whether the intra prediction mode of the current block is not a planar mode. For example, the mpm flag may be signaled first, and then the mpm index and the not planar flag may be signaled if the value of the mpm flag is 1. Additionally, the mpm index may be signaled if the value of the not planar flag is 1. Here, the reason why the above mpm list is configured not to include a planar mode as a candidate is not to mean that the planar mode is not mpm, but rather to first check whether it is a planar mode by signaling a flag (not planar flag) since the planar mode is always considered as mpm.
[0123] For example, whether the intra prediction mode applied to the current block is among the MPM candidates (and planar mode) or among the remaining mode can be indicated based on the mpm flag (e.g., intra_luma_mpm_flag). A value of 1 of the mpm flag can indicate that the intra prediction mode for the current block is among the mpm candidates (and planar mode), and a value of 0 of the mpm flag can indicate that the intra prediction mode for the current block is not among the mpm candidates (and planar mode). A value of 0 of the not planar flag (e.g., intra_luma_not_planar_flag) can indicate that the intra prediction mode for the current block is the planar mode, and a value of 1 of the not planar flag can indicate that the intra prediction mode for the current block is not the planar mode. The above mpm index may be signaled in the form of a mpm_idx or intra_luma_mpm_idx syntax element, and the remaining intra prediction mode information may be signaled in the form of a rem_intra_luma_pred_mode or intra_luma_mpm_remainder syntax element. For example, the remaining intra prediction mode information may indicate one of the remaining intra prediction modes that are not included in the mpm candidates (and planar modes) among all intra prediction modes by indexing them in the order of prediction mode numbers. The intra prediction mode may be an intra prediction mode for a luma component (sample). Hereinafter, the intra prediction mode information may include the mpm flag (e.g., intra_luma_mpm_flag), the not planar flag (e.g., intra_luma_not_planar_flag), and the mpm index (e.g., mpm_idx or intra_luma_mpm_idx) and may include at least one of the remaining intra prediction mode information (rem_intra_luma_pred_mode or intra_luma_mpm_remainder). In this document, the MPM list may be called by various terms such as MPM candidate list, candModeList, etc. When MIP is applied to the current block, a separate mpm flag for MIP (e.g., intra_mip_mpm_flag), mpm index (e.g., intra_mip_mpm_idx), and remaining intra prediction mode information (e.g., intra_mip_mpm_remainder) may be signaled, and the not planar flag is not signaled.
[0124] The intra prediction mode signaling procedure in the encoding device and the intra prediction mode determination procedure in the decoding device can be performed, for example, as follows.
[0125] Fig. 4 is a flowchart illustrating an example of an intra prediction mode signaling method in an encoding device.
[0126] Referring to FIG. 4, the encoding device constructs an MPM list for the current block (S400). The MPM list may include candidate intra-prediction modes (MPM candidates) likely to be applied to the current block. The MPM list may include intra-prediction modes of neighboring blocks, and may further include specific intra-prediction modes according to a predetermined method. A specific method for constructing the MPM list will be described below.
[0127] The encoding device determines the intra prediction mode of the current block (S410). The encoding device can perform prediction based on various intra prediction modes, and determine the optimal intra prediction mode based on rate-distortion optimization (RDO) based thereon. In this case, the encoding device may determine the optimal intra prediction mode using only the MPM candidates and planar mode configured in the MPM list, or may determine the optimal intra prediction mode using not only the MPM candidates and planar mode configured in the MPM list but also the remaining intra prediction modes. Specifically, for example, if the intra prediction type of the current block is a specific type (e.g., LIP, MRL, or ISP) rather than the normal intra prediction type, the encoding device may consider only the MPM candidates and the planar mode as intra prediction mode candidates for the current block to determine the optimal intra prediction mode. That is, in this case, the intra prediction mode for the current block can be determined only from among the MPM candidates and the planar mode, and in this case, the mpm flag may not be encoded / signaled. In this case, the decoding device can assume that the mpm flag is 1 without being separately signaled the mpm flag.
[0128] Meanwhile, in general, if the intra prediction mode of the current block is not a planar mode but one of the MPM candidates in the MPM list, the encoding device generates an MPM index (mpm idx) pointing to one of the MPM candidates. If the intra prediction mode of the current block is not in the MPM list either, the encoding device generates remaining intra prediction mode information pointing to a mode that is the same as the intra prediction mode of the current block among the remaining intra prediction modes that are not included in the MPM list (and the planar mode).
[0129] An encoding device can encode intra prediction mode information and output it in the form of a bitstream. The intra prediction mode information can include the above-described mpm flag, not-planar flag, mpm index, and / or remaining intra prediction mode information. In general, the mpm index and the remaining intra prediction mode information are not signaled simultaneously when indicating an intra prediction mode for a block because they have an alternative relationship. That is, the mpm flag value 1 and the not-planar flag or mpm index are signaled together, or the mpm flag value 0 and the remaining intra prediction mode information are signaled together. However, as described above, when a specific intra prediction type is applied to the current block, the mpm flag may not be signaled, and only the not-planar flag and / or mpm index may be signaled. That is, in this case, the intra prediction mode information may include only the not-planar flag and / or mpm index.
[0130] The decoding device can determine the intra prediction mode in response to intra prediction mode information determined and signaled by the encoding device.
[0131] Fig. 5 is a flowchart illustrating an example of a method for determining an intra prediction mode in a decoding device.
[0132] Referring to FIG. 5, the decoding device obtains intra prediction mode information from the bitstream (S500). The intra prediction mode information may include at least one of an MPM flag, a not-planar flag, an MPM index, and a remaining intra prediction mode, as described above.
[0133] The decoding device constructs an MPM list (S510). The MPM list is constructed identically to the MPM list constructed by the encoding device. That is, the MPM list may include intra prediction modes of surrounding blocks, or may further include specific intra prediction modes according to a predetermined method. A specific method for constructing the MPM list is described below.
[0134] Although S510 is shown as being performed after S500, this is an example, and S510 may be performed before S500 or may be performed concurrently.
[0135] The decoding device determines the intra prediction mode of the current block based on the MPM list and the intra prediction mode information (S520). For example, when the value of the mpm flag is 1, the decoding device may derive a planar mode as the intra prediction mode of the current block (based on the not planar flag) or derive a candidate indicated by the mpm index among MPM candidates in the MPM list as the intra prediction mode of the current block. As another example, when the value of the mpm flag is 0, the decoding device may derive an intra prediction mode indicated by the remaining intra prediction mode information among the remaining intra prediction modes not included in the MPM list and the planar mode as the intra prediction mode of the current block. Meanwhile, as another example, if the intra prediction type of the current block is a specific type (e.g., LIP, MRL, or ISP, etc.), the decoding device may derive the candidate indicated by the mpm index within the planar mode or the MPM list as the intra prediction mode of the current block without checking the mpm flag.
[0136] For example, the not planar flag may be signaled when the MRL is not applied to the current block (i.e., when intra_luma_ref_idx == 0), and the not planar flag may be omitted when the MRL is applied to the current block (i.e., when intra_luma_ref_idx != 0). When the not planar flag is omitted, its value may be assumed to be 1 by the decoding device.
[0137] Meanwhile, the intra prediction modes may include two directional intra prediction modes and 65 directional intra prediction modes. The non-directional intra prediction modes may include a planar intra prediction mode and a DC intra prediction mode, and the directional intra prediction modes may include intra prediction modes 2 to 66. Extended directional intra prediction may be applied to blocks of all sizes and may be applied to both luma components and chroma components.
[0138] Meanwhile, the intra prediction mode may further include a CCLM (cross-component linear model) mode for chroma samples in addition to the intra prediction modes described above. The CCLM mode may be divided into LT_CCLM, L_CCLM, and T_CCLM depending on whether left samples, upper samples, or both are considered for deriving LM parameters, and may only be applied to chroma components.
[0139] Intra prediction modes can be indexed, for example, as shown in Table 1 below.
[0140]
[0141]
[0142] Meanwhile, the intra prediction type (or additional intra prediction mode, etc.) may include at least one of the above-described LIP, PDPC, MRL, ISP, and MIP. The intra prediction type may be indicated based on intra prediction type information, and the intra prediction type information may be implemented in various forms. For example, the intra prediction type information may include intra prediction type index information indicating one of the intra prediction types. As another example, the intra prediction type information may include at least one of reference sample line information (e.g., intra_luma_ref_idx) indicating whether the MRL is applied to the current block and, if so, which reference sample line is used, ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block, ISP type information (e.g., intra_subpartitions_split_flag) indicating a split type of subpartitions if the ISP is applied, flag information indicating whether PDCP is applied, or flag information indicating whether LIP is applied. Additionally, the intra prediction type information may include a MIP flag (or may be called intra_mip_flag) indicating whether MIP is applied to the current block.
[0143] Meanwhile, when MIP is applied to the current block as described above (e.g., when the value of intra_mip_flag is 1), an MPM list for the MIP may be separately configured, and the MPM flag that may be included in the intra prediction mode information for the MIP may be called intra_mip_mpm_flag, the MPM index may be called intra_mip_mpm_idx, and the remaining intra prediction mode information may be called intra_mip_mpm_remainder.
[0144] In addition, various prediction modes can be used for MIP, and a matrix and offset for MIP can be derived according to the intra prediction mode for MIP. As described above, the matrix here can be called a (MIP) weight matrix, and the offset can be called a (MIP) offset vector or a (MIP) bias vector. The number of intra prediction modes for MIP can be set differently based on the size of the current block. For example, i) when the height and width of the current block (e.g., CB or TB) are each 4, 35 intra prediction modes (i.e., intra prediction modes 0 to 34) may be available, ii) when both the height and width of the current block are 8 or less, 19 intra prediction modes (i.e., intra prediction modes 0 to 18) may be available, and iii) in other cases, 11 intra prediction modes (i.e., intra prediction modes 0 to 10) may be available. For example, when the height and width of the current block are each 4, this is referred to as block size type 0, when both the height and width of the current block are 8 or less, this is referred to as block size type 1, and in other cases, this is referred to as block size type 2, the number of intra prediction modes for MIP may be organized as shown in the following table. However, this is an example, and the block size type and the number of available intra prediction modes may be changed. In this document, the intra prediction mode for MIP may be referred to as MIP intra prediction mode, MIP prediction mode, or MIP mode.
[0145]
[0146] Meanwhile, the enhanced compression model (ECM) introduces a secondary MPM list. The existing primary MPM (PMPM) list consists of 6 entries, and the secondary MPM (SMPM) list contains 16 entries. First, a general MPM list with 22 entries is constructed, and then the first 6 entries in the general MPM list are included in the PMPM list, and the remaining entries are included in the SMPM list. The first entry in the general MPM list is a planar mode, and the remaining entries are composed of intra modes of the left (L), top (A), bottom left (BL), top right (AR), and top left (AL) neighboring blocks, directional modes with offsets added from the first two available directional modes of the neighboring blocks, and a default mode, as illustrated in FIG. 6.
[0147] When the CU block is vertical, the order of the surrounding blocks can be A, L, BL, AR, and AL. Otherwise, the order is L, A, BL, AR, and AL.
[0148] The PMPM flag is parsed, and if its value is 1, the PMPM index is parsed to determine which entry in the PMPM list is selected. Otherwise, the SPMPM flag is parsed to determine whether to parse the SPMPM index or remaining modes.
[0149] Deriving peripheral reference samples
[0150] When intra prediction is applied to a current block, peripheral reference samples to be used for intra prediction of the current block can be derived. The peripheral reference samples of the current block may include a total of 2xnH samples adjacent to the left boundary and bottom-left neighbors of a current block of a size nWxnH, a total of 2xnW samples adjacent to the top boundary and top-right neighbors of the current block, and one sample adjacent to the top-left of the current block. Alternatively, the peripheral reference samples of the current block may include upper peripheral samples of multiple columns and left peripheral samples of multiple rows. In addition, the peripheral reference samples of the current block may include a total of nH samples adjacent to the right boundary of a current block of a size nWxnH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right of the current block.
[0151] Meanwhile, when the MRL described below is applied, reference samples may be located on lines 1 to 3, rather than line 0 adjacent to the current block on the left / upper side. In this case, the number of surrounding reference samples may increase. The specific areas and number of surrounding reference samples are described below.
[0152] Meanwhile, when the ISP described below is applied, the surrounding reference samples can be derived in subpartition units.
[0153] Some of the surrounding reference samples of the current block may not yet be decoded or available. In this case, the decoder can interpolate the available samples to construct the surrounding reference samples to be used for prediction.
[0154] Some of the surrounding reference samples of the current block may not yet be decoded or available. In this case, the decoder can construct the surrounding reference samples to be used for prediction by extrapolating the available samples. Starting from the lower left, the decoder can construct the surrounding reference samples by updating the reference samples with the most recent sample (the last available sample) and replacing or padding the pixels that are not yet decoded or available with the last available sample until it reaches the upper right reference sample.
[0155] Deriving prediction samples based on intra prediction mode / type
[0156] The prediction unit of the encoding device / decoding device can derive a reference sample according to the intra prediction mode of the current block among surrounding reference samples of the current block, and can generate a prediction sample of the current block based on the reference sample.
[0157] For example, (i) a prediction sample may be derived based on an average or interpolation of neighboring reference samples of the current block, and (ii) the prediction sample may be derived based on a reference sample that exists in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. The case of (i) may be called a non-directional mode or a non-angular mode, and the case of (ii) may be called a directional mode or an angular mode. In addition, the prediction sample may be generated through interpolation of the first neighboring sample and the second neighboring sample, which are located in the opposite direction of the prediction direction of the intra prediction mode of the current block with respect to the prediction sample of the current block among the neighboring reference samples. The above-described case may be called linear interpolation intra prediction (LIP). In addition, based on the filtered peripheral reference samples, a temporary prediction sample of the current block may be derived, and a prediction sample of the current block may be derived by weighting at least one reference sample derived according to the intra prediction mode among the existing peripheral reference samples, i.e., unfiltered peripheral reference samples, and the temporary prediction sample. The above-described case may be called Position dependent intra prediction (PDPC). In addition, intra prediction encoding may be performed by selecting a reference sample line with the highest prediction accuracy among the peripheral multiple reference sample lines of the current block, deriving a prediction sample using a reference sample located in the prediction direction in the corresponding line, and instructing (signaling) the used reference sample line to a decoding device. The above-described case may be called multi-reference line intra prediction (MRL) or MRL-based intra prediction.In addition, the current block can be divided into vertical or horizontal subpartitions, and intra prediction can be performed based on the same intra prediction mode, and surrounding reference samples can be derived and used for each subpartition. That is, in this case, the intra prediction mode for the current block is applied equally to the subpartitions, and surrounding reference samples can be derived and used for each subpartition, thereby improving the intra prediction performance in some cases. This prediction method can be called intra sub-partitions (ISP) or ISP-based intra prediction. The specific details will be described later. In addition, when the prediction direction based on the prediction sample points between the surrounding reference samples, that is, when the prediction direction points to the fractional sample position, the value of the prediction sample can also be derived through interpolation of a plurality of reference samples located around the corresponding prediction direction (around the corresponding fractional sample position).
[0158] The intra prediction methods described above may be referred to as intra prediction types, as distinguished from intra prediction modes. The intra prediction types may be referred to by various terms, such as intra prediction techniques or supplementary intra prediction modes. For example, the intra prediction types (or supplementary intra prediction modes, etc.) may include at least one of the above-described LIP, PDPC, MRL, and ISP. Information regarding the intra prediction types may be encoded in an encoding device, included in a bitstream, and signaled to a decoding device. Information regarding the intra prediction types may be implemented in various forms, such as flag information indicating whether each intra prediction type is applied or index information indicating one of multiple intra prediction types.
[0159] The MPM list for deriving the above-described intra prediction mode may be configured differently depending on the intra prediction type. Alternatively, the MPM list may be configured in common regardless of the intra prediction type.
[0160] Deriving peripheral reference samples
[0161] When intra prediction is applied to a current block, peripheral reference samples to be used for intra prediction of the current block can be derived. The peripheral reference samples of the current block may include a total of 2xnH samples adjacent to the left boundary and bottom-left neighbors of a current block of a size nWxnH, a total of 2xnW samples adjacent to the top boundary and top-right neighbors of the current block, and one sample adjacent to the top-left of the current block. Alternatively, the peripheral reference samples of the current block may include upper peripheral samples of multiple columns and left peripheral samples of multiple rows. In addition, the peripheral reference samples of the current block may include a total of nH samples adjacent to the right boundary of a current block of a size nWxnH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right of the current block.
[0162] Meanwhile, when MRL (Multiple Reference Line) is applied, reference samples may be located on lines 1 to 3, rather than line 0 adjacent to the current block on the left / upper side. In this case, the number of surrounding reference samples may increase. The specific areas and number of surrounding reference samples are described below.
[0163] Meanwhile, when ISP (Intra Sub-Partitions) is applied, the surrounding reference samples can be derived in units of sub-partitions.
[0164] DIMD(Decoder side intra mode derivation)
[0165] In DIMD, intra prediction can be derived as a weighted average between the planar and two derived directions. To this end, two angular modes are selected from the Histogram of Gradients (HoG) derived from the surrounding pixels of the current block. Once these two modes are selected, the predictors (prediction blocks) and the planar predictor are normally derived, and the weighted average can be used as the final predictor (final prediction block) of the current block. At this time, the corresponding amplitudes within the HoG are used for each of the two modes to determine the weights.
[0166] Since the derived intra modes are included in the primary list of the intra MPM, the DIMD process can be performed before constructing the MPM list. The primary derived intra modes of a DIMD block are stored with the block and can be used to construct the MPM lists of surrounding blocks.
[0167] The DIMD chroma mode can use the DIMD derivation method to derive the chroma intra prediction mode of the current block based on the reconstructed surrounding Y, Cb, and Cr samples in the second surrounding row and column illustrated in FIG. 12. Specifically, to build the HoG, horizontal gradients and vertical gradients can be computed for the reconstructed Cb and Cr samples as well as each collocated reconstructed luma sample of the current chroma block. Then, chroma intra prediction of the current chroma block can be performed using the intra prediction mode with the largest histogram amplitude value.
[0168] If the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the DM mode, the intra prediction mode with the second largest histogram amplitude value can be used as the DIMD chroma mode. A predetermined CU level flag can be signaled to indicate whether the above-described DIMD chroma mode is applied.
[0169] TIMD(Fusion for template-based intra mode derivation)
[0170] For each intra prediction mode within the MPM, the SATD between the template prediction sample and the reconstructed sample can be calculated. Then, the first two intra prediction modes with the smallest SATD can be selected as the TIMD modes. These two TIMD modes can be fused based on weights, and this weighted intra prediction can be used to code the current CU. The derivation of the TIMD mode can include the position-dependent intra prediction combination (PDPC) described above.
[0171] The costs of the two selected modes are compared with a predetermined threshold, and a cost factor 2 can be applied as in mathematical expression 1 below.
[0172]
[0173] If the condition of the above mathematical expression 1 is true, the above-described fusion can be applied. Conversely, if the condition of the above mathematical expression 1 is false, only mode 1 can be used.
[0174] Meanwhile, the weights of the above modes can be calculated from each SATD cost as in mathematical equation 2 below.
[0175]
[0176] Matrix-based intra prediction (MIP)
[0177] FIGS. 7 to 11 are diagrams for explaining a MIP (Matrix-based intra prediction) process that can be applied to the present disclosure.
[0178] MIP (Matrix-based intra prediction) can be referred to as ALWIP (Affine linear weighted intra prediction) or MIP (or MWIP) (Matrix weighted intra prediction). To predict samples of a rectangular block with width W and height H, MIP takes as input a line consisting of H reconstructed samples adjacent to the left boundary of the block and a line consisting of W reconstructed samples adjacent to the top boundary. If reconstructed samples are unavailable, they can be generated in the same way as those generated by conventional intra prediction.
[0179] Prediction signals can be generated essentially by the following steps:
[0180] 1. In the case where W=H=4 among the boundary samples, four samples can be extracted by averaging them, and in other cases, eight samples can be extracted by averaging them (averaging process).
[0181] 2. The matrix-vector multiplication performed after the offset addition can be performed using the averaged samples as input. As a result of this process, a reduced prediction signal for a subsampled sample set of the original block can be obtained (matrix-vector multiplication process).
[0182] 3. The prediction signals at the remaining positions can be generated through linear interpolation, which is a single step linear interpolation in each direction from the prediction signals of the subsampling set ((linear) interpolation process).
[0183] The matrices and offset vectors required to generate prediction signals (prediction blocks or prediction samples) are three sets of matrices, S 0 , S 1 , can be obtained from S2. The set S0 has 18 matrices , each matrix has 16 rows, 4 columns, and 18 offset vectors, each of size 16. can have. The matrices and offset vectors of the set can be used for blocks of size 4x4. Meanwhile, set S1 has 10 matrices , each matrix has 16 rows, 8 columns, and 10 offset vectors, each of size 16. can have. The matrices and offset vectors of the set can be used for blocks of sizes 4x8, 8x4, and 8x8. Finally, the set S2 has six matrices , each matrix has 64 rows, 8 columns, and 6 offset vectors, each of size 64. can have. The matrices and offset vectors of that set can be used for blocks of any size other than those described above.
[0184] The number of multiplications required for matrix-vector product operations can always be less than or equal to 4·W·H. In other words, in MIP mode, up to four multiplication operations may be required for each sample.
[0185] Overview of the entire MIP process
[0186] The overall process of averaging, matrix vector multiplication, and linear interpolation is described with reference to FIGS. 7 to 11. Meanwhile, even in the case of block shapes not shown in FIGS. 7 to 11, processing can be performed as disclosed in FIGS. 7 to 11.
[0187] 1. In a 4x4 block, MIP can take two average values along the axis of each boundary. The resulting four input samples can be input to matrix-vector multiplication. The matrix can be obtained from set S0. After offset addition, 16 final prediction samples can be obtained. Linear interpolation may not be necessary to generate the prediction signal. That is, (4·16) / (4·4) = 4 multiplications can be performed for each sample.
[0188] 2. In a 4x4 block, MIP can take four average values along the axis of each boundary. The resulting eight input samples can be input to a matrix-vector multiplication. The matrix can be obtained from set S1. Therefore, 16 samples can be obtained from odd positions of the prediction block. That is, (8·16) / (8·8) = 2 multiplications can be performed for each sample. After offset addition, these samples can be vertically interpolated using the reduced top boundary. Horizontal interpolation can be performed using the existing (original) left boundary. In this case, the interpolation process may not require multiplications. That is, only two multiplications per sample may be required to derive the MIP prediction.
[0189] 3. In an 8x4 block, MIP can take four existing boundary values at the left boundary and four average values along the horizontal axis of the boundary. The resulting eight input samples can be input to a matrix-vector multiplication. The matrix can be obtained from set S1. Therefore, 16 samples can be obtained at each vertical position and odd horizontal positions of the prediction block. That is, (8·16) / (8·4) = 4 multiplications can be performed for each sample. After offset addition, these samples can be horizontally interpolated using the existing (original) left boundary. In this case, the interpolation process may not require an additional multiplication step. That is, only a total of 4 multiplications per sample may be required to derive the MIP prediction. The transposed instances can also be handled accordingly.
[0190] 4. In a 16x16 block, MIP can take four averages along each axis of the boundary. The resulting eight input samples can be input to a matrix-vector multiplication. The matrix can be obtained from set S2. Therefore, 64 samples can be obtained from odd positions of the prediction block. That is, (8·64) / (16·16) = 2 multiplications can be performed for each sample. After offset addition, these samples can be vertically interpolated using the eight averages at the upper boundary. Horizontal interpolation can be performed using the original left boundary. In this case, no additional multiplication steps may be added in the interpolation process. That is, only two multiplications in total may be required for each sample to derive the MIP prediction.
[0191] The same process can be applied to larger blocks as well. In this case, the number of multiplications per sample is less than 4.
[0192] In a Wx8 block with W > 8, only horizontal interpolation may be necessary, since samples are provided at odd horizontal and vertical positions. In this case, (8·64) / (W·8) = 64 / W multiplications per sample may be performed to compute the reduced prediction. For W = 16, no additional multiplications may be required for linear interpolation, but for W > 16, the number of additional multiplications per sample required for linear interpolation may be less than 2. Therefore, the total number of multiplications per sample may be 4 or less.
[0193] Finally, for W×4 blocks where W>8, A k Let be a matrix generated by excluding all rows corresponding to odd entries along the horizontal axis of the downsampled block. In this case, the output size is 32, and only horizontal interpolation can be performed again. For the reduced prediction calculation, (8·32) / (W·4)=64 / W multiplications can be performed per sample. For W=16, no additional multiplications are required, but for W>16, linear interpolation may require fewer than two multiplications per sample. Therefore, the total number of multiplications can be 4 or less. The transposed case can also be handled accordingly.
[0194] Boundary averaging
[0195] Figure 12 is a diagram illustrating an example of a boundary averaging process that can be applied to the present disclosure. As an example, boundary averaging can be included in the MIP process.
[0196] Depending on the averaging process, averaging can be applied to each boundary (left boundary or upper boundary). Here, the boundary can mean a neighboring reference sample adjacent to the boundary of the current block as described with reference to FIGS. 7 to 11. For example, the left boundary (bdry left ) represents the left neighbor reference sample adjacent to the left boundary of the current block, and the upper boundary (bdry top) may represent the upper neighbor reference sample adjacent to the upper boundary of the current block. If the size of the current block is 4x4, each boundary size can be reduced to two samples through the averaging process. If the size of the current block is not 4x4, each boundary size can be reduced to four samples through the averaging process.
[0197] In the first step, the input boundary bdry top and bdry left is a smaller boundary and can be reduced to . Here, and All can consist of 2 samples for 4x4 blocks, and 4 samples for all other cases.
[0198] For a 4x4 block, it can be defined as follows for 0≤i<2:
[0199]
[0200] and, It can also be defined similarly to mathematical equation 3.
[0201] In other cases, the block width W is W = 4·2 k If , 0≤i<4, it can be defined as follows:
[0202]
[0203] and, It can also be defined similarly to mathematical equation 4.
[0204] These two reduced boundaries and is the reduced boundary vector bdry red Since it is connected to , the size can be 4 for blocks of 4x4 shape, and 8 for blocks of other shapes.
[0205] Meanwhile, if the variable mode represents the MIP mode, the following concatenation can be defined:
[0206]
[0207] Finally, in the interpolation of the subsampled prediction signal, for large blocks, a second version, i.e., a corrected version of the averaged boundary, may be needed. That is, min(W, H)>8 and W≥H, W=8*2. l , and for 0≤i<8, the following can be defined:
[0208]
[0209] If min(W, H)>8 and H>W, can be defined similarly.
[0210] Generation of reduced prediction signals by matrix-vector multiplication
[0211] Reduced input vector bdry red One reduced prediction signal pred red can be generated. The reduced prediction signal has a width W red and height H red It can be a signal for a downsampled block. Here, W red and H red can be defined as follows:
[0212]
[0213] Reduced prediction signal pred red can be derived by matrix-vector multiplication and offset addition:
[0214]
[0215] Here, A is W red ·H red It can be a matrix with 4 rows and 4 columns if W=H=4 and 8 columns otherwise. Here, b is W red ·H red It can be a vector of size.
[0216] The matrix A and vector b can be obtained from one of the sets S0, S1, S2 as follows. The index idx = idx(W, H) can be defined as follows:
[0217]
[0218] Afterwards, if idx<=1 or idx=2 and min(W, H)>4, and can be. When idx=2 and min(W, H)=4, A can correspond to the odd x-coordinate of the downsampled block for W=4 or to the odd y-coordinate of the downsampled block for H=4.
[0219] Finally, the reduced prediction signal can be replaced by a pre-signal if:
[0220]
[0221] pred red The total number of multiplications required to compute A can be 4 if W=H=4, since in this case A has 16 rows and 4 columns. Otherwise, A has 8 columns and W red ·H red It has a row of 8, in this case 8·W red ·H red <= We can see that 4·W·H multiplications are required. That is, in this case, pred red Up to four multiplications per sample may be required to compute .
[0222] linear interpolation
[0223] Fig. 13 is a diagram illustrating an example of a linear interpolation process that can be applied to the present disclosure. As an example, the linear interpolation process can be included in a MIP process. As an example, the interpolation process can be a linear interpolation or a bilinear interpolation process. The interpolation process can include two steps, 1) vertical interpolation and 2) horizontal interpolation, as illustrated in Fig. 13. If W>=H, vertical linear interpolation can be applied first, and horizontal linear interpolation can be applied thereafter. W <H이면, 수평 선형 보간이 먼저 적용될 수 있으며, 수직 선형 보간이 그 이후 적용될 수 있다. 4x4 블록에서, 보간 프로세스는 생략될 수 있다.
[0224] In W × H blocks where max(W, H)≥8, the prediction signal is linearly interpolated to W red ×H red Reduced prediction signal pred red Based on the block shape, linear interpolation can be applied in the vertical, horizontal, or both directions. When linear interpolation is applied in both directions, it can be applied in the horizontal direction first, and in the vertical direction first when W < H.
[0225] Suppose there is a WxH block where max(W, H) >= 8 and W >= H. Here, one-dimensional linear interpolation can be performed as follows. Considering generality, linear interpolation in the vertical direction can be performed. First, the reduced prediction signal can be extended to the top by the boundary signal. The vertical upsampling factor is U ver = H / H red Define it as, In this case, the extended reduced prediction signal can be defined as follows:
[0226]
[0227] In this case, a vertical linear interpolation prediction signal can be generated from the extended reduced prediction signal as follows.
[0228]
[0229] Example
[0230] The present disclosure relates to intra-screen prediction, and more specifically, to a technique for performing Matrix-based Intra Prediction (MIP) adaptively to various conditions. According to one embodiment of the present disclosure, Matrix-based Intra Prediction (MIP) may be performed using only a portion of reference samples, and MIP may be performed adaptively based on conditions related to the reference samples and the current block, etc.
[0231] FIG. 14 is a diagram illustrating a MIP prediction method applicable to the present disclosure. According to an example of FIG. 14, for prediction of the current block, surrounding reference samples are used as input values, and prediction values of the current block can be obtained as output values through a neural network.
[0232] Meanwhile, current MIP prediction can generate predicted blocks using both left and top reference samples surrounding the block. However, samples in the current block may have biased correlations with some reference samples, such as the top or left reference samples. Therefore, using partial reference samples to predict the current block can yield higher prediction accuracy.
[0233] Accordingly, the present disclosure proposes a method for adaptively performing MIP prediction on a reference sample region in various embodiments. That is, a method for performing MIP prediction by utilizing only a portion of the surrounding reference samples of the current prediction block is proposed.
[0234] In addition, the present disclosure proposes a method for adaptively performing MIP prediction based on conditions associated with reference samples and / or current blocks, etc., in various embodiments. More specifically, matters related to MIP application can be determined, such as adaptively determining a MIP prediction result value or determining a matrix for MIP prediction based on the size of the current block, the number of pixels of the current block, and / or the values of reference samples.
[0235] In addition, when describing various embodiments in the present disclosure, even if singular expressions such as sample, pixel, pixel, and block are used, they may mean samples, pixels, pixel, and blocks, and of course, they may also mean a single sample, pixel, pixel, or block.
[0236] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings.
[0237] FIG. 15 is a diagram for explaining a method for performing MIP prediction according to an embodiment of the present disclosure, and FIG. 16 is a diagram for explaining a method for MIP-based image encoding or decoding according to an embodiment of the present disclosure. In this embodiment, with reference to various drawings including FIG. 15 and FIG. 16, a method for performing MIP prediction using only some of the reference samples among the reference samples of the current block will also be explained.
[0238] Meanwhile, the image encoding method or image decoding method of FIG. 16 can be performed in an image encoder or an image decoder, respectively, and the image encoder and image decoder can include the devices described with reference to other drawings above.
[0239] First, a MIP-based image encoding / decoding method (S1610) according to one embodiment of the present disclosure may be performed. As an example, the S1610 method may be performed based on some selected reference samples.
[0240] As an example, in order to perform MIP-based image encoding or decoding, some of the left reference sample and the upper reference sample may be selected. As an example, the upper reference sample of the current block may be selected to be used for performing MIP (S1620). As described above, pixels of the current block may have a more biased correlation with some reference samples, such as the upper or left reference sample, and therefore, when prediction of the current block is performed using partial reference samples, higher prediction accuracy may be obtained. Here, according to an embodiment of the present disclosure according to FIG. 15, only the upper reference sample may be selected for MIP prediction and used as a MIP prediction input value. In this case, only the upper reference sample may be selected as a neural network input value for MIP prediction. Meanwhile, as illustrated in (a) of FIG. 15, the upper reference sample may correspond to a sample immediately adjacent to the current block. That is, according to the embodiment of FIG. 15 (a), only the reference sample immediately above the current block may be used. However, as another example, as illustrated in (b), adjacent samples that are adjacent to the top of the current block but have a greater number of adjacent samples (e.g., extended further to the right) than the width of the current block may be included in the upper reference sample, or reference samples of multiple reference sample lines may be included in the upper reference sample. According to the embodiment of (b) of Fig. 15, multiple lines of reference samples at the top and top-right of the current block may be used. As another example, as illustrated in (c), adjacent samples that are adjacent to the top of the current block but have a greater number of adjacent samples (e.g., extended further to the left and / or right) than the width of the current block may be included in the upper reference sample, or reference samples of multiple reference sample lines may be included. According to the embodiment of (c) of Fig. 15, reference samples at the top or top-right or top-left of the current block may be used.Meanwhile, the number of reference samples extended further to the left or right or the number of reference sample lines is not limited as disclosed in FIG. 15. In addition, for the method proposed in this embodiment, some samples within one or more reference sample lines may be utilized for MIP prediction, and only some areas of all upper reference samples of the current block may be utilized for the MIP prediction proposed in this embodiment. Meanwhile, a method of utilizing only some of the reference samples obtained through the examples listed above, as well as a case of using reference samples on the upper and upper-right sides of the current block for MIP prediction of the current block, may also be included as an embodiment of the present disclosure. For example, only reference samples derived through subsampling (wherein, the subsampling ratio may be any number) may be utilized, or a method of using only reference samples distinguished based on a specific threshold value may be used. For example, the reference sample may be used for MIP prediction only if its value is above or below a certain threshold, or within a certain range, and the specific threshold may be signaled from the bitstream or derived based on other parameters (e.g., the size of the current block, etc.).
[0241] Thereafter, a MIP mode may be selected (S1630) based on MIP mode information. The MIP mode may be applied to the current block or may be explicitly signaled via the bitstream. Meanwhile, the current step may be replaced by a step in which the video encoder selects a MIP mode and then determines MIP mode information.
[0242] Thereafter, based on the selected MIP mode, a neural network-based MIP may be applied (S1640). In other words, a MIP-based prediction may be performed. Meanwhile, as an example, based on the selected reference sample input value, a prediction block sample value, which is an output value, may be obtained through a neural network as shown in FIG. 14 and / or FIG. 15. Here, the neural network may be adaptively selected based on the reference sample input value. In other words, step S1640 may include a process of applying a MIP-based prediction based on the selected input sample value and the selected neural network. As an example, a neural network composed of a single matrix may be used to perform the MIP. Meanwhile, the prediction block sample value, which is an output value, may be derived as follows:
[0243]
[0244] In the above formula, r represents a reference sample input value selected through the method proposed in this embodiment, and pred represents a predicted block sample value, which is an output value. In other words, the predicted block sample value can be derived based on the selected reference sample value. As an example, A can represent a trained neural network matrix, and the Clip function is a clipping function that adjusts (clip) the output predicted sample (pred) value to a certain sample value range, and may be a function for limiting the lower value and / or the upper value. In this case, the clipping range may be predefined. That is, the output value pred can be obtained through an operation on one matrix A for the input value r. The size of the matrix A can be adaptively changed depending on the sizes of the input value and the output value.
[0245] Meanwhile, the configuration of the neural network for obtaining the output value pred for the input value r proposed in this embodiment is not limited to the example above. That is, in addition to the neural network using a single simple matrix, multiple matrix operations may be used through the configuration of multiple hidden layers, and bias terms and activation functions may be used for each neural network layer.
[0246] Additionally, when outputting a prediction sample, values may be output in the frequency domain in addition to the spatial domain, and in this case, an inverse transform may be performed on the prediction (pred) value to obtain the final prediction sample value.
[0247] Meanwhile, as in the previous example, the size of the neural network matrix that can be included in the matrix for MIP can be adaptively selected depending on the input and output values. When MIP is applied to blocks larger than a certain size, the size of the matrix and the size of the input vector (e.g., r) become very large, which may require great computational complexity for the image decoding and / or encoding device. Therefore, it may be required to reduce the sizes of the input vector and the MIP matrix (or kernel) for coding efficiency and computational complexity. Accordingly, in one embodiment of the present disclosure, the selected input values, i.e., the selected reference samples, may be sampled through a method such as appropriate down-sampling and then used as the inputs of the neural network. That is, the reference samples sampled through down-sampling of the selected reference samples at a ratio of 1:2, 1:4, or 1:8 can be used as the inputs of the neural network. In this case, since the size of the input values can be reduced, the size of the neural network matrix can be reduced, which can have the effect of significantly reducing the overall amount of MIP operations.
[0248] Meanwhile, the sampling method and ratio of the reference samples selected in this example are not limited to the example above, and the input value can be increased through upsampling in addition to downsampling. The sampling method and ratio of the reference samples can be adaptively selected according to the shape and size of the block and the number of selected reference samples. In addition, the selected reference samples can be preprocessed through various filters such as low-pass filtering and high-pass filtering and then used as input values for a neural network. In addition, it is possible to apply a transform to the selected reference samples to change them to the frequency domain and then input them to a neural network. In addition, after calculating the average value of the selected reference samples and subtracting the average value from all reference samples, the residual sample value can be used as the input value for the neural network.
[0249] Meanwhile, the filtering and sampling methods for input values described above can be similarly applied to output values. That is, if the number of output values of the neural network, i.e., the number of output samples, is less or more than the number of prediction samples required for the current block size, the output values can be upsampled or downsampled to match the number of prediction samples required for the current block size. In addition, if a transformation is applied to the input values so that frequency-domain input values are used, or residual samples excluding the mean value are used as input values, this can be compensated for in the output values.
[0250] As in the method of the previous example, the number of input samples and output samples for the neural network matrix operation proposed in the present disclosure is not limited to a specific number.
[0251] Additionally, post-processing filtering can be applied to the predicted blocks predicted using the method proposed in this disclosure. The Position Dependent Intra Prediction Combination (PDPC) method can also be applied to the predicted blocks predicted using the method proposed in this disclosure. For example, a smoothing filter can be applied to the predicted blocks predicted using the method proposed in this disclosure.
[0252] Meanwhile, as an example, when configuring a neural network for the method proposed in this embodiment, the coefficients of the matrix used in the operations within the neural network may have various bit precisions. For example, when performing a 10-bit precision matrix operation, the matrix coefficients have a coefficient value range of a specific value range (e.g., 0 to 1023 or -512 to 511), and the input and output values may be adaptively adjusted accordingly. These bit precision values are not limited to the above examples and may also be adaptively changed based on the number of layers of the neural network, etc.
[0253] Meanwhile, in the method proposed in this embodiment, the number of neural network matrix modes (the number of modes of the proposed MIP method) can be configured in various ways based on the mode (e.g., intra mode or inter mode, etc.), the width of the input block, the height of the input block, the number of pixels of the input block, the position of sub-blocks within the block, explicitly signaled syntax, statistical characteristics of surrounding pixels, whether or not a secondary transformation is used, etc. For example, the number of neural network matrix modes can be adaptively selected depending on the size of the current block. Alternatively, the number of neural network matrix modes can be equally selected for the sizes of all blocks.
[0254] As an example, the coefficients of the neural network matrix may be adaptively defined for all block shapes. For example, neural network coefficients corresponding to each block shape from 4x4 to 256x256 blocks may be defined. The number of input values and output values of the neural network matrix may be predetermined depending on the block shape. For example, as in the example of Fig. 15 (a), MIP prediction is performed using only one upper reference sample line, and 49 types of neural network matrices may be required that match the input values of the neural network matrix of 4, 8, 16, 32, 64, 128, 256 and the output values of 4, 8, 16, 32, 64, 128, 256. In addition, each of the 49 neural network matrices may have matrix coefficients according to the number of MIP modes proposed in this embodiment. For example, if the number of MIP modes of a neural network matrix with 8 inputs and 16 outputs (which may correspond to an 8x16 block shape) is 35, the neural network matrix may have a total of 8x16x35 matrix coefficients. Or, as another example, only neural network matrices suitable for a specific block shape may be defined, and may have limited input and output values. In this case, a neural network matrix of a specific block shape may be utilized through appropriate preprocessing or postprocessing of the input and output values. For example, if only neural network matrices suitable for a square block shape are defined, the input values of the neural network matrix of the square block can be corresponded to the input values through preprocessing of the reference samples for non-square blocks (e.g., downsampling or upsampling). The output values can also be corresponded to the predicted values of the non-square blocks through postprocessing (e.g., downsampling or upsampling).
[0255] Meanwhile, according to the method proposed in this embodiment, the neural network matrix mode (the mode of the proposed MIP method) can be selected based on at least one of the mode (e.g., inter mode or intra mode, etc.), the width of the input block, the height of the input block, the number of pixels of the input block, the position of sub-blocks within the block, explicitly signaled syntax, statistical characteristics of surrounding pixels, and whether or not a secondary transform is used. In this case, the encoder can select and apply an appropriate binarization method to transmit the selected mode information to the decoder. For example, the encoder can set a specific mode, such as the transmission of the intra luma mode, to the MPM mode, and then perform binarization after separately setting the remaining modes. In this case, binarization bits can be saved through appropriate context modeling during each binarization. Alternatively, as another example, the mode information can be transmitted to the decoder using binarization such as truncated binary / truncated unary / fixed length, taking into account the total number of MIP modes.
[0256] Meanwhile, in order to select the first and second transformation types in the prediction block predicted by the method proposed in this embodiment, the transformation mode may be selected based on at least one of the neural network matrix mode, the width of the input block, the height of the input block, the number of pixels in the input block, the position of sub-blocks within the block, explicitly signaled syntax, and statistical characteristics of surrounding pixels. For example, when selecting the first transformation type of the block to which the prediction method proposed in this disclosure is applied, the MTS (multiple transform set) transformation method may be applied. For example, when selecting the second transformation type of the block to which the prediction method proposed in this disclosure is applied, a second transformation kernel suitable for the planar intra mode may be selected.
[0257] Meanwhile, for the MIP mode selection proposed in this embodiment, a MIP prediction mode candidate of the current block, such as DIMD or TIMD, may be determined, and the MIP prediction mode of the current block may also be determined. That is, a MIP prediction mode candidate or MIP prediction mode of the corresponding block may be inferred by utilizing surrounding reference samples, such as DIMD or TIMD.
[0258] As another example, the mode candidate inference method used in DIMD may be applied to the MIP mode prediction of the method proposed in this embodiment. That is, the pixel gradient can be calculated using only the region of the upper reference sample, and the first or second MIP prediction mode obtained through the gradient can be set as the MIP prediction mode candidate of the current block.
[0259] As another example, the mode inference method used in TIMD can be applied to the MIP mode prediction of the method proposed in this embodiment. That is, TIMD template matching can be performed using only the region of the upper reference sample. The MIP prediction mode candidate to which TIMD template matching is applied can be selected, such as MPM, or all of the multiple MIP prediction modes can be selected. Alternatively, only a portion of any number of the multiple MIP prediction modes can be selected.
[0260] Meanwhile, whether the method proposed in this embodiment is applied or not can be signaled as information in HLS (high-level syntax) such as VPS, SPS, PPS, Picture Header, Slice Header, DCI, etc. For example, in order to determine whether the method proposed in this embodiment is applied or not on a PPS basis, information on whether the method proposed in this embodiment is applied or not can be signaled by including it in the PPS.
[0261] Additional information regarding the applicability of the proposed method in this embodiment may be explicitly signaled. Alternatively, the proposed method in this embodiment may be inferred or adaptively selected without additional information signaling. For example, the applicability of the proposed method in this embodiment on a CTU or CU basis may be signaled via a 1-bit flag.
[0262] As an example, the method proposed in this embodiment may be used only when the MIP mode is defined to be used in HLS (high level syntax). In addition, the method proposed in this embodiment may determine whether to apply or not by signaling additional information within the MIP mode. For example, if the value of the information on whether MIP is applied (e.g., MIP flag) is a specific value (e.g., 1, true), the information on whether the method proposed in this embodiment is applied or not may be signaled using a 1-bit flag.
[0263] Alternatively, if the method proposed in this embodiment can be applied depending on a specific block size or shape or the presence or absence of specific conditions, then whether or not the method proposed in this embodiment is applied may be signaled using a 1-bit flag. For example, if the block height is four times or more the block width, the method proposed in this embodiment may not be applied, and thus, the signaling of additional information regarding whether or not the method is applied accordingly may be omitted.
[0264] Meanwhile, as another example, under certain conditions, whether the method proposed in this embodiment is applied can be implicitly inferred. For example, if the left reference sample of the current block does not exist, such as a block at the left boundary of the image, or if the left reference sample is unavailable, signaling information on whether the method proposed in this embodiment is applied can be omitted. In addition, signaling information on whether the MIP mode is applied can also be omitted. In this case, it can be derived that the MIP method proposed in this embodiment is always applied. If the left reference sample of the current block is a CTU boundary / tile boundary / slice boundary / subpicture boundary, signaling information on whether the method proposed in this embodiment is applied can be omitted. In addition, signaling information on whether the MIP mode is applied can also be omitted. In this case, it can be derived that the MIP method proposed in this embodiment is always applied.
[0265] Meanwhile, as an example, whether or not to transmit information on the application of the proposed method of the present embodiment in a coding unit may be adaptively determined as the application of the proposed method of the present embodiment is determined at a higher level defined in HLS (high level syntax). For example, if the value of the information indicating whether or not to apply the proposed method of the present embodiment signaled in SPS is a specific value (e.g., 0, false) (i.e., indicating that the method proposed in the present embodiment is not used in the SPS unit), the method proposed in the present embodiment is not applied even at the coding unit level, and signaling of the usage information accordingly may not be performed.
[0266] FIG. 17 and FIG. 18 are diagrams for explaining a method for performing MIP prediction according to an embodiment of the present disclosure, and FIG. 19 is a diagram for explaining a method for MIP-based image encoding or decoding according to an embodiment of the present disclosure. In this embodiment, with reference to all of FIG. 17 to FIG. 19, a method for performing MIP prediction using only some of the reference samples among the reference samples of the current block will also be explained.
[0267] Meanwhile, the image encoding method or image decoding method of FIGS. 19a and 19b may be performed in an image encoder or an image decoder, respectively, and the image encoder and image decoder may include the devices described with reference to other drawings above. First, the image encoding method and / or image decoding method of FIG. 19a will be described.
[0268] First, a MIP-based image encoding / decoding method (S1910) according to one embodiment of the present disclosure may be performed. As an example, the S1910 method may be performed based on some selected reference samples.
[0269] For example, to perform MIP-based image encoding or decoding, some of the left reference samples and the upper reference samples may be selected. For example, the left reference sample of the current block may be selected (S1920) to be used for performing MIP.
[0270] As described above, pixels of the current block may have a more biased correlation with some reference samples, such as the upper or left reference samples, and therefore, when prediction of the current block is performed using partial reference samples, higher prediction accuracy can be obtained. Here, according to one embodiment of the present disclosure according to FIG. 18, only the left reference sample may be selected for MIP prediction and used as the MIP prediction input value. In this case, only the left reference sample may be selected as the neural network input value for MIP prediction. Meanwhile, as illustrated in (a) of FIG. 18, the left reference sample may correspond to a sample immediately adjacent to the current block. That is, according to the embodiment of FIG. 18 (a), only the reference sample immediately to the left of the current block may be used. However, as another example, as illustrated in (b), adjacent samples to the left of the current block but having a greater number of adjacent samples (e.g., extending downwards) than the height of the current block may be included in the left reference sample, or reference samples of multiple reference sample lines may be included in the left reference sample. According to the embodiment of Fig. 18 (b), multiple lines of reference samples at the top and left-bottom of the current block may be used. As another example, as illustrated in (c), adjacent samples (e.g., extended further to the top and / or bottom) adjacent to the left side of the current block but greater than the height of the current block may be included in the left reference sample, or reference samples of multiple reference sample lines may be included. According to the embodiment of Fig. 18 (c), reference samples at the left side or at the left-top or left-bottom of the current block may be used. Meanwhile, the number of reference samples extended further to the top or bottom or the number of reference sample lines is not limited as disclosed in Fig. 18.In addition, for the method proposed in this embodiment, some samples within one or more reference sample lines may be utilized for MIP prediction, and only some areas of all left reference samples of the current block may be utilized for the MIP prediction proposed in this embodiment. Meanwhile, the reference samples for MIP prediction may be reference samples on the left and / or upper left side of the current block, or may be reference samples obtained through the examples described above. For example, only reference samples obtained through subsampling (wherein the subsampling ratio may be any number) may be used, or a method of using only reference samples distinguished based on a specific threshold value may be used. Since the description of the specific threshold value is the same as that described with reference to FIGS. 14, 15, and 16 above, a redundant description thereof will be omitted.
[0271] Thereafter, a MIP mode may be selected (S1930) based on MIP mode information. The MIP mode may be applied to the current block or may be explicitly signaled via the bitstream. Meanwhile, the current step may be replaced by a step in which the video encoder selects a MIP mode and then determines MIP mode information.
[0272] Thereafter, based on the selected MIP mode, a neural network-based MIP may be applied (S1940). In other words, a MIP-based prediction may be performed. Meanwhile, as an example, based on the selected reference sample input value, a prediction block sample value, which is an output value, may be obtained through a neural network as shown in FIG. 14 and / or FIG. 17. Here, the neural network may be adaptively selected based on the reference sample input value. In other words, step S1940 may include a process of applying a MIP-based prediction based on the selected input sample value and the selected neural network.
[0273] As an example, embodiments related to the configuration of a neural network, a neural network matrix, input values of the neural network, processing of output values, signaling of information related to the embodiment, and / or a MIP prediction mode are the same as those described above with reference to FIGS. 14, 15, and 16, and therefore, redundant descriptions are omitted.
[0274] An embodiment of the image encoding method and / or image decoding method illustrated in FIG. 19B is described based on the embodiments described with reference to FIGS. 14 to 19A above. As an example, according to FIG. 19B, it can be determined whether to use only the reference samples located at the top of the current block based on the embodiments of FIGS. 14 to 16, or to use only the reference samples located at the left of the current block based on the embodiments of FIGS. 17 to 19A.
[0275] First, a MIP-based image encoding / decoding method (S1950) according to one embodiment of the present disclosure may be performed. As an example, the S1950 method may be performed based on some selected reference samples.
[0276] As an example, in order to perform MIP-based image encoding or decoding (S1950), some of the left reference sample and the upper reference sample may be selected. As an example, only either the left reference sample or the upper reference sample of the current block may be selected to be used for performing MIP (S1960). As an example, information regarding whether to use the upper reference sample or the left reference sample as a reference sample may be explicitly signaled, but may also be derived from an agreement between the encoder and the decoder without separate information signaling. As an example, in the absence of separate information signaling, the TIMD template matching method may be used, and TIMD template matching may be performed using only the upper reference sample region, or TIMD template matching may be performed using only the left reference sample region. In this case, the template matching error values of each reference sample may be compared, and a reference sample region with a smaller error value may be used as an input sample for MIP. At this time, the number of samples in each sample area may be different, and a process may be added to compare the error value of each reference sample area by dividing it by the number of samples in each sample area for accurate comparison.
[0277] Thereafter, a MIP mode may be selected (S1970) based on MIP mode information. The MIP mode may be applied to the current block or may be explicitly signaled via the bitstream. Meanwhile, the current step may be replaced by a step in which the video encoder selects a MIP mode and then determines MIP mode information.
[0278] Afterwards, based on the selected MIP mode, a neural network-based MIP can be applied (S1980).
[0279] Since the embodiment of Fig. 19b can be applied to the contents described in Figs. 14 to 19a, a duplicate description thereof will be omitted.
[0280] FIG. 20 and FIG. 21 are diagrams for explaining a method for performing MIP prediction according to an embodiment of the present disclosure, and FIG. 22 is a diagram for explaining a method for MIP-based image encoding or decoding according to an embodiment of the present disclosure. In this embodiment, with reference to all of FIG. 20, FIG. 21, and FIG. 22, a method for performing MIP prediction using only some of the reference samples among the reference samples of the current block will also be explained.
[0281] Meanwhile, the image encoding method or image decoding method of FIG. 22 can be performed in an image encoder or an image decoder, respectively, and the image encoder and image decoder can include the devices described with reference to other drawings above.
[0282] First, a MIP-based image encoding / decoding method (S2210) according to one embodiment of the present disclosure may be performed. As an example, the S2210 method may be performed based on some selected reference samples.
[0283] As an example, in order to perform MIP-based image encoding or decoding, some of the left reference samples and some of the upper reference samples may be selected. As an example, only some of the left and upper reference samples of the current block may be selected to be used for performing MIP (S2220). That is, in the embodiments of FIGS. 20 and 21, some or an extended left reference sample and some or an extended upper reference sample may be selected and used. As described above, pixels of the current block may have a more biased correlation with some reference samples, such as the upper or left reference sample, and therefore, when prediction of the current block is performed using partial or extended reference samples, higher prediction accuracy may be obtained. Here, according to one embodiment of the present disclosure according to FIG. 21, some or an extended reference sample among adjacent reference samples existing on the left or upper side may be selected for MIP prediction and used as MIP prediction input values. In this case, some left reference samples and some upper reference samples may be selected as neural network input values for MIP prediction. Meanwhile, as illustrated in (a) of FIG. 21, the left reference sample and the upper reference sample may correspond to samples immediately adjacent to the current block. That is, according to the embodiment of (a) of FIG. 21, only the reference sample immediately to the left of the current block may be used, and only the reference sample immediately above the current block may be used. In other words, only the reference samples adjacent to the current block may be used. However, as illustrated in (a) of FIG. 21, the reference sample immediately adjacent to the current block is used, but a block located at the upper part of the left reference samples corresponding to the height of the current block is excluded, and a left reference sample extended further downward may be included in the reference sample.In addition, as shown in (a) of Fig. 21, reference samples immediately adjacent to the current block are used, but blocks located on the left side of the upper reference samples corresponding to the width of the current block are excluded, and upper reference samples extended further to the right may be included in the reference samples. However, as another example, as shown in (b) of Fig. 21, adjacent reference samples adjacent to the left side of the current block but in a number less than the number of samples corresponding to the height of the current block may be included in the left reference samples, or reference samples of multiple reference sample lines may be included in the left reference samples. In addition, as shown in (b) of Fig. 21, adjacent reference samples adjacent to the top of the current block but in a number less than the number of samples corresponding to the width of the current block may be included in the upper reference samples, or reference samples of multiple reference sample lines may be included in the upper reference samples. In addition, the reference samples may further include samples located on the upper left side of the current block. That is, reference samples that are more extended to the left from the upper reference sample or more extended to the top from the left reference sample can be further selected. According to the embodiment of (b) of Fig. 21, multiple lines of reference samples on the left, top, and / or top-left of the current block may be used. Meanwhile, as another example, as shown in (c), the left reference sample of the current block may include reference samples corresponding to the height of the current block, but the upper reference sample may include only a smaller number of adjacent samples than the reference samples corresponding to the width of the current block. That is, only some of the upper adjacent reference samples may be included in the upper reference sample.
[0284] Meanwhile, the number of reference samples extended further upwards or downwards or the number of reference sample lines is not limited as disclosed in Fig. 21. In addition, for the method proposed in this embodiment, some samples within one or more reference sample lines may be utilized for MIP prediction, and only some areas of all left reference samples or upper reference samples of the current block may be utilized for MIP prediction proposed in this embodiment.
[0285] Thereafter, a MIP mode may be selected (S2230) based on MIP mode information. The MIP mode may be applied to the current block or may be explicitly signaled via the bitstream. Meanwhile, the current step may be replaced by a step in which the video encoder selects a MIP mode and then determines MIP mode information.
[0286] Thereafter, based on the selected MIP mode, a neural network-based MIP may be applied (S2240). In other words, a MIP-based prediction may be performed. Meanwhile, as an example, based on the selected reference sample input value, a prediction block sample value, which is an output value, may be obtained through a neural network as shown in FIG. 14 and / or FIG. 20. Here, the neural network may be adaptively selected based on the reference sample input value. In other words, step S2240 may include a process of applying a MIP-based prediction based on the selected input sample value and the selected neural network. As an example, embodiments related to the configuration of a neural network, a neural network matrix, input values of the neural network, processing of output values, signaling of information related to the embodiment, and / or a MIP prediction mode, etc. are the same as those described above with reference to FIGS. 14, 15, 16, 17, 18, 19a, 19b, 20, and 21, and therefore, redundant descriptions are omitted.
[0287] Meanwhile, additional information (e.g., 2 bits) may be signaled, and the information may be used to determine which of the embodiments of FIGS. 14 to 19B will be used. Alternatively, the method proposed in the present embodiment may be determined to be applied based on an agreement between the encoder and decoder without signaling. For example, the mode inference method used in TIMD may be applied. That is, TIMD template matching may be performed using only the upper reference sample region, or TIMD template matching may be performed using only the left reference samples. TIMD template matching may also be performed using only some of the upper reference samples and / or some of the lower reference sample regions. In addition, the template matching error values of each reference sample may be compared, and the reference sample region with the smaller error value may become the MIP input value. In this case, the number of samples in each sample region may be different, and a process may be added to compare the error value of each reference sample region by the number of samples in each sample region for accurate comparison.
[0288] Meanwhile, information related to MIP according to an embodiment of the present disclosure may be signaled. Below, information that may be signaled is described with reference to an exemplary table, but the names and order of the described information may be changed, and may be signaled at various levels (e.g., Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Slice Header, Picture Header, Coding Unit (CU), Prediction Unit (PU), etc.).
[0289] FIG. 23 is a drawing for explaining an image encoding method or an image decoding method that can be performed in an image encoding device or an image decoding device according to one embodiment of the present disclosure.
[0290] As an example, according to one embodiment of FIG. 23, MIP based on adaptive matrix selection can be performed (S2310). Below, a method for adaptively selecting and applying a neural network matrix, i.e., a neural network matrix, that can be applied during MIP prediction is proposed.
[0291] In order to perform S2310, first, the block size of the current block that can be the target of MIP prediction and a specific threshold value can be compared (S2320). As an example, MIP prediction can be performed using a neural network composed of a single matrix as in the above mathematical expression 12. In the above mathematical expression 12, r refers to a reference sample, which is an input value, and pred refers to a predicted block sample value, which is an output value. A can refer to a trained neural network matrix, and the Clip function can refer to a function that adjusts the output predicted value pred value to be within a certain range, i.e., clips it. That is, the output value pred can be obtained through a single matrix operation for the input value r. At this time, the size of the matrix, i.e., A, can be changed according to the sizes of the input value and the output value. In one embodiment of the present disclosure, a method is proposed for adaptively selecting and applying a neural network matrix applied during MIP prediction, i.e., A in the example of the above mathematical expression 12. For example, neural network matrices can be adaptively selected if certain conditions are met:
[0292] - Example 1: If the block size or the number of pixels in a block for which MIP prediction is to be performed is less than or equal to a certain size, a neural network matrix that does not apply sampling (e.g., downsampling and / or upsampling) to the input values (r) and output values (pred) may be applied. In other cases, a neural network matrix that applies sampling to the input values and output values may also be applied.
[0293] -Example 2: If the block size or number of pixels for which MIP prediction is to be performed is below or above a certain size, a more sophisticated and complex neural network matrix can be applied. A sophisticated neural network matrix may include a matrix that can process a larger number of input values and generate output values, has a large number of matrix factors, or requires additional processing to derive the corresponding matrix even when using a single matrix A. On the other hand, in other cases, that is, if the block size is above or below a certain size, a relatively simple neural network matrix may be used.
[0294] In the embodiment of Fig. 23, if the block size is less than a certain threshold, an enhanced matrix may be used for MIP (S2330). On the other hand, if the block size is greater than a certain threshold, a relatively simple or general matrix may be used for MIP (S2340). However, the present disclosure is not limited to Fig. 23, and the opposite case of Fig. 23 or other criteria described above may also be used to select a MIP matrix, and this will also be included in the present disclosure.
[0295] Meanwhile, in the above example, the size of the current block may mean the width (block width) or the height (block height) of the current block, or may be defined as the number of pixels in the current block. As another example, whether to apply sampling (e.g., downsampling or upsampling) or the factor value may be determined by considering cases where the block size is larger or smaller than a threshold value.
[0296] For example, even if the size of the current block is less than a certain size (e.g., 8x8), the size of the block may be relatively large compared to blocks adjacent to the current block with large changes in the distribution of pixel values, so a neural network matrix that does not apply sampling to the input and output values may be used in MIP. That is, for blocks of a certain size (e.g., 4x4, 4x8, 8x4, 8x8, etc.), sampling may not be applied to the input and output values. In this case, a neural network matrix corresponding to a block of the corresponding size (e.g., 4x4, 4x8, 8x4, 8x8 block, etc.) may be required.
[0297] As an example, assuming that MIP prediction is performed using only the upper reference sample, as in the embodiment of the present disclosure described with reference to FIGS. 14 to 16 and FIG. 19B above, and assuming that the size of the current block is 8x8, 8 input values and 64 output values may be required. In addition, a corresponding neural network matrix of 8x64 size may be required.
[0298] For example, if the number of pixels in a block is less than or equal to a certain value (e.g., 128), a neural network matrix that does not apply sampling to the input and output values can be used in MIP. That is, for blocks of a certain size (e.g., 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 16x4, 16x8, 32x4, etc.), sampling may not be applied to the input and output values, and a matrix corresponding to the specific size may be required.
[0299] As an example, assuming that MIP prediction is performed using only the left reference sample, as in the embodiment of the present disclosure described with reference to FIGS. 17 to 19B above, and assuming that the size of the current block is 8x16, 16 input values and 128 output values may be required. In addition, a corresponding neural network matrix of 16x128 size may be required.
[0300] Meanwhile, applying the method proposed in this embodiment can be expected to improve MIP prediction performance in small blocks. Furthermore, as block sizes increase, the required neural network matrix factors (i.e., neural network matrix coefficients) increase exponentially. Therefore, applying the method proposed in this embodiment can effectively suppress the increase in neural network matrix coefficient data, thereby reducing computational complexity and improving coding efficiency and quality.
[0301] Meanwhile, the matrix configuration for obtaining the output value pred for the input value r proposed in some embodiments of the present disclosure is not limited to the above example. That is, in addition to a neural network using a single simple matrix, multiple matrix operations may be used through a configuration of multiple hidden layers, and bias terms and / or activation functions may be used for each neural network layer. In addition, in addition to the spatial domain, a frequency domain value may be selected as the output value, and in this case, the final predicted sample values may be obtained by performing an inverse transform of the output value, that is, the predicted value, i.e., the pred value.
[0302] As in the previous example, the size of the neural network matrix can be adaptively selected according to the input and output values. In this embodiment, the input values selected through the method proposed in the previous example, i.e., the selected reference samples, can be sampled and used as the neural network input. That is, as described above, the selected reference samples can be sampled at a specific ratio (e.g., 1:2, 1:4, or 1:8). The sampling method and ratio of the selected reference samples in this example are not limited to the above example, and the input values can be increased / decreased through sampling, and the sampling method and ratio of the reference samples can be adaptively selected according to the shape and size of the block and the number of selected reference samples. In addition, as described above, the selected reference samples can be preprocessed through various filtering and used as the input values of the neural network. In addition, the values converted to the frequency domain by applying a transformation to the selected reference samples can be used as the matrix input values, or the residual sample values, etc. can be used as the input values. Since this is the same as described above, a duplicate explanation is omitted.
[0303] The filtering and sampling methods described above for input values can be similarly applied to output values. Furthermore, if frequency-domain input values are used or residual samples are used as input values, these can be compensated for in the output values. Since these are identical to the above descriptions, further explanation will be omitted.
[0304] As in the previous example, the number of input samples and output samples for matrix operations is not limited to a specific number.
[0305] Additionally, as described above, post-processing filtering can also be applied to the predicted prediction block. Since this is the same as described above, a detailed explanation will be omitted.
[0306] In addition, as explained above, when constructing a neural network, matrix factors, i.e. coefficients, can have various bit precisions. Since this is the same as explained above, a duplicate explanation will be omitted.
[0307] Meanwhile, as described above, the number of MIP prediction modes can be configured in various ways based on the prediction mode (e.g., intra mode, inter mode, etc.), width of the input block, height of the input block, number of pixels of the input block, position of sub-blocks within the block, explicitly signaled syntax elements, statistical characteristics of surrounding pixels, use of secondary transformation, etc., and a specific prediction mode can be selected based thereon. For example, the number of MIP prediction modes can be adaptively selected according to the size of the current block. Alternatively, the number of neural network matrix modes, i.e., MIP prediction modes, can be the same for all block sizes. Since this is the same as described above, redundant explanation is omitted.
[0308] Meanwhile, the encoder can apply an appropriate binarization method to the selected mode information to transmit it to the decoder after binarization. This is the same as described above, so a duplicate explanation is omitted.
[0309] In addition, the selection of the first and second transformation types, the selection of the MIP mode (e.g., TIMD or DIMD-based, etc.), the signaling of related information (e.g., HLS, etc.) and / or the derivation of related information in the predicted block predicted by the method proposed in this embodiment are the same as those described above, so a duplicate description is omitted.
[0310] Meanwhile, although the embodiment of FIG. 23 is such that the enhanced matrix is used for MIP when the block size is below a certain threshold, the present disclosure is not limited thereto, and the opposite case where the enhanced matrix is used for MIP when the block size is above a certain threshold is also possible, and this is also included in the embodiment of the present disclosure.
[0311] FIG. 24 is a diagram illustrating an image encoding method or a projection decoding method that can be performed in an image encoding device or an image decoding device according to an embodiment of the present disclosure. In this embodiment, an adaptive MIP method for a specific input value is proposed during MIP prediction.
[0312] As an example, according to one embodiment of FIG. 24, whether MIP will actually be performed depending on a specific input value can be further determined (S2410) regardless of whether MIP is allowed and / or applied. In other words, even if MIP allowance and / or application is determined based on other information or other parameters, a method is proposed in which MIP is not actually performed and the output value is limited to a specific value if the MIP input value satisfies a specific condition.
[0313] To perform S2410, first, a comparison may be made between input values (e.g., input values) for MIP prediction based on reference samples for MIP prediction (S2420). For example, it may be determined whether all input values are 0. Meanwhile, prior to S2420, a process may be performed in which a reference sample is selected and input values (e.g., input vectors) for MIP are determined.
[0314] As an example, MIP prediction can be performed using a neural network composed of a single matrix as in Equation 12 as described above, and a method for adaptively performing MIP prediction when an input value r is used for prediction is proposed.
[0315] - Example 1: If all reference sample values selected for MIP prediction are the same, or if all input values for MIP prediction are a specific value (e.g., 0) (e.g., all input vectors are 0), the result of multiplying the input values by a matrix (e.g., matrix A in Equation 12) always becomes a specific value (e.g., 0), so MIP prediction can be omitted without performing unnecessary complex matrix multiplication operations. In this case, all output values can be set to a specific value k. Here, the value of k can be adaptively changed according to the current MIP mode or other parameters. As an example, if all input values for MIP prediction are 0, the output value can be obtained through the following Equation 13.
[0316]
[0317] In mathematical expression 13, output is an output value, bitdepth is information related to bit depth, mode_num means the number of the current MIP mode, and max_mode_num can mean the total number of MIP modes.
[0318] As an example, if the MIP input values of a 16x16 block are all 0, the bitdepth is 10, the MIP modes are from 0 to 6 (max_mode_num = 6), and the current MIP mode is 3 (mode_num = 3), the output value by applying the formula in the example above is 512.
[0319] -Example 2: If all reference sample values selected for MIP prediction are the same, or if all input values for MIP prediction are specific values (for example, if the input vector is all 0, the first value of the input values can be set as in the following formula, and then MIP prediction can be performed.
[0320]
[0321]
[0322] In the above mathematical expression 14, input[x] means the xth value of the input value (e.g., input vector), bitdepth means the bit depth of the input pixel value, and ref[0] may mean the first value of the reference sample selected for MIP. The value A may mean a specific value set in advance, and according to mathematical expression 14, MIP prediction may be performed after replacing the first input value with a value other than 0. The value A may be a predefined value, or, as an example, may be fixed to a value of (1 << (bitdepth -1)), or may be adaptively changed based on the current MIP mode or other parameters.
[0323] - Example 3: If all the reference sample values selected for MIP prediction are the same, or if all the input values for MIP prediction are a specific value (for example, if the input vector is all 0 when it is 0), all the output values can be set to a specific value k. In this case, the MIP mode is limited to 1, and additional MIP mode signaling can be omitted. As an example, the output value k can be selected as the average of the selected reference sample values, or it can be set to the same as in other examples (for example, example 1 or example 2, etc.). In this example, since the MIP mode is unified to one, signaling of additional information about the mode can be omitted.
[0324] In the embodiment of Fig. 24, if all input values of the MIP prediction are 0, the MIP application rule based on the examples described above may be applied (S2430). In other words, if the MIP prediction input value is 0, the MIP is not performed and the output value may be determined as a specific value. On the other hand, in this case, the MIP based on the other embodiment described above may be applied (S2440), which may be one of the MIP embodiments described with reference to Figs. 14 to 23, etc.
[0325] However, the present disclosure is limited to FIG. 24, and other criteria for input values may also be used to determine whether to perform MIP, and these are also included in the present disclosure.
[0326] Hereinafter, an image decoding method that can be performed by an image decoding device (decoder) according to an embodiment of the present disclosure and a video encoding method that can be performed by an image encoding device (encoder) will be described with reference to FIGS. 25 and 26. As an example, FIGS. 25 and 26 may be based on the embodiment described with reference to FIGS. 14 and 24 above, and therefore, any description overlapping with the description of the above embodiment will be omitted below.
[0327] FIG. 25 is a diagram illustrating an image decoding method that can be performed by an image decoding device (decoder) according to one embodiment of the present disclosure.
[0328] As an example, in a video decoding method performed by an video decoding device, a prediction mode of a current block may be determined (S2510), and the step may include whether an intra prediction mode or an inter prediction mode is applied to the current block.
[0329] Based on the prediction mode applied to the current block being an intra prediction mode, information related to MIP (Matrix-based Intra Prediction) may be derived (S2520). This step may include determining whether to allow MIP, whether to apply MIP, etc., and, if it is determined that MIP is applicable, a process of deriving information related to the MIP to be applied may be included.
[0330] Meanwhile, as an example, MIP may be applied based on the size of the current block or the number of pixels included in the current block. More specifically, a matrix for MIP may be determined based on the size of the current block or the number of pixels included in the current block. Meanwhile, if it is determined that MIP is to be applied to the current block based on information related to MIP, the value of the prediction block of the current block may be limited to a specific value based on a specific condition, wherein the specific condition may be based on a reference sample value of the current block. As an example, if the specific condition is satisfied, application of MIP to the current block may be omitted regardless of the information related to MIP. Meanwhile, if it is determined that MIP is to be applied to the current block based on information related to MIP, the prediction block of the current block may be generated using only the upper reference sample or the left reference sample of the current block. Meanwhile, the upper reference sample may include the upper right reference sample of the current block, and the left reference sample may include the lower left reference sample of the current block. As another example, the prediction block of the current block is generated using a portion of the left reference sample and / or a portion of the upper reference sample of the current block, wherein at least one of the left reference sample or the upper reference sample may be partially adjacent to the current block. Meanwhile, if it is determined that MIP is applied to the current block based on information related to MIP, the position of the reference sample of the current block is determined, and the position of the reference sample of the current block may be derived based on information signaled from the bitstream.
[0331] Meanwhile, since FIG. 25 corresponds to one embodiment of the present disclosure, it is obvious that some steps may be changed, the order of some steps may be changed, some steps may be deleted, or some steps may be performed simultaneously, and this is also included in one embodiment of the present disclosure.
[0332] FIG. 26 is a diagram illustrating an image encoding method that can be performed by an image encoding device according to one embodiment of the present disclosure.
[0333] According to an image encoding method that can be performed by an image encoding device according to an embodiment of the present disclosure, whether to apply MIP (matrix-based intra prediction) to a current block can be determined (S2610). Prior to the current step S2610, a process of determining whether a prediction mode (e.g., intra prediction mode or inter prediction mode) is applied to the current block may be performed. Thereafter, a step of determining information related to MIP (S2620) can be performed.
[0334] Meanwhile, as described above, MIP can be applied based on the size of the current block or the number of pixels included in the current block, and more specifically, a matrix for MIP can be determined based on the size of the current block or the number of pixels included in the current block. Meanwhile, when it is determined that MIP is applied to the current block, the value of the prediction block of the current block can be limited to a specific value based on a specific condition, and the specific condition can be based on the MIP on a reference sample value of the current block. For example, when the specific condition is satisfied, application of MIP to the current block can be omitted regardless of whether MIP is applied. Meanwhile, when it is determined that MIP is applied to the current block, the prediction block of the current block can be generated using only the upper reference sample or the left reference sample of the current block. Meanwhile, the upper reference sample can include the upper right reference sample of the current block, and the left reference sample can include the lower left reference sample of the current block. As another example, the prediction block of the current block is generated using a portion of the left reference sample and / or a portion of the upper reference sample of the current block, wherein at least one of the left reference sample or the upper reference sample may be partially adjacent to the current block. Meanwhile, if it is determined that MIP is applied to the current block, the location of the reference sample of the current block is determined, and information about the location of the reference sample of the current block may be encoded and signaled in the bitstream.
[0335] Meanwhile, as an example, a bitstream generated by an image encoding method can be stored in a non-transitory computer-readable medium.
[0336] Additionally, as an example, a method of transmitting a bitstream generated by an image encoding method may be proposed.
[0337] Meanwhile, since FIG. 26 corresponds to one embodiment of the present disclosure, it is obvious that some steps may be changed, the order of some steps may be changed, some steps may be deleted, or some steps may be performed simultaneously, and this is also included in one embodiment of the present disclosure.
[0338] While the exemplary methods of this disclosure are presented as a series of operations for clarity of description, this is not intended to limit the order in which the steps are performed, and individual steps may be performed simultaneously or in different orders, if desired. To implement a method according to this disclosure, additional steps may be included in addition to the steps illustrated, some steps may be excluded and the remaining steps included, or some steps may be excluded and additional steps included.
[0339] In the present disclosure, a video encoding device or video decoding device performing a predetermined operation (step) may perform an operation (step) of checking the conditions or circumstances under which the operation (step) is performed. For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the video encoding device or video decoding device may perform an operation of checking whether the predetermined condition is satisfied and then perform the predetermined operation.
[0340] The various embodiments of the present disclosure are not intended to list all possible combinations but rather to illustrate representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.
[0341] Additionally, various embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, the embodiments may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.
[0342] In addition, the video decoding device and the video encoding device to which the embodiments of the present disclosure are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and may be used to process a video signal or a data signal. For example, the OTT (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), and the like.
[0343] FIG. 27 is a diagram illustrating an example of a content streaming system to which an embodiment according to the present disclosure can be applied.
[0344] As illustrated in FIG. 27, a content streaming system to which an embodiment of the present disclosure is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0345] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted.
[0346] The above bitstream can be generated by an image encoding method and / or an image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0347] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server may control commands / responses between each device within the content streaming system.
[0348] The streaming server can receive content from a media repository and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0349] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.
[0350] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0351] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium having such software or instructions stored thereon and executable on the device or computer.
[0352] Embodiments according to the present disclosure can be used to encode / decode images.
Claims
1. In a video decoding method performed by a video decoding device, a step of determining the prediction mode of the current block; and Based on the above prediction mode being an intra prediction mode, MIP (Matrix-based Intra A step of deriving information related to prediction; including, An image decoding method in which the above MIP is applied based on the size of the current block or the number of pixels included in the current block.
2. In paragraph 1, An image decoding method, wherein the matrix for the above MIP is determined based on the size of the current block or the number of pixels included in the current block.
3. In paragraph 1, A video decoding method, wherein when it is determined that the MIP is applied to the current block based on information related to the MIP, the value of the prediction block of the current block is limited to a specific value based on a specific condition.
4. In paragraph 3, A method for decoding an image, wherein the above specific condition is based on a reference sample value of the current block for the MIP.
5. In paragraph 4, A method for decoding an image, wherein application of the MIP to the current block is omitted regardless of information related to the MIP when the above specific condition is satisfied.
6. In paragraph 1, A method for decoding an image, wherein, when it is determined that the MIP is applied to the current block based on information related to the MIP, a prediction block of the current block is generated using only the upper reference sample of the current block.
7. In paragraph 6, A method for decoding an image, wherein the upper reference sample includes the upper right reference sample of the current block.
8. In paragraph 1, A method for decoding an image, wherein when it is determined that the MIP is applied to the current block based on information related to the MIP, a prediction block of the current block is generated using only the left reference sample of the current block.
9. In paragraph 8, A method for decoding an image, wherein the left reference sample includes a lower left reference sample of the current block.
10. In paragraph 1, The prediction block of the current block is generated using a portion of the left reference sample and the upper reference sample of the current block, A method for decoding an image, wherein at least one of the left reference sample or the upper reference sample is partially adjacent to the current block.
11. In paragraph 1, If it is determined that the MIP is to be applied to the current block based on information related to the MIP, the location of the reference sample of the current block is determined. A method for decoding an image, wherein the location of a reference sample of the current block is derived based on information signaled from a bitstream.
12. In the video encoding method, A step of determining whether to apply MIP (matrix-based intra prediction) to the current block; and A step of determining information related to the above MIP; including: An image encoding method in which the above MIP is applied based on the size of the current block or the number of pixels included in the current block.
13. In a non-transitory computer-readable medium storing a bitstream generated by a video encoding method, The above image encoding method is, A step of determining whether to apply MIP (matrix-based intra prediction) to the current block; and A step of determining information related to the above MIP; including: The above MIP is applied based on the size of the current block or the number of pixels included in the current block.
14. In the method of transmitting a bitstream, A step of transmitting a bitstream generated by a video encoding method; including: The above image encoding method is, A step of determining whether to apply MIP (matrix-based intra prediction) to the current block; and A step of determining information related to the above MIP; including: A method in which the above MIP is applied based on the size of the current block or the number of pixels included in the current block.