Image encoding / decoding method and apparatus, and recording medium storing bitstream
By adaptively selecting matrix kernels based on reference sample characteristics, the method improves the accuracy and reduces complexity of matrix-based intra prediction in image encoding, addressing inefficiencies in existing technologies for high-resolution images.
Patent Information
- Application Number
- PCT/KR2025/099131
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-23
- Filing Date
- 2025-01-22
- Publication Date
- 2025-07-31
AI Technical Summary
Existing image compression technologies face challenges in achieving high efficiency for high-resolution and high-quality images, particularly in accurately predicting pixel values within images using matrix-based intra prediction methods, which can be computationally complex.
The method involves determining reference samples for matrix-based intra prediction by selecting an appropriate matrix kernel or matrix set based on characteristics such as difference, variance, mean value, gradient, or distribution of reference samples, and applying adaptive matrix selection to reduce complexity and improve accuracy.
This approach enhances the accuracy of intra prediction while reducing the computational complexity of matrix-based intra prediction, thereby improving encoding performance for high-resolution images.
Smart Images

Figure KR2025099131_31072025_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device, and recording medium storing bitstream
[0001] The present invention relates to a video encoding / decoding method and device, and a recording medium storing a bitstream.
[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, is increasing in various application fields, and accordingly, high-efficiency image compression technologies are being discussed.
[0003] There are various technologies for image compression, such as inter prediction technology that predicts pixel values included in the current picture from pictures before or after the current picture, intra prediction technology that predicts pixel values included in the current picture using pixel information within the current picture, and entropy encoding technology that assigns short codes to values with high frequency of appearance and long codes to values with low frequency of appearance, and these technologies can be used to effectively compress and transmit or store image data.
[0004] The present disclosure provides a method and apparatus for determining a reference sample for matrix-based intra prediction.
[0005] The present disclosure provides a method and apparatus for determining a matrix kernel or a matrix set for matrix-based intra prediction.
[0006] The video decoding method and device according to the present disclosure can determine reference samples for MIP (matrix-based intra prediction) of a current block, generate a prediction block of the current block based on the reference samples and a matrix kernel, generate a residual block of the current block, and reconstruct the current block based on the prediction block and the residual block.
[0007] In the image decoding method and device according to the present disclosure, the matrix kernel may be determined as one of a plurality of matrix kernels based on the reference samples.
[0008] In the image decoding method and device according to the present disclosure, a matrix set for the current block can be determined from among a plurality of matrix sets based on the reference samples, and any one of a plurality of matrix kernels belonging to the determined matrix set can be determined as the matrix kernel of the current block.
[0009] In the image decoding method and device according to the present disclosure, the matrix set for the current block can be determined based on at least one of a difference, a variance, a mean value, a gradient value, or a distribution of frequency components of the reference samples.
[0010] In the image decoding method and device according to the present disclosure, if the values of the reference samples are the same, a first matrix set may be selected for the current block, and if not, a second matrix may be selected for the current block.
[0011] In the video decoding method and device according to the present disclosure, if at least one of the reference samples has a value of 0 and the value of the remaining reference samples corresponds to a non-zero value, a first matrix set may be selected for the current block, and otherwise, a second matrix may be selected for the current block.
[0012] In the video decoding method and device according to the present disclosure, the reference samples may be divided into one or more sample groups consisting of reference samples having the same value. If the number of the sample groups is T or less, a first matrix set may be selected for the current block, and otherwise, a second matrix may be selected for the current block.
[0013] In the image decoding method and device according to the present disclosure, when the values of the reference samples are the same, the value of at least one of the prediction samples, which is the output of the matrix kernel, can be set to a predetermined value.
[0014] In the image decoding method and device according to the present disclosure, when the value of the reference samples is 0, the value of at least one of the input vectors of the matrix kernel may be set to a value other than 0.
[0015] In the video decoding method and device according to the present disclosure, samples belonging to at least one of a first region or a second region for the current block may be determined as reference samples. Here, the first region may include at least one of an upper peripheral region or an upper-right peripheral region of the current block, and the second region may include at least one of a left peripheral region or a lower-left peripheral region of the current block.
[0016] The video encoding method and device according to the present disclosure can determine reference samples for MIP (matrix-based intra prediction) of a current block, generate a prediction block of the current block based on the reference samples and a matrix kernel, derive transform coefficients of the current block based on a residual block of the current block, and encode residual information regarding the transform coefficients.
[0017] A computer-readable digital storage medium is provided having encoded video / image information stored thereon, which causes a decoding device according to the present disclosure to perform a video decoding method.
[0018] A computer-readable digital storage medium storing video / image information generated by a video encoding method according to the present disclosure is provided.
[0019] A method and device for transmitting video / image information generated by a video encoding method according to the present disclosure are provided.
[0020] According to the present disclosure, the accuracy of intra prediction can be improved while reducing the complexity of matrix-based intra prediction.
[0021] According to the present disclosure, the encoding performance of intra prediction can be improved by adaptively determining a matrix kernel or a matrix set by considering the characteristics of reference samples for matrix-based intra prediction.
[0022] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0023] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0024] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0025] FIG. 4 illustrates a decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.
[0026] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.
[0027] FIG. 6 illustrates an encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.
[0028] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.
[0029] FIG. 8 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0030] The present disclosure may be modified in various ways and encompasses numerous embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.
[0031] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0032] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0033] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0034] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the versatile video coding (VVC) standard. In addition, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of audio video coding standard (AVS2), or the next generation of video / image coding standards (e.g., H.267 or H.268).
[0035] This specification presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0036] In this specification, a video may refer to a set of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs that has a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs that has a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged consecutively according to the CTU raster scan, while tiles within a picture may be arranged consecutively according to the tile raster scan. A slice may contain an integer number of complete tiles or an integer number of contiguous complete CTU rows within a picture, which may be exclusively contained within a single NAL unit. Meanwhile, a picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture.
[0037] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.
[0038] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0039] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."
[0040] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0041] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".
[0042] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”
[0043] Additionally, parentheses used herein may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be suggested as an example of "prediction." Furthermore, even when "prediction (i.e., intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction."
[0044] Technical features individually described in a single drawing in this specification may be implemented individually or simultaneously.
[0045] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0046] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device).
[0047] A source device can transmit encoded video / image information or data to a receiving device via a digital storage medium or a network in the form of a file or streaming. The source device may include a video source, an encoding device, and a transmitting device. The receiving device may include a receiving device, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.
[0048] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.
[0049] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0050] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and an element for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0051] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0052] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.
[0053] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0054] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The addition unit (250) may be called a reconstructor or a recontructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to an embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0055] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.
[0056] For example, a single coding unit may be split into multiple coding units with deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quad-tree structure. The coding procedure according to the present specification may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the largest coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively split into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.
[0057] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be split or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[0058] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).
[0059] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).
[0060] The prediction unit (220) can perform a prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit (220) can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit (220) can generate various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information related to prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0061] The intra prediction unit (222) can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one of a DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of detail in the prediction direction. However, this is only an example, and a greater or lesser number of directional modes may be used depending on the settings. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0062] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference pictures including the temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter prediction unit (221) may construct a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0063] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or a residual signal.
[0064] The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously restored pixels. In addition, the transform process can be applied to a pixel block having a square size and the same size, or can be applied to a block of a non-square variable size.
[0065] The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0066] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) can also encode information necessary for video / image restoration (e.g., values of syntax elements, etc.) together or separately from quantized transform coefficients.
[0067] Encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media, such as a USB, SD, CD, DVD, Blu-ray, HDD, or SSD. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240).
[0068] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as a reconstructed block. The addition unit (250) may be called a reconstructor or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0069] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0070] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device, and can also improve encoding efficiency.
[0071] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store restored samples of restored blocks within the current picture and transfer them to the intra prediction unit (222).
[0072] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0073] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (332) and an intra-prediction unit (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).
[0074] The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., a decoding device chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0075] When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Accordingly, the processing unit for decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output by the decoding device (300) can be reproduced through a reproduction device.
[0076] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this specification can be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of a decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310).
[0077] Meanwhile, a decoding device according to the present specification may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoding device (video / video / picture information decoding device) and a sample decoding device (video / video / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the addition unit (340), the filtering unit (350), the memory (360), the inter prediction unit (332), and the intra prediction unit (331).
[0078] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0079] In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0080] The prediction unit (320) can perform a prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit (320) can determine whether intra-prediction or inter-prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter-prediction mode.
[0081] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index may be included and signaled in the video / image information.
[0082] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The intra prediction unit (331) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0083] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode for the current block.
[0084] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter-prediction unit (332) and / or intra-prediction unit (331)). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the restoration block.
[0085] The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0086] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0087] The (corrected) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information is derived (or decoded) in the current picture and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (332) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (331).
[0088] In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (200) can be applied to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300) in the same or corresponding manner, respectively.
[0089] The matrix-based intra prediction (MIP) method according to the present disclosure can generate a prediction block of a current block based on predetermined reference samples and a matrix kernel. Specifically, a prediction block having the same size as the current block can be generated by applying a matrix kernel to the reference samples. Alternatively, downsampling can be performed on the reference samples, and the matrix kernel can be applied to the downsampled reference samples to derive prediction samples of the current block. A block composed of the derived prediction samples can have a smaller size than the current block. In this case, upsampling can be performed based on the derived prediction samples to generate a prediction block having the same dimension as the current block. Therefore, in the embodiments described below, reference samples to which the matrix kernel is applied (or reference samples that are input to the matrix kernel) can be understood as being replaced with downsampled reference samples. Furthermore, in the embodiments described below, prediction samples output from the matrix kernel can form a block having the same size as the current block, or can form a block having a smaller size than the current block. If the prediction samples output from the matrix kernel form a block with a smaller size than the current block, upsampling can be performed on the prediction samples to generate a prediction block.
[0090] The encoding device and the decoding device may define multiple matrix kernels for the MIP method, and one of the multiple matrix kernels may be selected based on a given MIP mode. The MIP mode may be derived based on mode information signaled through the bitstream. Hereinafter, with reference to FIG. 4, a method for predicting / recovering the current block based on the MIP method will be described.
[0091] FIG. 4 illustrates an image decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.
[0092] Referring to FIG. 4, reference samples for the MIP of the current block can be determined (S400).
[0093] The restored area adjacent to the current block can be divided into an area located at the top of the current block (hereinafter referred to as the first area) and an area located to the left of the current block (hereinafter referred to as the second area).
[0094] The first region according to the present disclosure may include at least one of an upper peripheral region, an upper right peripheral region, or an upper left peripheral region. Here, each peripheral region may be composed of one or more (horizontal) sample lines. The length of a sample line belonging to the upper peripheral region may be equal to the width of the current block. The length of a sample line belonging to the upper right peripheral region may be less than or equal to the width of the current block. Alternatively, the length of a sample line belonging to the upper right peripheral region may be greater than the width of the current block.
[0095] The second region according to the present disclosure may include at least one of a left peripheral region, a lower left peripheral region, or an upper left peripheral region. Here, each peripheral region may be composed of one or more (vertical) sample lines. The length of a sample line belonging to the left peripheral region may be equal to the height of the current block. The length of a sample line belonging to the lower left peripheral region may be less than or equal to the height of the current block. Alternatively, the length of a sample line belonging to the lower left peripheral region may be greater than the height of the current block.
[0096] Since samples in the current block may have a more biased correlation with some surrounding regions, higher prediction accuracy can be achieved by performing predictions based on some surrounding regions. Therefore, we propose a method for partially selecting reference regions or reference samples for MIP in the current block.
[0097] Method 1
[0098] According to the present disclosure, a first region may be determined as a reference region. In this case, one or more samples belonging to the first region may be used for MIP, and samples in the second region may not be used for MIP.
[0099] The first region may be composed of a single sample line adjacent to the current block. Here, the single sample line may belong to the upper peripheral region. The length of the single sample line may be equal to the width (W) of the current block. However, this is not limited thereto, and the first region may also be composed of multiple sample lines belonging to the upper peripheral region.
[0100] Alternatively, the first region may be composed of a plurality of sample lines. Here, the plurality of sample lines may belong to the upper and upper right peripheral regions. The length of the plurality of sample lines may be equal to the width (W) of the current block or may be (2*W). However, this is not limited thereto, and the first region may also be composed of a single sample line belonging to the upper and upper right peripheral regions.
[0101] Alternatively, the first region may be composed of a plurality of sample lines. Here, the plurality of sample lines may belong to the upper, upper right, and upper left peripheral regions. The length of the plurality of sample lines may be greater than (2*W). However, this is not limited thereto, and the first region may also be composed of a single sample line belonging to the upper, upper right, and upper left peripheral regions.
[0102] Method 2
[0103] A second region according to the present disclosure may be determined as a reference region. In this case, one or more samples belonging to the second region may be used for MIP, and samples from the first region may not be used for MIP.
[0104] The second region may be composed of a single sample line adjacent to the current block. Here, the single sample line may belong to the left peripheral region. The length of the single sample line may be equal to the height (H) of the current block. However, this is not limited thereto, and the second region may also be composed of multiple sample lines belonging to the left peripheral region.
[0105] Alternatively, the second region may be composed of a plurality of sample lines. Here, the plurality of sample lines may belong to the left and lower left peripheral regions. The length of the plurality of sample lines may be equal to the height (H) of the current block or may be (2*H). However, this is not limited thereto, and the second region may also be composed of a single sample line belonging to the left and lower left peripheral regions.
[0106] Alternatively, the second region may be composed of a plurality of sample lines. Here, the plurality of sample lines may belong to the left, lower left, and upper left peripheral regions. The length of the plurality of sample lines may be greater than (2*H). However, this is not limited thereto, and the second region may also be composed of a single sample line belonging to the left, lower left, and upper left peripheral regions.
[0107] The method of the present disclosure can share the coefficients of the matrix kernel available in Method 1. That is, when the number of inputs and the number of outputs are the same as in Method 1, the coefficients of the same matrix kernel can be used.
[0108] The method of the present disclosure can share the coefficients of the MIP mode and the corresponding matrix kernel available in Method 1. That is, when the number of inputs and the number of outputs are the same as in Method 1, the coefficients of the same MIP mode and matrix kernel can be used.
[0109] The encoding device and the decoding device may define methods 1 and 2. In this case, it is possible to determine whether method 1 is applied based on a flag. For example, if a first flag indicating whether MIP is applied is true, a second flag indicating whether a method of partially selecting reference samples for MIP is applied may be signaled. If the second flag is true, a third flag indicating either method 1 or method 2 described above may be signaled. The method indicated by the third flag may be used to determine the reference samples.
[0110] Alternatively, it may be implicitly determined whether Method 1 is applied, without signaling a flag indicating whether Method 1 is applied. For example, TIMD template matching may be performed based on reference samples of the first region according to Method 1, and TIMD template matching may be performed based on reference samples of the second region according to Method 2. Through this, a template matching error value may be calculated for each region. At this time, the number of samples of each region may be different, and a process of dividing the error value of each region by the number of samples of each region and comparing them may be added for accurate comparison. If the error value for the first region is smaller than the error value for the second region, it may be determined that Method 1 is applied. On the other hand, if the error value for the first region is greater than the error value for the second region, it may be determined that Method 2 is applied.
[0111] Method 3
[0112] According to the present disclosure, the first and second regions may be determined as reference regions. In this case, some samples belonging to the first region (hereinafter referred to as first reference samples) and some samples belonging to the second region (hereinafter referred to as second reference samples) may be used in MIP.
[0113] The first reference samples may belong to one or more sample lines belonging to the upper peripheral region. Here, the length of the one or more sample lines may be less than or equal to (W / 2). For example, if the current block is an 8x8 block and the coordinate of the upper left sample of the current block is (0, 0), the first reference samples may include a sample having at least one of the coordinates of (4, -1), (5, -1), (6, -1), (7, -1), (4, -2), (5, -2), (6, -2), or (7, -2). The second reference samples may belong to one or more sample lines belonging to the left peripheral region. Here, the length of the one or more sample lines may be less than or equal to (H / 2). For example, if the current block is an 8x8 block and the coordinate of the upper left sample of the current block is (0, 0), the second reference samples may include samples having coordinates of at least one of (-1, 4), (-1, 5), (-1, 6), (-1, 7), (-2, 4), (-2, 5), (-2, 6), or (-2, 7).
[0114] Alternatively, the first reference samples may belong to one or more sample lines belonging to the upper and upper right peripheral regions. Here, the length of the one or more sample lines may be less than or equal to the width (W) of the current block. For example, if the current block is an 8x8 block and the coordinate of the upper left sample of the current block is (0, 0), the first reference samples may include a sample having coordinates of at least one of (4, -1), (5, -1), (6, -1), (7, -1), (8, -1), (9, -1), (10, -1), or (11, -1). The second reference samples may belong to one or more sample lines belonging to the left and lower left peripheral regions. Here, the length of the one or more sample lines may be less than or equal to the height (H) of the current block. For example, if the current block is an 8x8 block and the coordinate of the upper left sample of the current block is (0, 0), the second reference samples may include samples having coordinates of at least one of (-1, 4), (-1, 5), (-1, 6), (-1, 7), (-1, 8), (-1, 9), (-1, 10), or (-1, 11).
[0115] Alternatively, the first reference samples may belong to one or more sample lines belonging to the upper and upper left peripheral areas. Here, the length of the one or more sample lines may be less than or equal to the width (W) or (W / 2) of the current block. For example, if the current block is an 8x8 block and the coordinate of the upper left sample of the current block is (0, 0), the first reference samples may include a sample having at least one of the coordinates of (-2, -1), (-1, -1), (0, -1), (1, -1), (2, -1), (3, -1), (4, -1), (-2, -2), (-1, -2), (0, -2), (1, -2), (2, -2), (3, -2), or (4, -2). The second reference samples may belong to one or more sample lines belonging to the left peripheral area. Here, the length of the one or more sample lines may be less than or equal to (W / 2). For example, if the current block is an 8x8 block and the coordinate of the upper left sample of the current block is (0, 0), the second reference samples may include samples having coordinates of at least one of (-1, 0), (-1, 1), (-1, 2), (-1, 3), (-2, 0), (-2, 1), (-2, 2), or (-2, 3).
[0116] Alternatively, the first reference samples may belong to one or more sample lines belonging to the upper peripheral region. When the width (W) of the current block is greater than the height (H), the length of one or more sample lines may be less than or equal to (W / 2). When the width (W) of the current block is less than the height (H), the length of one or more sample lines may be less than or equal to W. For example, when the current block is an 8x4 block and the coordinate of the upper left sample of the current block is (0, 0), the first reference samples may include a sample having at least one of the coordinates (4, -1), (5, -1), (6, -1), or (7, -1). The second reference samples may belong to one or more sample lines belonging to the left peripheral region. When the height (H) of the current block is greater than the width (W), the length of one or more sample lines may be less than or equal to (H / 2). If the height (H) of the current block is less than the width (W), the length of one or more sample lines may be less than or equal to H. For example, if the current block is an 8x4 block and the coordinate of the upper left sample of the current block is (0, 0), the second reference samples may include samples having coordinates of at least one of (-1, 0), (-1, 1), (-1, 2), or (-1, 3).
[0117] The method of the present disclosure can share the coefficients of the MIP mode and the corresponding matrix kernel available in Method 1 or 2. That is, when the number of inputs and the number of outputs are the same as in Method 1 or 2, the coefficients of the same MIP mode and matrix kernel used in Method 1 or 2 can be used.
[0118] The encoding device and the decoding device may have methods 1 to 3 defined. In this case, it is possible to determine which method among methods 1 to 3 is applied based on a 2-bit index. For example, if a first flag indicating whether MIP is applied is true, a second flag indicating whether a method of partially selecting reference samples for MIP is applied may be signaled. If the second flag is true, an index indicating any one of the aforementioned methods 1 to 3 may be signaled. The method indicated by the index can be used to determine the reference samples.
[0119] Alternatively, it is possible to implicitly determine which of methods 1 to 3 is applied without signaling the index. For example, TIMD template matching can be performed based on reference samples of the first region according to method 1, TIMD template matching can be performed based on reference samples of the second region according to method 2, and TIMD template matching can be performed based on reference samples of the first and second regions according to method 3. Through this, a template matching error value can be calculated for each region. At this time, the number of samples in each region may be different, and a process of dividing the error value of each region by the number of samples in each region and comparing them for accurate comparison can be added. It can be determined that the method with the smallest error value among the calculated error values is applied.
[0120] All samples belonging to the determined reference region may be used as reference samples for MIP. Alternatively, some samples belonging to the determined reference region may be used as reference samples for MIP. Here, some samples may be specified through subsampling of samples belonging to the corresponding reference region. Alternatively, some samples may be derived through downsampling of samples belonging to the corresponding reference region. Alternatively, some samples may be selected from among samples belonging to the reference region based on a predetermined threshold. For example, some samples may be defined as samples belonging to the reference region that are greater than or equal to a predetermined threshold. Alternatively, some samples may be defined as samples belonging to the reference region that are less than or equal to a predetermined threshold.
[0121] Referring to FIG. 4, a prediction block of the current block can be generated based on reference samples and a matrix kernel (S410).
[0122] Reference samples may be input into a neural network, and a prediction block of the current block may be output from the neural network. In the present disclosure, a neural network composed of a matrix kernel may be used.
[0123]
[0124] In mathematical expression 1, r denotes the values of reference samples selected through the aforementioned method, and pred may denote the values of prediction samples of the current block. A denotes a matrix kernel, and the Clip function may denote a function that corrects the output pred value to a value within the sample value range. That is, the output (pred) can be obtained by performing a matrix operation on the input (r). The size of the matrix kernel for the matrix operation can be derived based on the number of reference samples input to the matrix kernel (i.e., the number of inputs) and the number of prediction samples output from the matrix kernel (i.e., the number of outputs).
[0125] The neural network configuration for obtaining the output pred for the input r is not limited to the examples described above. In addition to a neural network using a single matrix kernel, multiple matrix operations can be performed through multiple neural network layers. At least one of a bias term or an activation function can be used for each neural network layer. The pred value may be output as a frequency domain value rather than a spatial domain value. In this case, a reverse transformation can be performed on the pred value to obtain the final prediction sample.
[0126] The present disclosure proposes a method for adaptively selecting matrix kernels for MIP.
[0127] A matrix kernel can be adaptively selected based on the size of the current block. Here, the size can be defined as the width, height, the product of width and height (i.e., the number of samples belonging to the current block), the sum of width and height, the minimum / minimum of width and height, or the ratio of width and height.
[0128] If the size of the current block is less than or equal to a specific size, a first matrix kernel that does not involve downsampling and upsampling for the input (r) and the output (pred) may be selected. Otherwise, a second matrix kernel that involves downsampling and upsampling for the input (r) and the output (pred) may be selected. That is, if the size of the current block is less than or equal to a specific size, a first matrix kernel may be selected from among a plurality of matrix kernels. In this case, non-downsampled reference samples may be input to the first matrix kernel, and upsampling may not be performed on prediction samples, which are outputs of the first matrix kernel. On the other hand, if the size of the current block is greater than a specific size, a second matrix kernel may be selected from among a plurality of matrix kernels. In this case, downsampled reference samples may be input to the second matrix kernel, and upsampling may be performed on prediction samples, which are outputs of the second matrix kernel.
[0129] Alternatively, if the size of the current block is greater than or equal to a specific size, a first matrix kernel that does not involve downsampling and upsampling for the input (r) and the output (pred) may be selected. Otherwise, a second matrix kernel that involves downsampling and upsampling for the input (r) and the output (pred) may be selected. That is, if the size of the current block is greater than or equal to a specific size, a first matrix kernel may be selected from among a plurality of matrix kernels. In this case, non-downsampled reference samples may be input to the first matrix kernel, and upsampling may not be performed on prediction samples, which are outputs of the first matrix kernel. On the other hand, if the size of the current block is less than a specific size, a second matrix kernel may be selected from among a plurality of matrix kernels. In this case, downsampled reference samples may be input to the second matrix kernel, and upsampling may be performed on prediction samples, which are outputs of the second matrix kernel.
[0130] Alternatively, if the size of the current block is less than or equal to a specific size, a first matrix kernel may be selected from among the plurality of matrix kernels, otherwise, a second matrix kernel may be selected from among the plurality of matrix kernels. Here, the first matrix kernel may be a more sophisticated matrix kernel than the second matrix kernel. For example, the number of inputs to the first matrix kernel may be greater than the number of inputs to the second matrix kernel. Alternatively, the number of outputs of the first matrix kernel may be greater than the number of outputs of the second matrix kernel. Alternatively, the number of inputs and outputs of the first matrix kernel may be greater than the number of inputs and outputs of the second matrix kernel, respectively.
[0131] Alternatively, if the size of the current block is greater than or equal to a specific size, a first matrix kernel among the plurality of matrix kernels may be selected, and otherwise, a second matrix kernel among the plurality of matrix kernels may be selected. Here, the first matrix kernel may be a more sophisticated matrix kernel than the second matrix kernel. For example, the number of inputs and outputs of the first matrix kernel may be greater than the number of inputs and outputs of the second matrix kernel.
[0132] Whether to apply downsampling (or upsampling) may be determined based on whether the width and / or height of the current block is greater than or less than a predetermined threshold. A downsampling ratio / factor (or upsampling ratio / factor) may be determined based on whether the width and / or height of the current block is greater than or less than a predetermined threshold.
[0133] For example, if the size of the current block is less than or equal to 8x8, the distribution change of the sample values adjacent to the current block may be relatively larger than that of a block larger than 8x8. In this case, a matrix kernel that does not involve downsampling and upsampling for the input (r) and output (pred) may be selected. That is, for 4x4, 4x8, 8x4, and 8x8 blocks, non-downsampled reference samples may be input to the matrix kernel, and upsampling may not be performed on the prediction samples, which are the output of the matrix kernel. For this purpose, matrix kernels for 4x4, 4x8, 8x4, and 8x8 blocks may be defined separately.
[0134] If the current block is an 8x8 block and the first and second regions are used as reference regions, the number of inputs to the matrix kernel can be 16, and the number of outputs to the matrix kernel can be 64. For this purpose, a matrix kernel of size 16x64 can be defined.
[0135] Alternatively, if the current block is an 8x8 block and the first region is used as a reference region, the number of inputs to the matrix kernel may be 8 and the number of outputs to the matrix kernel may be 64. For this purpose, a matrix kernel of size 8x64 may be defined.
[0136] If the current block size is not 4x4, 4x8, 8x4, or 8x8, the number of inputs and outputs of the matrix kernel can be set to the same as before. In this case, a newly derived matrix kernel can be used, or the existing matrix kernel can be applied in the same way.
[0137] If the number of samples belonging to the current block is less than or equal to a specific number, a first matrix kernel that does not involve downsampling and upsampling for the input (r) and the output (pred) may be selected. Otherwise, a second matrix kernel that involves downsampling and upsampling for the input (r) and the output (pred) may be selected. That is, if the number of samples belonging to the current block is less than or equal to a specific number, a first matrix kernel may be selected from among a plurality of matrix kernels. In this case, non-downsampled reference samples may be input to the first matrix kernel, and upsampling may not be performed on prediction samples which are outputs of the first matrix kernel. On the other hand, if the number of samples belonging to the current block is greater than a specific number, a second matrix kernel may be selected from among a plurality of matrix kernels. In this case, downsampled reference samples may be input to the second matrix kernel, and upsampling may be performed on prediction samples which are outputs of the second matrix kernel.
[0138] Alternatively, if the number of samples belonging to the current block is greater than or equal to a specific number, a first matrix kernel that does not involve downsampling and upsampling for the input (r) and the output (pred) may be selected. Otherwise, a second matrix kernel that involves downsampling and upsampling for the input (r) and the output (pred) may be selected. That is, if the number of samples belonging to the current block is greater than or equal to a specific number, a first matrix kernel may be selected from among a plurality of matrix kernels. In this case, non-downsampled reference samples may be input to the first matrix kernel, and upsampling may not be performed on prediction samples which are outputs of the first matrix kernel. On the other hand, if the number of samples belonging to the current block is less than a specific number, a second matrix kernel may be selected from among a plurality of matrix kernels. In this case, downsampled reference samples may be input to the second matrix kernel, and upsampling may be performed on prediction samples which are outputs of the second matrix kernel.
[0139] Alternatively, if the number of samples belonging to the current block is less than or equal to a specific number, a first matrix kernel may be selected from among the plurality of matrix kernels, and otherwise, a second matrix kernel may be selected from among the plurality of matrix kernels. Here, the first matrix kernel may be a more sophisticated matrix kernel than the second matrix kernel. For example, the number of inputs and outputs of the first matrix kernel may be greater than the number of inputs and outputs of the second matrix kernel.
[0140] Alternatively, if the number of samples belonging to the current block is greater than or equal to a specific number, a first matrix kernel may be selected from among the plurality of matrix kernels, and otherwise, a second matrix kernel may be selected from among the plurality of matrix kernels. Here, the first matrix kernel may be a more sophisticated matrix kernel than the second matrix kernel. For example, the number of inputs and outputs of the first matrix kernel may be greater than the number of inputs and outputs of the second matrix kernel.
[0141] For example, if the number of samples belonging to the current block is less than or equal to 128, a matrix kernel that does not involve downsampling and upsampling for the input (r) and output (pred) can be selected. That is, for 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 16x4, 16x8, and 32x4 blocks, non-downsampled reference samples can be input to the matrix kernel, and upsampling may not be performed on the prediction samples which are the output of the matrix kernel. For this purpose, matrix kernels for 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 16x4, 16x8, and 32x4 blocks can be defined separately.
[0142] If the current block is an 8x16 block and the first region and the second region are used as reference regions, the number of inputs to the matrix kernel may be 16 and the number of outputs to the matrix kernel may be 128. For this purpose, a matrix kernel of size 16x128 may be defined. Alternatively, if the current block is an 8x16 block and the second region is used as a reference region, the number of inputs to the matrix kernel may be 16 and the number of outputs to the matrix kernel may be 128. For this purpose, a matrix kernel of size 16x128 may be defined.
[0143] If the current block size is not 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 16x4, 16x8, and 32x4, the number of inputs and outputs of the matrix kernel can be set to the same as before. In this case, a newly derived matrix kernel can be used, or the existing matrix kernel can be applied in the same way.
[0144] By applying the aforementioned method, MIP performance can be improved in small blocks. Furthermore, since the coefficients (or weights) of the required matrix kernel increase exponentially as the block size increases, the aforementioned method can effectively suppress the increase in the coefficients of the matrix kernel.
[0145] The present disclosure proposes a method for performing MIP when a reference sample input to a matrix kernel has a specific value.
[0146] When the reference samples input to the matrix kernel have specific values, MIP can be performed adaptively as follows.
[0147] For example, if the values of the reference samples for MIP are all the same or if the values of the reference samples for MIP are all 0, the result of the matrix operation on the input can always be 0. In this case, the matrix operation can be omitted and the MIP for the current block can be omitted. Instead, the value of at least one of the prediction samples, which is the output of the matrix kernel, can be set to a predetermined value (k). Here, the value of k can be adaptively derived according to the MIP mode of the current block. For example, if the values of the reference samples for MIP are all 0, the predetermined value (k) can be derived as the output of the following mathematical expression 2.
[0148]
[0149] In mathematical expression 2, mode_num refers to the number of the MIP mode of the current block, and max_mode_num may refer to the maximum number of MIP modes available to the current block. In mathematical expression 2, when the bitdepth is 10, max_mode_num is 6, and mode_num is 3, the output of the matrix kernel may be 512. However, this is only an example, and the value of k may be a fixed value that is identically pre-defined for the encoding device and the decoding device, or may be derived based on the bitdepth of the image.
[0150] Alternatively, if the values of the reference samples for MIP are all the same or if the values of the reference samples for MIP are all 0, the value of at least one of the input vectors of the matrix kernel can be set to a value other than 0. For example, MIP can be performed by setting the values of the reference samples as in the following mathematical expression 3.
[0151]
[0152] In mathematical expression 3, input[x] may denote the value of the xth input vector. The N reference samples may be expressed as a reference sample array consisting of the 0th reference sample to the (N-1)th reference sample. Input vectors may be derived based on the reference samples constituting the reference sample array, respectively. At this time, input[x] may be an input vector corresponding to the xth reference sample. If the values of all reference samples are 0, the value of the 0th input vector (input[0]) among the input vectors of the matrix kernel may be set to the value of (A-ref[0]). ref[0] may denote the value of the 0th reference sample among the reference samples for MIP. A may be a value predefined identically in the encoding device and the decoding device. Alternatively, A may be set to a value of (1<<(bitdepth-1)). Alternatively, A may be adaptively derived based on the MIP mode of the current block.
[0153] In the present disclosure, the existing matrix kernel used in MIP may be used in the same manner, or another new matrix kernel that conforms to the present disclosure may be used.
[0154] Alternatively, if the values of the reference samples for MIP are all the same or if the values of the reference samples for MIP are all 0, the outputs of the matrix kernels can all be set to a specific value (k). In this case, the signaling of mode information for deriving the MIP mode can be omitted. The MIP mode of the current block can be derived as a pre-defined default mode. The specific value (k) can be derived as an average of the reference samples. Alternatively, the specific value (k) can be adaptively derived based on the MIP mode of the current block.
[0155] The present disclosure proposes a method for adaptively selecting a matrix kernel or a set of matrices by analyzing reference samples for MIP.
[0156] A single matrix set may be defined for the encoding device and the decoding device. Alternatively, multiple matrix sets may be defined to perform more sophisticated MIP depending on the characteristics of the image. Each matrix set may include one or more matrix kernels.
[0157] A matrix kernel or set of matrices for the current block can be adaptively determined based on the reference samples for MIP (or the reference samples input to the matrix kernel).
[0158] For example, a matrix kernel for a current block can be adaptively determined by analyzing the values of reference samples. Here, the analysis of the values of the reference samples can mean a difference, variance, or covariance between the reference samples, a histogram of gradient (HoG) analysis used in decoder-side intra mode derivation (DIMD), high-frequency / low-frequency component analysis, or an analysis of the average value of the reference samples. The determined matrix kernel can be any one of a plurality of matrix kernels belonging to a matrix set.
[0159] For example, a matrix set for a current block can be adaptively determined by analyzing the values of reference samples. Here, the analysis of the values of the reference samples can mean at least one of a difference between reference samples, a variance, a covariance, a histogram of gradient (HoG) analysis used in decoder-side intra mode derivation (DIMD), a high / low frequency component analysis, or an average value analysis of the reference samples. The determined matrix set can be any one of a plurality of matrix sets that are identically pre-defined for the encoding device and the decoding device. The determined matrix set can include a plurality of matrix kernels. MIP of the current block can be performed based on any one of the plurality of matrix kernels. Mode information specifying any one of the plurality of matrix kernels can be signaled.
[0160] For example, if the reference samples are concentrated on a small number of specific values, the image to which the current block belongs is likely to be a TGM (text and graphics with motion) image. In this case, a matrix kernel or matrix set optimized for TGM images can be selected.
[0161] Alternatively, different matrix kernels or matrix sets may be selected for cases where the reference samples have relatively more high-frequency components and cases where the reference samples have relatively more low-frequency components. In this way, the matrix kernels or matrix sets may be selected based on the distribution of frequency components for the reference samples.
[0162] Alternatively, different matrix kernels or sets of matrices may be selected for cases where the variance of the reference samples is relatively large and cases where the variance of the reference samples is relatively small.
[0163] Alternatively, as in DIMD, a gradient value may be calculated based on the aforementioned reference region (or reference samples), and a matrix kernel or a matrix set may be selected based on the calculated gradient value. The gradient value may be calculated for each of the pre-defined intra prediction modes, which may be referred to as a histogram of gradient (HoG). A matrix kernel or a matrix set may be selected based on the largest gradient value in the HoG or the intra prediction mode corresponding thereto.
[0164] The present disclosure proposes a method for adaptively selecting a set of matrices or matrix kernels based on whether reference samples satisfy a given condition.
[0165] For example, if the reference samples satisfy a predetermined condition, a first matrix set (or a first matrix kernel) may be selected for the current block, and if not, a second matrix set (or a second matrix kernel) may be selected for the current block. Here, the first matrix set (or the first matrix kernel) may be separately defined to apply MIP when the reference samples satisfy the predetermined condition. The first matrix set may include a different number of matrix kernels than the second matrix set. Alternatively, the first matrix set may include the same number of matrix kernels as the second matrix set, but at least one matrix kernel belonging to the first matrix set may have different coefficients than the matrix kernels belonging to the second matrix set. At least one of the size or coefficients of the first matrix kernel may be different from the second matrix kernel. The same meaning can be interpreted in the embodiments described below.
[0166] If the values of the reference samples are all the same, the first matrix set (or the first matrix kernel) may be selected for the current block, otherwise, the second matrix set (or the second matrix kernel) may be selected for the current block.
[0167] Alternatively, if all the values of the reference samples are 0 (or if all the input vectors are 0), the first matrix set (or the first matrix kernel) may be selected for the current block, otherwise, the second matrix set (or the second matrix kernel) may be selected for the current block.
[0168] Alternatively, if at least one of the reference samples has a value of 0 and the values of the remaining reference samples correspond to a specific value, the first matrix set (or the first matrix kernel) may be selected for the current block, otherwise, the second matrix set (or the second matrix kernel) may be selected for the current block.
[0169] Alternatively, if the values of all input vectors except the first input vector are 0, the first matrix set (or the first matrix kernel) may be selected for the current block, otherwise, the second matrix set (or the second matrix kernel) may be selected for the current block.
[0170] Alternatively, if the values of the reference samples have only a specific prime number of values, the first matrix set (or the first matrix kernel) may be selected for the current block, otherwise, the second matrix set (or the second matrix kernel) may be selected for the current block. The reference samples may be divided into one or more sample groups consisting of reference sample(s) having the same values. In this case, if the number of sample groups is T or less, it may be determined that the values of the reference samples have only a specific prime number of values. Here, T may be an integer of 1, 2, or a larger number.
[0171] For example, if the values of the reference samples are all the same, MIP can be performed as in the following mathematical expression 4.
[0172]
[0173] In Equation 4, r' can be a value derived based on the value of a 1x1 reference sample. Since the values of all reference samples are the same, the input of the matrix kernel can be derived based on a single value. A' can mean an Nx1 matrix kernel when the number of MIP outputs (i.e., the number of prediction samples) is N.
[0174] In Equation 4, the size of A' can be adaptively determined based on the size of r'. The value of r' can be derived to be the same value as the value of the reference sample. Alternatively, the value of r' can be derived as in Equation 5 below.
[0175]
[0176] In mathematical expression 5, bitdepth may mean the bit depth of the reference sample, and input may mean the value of the reference sample or the value of the input vector derived based on the value of the reference sample.
[0177] The above-described method for selecting an adaptive matrix kernel or matrix set is merely an example, and may be selected based on encoding information of a current block and / or neighboring blocks. Here, the encoding information may include at least one of a prediction mode (e.g., intra mode, inter mode), an intra prediction mode, a width, a height, a number of samples in a block, a position of a sub-block within a block, explicitly signaled syntax elements (e.g., mode information), statistical characteristics of samples within a block, or whether a secondary transform is applied.
[0178] Information specifying one of multiple matrix sets can be explicitly signaled. Based on the signaled information, the matrix set for the current block can be determined.
[0179] Information about a matrix kernel suitable for the corresponding image can be signaled in a high-level syntax (HLS) such as VPS, SPS, PPS, Picture Header, Slice Header, or DCI. Here, the information about the matrix kernel can include at least one of information about coefficients (or weights) of the matrix kernel, information about a matrix set, information about the number of available matrix kernels, or information about the number of available matrix sets. Based on the above-described information, a matrix kernel or a matrix set for the current block can be adaptively determined.
[0180] The size of the matrix kernel can be determined based on the number of inputs and the number of outputs of the matrix kernel. The reference samples determined through the above-described method can be input to the matrix kernel, or downsampled reference samples can be input to the matrix kernel. The downsampling can be performed based on a predefined downsampling factor (e.g., 2:1, 4:1, or 8:1). Since the number of inputs or the input size of the matrix kernel can be reduced through downsampling, the size of the matrix kernel can be reduced, and the overall amount of MIP computation can be greatly reduced. In the case of MIP, the larger the block, the larger the size of the matrix kernel and the size of the input vector (r), which causes great computational complexity to the encoding device and the decoding device. Therefore, it is very important to reduce the sizes of the input vector (r) and the matrix kernel.
[0181] The above downsampling method and ratio are not limited to the examples described above, and the number of inputs to the matrix kernel may also be increased through upsampling of reference samples. The sampling method and ratio for the reference samples may be determined based on at least one of the shape / size of the block or the number of reference samples.
[0182] Filtering (e.g., low-pass filtering, high-pass filtering) can be applied to the reference samples, and the filtered reference samples can be used as inputs to the matrix kernel. Alternatively, a transform can be applied to the reference samples to derive transform coefficients in the frequency domain, and the derived transform coefficients can be used as inputs to the matrix kernel. Alternatively, the average value of the reference samples can be calculated, and the values of the residual samples, which are the differences between the values of the reference samples and the average value, can be used as inputs to the matrix kernel. Alternatively, a transform can be applied to the residual samples to derive transform coefficients in the frequency domain, and the derived transform coefficients can be used as inputs to the matrix kernel.
[0183] The above-described method can be applied to the output of the matrix kernel in the same / similar manner. That is, the number of prediction samples may be less or more than the number of samples belonging to the current block, in which case upsampling or downsampling may be performed on the prediction samples so that they have the same size as the current block. Alternatively, if the transform coefficients in the frequency domain are used as input, the prediction samples can be derived based on an inverse transform on the output coefficients. Alternatively, if the residual samples are used as input, the prediction samples can be derived based on the output samples. Alternatively, if the transform coefficients for the residual samples are used as input, the prediction samples can be derived based on an inverse transform on the output coefficients.
[0184] Post-processing filtering may be applied to the prediction block according to the present disclosure. For example, position dependent prediction combination (PDPC) may be applied to the prediction block. Each sample constituting the prediction block (i.e., the prediction samples) may be corrected based on a weighted sum with at least one neighboring sample adjacent to the current block. Alternatively, a smoothing filter may be applied to the prediction block.
[0185] In the neural network according to the present disclosure, the coefficients of the matrix kernel used in matrix operations may have various bit precisions. For example, when performing a matrix operation with 10-bit precision, the coefficients of the matrix kernel may have coefficient value ranges of 0 to 1023 or -512 to 511, and input and / or output values may be adjusted accordingly. The bit precision may be adaptively determined based on the number of layers constituting the neural network, etc.
[0186] The coefficients of the matrix kernel according to the present disclosure may be defined to suit all block shapes. For example, for blocks ranging from 4x4 blocks to 256x256 blocks, coefficients of the matrix kernel corresponding to each block shape may be defined.
[0187] Depending on the block shape, the number of inputs and outputs of the matrix kernel may be defined. For example, if MIP is performed based on a single upper sample line, 49 kernel types may be required for inputs of 4, 8, 16, 32, 64, 128, and 256 and outputs of 4, 8, 16, 32, 64, 128, and 256. Multiple MIP modes may be defined for each of the 49 kernel types. If the number of inputs is 8, the number of outputs is 16, and the number of MIP modes is 35, the matrix kernel can have a total of (8x16x35) coefficients.
[0188] Alternatively, only matrix kernels that fit a specific block shape may be defined, which may have a limited number of inputs and outputs. In this case, matrix kernels for a specific block shape may be derived through preprocessing or postprocessing of the inputs and / or outputs. For example, if only matrix kernels that fit square blocks are defined, the number of inputs of the matrix kernels for square blocks may be adjusted through preprocessing (e.g., downsampling or upsampling) of the reference samples of non-square blocks. In addition, the number of samples of non-square blocks may be adjusted through postprocessing (e.g., downsampling or upsampling) of the output values of the matrix kernels.
[0189] A MIP mode according to the present disclosure can be derived based on signaled mode information. Here, the mode information can specify any one of a plurality of MIP modes that are predefined identically for an encoding device and a decoding device. Alternatively, a candidate list including at least two MIP mode candidates can be constructed for a current block. The at least two MIP mode candidates can be selected from the plurality of predefined MIP modes. In this case, the mode information can specify any one of the at least two MIP mode candidates belonging to the candidate list. The MIP mode candidates of the current block can be derived based on the MIP modes of neighboring blocks adjacent to the current block. Alternatively, the MIP mode candidates can be derived in a manner such as the DIMD method or the TIMD method described below.
[0190] The MIP mode according to the present disclosure can be derived in the same manner as the decoder-side intra-mode derivation (DIMD) method. Specifically, gradient values can be derived for pre-defined MIP modes based on a peripheral area of the current block (e.g., an upper peripheral area and / or a left peripheral area), and one or more MIP modes can be derived for the current block based on the derived gradient values. The top N MIP modes can be selected in descending order of the derived gradient values. Here, N can be an integer of 1, 2, or more. For example, the MIP mode with the largest gradient value can be set as the MIP mode of the current block. Alternatively, if two or more MIP modes are derived for the current block, the corresponding MIP modes can be used as the aforementioned MIP mode candidates.
[0191] The MIP mode according to the present disclosure can be derived in a manner similar to the template-based intra mode derivation (TIMD) method. Specifically, a cost for each of the predetermined MIP modes can be calculated. Here, the predetermined MIP modes may refer to the pre-defined MIP modes or may refer to MIP mode candidates belonging to a candidate list. One or more MIP modes can be selected based on the calculated cost. The MIP mode of the current block can be derived based on the selected one or more MIP modes. For example, the MIP mode with the lowest cost can be set as the MIP mode of the current block. Alternatively, if two or more MIP modes are derived for the current block, the corresponding MIP modes can be used as the aforementioned MIP mode candidates.
[0192] The MIP mode of the current block may be selected based on encoding information of the current block and / or surrounding blocks. The encoding information may include at least one of a prediction mode (e.g., intra mode, inter mode), a width, a height, a number of samples in the block, the position of a sub-block within the block, explicitly signaled syntax elements (e.g., mode information), statistical characteristics of samples within the block, or whether a secondary transform is applied.
[0193] The above mode information can be binarized using an appropriate binarization method. Context modeling can be applied to the bins of the mode information. Considering the total number of MIP modes, the mode information can be binarized using methods such as truncated binary, truncated unary, or fixed-length.
[0194] The number of MIP modes available to the current block may be selected based on encoding information of the current block and / or neighboring blocks. The encoding information may include at least one of a prediction mode (e.g., intra mode, inter mode), a width, a height, a number of samples in the block, a position of a sub-block within the block, explicitly signaled syntax elements, statistical characteristics of samples within the block, or whether a secondary transform is applied. For example, the number of MIP modes available to the current block may be adaptively determined based on the size of the current block. Alternatively, the number of available MIP modes may be the same for all block sizes.
[0195] Referring to Fig. 4, a residual block of the current block can be generated (S420).
[0196] Residual information of a current block can be obtained from a bitstream. Transform coefficients can be derived based on the residual information. A residual block of the current block can be generated based on at least one of a non-separable transform and a separable transform for the transform coefficients. The non-separable transform can represent a low frequency non-separable transform (LFNST) and / or a non-separable primary transform (NSPT).
[0197] The residual block of the current block can be generated based on a primary transform and / or a secondary transform for the above transform coefficients. The primary transform may correspond to a non-separable transform or a separable transform, and the secondary transform may correspond to a non-separable transform.
[0198] The transform type for the first transform can be determined based on the encoding information of the current block and / or surrounding blocks described above. The transform kernel for the second transform can be determined based on the encoding information of the current block and / or surrounding blocks. Here, the encoding information can include at least one of a width, a height, a number of samples in a block, a position of a sub-block within a block, an explicitly signaled syntax element, or a statistical characteristic of samples within a block.
[0199] Referring to FIG. 4, the current block can be restored based on the prediction block and the residual block (S430).
[0200] Whether or not the proposed method according to the present disclosure is applicable or applicable can be signaled in HLS (high-level syntax), such as VPS, SPS, PPS, Picture Header, Slice Header, DCI, etc. For example, in order to determine whether or not the proposed method is applicable or applicable on a PPS basis and to quickly perform an initialization process for a MIP process (e.g., MIP matrix loading and buffer management required for MIP) to increase the efficiency of the codec system, information about whether or not the proposed method is applicable or applicable can be signaled on a PPS basis. In addition, MIP can be adaptively performed in blocks of a specific block size or a specific number of samples or more, and for this purpose, information about the specific block size or the number of samples can be signaled in HLS.
[0201] Whether the proposed method according to the present disclosure applies can be determined without signaling additional information to the decoding device. Alternatively, additional information regarding whether the proposed method applies can be signaled. For example, a 1-bit flag indicating whether the proposed method applies can be signaled on a CTU or CU basis.
[0202] The proposed method according to the present disclosure may be used only when the MIP method or the regular MIP method is defined as applicable in HLS. Furthermore, the applicability of the proposed method may be determined through signaling of additional information within the MIP method.
[0203] If the proposed method is determined to be applicable, a 1-bit flag indicating whether the proposed method is applicable may be signaled. Here, whether the proposed method is applicable may be determined based on whether a specific block size / shape or specific conditions are satisfied. For example, if the height of the current block is more than four times the width of the current block, the proposed method may be determined to be inapplicable to the current block, and thus, additional information indicating whether the proposed method is applicable to the current block may not be signaled. Alternatively, the proposed method may be applied only when the width and height of the current block are less than or equal to 16.
[0204] Whether a proposed method is applied to the current block may be implicitly determined based on whether the current block satisfies certain conditions.
[0205] For example, if the left sample of the current block is not available (or the left sample does not exist), the signaling of the flag indicating whether the proposed method is applied may be omitted. In this case, it may be determined that the proposed method of the present disclosure is applied. The case where the left sample is not available may include the case where the left sample of the current block belongs to a different coding tree unit (CTU), tile, slice, or subpicture than the current block.
[0206] Information regarding whether the proposed method is allowed / applied can be signaled at higher levels, such as VPS, SPS, PPS, picture header, or slice header. Based on this information, information regarding whether the proposed method is allowed / applied can be adaptively signaled at the coding unit level. For example, if the information regarding whether the proposed method is allowed / applied in the SPS is false (i.e., the proposed method is not allowed / applied at the SPS level), information regarding whether the proposed method is allowed / applied is not signaled at the coding unit level, and the proposed method may not be applied to the corresponding coding unit.
[0207] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.
[0208] Referring to FIG. 5, the decoding device (300) may include a reference sample determination unit (500), a prediction block generation unit (510), a residual block generation unit (520), and a restoration unit (530). The reference sample determination unit (500) and the prediction block generation unit (510) may be provided in the intra prediction unit (331) of FIG. 3, and the residual block generation unit (520) may be provided in the residual processing unit (320) of FIG. 3.
[0209] The reference sample determination unit (500) can determine reference samples for the MIP of the current block. The reference sample determination method is as described with reference to FIG. 4, and any duplicate description will be omitted here.
[0210] The prediction block generation unit (510) can generate a prediction block of the current block based on reference samples and a matrix kernel, and the prediction block generation method is as described with reference to FIG. 4.
[0211] The residual block generation unit (520) can generate a residual block of the current block, and the residual block generation method is as described with reference to FIG. 4.
[0212] The restoration unit (530) can restore the current block based on the prediction block and residual block of the current block.
[0213] FIG. 6 illustrates an encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.
[0214] Referring to FIG. 6, reference samples for the MIP of the current block can be determined (S600). The reference sample determination method can be determined based on at least one of the aforementioned methods 1 to 3, and any redundant descriptions will be omitted herein.
[0215] Referring to Fig. 6, a prediction block of the current block can be generated based on reference samples and a matrix kernel (S610). The prediction block of the current block can be generated by applying the same method discussed with reference to Fig. 4, and any redundant description will be omitted here.
[0216] Referring to FIG. 6, the transformation coefficients of the current block can be derived based on the residual block of the current block (S620).
[0217] Specifically, a residual block of the current block can be generated based on a prediction block of the current block. Transform coefficients of the current block can be derived based on at least one of a non-separable transform or a separable transform for the residual block. The non-separable transform can represent a low frequency non-separable transform (LFNST) and / or a non-separable primary transform (NSPT).
[0218] Transform coefficients can be derived based on the primary transform and / or secondary transform of the residual block of the current block. Here, the primary transform may correspond to a non-separable transform or a separable transform, and the secondary transform may correspond to a non-separable transform. The method for determining the transform type for the primary and secondary transforms is as described with reference to Fig. 4, and a duplicate description will be omitted here.
[0219] Referring to FIG. 6, a bitstream can be generated by encoding residual information regarding the transform coefficients of the current block (S630).
[0220] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.
[0221] Referring to FIG. 7, the encoding device (200) may include a reference sample determination unit (700), a prediction block generation unit (710), a transform coefficient derivation unit (720), and a residual information encoding unit (730). The reference sample determination unit (700) and the prediction block generation unit (710) may be provided in the intra prediction unit (222) of FIG. 2, the transform coefficient derivation unit (720) may be provided in the residual processing unit (230) of FIG. 2, and the residual information encoding unit (730) may be provided in the entropy encoding unit (240).
[0222] The reference sample determination unit (700) can determine reference samples for the MIP of the current block. The reference sample determination method can be determined based on at least one of the aforementioned methods 1 to 3, and any duplicate descriptions will be omitted herein.
[0223] The prediction block generation unit (710) can generate a prediction block of the current block based on reference samples and a matrix kernel. The prediction block generation unit (710) can generate a prediction block of the current block by applying the same method as described with reference to FIG. 4.
[0224] The transform coefficient derivation unit (720) can derive transform coefficients of the current block based on the residual block of the current block. Specifically, the transform coefficient derivation unit (720) can derive transform coefficients based on at least one of a non-separable transform or a separable transform for the residual block. Alternatively, the transform coefficient derivation unit (720) can also derive transform coefficients based on a primary transform and / or a secondary transform for the residual block.
[0225] The residual information encoding unit (730) can encode residual information regarding transform coefficients.
[0226] In the embodiments described above, the methods are described based on a flowchart as a series of steps or blocks. However, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps depicted in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of this document.
[0227] The method according to the embodiments of the present document described above can be implemented in the form of software, and the encoding device and / or decoding device according to the present document can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0228] When the embodiments in this document are implemented as software, the above-described method can be implemented as a module (process, function, etc.) that performs the above-described function. The module can be stored in memory and executed by a processor. The memory can be internal or external to the processor and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), another chipset, logic circuit, and / or data processing device. The memory can include a read-only memory (ROM), a random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in each drawing can be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored on a digital storage medium.
[0229] In addition, the decoding device and encoding device to which the embodiment(s) of the present specification are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, a video phone video device, a transportation terminal (ex. a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, and may be used to process a video signal or a data signal. For example, the OTT (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0230] In addition, the processing method to which the embodiment(s) of the present specification are applied can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of the present specification can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0231] Additionally, the embodiments of the present disclosure may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0232] FIG. 8 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0233] Referring to FIG. 8, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0234] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted.
[0235] The above bitstream can be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0236] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, in which case the control server controls commands / responses between each device within the content streaming system.
[0237] The streaming server can receive content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0238] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.
[0239] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0240] The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a method.
Claims
1. A step of determining reference samples for MIP (matrix-based intra prediction) of the current block; A step of generating a prediction block of the current block based on the above reference samples and matrix kernel; A step of generating a residual block of the current block; and A method comprising the step of restoring the current block based on the prediction block and the residual block.
2. In paragraph 1, A method wherein the above matrix kernel is determined as one of a plurality of matrix kernels based on the above reference samples.
3. In paragraph 1, Based on the above reference samples, a matrix set for the current block is determined among a plurality of matrix sets, A method wherein any one of a plurality of matrix kernels belonging to the above-determined matrix set is determined as the matrix kernel.
4. In paragraph 3, A method wherein the matrix set for the current block is determined based on at least one of the difference, variance, mean, gradient, or frequency component distribution of the reference samples.
5. In paragraph 1, A method wherein, if the values of the above reference samples are the same, a first matrix set can be selected for the current block, otherwise, a second matrix set can be selected for the current block.
6. In paragraph 1, A method wherein, if at least one of the above reference samples has a value of 0 and the value of the remaining reference samples corresponds to a single non-zero value, a first matrix set may be selected for the current block, otherwise, a second matrix set may be selected for the current block.
7. In paragraph 1, The above reference samples are divided into one or more sample groups consisting of reference samples having the same value, A method wherein, if the number of the above sample groups is less than or equal to T, a first matrix set may be selected for the current block, otherwise, a second matrix set may be selected for the current block.
8. In paragraph 1, A method in which, when the values of the above reference samples are the same, the value of at least one of the prediction samples, which is the output of the matrix kernel, is set to a predetermined value.
9. In paragraph 1, A method in which, when the values of the above reference samples are 0, at least one of the input vectors of the matrix kernel is set to a value other than 0.
10. In paragraph 1, Samples belonging to at least one of the first region or the second region for the current block are determined as reference samples, A method wherein the first region includes at least one of the upper peripheral region or the upper right peripheral region of the current block, and the second region includes at least one of the left peripheral region or the lower left peripheral region of the current block.
11. A step of determining reference samples for MIP (matrix-based intra prediction) of the current block; A step of generating a prediction block of the current block based on the above reference samples and matrix kernel; A step of deriving transform coefficients of the current block based on a residual block of the current block; and A method comprising the step of encoding residual information regarding the above transformation coefficients.
12. A computer-readable storage medium storing a bitstream generated by the method according to Article 11.
13. A step of obtaining a bitstream for image information; wherein the bitstream is generated based on a step of determining reference samples for MIP (matrix-based intra prediction) of a current block, a step of generating a prediction block of the current block based on the reference samples and a matrix kernel, a step of deriving transform coefficients of the current block based on a residual block of the current block, and a step of encoding residual information about the transform coefficients, and A method comprising the step of transmitting data including the bitstream.
Citation Information
Patent Citations
Image decoding device, image encoding device, image processing system, image decoding method and program
JP2023015241A
Device for point of care molecular diagnosis
KR1020230101701A
Method, apparatus and computer program for generating surrounding environment information for automatic driving control of vehicle
KR102481084B1
Regenerative braking amount determine method for emergency braking of electric vehicles
KR102514400B1
KR20200108076A