Method and device for image encoding / decoding
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-05
- Publication Date
- 2026-08-13
Smart Images

Figure KR2026002170_13082026_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device
[0001] The present invention relates to an image encoding / decoding method and apparatus, and more specifically, to an image encoding / decoding method and apparatus for improving the efficiency of intra-prediction of chrominance components using intra-prediction information of luminance components.
[0002] Recently, the demand for multimedia data, such as video, has been increasing rapidly. In particular, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition), is growing across various application fields. High-resolution, high-quality video data involves a much larger volume compared to conventional video data. Consequently, the transmission and storage costs for storing and / or transmitting such high-resolution, high-quality video data increase compared to conventional video data.
[0003] To solve these problems, high-efficiency video encoding / decoding technology for videos with higher resolution and quality is required.
[0004] To encode images, various techniques are used, such as intra-prediction techniques that predict pixel values within the current picture using pixel information within the current picture, intra-prediction techniques that predict pixel values from previous or subsequent pictures, transformation and quantization techniques to compress the energy of residual signals—the difference between the predicted signal and the original signal—and entropy coding techniques that assign short codes to values with high frequency and long codes to values with low frequency. Furthermore, to improve image encoding efficiency, various tools are being developed to implement each of these techniques. Additionally, to decode the encoded image, the image can be restored and reproduced through image decoding techniques that utilize technologies and tools corresponding to the image encoding techniques.
[0005] By utilizing these video encoding and video decoding technologies, video data can be effectively compressed, transmitted, stored, and played back.
[0006] The present disclosure aims to provide an image encoding / decoding method and apparatus that improve the inefficiency of an intra-prediction method for a chrominance block based on intra-prediction information of a corresponding luminance block and improve the compression efficiency of image data.
[0007] The technical problems to be solved by the present disclosure are not limited to those mentioned above. In addition, other technical problems not mentioned in the present disclosure will be clearly understood by those skilled in the art from the present disclosure.
[0008] An image decoding method according to an embodiment of the present invention comprises the steps of: deriving a prediction mode of a current color difference block based on whether a luminance component-based color difference prediction is available; applying downsampling to samples of a corresponding luminance block corresponding to the current color difference block and a luminance region adjacent to the corresponding luminance block; deriving a convolutional filter based on samples of a color difference region adjacent to the current color difference block and a luminance region adjacent to the corresponding luminance block; and generating a predicted value of the current color difference block based on the downsampled corresponding luminance block and the convolutional filter, wherein the availability of the luminance component-based color difference prediction may be determined based on information of the current color difference block.
[0009] In the above image decoding method, the availability of the luminance component-based color difference prediction may be characterized by being determined based on the size of the current color difference block.
[0010] In the above image decoding method, the availability of the luminance component-based color difference prediction may be characterized by being determined based on information indicating the availability of the luminance component-based color difference prediction.
[0011] In the above image decoding method, when it is determined that the luminance component-based color difference prediction is available, the prediction mode of the current color difference block may be characterized by being derived based on information indicating whether to apply the luminance component-based color difference prediction.
[0012] In the above image decoding method, the color difference region adjacent to the current color difference block may be characterized by being determined based on the availability of samples adjacent to the current color difference block.
[0013] In the above image decoding method, the luminance region adjacent to the corresponding luminance block may be characterized by being determined based on the availability of samples adjacent to the corresponding luminance block.
[0014] In the above image decoding method, the step of applying downsampling to samples in a luminance region adjacent to the corresponding luminance block may be characterized by applying one downsampling filter determined among a plurality of downsampling filters to samples in a luminance region adjacent to the corresponding luminance block.
[0015] In the above image decoding method, the determined one downsampling filter may be characterized by being determined based on information indicating one downsampling filter among a plurality of downsampling filters.
[0016] In the above image decoding method, the convolutional filter may be characterized as being one convolutional filter determined from among a plurality of convolutional filters having different components.
[0017] In the above image decoding method, the determined convolutional filter may be characterized by being determined based on information indicating one convolutional filter among a plurality of convolutional filters.
[0018] In the above image decoding method, the step of deriving the convolutional filter may be characterized by deriving the coefficients of the convolutional filter using Gaussian elimination.
[0019] A video encoding method according to an embodiment of the present invention comprises the steps of: deriving a prediction mode of a current color difference block based on whether a luminance component-based color difference prediction is available; applying downsampling to samples of a corresponding luminance block corresponding to the current color difference block and a luminance region adjacent to the corresponding luminance block; deriving a convolutional filter based on samples of a color difference region adjacent to the current color difference block and a luminance region adjacent to the corresponding luminance block; and generating a predicted value of the current color difference block based on the downsampled corresponding luminance block and the convolutional filter, wherein the availability of the luminance component-based color difference prediction may be determined based on information of the current color difference block.
[0020] A non-transient computer-readable recording medium storing a bitstream generated by an image encoding method according to an embodiment of the present invention may store a bitstream generated by an image encoding method, comprising the steps of: deriving a prediction mode of a current color difference block based on whether a luminance component-based color difference prediction is available; applying downsampling to samples of a corresponding luminance block corresponding to the current color difference block and a luminance region adjacent to the corresponding luminance block; deriving a convolutional filter based on samples of a color difference region adjacent to the current color difference block and a luminance region adjacent to the corresponding luminance block; and generating a predicted value of the current color difference block based on the downsampled corresponding luminance block and the convolutional filter, wherein the availability of the luminance component-based color difference prediction is determined based on information of the current color difference block.
[0021] A method for transmitting a bitstream generated by an image encoding method according to an embodiment of the present invention includes the step of transmitting the bitstream, and, based on whether a luminance component-based color difference prediction is available, the step of deriving a prediction mode of a current color difference block, the step of applying downsampling to samples of a corresponding luminance block corresponding to the current color difference block and a luminance region adjacent to the corresponding luminance block, the step of deriving a convolutional filter based on samples of a color difference region adjacent to the current color difference block and a luminance region adjacent to the corresponding luminance block, and the step of generating a prediction value of the current color difference block based on the downsampled corresponding luminance block and the convolutional filter, wherein the availability of the luminance component-based color difference prediction is determined based on information of the current color difference block, and the bitstream generated by the image encoding method can be transmitted.
[0022] The present disclosure aims to provide an image encoding / decoding method and apparatus that improve the inefficiency of an intra-prediction method for a chrominance block based on intra-prediction information of a corresponding luminance block and improve the compression efficiency of image data.
[0023] Additionally, according to the present disclosure, a recording medium storing a bitstream generated by the image encoding method or device of the present invention may be provided.
[0024] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description below.
[0025] FIG. 1 is a block diagram showing an image encoding device according to one embodiment of the present invention.
[0026] FIG. 2 is a block diagram showing an image decoding device according to one embodiment of the present invention.
[0027] FIG. 3 is a schematic diagram showing a video coding system to which the present invention can be applied.
[0028] FIG. 4 is a diagram illustrating an exemplary content streaming system to which an embodiment according to the present invention can be applied.
[0029] FIG. 5 is a diagram illustrating a luminance component-based color difference component prediction according to one embodiment of the present disclosure.
[0030] FIG. 6 is a diagram illustrating a method for predicting color difference components based on luminance components according to one embodiment of the present disclosure.
[0031] Figure 7 is a diagram illustrating the template area used in the luminance component-based color difference component prediction method.
[0032] FIG. 8 is a diagram illustrating a template area used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0033] FIG. 9 is a diagram illustrating samples of template regions used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0034] FIG. 10 is a diagram illustrating a template area used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0035] FIG. 11 is a drawing illustrating a downsampling filter used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0036] FIG. 12 is a drawing illustrating a downsampling filter used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0037] FIG. 13 is a drawing illustrating a downsampling filter used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0038] FIG. 14 is a drawing illustrating a downsampling filter used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0039] FIG. 15 is a diagram illustrating the shape of a convolutional filter used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0040] FIG. 16 is a diagram illustrating a method for predicting color difference components based on luminance components using a convolutional filter according to one embodiment of the present disclosure.
[0041] FIG. 17 is a diagram illustrating a method for predicting color difference components based on luminance components using a convolutional filter according to one embodiment of the present disclosure.
[0042] FIG. 18 is a diagram illustrating a method for predicting color difference components based on luminance components using a convolutional filter according to one embodiment of the present disclosure.
[0043] FIG. 19 is a diagram illustrating a method for predicting color difference components based on luminance components using a convolutional filter according to one embodiment of the present disclosure.
[0044] FIG. 20 is a diagram illustrating a method for predicting color difference components based on luminance components using a convolutional filter according to one embodiment of the present disclosure.
[0045] FIG. 21 is a drawing illustrating a method for deriving convolutional filter coefficients according to one embodiment of the present disclosure.
[0046] FIG. 22 is a diagram illustrating a method for predicting color difference components based on luminance components according to one embodiment of the present disclosure.
[0047] The present invention is susceptible to various modifications and may have various embodiments; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the invention to specific embodiments, and it should be understood that the invention includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention. Similar reference numerals have been used for similar components in the description of each drawing.
[0048] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.
[0049] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0050] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0051] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. Hereinafter, the same reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted.
[0052] FIG. 1 is a block diagram showing an image encoding device according to one embodiment of the present invention.
[0053] Referring to FIG. 1, the image encoding device (100) may include an image segmentation unit (101), an intra prediction unit (102), an inter prediction unit (103), a subtraction unit (104), a conversion unit (105), a quantization unit (106), an entropy encoding unit (107), an inverse quantization unit (108), an inverse conversion unit (109), an addition unit (110), a filter unit (111), and a memory (112).
[0054] Each component shown in FIG. 1 is depicted independently to represent different characteristic functions of the image encoding device and does not imply that each component consists of separate hardware or a single software unit. That is, each component is listed and included as a separate component for the convenience of explanation, but at least two of the components may be combined to form a single component, or a single component may be divided into multiple components to perform functions, and such integrated and separated embodiments of each component are included within the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0055] Furthermore, some components may not be essential components performing an essential function in the present invention, but merely optional components for enhancing performance. The present invention may be implemented by including only the components essential for realizing the essence of the present invention, excluding components used solely for performance enhancement, and a structure including only the essential components, excluding optional components used solely for performance enhancement, is also included within the scope of the rights of the present invention.
[0056] The image segmentation unit (101) can divide the input image into at least one block. At this time, the input image may have various shapes and sizes, such as a sequence, picture, slice, tile, segment, tile group, coding tree unit, etc. According to another embodiment, the image segmentation unit (101) may divide an input picture into a plurality of sub-pictures defined as a group of rectangular slices, divide each sub-picture into the tile / slice, and divide the tile / slice into coding tree units.
[0057] Additionally, the image segmentation unit (101) can recursively segment the segmented coding tree unit. The terminal node segmented from the coding tree unit may be referred to as a coding unit (CU). A block may refer to a coding unit (CU), or a prediction unit (PU) or a transformation unit (TU) segmented from the coding unit (CU). The segmentation may be performed based on at least one of a quadtree, a binary tree, and a ternary tree. A quadtree is a method of dividing an upper block into lower blocks, each having a width and height that are half that of the upper block. A binary tree is a method of dividing an upper block into lower blocks, each having either a width or a height that is half that of the upper block. A ternary tree is a method of dividing an upper block into three lower blocks. For example, the three lower blocks may be obtained by dividing the width or height of the upper block in a ratio of 1:2:1. Through the aforementioned binary tree-based partitioning, blocks can have shapes that are not only square but also non-square. Blocks can first be partitioned into a quad tree. Blocks corresponding to the leaf nodes of the quad tree may not be partitioned, or may be partitioned into a binary tree or a terminal tree. The leaf nodes of the binary tree or terminal tree may be units of encoding, prediction, and / or transformation.
[0058] The image segmentation unit (101) can recursively segment the CTU into a quad tree (QT) as well as a multi-type tree (MTT). Here, the MTT can be composed of a binary tree (BT) and a triple tree (TT). For example, the MTT structure can be divided into a vertical binary tree segmentation mode (SPLIT_BT_VER), a horizontal binary tree segmentation mode (SPLIT_BT_HOR), a vertical binary tree segmentation mode (SPLIT_TT_VER), and a horizontal binary tree segmentation mode (SPLIT_TT_HOR).
[0059] Additionally, the image segmentation unit (101) can segment the CTU by applying a dual tree in which the CTU segmentation structures of the luminance and color difference components are used differently, or by applying a single tree in which the luminance and color difference CTBs (Coding Tree Blocks) within the CTU share a coding tree structure.
[0060] Alternatively, the image segmentation unit (101) may divide the input image into at least one block. At this time, the input image or frame may be divided into tiles. The tiles are divided into superblocks having a predetermined size, and each superblock may be divided into blocks. The image segmentation unit (101) may recursively divide the superblock into blocks. Here, the superblock may be divided into two or four vertically or horizontally, or recursively divided into four. Alternatively, the superblock may be divided once each in the vertical and horizontal directions to form three blocks. Additionally, the image segmentation unit (101) may divide the blocks into units of transform blocks, which are units of transformation and / or prediction.
[0061] The prediction unit (102, 103) may include an intra prediction unit (102) that performs intra prediction and an inter prediction unit (103) that performs inter prediction. The prediction unit (102, 103) may determine whether to use intra prediction or perform inter prediction for a prediction unit. Additionally, the prediction unit (102, 103) may determine specific information (e.g., intra prediction mode, inter prediction mode, motion vector, reference picture, etc.) according to the determined prediction method. At this time, the processing unit in which the prediction is performed and the processing unit in which the prediction method and specific details are determined may be different. For example, the prediction unit (102, 103) may determine the prediction method and prediction mode, etc. for each prediction unit and perform prediction according to the transformation unit.
[0062] According to another embodiment, the prediction unit may encode an input image using a third mode (e.g., IBC mode, Palette mode, etc.) that is a mode other than the intra mode and the inter mode. However, if the third mode has functional characteristics similar to the intra mode or the inter mode, the third mode may be classified as the intra mode or the inter mode. In this disclosure, the third mode will be described only when a specific description of the third mode is required.
[0063] Alternatively, the intra prediction mode used for intra prediction may include a directional prediction mode that uses reference pixel information according to the prediction direction, a non-directional mode that does not use directional information, a Recursive Intra prediction (RIP), and a Paeth intra prediction mode. Additionally, the mode for predicting luminance information and the mode for predicting chrominance information may be different, and the intra prediction mode information of the luminance component block or the predicted luminance signal information may be utilized to predict chrominance information.
[0064] In particular, when the current block is a block of chroma components, the intra prediction unit (102) can perform chroma intra prediction based on the CFL (chroma from luma) mode using a lumina block corresponding to the current block. In particular, the intra prediction unit (102) can generate an intra prediction block for the current block by performing a cross component prediction (CCP) using at least one of a corresponding lumina block, a block adjacent to the corresponding lumina block, and a block adjacent to the chroma block. Here, the cross component prediction may be a multi-hypothesis cross component prediction (CCP).
[0065] The intra prediction unit (102) can generate a prediction block of the current block based on the intra prediction mode of the current block and reference pixel information around the current block, which is pixel information within the current picture. If the surrounding blocks of the current block are predicted by inter prediction, the reference pixels included in the inter-predicted surrounding blocks can be replaced with reference pixels in other surrounding blocks that are intra-predicted. That is, if a reference pixel is not available, the intra prediction unit (102) can perform intra prediction of the current block by replacing the unavailable reference pixel with at least one of the available reference pixels.
[0066] The intra prediction unit (102) may use multiple reference pixel lines for intra prediction of the current block. When multiple reference pixel lines are available, information indicating the reference pixel line used for intra prediction among the multiple reference pixel lines may be signaled.
[0067] The intra prediction mode used for intra prediction may include a directional prediction mode that uses reference pixel information according to the prediction direction and a non-directional mode that does not use directional information. Here, the intra prediction mode may be derived using a list of most probable modes (MPM), and in particular, may be derived using a first MPM and / or second MPM list. Additionally, the mode for predicting luminance information and the mode for predicting chrominance information may be different, and the intra prediction mode information of the luminance component block or the predicted luminance signal information may be utilized to predict chrominance information.
[0068] Alternatively, the intra prediction unit (102) may perform intra prediction for the current block by applying at least one of the following modes: DIMD (decoder-side intra mode derivation), BVG-DIMD (block-vector guided DIMD), OBIC (Occurrence-based intra coding), EIP (extrapolation filter based intra prediction mode), BVG-DIMD (block-vector guided EIP), MM-EIP (multi-model EIP), TIMD (template based intra mode derivation), TIMD merge mode, SGPM (spatial geometric partitioning mode) mode, intra template matching prediction (IntraTMP), and intra block copy. When the intra prediction mode of the block is a predetermined mode, the intra prediction unit (102) may perform intra prediction for the current block using a block vector indicating a block other than the current block.
[0069] In particular, when the current block is a block of chroma components, the intra prediction unit (102) can perform intra prediction for the current block by applying at least one of DM (direct mode), DBV (direct block-vector), DIMD chroma mode, CCLM (cross-component linear model), CCCM (convolutional cross-component model), BVG-CCCM (block-vector guided CCCM), and GL-CCCM (gradient and location based CCCM).
[0070] The intra prediction unit (102) may include a reference sample filter, an interpolation filter, and a DC filter. The reference sample filter is a filter that performs filtering on the reference pixel of the current block and may be adaptively applied depending on the prediction mode, size, shape, and / or whether the reference pixel is included in the reference pixel line immediately adjacent to the current block. If the prediction mode of the current block is a mode that does not perform reference sample filtering, the reference sample filter may not be applied.
[0071] The interpolation filter is a filter that interpolates and filters the prediction samples of the current block, and can be applied adaptively depending on the prediction mode, size, shape, and / or whether the reference pixel is included in the reference pixel line immediately adjacent to the current block.
[0072] If the prediction mode of the current block is DC mode, a prediction block can be generated by applying a DC filter.
[0073] According to one embodiment, the intra prediction unit (102) can perform intra prediction using a pre-trained neural network (NN) model. For example, the intra prediction unit (102) can induce an intra prediction mode of a block to which a DIMD mode is applied and use a pre-trained neural network model to perform intra prediction.
[0074] The inter prediction unit (103) generates a prediction block using a previously restored reference image stored in memory (112), an inter prediction mode, and motion information. Here, inter prediction may mean motion prediction or motion compensation.
[0075] The inter prediction unit (103) can set the inter mode of the prediction unit included in the encoding unit to one of Skip Mode, Merge Mode, or Advanced Motion Vector Prediction (AMVP) Mode in order to perform motion prediction and / or motion compensation. In addition, the inter prediction unit (103) can perform motion prediction and / or motion compensation for the prediction unit according to the set mode.
[0076] Additionally, the inter prediction unit (103) can perform motion prediction and / or motion compensation for the prediction unit by applying the AFFINE mode of sub-PU-based prediction, the SbTMVP (Subblock-based Temporal Motion Vector Prediction) mode, and the MMVD (Merge with MVD) mode and GPM (Geometric Partitioning Mode) mode of PU-based prediction based on the inter prediction mode. In addition, the inter prediction unit (103) can perform motion prediction and / or motion compensation for the prediction unit by applying HMVP (History based MVP), PAMVP (Pairwise Average MVP), CIIP (Combined Intra / Inter Prediction), AMVR (Adaptive Motion Vector Resolution), DMVR (Decoder side Motion Vector Refinement), BDOF (Bi-Directional Optical-Flow), PROF (Prediction Refinement With Optical Flow), BCW (Bi-predictive with CU Weights), LIC (Local Illumination Compensation), TM (Template Matching), OBMC (Overlapped Block Motion Compensation), etc., to improve the performance of each mode.
[0077] Here, AFFINE mode can be used in both AMVP and MERGE modes. It is a technique with high encoding efficiency. AFFINE mode can be a prediction mode using a 4-parameter affine motion model using two control point motion vectors (CPMV) and a 6-parameter affine motion model using three control point motion vectors. Here, CPMV can be a vector representing any one of the top-left, top-right, or bottom-left affine motion models of the current block.
[0078] Motion information may include, for example, motion vectors, reference picture indices, List 1 prediction flags, List 0 prediction flags, half-sample interpolation filter indices, bidirectional prediction weight indices, etc.
[0079] Alternatively, the inter prediction unit (103) may perform motion vector prediction (MVP) to perform motion prediction and / or motion compensation, and apply a simple inter prediction mode, OBMC (Overlapped Block Motion Compensation), and a local warp mode using modified motion information based on an affine model to the encoding unit to perform motion prediction and / or motion compensation for the prediction unit.
[0080] Alternatively, the inter prediction unit (103) may perform motion prediction and / or motion compensation for the prediction unit by applying a compound prediction mode that synthesizes different prediction values. For example, the inter prediction unit (103) may perform motion prediction and / or motion compensation for the prediction unit by applying Compound Wedge Prediction, Frame distance based compound prediction, Inter-Intra Prediction, etc.
[0081] According to one embodiment, the inter prediction unit (103) can perform inter prediction using a pre-trained neural network model. For example, the inter prediction unit (103) can synthesize a reference frame using a pre-trained neural network model and perform inter prediction based on the synthesized reference frame.
[0082] A residual block containing residual value information, which is the difference between the prediction unit generated in the prediction unit (102, 103) and the original block of the prediction unit, can be generated. The generated residual block can be input to the conversion unit (130) and converted.
[0083] The subtraction unit (104) subtracts the prediction block generated by the intra prediction unit (102) or the inter prediction unit (103) from the block currently to be encoded to generate a residual block of the current block. The residual value (residual block) between the generated prediction block and the original block can be input to the conversion unit (105).
[0084] In addition, the prediction mode information and motion vector information used for prediction can be encoded in the entropy encoding unit (107) along with the residual value and transmitted to the decoder. When a specific encoding mode is used, it is also possible to encode the original block as is and transmit it to the decoder without generating a prediction block through the prediction unit (102, 103).
[0085] The transformation unit (105) can perform a transformation on a residual block containing residual data to generate and output a transformation coefficient. Here, the transformation coefficient may be a coefficient value generated by performing a transformation on the residual block. When a transform skip mode is applied, the transformation unit (105) may skip the transformation on the residual block.
[0086] The conversion unit (105) can determine a conversion type and a conversion kernel based on at least one of encoding parameters such as the size of the conversion block, color component, and prediction mode, and perform a conversion on the conversion block using the determined conversion type and conversion kernel.
[0087] According to one embodiment, the transformation unit (105) can perform a transformation on the 4x4 luminance residual block generated from the intra prediction result using a transformation type and transformation kernel according to the Discrete Sine Transform (DST), and for the remaining residual block, perform a transformation using a transformation type and transformation kernel according to the Discrete Cosine Transform (DCT).
[0088] According to another embodiment, the transformation unit (105) may apply MTS (Multiple Transform Selection) technology, which performs transformations by selectively using various transformation types and transformation kernels. That is, the transformation unit (105) may perform transformations on a sub-block basis through SBT (Sub-block Transform) technology. Specifically, SBT may be applied only to inter-predicted blocks, and the current block may be divided into ½ or ¼ sizes in the vertical or horizontal direction, and transformation may be performed on only one of the blocks. For example, the transformation unit (105) may perform transformation on the leftmost or rightmost block among the current blocks divided vertically, and perform transformation on the topmost or bottommost block among the current blocks divided horizontally.
[0089] According to another embodiment, the transformation unit (105) may apply a non-separable primary transform (NSPT) technique that performs transformations by selectively using multiple transformation kernels based on an intra-prediction mode or the size and / or shape of the block.
[0090] According to another embodiment, the conversion unit (105) may apply a Low Frequency Non-Separable Transform (LFNST), which is a technique for applying a secondary transform to a residual signal converted into the frequency domain through a DCT or DST. LFNST can additionally perform a transformation on a 4x4 or 8x8 low-frequency region in the upper left to concentrate the residual coefficients to the upper left.
[0091] According to another embodiment, the transformation unit (105) can perform a transformation using a transformation type and a transformation kernel according to at least one of DCT (Discrete Cosine Transform), ADST (Asymmetric Discrete Sine Transform), IDTX (Identity Transform), and WHT (Walsh Hadamard Transform). That is, the transformation unit (105) can derive a transform coefficient using a transformation type and a transformation kernel determined in the transformation block. The transform coefficient derived using a transformation type and a transformation kernel according to at least one of DCT, ADST, IDTX, and WHT can be referred to as a first-order transform coefficient. Additionally, the transformation unit (105) can derive a second-order transform coefficient by applying a second-order transform to the first-order transform coefficient. Furthermore, the transformation unit can rearrange the second-order transform coefficient according to a predetermined scan direction.
[0092] The quantization unit (106) can quantize the conversion coefficient or residual signal converted into the frequency domain by the conversion unit (105) according to a quantization parameter (QP). The quantization parameter may vary depending on the block or the importance of the image. The value calculated by the quantization unit (106) may be provided to the inverse quantization unit (108) and the entropy encoding unit (107).
[0093] The above-mentioned conversion unit (105) and / or quantization unit (106) may be optionally included in the image encoding device (100). That is, the image encoding device (100) may perform at least one of conversion or quantization on the residual data of the residual block, or may encode the residual block by skipping both conversion and quantization. Even if neither conversion nor quantization is performed in the image encoding device (100), or if neither conversion nor quantization is performed, the block that enters as input to the entropy encoding unit (107) is typically referred to as a conversion block.
[0094] The entropy encoding unit (107) can generate and output a bitstream by performing entropy encoding according to a probability distribution on values output by the quantization unit (106), coding parameter values output during the encoding process, information for decoding an image, etc. Here, the information for decoding an image may include syntax elements, etc.
[0095] Coding parameters may include information (flags, indexes, etc.) that is encoded in the encoding device (100) and signaled to the decoding device (200), such as syntax elements, as well as information derived during the encoding process or decoding process, and may refer to information required when encoding or decoding images.
[0096] The entropy encoding unit (107) can encode various information such as coefficient information of a conversion block, block type information, prediction mode information, division unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information. The coefficients of the conversion block can be encoded in units of sub-blocks within the conversion block.
[0097] For encoding the coefficients of a transform block, various syntax elements may be encoded, such as Last_sig, a syntax element indicating the location of the first non-zero coefficient in backscan order; Coded_sub_blk_flag, a flag indicating whether there is at least one non-zero coefficient in the subblock; Sig_coeff_flag, a flag indicating whether it is a non-zero coefficient; Abs_greater1_flag, a flag indicating whether the absolute value of the coefficient is greater than 1; Abs_greater2_flag, a flag indicating whether the absolute value of the coefficient is greater than 2; and Sign_flag, a flag indicating the sign of the coefficient. The remaining values of the coefficients that are not encoded by the above syntax elements alone may be encoded through the syntax element remaining_coeff.
[0098] When entropy encoding is applied, the size of the bit sequence for the symbols to be encoded can be reduced by allocating a small number of bits to symbols with a high probability of occurrence and a large number of bits to symbols with a low probability of occurrence when representing symbols. Input data is entropied. For example, entropy encoding can utilize various encoding methods such as Exponential Golomb and CABAC (Context-Adaptive Binary Arithmetic Coding).
[0099] The inverse quantization unit (108) and the inverse transformation unit (109) can inverse quantize the values quantized in the quantization unit (106) and inverse transform the values transformed in the transformation unit (105). The residual value generated in the inverse quantization unit (108) and the inverse transformation unit (109) can be combined with the prediction unit predicted through the motion estimation unit, motion compensation unit, and intra prediction unit (102) included in the prediction unit (102, 103) to generate a reconstructed block. The addition unit (110) generates a reconstructed block by adding the prediction block generated in the prediction unit (102, 103) and the residual block generated through the inverse transformation unit (109).
[0100] The filter section (111) can apply deblocking filter, Sample Adaptive Offset (SAO), Adaptive Loop Filter (ALF), Bilateral filter (BIF), LMCS (Luma Mapping with Chroma Scaling), etc., to the reconstructed sample, reconstructed block, or reconstructed image as a whole or part of the filtering technique.
[0101] The deblocking filter can remove block distortion caused by boundaries between blocks in the restored picture. To determine whether to perform deblocking, the decision to apply the deblocking filter to the current block can be made based on the pixels contained in a certain number of columns or rows within the block. When applying the deblocking filter to a block, a Strong Filter or a Weak Filter can be applied depending on the required deblocking filtering strength. Additionally, when applying the deblocking filter, horizontal and vertical filtering can be processed in parallel.
[0102] Sample adaptive offset may be a method of correcting the offset from the original image on a sample-by-sample basis for an image that has undergone deblocking. The filter unit (111) may use a method of dividing the samples included in the image into a certain number of regions, determining the region to perform the offset on, and applying the offset to that region, or a method of applying the offset by considering the edge information of each sample. Here, the sample adaptive offset may be at least one of a general sample adaptive offset, a bilateral filter, and a cross-component sample adaptive offset (CCSAO).
[0103] Adaptive Loop Filtering (ALF) can be performed based on a comparison between the filtered restored image and the original image. After dividing the pixels included in the image into predetermined groups, a single filter to be applied to each group can be determined, allowing for differential filtering for each group. Information regarding whether to apply ALF can be transmitted per coding unit (CU), and the shape and filter coefficients of the ALF filter to be applied may vary depending on each block. Additionally, an ALF filter of the same form (fixed form) may be applied regardless of the characteristics of the block to be applied.
[0104] An adaptive loop filter can perform filtering based on a comparison of the reconstructed image and the original image. After dividing the samples included in the image into predetermined groups, a filter to be applied to each group can be determined, thereby performing filtering differently for each group. Information regarding whether to apply an adaptive loop filter can be signaled per coding unit (CU), and the shape and filter coefficients of the adaptive loop filter to be applied may vary depending on each block.
[0105] Alternatively, the filter unit (111) may apply an edge loop filter, an adaptive loop filter (ALF), a CDEF (Constrained Directional Enhancement Filter), a loop restoration filter, etc., to a restored sample, a restored block, or a restored image as a whole or in part filtering technique.
[0106] According to one embodiment, the filter unit (111) can filter all or part of a restored sample, restored block, or restored image using a pre-trained neural network model. Specifically, the filter unit (111) can apply an adaptive loop filter to all or part of a restored sample, restored block, or restored image using a pre-trained neural network model.
[0107] The memory (112) can store a restored block or picture calculated through the filter unit (111). The memory (112) may include a reference picture buffer. Additionally, the stored restored block or picture in the memory (112) may be provided to the prediction unit (102, 103) when performing inter-prediction.
[0108] Next, an image decoding device according to one embodiment of the present invention will be described with reference to the drawings.
[0109] FIG. 2 is a block diagram showing an image decoding device (200) according to one embodiment of the present invention.
[0110] Referring to FIG. 2, the image decoding device (200) may include an entropy decoding unit (201), an inverse quantization unit (202), an inverse transformation unit (203), a prediction unit (204, 205), an adder unit (206), a filter unit (207), and a memory (208).
[0111] The video decoding device (200) can receive a bitstream output by the video encoding device (100). The video decoding device (200) can receive a bitstream stored in a computer-readable recording medium or receive a bitstream stream streamed through a wired / wireless transmission medium. The video decoding device (200) can decode the bitstream to generate a restored video or a decoded video, and can output the restored video or the decoded video.
[0112] The entropy decoding unit (201) can generate symbols by performing entropy decoding according to the probability distribution of the bitstream. The generated symbols may include symbols in the form of quantized levels. Here, the entropy decoding method may be the inverse process of the entropy encoding method described above.
[0113] The entropy decoding unit (201) can convert a one-dimensional vector-shaped coefficient into a two-dimensional block-shaped coefficient through a conversion coefficient scanning method to decode a conversion coefficient level (quantized level).
[0114] The entropy decoding unit (201) can perform entropy decoding in the opposite procedure to that which the entropy encoding unit (107) of the image encoding device (100) performed. For example, various methods such as Exponential Golomb and CABAC (Context-Adaptive Binary Arithmetic Coding) can be applied in correspondence with the method performed in the image encoder.
[0115] The entropy decoding unit (201) can obtain various information such as coefficient information of a conversion block as described above, block type information, prediction mode information, division unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information by decoding.
[0116] The inverse quantization unit (202) generates a conversion block by performing inverse quantization on the quantized conversion block. It operates substantially the same as the inverse quantization unit (108) of FIG. 1.
[0117] The inverse transformation unit (203) performs an inverse transformation on the transformation block to generate a residual block. At this time, the transformation method may be determined based on information regarding the prediction method (inter or intra prediction), the size and / or shape of the block, the intra prediction mode, etc. It operates substantially the same as the inverse transformation unit (109) of FIG. 1.
[0118] According to one embodiment, the inverse transformation unit (203) can perform an inverse transformation using a transformation type and a transformation kernel according to the Discrete Sine Transform (DST) on the transformation coefficient level of the 4x4 luminance component generated from the intra prediction result, and perform an inverse transformation using a transformation type and a transformation kernel according to the Discrete Cosine Transform (DCT) on the remaining transformation coefficient level.
[0119] According to another embodiment, the inverse transformation unit (203) may apply MTS (Multiple Transform Selection) technology, which performs transformations by selectively using multiple transformation kernels.
[0120] According to another embodiment, the inverse transform unit (203) may apply LFNST (Low Frequency Non-Separable Transform), which is a technique for applying a secondary inverse transform to the inverse transform coefficient level through a DCT or DST-based transform type and a transform kernel.
[0121] According to another embodiment, the inverse transform unit (203) may apply a non-separable primary transform (NSPT) technique that performs the inverse transform by selectively using a transform kernel based on an intra-prediction mode or the size and / or shape of the block.
[0122] The prediction unit (204, 205) can generate a prediction block based on the prediction block generation information provided by the entropy decoding unit (201) and the previously decoded block or picture information provided by the memory (208).
[0123] The prediction unit (204, 205) may include an intra prediction unit (204) and an inter prediction unit (205). The prediction unit (204, 205) receives various information, such as prediction unit information input from the entropy decoding unit (201), prediction mode information of the intra prediction method, and motion prediction related information of the inter prediction method, distinguishes the prediction unit from the current encoding unit, and determines the prediction mode of the prediction unit.
[0124] The intra prediction unit (204) can generate a prediction block of the current block based on the intra prediction mode of the current block and reference pixel information around the current block, which is pixel information within the current picture.
[0125] The intra prediction unit (204) can generate a prediction block based on reference pixel information around the current block, which is pixel information within the current picture. The intra prediction unit (204) can use multiple reference pixel lines for intra prediction. If multiple reference pixel lines are available, the intra prediction unit (204) can obtain information indicating a reference pixel line used for intra prediction among the multiple reference pixel lines.
[0126] The intra prediction mode used for intra prediction may be a directional prediction mode or a non-directional mode. Additionally, the mode for predicting luminance information and the mode for predicting chrominance information may be different, and the intra prediction mode information of the luminance component block or the predicted luminance signal information may be utilized to predict chrominance information.
[0127] Alternatively, the intra prediction mode used for intra prediction may be one of a directional prediction mode, a non-directional mode, RIP, or Paeth intra prediction mode. Additionally, the mode for predicting luminance information and the mode for predicting chrominance information may be different, and the intra prediction mode information of the luminance component block or the predicted luminance signal information may be utilized to predict chrominance information.
[0128] According to one embodiment, the intra prediction unit (204) can perform intra prediction using a pre-trained neural network model. The intra prediction unit (204) can perform intra prediction using the same neural network model as the intra prediction unit (102) of FIG. 1.
[0129] The intra prediction unit (204) operates substantially the same as the intra prediction unit (102) of FIG. 1.
[0130] The inter prediction unit (205) can perform an inter prediction for the current prediction unit based on information included in at least one of the previous or subsequent pictures of the current picture containing the current prediction unit, using information necessary for inter prediction of the current prediction unit provided by the video encoding device (100). Alternatively, the inter prediction may be performed based on information of a partially restored area within the current picture containing the current prediction unit.
[0131] The inter prediction unit (205) can set the inter mode of the prediction unit included in the encoding unit to one of Skip Mode, Merge Mode, or Advanced Motion Vector Prediction (AMVP) Mode in order to perform motion prediction and / or motion compensation. Additionally, the inter prediction unit (205) can perform motion compensation for the prediction unit according to the set mode.
[0132] Additionally, the inter prediction unit (205) can perform motion compensation for the prediction unit by applying the AFFINE mode of sub-PU-based prediction, the SbTMVP (Subblock-based Temporal Motion Vector Prediction) mode, and the MMVD (Merge with MVD) mode and GPM (Geometric Partitioning Mode) mode of PU-based prediction based on the inter prediction mode. Additionally, the inter prediction unit (205) can perform motion compensation for the prediction unit by applying HMVP (History based MVP), PAMVP (Pairwise Average MVP), CIIP (Combined Intra / Inter Prediction), AMVR (Adaptive Motion Vector Resolution), BDOF (Bi-Directional Optical-Flow), BCW (Bi-predictive with CU Weights), LIC (Local Illumination Compensation), TM (Template Matching), OBMC (Overlapped Block Motion Compensation), etc., to improve the performance of each mode.
[0133] Motion information may include, for example, motion vectors, reference picture indices, List 1 prediction flags, List 0 prediction flags, half-sample interpolation filter indices, bidirectional prediction weight indices, etc.
[0134] According to one embodiment, the inter prediction unit (205) can perform inter prediction using a pre-trained neural network model. The inter prediction unit (205) can perform inter prediction using the same neural network model as the inter prediction unit (103) of FIG. 1.
[0135] The inter prediction unit (205) can operate substantially the same as the inter prediction unit (103) of FIG. 1.
[0136] The adder (206) generates a restoration block by adding the prediction block generated in the intra prediction unit (204) or the inter prediction unit (205) and the residual block generated through the inverse transformation unit (203). It operates substantially the same as the adder (110) of FIG. 1.
[0137] The filter section (207) can reduce various types of noise occurring in the restored blocks. The filter section (207) may include a deblocking filter, a sample adaptive offset, an adaptive loop filter, a bidirectional filter, and an LMCS.
[0138] The filter unit (207) may receive information regarding whether each filter is applied, information regarding the filter strength, etc. from the image encoding device (100). The filter unit (207) of the image decoding device (200) may receive filter-related information provided by the image encoding device (100) and perform filtering on the corresponding block in the image decoding device (200).
[0139] According to one embodiment, the filter unit (207) can filter all or part of a restored sample, restored block, or restored image using a pre-trained neural network model. Specifically, the filter unit (207) can apply an adaptive loop filter to all or part of a restored sample, restored block, or restored image using a pre-trained neural network model.
[0140] The filter section (207) can operate substantially the same as the filter section (111) of FIG. 1.
[0141] The memory (208) can store a restoration block generated by the adder (206). For example, the memory (208) may include a reference picture buffer. The memory (208) may operate substantially the same as the memory (112) of FIG. 1.
[0142]
[0143] FIG. 3 is a schematic diagram showing a video coding system to which the present invention can be applied.
[0144] A video coding system according to one embodiment may include an encoding device (10) and a decoding device (20). The encoding device (10) may transmit encoded video and / or image information or data to the decoding device (20) via a digital storage medium or network in the form of a file or streaming.
[0145] An encoding device (10) according to one embodiment may include an image generation unit (11), an encoding unit (12), and a transmission unit (13). A decoding device (20) according to one embodiment may include a receiving unit (21), a decoding unit (22), and an image playback unit (23). The encoding unit (12) may be called a video / image encoding unit, and the decoding unit (22) may be called a video / image decoding unit. The transmission unit (13) may be included in the encoding unit (12). The receiving unit (21) may be included in the decoding unit (22). The image playback unit (23) may include a display unit, and the display unit may be composed of a separate device or an external component.
[0146] The image generation unit (11) can acquire video / image through a process of capturing, synthesizing, or generating video / image. The image generation unit (11) may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / image, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and can generate video / image (electronically). For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.
[0147] The encoding unit (12) can encode the input video / image. The encoding unit (12) can perform a series of procedures such as prediction, conversion, and quantization for compression and encoding efficiency. The encoding unit (12) can output the encoded data (encoded video / image information) in the form of a bitstream. The detailed configuration of the encoding unit (12) can be configured in the same way as the encoding device (100) of FIG. 1 described above.
[0148] The transmission unit (13) can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit (21) of the decoding device (20) via a digital storage medium or network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit (21) can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit (22).
[0149] The decoding unit (22) can decode a video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding unit (12). The detailed configuration of the decoding unit (22) can be configured in the same way as the decoding device (200) of FIG. 2 described above.
[0150] The image playback unit (23) can render the decoded video / image. The rendered video / image can be displayed through the display unit.
[0151]
[0152] FIG. 4 is a diagram illustrating an exemplary content streaming system to which an embodiment according to the present invention can be applied.
[0153] As illustrated in FIG. 4, a content streaming system to which an embodiment of the present invention is applied may largely include a multimedia input device, a media storage, an encoding server, a streaming server, a web server, and a user device.
[0154] The above encoding server plays the role of compressing content input from multimedia input devices, such as smartphones, cameras, and CCTVs, into digital data to generate a bitstream and transmitting it to the above streaming server. Alternatively, the encoding server plays the role of compressing content previously stored in a media storage into digital data to generate a bitstream and transmitting it to the above streaming server.
[0155] As another example, when multimedia input devices such as smartphones, cameras, and CCTVs directly generate bitstreams, the encoding server may be omitted.
[0156] The bitstream above may be generated by a video encoding method and / or video encoding device to which an embodiment of the present invention is applied, and the streaming server may temporarily or non-temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0157] The streaming server transmits multimedia data to a user device based on a user request through a web server, and the web server can act as a medium to inform the user of available services. When a user device requests a desired service from the web server, the web server forwards it to the streaming server, and the streaming server can transmit multimedia data to the user device. At this time, the content streaming system may include a separate control server, and in this case, the control server can perform the role of controlling commands and responses between each device within the content streaming system.
[0158] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a seamless streaming service, the streaming server can store the bitstream for a certain period of time.
[0159] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.
[0160] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.
[0161]
[0162] According to existing video coding methods, color difference components were predicted using simple prediction techniques. However, with the establishment of subsequent standard specifications, more precise color difference prediction technologies have been adopted. For example, the CCLM (Cross-Component Linear Model) method is one of the color difference prediction techniques that predicts color difference components based on luminance components.
[0163] CCLM has significantly improved the precision of color difference component predictions. However, given that it performs predictions linearly by utilizing only the predicted values of the reconstructed luminance component, there is potential to improve accuracy.
[0164]
[0165] FIG. 5 is a diagram illustrating a luminance component-based color difference component prediction according to one embodiment of the present disclosure.
[0166] Referring to FIG. 5, a current color difference block and a restored luminance block corresponding to the current color difference block can be set. Additionally, a current color difference template including samples adjacent to the current color difference block and a luminance template including samples adjacent to the restored luminance block can be set.
[0167] A correlation between the color difference template and the luminance template is derived, and filter coefficients can be generated based on the derived correlation. Then, by applying the filter coefficients to the predicted value of the restored luminance block, a predicted value for the current color difference block can be generated.
[0168]
[0169] And, below, a method for predicting color difference components based on luminance components is explained.
[0170]
[0171] FIG. 6 is a diagram illustrating a luminance component-based chrominance component prediction method according to one embodiment of the present disclosure. The luminance component-based chrominance component prediction method can be performed by an image decoder or an image encoding device.
[0172] Referring to FIG. 6, in step S610, the image decoder can set the current color difference block.
[0173] And, in step S620, the image decoder may set a color difference template and a padding region adjacent to the current color difference block. Here, the color difference template may include at least one sample adjacent to the color difference block. And, the padding region may be a region including the samples adjacent to the color difference template.
[0174] In step S630, the image decoder may apply downsampling to a corresponding luminance block corresponding to the current chrominance block and to a previously restored value of an area adjacent to the corresponding luminance block. Here, the area adjacent to the corresponding luminance block may be referred to as a luminance template. The luminance template may include at least one sample adjacent to the corresponding luminance block. That is, downsampling may be applied to a previously restored luminance sample value corresponding to the current chrominance block and the chrominance template.
[0175] In step S640, the image decoder can derive a convolutional filter based on a chrominance template and a downsampled luminance template. Here, the image decoder can derive the type of the filter, the elements of the filter, and the coefficients of the filter. According to one embodiment, the filter may be a convolutional filter.
[0176] In addition, at step S650, the image decoder can generate a predicted value of the current chrominance block by applying a convolutional filter to the downsampled corresponding luminance block. According to one embodiment, the image decoder can subtract the average value of the downsampled luminance sample values from the sample values of the downsampled corresponding luminance block. The image decoder can derive an initial predicted value by applying a convolutional filter to the luminance value derived from the result of the subtraction operation. In addition, the image decoder can generate a predicted value of the current chrominance block based on a value obtained by adding a predicted value according to the non-directional mode for the current chrominance block to the initial predicted value.
[0177]
[0178] When predicting color difference components by utilizing the restored luminance component values and surrounding context, the precision of the color difference prediction values can be improved, but the complexity of the prediction may also increase.
[0179] Accordingly, the present disclosure proposes a method and apparatus for predicting color difference components based on luminance components that are implemented with low complexity.
[0180]
[0181] According to one embodiment of the present disclosure, whether a luminance component-based color difference component prediction method is applied may be determined through an explicit and / or implicit method.
[0182] For example, whether to apply a luminance component-based color difference component prediction method can be determined through explicit signaling as follows.
[0183] A video encoding device can explicitly signal information indicating whether to apply a luminance component-based chrominance component prediction method at various units. Here, the various units may be one of units such as a sequence, a group of picture (GOP), a frame, a picture, a slice, a CTU, a CU, or a TU. Here, the information indicating whether to apply a luminance component-based chrominance component prediction method may be transmitted in stages. Furthermore, the information indicating whether to apply a luminance component-based chrominance component prediction method may be transmitted over multiple stages.
[0184] For example, if it is determined that a luminance component-based color difference component prediction method is used at an upper stage, whether or not to apply the luminance component-based color difference component prediction method at a lower stage may be determined dependently on the information at the upper stage. Alternatively, if it is determined that a luminance component-based color difference component prediction method is used at an upper stage, additional information regarding the luminance component-based color difference component prediction method may be conveyed at the lower stage.
[0185] For example, if it is determined by explicit information that a luminance component-based chrominance component prediction method is used at the sequence level, additional information indicating whether to apply the luminance component-based chrominance component prediction method at the lower-level CU level may be transmitted. In this case, the information indicating whether to apply the luminance component-based chrominance component prediction method at the lower-level CU level may be signaled dependently on the information indicating whether to apply the luminance component-based chrominance component prediction method at the sequence level.
[0186] The application of a luminance-based chrominance prediction method at the lowest level can be determined through a Rate-Distortion Optimization (RDO) method that simultaneously considers prediction accuracy and bit usage in the video encoding device. Through the RDO method, the video encoding device can optimize the application of the luminance-based chrominance prediction method while maximizing encoding efficiency.
[0187]
[0188] Meanwhile, according to another embodiment of the present disclosure, whether or not to apply a luminance component-based color difference component prediction method may be implicitly determined as follows.
[0189] To determine whether to apply a luminance component-based chrominance component prediction method, an image encoding device and an image decoder may use predefined information. Alternatively, the information used to determine whether to apply a luminance component-based chrominance component prediction method may be signaled in various units. Here, the information used to determine whether to apply a luminance component-based chrominance component prediction method may include information such as the width, height, number of pixels, ratio of width to height, minimum length, slice type, block type, prediction method used, motion vector value, and surrounding block information of the current block. And, the various units may be one of units such as sequence, GOP (group of picture), frame, picture, slice, CTU, CU, TU, etc.
[0190] According to one embodiment, whether to apply a luminance component-based color difference component prediction method may be implicitly determined based on the width and height of the current block.
[0191] Parameters such as Th1_width, Th1_height, Th2_width, and Th2_height may be used to implicitly determine whether to apply the luminance component-based chrominance component prediction method based on the width and height of the block. In this case, if the current block width is greater than Th1_width and less than Th2_width, and the current block height is greater than Th1_height and less than Th2_height, the luminance component-based chrominance component prediction method may be applied.
[0192] According to another embodiment, whether to apply the luminance component-based color difference component prediction method may be implicitly determined based on the number of pixels in the current block.
[0193] Parameters such as Th1_pixels and Th2_pixels may be used to implicitly determine whether to apply a luminance-based chrominance prediction method based on the number of pixels in a block. In this case, the luminance-based chrominance prediction method may be applied if the number of pixels in the current block is greater than Th1_pixels and less than Th2_pixels.
[0194] In an embodiment using threshold values, instead of using two threshold values as described above, only one threshold value may be used.
[0195] According to another embodiment, whether the luminance component-based color difference component prediction method is applied may be implicitly determined based on the prediction direction. That is, whether the luminance component-based color difference component prediction method is applied may be determined based on whether uni-prediction or bi-prediction is applied. The luminance component-based color difference component prediction method may be applied to the current block only when uni-prediction is applied to the current block. Alternatively, the luminance component-based color difference component prediction method may be applied to the current block only when bi-prediction is applied to the current block. Or, the luminance component-based color difference component prediction method may be applied to the current block when uni-prediction or bi-prediction is applied to the current block.
[0196] Alternatively, through a method similar to the embodiment described above, the luminance component-based color difference component prediction method can be implicitly determined using other parameters.
[0197]
[0198] Meanwhile, whether to apply a luminance component-based color difference component prediction method can be determined through a combination of explicit and implicit methods as follows.
[0199] According to one embodiment, information indicating whether to apply a luminance component-based chrominance component prediction method at the sequence and CU levels may be explicitly signaled. And, at lower levels, whether to apply a luminance component-based chrominance component prediction method may be implicitly determined based on the block size of the current block.
[0200] First, whether to apply a luminance component-based chrominance component prediction method can be determined through information signaled at the sequence and CU levels. Then, if the size of the current block falls within a predetermined range, the luminance component-based chrominance component prediction method may be applied to the current block. For example, the luminance component-based chrominance component prediction method may be applied only when the size of the current chrominance block is 4x4 or larger and 32x32 or smaller, and when the size of the luminance block corresponding to the current chrominance block is 8x8 or larger and 64x64 or smaller. In this case, the video encoding device may determine whether to apply the luminance component-based chrominance component prediction method by applying RDO to an arbitrarily selected block. Then, the video encoding device may transmit information indicating whether to apply the luminance component-based chrominance component prediction method to the video decoder.
[0201] Alternatively, the availability of a luminance component-based chrominance component prediction method may be determined through information signaled at the sequence and CU levels. Furthermore, if the size of the current block falls within a predetermined range, it may be determined that the luminance component-based chrominance component prediction method is available for the current block. For example, if information indicating the availability of a luminance component-based chrominance component prediction method is signaled at the sequence and CU levels, and the size of the current chrominance block is 64x64 or less, it may be determined that the luminance component-based chrominance component prediction method is available for the current block. In the process of determining the availability of the chrominance component prediction method, the partition type, slice type, etc. of the current block may be further considered. In this case, based on information indicating whether the luminance component-based chrominance component prediction method is applied, the prediction mode of the current block may be determined to be a luminance component-based chrominance component prediction method.
[0202] According to another embodiment, whether the luminance component-based color difference component prediction method is applied is explicitly signaled through information at the sequence and CU levels, and can be implicitly determined based on the number of pixels in the current block.
[0203] First, whether to apply the luminance component-based chrominance component prediction method can be determined through information signaled at the sequence level. Then, it can be determined that the luminance component-based chrominance component prediction method is applied when the number of pixels in the current block falls within a predetermined range. For example, the luminance component-based chrominance component prediction method may be applied only when the number of pixels in the current chrominance block is 32 or more. In this case, the video encoding device may determine whether to apply the luminance component-based chrominance component prediction method by applying RDO to an arbitrarily selected block. Then, the video encoding device may transmit information indicating whether to apply the luminance component-based chrominance component prediction method to the video decoder.
[0204] According to another embodiment, whether a luminance component-based color difference component prediction method is applied is explicitly signaled through information at the sequence and CU levels, and can be implicitly determined according to the prediction direction.
[0205] For example, first, whether to apply a luminance component-based chrominance component prediction method can be determined through information signaled at the sequence level. Then, a luminance component-based chrominance component prediction method can be applied to the current block only when unidirectional prediction is applied to the current block. In this case, the video encoding device can apply RDO to the unidirectionally predicted block to determine whether to apply the luminance component-based chrominance component prediction method at the CU level and transmit the determined information to the decoder.
[0206]
[0207] When a luminance component-based color difference component prediction method is used, a current color difference template including samples adjacent to the current color difference block and a luminance template including samples adjacent to a corresponding luminance block corresponding to the current color difference block may be defined. Then, using the current color difference template and the luminance template, the current color difference block can be predicted.
[0208]
[0209] Figure 7 is a diagram illustrating the template area used in the luminance component-based color difference component prediction method.
[0210] Referring to FIG. 7, the reference region may consist of samples of two or six reference sample lines located at the top and left of the prediction target unit. Meanwhile, the reference region may include an integer number of reference sample lines adjacent to the prediction target unit. For example, the reference region may include one, three, or more reference sample lines.
[0211] In CCCM forecasting, the number of reference sample lines used to derive CCCM model parameters can be determined by the template cost of the forecast based on the number of reference sample lines. For example, in CCCM forecasting, whether to use six lines or two lines of adjacent samples can be determined by the template cost of each reference region.
[0212] The template cost can be derived by calculating the SATD between the predicted value of the reference region and the restored value of the reference region, which is generated using two or six reference sample lines adjacent to the unit to be predicted.
[0213] In addition, a padding region may be defined that includes one sample line adjacent to the reference region and one sample line adjacent to the right and bottom of the prediction target unit. The padding region may be adjusted to include only available reference samples. The samples in the padding region may be used in the process of applying a cross-shaped spatial filter. Meanwhile, if the samples in the padding region are not available, they may be replaced with samples from the adjacent reference region.
[0214]
[0215] To reduce the complexity of the luminance component-based color difference component prediction method, the template area can be reduced.
[0216] Below, we describe the reduced template area used in the low-complexity luminance component-based chrominance component prediction method.
[0217]
[0218] FIG. 8 is a diagram illustrating a template area used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0219] Referring to Fig. 8, in a luminance component-based color difference component prediction method, the template regions described below can be used.
[0220] For example, if samples adjacent to the left of the current block and samples adjacent to the top of the current block are all available, the template area may include samples adjacent to the left of the current block and samples adjacent to the top of the current block. A template area including samples adjacent to the left of the current block and samples adjacent to the top of the current block may be referred to as an L-shaped template area.
[0221] If samples adjacent to the left of the current block and samples adjacent to the top of the current block are all available, according to the luminance component-based color difference component prediction method, a prediction block for the color difference component can be generated by utilizing both the pixels of the left template area and the top template area.
[0222] Alternatively, if samples adjacent to the left of the current block are not available, the template area may include samples adjacent to the top of the current block. The template area including samples adjacent to the top of the current block may be referred to as the top template area.
[0223] On the other hand, if samples from the left template area are not available, according to the luminance component-based color difference component prediction method, pixels from the top template area can be utilized to generate prediction blocks for color difference components.
[0224] Meanwhile, if samples adjacent to the top of the current block are not available, the template area may include samples adjacent to the left of the current block. The template area including samples adjacent to the left of the current block may be referred to as the left template area.
[0225] Alternatively, if samples from the upper template area are not available, according to the luminance component-based color difference component prediction method, pixels from the left template area can be utilized to generate prediction blocks for color difference components.
[0226] Among the template regions described above, the template region used can be determined based on the coding parameters of the current block. For example, the coding parameters may include information such as the block's prediction information, block size, QP value, slice type, MPM value, the aspect ratio of the current block, intra / inter prediction mode information used in surrounding blocks, the current block's MRL indicator, the reference value of the block vector, and the size value of the corresponding L type.
[0227]
[0228] Alternatively, in a low-complexity luminance component-based chrominance component prediction method, samples from the template region described below may be adaptively used. Here, at least some of the samples from the template region may be used in the low-complexity luminance component-based chrominance component prediction method.
[0229]
[0230] FIG. 9 is a diagram illustrating samples of template regions used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0231] Referring to Fig. 9, all samples of the template area or some samples of the template area can be used in the luminance component-based color difference component prediction method.
[0232] For example, all samples in the template area can be used in a low-complexity luminance component-based chrominance component prediction method.
[0233] Alternatively, subsampled samples from the template region can be used in a low-complexity luminance component-based chrominance component prediction method. For example, the subsampling ratio can be 2:1. Alternatively, the subsampling ratio can be determined according to the color format. Furthermore, the subsampling ratio applied to the samples in the template region may differ for each reference line in the template region. For example, different subsampling filters may be applied alternately to the reference sample lines.
[0234] Alternatively, among the sample lines in the template region, only samples from a specific sample line may be used in the low-complexity luminance component-based chrominance component prediction method. For example, among the sample lines in the template region, the low-complexity luminance component-based chrominance component prediction method can be performed using only samples from the second sample line.
[0235] Alternatively, subsampled samples of the template region and samples of a specific sample line may be used in a low-complexity luminance component-based chrominance component prediction method. For example, among the sample lines of the template region, the low-complexity luminance component-based chrominance component prediction method may be performed using only the subsampled samples of the first sample lines and the samples of the second sample line.
[0236]
[0237] The number of reference samples used in the low-complexity luminance component-based chrominance component prediction method can be 2n. Therefore, multiplication and / or shift operations can be used in the low-complexity luminance component-based chrominance component prediction method.
[0238] Subsampled samples of the template area can be determined based on the coding parameters of the current block. For example, the coding parameters may include information such as the block's prediction information, block size, QP value, slice type, MPM value, the aspect ratio of the current block, the MRL indicator of the current block, intra / inter prediction mode information used in neighboring blocks, availability of neighboring blocks, reference values of block vectors, and the size values of the corresponding L type.
[0239]
[0240] According to another embodiment, the template area used in the luminance component-based color difference component prediction method can be adaptively set as follows based on the processing unit.
[0241]
[0242] FIG. 10 is a diagram illustrating a template area used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0243] Referring to FIG. 10, the template area used in the luminance component-based color difference component prediction method can be adaptively changed considering the processing unit in the hardware. For example, the decoding process at the VDPU (Virtual Decode Processing Unit) level can be implemented on the hardware based on a 64x64 decoder pipeline. In this case, in order to process the luminance component-based color difference component prediction process within the processing unit, the template area used in the luminance component-based color difference component prediction method can be adaptively changed. Thus, the delay of the luminance component-based color difference component prediction process can be minimized.
[0244] For example, when applying a luminance component-based chrominance component prediction method to a block of size 64x128, the required size of the left template area may be width x 128. However, if the bottom-left portion is not decoded during the hardware pipeline process, the values of some samples in the left template area may not be restored. Therefore, the template size can be adaptively adjusted to width x 64. Thus, the delay in the luminance component-based chrominance component prediction process can be minimized.
[0245] According to another example, when using a luminance component-based chrominance component prediction method on a block of size 64x128 or 32x128, the required left template area size may be width x 128. The unrecovered left area may be excluded from the template area. Therefore, the template area may ultimately be set to width x 64. In this case, the width of the template area may be determined according to the processing unit.
[0246]
[0247] Alternatively, the template area used in the luminance component-based color difference component prediction method can be set through information explicitly signaled as follows.
[0248] For example, the video encoding device may select one template region from among the candidates included in the candidate group of template regions. Then, the video encoding device may transmit information indicating the selected template region to a decoder. The information indicating the template region may be explicitly signaled through various units. Here, the various units may be one of units such as sequence, GOP (group of picture), frame, picture, slice, CTU, CU, TU, etc. Specifically, the video encoding device may signal information indicating one template region among the top-left template region, the left template region, and the top template region through a predetermined unit.
[0249] Here, the order of candidates included in the candidate pool of the template area can be determined based on information regarding blocks adjacent to the current block. For example, the order of the candidate pool of the template area can be set by default to the order of the L-shaped template area, the left template area, and the top template area. Meanwhile, if the template area used for predicting blocks adjacent to the current block is the top template area, the order of the candidate pool of the template area can be set to the order of the top template area, the L-shaped template area, and the left template area. By changing the order of the candidate pool based on information regarding blocks adjacent to the current block and encoding based on the changed order, a bit reduction effect can be achieved through variable-length coding.
[0250] Alternatively, the candidate group of template regions may include only some candidates. For example, if the template regions used for predicting blocks adjacent to the current block are the left template region and the top template region, respectively, the candidate group of template regions may include only the left template region and the top template region. Furthermore, the video encoding device may transmit information indicating one of the template regions among the left template region and the top template region to the video decoder.
[0251] Alternatively, if the template regions used for predicting blocks adjacent to the current block are each left template regions, the candidate template regions may include left template regions and L-shaped template regions. Furthermore, the image encoding device may transmit information indicating one of the left template regions and L-shaped template regions to the image decoder.
[0252]
[0253] Alternatively, samples of the template region used in the luminance component-based color difference component prediction method can be set through information that is explicitly signaled as follows.
[0254] For example, the video encoding device may select samples of a specific template region from among the candidates included in the candidate group of samples of the template region being used. Then, the video encoding device may transmit information indicating the samples of the specific template region to a decoder. The information indicating the samples of the specific template region may be explicitly signaled through various units. Here, the various units may be one of units such as sequence, group of picture (GOP), frame, picture, slice, CTU, CU, TU, etc. Specifically, the video encoding device may signal information indicating samples of a specific template region through a predetermined unit from among candidates including all reference samples, candidates including downsampled samples of the template region, and candidates including samples of a partial reference sample line of the template region.
[0255] Here, the order of candidates included in the candidate pool of samples in the template area can be determined based on information from blocks adjacent to the current block. By default, the order of candidates included in the candidate pool of samples in the template area can be set in the order of template candidates containing all reference samples, template candidates containing downsampled samples of the template area, and template candidates containing samples from partial reference sample lines of the template area. Meanwhile, if the samples in the template area used for predicting blocks adjacent to the current block are samples from partial reference sample lines of the template area, the order of the candidate pool in the template area can be set in the order of candidates containing samples from partial reference sample lines of the template area, candidates containing all reference samples, and candidates containing downsampled samples of the template area. By changing the order of the candidate pool based on information from blocks adjacent to the current block and encoding based on the changed order, a bit reduction effect can be achieved through variable-length coding.
[0256] Alternatively, the candidate pool of samples in the template area may include only some candidates. For example, if the samples in the template area used for predicting blocks adjacent to the current block are, respectively, downsampled samples of the template area and samples of a partial reference sample line of the template area, the candidate pool of samples in the template area may include only candidates including downsampled samples of the template area and candidates including samples of a partial reference sample line of the template area. And, the image encoding device may transmit information indicating a specific candidate among the candidates included in the candidate pool of samples in the template area to the image decoder.
[0257] Alternatively, if the samples in the template region used for predicting blocks adjacent to the current block are downsampled samples of the template region, the candidate group of samples in the template region may include candidates including all reference samples having a relatively high priority and candidates including downsampled samples of the template region. And, the image encoding device may transmit information indicating a specific candidate among the candidates included in the candidate group of samples in the template region to the image decoder.
[0258]
[0259] Alternatively, the template area used in the luminance component-based color difference component prediction method can be set through information explicitly signaled as follows.
[0260] When adaptively setting the template area considering the processing unit in the hardware, the video encoding device can transmit information about the reduced left template area and the top template area to the decoder. Information about the reduced left template area and the top template area can be explicitly signaled through various units. Here, the various units may be one of units such as sequence, group of picture (GOP), frame, picture, slice, CTU, CU, TU, etc.
[0261] For example, information regarding the collapsed left template area may be information for configuring the template area, such as width and height values. Additionally, information regarding the collapsed top template area may be information for configuring the template area, such as width and height values.
[0262] If at least part of the width or height of the collapsed template area is implicitly determined, information indicating the element that is not implicitly determined may be explicitly transmitted. For example, if the width value of the collapsed left template is fixed at 3 and the height value of the collapsed top template is fixed at 3, information indicating the height value of the collapsed left template and the width value of the collapsed top template may be explicitly transmitted.
[0263]
[0264] To apply a luminance component-based color difference component prediction method, it is necessary to set the number of pixels for the luminance component and the number of pixels for the color difference component to be equal. Therefore, downsampling can be applied to the values of samples in the luminance block at the same location as the current color difference block, as well as to samples adjacent to the luminance block. Here, the downsampling ratio may vary depending on the color format.
[0265] For example, if the color format is 4:4:4, downsampling may not be applied to the luminance samples. Meanwhile, if the color format is 4:2:2, the downsampling ratio applied to the luminance samples may be 2:1. Also, if the color format is 4:2:0 or 4:1:1, the downsampling ratio applied to the luminance samples may be 4:1.
[0266]
[0267] Below, we explain the downsampling methods applied to luminance samples according to color components.
[0268] FIG. 11 is a drawing illustrating a downsampling filter used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0269] Referring to Fig. 11, when the color format is 4:2:2, the luminance component sample can be downsampled using a filter as shown below.
[0270] According to one embodiment, a [m, n] filter may be applied to the values of two luminance samples. Here, the values of the downsampling filter coefficients m and n may be arbitrary integers set through various methods. For example, the values of the downsampling filter coefficients may be preset values. Alternatively, the values of the downsampling filter coefficients may be set based on signaling information. Alternatively, the values of the downsampling filter coefficients may be set based on the coding parameters of the current block. For example, a [1, 1] filter may be applied to the values of two luminance samples to downsample to the average value of the two luminance sample values.
[0271] According to another embodiment, among the values of two luminance samples, the value of the luminance sample at a specific location may be downsampled. For example, a [1, 0] filter may be applied to the values of two luminance samples to downsample them to the value of one luminance sample.
[0272]
[0273] FIG. 12 is a drawing illustrating a downsampling filter used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0274] Referring to Fig. 12, when the color format is 4:2:0 or 4:1:1, a luminance component sample can be downsampled using a filter as shown below.
[0275] According to one embodiment, a [[m, n], [k, p]] filter may be applied to the values of four luminance samples. Here, the values of the downsampling filter coefficients m, n, k, and p may be arbitrary integers set through various methods. For example, the values of the downsampling filter coefficients may be preset values. Alternatively, the values of the downsampling filter coefficients may be set based on signaling information. Alternatively, the values of the downsampling filter coefficients may be set based on the coding parameters of the current block. For example, a [[1, 1], [1, 1]] filter may be applied to the values of four luminance samples to downsample to the average value of the four luminance sample values.
[0276] According to another embodiment, among the values of four luminance samples, the values can be downsampled to the value of a luminance sample at a specific location. For example, a [[1, 0], [0, 0]] filter can be applied to the values of four luminance samples to downsample them to the value of a single luminance sample.
[0277]
[0278] FIG. 13 is a drawing illustrating a downsampling filter used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0279] Referring to FIG. 13, a [[m, n, k], [p, q, r]] filter can be applied to the values of six luminance samples. Here, the values of the downsampling filter coefficients m, n, k, p, q, and r can be arbitrary integers set through various methods. For example, the values of the downsampling filter coefficients can be preset values. Or, the values of the downsampling filter coefficients can be set based on signaling information. Or, the values of the downsampling filter coefficients can be set based on the coding parameters of the current block. For example, a [[1, 2, 1], [1, 2, 1]] filter can be applied to the values of six luminance samples. Or, a [[1, 4, 1], [1, 4, 1]] filter can be applied to the values of six luminance samples. Or, a [[1, 0, -1], [1, 0, -1]] filter can be applied to the values of six luminance samples. Alternatively, a [[1, 2, 1], [-1, -2, -1]] filter may be applied to the values of the six luminance samples. Alternatively, a [[-1, 1, 2], [-2, -1, 1]] filter may be applied to the values of the six luminance samples. Thus, the luminance samples may be downsampled to the average value or weighted average value of the six luminance sample values.
[0280]
[0281] FIG. 14 is a drawing illustrating a downsampling filter used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0282] Referring to FIG. 14, a [[m, n, k], [p, q, r], [x, y, z]] filter can be applied to the values of nine luminance samples. Here, the values of the downsampling filter coefficients m, n, k, p, q, r, x, y, and z can be arbitrary integers set through various methods. For example, the values of the downsampling filter coefficients can be preset values. Alternatively, the values of the downsampling filter coefficients can be set based on signaling information. Alternatively, the values of the downsampling filter coefficients can be set based on the coding parameters of the current block. For example, a [[1, 1, 1], [1, 8, 1], [1, 1, 1]] filter can be applied to the values of nine luminance samples.
[0283] Alternatively, a filter may be applied to only some of the values of the nine luminance samples. In this case, the [[0, m, 0], [n, k, p], [0, q, 0]] filter may be applied to the values of the luminance samples. For example, the [[0, 1, 0], [1, 4, 1], [0, 1, 0]] filter may be applied to the values of the luminance samples.
[0284] In addition, downsampled luminance samples can be derived using downsampling filters that use different samples depending on the color format. For example, if the color format is 4:2:0, one of the downsampling filters [[0, 0, 0], [0, 2, 2], [0, 2, 2]], [[0, 0, 0], [1, 2, 1], [1, 2, 1]], [[0, 1, 0], [1, 4, 1], [0, 1, 0]] can be used. Or, if the color format is 4:2:2, the downsampling filters [[0, 4, 4], [2, 4, 2], [0, 8, 0]] can be used.
[0285]
[0286] The downsampling method applied to luminance samples can significantly affect the prediction performance of color difference components. Therefore, the downsampling method can be adaptively determined through an implicit method and applied to luminance samples.
[0287] According to one embodiment, the downsampling method may be determined implicitly based on the current block size. For example, as the current block size increases, the downsampling method applied to the luminance sample may be set to use a large value filter tap. On the other hand, as the block size decreases, the downsampling method applied to the luminance component may be set to use a small value filter tap.
[0288] According to another embodiment, the downsampling method may be implicitly determined based on the shape of the current block. For example, if the current block is a non-square block, the downsampling method may be configured to use a filter of a shape corresponding to the shape of the current block. For example, if the width of the block is greater than its height, the shape of the filter used for downsampling may be a shape that is wide to the left.
[0289] According to another embodiment, the downsampling method can be configured according to the slice type containing the current block. Depending on whether the slice type is an intra-slice or an inter-slice, the characteristics of the reconstructed pixel may differ. Accordingly, the downsampling method can be adaptedly configured according to the slice type.
[0290] According to another embodiment, the downsampling method may be determined based on the coding parameters of the current block. For example, the coding parameters of the current block may include the mode type of the current block, prediction information, BS value, segmentation information, information related to chrominance components, etc. Here, the information related to the chrominance components of the current block may include video color format information of the picture of the current block, information related to the positional assignment between the samples of the current block and the luminance block corresponding to the current block of the chrominance components.
[0291] According to another embodiment, the downsampling method can be set based on the QP value of the current block. As the QP value decreases, the quality of the reconstructed pixels is similar to the original, whereas as the QP value increases, the quality may degrade. Therefore, if the downsampling method is adaptively set to correspond to the QP value, prediction accuracy can be improved. For example, when the QP is low, the downsampling method applied to the luminance component can be set to use a filter tab with a large value. On the other hand, when the QP is high, the downsampling method applied to the luminance component can be set to use a filter tab with a small value.
[0292] According to another embodiment, the downsampling method may be set based on information of blocks adjacent to the left and top of the current block. For example, the downsampling method may be adaptively determined by utilizing coding parameters such as QP values, mode types, prediction information, BS values, and splitting information of blocks adjacent to the left and top of the current block.
[0293] The effect of downsampling can be maximized when the downsampling method for the luminance component and the sampling method for the chrominance component are the same. Therefore, if the sampling method for the chrominance component is pre-configured, the number of filter taps and the number of reference pixels used for downsampling the luminance component can be explicitly set. In this case, the location of the downsampled luminance sample can coincide with the location of the chrominance sample.
[0294]
[0295] Alternatively, the filter used for downsampling the luminance component can be explicitly determined and applied to the luminance component.
[0296] For example, if the color format is 4:2:2, a candidate group including candidates for filters applied to luminance samples may be formed. Then, the image encoding device may determine one filter candidate from among the filter candidates included in the candidate group and transmit information indicating the determined filter candidate to the decoder. At this time, the information indicating the determined filter candidate may be transmitted in units such as sequence, group of picture (GOP), frame, picture, slice, CTU, CU, TU, etc.
[0297] Alternatively, if the color format is 4:2:0 or 4:1:1, a candidate group including candidates for filters applied to luminance samples may be formed. Then, the image encoding device may determine one filter candidate among the filter candidates included in the candidate group and transmit information indicating the determined filter candidate to the decoder. At this time, the information indicating the determined filter candidate may be transmitted in units such as sequence, group of picture (GOP), frame, picture, slice, CTU, CU, TU, etc.
[0298] When the downsampling method for the luminance component and the sampling method for the chrominance component are the same, the effect of downsampling can be maximized. Therefore, when the sampling method for the chrominance component is pre-set, the number of taps of the filter used for downsampling the luminance component and the number of reference pixels can be explicitly set. In this case, multiple sampling methods for the chrominance component are formed into a candidate group, and after selecting the optimal method in the image encoding device, the information can be transmitted to the decoder in units such as sequence, group of picture (GOP), frame, picture, slice, CTU, CU, TU, etc.
[0299]
[0300] In a luminance component-based chrominance component prediction method, a filter can be applied to the reconstructed value of a luminance sample to generate a predicted value of the chrominance component. The filter applied to the reconstructed value of the luminance sample to generate the predicted value of the chrominance component can be referred to as a convolution filter.
[0301] The shape of the convolutional filter can significantly affect the prediction performance of the chrominance component. Therefore, the shape of the convolutional filter can be adaptively determined through an implicit method and applied to the luminance component.
[0302] Below, we describe the form of the convolutional filter applied to the restored value of the luminance sample in the luminance component-based chrominance component prediction method.
[0303]
[0304] FIG. 15 is a diagram illustrating the shape of a convolutional filter used in a luminance component-based color difference component prediction method according to one embodiment of the present disclosure.
[0305] Referring to FIG. 15, the convolutional filter may have a form generated through a combination of various coefficients and components based on 3x3. For example, the convolutional filter may have coefficients c1 through c9, each coefficient may be associated with the components of the luminance sample corresponding to the chrominance sample. Here, coefficient c5 may be associated with the luminance sample corresponding to the prediction target sample of the current chrominance block.
[0306] Specifically, to reduce the complexity of the luminance component-based color difference component prediction method, the convolutional filter can be configured to use five or fewer components.
[0307] For example, the convolutional filter may be configured to use the target sample and samples adjacent to the target sample in the horizontal and vertical directions. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the target sample in the diagonal direction. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the target sample in the horizontal direction. The convolutional filter may be configured to use the target sample and samples adjacent to the target sample in the vertical direction. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the left, top, and top-left of the target sample. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the right, top, and top-right of the target sample. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the left, bottom, and bottom-left of the target sample. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the right, bottom, and bottom-right of the target sample.
[0308] Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the left of the target sample. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the right of the target sample. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the top of the target sample. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the bottom of the target sample.
[0309] Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the bottom-left of the target sample. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the top-right of the target sample. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the top-left of the target sample. Alternatively, the convolutional filter may be configured to use the target sample and samples adjacent to the bottom-right of the target sample.
[0310] Alternatively, the convolutional filter may be configured to use one of the samples adjacent to the target sample or the target sample. For example, the convolutional filter may be configured to use the target sample. Alternatively, the convolutional filter may be configured to use the sample adjacent to the left of the target sample. Alternatively, the convolutional filter may be configured to use the sample adjacent to the right of the target sample. Alternatively, the convolutional filter may be configured to use the sample adjacent to the top of the target sample. Alternatively, the convolutional filter may be configured to use the sample adjacent to the bottom of the target sample.
[0311]
[0312] According to one embodiment, a convolutional filter can be implemented through a linear combination of the components used.
[0313] On the other hand, according to another embodiment, the convolutional filter can be implemented through a linear or non-linear combination of the components used. That is, the convolutional filter may include additional non-linear elements. If additional non-linear elements are included, the convolutional filter may include additional components that are not generated by a linear combination of surrounding samples. Therefore, the convolutional filter can generate predicted values without being limited to a linear model. The non-linear elements used in the convolutional filter are described below.
[0314]
[0315] [Mathematical Formula 1]
[0316] P = (C 2 +midVal)>>bitDepth
[0317] Here, P may indicate a non-linear element. C may indicate a luminance sample corresponding to the target sample for prediction of the current chrominance block. And, bitDepth may indicate a bit depth.
[0318] And, according to one embodiment, midVal may indicate the middle value of a range of sample values. According to another embodiment, midVal may indicate the middle value of the samples constituting the current block. According to yet another embodiment, midVal may indicate the middle value of the samples used in a convolutional filter.
[0319]
[0320] [Mathematical Formula 2]
[0321] P = (C 2 +avgVal)>>bitDepth
[0322] Here, P may indicate a non-linear element. C may indicate a luminance sample corresponding to the target sample for prediction of the current chrominance block. And, bitDepth may indicate a bit depth.
[0323] And, according to one embodiment, avgVal may indicate the average value of the samples constituting the current block. According to another embodiment, midVal may indicate the average value of the samples used in the convolutional filter.
[0324]
[0325] [Mathematical Formula 3]
[0326] P = midVal
[0327] Here, according to one embodiment, midVal may indicate the middle value of a range of sample values. According to another embodiment, midVal may indicate the middle value of the samples constituting the current block. According to yet another embodiment, midVal may indicate the middle value of the samples used in a convolutional filter.
[0328]
[0329] [Mathematical Formula 4]
[0330] P = avgVal
[0331] Here, according to one embodiment, avgVal may indicate the average value of the samples constituting the current block. According to another embodiment, midVal may indicate the average value of the samples used in the convolutional filter.
[0332]
[0333] In a method for predicting color difference components based on luminance components, a convolutional filter can be configured to include additional non-linear elements in a linear filter. Here, the convolutional filter may include one or more non-linear elements.
[0334] Below, a method for predicting chrominance components based on luminance components using a convolutional filter containing linear and non-linear elements is described.
[0335]
[0336] FIG. 16 is a diagram illustrating a method for predicting color difference components based on luminance components using a convolutional filter according to one embodiment of the present disclosure.
[0337] Referring to FIG. 16, the convolutional filter can be defined by the coefficients and components of the convolutional filter. Here, C1 to C11 indicate the coefficients of the convolutional filter, and A1 to A11 indicate luminance samples used in the convolutional filter. Here, the luminance samples may be downsampled samples.
[0338] According to one embodiment, the convolutional sample may be configured to use non-linear elements with the target sample and samples adjacent to the target sample in the horizontal direction. The value of the color difference sample derived using the convolutional filter can be expressed as shown in the following mathematical formula.
[0339] [Mathematical Formula 5]
[0340] Pred chroma =C4*A4+C5*A5+C6*A6+C 10 *A 10 +C 11 *A 11
[0341] A 10 =(A5*A5+midVal)>>Bitdepth
[0342] A 11 =midVal
[0343] Here, midVal can indicate the median value of a pixel range.
[0344]
[0345] FIG. 17 is a diagram illustrating a method for predicting color difference components based on luminance components using a convolutional filter according to one embodiment of the present disclosure.
[0346] Referring to FIG. 17, the convolutional filter can be defined by the coefficients and components of the convolutional filter. Here, C1 to C11 indicate the coefficients of the convolutional filter, and A1 to A11 indicate luminance samples used in the convolutional filter. Here, the luminance samples may be downsampled samples.
[0347] According to one embodiment, the convolutional sample may be configured to use non-linear elements with the target sample and samples adjacent to the target sample in the vertical direction. The value of the color difference sample derived using the convolutional filter can be expressed as shown in the following mathematical formula.
[0348]
[0349] [Mathematical Formula 6]
[0350] Pred chroma =C2*A2+C5*A5+C7*A7+C 10 *A 10 +C 11 *A 11
[0351] A 10 =(A5*A5+midVal)>>Bitdepth
[0352] A 11 =midVal
[0353] Here, midVal can indicate the median value of a pixel range.
[0354]
[0355] FIG. 18 is a diagram illustrating a method for predicting color difference components based on luminance components using a convolutional filter according to one embodiment of the present disclosure.
[0356] Referring to FIG. 18, the convolutional filter can be defined by the coefficients and components of the convolutional filter. Here, C1 to C11 indicate the coefficients of the convolutional filter, and A1 to A11 indicate luminance samples used in the convolutional filter. Here, the luminance samples may be downsampled samples.
[0357] According to one embodiment, the convolutional sample may be configured to use non-linear elements with the target sample and samples adjacent to the target sample in the vertical direction. The value of the color difference sample derived using the convolutional filter can be expressed as shown in the following mathematical formula.
[0358]
[0359] [Mathematical Formula 7]
[0360] Pred chroma =C5*A5+C 10 *A 10 +C 11 *A 11
[0361] A 10 =(A5*A5+midVal)>>Bitdepth
[0362] A 11 =midVal
[0363] Here, midVal can indicate the median value of a pixel range.
[0364]
[0365] FIG. 19 is a diagram illustrating a method for predicting color difference components based on luminance components using a convolutional filter according to one embodiment of the present disclosure.
[0366] Referring to FIG. 19, the convolutional filter can be defined by the coefficients and components of the convolutional filter. Here, C1 to C11 indicate the coefficients of the convolutional filter, and A1 to A11 indicate luminance samples used in the convolutional filter. Here, the luminance samples may be downsampled samples.
[0367] According to one embodiment, the convolutional sample may be configured to use non-linear elements with the target sample and samples adjacent to the target sample in the vertical direction. The value of the color difference sample derived using the convolutional filter can be expressed as shown in the following mathematical formula.
[0368]
[0369] [Mathematical Formula 8]
[0370] Pred chroma =C4*A4+C 10 *A 10 +C 11 *A 11
[0371] A 10 =(A5*A5+midVal)>>Bitdepth
[0372] A11 =midVal
[0373] Here, midVal can indicate the median value of a pixel range.
[0374]
[0375] FIG. 20 is a diagram illustrating a method for predicting color difference components based on luminance components using a convolutional filter according to one embodiment of the present disclosure.
[0376] Referring to FIG. 20, the convolutional filter can be defined by the coefficients and components of the convolutional filter. Here, C1 to C11 indicate the coefficients of the convolutional filter, and A1 to A11 indicate luminance samples used in the convolutional filter. Here, the luminance samples may be downsampled samples.
[0377] According to one embodiment, the convolutional sample may be configured to use non-linear elements with the target sample and samples adjacent to the target sample in the vertical direction. The value of the color difference sample derived using the convolutional filter can be expressed as shown in the following mathematical formula.
[0378]
[0379] [Mathematical Formula 9]
[0380] Pred chroma =C2*A2+C 10 *A 10 +C 11 *A 11
[0381] A 10 =(A5*A5+midVal)>>Bitdepth
[0382] A 11 =midVal
[0383] Here, midVal can indicate the median value of a pixel range.
[0384]
[0385] Meanwhile, according to another embodiment, different downsampling filters can be applied to samples in the restored luminance block and luminance template area to derive downsampled luminance samples. For example, A1 to A9 in Equations 5 to 6 may be downsampled luminance samples derived by applying different downsampling filters to samples in the restored luminance block and luminance template area. Furthermore, a predicted value for the chrominance block can be generated by applying a convolutional filter to the luminance samples derived as a result of applying different downsampling filters.
[0386]
[0387] The shape of the convolutional filter can significantly affect the prediction performance for chrominance components. Therefore, the shape of the convolutional filter applied to the luminance sample can be implicitly determined.
[0388] According to one embodiment, the shape of the convolutional filter may be implicitly determined based on the size of the current block. For example, as the block size increases, the convolutional filter may be configured to use a larger number of luminance samples. On the other hand, as the block size decreases, the convolutional filter may be configured to use a smaller number of luminance samples.
[0389] According to another embodiment, the shape of the convolutional filter may be implicitly determined based on the shape of the current block. For example, if the current block is a non-square block, the shape of the convolutional filter may have a shape corresponding to the shape of the current block. For example, if the width of the block is greater than its height, the shape of the convolutional filter may be wide to the left.
[0390] According to another embodiment, the shape of the convolutional filter can be set according to the slice type containing the current block. Depending on whether the slice type is an intra-slice or an inter-slice, the characteristics of the reconstructed pixel may differ. Accordingly, the shape of the convolutional filter can be adaptively set according to the slice type.
[0391] According to another embodiment, the shape of the convolutional filter can be set according to the QP value of the current block. As the QP value decreases, the quality of the reconstructed pixels is similar to the original, whereas as the QP value increases, the quality may degrade. Therefore, by adaptively setting the convolutional filter to correspond to the QP value, prediction accuracy can be improved. For example, when the QP is low, the convolutional filter can be set to use a filter tab with a large value. On the other hand, when the QP is high, the convolutional filter can be set to use a filter tab with a small value.
[0392] According to another embodiment, the shape of the convolutional filter can be set based on information from blocks adjacent to the left and top of the current block. For example, the shape of the convolutional filter can be adaptively determined by utilizing coding parameters such as QP values, mode types, prediction information, BS values, and splitting information from blocks adjacent to the left and top of the current block.
[0393] According to another embodiment, the shape of the convolutional filter may be set based on the location of a block or pixel. For example, the shape of the convolutional filter may be adaptively determined based on the location of a block or pixel within a predetermined unit. Here, the predetermined unit may be a unit such as a CU, CTU, slice, picture, etc.
[0394]
[0395] The nonlinear elements of a convolutional filter can significantly affect the prediction performance of the chrominance component. Therefore, the type and number of nonlinear elements of the convolutional filter can be adaptively determined through implicit methods and applied to the luminance component.
[0396] According to one embodiment, the type and number of nonlinear elements of a convolutional filter may be implicitly determined based on the current block size. For example, as the block size increases, the filter may be configured to use a greater number of nonlinear elements. On the other hand, as the block size decreases, the filter may be configured to use a smaller number of nonlinear elements.
[0397] According to another embodiment, the type and number of nonlinear elements of the convolutional filter can be determined implicitly based on the shape of the current block.
[0398] According to another embodiment, the type and number of nonlinear elements of the convolutional filter can be set according to the slice type containing the current block. Depending on whether the slice type is an intra-slice or an inter-slice, the characteristics of the reconstructed pixels may differ. Accordingly, the type and number of nonlinear elements of the convolutional filter can be adaptively set according to the slice type.
[0399] According to another embodiment, the type and number of nonlinear elements of the convolutional filter can be set according to the QP value of the current block. As the QP value is low, the quality of the reconstructed pixels is similar to the original, whereas as the QP value is high, the quality may degrade. Therefore, prediction accuracy can be improved by adaptively setting the downsampling method according to the QP value. For example, when the QP is low, the filter can be set to use fewer nonlinear elements. On the other hand, when the QP is high, the filter can be set to use more nonlinear elements.
[0400] According to another embodiment, the type and number of non-linear elements of a convolutional filter can be set based on information from blocks adjacent to the left and top of the current block. For example, the type and number of non-linear elements of a convolutional filter can be adaptively determined by utilizing coding parameters such as QP values, mode types, prediction information, BS values, and splitting information from blocks adjacent to the left and top of the current block.
[0401] According to another embodiment, the type and number of nonlinear elements of a convolutional filter may be set based on the location of a block or pixel. For example, the type and number of nonlinear elements of a convolutional filter may be adaptively determined based on the location of a block or pixel within a predetermined unit. Here, the predetermined unit may be a unit such as a CU, CTU, slice, picture, etc.
[0402]
[0403] The shape of the convolutional filter can significantly affect the prediction performance of chrominance components. Therefore, the shape of the convolutional filter applied to the luminance sample can be determined based on explicitly signaled information.
[0404] For example, among various types of filters, a candidate group including at least some of the filters may be established. Then, the image encoding device may determine one filter candidate among the filter candidates included in the candidate group and transmit information indicating the determined filter candidate to the image decoder. At this time, the information indicating the determined filter candidate may be transmitted in units such as sequence, GOP (group of picture), frame, picture, slice, CTU, CU, TU, etc. For example, if the candidate group includes vertical filters and horizontal filters, the image encoding device may determine the optimal filter among the vertical filters and horizontal filters by considering RD performance. Then, the image encoding device may transmit information indicating the optimal filter to the image decoder.
[0405]
[0406] The nonlinear elements of a convolutional filter can significantly affect the prediction performance of chrominance components. Therefore, the nonlinear elements of the convolutional filter applied to luminance samples can be determined based on explicitly signaled information.
[0407] For example, among filters containing different types and different numbers of non-linear elements, a candidate group including at least some of the filters may be established. Then, the image encoding device may determine one filter candidate among the filter candidates included in the candidate group and transmit information indicating the determined filter candidate to a decoder. At this time, the information indicating the determined filter candidate may be transmitted in units such as sequence, GOP (group of picture), frame, picture, slice, CTU, CU, TU, etc. For example, if the filters constituting the candidate group contain different non-linear elements, the image encoding device may determine the optimal filter among the filters constituting the candidate group by considering RD performance. Then, the image encoding device may transmit information indicating the optimal filter to an image decoder.
[0408]
[0409] Alternatively, the type of convolutional filter applied to the luminance sample and the combination of nonlinear elements can be determined based on explicitly signaled information.
[0410] For example, filter candidates constituting the candidate group may have different forms and include different non-linear elements. Then, the video encoding device may determine one filter candidate among the filter candidates included in the candidate group and transmit information indicating the determined filter candidate to the decoder. At this time, the information indicating the determined filter candidate may be transmitted in units such as sequence, GOP (group of picture), frame, picture, slice, CTU, CU, TU, etc. For example, if the filters constituting the candidate group include different non-linear elements, the video encoding device may determine the optimal filter among the filters constituting the candidate group by considering RD performance. Then, the video encoding device may transmit information indicating the optimal filter to the decoder.
[0411]
[0412] Meanwhile, the shape of the convolutional filter applied to the luminance sample and the combination of nonlinear elements can be determined through a combination of explicit and implicit methods as follows.
[0413] For example, at the sequence and CU levels, information indicating whether various filter candidates are available in luminance component-based chrominance component prediction may be signaled. And, at a lower level, if the size of the current block falls within a predetermined range, information indicating whether to apply luminance component-based chrominance component prediction using various filter candidates may be signaled. Here, information indicating whether various filter candidates are available in luminance component-based chrominance component prediction may be signaled dependently on information indicating the availability of luminance component-based chrominance component prediction.
[0414] For example, if there is information indicating whether luminance component-based chrominance component prediction is available, and the current block size is greater than 4x4 and less than or equal to 32x32, information indicating whether various filter candidates are available in luminance component-based chrominance component prediction may be signaled.
[0415] When information indicating whether to apply luminance component-based chrominance component prediction using various filter candidates is signaled, an index indicating the type of luminance component-based chrominance component prediction may be defined. Additionally, among the various filter candidates supported by luminance component-based chrominance component prediction, information indicating one filter candidate may be signaled.
[0416]
[0417] In a luminance component-based color difference component prediction method, the coefficients of a convolutional filter can be derived through correlation analysis between a color difference template region adjacent to the current color difference block and a luminance template region adjacent to the corresponding luminance block. Specifically, a color difference prediction value can be derived by applying a filter to the reconstructed value of the luminance template region. Furthermore, to optimize the coefficients of the convolutional filter, the color difference prediction value can be optimized to be similar to the reconstructed value of the color difference template. Various correlation analysis methods can be used to derive the filter coefficients; for example, the following methods may be utilized.
[0418] According to one embodiment, filter coefficients can be derived using the Cholesky decomposition method. The Cholesky decomposition method may be a method of decomposing a symmetric positive definite matrix A into two matrices A=LLT. Here, L is a lower triangular matrix. LT is the transpose of L and is an upper triangular matrix.
[0419] According to another embodiment, filter coefficients can be derived using the LDL decomposition method. The LDL decomposition method may be a method of decomposing a symmetric matrix A into two matrices A = LDLT. Here, L is a lower triangular matrix with a unit diagonal. D is a diagonal matrix.
[0420] According to another embodiment, filter coefficients can be derived using Gaussian elimination consisting of two steps: forward elimination and back substitution.
[0421] In the process of deriving coefficients according to the method described above, square root operations can be avoided and division can be set not to be applied. Also, in the process of deriving coefficients, only multiplication and shift operations can be used.
[0422] Among the coefficient derivation methods described above, one coefficient derivation method can be established based on the current block's block size, block shape, slice type, QP value, left and upper block information, and block (pixel) position information. Then, filter coefficients can be derived by the established coefficient derivation method.
[0423] For example, to derive filter coefficients, you can set it to always use LDL decomposition. Alternatively, to derive filter coefficients, you can set it to always use Gaussian elimination. Or, to derive filter coefficients, you can use both methods interchangeably depending on the slice type.
[0424] Alternatively, a candidate group may be established using multiple candidate coefficient derivation methods. Then, the video encoding device may determine a candidate coefficient derivation method from among the candidate coefficient derivation methods and transmit information indicating the determined candidate coefficient derivation method to the decoder. At this time, the information indicating the determined candidate coefficient derivation method may be transmitted in units such as sequence, GOP (group of picture), frame, picture, slice, CTU, CU, TU, etc. For example, the video encoding device may determine a coefficient derivation method from among the candidate coefficient derivation methods by considering RD performance. Then, the video encoding device may transmit information indicating the coefficient derivation method to the decoder.
[0425]
[0426] Meanwhile, according to one embodiment of the present disclosure, in deriving the coefficients of a convolutional filter, instead of calculating the coefficients of each term each time, the coefficients applied to pixels located at a 3x3 position can be calculated at once, and the filter coefficients actually used can be applied to the restored luminance sample to generate a predicted value of the color difference sample.
[0427] Below, a method for generating coefficients of a convolutional filter according to one embodiment of the present disclosure is described.
[0428]
[0429] FIG. 21 is a drawing illustrating a method for deriving convolutional filter coefficients according to one embodiment of the present disclosure.
[0430] Referring to FIG. 21, filter candidates constituting the candidate group may use a central sample and / or one or more samples adjacent to the central sample. Here, an image encoding device or an image decoder may derive filter coefficients applied to the central sample and samples adjacent to the central sample.
[0431] In addition, among various types of convolutional filter candidates, one or more convolutional filter candidates may be selected. For example, the selected convolutional filter candidates may include filter candidates composed of linear and non-linear elements utilizing the top sample, left sample, center sample, right sample, and bottom sample, and filter candidates composed of linear and non-linear elements utilizing the top sample, left sample, center sample, right sample, and bottom sample. In this case, each filter may utilize coefficient values of [C4, C5, C6, C10, C11] and [C2, C5, C8, C10, C11]. Furthermore, the coefficients of the filters may each be set to the filter coefficients derived earlier.
[0432] Accordingly, according to one embodiment of the present disclosure, instead of calculating coefficients for a specific filter individually, all coefficients corresponding to C1 to C11 can be calculated. And, by setting the coefficients of each filter candidate based on the calculated coefficient values, the complexity of the filter candidate generation process can be reduced.
[0433]
[0434] Meanwhile, some samples in the template area adjacent to the current chrominance block may not be available. Alternatively, some samples in the template area adjacent to the corresponding luminance block of the current chrominance block may not be available. For example, if the current chrominance block or the corresponding luminance block corresponds to the boundary of a predetermined area, some samples in the template area may not be available. The case described above may be referred to as a boundary condition.
[0435] A block corresponding to the boundary condition may be a block located at the left boundary, top-left boundary, bottom-left boundary, top boundary, right boundary, top-right boundary, or bottom-right boundary of a frame, slice, picture, tile, etc. Alternatively, a block corresponding to the boundary condition may be a block adjacent to a coding tree unit and / or coding unit that is not managed in the buffer. In addition, a block corresponding to the boundary condition may be a block referencing an unavailable sample.
[0436] The template region of the block corresponding to the boundary condition may contain fewer samples than the number of reference samples required to apply the luminance component-based chrominance component prediction method. Therefore, the chrominance component prediction method for the block corresponding to the boundary condition can be performed as follows.
[0437] According to one embodiment, the luminance component-based chrominance component prediction method may not be applied to a block corresponding to a boundary condition. For example, if a block corresponds to a boundary condition and coefficients cannot be generated using samples from a template region, the image encoding device and the image decoder may be configured not to apply the luminance component-based chrominance component prediction method to the block.
[0438] Alternatively, if at least some of the samples in the template region are not available, the image encoding device may not apply the luminance component-based chrominance component prediction method to the block. Furthermore, the image encoding device may omit encoding information regarding the luminance component-based chrominance component prediction method. Additionally, if at least some of the samples in the template region are not available, the image decoder may not apply the luminance component-based chrominance component prediction method to the block.
[0439] According to another embodiment, a luminance component-based color difference component prediction method can be applied using samples that are referenced in a template area of a block corresponding to a boundary condition. For example, if at least some of the samples in the template area are not available, available samples can be combined, and the luminance component-based color difference component prediction method can be performed using the combined samples.
[0440] Alternatively, if at least some of the samples in the template area are not available, the luminance component-based color difference component prediction method may be performed only when the number of available samples is greater than or equal to a predetermined number. For example, in a template containing 16 samples, if the number of available samples is less than 8, the luminance component-based color difference component prediction method may not be performed. On the other hand, if the number of available samples in the template is 8 or more, the luminance component-based color difference component prediction method may be performed.
[0441] Alternatively, if at least some of the samples in the template region are unavailable, a luminance component-based color difference component prediction method may be performed using only the available samples in the region. For example, if the upper template region is unavailable, a luminance component-based color difference component prediction method may be performed using only the left template region. Conversely, if the left template region is unavailable, a luminance component-based color difference component prediction method may be performed using only the upper template region.
[0442] Alternatively, if at least some of the samples in the template area are not available, a luminance component-based color difference component prediction method may be performed using extended samples based on the available samples.
[0443] For example, samples used in a luminance component-based chrominance component prediction method may further include interpolated pixels obtained by applying upsampling to available samples. Alternatively, unavailable samples may be replaced with the values of adjacent available samples. Alternatively, unavailable samples may be replaced with the average value of available samples. Alternatively, unavailable samples may be replaced with values obtained by applying a transpose operation to available samples. Alternatively, unavailable samples may be replaced with samples derived by applying a template matching method based on available samples.
[0444] Meanwhile, if the left template is not available, the sample values of the left template can be replaced with the sample values of the top template. Alternatively, if the left template is not available, the area of the top template can be expanded by a factor of two, and samples from the expanded top template area can be used in the luminance component-based color difference component prediction method.
[0445] Meanwhile, if the top template is not available, the sample values of the top template may be replaced with the sample values of the left template. Alternatively, if the top template is not available, the area of the left template may be expanded by a factor of two, and samples from the expanded left template area may be used in the luminance component-based color difference component prediction method.
[0446] Alternatively, if at least some of the samples in the template area are unavailable and wrap-around pixels are available, the unavailable samples may be replaced with the wrap-around pixels. Furthermore, the available samples and the replaced samples may be used in the luminance component-based chrominance component prediction method. For example, if the current block is located at the left boundary of a predetermined area and wrap-around pixels are available, the samples in the left template area may be replaced with the values of the samples located to the right at the same height.
[0447]
[0448] Meanwhile, when applying a luminance component-based chrominance component prediction method to a block located at the boundary of a predetermined area, the convolutional filter can be adaptively determined.
[0449] For example, the shape of the convolutional filter can be determined based on the available template area of the current block. Then, the convolutional filter can be generated by adding non-linear elements to the filter of the determined shape. For example, if the top template of the current block is not available, only the left template can be used in the luminance component-based chrominance component prediction method. Therefore, a vertical filter can be applied to the reconstructed value of the luminance component. On the other hand, if the left template of the current block is not available, only the top template can be used in the luminance component-based chrominance component prediction method. Therefore, a horizontal filter can be applied to the reconstructed value of the luminance component.
[0450]
[0451] In addition, information regarding a luminance component-based chrominance component prediction method for a block located at the boundary of a predetermined area can be explicitly signaled. That is, the image encoding device determines a luminance component-based chrominance component prediction method applied to a block located at the boundary of a predetermined area and can transmit information regarding the determined method to an image decoder.
[0452] Here, information regarding the determined method may be information indicating whether luminance component-based chrominance component prediction is available. Alternatively, information regarding the determined method may indicate a method for constructing a template used in luminance component-based chrominance component prediction. For example, information regarding the determined method may be information instructing to use only referenceable pixels or to expand pixels within the template.
[0453] Alternatively, information regarding the determined method may be information regarding a convolutional filter. For example, among filters containing different types and different numbers of non-linear elements, a candidate group including at least some of the filters may be established. Then, the image encoding device may determine one filter candidate among the filter candidates included in the candidate group and transmit information indicating the determined filter candidate to a decoder. At this time, the information indicating the determined filter candidate may be transmitted in units such as sequence, group of picture (GOP), frame, picture, slice, CTU, CU, TU, etc. The image encoding device may determine the optimal filter among the filters constituting the candidate group by considering RD performance. Then, the image encoding device may transmit information indicating the optimal filter to an image decoder.
[0454]
[0455] FIG. 22 is a diagram illustrating a luminance component-based chrominance component prediction method according to one embodiment of the present disclosure. The luminance component-based chrominance component prediction method can be performed by an image decoding device.
[0456] Referring to FIG. 22, in step S2210, the image decoder can derive a prediction mode of the current color difference block based on whether a luminance component-based color difference prediction is available. Here, the availability of a luminance component-based color difference prediction can be determined based on the size of the current color difference block.
[0457] Alternatively, the availability of luminance component-based color difference prediction may be determined based on information indicating the availability of luminance component-based color difference prediction. If it is determined that luminance component-based color difference prediction is available, the prediction mode of the current color difference block may be derived based on information indicating whether to apply luminance component-based color difference prediction.
[0458] In step S2220, the image decoder may apply downsampling to samples in a corresponding luminance block corresponding to the current chrominance block and in a luminance region adjacent to the corresponding luminance block. The luminance region adjacent to the corresponding luminance block may be determined based on the availability of samples adjacent to the corresponding luminance block.
[0459] In the step of applying downsampling to samples in a luminance region adjacent to a corresponding luminance block, one downsampling filter determined from among a plurality of downsampling filters may be applied to the samples in the luminance region adjacent to the corresponding luminance block. Here, the one downsampling filter may be determined based on information indicating one downsampling filter among the plurality of downsampling filters.
[0460] In step S2230, the image decoder can derive a convolutional filter based on samples of a chrominance region adjacent to the current chrominance block and a luminance region adjacent to the corresponding luminance block. The chrominance region adjacent to the current chrominance block can be determined based on the availability of samples adjacent to the current chrominance block. Furthermore, in the step of deriving the convolutional filter, the coefficients of the convolutional filter can be derived by utilizing Gaussian elimination.
[0461] Here, the convolutional filter may be a single convolutional filter determined from among multiple convolutional filters having different components. And, the single convolutional filter may be determined based on information indicating one convolutional filter among the multiple convolutional filters.
[0462] In step S2240, the image decoder can generate a predicted value of the current color difference block based on the downsampled corresponding luminance block and the convolutional filter.
[0463]
[0464] Meanwhile, the steps described in FIG. 22 can be performed in the same way in a video encoding method. Additionally, a bitstream can be generated by a video encoding method including the steps described in FIG. 22. The bitstream can be stored on a non-transient computer-readable recording medium and can also be transmitted (or streamed).
[0465]
[0466] The exemplary methods of the present disclosure are described as a series of operations for clarity of description, but this is not intended to limit the order in which the steps are performed, and if necessary, each step may be performed simultaneously or in a different order. To implement the method according to the present disclosure, additional steps may be included in addition to the steps exemplified, steps excluding some steps and including the remaining steps, or steps excluding some steps and including additional steps.
[0467] The various embodiments of the present disclosure are not intended to list all possible combinations but to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.
[0468] Various embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof. In the case of implementation by hardware, it may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.
[0469] Alternatively, various embodiments of the present disclosure may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. And, a bitstream generated by the encoding method according to the embodiment may be stored on a non-transient computer-readable recording medium.
[0470] The above-mentioned computer-readable recording medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the above-mentioned computer-readable recording medium may be those specifically designed and configured for the present disclosure, or they may be those known and available to those skilled in the art of computer software.
[0471] In the foregoing, the present disclosure is described based on specific details, such as specific components, limited embodiments, and drawings. However, the embodiments of the present disclosure are provided merely to aid in the overall understanding of the present disclosure and are not intended to limit the present disclosure to the above embodiments. Accordingly, a person skilled in the art can make various modifications and variations from this description.
[0472] Accordingly, the scope of the present invention should not be limited to the embodiments described above, and all modifications equivalent to or equivalent to the claims set forth below, as well as the claims described below, shall be considered to fall within the scope of the concept of the present invention.
[0473] The present invention can be used in a device for encoding an image, a device for decoding an image, and a recording medium for storing a bitstream.
Claims
1. In a video decoding method, A step of deriving a prediction mode of the current color difference block based on the availability of luminance component-based color difference prediction; A step of applying downsampling to samples of a corresponding luminance block corresponding to the current color difference block and a luminance region adjacent to the corresponding luminance block; A step of deriving a convolutional filter based on samples of a color difference region adjacent to the current color difference block and a luminance region adjacent to the corresponding luminance block; and The method includes the step of generating a predicted value of the current color difference block based on the downsampled corresponding luminance block and the convolutional filter. An image decoding method in which the availability of the above-mentioned luminance component-based color difference prediction is determined based on the information of the above-mentioned current color difference block.
2. In Paragraph 1, The availability of the above-mentioned luminance component-based color difference prediction is, An image decoding method characterized by being determined based on the size of the current color difference block.
3. In Paragraph 1, The availability of the above-mentioned luminance component-based color difference prediction is, An image decoding method characterized by being determined based on information indicating whether a luminance component-based color difference prediction is available.
4. In Paragraph 1, If it is determined that the above luminance component-based color difference prediction is available, An image decoding method characterized in that the prediction mode of the current color difference block is derived based on information indicating whether to apply the luminance component-based color difference prediction.
5. In Paragraph 1, The color difference area adjacent to the above current color difference block is, An image decoding method characterized by being determined based on the availability of samples adjacent to the current color difference block.
6. In Paragraph 1, The luminance area adjacent to the above-mentioned corresponding luminance block is An image decoding method characterized by being determined based on the availability of samples adjacent to the corresponding luminance block.
7. In Paragraph 1, The step of applying downsampling to samples in a luminance region adjacent to the corresponding luminance block is: An image decoding method characterized by applying one downsampling filter determined from among a plurality of downsampling filters to samples in a luminance region adjacent to the corresponding luminance block.
8. In Paragraph 7, The one downsampling filter determined above is, An image decoding method characterized by determining one downsampling filter among a plurality of downsampling filters based on information indicating the downsampling filter.
9. In Paragraph 1, The above convolutional filter is, An image decoding method characterized by being one convolutional filter determined from among a plurality of convolutional filters having different components.
10. In Paragraph 9, The one convolutional filter determined above is, An image decoding method characterized by determining one convolutional filter among a plurality of convolutional filters based on information indicating the convolutional filter.
11. In Paragraph 1, The step of inducing the above convolutional filter is, An image decoding method characterized by deriving the coefficients of a convolutional filter using Gaussian elimination.
12. In a video encoding method, A step of deriving a prediction mode of the current color difference block based on the availability of luminance component-based color difference prediction; A step of applying downsampling to samples of a corresponding luminance block corresponding to the current color difference block and a luminance region adjacent to the corresponding luminance block; A step of deriving a convolutional filter based on samples of a color difference region adjacent to the current color difference block and a luminance region adjacent to the corresponding luminance block; and The method includes the step of generating a predicted value of the current color difference block based on the downsampled corresponding luminance block and the convolutional filter. An image encoding method in which the availability of the above-mentioned luminance component-based color difference prediction is determined based on the information of the above-mentioned current color difference block.
13. A non-transient computer-readable recording medium storing a bitstream generated by a video encoding method, The above image encoding method is, A step of deriving a prediction mode of the current color difference block based on the availability of luminance component-based color difference prediction; A step of applying downsampling to samples of a corresponding luminance block corresponding to the current color difference block and a luminance region adjacent to the corresponding luminance block; A step of deriving a convolutional filter based on samples of a color difference region adjacent to the current color difference block and a luminance region adjacent to the corresponding luminance block; and The method includes the step of generating a predicted value of the current color difference block based on the downsampled corresponding luminance block and the convolutional filter. A non-transient computer-readable recording medium in which the availability of the above-mentioned luminance component-based color difference prediction is determined based on the information of the above-mentioned current color difference block.
14. A method for transmitting a bitstream generated by a video encoding method, The above transmission method includes the step of transmitting the bitstream, and The above image encoding method is, A step of deriving a prediction mode of the current color difference block based on the availability of luminance component-based color difference prediction; A step of applying downsampling to samples of a corresponding luminance block corresponding to the current color difference block and a luminance region adjacent to the corresponding luminance block; A step of deriving a convolutional filter based on samples of a color difference region adjacent to the current color difference block and a luminance region adjacent to the corresponding luminance block; and The method includes the step of generating a predicted value of the current color difference block based on the downsampled corresponding luminance block and the convolutional filter. A transmission method in which the availability of the above luminance component-based color difference prediction is determined based on the information of the above current color difference block.