Tranform based image coding method and apparatus
Patent Information
- Application Number
- KR1020237014791
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-29
- Filing Date
- 2020-10-29
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2040-10-29
Smart Images

Figure R1020237014791_ABST
Abstract
Description
Technology Field
[0001] This document relates to image coding technology, and more specifically, to an image coding method and apparatus based on transform in an image coding system. Background Technology
[0002] Recently, the demand for high-resolution, high-quality video, such as 4K or 8K or higher UHD (Ultra High Definition) video, is increasing across various fields. As video data becomes higher resolution and higher quality, the amount of information or bits transmitted increases relative to existing video data; therefore, when transmitting video data using media such as existing wired or wireless broadband lines or storing video data using existing storage media, transmission and storage costs increase.
[0003] In addition, interest in and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have recently been increasing, and the broadcasting of video content with characteristics different from reality, such as game footage, is on the rise.
[0004] Accordingly, high-efficiency image / video compression technology is required to effectively compress, transmit, store, and play back high-resolution, high-quality image / video information having various characteristics as described above. The problem to be solved
[0005] The technical objective of this document is to provide a method and device for increasing image coding efficiency.
[0006] Another technical objective of this document is to provide a method and apparatus for increasing the efficiency of LFNST index coding.
[0007] Another technical objective of this document is to provide a method and apparatus for increasing the efficiency of a second transformation through the coding of an LFNST index.
[0008] Another technical objective of this document is to provide an image coding method and apparatus for deriving an LFNST transform set by borrowing the intra mode of the luminance block in CCLM mode. means of solving the problem
[0009] According to one embodiment of the present document, a video decoding method performed by a decoding device is provided. The method comprises the steps of: deriving an intra prediction mode of a chroma block into a CCLM (cross-component linear model) mode; updating the intra prediction mode of the chroma block based on an intra prediction mode of a luminance block corresponding to the chroma block; and determining an LFNST set including LFNST matrices based on the updated intra prediction mode, wherein the updated intra prediction mode is derived into an intra prediction mode corresponding to a specific location within the luminance block, and the specific location is set based on the color format of the chroma block.
[0010] The above specific location may be the center location of the above-mentioned Luma block.
[0011] The above specific location is set to ((xTbY + ( nTbW * SubWidthC ) / 2), (yTbY + ( nTbH * SubHeightC ) / 2)), where xTbY and yTbY represent the top-left coordinates of the above luma block, nTbW and nTbH represent the width and height of the above chroma block, and SubWidthC and SubHeightC may be variables corresponding to the above color format.
[0012] If the above color format is 4:2:0, SubWidthC and SubHeightC are 2, and if the above color format is 4:2:2, SubWidthC is 2 and SubHeightC can be 1.
[0013] If the intra prediction mode corresponding to the specific location above is an MIP mode, the updated intra prediction mode may be an intra planner mode.
[0014] If the intra prediction mode corresponding to the specific location above is an IBC mode, the updated intra prediction mode may be an intra DC mode.
[0015] If the intra prediction mode corresponding to the specific location above is a palette mode, the updated intra prediction mode may be an intra DC mode.
[0016] According to one embodiment of the present document, a video encoding method performed by an encoding device is provided. The method comprises the steps of: deriving an intra prediction mode for a chroma block as a CCLM (cross-component linear model) mode; deriving prediction samples for the chroma block based on the CCLM mode; deriving residual samples for the chroma block based on the prediction samples; updating the intra prediction mode of the chroma block based on an intra prediction mode of a luminance block corresponding to the chroma block; determining an LFNST set including LFNST matrices based on the updated intra prediction mode; and deriving modified transformation coefficients for the chroma block based on the residual samples and the LFNST matrices, wherein the updated intra prediction mode is derived as an intra prediction mode corresponding to a specific location within the luminance block, and the specific location may be set based on the color format of the chroma block.
[0017] According to another embodiment of the present document, a digital storage medium may be provided that stores video data including encoded video information and a bitstream generated according to a video encoding method performed by an encoding device.
[0018] According to another embodiment of the present document, a digital storage medium may be provided that stores image data including encoded image information and a bitstream that causes the image decoding method to be performed by a decoding device. Effects of the invention
[0019] According to this document, overall video compression efficiency can be increased.
[0020] According to this document, the efficiency of LFNST index coding can be improved.
[0021] According to this document, the efficiency of the second transformation can be improved through the coding of the LFNST index.
[0022] Another technical challenge of this document is to provide an image coding method and apparatus for deriving an LFNST transform set by borrowing the intra mode of the luminance block in CCLM mode.
[0023] The effects obtainable through the specific examples of this specification are not limited to those listed above. For example, there may be various technical effects that a person having ordinary skill in the related art can understand or derive from this specification. Accordingly, the specific effects of this specification are not limited to those explicitly described herein, but may include various effects that can be understood or derived from the technical features of this specification. Brief explanation of the drawing
[0024] Figure 1 is a diagram schematically illustrating the configuration of a video / image encoding device to which the present document can be applied. Figure 2 is a diagram schematically illustrating the configuration of a video / image decoding device to which the present document can be applied. FIG. 3 schematically illustrates a multiple conversion technique according to one embodiment of the present document. Figure 4 exemplarily shows 65 intra-directional modes of prediction directions. FIG. 5 is a drawing for explaining an RST according to one embodiment of the present document. Figure 6 is a diagram illustrating the sequence of arranging output data of a forward first-order transformation into a one-dimensional vector according to one example. Figure 7 is a diagram illustrating the sequence of arranging output data of a forward quadratic transformation into two-dimensional blocks according to one example. Figure 8 is a drawing illustrating the block shape to which LFNST is applied. Figure 9 is a diagram illustrating the arrangement of output data of a forward LFNST according to one example. Figure 10 is a drawing illustrating zero out in a block to which a 4x4 LFNST is applied according to one example. Figure 11 is a drawing illustrating zero out in a block to which an 8x8 LFNST is applied according to one example. FIG. 12 is a diagram illustrating a CCLM that can be applied when deriving an intra-prediction mode of a chroma block according to one embodiment. Figure 13 is a diagram illustrating a method for decoding an image according to one example. Figure 14 is a diagram illustrating an image encoding method according to one example. Figure 15 schematically illustrates an example of a video / image coding system to which the present document can be applied. Figure 16 illustrates an exemplary structural diagram of a content streaming system to which this document applies. Specific details for implementing the invention
[0025] As this document is subject to various modifications and may have various embodiments, specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. Terms used in this specification are used merely to describe specific embodiments and are not intended to limit the technical scope of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "comprising" or "having" in this specification are intended to indicate the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0026] Meanwhile, each component in the drawings described in this document is depicted independently for the convenience of explaining different characteristic functions and does not imply that each component is implemented in separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of the rights of this document, provided that they do not deviate from the essence of this document.
[0027] Hereinafter, preferred embodiments of the present document will be described in more detail with reference to the attached drawings. Hereinafter, the same reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted.
[0028] This document relates to video / image coding. For example, the methods / executions disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), next-generation video / image coding standards following VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), EVC (essential video coding) standard, AVS2 standard, etc.).
[0029] This document presents various embodiments regarding video / image coding, and unless otherwise noted, the embodiments may be performed in combination with one another.
[0030] In this document, "video" may refer to a set of images over time. "Picture" generally refers to a unit representing a single image at a specific time, and "slice" or "tile" are units that constitute a part of a picture in coding. A slice or tile may contain one or more CTUs (coding tree units). A single picture may consist of one or more slices or tiles. A single picture may consist of one or more tile groups. A tile group may contain one or more tiles.
[0031] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. Generally, a sample can represent a pixel or its value; it may represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chroma component. Alternatively, a sample may refer to a pixel value in the spatial domain, or it may refer to the transformation coefficient in the frequency domain when such pixel values are converted to the frequency domain.
[0032] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0033] In this document, “ / ” and “,” are interpreted as “and / or.” For example, “A / B” is interpreted as “A and / or B,” and “A, B” is interpreted as “A and / or B.” Additionally, “A / B / C” means “at least one of A, B and / or C.” Furthermore, “A, B, C” also means “at least one of A, B and / or C.”
[0034] Additionally, “or” in this document is interpreted as “and / or.” For example, “A or B” may mean 1) “A” only, 2) “B” only, or 3) “A and B.” Alternatively, “or” in this document may mean “additionally or alternatively.”
[0035] In this specification, “at least one of A and B” may mean “only A,” “only B,” or “both A and B.” Additionally, in this specification, the expressions “at least one of A or B” or “at least one of A and / or B” may be interpreted as synonymous with “at least one of A and B.”
[0036] Additionally, in this specification, “at least one of A, B and C” may mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” may mean “at least one of A, B and C.”
[0037] Additionally, parentheses used in this specification may mean “for example.” Specifically, where indicated as “prediction (intra-prediction),” “intra-prediction” may be proposed as an example of “prediction.” In other words, “prediction” in this specification is not limited to “intra-prediction,” and “intra-prediction” may be proposed as an example of “prediction.” Furthermore, even when indicated as “prediction (i.e., intra-prediction),” “intra-prediction” may be proposed as an example of “prediction.”
[0038] Technical features described individually within a single drawing in this specification may be implemented individually or simultaneously.
[0039] FIG. 1 is a diagram schematically illustrating the configuration of a video / image encoding device that can be applied to embodiments of the present document. Hereinafter, the term "video encoding device" may include an image encoding device.
[0040] Referring to FIG. 1, the encoding device (100) may be configured to include an image partitioner (110), a predictor (120), a residual processor (130), an entropy encoder (140), an adder (150), a filter (160), and a memory (170). The predictor (120) may include an inter-predictor (121) and an intra-predictor (122). The residual processor (130) may include a transformer (132), a quantizer (133), a dequantizer (134), and an inverse transformer (135). The residual processor (130) may further include a subtractor (131). The addition unit (150) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (110), prediction unit (120), residual processing unit (130), entropy encoding unit (140), addition unit (150), and filtering unit (160) may be configured by one or more hardware components (e.g., an encoder chipset or processor) according to the embodiment. Additionally, the memory (170) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (170) as an internal / external component.
[0041] The image segmentation unit (110) can divide an input image (or picture, frame) input to an encoding device (100) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit may be divided into multiple coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first and the binary-tree structure and / or ternary structure may be applied later. Or the binary-tree structure may be applied first. A coding procedure according to this document may be performed based on the final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively divided into lower-depth coding units so that a coding unit of the optimal size is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the final coding unit described above.The above prediction unit may be a unit of sample prediction, and the above transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.
[0042] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.
[0043] The encoding device (100) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (121) or an intra prediction unit (122) from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (132). In this case, as illustrated, the unit that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device (100) may be called a subtraction unit (131). The prediction unit performs a prediction for a block to be processed (hereinafter referred to as the current block) and can generate a predicted block containing prediction samples for the current block. The prediction unit can determine whether intra prediction is applied or inter prediction is applied at the current block or CU level. The prediction unit can generate various information regarding prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (140). The information regarding prediction can be encoded in the entropy encoding unit (140) and output in the form of a bitstream.
[0044] The intra prediction unit (122) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional prediction modes may be used. The intra prediction unit (122) may also determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.
[0045] The inter prediction unit (121) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different. The above temporal surrounding blocks may be referred to by names such as collocated reference block, collocated CU (colCU), etc., and the reference picture containing the above temporal surrounding blocks may be referred to as a collocated picture (colPic). For example, the inter prediction unit (121) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (121) may use the motion information of surrounding blocks as motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0046] The prediction unit (120) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, sample values within the picture can be signaled based on information regarding the palette table and palette index.
[0047] The prediction signal generated through the prediction unit (including the inter prediction unit (121) and / or the intra prediction unit (122)) may be used to generate a restored signal or to generate a residual signal. The transformation unit (132) may generate transform coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of the Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), Graph-Based Transform (GBT), or Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on generating a prediction signal using all previously reconstructed pixels. In addition, the transformation process can be applied to pixel blocks of the same square size, or to non-square blocks of variable size.
[0048] The quantization unit (133) quantizes the transformation coefficients and transmits them to the entropy encoding unit (140), and the entropy encoding unit (140) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (133) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients. The entropy encoding unit (140) can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (140) may encode information necessary for video / image restoration (e.g., values of syntax elements) together or separately, in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream.The above bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits the signal output from the entropy encoding unit (140) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the encoding device (100), or the transmission unit may be included in the entropy encoding unit (140).
[0049] Quantized transformation coefficients output from the quantization unit (133) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (134) and the inverse transformation unit (135). An adder (155) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter-prediction unit (121) or the intra-prediction unit (122). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (150) may be called a reconstruction unit or a reconstruction block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.
[0050] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0051] The filtering unit (160) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (160) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (170), specifically in the DPB of memory (170). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (160) can generate various information regarding filtering and transmit it to the entropy encoding unit (140), as described below in the description of each filtering method. The information regarding filtering can be encoded in the entropy encoding unit (140) and output in the form of a bitstream.
[0052] The modified restored picture transmitted to memory (170) can be used as a reference picture in the inter-prediction unit (121). Through this, when inter-prediction is applied, the encoding device can avoid prediction mismatches between the encoding device (100) and the decoding device, and can also improve encoding efficiency.
[0053] The memory (170) DPB can store the modified restored picture to be used as a reference picture in the inter-prediction unit (121). The memory (170) can store motion information of blocks from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (121) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (170) can store restoration samples of the blocks restored within the current picture and transmit them to the intra-prediction unit (122).
[0054] FIG. 2 is a diagram schematically illustrating the configuration of a video / image decoding device that can be applied to embodiments of the present document.
[0055] Referring to FIG. 2, the decoding device (200) may be configured to include an entropy decoder (210), a residual processor (220), a predictor (230), an adder (240), a filter (250), and a memory (260). The predictor (230) may include an inter-predictor (231) and an intra-predictor (232). The residual processor (220) may include a dequantizer (221) and an inverse transformer (221). The aforementioned entropy decoding unit (210), residual processing unit (220), prediction unit (230), addition unit (240), and filtering unit (250) may be configured by a single hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Additionally, the memory (260) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (260) as an internal / external component.
[0056] When a bitstream containing video / image information is input, the decoding device (200) can restore the image in correspondence with the process in which the video / image information is processed in the encoding device of FIG. 1. For example, the decoding device (200) can derive units / blocks based on block division information obtained from the bitstream. The decoding device (200) can perform decoding using a processing unit applied in the encoding device. Accordingly, the processing unit for decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a binary tree structure. One or more conversion units may be derived from the coding unit. And, the restored image signal decoded and output through the decoding device (200) can be played back through a playback device.
[0057] A decoding device (200) can receive a signal output from the encoding device of FIG. 1 in the form of a bitstream, and the received signal can be decoded through an entropy decoding unit (210). For example, the entropy decoding unit (210) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device can decode the picture based on information regarding the parameter sets and / or the general constraint information. The signaling / receiving information and / or syntax elements described below in this document can be obtained from the bitstream by being decoded through the decoding procedure. For example, the entropy decoding unit (210) can decode information within a bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in a bitstream, determine a context model using information of the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (210), information regarding prediction is provided to the prediction unit (inter prediction unit (232) and intra prediction unit (231)), and the residual value for which entropy decoding was performed in the entropy decoding unit (210), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (220). The residual processing unit (220) can derive residual signals (residual blocks, residual samples, residual sample array). Additionally, among the information decoded in the entropy decoding unit (210), information regarding filtering can be provided to the filtering unit (250). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (200), or the receiving unit may be a component of the entropy decoding unit (210). Meanwhile, the decoding device according to the present document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit (210), and the sample decoder may include at least one of the inverse quantization unit (221), inverse transform unit (222), adder (240), filtering unit (250), memory (260), inter prediction unit (232), and intra prediction unit (231).
[0058] In the inverse quantization unit (221), the quantized transform coefficients can be inversely quantized to output transform coefficients. The inverse quantization unit (221) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit (221) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0059] In the inverse conversion unit (222), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0060] The prediction unit performs a prediction for the current block and can generate a predicted block containing prediction samples for the current block. Based on information regarding the prediction output from the entropy decoding unit (210), the prediction unit can determine whether an intra prediction or an inter prediction is applied to the current block and can determine a specific intra / inter prediction mode.
[0061] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and palette index can be included in the above video / image information and signaled.
[0062] The intra prediction unit (231) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit (231) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.
[0063] The inter prediction unit (232) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (232) may construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes, and information regarding the prediction may include information indicating the mode of inter-prediction for the current block.
[0064] The adder (240) can generate a restoration signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit (232) and / or the intra prediction unit (231)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the restoration block.
[0065] The addition unit (240) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture.
[0066] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0067] The filtering unit (250) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (250) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to memory (260), specifically to the DPB of memory (260). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0068] The (modified) restored picture stored in the DPB of the memory (260) can be used as a reference picture in the inter prediction unit (232). The memory (260) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter prediction unit (232) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (260) can store restoration samples of blocks restored within the current picture and transmit them to the intra prediction unit (231).
[0069] In this document, the embodiments described in the filtering unit (160), inter prediction unit (121), and intra prediction unit (122) of the encoding device (100) may be applied to the filtering unit (250), inter prediction unit (232), and intra prediction unit (231) of the decoding device (200) in the same or corresponding manner.
[0070] As described above, prediction is performed to increase compression efficiency during video coding. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived identically by both the encoding device and the decoding device. The encoding device can increase video coding efficiency by signaling information regarding the residuals between the original block and the predicted block (residual information) to the decoding device, rather than the original sample values of the original block itself. Based on the residual information, the decoding device derives a residual block containing residual samples, combines the residual block and the predicted block to generate a restored block containing restored samples, and can generate a restored picture containing the restored blocks.
[0071] The above residual information can be generated through transformation and quantization procedures. For example, an encoding device may derive a residual block between the original block and the predicted block, perform a transformation procedure on the residual samples (residual sample array) included in the residual block to derive transformation coefficients, perform a quantization procedure on the transformation coefficients to derive quantized transformation coefficients, and signal the related residual information to a decoding device (via a bitstream). Here, the residual information may include information such as value information of the quantized transformation coefficients, position information, transformation technique, transformation kernel, and quantization parameters. The decoding device may perform inverse quantization / inverse transformation procedures based on the residual information and derive residual samples (or residual blocks). The decoding device may generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also derive a residual block by inversely quantizing / inversely transforming the quantized transform coefficients for reference to inter-predicting of the picture, and generate a restored picture based thereon.
[0072] Figure 3 schematically illustrates a multiple conversion technique according to the present document.
[0073] Referring to FIG. 3, the conversion unit may correspond to the conversion unit within the encoding device of FIG. 1 described above, and the inverse conversion unit may correspond to the inverse conversion unit within the encoding device of FIG. 1 described above or the inverse conversion unit within the decoding device of FIG. 2.
[0074] The transformation unit can derive (primary) transformation coefficients by performing a primary transformation based on the residual samples (residual sample array) within the residual block (S310). This primary transformation can be referred to as a core transformation. Here, the primary transformation can be based on Multiple Transform Selection (MTS), and if multiple transformations are applied as the primary transformation, it can be referred to as multiple core transformations.
[0075] A multi-core transformation may represent a method of transformation using additional Discrete Cosine Transform (DCT) Type 2, Discrete Sine Transform (DST) Type 7, DCT Type 8, and / or DST Type 1. That is, the multi-core transformation may represent a transformation method that transforms a residual signal (or residual block) in the spatial domain into transformation coefficients (or first-order transformation coefficients) in the frequency domain based on a plurality of transformation kernels selected from DCT Type 2, DST Type 7, DCT Type 8, and DST Type 1. Here, the first-order transformation coefficients may be referred to as temporary transformation coefficients from the perspective of the transformation unit.
[0076] In other words, when the existing transformation method is applied, transformation coefficients could be generated by applying a transformation from the spatial domain to the frequency domain for a residual signal (or residual block) based on DCT Type 2. In contrast, when the above-mentioned multi-core transformation is applied, transformation coefficients (or first-order transformation coefficients) can be generated by applying a transformation from the spatial domain to the frequency domain for a residual signal (or residual block) based on DCT Type 2, DST Type 7, DCT Type 8, and / or DST Type 1, etc. Here, DCT Type 2, DST Type 7, DCT Type 8, and DST Type 1, etc., may be referred to as transformation types, transformation kernels, or transformation cores. These DCT / DST transformation types can be defined based on basis functions.
[0077] When the above multiple core transformations are performed, a vertical transformation kernel and a horizontal transformation kernel for the target block may be selected among the transformation kernels, a vertical transformation for the target block may be performed based on the vertical transformation kernel, and a horizontal transformation for the target block may be performed based on the horizontal transformation kernel. Here, the horizontal transformation may represent a transformation for the horizontal components of the target block, and the vertical transformation may represent a transformation for the vertical components of the target block. The vertical transformation kernel / horizontal transformation kernel may be adaptively determined based on the prediction mode and / or transformation index of the target block (CU or subblock) containing the residual block.
[0078] In addition, according to one example, when performing a first-order transformation by applying MTS, specific basis functions can be set to predetermined values, and a mapping relationship for the transformation kernel can be established by combining whether the basis functions are applied when it is a vertical transformation or a horizontal transformation. For example, if the horizontal direction transformation kernel is represented as trTypeHor and the vertical direction transformation kernel is represented as trTypeVer, a value of 0 for trTypeHor or trTypeVer can be set to DCT2, a value of 1 for trTypeHor or trTypeVer can be set to DST7, and a value of 2 for trTypeHor or trTypeVer can be set to DCT8.
[0079] In this case, MTS index information may be encoded and signaled to a decoding device to indicate one of a plurality of conversion kernel sets. For example, if the MTS index is 0, it indicates that both trTypeHor and trTypeVer values are 0; if the MTS index is 1, it indicates that both trTypeHor and trTypeVer values are 1; if the MTS index is 2, it indicates that trTypeHor value is 2 and trTypeVer value is 1; if the MTS index is 3, it indicates that trTypeHor value is 1 and trTypeVer value is 2; and if the MTS index is 4, it indicates that both trTypeHor and trTypeVer values are 2.
[0080] For example, the conversion kernel set based on MTS index information is shown in the table below.
[0081]
[0082] The transformation unit can derive modified (secondary) transformation coefficients by performing a secondary transformation based on the above (firstary) transformation coefficients (S320). The above firstary transformation is a transformation from the spatial domain to the frequency domain, and the above secondary transformation means transforming into a more compressed representation using the correlation existing between the (firstary) transformation coefficients. The above secondary transformation may include a non-separable transform. In this case, the above secondary transformation may be called a non-separable secondary transform (NSST) or MDNSST (mode-dependent non-separable secondary transform). The above non-separable secondary transform may represent a transformation that generates modified transformation coefficients (or secondary transformation coefficients) for a residual signal by performing a secondary transformation on the (firstary) transformation coefficients derived through the above firstary transformation based on a non-separable transform matrix. Here, based on the inseparable transformation matrix, the transformation can be applied to the (first-order) transformation coefficients in a single step without separating the vertical and horizontal transformations (or applying horizontal and vertical transformations independently). In other words, the inseparable second-order transformation is not applied separately to the (first-order) transformation coefficients in the vertical and horizontal directions; instead, a transformation method can be described in which, for example, two-dimensional signals (transformation coefficients) are rearranged into one-dimensional signals through a specific defined direction (e.g., row-first or column-first direction), and then modified transformation coefficients (or second-order transformation coefficients) are generated based on the inseparable transformation matrix. For example, row-first ordering involves arranging the elements in a linear sequence of the 1st row, 2nd row, ..., Nth row for an MxN block, and column-first ordering involves arranging the elements in the 1st column, 2nd column, ......arranged in a row in the order of the Mth column. The above non-separable second transformation can be applied to the top-left area of a block composed of (first) transformation coefficients (hereinafter referred to as a transformation coefficient block). For example, if the width (W) and height (H) of the transformation coefficient block are both 8 or greater, an 8×8 non-separable second transformation can be applied to the top-left 8×8 area of the transformation coefficient block. Additionally, if the width (W) and height (H) of the transformation coefficient block are both 4 or greater, and the width (W) or height (H) of the transformation coefficient block is less than 8, a 4×4 non-separable second transformation can be applied to the top-left min(8,W)×min(8,H) area of the transformation coefficient block. However, the embodiments are not limited thereto, and for example, even if only the condition that the width (W) or height (H) of the transformation factor block is 4 or greater is satisfied, a 4×4 non-separable second transformation may be applied to the upper-left min(8,W)×min(8,H) area of the transformation factor block.
[0083] Specifically, for example, when a 4×4 input block is used, the non-separable second transformation can be performed as follows.
[0084] The above 4×4 input block X can be represented as follows.
[0085]
[0086] When the above X is represented in vector form, the vector It can be expressed as follows.
[0087]
[0088] As in mathematical equation 2, vector It rearranges the 2D block of X in Equation 1 into a 1D vector according to row-first order.
[0089] In this case, the above second-order inseparable transformation can be calculated as follows.
[0090]
[0091] Here, represents the transformation coefficient vector, and T represents the 16×16 (inseparable) transformation matrix.
[0092] 16×1 transformation coefficient vector through the above mathematical formula 3 ... can be derived, and the above It can be reorganized into 4×4 blocks through a scan order (horizontal, vertical, diagonal, etc.). However, the above calculation is merely an example, and to reduce the computational complexity of the inseparable quadratic transformation, the Hypercube-Givens Transform (HyGT) may be used for the calculation of the inseparable quadratic transformation.
[0093] Meanwhile, the above-mentioned non-separable second-order transformation may select a transformation kernel (or transformation core, transformation type) based on a mode-dependent basis. Here, the mode may include an intra-prediction mode and / or an inter-prediction mode.
[0094] As described above, the inseparable second transformation can be performed based on an 8×8 transformation or a 4×4 transformation determined based on the width (W) and height (H) of the transformation factor block. An 8x8 transformation refers to a transformation that can be applied to an 8x8 area contained within the transformation factor block when both W and H are greater than or equal to 8, and the 8x8 area may be the top-left 8x8 area within the transformation factor block. Similarly, a 4x4 transformation refers to a transformation that can be applied to a 4x4 area contained within the transformation factor block when both W and H are greater than or equal to 4, and the 4x4 area may be the top-left 4x4 area within the transformation factor block. For example, the 8x8 transformation kernel matrix may be a 64x64 / 16x64 matrix, and the 4x4 transformation kernel matrix may be a 16x16 / 8x16 matrix.
[0095] At this time, for mode-based conversion kernel selection, two inseparable second-order conversion kernels may be configured per conversion set for both 8×8 conversions and 4×4 conversions, and there may be four conversion sets. That is, four conversion sets may be configured for 8×8 conversions and four conversion sets may be configured for 4×4 conversions. In this case, each of the four conversion sets for 8×8 conversions may include two 8×8 conversion kernels, and each of the four conversion sets for 4×4 conversions may include two 4×4 conversion kernels.
[0096] However, the size of the transformation, that is, the size of the area to which the transformation is applied, may be a size other than 8×8 or 4×4 as an example, and the number of sets may be n, and the number of transformation kernels within each set may be k.
[0097] The above set of transformations may be referred to as the NSST set or the LFNST set. The selection of a specific set among the above transformation sets may be performed, for example, based on the intra-prediction mode of the current block (CU or subblock). LFNST (Low-Frequency Non-Separable Transform) may be an example of the reduced non-separable transform described later, and represents a non-separable transform for low-frequency components.
[0098] For reference, for example, the intra prediction modes may include two non-directional (or non-angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non-directional intra prediction modes may include a planar intra prediction mode number 0 and a DC intra prediction mode number 1, and the directional intra prediction modes may include 65 intra prediction modes number 2 through 66. However, this is merely an example, and the present document may apply even if the number of intra prediction modes is different. Meanwhile, in some cases, an additional 67th intra prediction mode may be used, and the 67th intra prediction mode may represent a linear model (LM) mode.
[0099] Figure 4 exemplarily shows 65 intra-directional modes of prediction directions.
[0100] Referring to Fig. 4, intra prediction modes with horizontal directionality and intra prediction modes with vertical directionality can be distinguished, centering on intra prediction mode 34, which has a rightward diagonal prediction direction. In Fig. 4, H and V represent horizontal and vertical directionality, respectively, and the numbers -32 to 32 represent displacements in units of 1 / 32 on the sample grid position. This can represent an offset for the mode index value. Intra prediction modes 2 through 33 have horizontal directionality, and intra prediction modes 34 through 66 have vertical directionality. Meanwhile, strictly speaking, intra prediction mode 34 can be considered neither horizontal nor vertical, but it can be classified as belonging to horizontal directionality from the perspective of determining the transformation set of the second transformation. This is because the input data is transposed for the vertical mode symmetrical to the 34th intra prediction mode, and the input data alignment method for the horizontal mode is used for the 34th intra prediction mode. Transposing the input data means that for 2D block data MxN, rows become columns and columns become rows to form NxM data. The 18th intra prediction mode and the 50th intra prediction mode represent the horizontal intra prediction mode and the vertical intra prediction mode, respectively. The 2nd intra prediction mode predicts in the upward-right direction using the left reference pixel, so it can be called the upward-right diagonal intra prediction mode; in the same context, the 34th intra prediction mode can be called the downward-right diagonal intra prediction mode, and the 66th intra prediction mode can be called the downward-left diagonal intra prediction mode.
[0101] For example, depending on the intra prediction mode, the mapping of four sets of transformations can be represented as shown in the following table.
[0102]
[0103] As shown in Table 2, depending on the intra prediction mode, any one of the four transformation sets, i.e., lfnstTrSetIdx, can be mapped to any one of the four, from 0 to 3.
[0104] Meanwhile, if it is determined that a specific set is used for the inseparable transformation, one of the k transformation kernels within the specific set may be selected through the inseparable second-order transformation index. The encoding device may derive an inseparable second-order transformation index pointing to a specific transformation kernel based on a rate-distortion (RD) check, and may signal the inseparable second-order transformation index to the decoding device. The decoding device may select one of the k transformation kernels within the specific set based on the inseparable second-order transformation index. For example, an lfnst index value of 0 may point to the first inseparable second-order transformation kernel, an lfnst index value of 1 may point to the second inseparable second-order transformation kernel, and an lfnst index value of 2 may point to the third inseparable second-order transformation kernel. Alternatively, an lfnst index value of 0 may indicate that the first non-separable second transformation is not applied to the target block, and lfnst index values 1 to 3 may indicate the three transformation kernels.
[0105] The transformation unit can perform the non-separable second transformation based on the selected transformation kernels and obtain modified (second) transformation coefficients. The modified transformation coefficients can be derived into quantized transformation coefficients through the quantization unit as described above, and can be encoded and transmitted to the decoding unit, signaling and encoding unit, and the inverse quantization / inverse transformation unit within the encoding unit.
[0106] Meanwhile, as described above, when the second transformation is omitted, the (first) transformation coefficients, which are the outputs of the first (separation) transformation, can be derived as quantized transformation coefficients through the quantization unit as described above, and can be encoded and transmitted to the decoding device, and to the inverse quantization / inverse transformation unit within the signaling and encoding device.
[0107] The inverse transform unit may perform a series of procedures in the reverse order of the procedures performed in the transform unit described above. The inverse transform unit receives (inversely quantized) transform coefficients, performs a second (inverse) transform to derive (first) transform coefficients (S350), and performs a first (inverse) transform on the (first) transform coefficients to obtain residual blocks (residual samples) (S360). Here, the first transform coefficients may be referred to as modified transform coefficients from the perspective of the inverse transform unit. As described above, the encoding device and the decoding device may generate a restored block based on the residual block and the predicted block, and generate a restored picture based thereon.
[0108] Meanwhile, the decoding device may further include a second-order inverse transform application determination unit (or an element determining whether to apply the second-order inverse transform) and a second-order inverse transform determination unit (or an element determining the second-order inverse transform). The second-order inverse transform application determination unit may determine whether to apply the second-order inverse transform. For example, the second-order inverse transform may be NSST, RST, or LFNST, and the second-order inverse transform application determination unit may determine whether to apply the second-order inverse transform based on a second-order transform flag parsed from the bitstream. As another example, the second-order inverse transform application determination unit may determine whether to apply the second-order inverse transform based on the transform coefficients of the residual block.
[0109] The second inverse transform determining unit can determine a second inverse transform. In this case, the second inverse transform determining unit can determine the second inverse transform applied to the current block based on a set of LFNST (NSST or RST) transforms specified according to the intra-prediction mode. Additionally, as an example, the second inverse transform determining method may be determined dependently on the first transform determining method. Various combinations of the first transform and the second transform may be determined according to the intra-prediction mode. Additionally, as an example, the second inverse transform determining unit may determine the area to which the second inverse transform is applied based on the size of the current block.
[0110] Meanwhile, as described above, in the case where the second (inverse) transform is omitted, the (inversely quantized) transform coefficients are received and the first (separated) inverse transform is performed to obtain a residual block (residual samples). As described above, the encoding device and the decoding device can generate a restored block based on the residual block and the predicted block, and generate a restored picture based thereon.
[0111] Meanwhile, in order to reduce the amount of computation and memory requirements associated with non-separable secondary transformations, the reduced secondary transform (RST) can be applied in the concept of NSST, in which the size of the transformation matrix (kernel) is reduced.
[0112] Meanwhile, the conversion kernel, conversion matrix, and coefficients constituting the conversion kernel matrix described in this document—that is, kernel coefficients or matrix coefficients—can be represented in 8 bits. This may be a condition for implementation in decoding and encoding devices, and can reduce the memory requirements for storing the conversion kernel while entailing a reasonably acceptable performance degradation compared to existing 9-bit or 10-bit methods. Additionally, by representing the kernel matrix in 8 bits, a smaller multiplier can be used, and it may be more suitable for SIMD (Single Instruction Multiple Data) instructions used for optimal software implementation.
[0113] In this specification, RST may refer to a transformation performed on residual samples of a target block based on a transform matrix whose size is reduced according to a simplification factor. When performing a simplification transformation, the amount of computation required during the transformation may be reduced due to the reduction in the size of the transform matrix. That is, RST can be used to resolve computational complexity issues that arise during the transformation of large blocks or non-separable transformations.
[0114] RST may be referred to by various terms such as reduced transform, reduced secondary transform, reduction transform, simplified transform, and simple transform, and the names by which RST can be referred are not limited to the examples listed. Alternatively, since RST is primarily performed in the low-frequency region containing non-zero coefficients in the transform block, it may also be referred to as LFNST (Low-Frequency Non-Separable Transform). The above transform index may be named the LFNST index.
[0115] Meanwhile, when the second inverse transformation is performed based on RST, the inverse transformation unit (135) of the encoding device (100) and the inverse transformation unit (222) of the decoding device (200) may include an inverse RST unit that derives modified transformation coefficients based on the inverse RST for the transformation coefficients, and an inverse first transformation unit that derives residual samples for the target block based on the inverse first transformation for the modified transformation coefficients. The inverse first transformation refers to the inverse transformation of the first transformation that was applied to the residual. In this document, deriving transformation coefficients based on a transformation may mean deriving transformation coefficients by applying the transformation.
[0116] FIG. 5 is a drawing for explaining an RST according to one embodiment of the present document.
[0117] In this specification, “target block” may mean the current block, residual block, or conversion block where coding is performed.
[0118] In an RST according to one embodiment, an N-dimensional vector may be mapped to an R-dimensional vector located in a different space to determine a reduced transformation matrix, wherein R is smaller than N. N may represent the square of the length of one side of the block to which the transformation is applied or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor may represent the value of R / N. The simplification factor may be referred to by various terms such as reduced factor, reduction factor, simplified factor, simple factor, etc. Meanwhile, R may be referred to as the reduced coefficient, but in some cases, the simplification factor may represent R. Additionally, in some cases, the simplification factor may represent the value of N / R.
[0119] In one embodiment, the simplification factor or simplification factor may be signaled via a bitstream, but the embodiment is not limited thereto. For example, a predefined value for the simplification factor or simplification factor may be stored in each encoding device (100) and decoding device (200), in which case the simplification factor or simplification factor may not be signaled separately.
[0120] The size of the simplified transformation matrix according to one embodiment is RxN, which is smaller than the size NxN of a conventional transformation matrix, and can be defined as shown in Equation 4 below.
[0121]
[0122] The matrix T in the Reduced Transform block shown in FIG. 5(a) is the matrix T of Equation 4. RxN It can mean that the simplified transformation matrix T for the residual samples of the target block as shown in Fig. 5(a). RxN When multiplied, transformation coefficients for the target block can be derived.
[0123] In one embodiment, when the size of the block to which the transformation is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), the RST according to FIG. 5(a) can be expressed as a matrix operation as shown in Equation 5 below. In this case, memory and multiplication operations can be reduced to approximately 1 / 4 by a simplification factor.
[0124] In this document, matrix operation can be understood as an operation in which a matrix is placed to the left of a column vector and the matrix is multiplied to obtain a column vector.
[0125]
[0126] In mathematical formula 5, r1 to r 64may represent residual samples for the target block, and more specifically, may be transformation coefficients generated by applying a linear transformation. Transformation coefficients c for the target block resulting from the operation of Equation 5 i can be derived, and c i The derivation process of can be as shown in mathematical formula 6.
[0127]
[0128] The result of the operation of mathematical formula 6, transformation coefficients c1 to c for the target block R This can be derived. That is, when R=16, the transformation coefficients c1 to c for the target block 16 This can be derived. If a regular transformation, rather than RST, were applied and a transformation matrix of size 64x64 (NxN) were multiplied by a 64x1 (Nx1) regular sample, 64 (N) transformation coefficients for the target block would have been derived, but because RST was applied, only 16 (R) transformation coefficients for the target block are derived. Since the total number of transformation coefficients for the target block is reduced from N to R, the amount of data transmitted by the encoding device (100) to the decoding device (200) is reduced, so the transmission efficiency between the encoding device (100) and the decoding device (200) can be increased.
[0129] When considering the size of the transformation matrix, the size of a standard transformation matrix is 64x64 (NxN), but the size of a simplified transformation matrix is reduced to 16x64 (RxN). Therefore, compared to performing a standard transformation, memory usage when performing RST can be reduced by the ratio of R / N. In addition, compared to the number of multiplication operations NxN when using a standard transformation matrix, using a simplified transformation matrix can reduce the number of multiplication operations by the ratio of R / N (RxN).
[0130] In one embodiment, the transformation unit (132) of the encoding device (100) can derive transformation coefficients for the target block by performing a first transformation and a second transformation based on RST on the residual samples for the target block. These transformation coefficients can be transmitted to the inverse transformation unit of the decoding device (200), and the inverse transformation unit (222) of the decoding device (200) can derive modified transformation coefficients based on the inverse RST (reduced secondary transform) for the transformation coefficients and derive residual samples for the target block based on the inverse first transformation for the modified transformation coefficients.
[0131] The size of the inverse RST matrix TNxR according to one embodiment is NxR, which is smaller than the size NxN of a conventional inverse transformation matrix, and is in a transpose relationship with the simplified transformation matrix TRxN shown in Equation 4.
[0132] Matrix T within the Reduced Inv. Transform block shown in Fig. 5 (b). t is the inverse RST matrix T RxN T It can mean (the superscript T signifies transpose). As shown in Fig. 5(b), the inverse RST matrix T for the transformation coefficients for the target block. RxN T When multiplied, modified transformation coefficients for the target block or residual samples for the target block can be derived. Inverse RST matrix T RxN T is (T RxN ) T NxR It can also be expressed as.
[0133] More specifically, when the inverse RST is applied as a second-order inverse transform, the inverse RST matrix T for the transform coefficients of the target block RxN TWhen multiplied, modified transformation coefficients for the target block can be derived. Meanwhile, inverse RST can be applied as an inverse first-order transformation, and in this case, when the inverse RST matrix TRxNT is multiplied for the transformation coefficients for the target block, residual samples for the target block can be derived.
[0134] In one embodiment, when the size of the block to which the inverse transformation is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), the RST according to FIG. 5 (b) can be expressed by a matrix operation as shown in Equation 7 below.
[0135]
[0136] In mathematical formula 7, c1 to c 16 can represent the transformation coefficients for the target block. r represents the modified transformation coefficients for the target block or the residual samples for the target block as a result of the operation of Equation 7. i can be derived, and r i The derivation process of can be as shown in mathematical formula 8.
[0137]
[0138] r1 to r representing the result of the operation of Equation 8, modified transformation coefficients for the target block or residual samples for the target block N This can be derived. When examined from the perspective of the size of the inverse transform matrix, the size of a standard inverse transform matrix is 64x64 (NxN), but the size of a simplified inverse transform matrix is reduced to 64x16 (NxR). Therefore, compared to performing a standard inverse transform, memory usage when performing an inverse RST can be reduced by the ratio of R / N. In addition, compared to the number of multiplication operations NxN when using a standard inverse transform matrix, using a simplified inverse transform matrix can reduce the number of multiplication operations by the ratio of R / N (NxR).
[0139] Meanwhile, for the 8x8 RST, a transformation set configuration as shown in Table 2 can also be applied. That is, the corresponding 8x8 RST can be applied according to the transformation set in Table 2. Since a single transformation set consists of two or three transformations (kernels) depending on the in-frame prediction mode, it can be configured to select one of up to four transformations, including the case where a second transformation is not applied. The transformation when a second transformation is not applied can be considered as the identity matrix being applied. Assuming indices 0, 1, 2, and 3 are assigned to the four transformations respectively (for example, index 0 can be assigned to the identity matrix, i.e., the case where a second transformation is not applied), the transformation to be applied can be specified by signaling a syntax element called the transformation index or lfnst index for each transformation coefficient block. That is, for an 8x8 top-left block via the transformation index, an 8x8 RST can be specified in the RST configuration, or an 8x8 lfnst can be specified if LFNST is applied. The 8x8 lfnst and 8x8 RST point to a transformation that can be applied to an 8x8 area contained within the transformation factor block when both W and H of the target block are greater than or equal to 8, and that 8x8 area may be the top-left 8x8 area within the transformation factor block. Similarly, the 4x4 lfnst and 4x4 RST point to a transformation that can be applied to a 4x4 area contained within the transformation factor block when both W and H of the target block are greater than or equal to 4, and that 4x4 area may be the top-left 4x4 area within the transformation factor block.
[0140] Meanwhile, according to one embodiment of the present document, in the transformation of the encoding process, instead of a 16 x 64 transformation kernel matrix for the 64 data constituting an 8 x 8 area, only 48 data can be selected and a maximum 16 x 48 transformation kernel matrix can be applied. Here, “maximum” means that for an m x 48 transformation kernel matrix capable of generating m coefficients, the maximum value of m is 16. That is, when performing RST by applying an m x 48 transformation kernel matrix (m ≤ 16) to an 8 x 8 area, 48 data can be input and m coefficients can be generated. When m is 16, 48 data are input and 16 coefficients are generated. That is, assuming that 48 data form a 48 x 1 vector, a 16 x 1 vector can be generated by multiplying the 16 x 48 matrix and the 48 x 1 vector in order. At this time, a 48 x 1 vector can be constructed by appropriately arranging 48 data points that make up an 8 x 8 area. For example, a 48 x 1 vector can be constructed based on 48 data points that make up the area excluding the bottom-right 4 x 4 area of the 8 x 8 area. At this time, if matrix operations are performed by applying a maximum 16 x 48 transformation kernel matrix, 16 modified transformation coefficients are generated, and the 16 modified transformation coefficients can be placed in the top-left 4 x 4 area according to the scanning order, while the top-right 4 x 4 area and the bottom-left 4 x 4 area can be filled with 0.
[0141] For the inverse transformation of the decoding process, the transposed matrix of the transformation kernel matrix described above may be used. That is, when an inverse RST or LFNST is performed as an inverse transformation process performed by a decoding device, the input coefficient data to which the inverse RST is to be applied is configured as a one-dimensional vector according to a predetermined array order, and the modified coefficient vector obtained by multiplying the one-dimensional vector by the corresponding inverse RST matrix from the left can be arranged in a two-dimensional block according to a predetermined array order.
[0142] In summary, during the transformation process, when RST or LFNST is applied to an 8x8 area, matrix operations are performed between the 48 transformation coefficients in the top-left, top-right, and bottom-left areas (excluding the bottom-right area of the 8x8 area) and a 16x48 transformation kernel matrix. For the matrix operations, the 48 transformation coefficients are input as a one-dimensional array. When these matrix operations are performed, 16 modified transformation coefficients are derived, and the modified transformation coefficients can be arranged in the top-left area of the 8x8 area.
[0143] Conversely, during the inverse transformation process, when an inverse RST or LFNST is applied to an 8x8 area, the 16 transformation coefficients corresponding to the top-left corner of the 8x8 area are input in the form of a one-dimensional array according to the scanning order and can be matrix-operated with a 48 x 16 transformation kernel matrix. That is, the matrix operation in this case can be represented as (48 x 16 matrix) * (16 x 1 transformation coefficient vector) = (48 x 1 modified transformation coefficient vector). Here, since an nx1 vector can be interpreted as having the same meaning as an nx1 matrix, it may also be denoted as an nx1 column vector. Additionally, * signifies matrix multiplication. When this matrix operation is performed, 48 modified transformation coefficients can be derived, and these 48 modified transformation coefficients can be arranged in the top-left, top-right, and bottom-left areas, excluding the bottom-right area of the 8x8 area.
[0144] Meanwhile, when the second inverse transformation is performed based on RST, the inverse transformation unit (135) of the encoding device (100) and the inverse transformation unit (222) of the decoding device (200) may include an inverse RST unit that derives modified transformation coefficients based on the inverse RST for the transformation coefficients, and an inverse first transformation unit that derives residual samples for the target block based on the inverse first transformation for the modified transformation coefficients. The inverse first transformation refers to the inverse transformation of the first transformation that was applied to the residual. In this document, deriving transformation coefficients based on a transformation may mean deriving transformation coefficients by applying the transformation.
[0145] The aforementioned non-separable transformation, LFNST, is examined in detail as follows. LFNST may include a forward transformation by an encoding device and an inverse transformation by a decoding device.
[0146] The encoding device applies a primary (core) transform and then takes the resulting product (or part of the result) as input to apply a secondary transform.
[0147]
[0148] In the above mathematical equation 9, x and y are the input and output of the quadratic transformation, respectively, and G is the matrix representing the quadratic transformation, where the transform basis vectors consist of column vectors. In the case of the inverse LFNST, when the dimension of the transformation matrix G is denoted as [number of rows x number of columns], in the case of the forward LFNST, G is obtained by taking the transform of matrix G. T It becomes the dimension of.
[0149] For the inverse LFNST, the dimensions of matrix G are [ 48 x 16 ], [ 48 x 8 ], [ 16 x 16 ], and [16 x 8 ], and the [48 x 8] matrix and the [16 x 8 ] matrix are submatrices obtained by sampling the 8 transformation basis vectors from the left of the [ 48 x 16 ] matrix and the [ 16 x 16 ] matrix, respectively.
[0150] On the other hand, in the case of forward LFNST, matrix G T The dimensions of are [ 16 x 48 ], [ 8 x 48 ], [ 16 x 16 ], and [ 8 x 16 ], and the [ 8 x 48 ] matrix and the [ 8 x 16 ] matrix are submatrices obtained by sampling the top 8 transformation basis vectors of the [ 16 x 48 ] matrix and the [ 16 x 16 ] matrix, respectively.
[0151] Therefore, for a forward LFNST, the input x can be a [ 48 x 1 ] vector or a [ 16 x 1 ] vector, and the output y can be a [ 16 x 1 ] vector or an [ 8 x 1 ] vector. Since the output of a forward first-order transformation in video coding and decoding is two-dimensional (2D) data, in order to construct a [ 48 x 1 ] vector or a [ 16 x 1 ] vector as the input x, the 2D data output of the forward transformation must be appropriately arranged to construct a one-dimensional vector.
[0152] FIG. 6 illustrates the sequence of arranging output data of a forward linear transformation into a one-dimensional vector according to one example. The left diagram in FIG. 6 (a) and (b) shows the sequence for creating a [ 48 x 1 ] vector, and the right diagram in FIG. 6 (a) and (b) shows the sequence for creating a [ 16 x 1 ] vector. In the case of LFNST, a one-dimensional vector x can be obtained by sequentially arranging 2D data in the same order as FIG. 6 (a) and (b).
[0153] The arrangement direction of the output data of this forward first-order transformation can be determined according to the intra-prediction mode of the current block. For example, if the intra-prediction mode of the current block is horizontal with respect to the diagonal direction, the output data of the forward first-order transformation can be arranged in the order of (a) in FIG. 6, and if the intra-prediction mode of the current block is vertical with respect to the diagonal direction, the output data of the forward first-order transformation can be arranged in the order of (b) in FIG. 6.
[0154] For example, a different ordering from that of FIGS. 6 (a) and (b) can be applied, and to obtain the same result (y vector) as when the ordering of FIGS. 6 (a) and (b) is applied, the column vectors of matrix G can be rearranged according to that ordering. That is, the column vectors of G can be rearranged so that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0155] Since the output y derived from Equation 9 is a one-dimensional vector, if a configuration that processes the result of a forward quadratic transformation as input, for example, a configuration that performs quantization or residual coding, requires two-dimensional data as input data, the output y vector of Equation 9 must be appropriately reconfigured into 2D data.
[0156] Figure 7 is a diagram illustrating the sequence of arranging output data of a forward quadratic transformation into two-dimensional blocks according to one example.
[0157] In the case of LFNST, it can be placed in a 2D block according to a predetermined scan order. FIG. 7(a) shows that when the output y is a [ 16 x 1 ] vector, the output value is placed in 16 positions of the 2D block according to the diagonal scan order. FIG. 7(b) shows that when the output y is an [ 8 x 1 ] vector, the output value is placed in 8 positions of the 2D block according to the diagonal scan order, and the remaining 8 positions are filled with 0. X in FIG. 7(b) indicates that it is filled with 0.
[0158] In other examples, the order in which the output vector y is processed by a configuration that performs quantization or residual coding may be performed according to a preset order, so the output vector y may not be placed in a 2D block as shown in FIG. 7. However, in the case of residual coding, data coding may be performed in units of 2D blocks (e.g., 4x4) such as CG (Coefficient Group), and in this case, the data may be arranged according to a specific order, such as the diagonal scan order in FIG. 7.
[0159] Meanwhile, the decoding device can construct a one-dimensional input vector y by arranging the two-dimensional data output through an inverse quantization process, etc., according to a preset scan order for reverse conversion. The input vector y can be output as an input vector x by the following mathematical formula.
[0160]
[0161] For the inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [ 16 x 1 ] vector or an [ 8 x 1 ] vector, by the matrix G. For the inverse LFNST, the output vector x can be a [ 48 x 1 ] vector or a [ 16 x 1 ] vector.
[0162] The output vector x is placed in a two-dimensional block according to the order shown in FIG. 6 and arranged as two-dimensional data, and this two-dimensional data becomes the input data (or part of the input data) for the inverse first-order transformation.
[0163] Therefore, the inverse second-order transformation is the opposite of the forward second-order transformation process overall, and in the case of the inverse transformation, unlike in the forward direction, the inverse second-order transformation is applied first, followed by the inverse first-order transformation.
[0164] In the reverse LFNST, one of eight [48 x 16] matrices and eight [16 x 16] matrices can be selected as the transformation matrix G. Whether to apply the [48 x 16] matrices or the [16 x 16] matrices is determined by the size and shape of the block.
[0165] Additionally, eight matrices can be derived from four transformation sets as described in Table 2 above, and each transformation set can consist of two matrices. Which of the four transformation sets to use is determined by the intra-prediction mode, and more specifically, the transformation set is determined based on the extended intra-prediction mode value, taking into account the Wide Angle Intra Prediction (WAIP) mode. Which of the two matrices constituting the selected transformation set to choose is determined through index signaling. More specifically, the transmitted index values can be 0, 1, or 2, where 0 indicates that LFNST is not applied, and 1 or 2 can indicate either of the two transformation matrices constituting the selected transformation set based on the intra-prediction mode value.
[0166] Meanwhile, as described above, whether to apply the [ 48 x 16 ] matrix or the [ 16 x 16 ] matrix to LFNST is determined by the size and shape of the block to be transformed.
[0167] FIG. 8 is a diagram illustrating the shape of a block to which LFNST is applied. FIG. 8 shows a 4 x 4 block, (b) 4 x 8 and 8 x 4 blocks, (c) 4 x N or N x 4 blocks where N is 16 or greater, (d) 8 x 8 blocks, and (e) M x N blocks where M ≥ 8, N ≥ 8, and N > 8 or M > 8.
[0168] In FIG. 8, the blocks with thick borders indicate the areas where LFNST is applied. For the blocks in FIG. 8 (a) and (b), LFNST is applied to the top-left 4x4 area, and for the block in FIG. 8 (c), LFNST is applied to each of the two consecutively arranged top-left 4x4 areas. Since LFNST is applied in units of 4x4 areas in FIG. 8 (a), (b), and (c), this LFNST will be referred to as “4x4 LFNST” below, and the transformation matrix may be a [ 16 x 16 ]. or [ 16 x 8 ]. based on the matrix dimensions of G in Equations 9 and 10, a [ 16 x 16 ].
[0169] More specifically, for the 4x4 block (4x4 TU or 4x4 CU) in Fig. 8 (a), a [ 16 x 8 ] matrix is applied, and for the blocks in Fig. 8 (b) and (c), a [ 16 x 16 ] matrix is applied. This is to match the computational complexity for the worst case to 8 multiplications per sample.
[0170] For Figures 8 (d) and (e), an LFNST is applied to the top-left 8x8 region, and this LFNST will be referred to as the “8x8 LFNST” below. A [ 48 x 16 ] or [ 48 x 8 ] matrix may be applied as the corresponding transformation matrix. In the case of the forward LFNST, since the [ 48 x 1 ] vector (the x vector of Equation 9) is input as input data, not all sample values of the top-left 8x8 region are used as input values for the forward LFNST. That is, as can be seen in the left-hand order of Figure 6 (a) or the left-hand order of Figure 6 (b), the bottom-right 4x4 block is left as is, and the [ 48 x 1 ] vector can be constructed based on the samples belonging to the remaining three 4x4 blocks.
[0171] A [ 48 x 8 ] matrix can be applied to the 8x8 block (8x8 TU or 8x8 CU) in Fig. 8 (d), and a [ 48 x 16 ] matrix can be applied to the 8x8 block in Fig. 8 (e). This is also to match the computational complexity for the worst case to 8 multiplications per sample.
[0172] Depending on the block shape, when the corresponding forward LFNST (4x4 LFNST or 8x8 LFNST) is applied, 8 or 16 output data (y vector in Equation 9, [ 8 x 1 ] or [ 16 x 1 ] vector) are generated, and in the forward LFNST, due to the characteristics of the matrix GT, the number of output data is equal to or less than the number of input data.
[0173] FIG. 9 is a diagram illustrating the arrangement of output data of a forward LFNST according to one example, showing blocks in which output data of a forward LFNST is arranged according to the block shape.
[0174] The shaded area at the top-left of the block shown in Fig. 9 corresponds to the region where the output data of the forward LFNST is located; the locations marked with 0 represent samples filled with zero values, and the remaining region represents an area that is not modified by the forward LFNST. In the region not modified by the LFNST, the output data of the forward first-order transformation remains unchanged.
[0175] As described above, since the dimension of the transformation matrix applied varies depending on the block shape, the number of output data also varies. As shown in FIG. 9, the output data of the forward LFNST may not completely fill the top-left 4x4 block. In the case of FIG. 9 (a) and (d), the [ 16 x 8 ] matrix and the [ 48 x 8 ] matrix are applied to the block or a portion of the block indicated by the bold line, respectively, to generate an [ 8 x 1 ] vector as the output of the forward LFNST. That is, according to the scan order shown in FIG. 7 (b), only 8 output data points are filled as in FIG. 9 (a) and (d), and the remaining 8 positions may be filled with 0. In the case of the LFNST applied block in FIG. 8 (d), the two 4x4 blocks in the top-right and bottom-left adjacent to the top-left 4x4 block are also filled with 0 values, as in FIG. 9 (d).
[0176] As described above, the LFNST index is basically signaled to specify whether to apply LFNST and the transformation matrix to be applied. As illustrated in Fig. 9, when LFNST is applied, since the number of output data of the forward LFNST may be equal to or less than the number of input data, an area filled with zero values occurs as follows.
[0177] 1) Positions after the 8th in the scan order within the top-left 4x4 block as shown in Fig. 9(a), i.e., samples from the 9th to the 16th
[0178] 2) As in FIG. 9 (d) and (e), a [ 16 x 48 ] matrix or an [ 8 x 48 ] matrix is applied to two 4x4 blocks adjacent to the top-left 4x4 block, or the second and third 4x4 blocks in scan order.
[0179] Therefore, by checking the areas of 1) and 2) above, if non-zero data exists, it is certain that LFNST has not been applied, so the signaling of the corresponding LFNST index can be omitted.
[0180] For example, in the case of LFNST adopted in the VVC standard, signaling of the LFNST index is performed after residual coding, so the encoding device can determine whether non-zero data (effective count) exists for all locations within the TU or CU block through residual coding. Therefore, the encoding device can determine whether to perform signaling for the LFNST index based on the existence of non-zero data, and the decoding device can determine whether to parse the LFNST index. If non-zero data does not exist in the area specified in 1) and 2) above, signaling of the LFNST index is performed.
[0181] Meanwhile, the following simplification methods can be applied to the adopted LFNST.
[0182] (i) Depending on the example, the number of output data for the forward LFNST can be limited to a maximum of 16.
[0183] In the case of Fig. 8(c), 4x4 LFNSTs can be applied to each of the two 4x4 regions adjacent to the top-left corner, and up to 32 LFNST output data can be generated. If the number of output data for the forward LFNST is limited to a maximum of 16, then for the 4xN / Nx4 (N≥16) blocks (TU or CU), 4x4 LFNSTs are applied only to the one 4x4 region located at the top-left corner, and LFNSTs can be applied only once to all blocks in Fig. 8. This simplifies the implementation of image coding.
[0184] (ii) Depending on the example, zero-out may be additionally applied to regions where LFNST is not applied. In this document, zero-out may mean filling the values of all locations belonging to a specific region with zero values. That is, zero-out may be applied even to regions that retain the results of the forward first-order transformation and are not changed by LFNST. As mentioned above, since LFNST is divided into 4x4 LFNST and 8x8 LFNST, zero-out can be classified into two types ((ii)-(A) and (ii)-(B)) as follows.
[0185] (ii)-(A) When 4x4 LFNST is applied, the area where 4x4 LFNST is not applied can be zeroed out. FIG. 10 is a drawing illustrating zeroing out in a block where 4x4 LFNST is applied according to one example.
[0186] As shown in FIG. 10, for the blocks to which 4x4 LFNST is applied, that is, for the blocks of FIG. 9 (a), (b) and (c), all areas where LFNST is not applied can be filled with 0.
[0187] Meanwhile, (d) of FIG. 10 shows that when the maximum value of the number of output data of the forward LFNST is limited to 16 according to one example, zero-out is performed on the remaining blocks to which the 4x4 LFNST is not applied.
[0188] (ii)-(B) When 8x8 LFNST is applied, the area where 8x8 LFNST is not applied can be zeroed out. FIG. 11 is a drawing illustrating zeroing out in a block where 8x8 LFNST is applied according to one example.
[0189] As shown in FIG. 11, for the blocks to which 8x8 LFNST is applied, that is, for the blocks of FIG. 9 (d) and (e), all areas where LFNST is not applied can be filled with 0.
[0190] (iii) Due to the zero-out proposed in (ii) above, the area filled with zeros when LFNST is applied may differ. Therefore, depending on the zero-out proposed in (ii) above, it is possible to check whether non-zero data exists over a wider area than in the case of LFNST in Fig. 9.
[0191] For example, when applying (ii)-(B), after checking whether non-zero data exists in the area filled with zero values in (d) and (e) of FIG. 9 and additionally in the area filled with zeros in FIG. 11, signaling for the LFNST index can be performed only if there is no non-zero data.
[0192] Of course, even if the zero-out proposed in (ii) above is applied, it is possible to check whether non-zero data exists in the same way as the existing LFNST index signaling. That is, for the blocks filled with zeros in FIG. 9, it is possible to check whether non-zero data exists and apply LFNST index signaling. In this case, zero-out is performed only in the encoding device, and the decoding device does not assume the zero-out, that is, it is possible to perform LFNST index parsing by checking only whether non-zero data exists for the areas explicitly marked as zeros in FIG. 9.
[0193] Various embodiments can be derived by applying combinations of the simplification methods ((i), (ii)-(A), (ii)-(B), (iii)) for the above LFNST. Of course, the combinations of the above simplification methods are not limited to the embodiments below, and any combination can be applied to the LFNST.
[0194] Examples
[0195] - Limit the number of output data for the forward LFNST to a maximum of 16 → (i)
[0196] - When 4x4 LFNST is applied, zero out all areas where 4x4 LFNST is not applied → (ii)-(A)
[0197] - When 8x8 LFNST is applied, zero out all areas where 8x8 LFNST is not applied → (ii)-(B)
[0198] - Check whether non-zero data exists in the areas filled with existing zero values and in the areas filled with zeros due to additional zero-outs ((ii)-(A), (ii)-(B)), and signal LFNST indexing only if no non-zero data exists → (iii)
[0199] In the above embodiment, when LFNST is applied, the area where non-zero output data may exist is limited to within the upper-left 4x4 area. More specifically, in the case of FIG. 10 (a) and FIG. 11 (a), the 8th position in the scan order becomes the last position where non-zero data may exist, and in the case of FIG. 10 (b) and (d) and FIG. 11 (b), the 16th position in the scan order (i.e., the lower-right edge position of the upper-left 4x4 block) becomes the last position where non-zero data may exist.
[0200] Therefore, when LFNST is applied, the decision to signal the LFNST index can be made after checking whether there is non-zero data at a location where the residual coding process is not allowed (a location beyond the last location).
[0201] In the case of the zero-out method proposed in (ii), the amount of computation required to perform the entire transformation process can be reduced because the number of data points ultimately generated when both the first-order transformation and LFNST are applied is reduced. That is, when LFNST is applied, zero-out is applied even to the forward first-order transformation output data existing in the region where LFNST is not applied; therefore, there is no need to generate data for the zero-out region from the beginning of the forward first-order transformation. Consequently, the amount of computation required for generating such data can be saved. The additional effects of the zero-out method proposed in (ii) can be summarized as follows.
[0202] First, as mentioned above, the amount of computation required to perform the entire conversion process is reduced.
[0203] In particular, applying (ii)-(B) can reduce the amount of computation for the worst-case scenario, thereby making the transformation process lighter. To elaborate, while a large amount of computation is generally required to perform large-scale first transformations, applying (ii)-(B) can reduce the number of data points resulting from forward LFNST execution to 16 or fewer, and the reduction in transformation computation is further enhanced as the total block (TU or CU) size increases.
[0204] Second, the amount of computation required for the entire conversion process is reduced, which can lower the power consumption required to perform the conversion.
[0205] Third, it reduces the latency associated with the conversion process.
[0206] Secondary transformations such as LFNST add computational load to existing primary transformations, thereby increasing the total latency associated with performing the transformation. In particular, in the case of intra prediction, since the reconstruction data of neighboring blocks is used during the prediction process, the increase in latency caused by secondary transformations during encoding leads to an increase in latency until reconstruction, which can result in an overall increase in latency for intra prediction encoding.
[0207] However, by applying the zero-out method presented in (ii), the delay time for the first conversion when applying LFNST can be significantly reduced, so the delay time for the entire conversion process remains the same or is even reduced, allowing the encoding device to be implemented more simply.
[0208] Meanwhile, conventional intra prediction treated the block currently to be encoded as a single encoding unit and performed encoding without division. However, Intra Sub-Paritions (ISP) coding refers to performing intra prediction coding by dividing the block currently to be encoded in a horizontal or vertical direction. At this time, encoding / decoding is performed on the divided block units to generate a restored block, and the restored block can be used as a reference block for the next divided block. For example, during ISP coding, a single coding block may be divided into two or four sub-blocks for coding, and in ISP, intra prediction is performed by referencing the restored pixel value of a sub-block located to the adjacent left or adjacent above for one sub-block. Hereinafter, the term “coding” used may be used as a concept that includes both coding performed by an encoding device and decoding performed by a decoding device.
[0209] ISP divides blocks predicted by Luma Intra into 2 or 4 sub-partitions in the vertical or horizontal direction, depending on the block size. For example, the minimum block size to which ISP can be applied is 4 x 8 or 8 x 4. If the block size is larger than 4 x 8 or 8 x 4, the block is divided into 4 sub-partitions.
[0210] When applying ISP, sub-blocks are coded sequentially according to the partitioning pattern, for example, horizontally or vertically, or from left to right or top to bottom. Coding for the next sub-block can proceed only after the restoration process is performed following the inverse transformation and intra prediction of a single sub-block. For the leftmost or topmost sub-block, the restored pixels of an already coded block are referenced, as in conventional intra prediction methods. Additionally, for each edge of a subsequent internal sub-block that is not adjacent to the previous sub-block, the restored pixels of an already coded adjacent block are referenced, as in conventional intra prediction methods, to derive the reference pixels adjacent to that edge.
[0211] In ISP coding mode, all sub-blocks can be coded with the same intra-predicted mode, and a flag indicating whether to use ISP coding and a flag indicating in which direction (horizontal or vertical) to split can be signaled. At this time, the number of sub-blocks can be adjusted to 2 or 4 depending on the block shape, and if the size (width x height) of a single sub-block is less than 16, splitting into that sub-block is not allowed, or the application of ISP coding itself can be restricted.
[0212] Meanwhile, in ISP prediction mode, one coding unit is divided into two or four partition blocks, i.e., sub-blocks, and the same intra-frame prediction mode is applied to the two or four divided partition blocks.
[0213] As described above, the division direction can be either horizontal (when an MxN coding unit with horizontal and vertical lengths M and N is divided horizontally, it is divided into Mx(N / 2) blocks if divided into 2 and into Mx(N / 4) blocks if divided into 4) or vertical (when an MxN coding unit is divided vertically, it is divided into (M / 2)xN blocks if divided into 2 and into (M / 4)xN blocks if divided into 4). When divided horizontally, the partition blocks are coded in order from top to bottom, and when divided vertically, the partition blocks are coded in order from left to right. The partition block currently being coded can be predicted by referencing the restored pixel value of the upper (left) partition block in the case of horizontal (vertical) division.
[0214] Transformation can be applied to the residual signal generated by the ISP prediction method in partition block units. Based on the forward direction, MTS (Multiple Transform Selection) technology based on the DST-7 / DCT-8 combination as well as the existing DCT-2 can be applied to the primary transform (core transform or primary transform), and the forward LFNST (Low Frequency Non-Separable Transform) can be applied to the transformation coefficients generated according to the primary transform to generate the final modified transformation coefficients.
[0215] In other words, LFNST can be applied to partition blocks that have been divided by applying the ISP prediction mode, and as mentioned above, the same intra prediction mode is applied to the divided partition blocks. Therefore, when selecting the LFNST set derived based on the intra prediction mode, the derived LFNST set can be applied to all partition blocks. That is, since the same intra prediction mode is applied to all partition blocks, the same LFNST set can be applied to all partition blocks.
[0216] Meanwhile, depending on the example, LFNST can be applied only to transformation blocks where both the width and height are 4 or greater. Therefore, if the width or height of a partition block divided according to the ISP prediction method is less than 4, LFNST is not applied and the LFNST index is not signaled. Additionally, when LFNST is applied to each partition block, that partition block can be considered as a single transformation block. Of course, if the ISP prediction method is not applied, LFNST can be applied to the coding block.
[0217] The specifics of applying LFNST to each partition block are as follows.
[0218] According to one example, after applying forward LFNST to individual partition blocks, zero-out may be applied to the top-left 4x4 area, leaving only up to 16 (8 or 16) coefficients according to the scanning order of the transformation coefficients, and filling all remaining positions and areas with 0 values.
[0219] Alternatively, depending on the example, if the length of one side of the partition block is 4, LFNST may be applied only to the top-left 4x4 area, and if the length of all sides of the partition block, i.e., width and height, is 8 or more, LFNST may be applied to the remaining 48 coefficients excluding the bottom-right 4x4 area inside the top-left 8x8 area.
[0220] Alternatively, depending on one example, to match the worst-case computational complexity to 8 multiplications per sample, only 8 transformation factors can be output after applying forward LFNST when each partition block is 4x4 or 8x8. That is, if the partition block is 4x4, an 8x16 matrix can be applied as the transformation matrix, and if the partition block is 8x8, an 8x48 matrix can be applied as the transformation matrix.
[0221] Meanwhile, in the current VVC standard, LFNST index signaling is performed at the coding unit level. Therefore, if the ISP prediction mode is active and LFNST is applied to all partition blocks, the same LFNST index value can be applied to those partition blocks. That is, once an LFNST index value is transmitted at the coding unit level, that LFNST index can be applied to all partition blocks within the coding unit. As described above, the LFNST index value can have values of 0, 1, or 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate two transformation matrices existing within a single LFNST set when LFNST is applied.
[0222] As described above, the LFNST set is determined by the intra prediction mode, and in the case of the ISP prediction mode, since all partition blocks within the coding unit are predicted using the same intra prediction mode, the partition blocks can refer to the same LFNST set.
[0223] As another example, LFNST index signaling is still performed at the coding unit level, but in the case of ISP prediction mode, instead of uniformly deciding whether to apply LFNST to all partition blocks, it is possible to determine whether to apply the LFNST index value signaled at the coding unit level to each partition block or not to apply LFNST through a separate condition. Here, the separate condition can be signaled in the form of a flag for each partition block through a bitstream, and if the flag value is 1, the LFNST index value signaled at the coding unit level is applied, and if the flag value is 0, LFNST is not applied.
[0224] The following describes a method to maintain computational complexity for the worst case when applying LFNST to ISP mode.
[0225] In ISP mode, the application of LFNST can be limited to keep the number of multiplications per sample (or per coefficient, per position) below a certain value. Depending on the size of the partition block, LFNST can be applied as follows to keep the number of multiplications per sample (or per coefficient, per position) below 8.
[0226] 1. If both the width and height of the partition block are 4 or greater, the same method as the worst-case computational complexity control method for LFNST in the current VVC standard can be applied.
[0227] That is, when the partition block is a 4x4 block, instead of a 16x16 matrix, an 8x16 matrix obtained by sampling the top 8 rows from a 16x16 matrix can be applied in the forward direction, and a 16x8 matrix obtained by sampling the left 8 columns from a 16x16 matrix can be applied in the reverse direction. Also, when the partition block is an 8x8 block, instead of a 16x48 matrix, an 8x48 matrix obtained by sampling the top 8 rows from a 16x48 matrix can be applied in the forward direction, and instead of a 48x16 matrix, a 48x8 matrix obtained by sampling the left 8 columns from a 48x16 matrix can be applied in the reverse direction.
[0228] For 4xN or Nx4 (N > 4) blocks, when performing a forward transformation, a 16x16 matrix is applied only to the top-left 4x4 block, and the resulting 16 coefficients are placed in the top-left 4x4 area, while the remaining areas can be filled with 0 values. Additionally, when performing a reverse transformation, the 16 coefficients located in the top-left 4x4 block are placed according to scanning order to form an input vector, and then a 16x16 matrix is multiplied to generate 16 output data. The generated output data is placed in the top-left 4x4 area, and the remaining areas excluding the top-left 4x4 area can be filled with 0.
[0229] For 8xN or Nx8 (N > 8) blocks, when performing a forward transformation, a 16x48 matrix can be applied only to the ROI area within the top-left 8x8 block (the area remaining after excluding the bottom-right 4x4 block from the top-left 8x8 block), and the 16 resulting coefficients can be placed in the top-left 4x4 area, while the remaining areas can be filled with zero values. Additionally, when performing a reverse transformation, the 16 coefficients located in the top-left 4x4 block can be placed according to the scanning order to form an input vector, and then a 48x16 matrix can be multiplied to generate 48 output data. The generated output data can be filled in the aforementioned ROI area, and the remaining areas can be filled with zero values.
[0230] As another example, to keep the number of multiplications per sample (or per coefficient, per position) below a certain value, the number of multiplications per sample (or per coefficient, per position) can be kept at 8 or less based on the ISP coding unit size rather than the ISP partition block size. If there is only one block among the ISP partition blocks that satisfies the condition for applying LFNST, the complexity calculation for the worst-case LFNST can be applied based on the corresponding coding unit size rather than the partition block size. For example, if a lumens coding block for a coding unit is divided into four 4x4 partition blocks and coded by ISP, and two of the partition blocks do not have non-zero transformation coefficients, the other two partition blocks can be configured to generate 16 transformation coefficients (based on the encoder) instead of 8 each.
[0231] Below, we examine the method of signaling the LFNST index in ISP mode.
[0232] As described above, the LFNST index can have values of 0, 1, or 2, where 0 indicates that LFNST is not applied, and 1 or 2 indicates either of the two LFNST kernel matrices included in the selected LFNST set. LFNST is applied based on the LFNST kernel matrix selected by the LFNST index. The method by which the LFNST index is transmitted in the current VVC standard is described as follows.
[0233] 1. An LFNST index can be transmitted once per coding unit (CU), and in the case of a dual-tree, individual LFNST indices can be signaled for the luminance block and the chroma block, respectively.
[0234] 2. If the LFNST index is not signaled, the LFNST index value is set to the default value of 0 (infer). The cases in which the LFNST index value is inferred as 0 are as follows.
[0235] A. In modes where transformations are not applied (e.g., transform skip, BDPCM, lossless coding, etc.)
[0236] B. Cases where the primary transformation is not DCT-2 (DST7 or DCT8), i.e., cases where the horizontal or vertical transformation is not DCT-2
[0237] C. LFNST cannot be applied if the horizontal or vertical length of the coding unit's luma block exceeds the size of the maximum luma conversion that can be converted, for example, if the size of the coding block's luma block is 128x16 and the maximum luma conversion that can be converted is 64.
[0238] In the case of a dual tree, it is determined whether the maximum luminance transform size is exceeded for each coding unit for the luminance component and the chroma component, respectively. That is, for the luminance block, it is checked whether the maximum convertible luminance transform size is exceeded, and for the chroma block, it is checked whether the width and height of the corresponding luminance block for the color format and the maximum convertible luminance transform size are exceeded. For example, if the color format is 4:2:0, the width and height of the corresponding luminance block become twice the size of the corresponding chroma block, and the transform size of the corresponding luminance block becomes twice the size of the corresponding chroma block. As another example, if the color format is 4:4:4, the width and height of the corresponding luminance block and the transform size are the same as those of the corresponding chroma block.
[0239] The 64-length transformation or the 32-length transformation means a transformation applied to a width or height having a length of 64 or 32, respectively, and "transformation size" may mean the corresponding length of 64 or 32.
[0240] In the case of a single tree, for a luma block, check whether the horizontal or vertical length exceeds the maximum convertible luma conversion block size; if it does, LFNST index signaling can be omitted.
[0241] D. An LFNST index can be transmitted only when both the width and height of the coding unit are 4 or greater.
[0242] In the case of a dual tree, the LFNST index can be signaled only when both the horizontal and vertical lengths for the corresponding component (i.e., the luminance or chroma component) are 4 or greater.
[0243] In the case of a single tree, the LFNST index can be signaled when both the horizontal and vertical lengths of the luma component are 4 or greater.
[0244] E. If the last non-zero coefficient position is not the DC position (top-left position of the block), if it is a dual-tree type lumens block, if the last non-zero coefficient position is not the DC position, the LFNST index is transmitted. If it is a dual-tree type chroma block, if either the last non-zero coefficient position for Cb or the last non-zero coefficient position for Cr is not the DC position, the corresponding LNFST index is transmitted.
[0245] For single tree type, if the position of the last non-zero coefficient of any of the luma component, Cb component, or Cr component is not a DC position, the LFNST index is transmitted.
[0246] Here, if the CBF (coded block flag) value indicating whether a transformation factor exists for a single transformation block is 0, the position of the last non-zero factor for that transformation block is not checked to determine whether to perform LFNST index signaling. In other words, if the CBF value is 0, the transformation is not applied to that block, so the position of the last non-zero factor may not be considered when checking the condition for LFNST index signaling.
[0247] For example, 1) if it is a dual tree type and is a luminance component, if the corresponding CBF value is 0, the LFNST index is not signaled; 2) if it is a dual tree type and is a chroma component, if the CBF value for Cb is 0 and the CBF value for Cr is 1, only the location of the last non-zero coefficient for Cr is checked and the corresponding LFNST index is transmitted; and 3) if it is a single tree type, the location of the last non-zero coefficient is checked only for the components where each CBF value is 1 for luminance, Cb, and Cr.
[0248] F. If it is confirmed that an LFNST transform factor exists in a location where it is not possible for an LFNST transform factor to exist, LFNST index signaling may be omitted. For 4x4 transform blocks and 8x8 transform blocks, according to the transform factor scanning order in the VVC standard, LFNST transform factors may exist in 8 locations starting from the DC position, and all remaining locations are filled with 0. Additionally, for non-4x4 transform blocks and 8x8 transform blocks, according to the transform factor scanning order in the VVC standard, LFNST transform factors may exist in 16 locations starting from the DC position, and all remaining locations are filled with 0.
[0249] Therefore, if there is a non-zero conversion factor in the area where the above zero value must be filled after residual coding, LFNST index signaling can be omitted.
[0250] Meanwhile, the ISP mode may be applied only to the luminance block, or to both the luminance and chroma blocks. As mentioned above, when ISP prediction is applied, the corresponding coding unit is divided into two or four partition blocks for prediction, and transformations can be applied to each of these partition blocks. Therefore, when determining the conditions for signaling the LFNST index at the coding unit level, the fact that LFNST can be applied to each of these partition blocks must be taken into account. Furthermore, if the ISP prediction mode is applied only to a specific component (e.g., the luminance block), the LFNST index should be signaled considering that it is divided into partition blocks only for that specific component. The possible LFNST index signaling methods when in ISP mode are summarized as follows.
[0251] 1. An LFNST index can be transmitted once per coding unit (CU), and in the case of a dual-tree, individual LFNST indices can be signaled for the luminance block and the chroma block, respectively.
[0252] 2. If the LFNST index is not signaled, the LFNST index value is set to the default value of 0 (infer). The cases in which the LFNST index value is inferred as 0 are as follows.
[0253] A. In modes where transformations are not applied (e.g., transform skip, BDPCM, lossless coding, etc.)
[0254] B. LFNST cannot be applied if the horizontal or vertical length of the coding unit's luma block exceeds the size of the maximum luma conversion that can be converted, for example, if the size of the coding block's luma block is 128x16 and the maximum luma conversion size that can be converted is 64.
[0255] Instead of coding units, the signaling of the LFNST index can be determined based on the size of the partition block. That is, if the width or height of the partition block for the corresponding luma block exceeds the size of the maximum luma transformation that can be transformed, the LFNST index signaling can be omitted and the LFNST index value can be inferred as 0.
[0256] In the case of a dual tree, it is determined whether the maximum transform block size is exceeded for each coding unit or partition block for the luminance component and each for the chroma component. Specifically, the width and height of the coding unit or partition block for luminance are compared with the maximum luminance transform size; if either is larger than the maximum luminance transform size, LFNST is not applied. For the coding unit or partition block for chroma, the width and height of the corresponding luminance block for the color format are compared with the maximum possible luminance transform size. For example, if the color format is 4:2:0, the width and height of the corresponding luminance block become twice the size of the corresponding chroma block, and the transform size of the corresponding luminance block becomes twice the size of the corresponding chroma block. As another example, if the color format is 4:4:4, the width and height of the corresponding luminance block and the transform size are equal to those of the corresponding chroma block.
[0257] In the case of a single tree, for a luma block (coding unit or partition block), check whether the horizontal or vertical length exceeds the maximum convertible luma conversion block size; if it does, LFNST index signaling can be omitted.
[0258] C. If LFNST included in the current VVC standard is applied, LFNST indexes can be transmitted only when both the width and height of the partition block are 4 or greater.
[0259] If LFNST is applied to 2xM (1xM) or Mx2 (Mx1) blocks in addition to the LFNST included in the current VVC standard, LFNST indexes can be transmitted only when the size of the partition block is equal to or greater than 2xM (1xM) or Mx2 (Mx1) blocks. Here, the fact that the PxQ block is equal to or greater than the RxS block means that P≥R and Q≥S.
[0260] In summary, an LFNST index can be transmitted only when the partition block is equal to or larger than the minimum size for which LFNST is applicable. For dual trees, an LFNST index can be signaled only when the partition block for the luminance or chroma component is equal to or larger than the minimum size for which LFNST is applicable. For single trees, an LFNST index can be signaled only when the partition block for the luminance component is equal to or larger than the minimum size for which LFNST is applicable.
[0261] In this document, an MxN block being greater than or equal to a KxL block means that M is greater than or equal to K and N is greater than or equal to L. An MxN block being greater than a KxL block means that M is greater than or equal to K and N is greater than or equal to L, and that M is greater than K or N is greater than L. An MxN block being less than or equal to a KxL block means that M is less than or equal to K and N is less than or equal to L, and an MxN block being smaller than a KxL block means that M is less than or equal to K and N is less than or equal to L, and that M is less than K or N is less than L.
[0262] D. If the last non-zero coefficient position is not the DC position (top-left position of the block), for a dual-tree type chroma block, LFNST can be transmitted if at least one of the partition blocks has the last non-zero coefficient position not as the DC position. For a dual-tree type chroma block, if at least one of the last non-zero coefficient positions of all partition blocks for Cb (assuming the number of partition blocks is one if ISP mode is not applied to the chroma component) and the last non-zero coefficient positions of all partition blocks for Cr (assuming the number of partition blocks is one if ISP mode is not applied to the chroma component) is not the DC position, the corresponding LNFST index can be transmitted.
[0263] In the case of a single tree type, if the location of the last non-zero coefficient in any of the partition blocks for the luma component, Cb component, or Cr component is not a DC location, the corresponding LFNST index can be transferred.
[0264] Here, if the CBF (coded block flag) value indicating whether a transformation factor exists for each partition block is 0, the location of the last non-zero factor for that partition block is not checked to determine whether LFNST index signaling is performed. In other words, if the CBF value is 0, the transformation is not applied to the block, so the location of the last non-zero factor for that partition block is not considered when checking the condition for LFNST index signaling.
[0265] For example, 1) if it is a dual tree type and has a chroma component, if the CBF value for each partition block is 0, the corresponding partition block is excluded when determining whether to signal the LFNST index; 2) if it is a dual tree type and has a chroma component, if the CBF value for Cb is 0 and the CBF value for Cr is 1 for each partition block, only the location of the last non-zero coefficient for Cr is checked to determine whether to signal the LFNST index; and 3) if it is a single tree type, for all partition blocks of the chroma, Cb, and Cr components, only the location of the last non-zero coefficient for the blocks with a CBF value of 1 is checked to determine whether to signal the LFNST index.
[0266] In the case of ISP mode, image information may be configured so as not to check the position of the last non-zero count, and an example thereof is as follows.
[0267] i. In ISP mode, the check for the location of the last non-zero coefficient for both the luminance block and the chroma block can be omitted and LFNST index signaling can be allowed. That is, for all partition blocks, even if the location of the last non-zero coefficient is a DC location or the corresponding CBF value is 0, the corresponding LFNST index signaling can be allowed.
[0268] ii. In the case of ISP mode, the check for the position of the last non-zero coefficient may be omitted only for the luminance block, and the check for the position of the last non-zero coefficient in the manner described above may be performed for the chroma block. For example, in the case of a dual tree type and a luminance block, LFNST index signaling may be allowed without checking the position of the last non-zero coefficient, and in the case of a dual tree type and a chroma block, the existence of a DC position for the position of the last non-zero coefficient may be checked in the manner described above to determine whether to signal the corresponding LFNST index.
[0269] iii. In the case of ISP mode and single tree type, the above method i or ii may be applied. That is, if method i is applied to ISP mode and single tree type, the check for the position of the last non-zero coefficient for both the luminance block and the chroma block may be omitted and LFNST index signaling may be allowed. Alternatively, by applying method ii, the check for the position of the last non-zero coefficient for the partition blocks for the luminance component may be omitted, and for the partition blocks for the chroma component (if ISP is not applied to the chroma component, the number of partition blocks may be considered as 1), the check for the position of the last non-zero coefficient may be performed in the manner described above to determine whether to perform LFNST index signaling.
[0270] E. If it is confirmed that a transformation factor exists in a location other than where an LFNST transformation factor can exist for at least one of the partition blocks, LFNST index signaling may be omitted.
[0271] For example, in the case of 4x4 partition blocks and 8x8 partition blocks, LFNST transform factors may exist in 8 positions starting from the DC position according to the transform factor scanning order in the VVC standard, and all remaining positions are filled with 0. Also, in the case of a partition block that is greater than or equal to 4x4 but is not a 4x4 partition block or an 8x8 partition block, LFNST transform factors may exist in 16 positions starting from the DC position according to the transform factor scanning order in the VVC standard, and all remaining positions are filled with 0.
[0272] Therefore, if there is a non-zero conversion factor in the area where the above zero value must be filled after residual coding, LFNST index signaling can be omitted.
[0273] Meanwhile, in ISP mode, the current VVC standard independently checks length conditions for the horizontal and vertical directions and applies DST-7 instead of DCT-2 without signaling for the MTS index. It is determined whether the horizontal or vertical length is greater than or equal to 4 and less than or equal to 16, and the primary transformation kernel is determined based on the result of this determination. Therefore, for cases where LFNST can be applied in ISP mode, the following transformation combination configurations are possible.
[0274] 1. For cases where the LFNST index is 0 (including cases where the LFNST index is inferred to be 0), the first-order transformation determination condition for the ISP included in the current VVC standard may be followed. That is, the length condition (greater than or equal to 4 and less than or equal to 16) is checked independently for the horizontal and vertical directions, and if satisfied, DST-7 is applied instead of DCT-2 for the first-order transformation, and if not satisfied, DCT-2 is applied.
[0275] 2. For cases where the LFNST index is greater than 0, the following two configurations are possible as a first-order transformation.
[0276] A. DCT-2 can be applied to both horizontal and vertical directions.
[0277] B. The first conversion determination condition for ISP included in the current VVC standard can be followed. That is, the length condition (greater than or equal to 4 and less than or equal to 16) is checked independently for the horizontal and vertical directions, and if satisfied, DST-7 is applied instead of DCT-2, and if not satisfied, DCT-2 is applied.
[0278] When in ISP mode, the image information can be configured so that the LFNST index is transmitted per partition block rather than per coding unit. In this case, in the LFNST index signaling method described above, it can be determined whether to signal the LFNST index by assuming that there is only one partition block within the unit in which the LFNST index is transmitted.
[0279] Meanwhile, the signaling order of the LFNST index and MTS index is examined below.
[0280] In one example, the LFNST index signaled in residual coding may be coded after the coding position for the last non-zero count position, and the MTS index may be coded immediately after the LFNST index. In this configuration, the LFNST index may be signaled for each transformation unit. Alternatively, even if not signaled in residual coding, the LFNST index may be coded after the coding for the last effective count position, and the MTS index may be coded after the LFNST index.
[0281] The syntax of residual coding according to one example is as follows.
[0282]
[0284] *
[0285] The meanings of the major variables shown in Table 3 are as follows.
[0286] 1. cbWidth, cbHeight: Width and height of the current coding block
[0287] 2. log2TbWidth, log2TbHeight: Base-2 log values for the width and height of the current Transform Block, which can be reduced to the top-left area where non-zero coefficients may exist due to zero-out reflection.
[0288] 3. sps_lfnst_enabled_flag: A flag indicating whether LFNST is enabled. If the flag value is 0, it indicates that LFNST is not enabled, and if the flag value is 1, it indicates that LFNST is enabled. It is defined in the Sequence Parameter Set (SPS).
[0289] 4. CuPredMode[ chType ][ x0 ][ y0 ]: The prediction mode corresponding to the variable chType and the position (x0, y0). chType can have values of 0 or 1, where 0 represents the luminance component and 1 represents the chroma component. The position (x0, y0) represents the position on the picture, and the values for CuPredMode[ chType ][ x0 ][ y0 ] allow MODE_INTRA (intra prediction) and MODE_INTER (inter prediction).
[0290] 5. IntraSubPartitionsSplit[ x0 ][ y0 ]: The details for the (x0, y0) location are the same as in number 4. It indicates which ISP splitting has been applied at the (x0, y0) location, and ISP_NO_SPLIT indicates that the coding unit corresponding to the (x0, y0) location has not been split into partition blocks.
[0291] 6. intra_mip_flag[ x0 ][ y0 ]: The details regarding the (x0, y0) location are the same as those in item 4 above. intra_mip_flag is a flag indicating whether the Matrix-based Intra Prediction (MIP) prediction mode is applied. If the flag value is 0, it indicates that MIP cannot be applied, and if the flag value is 1, it indicates that MIP is applied.
[0292] 7. cIdx: A value of 0 represents chroma, and values of 1 and 2 represent the chroma components Cb and Cr, respectively.
[0293] 8. treeType: Refers to single-tree and dual-tree, etc. (SINGLE_TREE: Single tree, DUAL_TREE_LUMA: Dual tree for luminar components, DUAL_TREE_CHROMA: Dual tree for chroma components)
[0294] 9. tu_cbf_cb[ x0 ][ y0 ]: The content regarding the (x0, y0) position is the same as in 4. It represents the CBF (Coded Block Flag) for the Cb component. If the value is 0, it means that there are no non-zero coefficients in the corresponding transformation unit for the Cb component, and if it is 1, it means that there are non-zero coefficients in the corresponding transformation unit for the Cb component.
[0295] 10. lastSubBlock: Indicates the position in the scan order of the sub-block (Coefficient Group (CG)) where the last non-zero coefficient is located. 0 indicates a sub-block containing a DC component, and if greater than 0, it is not a sub-block containing a DC component.
[0296] 11. lastScanPos: Indicates the position of the last valid count in the scan order within a sub-block. If a sub-block consists of 16 positions, values from 0 to 15 are possible.
[0297] 12. lfnst_idx[ x0 ][ y0 ]: This is the LFNST index syntax element to be parsed. If not parsed, it is inferred to be a value of 0. In other words, the default value is set to 0, indicating that LFNST is not applied.
[0298] 13. LastSignificantCoeffX, LastSignificantCoeffY: Represents the x and y coordinates where the last significant coefficient is located within the transformation block. The x-coordinates start at 0 and increase from left to right, while the y-coordinates start at 0 and increase from top to bottom. If both variables are 0, it means the last significant coefficient is located in the DC.
[0299] 14. cu_sbt_flag: A flag indicating whether the SubBlock Transform (SBT) included in the current VVC standard is applicable. If the flag value is 0, it indicates that the SBT is not applicable, and if the flag value is 1, it indicates that the SBT is applicable.
[0300] 15. sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: Flags indicating whether explicit MTS is applied to the inter-CU and intra-CU, respectively. A flag value of 0 indicates that MTS is not applicable to the inter-CU or intra-CU, and 1 indicates that it is applicable.
[0301] 16. tu_mts_idx[ x0 ][ y0 ]: This is the MTS index syntax element to be parsed. If not parsed, it is inferred to be a value of 0. In other words, the default value is set to 0, indicating that DCT-2 is applied to both the horizontal and vertical directions.
[0302] As shown in Table 3, in the case of a single tree, the signaling of the LFNST index can be determined based solely on the condition of the last effective coefficient location for the luma. That is, the LFNST index is signaled if the last effective coefficient location is not a DC and the last effective coefficient exists inside the top-left sub-block (CG), for example, a 4x4 block. In this case, for 4x4 and 8x8 transformation blocks, the LFNST index is signaled only if the last effective coefficient exists at a position from 0 to 7 inside the top-left sub-block.
[0303] In the case of a dual tree, the LFNST index is signaled independently for both the luminance and the chroma, and for the chroma, the LFNST index can be signaled by applying the last effective coefficient position condition only to the Cb component. The condition may not be checked for the Cr component, and if the CBF value for Cb is 0, the LFNST index can be signaled by applying the last effective coefficient position condition to the Cr component.
[0304] 'Min( log2TbWidth, log2TbHeight ) >= 2' in Table 3 can be expressed as “Min( tbWidth, tbHeight ) >= 4”, and 'Min( log2TbWidth, log2TbHeight ) >= 4’ can be expressed as “Min( tbWidth, tbHeight ) >= 16”.
[0305] In Table 3, log2ZoTbWidth and log2ZoTbHeight represent the base-2 log values of the width and height for the top-left area where the last valid coefficient may exist due to zero-out, respectively.
[0306] As shown in Table 3, the log2ZoTbWidth and log2ZoTbHeight values can be updated in two places. The first is before the MTS index or LFNST index value is parsed, and the second is after the MTS index is parsed.
[0307] Since the first update is before the MTS index (tu_mts_idx[ x0 ][ y0 ]) value is parsed, log2ZoTbWidth and log2ZoTbHeight can be set regardless of the MTS index value.
[0308] After the MTS index is parsed, log2ZoTbWidth and log2ZoTbHeight are set for cases where the MTS index value is greater than 0 (in the case of a DST-7 / DCT-8 combination). When DST-7 / DCT-8 is applied independently to the horizontal and vertical directions in the first transformation, there may be up to 16 valid coefficients per row or column for each direction. That is, after applying DST-7 / DCT-8 of length 32 or more, up to 16 transformation coefficients can be derived per row or column starting from the left or top. Therefore, for a 2D block, when DST-7 / DCT-8 is applied to both the horizontal and vertical directions, valid coefficients may exist only up to a maximum top-left 16x16 area.
[0309] In addition, when DCT-2 is applied independently to the horizontal and vertical directions in the current first transformation, there may be up to 32 effective coefficients per row or column for each direction. That is, when applying DCT-2 of length 64 or more, up to 32 transformation coefficients can be derived per row or column starting from the left or top. Therefore, for a 2-dimensional block, when DCT-2 is applied to both the horizontal and vertical directions, there may only be effective coefficients up to a maximum top-left 32x32 area.
[0310] Additionally, when DST-7 / DCT-8 is applied in one direction and DCT-2 is applied in the other direction, there may be 16 effective coefficients in the former direction and 32 effective coefficients in the latter direction. For example, in the case of a 64x8 transformation block where DCT-2 is applied in the horizontal direction and DST-7 is applied in the vertical direction (which may occur in a situation where implicit MTS is applied), there may be effective coefficients in the maximum top-left 32x8 area.
[0311] If log2ZoTbWidth and log2ZoTbHeight are updated in two places as shown in Table 3, that is, before parsing the MTS index, the ranges of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be determined by log2ZoTbWidth and log2ZoTbHeight as shown in the table below.
[0312]
[0313] In addition, in such cases, the maximum values of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be set by reflecting the log2ZoTbWidth and log2ZoTbHeight values during the binarization process for last_sig_coeff_x_prefix and last_sig_coeff_y_prefix.
[0314]
[0315] Meanwhile, according to one example, when LFNST is applied in ISP mode, the spec text can be configured as shown in Table 6 when the signaling in Table 3 is applied. Compared to Table 3, the condition that signaled the LFNST index only when not in ISP mode (IntraSubPartitionsSplit[ x0 ][ y0 ] = ISP_NO_SPLIT in Table 3) has been removed.
[0316] In the case of a single tree, if the LFNST index transmitted when it is a luminance (when cIdx = 0) is reused when it is a chroma, the LFNST index transmitted for the first ISP partition block with valid coefficients can be applied to the chroma transformation block. Alternatively, even in the case of a single tree, the LFNST index can be signaled separately from the luminance component for the chroma component. The descriptions of the variables listed in Table 6 are the same as those in Table 3.
[0317]
[0318] Meanwhile, according to one example, the LFNST index or / and MTS index may be signaled at the coding unit level. The LFNST index may have three values, 0, 1, and 2 as described above, where 0 indicates that no LFNST is applied, and 1 and 2 indicate the first and second candidates, respectively, of the two LFNST kernel candidates included in the selected LFNST set. The LFNST index is coded through truncated unary binarization, and the values 0, 1, and 2 may be coded as empty strings 0, 10, and 11, respectively.
[0319] According to one example, LFNST can be applied only when DCT-2 is applied to both the horizontal and vertical directions as a first-order transformation. Therefore, if the MTS index is signaled after the LFNST index signaling, the MTS index can be signaled only when the LFNST index value is 0, and when the LFNST index is not 0, the MTS index is not signaled, and the first-order transformation can be performed by applying DCT-2 to both the horizontal and vertical directions.
[0320] The MTS index value can have values of 0, 1, 2, 3, and 4, and 0, 1, 2, 3, and 4 can indicate that DCT-2 / DCT-2, DST-7 / DST-7, DCT-8 / DST-7, DST-7 / DCT-8, and DCT-8 / DCT-8 are applied for the horizontal and vertical directions, respectively. Additionally, the MTS index can be coded through truncated binary conversion, and the values 0, 1, 2, 3, and 4 can be coded as empty strings 0, 10, 110, 1110, and 1111, respectively.
[0321] The signaling of an LFNST index at the coding unit level can be represented as shown in the table below. An LFNST index can be signaled in the latter part of the coding unit syntax table.
[0322]
[0323] The variables LfnstDcOnly and LfnstZeroOutSigCoeffFlag in Table 7 can be set as shown in Table 10 below.
[0324] The variable LfnstDcOnly becomes 1 if the last valid coefficient is located at the DC position (top-left position) for all transformation blocks where the CBF (Coded Block Flag, 1 if there is at least one valid coefficient in the block, 0 otherwise) value is 1, and becomes 0 otherwise. More specifically, in the case of dual-tree luminance, the position of the last valid coefficient is checked for a single luminance transformation block, and in the case of dual-tree chroma, the position of the last valid coefficient is checked for both the transformation block for Cb and the transformation block for Cr. In the case of single-tree, the position of the last valid coefficient can be checked for the transformation blocks for luminance, Cb, and Cr.
[0325] The variable LfnstZeroOutSigCoeffFlag is 0 if there is an effective coefficient at the zero-out position when LFNST is applied, and 1 otherwise.
[0326] lfnst_idx[ x0 ][ y0 ] included in Table 7 and the following tables represent the LFNST index for the corresponding coding unit, and tu_mts_idx[ x0 ][ y0 ] represents the MTS index for the corresponding coding unit.
[0327] For example, if you want to code an MTS index consecutively after an LFNST index at the coding unit level, a coding unit syntax table can be configured as shown in Table 8.
[0328]
[0329] When comparing Table 8 with Table 7, the condition of checking whether the value of tu_mts_idx[ x0 ][ y0 ] is 0 in the condition of signaling lfnst_idx[ x0 ][ y0 ] (i.e., checking whether it is DCT-2 for both horizontal and vertical directions) has been changed to the condition of checking whether the value of transform_skip_flag[ x0 ][ y0 ] is 0 (!transform_skip_flag[ x0 ][ y0 ]). transform_skip_flag[ x0 ][ y0 ] indicates whether the coding unit is coded in a transform skip mode where the transformation is omitted, and the above flag is signaled before the MTS index and LFNST index. That is, since lfnst_idx[ x0 ][ y0 ] is signaled before tu_mtx_idx[ x0 ][ y0 ] values, only the condition for transform_skip_flag[ x0 ][ y0 ] values can be checked.
[0330] As shown in Table 8, several conditions are checked when coding tu_mts_idx[ x0 ][ y0 ], and as described above, tu_mts_idx[ x0 ][ y0 ] is signaled only when the value of lfnst_idx[ x0 ][ y0 ] is 0.
[0331] In addition, tu_cbf_luma[ x0 ][ y0 ] is a flag indicating whether there is an effective coefficient for the luminance component, and cbWidth and cbHeight represent the width and height of the coding unit for the luminance component, respectively.
[0332] Also, in Table 8, ( IntraSubPartitionsSplit[ x0 ][ y0 ] = = ISP_NO_SPLIT ) indicates the case where it is not in ISP mode, and ( !cu_sbt_flag ) indicates the case where SBT is not applied.
[0333] According to Table 8, tu_mts_idx[ x0 ][ y0 ] is signaled when both the width and height of the coding unit for the luminance component are 32 or less, that is, whether MTS is applied is determined by the width and height of the coding unit for the luminance component.
[0334] In other examples, when transformation block tiling (TU tiling) occurs (e.g., when the maximum transformation size is set to 32, a 64x64 coding unit is divided into four 32x32 transformation blocks for coding), the MTS index can be signaled based on the size of each transformation block. For example, when both the width and height of a transformation block are 32 or less, the same MTS index value can be applied to all transformation blocks within the coding unit, so that the same first transformation can be applied. Also, when transformation block tiling occurs, the tu_cbf_luma[ x0 ][ y0 ] value in Table 8 can be set to 1 if the CBF value for the top-left transformation block is 1 for all transformation blocks, or if at least one transformation block has a corresponding CBF value of 1.
[0335] For example, if the ISP mode is applied to the current block, LFNST can be applied, and in this case, Table 8 can be changed as shown in Table 9.
[0336]
[0337] As shown in Table 9, even in ISP mode (IntraSubPartitionsSplitType ! = ISP_NO_SPLIT), it can be configured to signal lfnst_idx[ x0 ][ y0 ], and the same LFNST index value can be applied to all ISP partition blocks.
[0338] In addition, as shown in Table 9, tu_mts_idx[ x0 ][ y0 ] can be signaled only when not in ISP mode, so the MTS index coding part is the same as in Table 8.
[0339] As shown in Tables 8 and 9, when the MTS index is signaled immediately after the LFNST index, information about the first transformation cannot be known when performing residual coding. That is, the MTS index is signaled after residual coding. Therefore, the part of the residual coding that performs zero-out leaving only 16 coefficients for a 32-length DST-7 or DCT-8 can be changed as shown in Table 10 below.
[0340]
[0341]
[0342] As shown in Table 10, in the process of determining log2ZoTbWidth and log2ZoTbHeight (where log2ZoTbWidth and log2ZoTbHeight represent the base-2 log values of the width and height of the top-left area remaining after zero-out, respectively), the part checking the tu_mts_idx[ x0 ][ y0 ] value may be omitted.
[0343] Binarization for last_sig_coeff_x_prefix and last_sig_coeff_y_prefix in Table 10 can be determined based on log2ZoTbWidth and log2ZoTbHeight as shown in Table 5.
[0344] In addition, as shown in Table 10, a condition to check sps_mts_enable_flag can be added when determining log2ZoTbWidth and log2ZoTbHeight in residual coding.
[0345] Meanwhile, according to another example, the coding unit syntax table, the transformation unit syntax table, and the residual coding syntax table are as shown in the following table. According to Table 11, the MTS index moves from the transformation unit level to the coding unit level syntax and is signaled after the LFNST index signaling. Additionally, the restriction that disallows LFNST when an ISP is applied to the coding unit has been removed. Since the restriction that disallows LFNST when an ISP is applied to the coding unit is removed, LFNST can be applied to all intra-prediction blocks. Furthermore, both the MTS index and the LFNST index are conditionally signaled at the end of the coding unit level.
[0346]
[0347]
[0348]
[0349] In Table 11, MtsZeroOutSigCoeffFlag is initially set to 1, and this value can be changed in the residual coding of Table 13. The variable MtsZeroOutSigCoeffFlag is changed from 1 to 0 if there is a valid coefficient in the area that should be filled with zeros due to zeroing (LastSignificantCoeffX > 15 | LastSignificantCoeffY > 15), in which case, as in Table 11, the MTS index is not signaled.
[0350] Meanwhile, as shown in Table 11, when tu_cbf_luma[ x0 ][ y0 ] is 0, the coding of mts_idx[ x0 ][ y0 ] can be omitted. That is, since the transformation is not applied when the CBF value of the luma component is 0, there is no need to signal the MTS index, so the coding of the MTS index can be omitted.
[0351] According to one example, the above technical feature may be implemented with other conditional statements. For example, after MTS is performed, a variable indicating whether there is a valid coefficient in the area excluding the DC area of the current block may be derived, and if the variable indicates that there is a valid coefficient in the area excluding the DC area, the MTS index may be signaled. That is, the existence of a valid coefficient in the area excluding the DC area of the current block indicates that the value of tu_cbf_luma[ x0 ][ y0 ] is 1, and in this case, the MTS index may be signaled.
[0352] The above variable can be represented as MtsDcOnly, and the variable MtsDcOnly is initially set to 1 at the coding unit level, and then its value can be changed to 0 at the residual coding level if it indicates that there is a valid coefficient in the area excluding the DC area of the current block. When the variable MtsDcOnly is 0, image information can be configured so that the MTS index is signaled.
[0353] If tu_cbf_luma[ x0 ][ y0 ] is 0, the variable MtsDcOnly retains its initial value of 1 because the resident coding syntax is not called at the transform unit level of Table 12. In this case, since the variable MtsDcOnly is not changed to 0, the image information can be configured so that the MTS index is not signaled. That is, the MTS index is not parsed or signaled.
[0354] Meanwhile, the decoding device can determine the color index (cIdx) of the conversion coefficient to derive the variable MtsZeroOutSigCoeffFlag in Table 13. A color index (cIdx) of 0 indicates a luminance component.
[0355] For example, since MTS can be applied only to the luminance component of the current block, the decoding device can determine whether the color index is luminance when deriving the variable MtsZeroOutSigCoeffFlag, which determines whether to parse the MTS index (if cIdx = 0, MtsZeroOutSigCoeffFlag = 0).
[0356] The variable MtsZeroOutSigCoeffFlag indicates whether zero-out was performed when MTS was applied, and indicates whether transformation coefficients exist in an area other than the top-left 16x16 area, i.e., the top-left area where the last valid coefficient may exist due to zero-out after MTS is performed. As shown in Table 11, the variable MtsZeroOutSigCoeffFlag is initially set to 1 at the coding unit level (MtsZeroOutSigCoeffFlag = 1), and if transformation coefficients exist in an area other than the 16x16 area, its value can be changed from 1 to 0 at the residual coding level (MtsZeroOutSigCoeffFlag = 0) as shown in Table 13. If the value of the variable MtsZeroOutSigCoeffFlag is 0, the MTS index is not signaled.
[0357] As shown in Table 13, at the residual coding level, a non-zero-out area may be set where a non-zero conversion factor may exist depending on whether zero-out accompanying MTS is performed, and in this case, when the color index (cIdx) is 0, the non-zero-out area may be set to the top-left 16x16 area of the current block.
[0358] As such, when deriving the variable determining whether to parse the MTS index, it is determined whether the color component is luminance or chroma; however, since LFNST can be applied to both the luminance and chroma components of the current block, the color component is not determined when deriving the variable determining whether to parse the LFNST index.
[0359] For example, Table 11 shows the variable LfnstZeroOutSigCoeffFlag, which indicates whether zero-out has been performed when LFNST is applied. The variable LfnstZeroOutSigCoeffFlag indicates whether valid coefficients exist in the second region, excluding the first region at the top left of the current block; this value is initially set to 1, and if valid coefficients exist in the second region, the value can be changed to 0. The LFNST index can be parsed only if the initially set value of the variable LfnstZeroOutSigCoeffFlag is maintained at 1. When determining and deriving whether the value of the variable LfnstZeroOutSigCoeffFlag is 1, the color index of the current block is not determined because LFNST can be applied to both the luminance component and the chroma component of the current block.
[0360] FIG. 12 is a diagram illustrating a CCLM that can be applied when deriving an intra-prediction mode of a chroma block according to one embodiment.
[0361] In this specification, “reference sample template” may mean a set of reference samples around the current chroma block for predicting the current chroma block. The reference sample template may be predefined, and information regarding the reference sample template may be signaled from the encoding device (100) to the decoding device (200).
[0362] Referring to FIG. 12, the set of samples shaded as one line around the current chroma block, a 4x4 block, represents a reference sample template. While the reference sample template consists of one line of reference samples, FIG. 12 can be seen that the reference sample area within the luminance area corresponding to the reference sample template consists of two lines.
[0363] In one embodiment, when performing intra-frame encoding of a chroma image in the Joint Explolation Test Model (JEM) used by the Joint Video Exploration Team (JVET), a Cross Component Linear Model (CCLM) may be used. CCLM is a method that predicts pixel values of a chroma image from pixel values of a restored luminance image, and is based on the characteristic of high correlation between the luminance image and the chroma image.
[0364] CCLM prediction of Cb and Cr chroma images can be based on the following mathematical formula.
[0365]
[0366] Here, Pred c (i,j) is the predicted Cb or Cr chroma image, Rec L (i,j) represents the reconstructed luminance image adjusted to the chroma block size, and (i,j) represents the pixel coordinates. In the 4:2:0 color format, since the size of the luminance image is twice that of the color image, Rec of the chroma block size is obtained through downsampling. L ' must be generated, and therefore chroma image Pred c The pixels of the luminance image to be used at (i,j) are Rec L In addition to (2i, 2j), it can be used by considering all surrounding pixels. The above Rec L (i,j) can be represented as a downsampled luma sample.
[0367] For example, the above Rec L '(i,j) can be derived using 6 surrounding pixels as shown in the following mathematical formula.
[0368]
[0369] In addition, α and β represent the difference in cross-correlation and average values between the template surrounding the Cb or Cr chroma block and the template surrounding the luminance block, as shown in the shaded area of Fig. 12, and are, for example, as in Equation 13 below.
[0370]
[0371] Here, L(n) represents the surrounding reference samples and / or left surrounding samples of the luminance block corresponding to the current chroma image, C(n) represents the surrounding reference samples and / or left surrounding samples of the current chroma block to which the current encoding is applied, and (i,j) represents the pixel location. Additionally, L(n) may represent the down-sampled upper surrounding samples and / or left surrounding samples of the current luminance block. Additionally, N may represent the total number of pixel pairs (luminance and chroma) used in the CCLM parameter calculation, and may represent a value that is twice the smaller of the width and height of the current chroma block.
[0372] Meanwhile, pictures can be divided into a sequence of Coding Tree Units (CTUs). A CTU may correspond to a Coding Tree Block (CTB). Alternatively, a CTU may include a Coding Tree Block of Luma samples and a Coding Tree Block of corresponding Chroma samples. The tree type can be classified as a Single Tree or a Dual Tree depending on whether the Luma block and the corresponding Chroma block have separate partitioning structures. If the Chroma block has the same partitioning structure as the Luma block, it can be represented as a Single Tree; if the Chroma component block has a different partitioning structure from the Luma component block, it can be represented as a Dual Tree.
[0373] Meanwhile, according to one example, when applying LFNST to a chroma transform block, it is necessary to refer to information about the collocated Luma transform block at the same location.
[0374] The existing specification text for that section is presented in a table as follows.
[0375]
[0376] As shown in Table 14, when the current intra prediction mode is CCLM mode, the intra prediction mode value for the same positional luminance transform block is taken to determine the value of the predModeIntra variable for the corresponding chroma transform block (part shown in italics). Thus, the intra prediction mode value (predModeIntra value) of the luminance transform block can be used to determine the LFNST set.
[0377] However, variables nTbW and nTbH, which are input values for this transformation process, represent the width and height of the current transform block; thus, if the current block is a luminance transform block, variables nTbW and nTbH represent the width and height of the luminance transform block, and if the current block is a chroma transform block, variables nTbW and nTbH represent the width and height of the chroma transform block.
[0378] At this time, the variables nTbW and nTbH within the italicized portion of Table 14 represent the width and height of the chroma conversion block without reflecting the color format, and thus fail to accurately indicate the reference position of the luminance conversion block corresponding to the chroma conversion block. Therefore, the italicized portion of Table 14 can be modified as shown in the table below.
[0379]
[0380] As shown in Table 15, nTbW and nTbH were changed to (nTbW * SubWidthC) / 2 and (nTbH * SubHeightC) / 2, respectively. xTbY and yTbY represent the positions of the luma within the current picture (the top-left sample of the current luma transform block relative to the top-left luma sample of the current picture), respectively, and nTbW and nTbH can represent the width and height of the transform block currently being coded (a variable nTbW specifying the width of the current transform block, a variable nTbH specifying the height of the current transform block).
[0381] If the transform block currently being coded is a transform block for chroma (for Cb or Cr), then nTbW and nTbH are the width and height of the chroma transform block, respectively. Therefore, when the transform block currently being coded is a chroma transform block (cIdx > 0), the reference position for a collocated Luma transform block must be determined using the width and height of that Luma transform block. SubWidthC and SubHeightC in Table 15 are values set according to the color format (Chroma format, e.g., 4:2:0, 4:2:2, 4:4:4), and more specifically, represent the ratio of the width and height of the Luma component and the Chroma component, respectively (see Table 16 below). Thus, for a chroma transform block, (nTbW * SubWidthC) and (nTbH * SubHeightC) can be the values for the width and height of the collocated Luma transform block, respectively.
[0382] Consequently, the xTbY + (nTbW * SubWidthC) / 2 and yTbY + (nTbH * SubHeightC) / 2 values indicate the center position value within the same-position luma transformation block based on the top-left position of the current picture, so the same-position luma transformation block can be indicated more clearly.
[0383]
[0384] In Table 15, the predModeIntra variable indicates the intra prediction mode value, and when the predModeIntra variable value is INTRA_LT_CCLM, INTRA_L_CCLM, or INTRA_T_CCLM, it indicates that the current transformation block is a transformation block for chroma. For example, in the current VVC standard, INTRA_LT_CCLM, INTRA_L_CCLM, and INTRA_T_CCLM correspond to mode values 81, 82, and 83, respectively, among the intra prediction mode values. Therefore, as shown in Table 15, the reference position in the same location luminance transformation block must be obtained using the xTbY + (nTbW * SubWidthC) / 2 value and the yTbY + (nTbH * SubHeightC) / 2 value.
[0385] As shown in Table 15, the value of the predModeIntra variable is updated by considering the intra_mip_flag[ xTbY + ( nTbW * SubWidthC ) / 2 ][ yTbY + ( nTbH * SubHeightC ) / 2 ] variable and the CuPredMode
[0000] [ xTbY + ( nTbW * SubWidthC ) / 2 ][ yTbY + ( nTbH * SubHeightC ) / 2 ] variable together.
[0386] The variable intra_mip_flag indicates whether the current transformation block (or coding unit) is coded using the Matrix-based Intra Prediction (MIP) method. intra_mip_flag[ x ][ y ] represents a flag value indicating whether MIP is applied to the position corresponding to the (x, y) coordinates relative to the luminance component, with the top-left position within the current picture set to (0, 0). The x and y coordinates increase from left to right and from top to bottom, respectively. A flag value of 1 indicates that MIP has been applied, while a flag value of 0 indicates that MIP has not been applied. MIP can only be applied to the luminance block.
[0387] According to the modified section of Table 15, when the value of intra_mip_flag[ xTbY + ( nTbW * SubWidthC ) / 2 ][ yTbY + ( nTbH * SubHeightC ) / 2 ] inside the same position luma transform block is 1, the predModeIntra value is set to planar mode (INTRA_PLANAR).
[0388] The variable value of CuPredMode
[0000] [ xTbY + ( nTbW * SubWidthC ) / 2 ][ yTbY + ( nTbH * SubHeightC ) / 2 ] represents the prediction mode value corresponding to the coordinates ( xTbY + ( nTbW * SubWidthC ) / 2, yTbY + ( nTbH * SubHeightC ) / 2 ) when the top-left position of the current picture is set to (0, 0) for the luma component. The prediction mode value can have MODE_INTRA, MODE_IBC, MODE_PLT, or MODE_INTER values, representing the intra prediction mode, IBC (Intra Block Copy) prediction mode, PLT (Palette) coding mode, and inter prediction mode, respectively. According to Table 15, if the variable value of CuPredMode
[0000] [ xTbY + ( nTbW * SubWidthC ) / 2 ][ yTbY + ( nTbH * SubHeightC ) / 2 ] is MODE_IBC or MODE_PLT, the variable predModeIntra is set to DC mode. If neither of the above two cases applies, the variable predModeIntra is set to IntraPredModeY[ xTbY + ( nTbW * SubWidthC ) / 2 ][ yTbY + ( nTbH * SubHeightC ) / 2 ] (the intra prediction mode value corresponding to the center position inside the same position luminance transform block).
[0389] For example, based on the updated predModeIntra value in Table 15, the variable predModeIntra value can be updated once more by considering whether wide angle intra prediction is performed as shown in the following table.
[0390]
[0391] The input values of the mapping process presented in Table 17, predModeIntra, nTbW, and nTbH, are equal to the updated variable predModeIntra value in Table 15 and nTbW and nTbH referenced in Table 15, respectively.
[0392] In Table 17, nCbW and nCbH represent the width and height of the coding block corresponding to the respective transformation block, and the IntraSubPartitionsSplitType variable indicates whether ISP mode is applied; if IntraSubPartitionsSplitType is ISP_NO_SPLIT, it indicates that the coding unit is not split due to ISP (i.e., ISP mode is not applied). If the value of the variable IntraSubPartitionsSplitType is not ISP_NO_SPLIT, it indicates that ISP mode is applied and the coding unit is split into 2 or 4 partition blocks. In Table 17, cIdx is an index indicating the color component; if the cIdx value is 0, it represents the luminance block, and if the cIdx value is not 0, it represents the chroma block. The predModeIntra value output through the mapping process in Table 17 is an updated value considering whether Wide Angle Intra Prediction (WAIP) mode is applied.
[0393] For the predModeIntra values updated through Table 17, the LFNST set can be determined through the mapping relationship shown in the table below.
[0394]
[0395] In the table above, lfnstTrSetIdx represents an index pointing to an LFNST set; since it has values from 0 to 3, it can be confirmed that a total of four LFNST sets are configured. Each LFNST set can be composed of two transformation kernels, i.e., LFNST kernels (depending on the region where LFNST is applied, the corresponding transformation kernel can be a 16x16 matrix or a 16x48 matrix in the forward direction), and which of the two transformation kernels is applied can be specified through the signaling of the LFNST index. Additionally, whether or not to apply LFNST can be specified through the LFNST index. In the current VVC standard, the LFNST index can have values of 0, 1, or 2; 0 indicates that LFNST is not applied, while 1 and 2 represent the respective two transformation kernels.
[0396] The following drawings are made to illustrate a specific example of the present specification. The names of specific devices or specific signals / messages / fields described in the drawings are presented as examples, and therefore the technical features of the present specification are not limited to the specific names used in the following drawings.
[0397] FIG. 13 is a flowchart illustrating the operation of a video decoding device according to one embodiment of the present document.
[0398] Each step disclosed in FIG. 13 is based on some of the contents described in FIG. 3 through 12. Therefore, specific details that overlap with the contents described in FIG. 2 through 12 will be omitted or simplified.
[0399] A decoding device (200) according to one embodiment can obtain intra prediction mode information and an LFNST index from a bitstream (S1310).
[0400] Intra-predicted mode information may include an MPM index indicating one of the MPM candidates in a list of MPMs (most probable modes) derived based on the intra-predicted modes of the surrounding blocks of the current block (e.g., left and / or upper surrounding blocks) and additional candidate modes, or remaining intra-predicted mode information indicating one of the remaining intra-predicted modes not included in the said MPM candidates.
[0401] Additionally, intra-mode information may include sps_cclm_enabled_flag, a flag indicating whether CCLM is applied to the current block, and intra_chroma_pred_mode, information regarding the intra-prediction mode for the chroma component.
[0402] LFNST index information is received as syntax information, and syntax information is received as a binary empty string containing 0 and 1.
[0403] The syntax element of the LFNST index according to the present embodiment may indicate whether an inverse LFNST or an inverse non-separable transformation is applied and one of the transformation kernel matrices included in the transformation set, and if the transformation set includes two transformation kernel matrices, the value of the syntax element of the transformation index may be three.
[0404] That is, according to one embodiment, the syntax element value for the LFNST index may include 0, indicating that the inverse LFNST is not applied to the target block; 1, indicating the first transformation kernel matrix among the transformation kernel matrices; and 2, indicating the second transformation kernel matrix among the transformation kernel matrices.
[0405] Additionally, the decoding device (200) can decode information regarding quantized transformation coefficients for the current block from the bitstream and can derive quantized transformation coefficients for the target block based on information regarding quantized transformation coefficients for the current block. Information regarding quantized transformation coefficients for the target block may be included in a Sequence Parameter Set (SPS) or a slice header and may include at least one of information regarding whether a simplified transformation (RST) is applied, information regarding a simplification factor, information regarding a minimum transformation size for applying the simplified transformation, information regarding a maximum transformation size for applying the simplified transformation, a simplified inverse transformation size, and information regarding a transformation index indicating any one of the transformation kernel matrices included in the transformation set.
[0406] The decoding device (200) can derive transformation coefficients by performing inverse quantization on residual information for the current block, i.e., quantized transformation coefficients, and can arrange the derived transformation coefficients in a predetermined scanning order.
[0407] Specifically, the derived transform coefficients can be arranged in reverse diagonal scan order in 4 x 4 block units, and the transform coefficients within the 4 x 4 blocks can also be arranged in reverse diagonal scan order. That is, the transform coefficients after inverse quantization can be arranged according to the reverse scan order applied in video codecs such as VVC or HEVC.
[0408] The transformation coefficients derived based on this residual information may be inversely quantized transformation coefficients as described above, or they may be quantized transformation coefficients. That is, the transformation coefficients only need to be data that can check whether the current block is non-zero data, regardless of whether it is quantized or not.
[0409] The decoding device can derive the intra prediction mode of the chroma block as a CCLM mode based on the intra prediction mode information (S1320).
[0410] For example, a decoding device can receive information about the intra-predicted mode of the current chroma block through a bitstream, and can derive a CCLM mode as the intra-predicted mode of the current chroma block based on the information about the intra-predicted mode.
[0411] CCLM modes may include upper-left based CCLM mode, upper-side based CCLM mode, or left-side based CCLM mode.
[0412] As described above, the decoding device can derive residual samples by applying an inseparable transformation, LFNST, or a separable transformation, MTS, and these transformations can be performed based on an LFNST kernel, that is, an LFNST index indicating an LFNST matrix and an MTS index indicating an MTS kernel, respectively.
[0413] Meanwhile, for LFNST, an LFNST set must be determined, and the LFNST set has a mapping relationship with the intra-prediction mode of the current block.
[0414] The decoding device can update the intra prediction mode of the chroma block for the inverse LFNST of the chroma block based on the intra prediction mode of the lumina block corresponding to the chroma block (S1330).
[0415] For example, the updated intra prediction mode can be derived as an intra prediction mode corresponding to a specific location within a luminance block, wherein the specific location can be set based on the color format of a chroma block.
[0416] A specific location can be the center location of the luma block and can be expressed as ((xTbY + ( nTbW * SubWidthC ) / 2), (yTbY + ( nTbH * SubHeightC ) / 2)).
[0417] At the above center position, xTbY and yTbY represent the top-left coordinates of the luminance block, that is, the top-left position relative to the luminance sample for the current transformation block, nTbW and nTbH represent the width and height of the chroma block, and SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY + ( nTbW * SubWidthC ) / 2), (yTbY + ( nTbH * SubHeightC ) / 2)) represents the middle position of the luminance transformation block, and IntraPredModeY[xTbY + ( nTbW * SubWidthC ) / 2 ][ yTbY + ( nTbH * SubHeightC) / 2] indicates the intra-prediction mode in the luminance block for that position.
[0418] SubWidthC and SubHeightC can be derived as shown in Table 16. That is, if the color format is 4:2:0, SubWidthC and SubHeightC are 2, and if the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0419] As shown in Table 15, in order to specify a specific location of the luminance block corresponding to the chroma block regardless of the color format, the color format was reflected in the variable indicating the specific location.
[0420] According to one example, if the intra prediction mode of a luma block corresponding to a specific location is a matrix-based intra prediction (hereinafter MIP) mode, the decoding device can set the updated intra prediction mode to an intra planner mode.
[0421] The MIP mode may be referred to as Affine linear weighted intra prediction (ALWIP) or Matrix weighted intra prediction (MWIP). When the MIP is applied to the current block, prediction samples for the current block can be derived by i) using surrounding reference samples from which an averaging procedure has been performed, ii) performing a matrix-vector-multiplication procedure, and iii) further performing a horizontal / vertical interpolation procedure as needed.
[0422] Alternatively, depending on one example, if the intra prediction mode corresponding to a specific location is an intra block copy (IBC) mode or a palette mode, the decoding device may set the updated intra prediction mode to an intra DC mode.
[0423] IBC prediction mode or palette mode can be used for coding content images / videos, such as games, for example, in screen content coding (SCC). Although IBC fundamentally performs prediction within the current picture, it can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be viewed as an example of intra-coding or intra-prediction. When palette mode is applied, sample values within the picture can be signaled based on information regarding palette tables and palette indices.
[0424] In summary, if the intra prediction mode for the center position is MIP mode, IBC mode, and Palette mode, the intra prediction mode of the chroma block can be updated to a specific mode such as intra planner mode or intra DC mode.
[0425] Of course, if the intra prediction mode of the center position is not MIP mode, IBC mode, or Palette mode, the intra prediction mode of the chroma block can be updated to the intra prediction mode of the lumina block for the center position to reflect the association between the chroma block and the lumina block.
[0426] The decoding device determines an LFNST set including LFNST matrices based on an updated intra-prediction mode (S1340), and can derive transformation coefficients for a chroma block based on the LFNST matrix derived from the LFNST set (S1350).
[0427] Any one of the multiple LFNST matrices can be selected based on the LFNST set and LFNST index.
[0428] As shown in Table 18, the LFNST transformation set is derived according to the intra prediction mode, and 81 to 83, which represent the CCLM mode in the intra prediction mode, are omitted because the LFNST transformation set is derived using the intra mode value for the corresponding luma block in the case of CCLM mode.
[0429] According to one example, as shown in Table 18, any one of the four LFNST sets can be determined according to the intra prediction mode of the current block, and at this time, the LFNST set to be applied to the current chroma block can also be determined.
[0430] Then, the decoding device can derive modified transform coefficients for the current chroma block by applying an LFNST matrix to the inverse quantized transform coefficients and performing an inverse RST, e.g., an inverse LFNST.
[0431] The decoding device can derive residual samples from the transformation coefficients through a first-order inverse transformation (S1360). MTS can be used for the first-order inverse transformation.
[0432] Additionally, the decoding device can generate reconstructed samples based on residual samples for the current block and predicted samples for the current block. The current block may be the current luminance block or the current chroma block.
[0433] The following drawings are made to illustrate a specific example of the present specification. The names of specific devices or specific signals / messages / fields described in the drawings are presented as examples, and therefore the technical features of the present specification are not limited to the specific names used in the following drawings.
[0434] FIG. 14 is a flowchart illustrating the operation of a video encoding device according to one embodiment of the present document.
[0435] Each step disclosed in FIG. 14 is based on some of the contents described in FIG. 3 through 12. Therefore, specific details that overlap with the contents described in FIG. 1 and FIG. 3 through 12 will be omitted or simplified.
[0436] An encoding device (100) according to one embodiment can derive an intra prediction mode for a chroma block as a CCLM mode (S1410).
[0437] For example, the encoding device can determine the intra-prediction mode of the current chroma block based on the Rate-Distortion Cost (RD cost) (or RDO). Here, the RD cost can be derived based on the Sum of Absolute Difference (SAD). The encoding device can determine the CCLM mode as the intra-prediction mode of the current chroma block based on the RD cost.
[0438] CCLM modes may include upper-left based CCLM mode, upper-side based CCLM mode, or left-side based CCLM mode.
[0439] Additionally, the encoding device can encode information regarding the intra prediction mode of the current chroma block, and the information regarding the intra prediction mode can be signaled through a bitstream. The prediction-related information of the current chroma block may include information regarding the intra prediction mode.
[0440] The encoding device can derive predicted samples for the chroma block based on CCLM mode (S1420).
[0441] An encoding device according to one embodiment can derive residual samples for a chroma block based on predicted samples (S1430).
[0442] An encoding device according to one embodiment can derive transformation coefficients for a chroma block based on a first transformation for a residual sample.
[0443] The first transformation can be performed through multiple transformation kernels, and in this case, the transformation kernel can be selected based on the intra prediction mode.
[0444] The encoding device can update the intra prediction mode of the chroma block for the LFNST of the chroma block based on the intra prediction mode of the lumina block corresponding to the chroma block (S1440).
[0445] As shown in Table 15, the encoding device can update the CCLM mode for the chroma block based on the intra prediction mode of the luminance block corresponding to the chroma block (When predModeIntra is equal to either INTRA_LT_CCLM, INTRA_L_CCLM, or INTRA_T_CCLM, predModeIntra is derived as follows:).
[0446] For example, the updated intra prediction mode can be derived as an intra prediction mode corresponding to a specific location within a luminance block, wherein the specific location can be set based on the color format of a chroma block.
[0447] A specific location can be the center location of the luma block and can be expressed as ((xTbY + ( nTbW * SubWidthC ) / 2), (yTbY + ( nTbH * SubHeightC ) / 2)).
[0448] At the above center position, xTbY and yTbY represent the top-left coordinates of the luminance block, that is, the top-left position relative to the luminance sample for the current transformation block, nTbW and nTbH represent the width and height of the chroma block, and SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY + ( nTbW * SubWidthC ) / 2), (yTbY + ( nTbH * SubHeightC ) / 2)) represents the middle position of the luminance transformation block, and IntraPredModeY[xTbY + ( nTbW * SubWidthC ) / 2 ][ yTbY + ( nTbH * SubHeightC) / 2 ] indicates the intra-prediction mode in the luminance block for the corresponding position.
[0449] SubWidthC and SubHeightC can be derived as shown in Table 16. That is, if the color format is 4:2:0, SubWidthC and SubHeightC are 2, and if the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0450] As shown in Table 15, in order to specify a specific location of the luminance block corresponding to the chroma block regardless of the color format, the color format was reflected in the variable indicating the specific location.
[0451] According to one example, if the intra prediction mode of a luma block corresponding to a specific location is a matrix-based intra prediction (hereinafter MIP) mode, the encoding device can set the updated intra prediction mode to an intra planner mode.
[0452] The MIP mode may be referred to as Affine linear weighted intra prediction (ALWIP) or Matrix weighted intra prediction (MWIP). When the MIP is applied to the current block, prediction samples for the current block can be derived by i) using surrounding reference samples from which an averaging procedure has been performed, ii) performing a matrix-vector-multiplication procedure, and iii) further performing a horizontal / vertical interpolation procedure as needed.
[0453] Alternatively, depending on one example, if the intra prediction mode corresponding to a specific location is an intra block copy (IBC) mode or a palette mode, the encoding device may set the updated intra prediction mode to an intra DC mode.
[0454] IBC prediction mode or palette mode can be used for coding content images / videos, such as games, for example, in screen content coding (SCC). Although IBC fundamentally performs prediction within the current picture, it can be performed similarly to inter-prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be viewed as an example of intra-coding or intra-prediction. When palette mode is applied, sample values within the picture can be signaled based on information regarding palette tables and palette indices.
[0455] In summary, if the intra prediction mode for the center position is MIP mode, IBC mode, and Palette mode, the intra prediction mode of the chroma block can be updated to a specific mode such as intra planner mode or intra DC mode.
[0456] Of course, if the intra prediction mode of the center position is not MIP mode, IBC mode, or Palette mode, the intra prediction mode of the chroma block can be updated to the intra prediction mode of the lumina block for the center position to reflect the association between the chroma block and the lumina block.
[0457] The encoding device determines an LFNST set including LFNST matrices based on an updated intra-prediction mode (S1450), and can derive modified transformation coefficients for a chroma block based on residual samples and LFNST matrices (S1460).
[0458] The encoding device determines a transformation set based on the mapping relationship according to the intra prediction mode applied to the current block, and can perform LFNST, i.e., inseparable transformation, based on one of the two LFNST matrices included in the transformation set.
[0459] As described above, multiple sets of transformations can be determined according to the intra-prediction mode of the transformation block to be transformed. The matrix applied to the LFNST is in a transpose relationship with the matrix used for the inverse LFNST.
[0460] In one example, the LFNST matrix can be a non-square matrix where the number of rows is less than the number of columns.
[0461] The encoding device can perform quantization based on modified transform coefficients for the current chroma block to derive quantized transform coefficients, and output image information including information regarding the quantized transform coefficients, intra prediction mode information, and an LFNST index indicating an LFNST matrix after encoding (S1470).
[0462] More specifically, the encoding device (100) can generate information regarding quantized conversion coefficients and encode the generated information regarding quantized conversion coefficients.
[0463] In one example, information regarding quantized transform coefficients may include at least one of information regarding whether LFNST is applied, information regarding a simplification factor, information regarding a minimum transform size for applying LFNST, and information regarding a maximum transform size for applying LFNST.
[0464] The encoding device can encode sps_cclm_enabled_flag, a flag indicating whether CCLM is applied to the current block, and intra_chroma_pred_mode, information regarding the intra prediction mode for the chroma component, as intra-mode information.
[0465] intra_chroma_pred_mode, which is information about the CCLM mode, can indicate an upper-left based CCLM mode, an upper-left based CCLM mode, or a left-left based CCLM mode.
[0466] In this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If the quantization / inverse quantization is omitted, the quantized transformation coefficients may be called transformation coefficients. If the transformation / inverse transformation is omitted, the transformation coefficients may be called coefficients or residual coefficients, or for the sake of consistency of expression, they may still be called transformation coefficients.
[0467] Additionally, in this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information regarding the transform coefficient(s), and information regarding said transform coefficient(s) may be signaled through residual coding syntax. Transform coefficients may be derived based on said residual information (or information regarding said transform coefficient(s), and scaled transform coefficients may be derived through inverse transform (scaling) of said transform coefficients. Residual samples may be derived based on the inverse transform (transform) of said scaled transform coefficients. This may be similarly applied / expressed in other parts of this document.
[0468] In the embodiments described above, methods are described based on flowcharts as a series of steps or blocks; however, this document is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps as described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and that other steps may be included, or that one or more steps of the flowcharts may be omitted without affecting the scope of this document.
[0469] The method according to the above-described document may be implemented in the form of software, and the encoding device and / or decoding device according to the above document may be included in a device that performs image processing, such as a TV, computer, smartphone, set-top box, display device, etc.
[0470] When the embodiments described in this document are implemented in software, the method described above may be implemented as a module (process, function, etc.) that performs the function described above. The module may be stored in memory and executed by a processor. The memory may be located inside or outside the processor and may be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0471] In addition, the decoding and encoding devices to which this document applies may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, Over-the-top video (OTT) devices, internet streaming service providers, 3D video devices, video phone video devices, and medical video devices, and may be used to process video signals or data signals. For example, Over-the-top video (OTT) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, Digital Video Recorders (DVRs), etc.
[0472] Furthermore, the processing method to which this document applies may be produced in the form of a program that is executed by a computer and may be stored on a computer-readable recording medium. Multimedia data having a data structure according to this document may also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, a floppy disk, and an optical data storage device. Additionally, the computer-readable recording medium includes media implemented in the form of a carrier wave (e.g., transmission over the Internet). Furthermore, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network. Additionally, the embodiments of this document may be implemented as a computer program product by program code, and said program code may be executed on a computer by the embodiments of this document. The above program code can be stored on a computer-readable carrier.
[0473] Figure 15 schematically illustrates an example of a video / image coding system to which this document can be applied.
[0474] Referring to FIG. 15, a video / image coding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data in the form of a file or streaming to the receiving device via a digital storage medium or a network.
[0475] The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. A transmitter may be included in the encoding device. A receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0476] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and may generate video / images (electronically). For example, virtual video / images may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.
[0477] The encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0478] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0479] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.
[0480] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.
[0481] Figure 16 illustrates an exemplary structural diagram of a content streaming system to which this document applies.
[0482] In addition, the content streaming system to which this document applies may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0483] The above-mentioned encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the above-mentioned streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, directly generate the bitstream, the above-mentioned encoding server may be omitted. The above-mentioned bitstream may be generated by an encoding method or a bitstream generation method to which this document applies, and the above-mentioned streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0484] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server plays the role of controlling commands and responses between each device within the content streaming system.
[0485] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.
[0486] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc. Each server within the above content streaming system may be operated as a distributed server, and in this case, data received from each server may be processed in a distributed manner.
[0487] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined to be implemented as a device, and the technical features of the device claims in this specification may be combined to be implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a device, and the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a method.
Claims
Claim 1 A video decoding method performed by a decoding device comprises: a step of obtaining intra prediction mode information and an LFNST index from a bitstream; a step of deriving an intra prediction mode of a chroma block as a CCLM (cross-component linear model) mode based on the intra prediction mode information; a step of updating the intra prediction mode of the chroma block based on an intra prediction mode of a luminance block corresponding to the chroma block; a step of determining an LFNST set including LFNST matrices based on the updated intra prediction mode; a step of deriving transformation coefficients for the chroma block based on the LFNST index and the LFNST matrix derived from the LFNST set; and a step of deriving residual samples for the chroma block based on the transformation coefficients, wherein the updated intra prediction mode is derived as an intra prediction mode corresponding to a specific location within the luminance block, the color format is 4:2:2, the specific location is set as ((xTbY + nTbW), (yTbY + nTbH / 2)), and xTbY An image decoding method characterized in that yTbY represents the top-left coordinates of the lumina block, and nTbW and nTbH represent the width and height of the chroma block. Claim 2 A video decoding method according to claim 1, characterized in that the specific location is the center location of the luma block. Claim 3 A video decoding method according to claim 1, characterized in that the color format is 4:2:0 and the specific position is set to ((xTbY + nTbW), (yTbY + nTbH)). Claim 4 A video encoding method performed by a video encoding device comprises: a step of deriving an intra prediction mode for a chroma block as a CCLM (cross-component linear model) mode; a step of deriving prediction samples for the chroma block based on the CCLM mode; a step of deriving residual samples for the chroma block based on the prediction samples; a step of updating the intra prediction mode of the chroma block based on the intra prediction mode of a luminance block corresponding to the chroma block; a step of determining an LFNST set including LFNST matrices based on the updated intra prediction mode; and a step of deriving modified transformation coefficients for the chroma block, wherein the updated intra prediction mode is derived as an intra prediction mode corresponding to a specific location within the luminance block, the color format is 4:2:2, the specific location is set as ((xTbY + nTbW), (yTbY + nTbH / 2)), xTbY and yTbY represent the top-left coordinates of the luminance block, and nTbW and A video encoding method characterized in that nTbH represents the width and height of the chroma block. Claim 5 A video encoding method characterized in that, in paragraph 4, the specific location is the center location of the luma block. Claim 6 A video encoding method according to claim 4, characterized in that the color format is 4:2:0 and the specific location is set to ((xTbY + nTbW), (yTbY + nTbH)). Claim 7 A computer-readable digital storage medium, wherein a bitstream generated by the image encoding method of claim 4 is stored. Claim 8 A method for transmitting image data, wherein the method comprises the steps of: acquiring a bitstream for the image, wherein the bitstream derives an intra prediction mode for a chroma block as a CCLM (cross-component linear model) mode; deriving prediction samples for the chroma block based on the CCLM mode; deriving residual samples for the chroma block based on the prediction samples; updating the intra prediction mode of the chroma block based on the intra prediction mode of a luminance block corresponding to the chroma block; determining an LFNST set including LFNST matrices based on the updated intra prediction mode; and deriving modified transformation coefficients for the chroma block. A transmission method characterized by comprising the step of generating the bitstream based on encoding residual information related to the modified conversion coefficient to generate the bitstream, and the step of transmitting the data including the bitstream, wherein the updated intra prediction mode is derived as an intra prediction mode corresponding to a specific location within the luminance block, the color format is 4:2:2, the specific location is set as ((xTbY + nTbW), (yTbY + nTbH / 2)), xTbY and yTbY represent the top-left coordinates of the luminance block, and nTbW and nTbH represent the width and height of the chroma block.