Image encoding and decoding apparatuses and apparatuses for transmitting image data
By deriving the CCLM mode of chroma blocks and updating the intra-frame prediction mode in image coding, and using LFNST indexes and matrices for coding, the problem of efficient compression of high-resolution images/videos is solved, improving coding efficiency and device performance.
Patent Information
- Application Number
- CN202310610728.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-29
- Filing Date
- 2020-10-29
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2040-10-29
AI Technical Summary
Existing technologies, when transmitting and storing high-resolution, high-quality images/videos, lead to increased costs due to the increased amount of information, and the demand for immersive media and broadcasting images/videos with different image characteristics is not effectively met.
By deriving the intra-frame prediction mode of the chroma block as the cross-component linear model (CCLM) mode, and updating the intra-frame prediction mode of the chroma block based on the luma block, the image coding efficiency is improved by using the LFNST index and LFNST matrix for encoding.
It increases image/video compression efficiency, improves the efficiency of encoding LFNST index and secondary transformation, and optimizes the performance of image encoding devices.
Smart Images

Figure CN116600140B_ABST
Abstract
Description
[0001] This application is a divisional application of the original application No. 202080089314.7 (International Application No. PCT / KR2020 / 014929, filed on October 29, 2020, entitled "Image encoding and decoding method, storage medium, and method of transmitting image data") for an invention patent application. TECHNICAL FIELD
[0002] The disclosure relates generally to an image encoding technology, and more particularly, to a transform-based image encoding method in an image encoding system and an apparatus thereof. BACKGROUND
[0003] Nowadays, there is a growing demand for high-resolution and high-quality images / videos such as ultra-high definition (UHD) images / videos of 4K, 8K or more in various fields. As image / video data becomes higher in resolution and quality, the amount of information or bits transmitted increases compared to conventional image data. Therefore, when transmitting image data using a medium such as a conventional wired / wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0004] In addition, nowadays, interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and broadcasting of images / videos having different image characteristics from real images such as game images is increasing.
[0005] Therefore, there is a need for an efficient image / video compression technology that effectively compresses and transmits or stores, and reproduces information of high-resolution and high-quality images / videos having various characteristics as described above. SUMMARY
[0006] TECHNICAL PROBLEM
[0007] A technical aspect of the disclosure is to provide a method and apparatus for increasing image encoding efficiency.
[0008] Another technical aspect of the disclosure is to provide a method and apparatus for increasing the efficiency of encoding LFNST index.
[0009] Still another technical aspect of the disclosure is to provide a method and apparatus for improving the efficiency of secondary transform by encoding LFNST index.
[0010] Still another technical aspect of the disclosure is to provide an image encoding method and image encoding apparatus for deriving a LFNST transform set using an intra mode of a luma block in a CCLM mode.
[0011] TECHNICAL SOLUTION
[0012] According to embodiments of the present disclosure, an image decoding method performed by a decoding device is provided. The method includes deriving an intra prediction mode of a chroma block as a cross-component linear model (CCLM) mode based on intra prediction mode information, updating the intra prediction mode of the chroma block based on an intra prediction mode of a luma block corresponding to the chroma block, when the chroma block is not square, remapping the updated intra prediction mode to a wide-angle intra prediction mode, determining an LFNST set including an LFNST matrix based on the remapped intra prediction mode, deriving transform coefficients of the chroma block based on the LFNST matrix derived from an LFNST index and the LFNST set, and deriving residual samples of the chroma block based on the transform coefficients.
[0013] The updated intra prediction mode is derived as an intra prediction mode corresponding to a specific position in the luma block, and the specific position is set based on a color format of the chroma block.
[0014] The specific position is set as ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)), xTbY and yTbY denote top-left coordinates of the luma block, nTbW and nTbH denote a width and a height of the chroma block, and SubWidthC and SubHeightC denote variables corresponding to the color format.
[0015] SubWidthC and SubHeightC are 2 when the color format is 4:2:0, and SubWidthC is 2 and SubHeightC is 1 when the color format is 4:2:2.
[0016] When the intra prediction mode corresponding to the specific position is a palette mode or an IBC mode, the intra prediction mode of the chroma block is updated as an intra DC mode.
[0017] When the intra prediction mode corresponding to the specific position is a MIP mode, the intra prediction mode of the chroma block is updated as an intra planar mode.
[0018] When a width of the chroma block is greater than a height, the updated intra mode is 2 or greater and the updated intra mode is less than a variable (whRatio>1)? (8+2*whRatio):8 [where whRatio is Abs(Log2(nW / nH))], the updated intra mode is remapped to “updated intra mode+65”.
[0019] When the height of the chroma block is greater than the width, the updated intra mode is 66 or less, and the updated intra mode is greater than a variable (whRatio > 1)? (60 - 2*whRatio):60 [where whRatio is Abs(Log2(nW / nH))], the updated intra mode is remapped to "updated intra mode - 67".
[0020] According to another embodiment of the disclosure, an image encoding method performed by an encoding apparatus is provided. The method includes deriving an intra prediction mode of a chroma block as a cross-component linear model (CCLM) mode, deriving prediction samples of the chroma block based on the CCLM mode, deriving residual samples of the chroma block based on the prediction samples, updating the intra prediction mode of the chroma block based on an intra prediction mode of a luma block corresponding to the chroma block, when the chroma block is not square, remapping the updated intra prediction mode to a wide-angle intra prediction mode, determining an LFNST set including an LFNST matrix based on the remapped intra prediction mode, and deriving modified transform coefficients of the chroma block based on the residual samples and the LFNST matrix.
[0021] According to still another embodiment of the disclosure, a digital storage medium storing image data including a bitstream and encoded image information generated according to an image encoding method performed by an encoding apparatus can be provided.
[0022] According to still another embodiment of the disclosure, a digital storage medium storing image data including encoded image information and a bitstream to enable a decoding apparatus to perform an image decoding method can be provided.
[0023] Technical Effects
[0024] According to the disclosure, overall image / video compression efficiency can be increased.
[0025] According to the disclosure, efficiency of encoding an LFNST index can be increased.
[0026] According to the disclosure, efficiency of a secondary transform can be increased by encoding an LFNST index.
[0027] According to the disclosure, an image encoding method and an image encoding apparatus for deriving an LFNST transform set using an intra mode of a luma block in a CCLM mode can be provided.
[0028] Effects that can be obtained through specific examples of the disclosure are not limited to those listed above. For example, there can be various technical effects that can be understood or derived by those of ordinary skill in the related art from the disclosure. Therefore, specific effects of the disclosure are not limited to those explicitly described in the disclosure, and can include various effects that can be understood or derived from technical features of the disclosure. Attached Figure Description
[0029] Figure 1 Examples of video / image coding systems to which this disclosure can be applied are illustrated schematically.
[0030] Figure 2 This is a diagram that schematically illustrates the configuration of a video / image encoding device to which this disclosure can be applied.
[0031] Figure 3 This is a diagram that schematically illustrates the configuration of a video / image decoding device to which this disclosure can be applied.
[0032] Figure 4 The structure of a content streaming system applying this disclosure is illustrated.
[0033] Figure 5 Multiple transformation techniques according to embodiments of the present disclosure are illustrated schematically.
[0034] Figure 6 The intra-frame orientation patterns for 65 predicted directions are schematically shown.
[0035] Figure 7 This is a diagram used to illustrate an embodiment of the RST according to the present disclosure.
[0036] Figure 8 This is a diagram illustrating the order in which the output data of a forward first transformation is arranged into a one-dimensional vector, based on the example.
[0037] Figure 9 This is a diagram illustrating the order in which the output data of the forward quadratic transform is arranged into two-dimensional blocks, based on the example.
[0038] Figure 10 This is a diagram illustrating a wide-angle intra-frame prediction mode according to an implementation method described in this document.
[0039] Figure 11 This is a diagram illustrating the block shape to which LFNST is applied.
[0040] Figure 12 This is a diagram illustrating the arrangement of the output data of the positive LFNST according to the example.
[0041] Figure 13 This example illustrates zeroing in a block where 4×4LFNST has been applied, based on the example.
[0042] Figure 14 This example illustrates zeroing in a block where 8×8 LFNST has been applied, based on the example.
[0043] Figure 15 An example of a CCLM that can be applied when deriving the intra-prediction mode of a chroma block according to an embodiment is shown.
[0044] Figure 16 FIG. 2 is a flowchart illustrating an operation of a video decoding apparatus according to an embodiment of the disclosure.
[0045] Figure 17 FIG. 3 is a flowchart illustrating an operation of a video encoding apparatus according to an embodiment of the disclosure. DETAILED DESCRIPTION
[0046] Although the disclosure can be susceptible to various modifications and alternative forms, specific embodiments thereof have been shown by way of example in the drawings and will herein be described in detail. The disclosure is not intended to be limited to the particular forms shown. Rather, the disclosure is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the disclosure. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises", "comprising", "includes", "including" and the like are to be construed to be open-ended and to mean "including, but not limited to".
[0047] Further, the components on the drawings described herein are independently illustrated for the convenience of describing different characteristic functions from each other, however, it is not intended that the components are implemented by separate hardware or software. For example, any two or more of the components can be combined to form a single component, and any single component can be divided into a plurality of components. Embodiments in which components are combined and / or divided are within the scope of the patent right of the disclosure as long as they do not depart from the essence of the disclosure.
[0048] Hereinafter, preferred embodiments of the disclosure will be described in greater detail with reference to the accompanying drawings. Also, in the drawings, the same drawing reference numerals are used for the same components, and overlapping descriptions for the same components will be omitted.
[0049] This document relates to video / image encoding. For example, the methods / examples disclosed in this document can relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next generation video / image encoding standard after VVC, or other video encoding related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (Elementary Video Coding) standard, the AVS2 standard, etc.).
[0050] In this document, various embodiments related to video / image encoding can be provided, and unless specified to the contrary, these embodiments can be combined with and executed with each other.
[0051] In this document, a video can refer to a set of a series of pictures over time. Typically, a picture refers to a unit of an image representing a specific time region, while a slice / tile refers to a unit of a part constituting a picture. A slice / tile can include one or more coding tree units (CTUs). A picture can be composed of one or more slices / tiles. A picture can be composed of one or more tile groups. One tile group can include one or more tiles.
[0052] A pixel or a pel can refer to a minimum unit constituting a picture (or an image). Also, a "sample" can be used as a term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. Alternatively, a sample can mean a pixel value in a spatial domain, or when the pixel value is transformed into a frequency domain, it can mean a transform coefficient in the frequency domain.
[0053] A unit can represent a basic unit of image processing. A unit can include at least one of a specific region and information related to the region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. Depending on the situation, a unit and terms such as a block, a region, and the like can be used interchangeably. In general cases, an MxN block can include a set (or an array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0054] In this document, the terms “ / ” and “,” should be interpreted to mean “and / or.” For example, the expression “A / B” can mean “A and / or B.” Also, “A, B” can mean “A and / or B.” Also, “A / B / C” can mean “at least one of A, B, and / or C.” Also, “A / B / C” can mean “at least one of A, B, and / or C.”
[0055] Also, in this document, the term “or” should be interpreted to mean “and / or.” For example, the expression “A or B” can include 1) only A, 2) only B, and / or 3) both A and B. In other words, a term “or” in this document should be interpreted as indicating “additionally or alternatively.”
[0056] In this disclosure, “at least one of A and B” can mean “only A”, “only B”, or “both A and B”. Also, in this disclosure, the expression “at least one of A or B” or “at least one of A and / or B” can be interpreted as “at least one of A and B”.
[0057] Also, in the disclosure, "at least one of A, B and C" can mean "only A", "only B", "only C", or "any combination of A, B and C". Also, "at least one of A, B or C" or "at least one of A, B and / or C" can mean "at least one of A, B and C".
[0058] In addition, the bracket used in the disclosure can mean "for example". Specifically, when indicated as "prediction (intra prediction)", it can mean that "intra prediction" is proposed as an example of "prediction". In other words, the "prediction" of the disclosure is not limited to "intra prediction", and "intra prediction" is proposed as an example of "prediction". In addition, when indicated as "prediction (i.e., intra prediction)", this can also mean that "intra prediction" is proposed as an example of "prediction".
[0059] The technical features described separately in one drawing in the disclosure can be implemented separately or can be implemented simultaneously.
[0060] Figure 1 An example of a video / image encoding system to which the disclosure is applicable is schematically illustrated.
[0061] Referring to Figure 1 , the video / image encoding system can include a first apparatus (a source apparatus) and a second apparatus (a receiving apparatus). The source apparatus can deliver encoded video / image information or data in the form of a file or a stream to the receiving apparatus via a digital storage medium or a network.
[0062] The source apparatus can include a video source, an encoding device, and a transmitter. The receiving apparatus can include a receiver, a decoding device, and a renderer. The encoding device can be referred to as a video / image encoding device, and the decoding device can be referred to as a video / image decoding device. The transmitter can be included in the encoding device. The receiver can be included in the decoding device. The renderer can include a display, and the display can be configured as a separate apparatus or an external component.
[0063] The video source can obtain a video / image through a process of capturing, synthesizing, or generating a video / image. The video source can include a video / image capturing apparatus and / or a video / image generating apparatus. The video / image capturing apparatus can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generating apparatus can include, for example, a computer, a tablet, and a smart phone, and can (electronically) generate a video / image. For example, a virtual video / image can be generated through a computer, etc. In this case, the video / image capturing process can be replaced by a process of generating related data.
[0064] An encoding apparatus can encode an input video / image. The encoding apparatus can perform a series of processes such as prediction, transform, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0065] A transmitter can transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving apparatus in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include an element for generating a media file through a predetermined file format, and can include an element for transmission through a broadcasting / communication network. The receiver can receive / extract a bitstream, and transmit the received / extracted bitstream to a decoding apparatus.
[0066] A decoding apparatus can decode a video / image by performing a series of processes such as dequantization, inverse transform, prediction, etc. corresponding to the operations of the encoding apparatus.
[0067] A renderer can render the decoded video / image. The rendered video / image can be displayed through a display.
[0068] Figure 2 FIG. 1 is a diagram schematically illustrating a configuration of a video / image encoding apparatus to which the present disclosure is applicable. Hereinafter, the so-called video encoding apparatus can include an image encoding apparatus.
[0069] Referring to Figure 2 , the encoding apparatus 200 can include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be constituted by one or more hardware components (e.g., an encoder chipset or a processor). In addition, the memory 270 can include a decoded picture buffer (DPB), and can be constituted by a digital storage medium. The hardware components can further include the memory 270 as an internal / external component.
[0070] The picture partitioner 210 can partition an input picture (or a picture or a frame) input to the encoding apparatus 200 into one or more processing units. As one example, the processing unit can be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a largest coding unit (LCU), the coding units can be recursively partitioned according to a quad tree binary tree ternary tree (QTBTTT) structure. For example, based on a quad tree structure, a binary tree structure, and / or a ternary tree structure, one coding unit can be partitioned into a plurality of coding units at a deeper depth. In this case, for example, the quad tree structure can be first applied, and the binary tree structure and / or the ternary tree structure can be later applied. Alternatively, the binary tree structure can be first applied. The encoding process according to the disclosure can be performed based on a final coding unit that is not further partitioned. In this case, based on the coding efficiency according to the characteristics of the picture, the largest coding unit can be directly used as the final coding unit. Alternatively, the coding unit can be recursively partitioned into a coding unit at a deeper depth as necessary, whereby a coding unit of an optimal size can be used as the final coding unit. Here, the encoding process can include processes such as prediction, transform, and reconstruction, which will be described later. As another example, the processing unit can further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be separated or partitioned from the above-described final coding unit. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0071] Depending on the situation, a unit and a term such as a block, a region, and the like can be used instead of each other. In general cases, an MxN block can denote a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally denote a pixel or a pixel value, and can denote only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component. A sample can be used as a term corresponding to a pixel or a pel of one picture (or image).
[0072] The subtractor 231 subtracts the prediction signal (prediction block, prediction sample array) output from the predictor 220 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the transformer 232. The predictor 220 can perform prediction on a processing target block (hereinafter referred to as a "current block"), and can generate a prediction block including predicted samples of the current block. The predictor 220 can determine whether to apply intra prediction or to apply inter prediction on a basis of the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various information related to prediction, such as prediction mode information, and transmit the generated information to the entropy encoder 240. The information about prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0073] The intra predictor 222 can predict the current block by referring to samples in the current picture. The reference samples can be located in the vicinity of the current block or separated from the current block according to the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. The directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of detail of the prediction direction. However, this is merely an example, and more or less directional prediction modes can be used according to the settings. The intra predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0074] The inter predictor 221 can derive a prediction block for a current block based on a reference block (a reference sample array) designated by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted on a block, sub-block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same as or different from each other. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), or the like, and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. The inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter predictor 221 can use the motion information of the neighboring blocks as the motion information of the current block. In the skip mode, unlike the merge mode, a residual signal cannot be transmitted. In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling a motion vector difference.
[0075] The predictor 220 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction to the prediction of one block, and can also simultaneously apply intra prediction and inter prediction. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode can be used for content image / video encoding of a game or the like such as screen content coding (SCC). Although the IBC basically performs prediction in the current block, it is similar to the inter prediction in that it derives a reference block in the current block. That is, the IBC can use at least one of the inter prediction techniques described in the present disclosure.
[0076] The prediction signal generated by the inter-predictor 221 and / or the intra-predictor 222 can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, the GBT means a transform obtained from a graph when relationship information between pixels is expressed in a graph. The CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to a square pixel block having the same size, or can be applied to a block having a variable size other than a square block.
[0077] The quantizer 233 can quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 can encode the quantized signals (information about the quantized transform coefficients) and output the encoded signals in a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the quantized transform coefficients of a block type into a one-dimensional vector form based on a coefficient scan order, and generate the information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 can perform various encoding methods such as, for example, exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), and the like. The entropy encoder 240 can encode information required for video / image reconstruction, other than the quantized transform coefficients (e.g., values of syntax elements, etc.), together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream on a unit basis of a network abstraction layer (NAL). The video / image information can further include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), and the like. In addition, the video / image information can further include regular constraint information. In the disclosure, information and / or syntax elements transmitted from the encoding apparatus to / signaled to the decoding apparatus can be included in the video / image information. The video / image information can be encoded through the encoding process described above and included in the bitstream. The bitstream can be transmitted through a network, or stored in a digital storage medium. Here, the network can include a broadcasting network, a communication network, and / or the like, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, and the like. A transmitter (not shown) that transmits the signals output from the entropy encoder 240 or a memory (not shown) that stores the same can be configured as an internal / external element of the encoding apparatus 200, or the transmitter can be included in the entropy encoder 240.
[0078] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transform to the quantized transform coefficients using the dequantizer 234 and the inverse transformer 235, a residual signal (a residual block or residual samples) can be reconstructed. The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-predictor 221 or the intra-predictor 222, so that a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array) can be generated. When there is no residual for the processing target block as in the case of applying a skip mode, the prediction block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of the next processing target block in the target picture, and as subsequently described, can be used for inter-prediction of the next picture by filtering.
[0079] Further, in the picture encoding and / or reconstruction processing, luma mapping with chroma scaling (LMCS) can be applied.
[0080] The filter 260 can improve subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can store the modified reconstructed picture in the memory 270, especially in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filter, sample adaptive offset, adaptive loop filter, bilateral filter, etc. As discussed subsequently in the description of each filtering method, the filter 260 can generate various information related to the filtering, and transmit the generated information to the entropy encoder 240. The information about the filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0081] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter-predictor 221. Thereby, the encoding apparatus can avoid prediction mismatch in the encoding apparatus 100 and the decoding apparatus when applying inter-prediction, and can also improve encoding efficiency.
[0082] The memory 270 DPB can store the modified reconstructed picture in order to use it as a reference picture in the inter-predictor 221. The memory 270 can store motion information of a block in a current picture from which motion information has been derived (or encoded) and / or motion information of a block in a picture that has been reconstructed. The stored motion information can be transmitted to the inter-predictor 221 to be used as motion information of a neighboring block or motion information of a temporally neighboring block. The memory 270 can store reconstructed samples of a reconstructed block in the current picture, and transmit them to the intra-predictor 222.
[0083] Figure 3FIG. 1 is a diagram schematically illustrating a configuration of a video / image decoding apparatus to which the present disclosure is applicable.
[0084] Referring to Figure 3 , the video decoding apparatus 300 can include an entropy decoder 310, a residue processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an intra predictor 331 and an inter predictor 332. The residue processor 320 can include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residue processor 320, the predictor 330, the adder 340, and the filter 350 described above can be constituted by one or more hardware components (e.g., a decoder chipset or a processor). In addition, the memory 360 can include a decoded picture buffer (DPB), and can be constituted by a digital storage medium. The hardware components can further include the memory 360 as an internal / external component.
[0085] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image in correspondence with a process by which the video / image information has been processed in the encoding apparatus from which the bitstream has been obtained. Figure 2 For example, the decoding apparatus 300 can derive a unit / block based on information related to block partitioning obtained from the bitstream. The decoding apparatus 300 can perform decoding by using a processing unit to which a process applied in the encoding apparatus is applied. Accordingly, the decoded processing unit can be, for example, an encoding unit, which can be partitioned with a coding tree unit or a largest coding unit following a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived with the encoding unit. Also, a reconstructed image signal decoded and output by the decoding apparatus 300 can be reproduced through a reproducer.
[0086] The decoding apparatus 300 can receive a bitstream from Figure 2The signal output from the encoding apparatus can be received and decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can further include information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc. In addition, the video / image information can further include general constraint information. The decoding apparatus can further decode a picture based on the information on the parameter sets and / or the general constraint information. In the disclosure, the information and / or syntax elements to be signaled / received, which will be described later, can be decoded by the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode information in the bitstream based on an encoding method such as an exponential Golomb coding, a CAVLC, a CABAC, etc., and can output values of syntax elements required for image reconstruction and quantized values of transform coefficients on a residual. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using decoded information of a target syntax element and decoded information of neighboring and a decoding target block or a symbol / bin decoded in a previous step, predict a bin generation probability according to the determined context model, and perform arithmetic decoding on the bin to generate a symbol corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model using information of a symbol / bin decoded for a next context model after determining the context model. Information on prediction among the information decoded in the entropy decoder 310 can be provided to the predictor (inter-predictor 332 and intra-predictor 331), and residual values (i.e., quantized transform coefficients) and associated parameter information on which entropy decoding has been performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (a residual block, a residual sample, a residual sample array). In addition, information on filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) that receives a signal output from the encoding apparatus can also constitute the decoding apparatus 300 as an internal / external element, and the receiver can be a component of the entropy decoder 310. Furthermore, the decoding apparatus according to the disclosure can be referred to as a video / image / picture encoding apparatus, and the decoding apparatus can be divided into an information decoder (a video / image / picture information decoder) and a sample decoder (a video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of the dequantizer 321, the inverse transformer 322, the adder 340, the filter 350, the memory 360, the inter-predictor 332, and the intra-predictor 331.
[0087] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into the form of a two-dimensional block. In this case, the rearrangement can be performed based on the order of coefficient scanning that has been performed in the encoding apparatus. The dequantizer 321 can perform dequantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step information), and obtain transform coefficients.
[0088] The dequantizer 321 obtains a residual signal (a residual block, a residual sample array) by inverse-transforming the transform coefficients.
[0089] The predictor can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on information about prediction output from the entropy decoder 310, and specifically can determine an intra / inter prediction mode.
[0090] The predictor can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction to prediction of one block, and can also simultaneously apply intra prediction and inter prediction. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can perform intra block copy (IBC) for prediction of a block. Intra block copy can be used for content image / video encoding of contents such as games, etc., screen content coding (SCC). Although IBC basically performs prediction in the current block, it is similar to inter prediction in that it derives a reference block in the current block. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.
[0091] The intra predictor 331 can predict the current block by referring to samples in the current picture. The reference samples can be located in the vicinity of the current block or separated from the current block according to the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0092] The inter predictor 332 can derive a prediction block for the current block based on a reference block (a reference sample array) designated by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted on a block, sub-block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter predictor 332 can configure a motion information candidate list based on the neighboring blocks, and derive a motion vector and / or a reference picture index of the current block based on received candidate selection information. The inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating a mode of inter prediction for the current block.
[0093] The adder 340 can generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array) by adding the obtained residual signal to a prediction signal (a prediction block, a prediction sample array) output from the predictor 330. When there is no residual for the processing target block as in the case of applying a skip mode, the prediction block can be used as the reconstructed block.
[0094] The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra prediction of a next processing target block in the current block, and as subsequently described, can be output through filtering or used for inter prediction of a next picture.
[0095] Further, in the picture decoding process, luma mapping with chroma scaling (LMCS) can be applied.
[0096] The filter 350 can improve subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can transmit the modified reconstructed picture to the memory 360, especially to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0097] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction 332. The memory 360 can store motion information of a block in a current picture from which motion information has been derived (or decoded) and / or motion information of a block in a picture that has been reconstructed. The stored motion information can be sent to the inter prediction 332 to be used as motion information of a neighboring block or motion information of a temporally neighboring block. The memory 360 can store reconstructed samples of a reconstructed block in a current picture and send them to the intra prediction 331.
[0098] In this specification, the examples described in the predictors 330, the dequantizer 321, the inverse transformer 322, and the filter 350 of the decoding device 300 can be similarly or correspondingly applied to the predictors 220, the dequantizer 234, the inverse transformer 235, and the filter 260 of the encoding device 200, respectively.
[0099] As described above, prediction is performed in order to improve compression efficiency when performing video encoding. Accordingly, a prediction block including prediction samples for a current block that is an encoding target block can be generated. Here, the prediction block includes prediction samples in a spatial domain (or pixel domain). The prediction block can be derived identically in the encoding device and the decoding device, and the encoding device can improve image encoding efficiency by signaling, to the decoding device, information (residual information) on a residual between the original block and the prediction block, rather than original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed picture including the reconstructed block.
[0100] The residual information can be generated through a transform process and a quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on residual samples (a residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients, so that it can signal associated residual information to the decoding device (through a bitstream). Here, the residual information can include value information of the quantized transform coefficients, position information, a transform technique, a transform kernel, a quantization parameter, and the like. The decoding device can perform a quantization / dequantization process based on the residual information and derive residual samples (or a residual sample block). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can derive the residual block by dequantizing / inverse-transforming the quantized transform coefficients in order to be used as a reference for inter prediction of a next picture, and can generate a reconstructed picture based thereon.
[0101] Figure 4 A structure of a content streaming system to which the present disclosure is applied is exemplified.
[0102] Further, a content streaming system applying the present disclosure can generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0103] The encoding server serves to compress content input from a multimedia input device such as a smart phone, a camera, a camcorder, etc. into digital data to generate a bitstream, and transmit it to the streaming server. As another example, in the case where a multimedia input device such as a smart phone, a camera, a camcorder, etc. directly generates a bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or the bitstream generation method of the present disclosure. And the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0104] The streaming server transmits multimedia data to a user device through a web server based on a request of a user, and the web server serves as a means for informing a user of what services exist. When a user requests a service that the user wants, the web server transmits the request to the streaming server, and the streaming server transmits multimedia data to the user. In this regard, the content streaming system can include a separate control server, and in this case, the control server serves to control commands / responses between respective devices in the content streaming system.
[0105] The streaming server can receive content from the media storage and / or the encoding server. For example, in the case of receiving content from the encoding server, the content can be received in real time. In this case, in order to smoothly provide a streaming service, the streaming server can store a bitstream for a predetermined time.
[0106] For example, the user device can include a mobile phone, a smart phone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation, a board PC, a tablet PC, an ultrabook, a wearable device (e.g., a watch-type terminal (smart watch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, a digital signage, etc. The respective servers in the content streaming system can operate as distributed servers, and in this case, data received by the respective servers can be processed in a distributed manner.
[0107] Figure 5 A multi-transform technique according to an embodiment of the present disclosure is schematically illustrated.
[0108] Referring to Figure 5 , the transformer can correspond to the transformer in the encoding apparatus of the foregoing Figure 2 , and the inverse transformer can correspond to the inverse transformer in the encoding apparatus of the foregoing Figure 2 , or the inverse transformer in the decoding apparatus of the foregoing Figure 3 .
[0109] The transformer can derive (primary) transform coefficients by performing a primary transform based on residual samples (an array of residual samples) in a residual block (S510). The primary transform can be referred to as a core transform. In this document, the primary transform can be based on a multiple transform selection (MTS), and when multiple transforms are used as the primary transform, it can be referred to as a multi-core transform.
[0110] The multi-core transform can denote a method of performing a transform using a discrete cosine transform (DCT) type 2 and a discrete sine transform (DST) type 7, a DCT type 8, and / or a DST type 1 in addition. That is, the multi-core transform can denote a transform method of transforming a residual signal (or a residual block) in a spatial domain into transform coefficients (or primary transform coefficients) in a frequency domain based on multiple transform kernels selected from among a DCT type 2, a DST type 7, a DCT type 8, and a DST type 1. In this document, the primary transform coefficients can be referred to as temporary transform coefficients from the perspective of the transformer.
[0111] In other words, when a conventional transform method is applied, transform coefficients can be generated by applying a transform from a spatial domain to a frequency domain based on a DCT type 2 to a residual signal (or a residual block). Unlike this, when a multi-core transform is applied, transform coefficients (or primary transform coefficients) can be generated by applying a transform from a spatial domain to a frequency domain based on a DCT type 2, a DST type 7, a DCT type 8, and / or a DST type 1 to a residual signal (or a residual block). In this document, the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1 can be referred to as a transform type, a transform kernel, or a transform core. These DCT / DST transform types can be defined based on a basis function.
[0112] When the multi-core transform is performed, a vertical transform kernel and a horizontal transform kernel for a target block can be selected from among transform kernels, a vertical transform can be performed on the target block based on the vertical transform kernel, and a horizontal transform can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform can indicate a transform on a horizontal component of the target block, and the vertical transform can indicate a transform on a vertical component of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on a prediction mode of a target (CU or sub-block) including the residual block and / or a transform index.
[0113] Further, according to an example, if one transform is performed by applying the MTS, a mapping relationship of the transform kernel can be set by setting a certain basis function to a predetermined value and combining basis functions to be applied in the vertical transform or the horizontal transform. For example, when a horizontal transform kernel is denoted as trTypeHor and a vertical direction transform kernel is denoted as trTypeVer, trTypeHor or trTypeVer having a value of 0 can be set to DCT2, trTypeHor or trTypeVer having a value of 1 can be set to DST7, and trTypeHor or trTypeVer having a value of 2 can be set to DCT8.
[0114] In this case, MTS index information can be encoded and signaled to a decoding device to indicate any one of a plurality of transform kernel sets. For example, MTS index 0 can indicate that both trTypeHor and trTypeVer values are 0, MTS index 1 can indicate that both trTypeHor and trTypeVer values are 1, MTS index 2 can indicate that trTypeHor value is 2 and trTypeVer value is 1, MTS index 3 can indicate that trTypeHor value is 1 and trTypeVer value is 2, and MTS index 4 can indicate that both trTypeHor and trTypeVer values are 2.
[0115] In one example, transform kernel sets according to MTS index information are shown in the following table.
[0116] [Table 1]
[0117] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2
[0118] The transformer can perform a secondary transform based on the (primary) transform coefficients to derive modified (secondary) transform coefficients (S520). The primary transform is a transform from a spatial domain to a frequency domain, and the secondary transform refers to a transform into a more compact representation using a correlation existing between the (primary) transform coefficients. The secondary transform can include a non-separable transform. In this case, the secondary transform can be referred to as a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The NSST can denote a transform that performs a secondary transform on the (primary) transform coefficients derived through the primary transform based on a non-separable transform matrix to generate modified transform coefficients (or secondary transform coefficients) for the residual signal. Here, based on the non-separable transform matrix, the transform can be applied to the (primary) transform coefficients without separating (or independently applying the horizontal / vertical transform) the vertical transform and the horizontal transform. In other words, the NSST is not applied to the (primary) transform coefficients in the vertical direction and the horizontal direction separately, and can denote, for example, a transform method that rearranges a two-dimensional signal (transform coefficients) into a one-dimensional signal through a specific predetermined direction (e.g., a row-major direction or a column-major direction) and then generates modified transform coefficients (or secondary transform coefficients) based on a non-separable transform matrix. For example, the row-major order is to arrange M×N blocks in the order of the first row, the second row,..., and the Nth row, and the column-major order is to arrange M×N blocks in the order of the first column, the second column,..., and the Mth column. The NSST can be applied to a left-top region of a block (hereinafter, referred to as a transform coefficient block) configured with the (primary) transform coefficients. For example, when both a width W and a height H of the transform coefficient block are 8 or more, 8×8 NSST can be applied to a left-top 8×8 region of the transform coefficient block. Also, while both the width (W) and the height (H) of the transform coefficient block are 4 or more, when the width (W) or the height (H) of the transform coefficient block is less than 8, 4×4 NSST can be applied to a left-top min(8, W)×min(8, H) region of the transform coefficient block. However, embodiments are not limited thereto, and for example, even if only the condition that the width W or the height H of the transform coefficient block is 4 or more is satisfied, 4×4 NSST can be applied to the left-top min(8, W)×min(8, H) region of the transform coefficient block.
[0119] Specifically, for example, if a 4×4 input block is used, the non-separable secondary transform can be performed as follows.
[0120] The 4×4 input block X can be expressed as follows.
[0121] [Equation 1]
[0122]
[0123] If X is expressed in the form of a vector, the vector may be expressed as follows.
[0124] [Equation 2]
[0125]
[0126] In Equation 2, the vector is a one-dimensional vector obtained by rearranging the two-dimensional block X of Equation 1 according to row-major order.
[0127] In this case, the non-separable quadratic transform can be calculated as follows.
[0128] [Equation 3]
[0129]
[0130] In this equation, denotes a transform coefficient vector, and T denotes a 16x16 (non-separable) transform matrix.
[0131] Through Equation 3 described above, the 16x1 transform coefficient vector may be derived, and the vector may be reorganized into 4x4 blocks through a scan order (horizontal, vertical, and diagonal, etc.). However, the above calculation is an example, and a hypercube-Givens transform (HyGT) or the like can also be used for the calculation of the non-separable quadratic transform in order to reduce the calculation complexity of the non-separable quadratic transform.
[0132] Further, in the non-separable quadratic transform, a transform kernel (or transform core, transform type) can be selected to be mode-dependent. In this case, the mode can include an intra-prediction mode and / or an inter-prediction mode.
[0133] As described above, the non-separable quadratic transform can be performed based on an 8x8 transform or a 4x4 transform determined based on a width (W) and a height (H) of a transform coefficient block. The 8x8 transform refers to a transform applicable to an 8x8 region included in the transform coefficient block when both W and H are equal to or greater than 8, and the 8x8 region can be a top-left 8x8 region in the transform coefficient block. Similarly, the 4x4 transform refers to a transform applicable to a 4x4 region included in the transform coefficient block when both W and H are equal to or greater than 4, and the 4x4 region can be a top-left 4x4 region in the transform coefficient block. For example, an 8x8 transform kernel matrix can be a 64x64 / 16x64 matrix, and a 4x4 transform kernel matrix can be a 16x16 / 8x16 matrix.
[0134] Here, to select a mode-dependent transform kernel, two non-separable secondary transform kernels for each transform set for non-separable secondary transform can be configured for both 8x8 transform and 4x4 transform, and there can be four transform sets. That is, four transform sets can be configured for 8x8 transform, and four transform sets can be configured for 4x4 transform. In this case, each of the four transform sets for 8x8 transform can include two 8x8 transform kernels, and each of the four transform sets for 4x4 transform can include two 4x4 transform kernels.
[0135] However, as the size of the transform (i.e., the size of the region to which the transform is applied) can be a size other than 8x8 or 4x4, for example, the number of sets can be n, and the number of transform kernels in each set can be k.
[0136] The transform set can be referred to as an NSST set or an LFNST set. A particular set among the transform sets can be selected, for example, based on an intra prediction mode of a current block (CU or sub-block). Low-frequency non-separable transform (LFNST) can be an example of a reduced non-separable transform, which will be described later, and denotes a non-separable transform for a low-frequency component.
[0137] For reference, for example, the intra prediction mode can include two non-directional (or non-angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non-directional intra prediction mode can include a planar intra prediction mode of 0 number and a DC intra prediction mode of 1 number, and the directional intra prediction mode can include 65 intra prediction modes of 2 number to 66 number. However, this is an example, and the present document can be applied even if the number of intra prediction modes is different. Also, in some cases, a 67th intra prediction mode can also be used, and the 67th intra prediction mode can denote a linear model (LM) mode.
[0138] Figure 6 An intra directional mode of 65 prediction directions is schematically shown.
[0139] Referring to Figure 6 , based on the intra prediction mode 34 having an upper-left diagonal prediction direction, the intra prediction mode can be divided into an intra prediction mode having horizontal directionality and an intra prediction mode having vertical directionality. In Figure 6In this context, H and V denote horizontal and vertical directionality, respectively, and the numbers -32 to 32 indicate a displacement in 1 / 32 units on the sample grid position. These numbers can represent an offset for the mode index value. Intra prediction modes 2 to 33 have horizontal directionality and intra prediction modes 34 to 66 have vertical directionality. Strictly speaking, intra prediction mode 34 can be considered neither horizontal nor vertical, but can be classified as belonging to horizontal directionality when determining the transform set of the secondary transform. This is because the input data is transposed for the vertically oriented mode that is symmetric about intra prediction mode 34, and the input data alignment method for horizontal modes is used for intra prediction mode 34. Transposing the input data means switching the rows and columns of a two-dimensional M x N block of data into N x M data. Intra prediction mode 18 and intra prediction mode 50 can represent a horizontal intra prediction mode and a vertical intra prediction mode, respectively, and intra prediction mode 2 can be referred to as a top-right diagonal intra prediction mode because it has a left reference pixel and performs prediction in a top-right direction. Similarly, intra prediction mode 34 can be referred to as a bottom-right diagonal intra prediction mode, and intra prediction mode 66 can be referred to as a bottom-left diagonal intra prediction mode.
[0140] According to an example, four transform sets according to intra prediction modes can be mapped, for example, as shown in the following table.
[0141] [Table 2]
[0142] lfnstPredModeIntra lfnstTrSetIdx lfnstPredModeIntra < 0 1 0 <= lfnstPredModeIntra <= 1 0 2 <= lfnstPredModeIntra <= 12 1 13 <= lfnstPredModeIntra <= 23 2 24 <= lfnstPredModeIntra <= 44 3 45 <= lfnstPredModeIntra <= 55 2 56 <= lfnstPredModeIntra <= 80 1 81 <= lfnstPredModeIntra <= 83 0
[0143] As shown in Table 2, according to intra prediction modes, any one of the four transform sets, i.e., lfnstTrSetldx, can be mapped to any one of the four indices, i.e., 0 to 3.
[0144] When a particular set is determined to be used for a non-separable transform, one of the k transform kernels in the particular set can be selected by a non-separable secondary transform index. The encoding device can derive the non-separable secondary transform index indicating the particular transform kernel based on rate-distortion (RD) checks, and can signal the non-separable secondary transform index to the decoding device. The decoding device can select one of the k transform kernels in the particular set based on the non-separable secondary transform index. For example, lfnst index value 0 can refer to a first non-separable secondary transform kernel, lfnst index value 1 can refer to a second non-separable secondary transform kernel, and lfnst index value 2 can refer to a third non-separable secondary transform kernel. Alternatively, lfnst index value 0 can indicate that a first non-separable secondary transform is not applied to the target block, and lfnst index values 1 to 3 can indicate three transform kernels.
[0145] The transformer can perform a non-separable secondary transform based on the selected transform kernel, and can obtain modified (secondary) transform coefficients. As described above, the modified transform coefficients can be derived as transform coefficients quantized by the quantizer, and can be encoded and signaled to the decoding device, and delivered to the dequantizer / inverse transformer in the encoding device.
[0146] Further, as described above, if the secondary transform is omitted, the (primary) transform coefficients that are output as the primary (separable) transform can be derived as transform coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device, and delivered to the dequantizer / inverse transformer in the encoding device.
[0147] The inverse transformer can perform a series of processes in an order opposite to the order that has been performed in the above-described transformer. The inverse transformer can receive the (dequantized) transform coefficients, and derive the (primary) transform coefficients by performing a secondary (inverse) transform (S550), and can obtain the residual block (residual samples) by performing a primary (inverse) transform on the (primary) transform coefficients (S560). In this regard, the primary transform coefficients can be referred to as modified transform coefficients from the perspective of the inverse transformer. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0148] The decoding device can further include a secondary inverse transform application determiner (or an element for determining whether to apply a secondary inverse transform) and a secondary inverse transform determiner (or an element for determining a secondary inverse transform). The secondary inverse transform application determiner can determine whether to apply a secondary inverse transform. For example, the secondary inverse transform can be NSST, RST, or LFNST, and the secondary inverse transform application determiner can determine whether to apply the secondary inverse transform based on a secondary transform flag obtained by parsing a bitstream. In another example, the secondary inverse transform application determiner can determine whether to apply the secondary inverse transform based on transform coefficients of a residual block.
[0149] The secondary inverse transform determiner can determine a secondary inverse transform. In this case, the secondary inverse transform determiner can determine a secondary inverse transform applied to a current block based on a LFNST (NSST or RST) transform set specified according to an intra prediction mode. In an embodiment, a secondary transform determination method can be determined depending on a primary transform determination method. Various combinations of primary and secondary transforms can be determined according to an intra prediction mode. Further, in an example, the secondary inverse transform determiner can determine a region to which a secondary inverse transform is applied based on a size of the current block.
[0150] Further, as described above, if the secondary (inverse) transform is omitted, the (dequantized) transform coefficients can be received, the primary (separable) inverse transform can be performed, and the residual block (residual samples) can be obtained. As described above, the encoding device and the decoding device can generate the reconstructed block based on the residual block and the prediction block, and can generate the reconstructed picture based on the reconstructed block.
[0151] Further, in the disclosure, a reduced secondary transform (RST) in which the size of a transform matrix (kernel) is reduced can be applied in the concept of the NSST in order to reduce the amount of calculation and the amount of storage required for the non-separable secondary transform.
[0152] Further, the transform kernel, the transform matrix, and the coefficients constituting the transform kernel matrix, i.e., the kernel coefficients or the matrix coefficients, described in the disclosure can be represented in 8 bits. This can be a condition that is implemented in the decoding device and the encoding device, and compared to the existing 9 bits or 10 bits, the amount of storage required to store the transform kernel can be reduced, and performance degradation can be reasonably accommodated. In addition, representing the kernel matrix in 8 bits can allow the use of a small multiplier, and can be more suitable for single instruction multiple data (SIMD) instructions used for optimal software implementation.
[0153] In the present specification, the term "RST" can refer to a transform performed on residual samples of a target block based on a transform matrix whose size is reduced according to a reduction factor. In the case of performing a reduced transform, the amount of calculation required for the transform can be reduced due to the reduction in the size of the transform matrix. That is, the RST can be used to solve the problem of computational complexity that occurs when transforming a block of large size or a non-separable transform.
[0154] The RST can be referred to as various terms such as reduced transform, reduced secondary transform, downsize transform, simplified transform, and simple transform, and the name by which the RST can be referred to is not limited to the listed examples. Alternatively, since the RST is mainly performed in a low frequency region including non-zero coefficients in a transformed block, it can be referred to as a low frequency non-separable transform (LFNST). The transform index can be referred to as an LFNST index.
[0155] Further, when performing a secondary inverse transform based on the RST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 can include an inverse reduced secondary transformer that derives modified transform coefficients based on an inverse RST of the transform coefficients, and an inverse primary transformer that derives residual samples of the target block based on an inverse primary transform of the modified transform coefficients. The inverse primary transform refers to an inverse transform of the primary transform applied to the residual. In the disclosure, deriving the transform coefficients based on the transform can refer to deriving the transform coefficients by applying the transform.
[0156] Figure 7is a diagram illustrating an RST according to an embodiment of the disclosure.
[0157] In the disclosure, a "target block" can refer to a current block to be encoded, a residual block, or a transform block.
[0158] In the RST according to an example, an N-dimensional vector can be mapped to an R-dimensional vector located in another space, and thus a reduced transform matrix can be determined, where R is smaller than N. N can refer to the square of the length of a side of a block to which a transform is applied, or the total number of transform coefficients corresponding to a block to which a transform is applied, and a reduction factor can refer to an R / N value. The reduction factor can be referred to as a reduction factor, a downsize factor, a simplification factor, a simple factor, or other various terms. Also, R can be referred to as a reduction coefficient, but the reduction factor can refer to R according to circumstances. Also, the reduction factor can refer to an N / R value according to circumstances.
[0159] In an example, the reduction factor or the reduction coefficient can be signaled through a bitstream, but the example is not limited thereto. For example, a predetermined value for the reduction factor or the reduction coefficient can be stored in each of the encoding apparatus 200 and the decoding apparatus 300, and in this case, the reduction factor or the reduction coefficient can not be separately signaled.
[0160] The size of the reduced transform matrix according to an example can be R×N, which is smaller than N×N (the size of a regular transform matrix), and can be defined as in Equation 4 below.
[0161] [Equation 4]
[0162]
[0163] Figure 7 The matrix T in the reduced transform block shown in (a) of FIG. 1 can refer to the matrix T of Equation 4 R×N . As Figure 7 indicated in (a) of FIG. 1, when the reduced transform matrix T R×N is multiplied by the residual samples of the target block, the transform coefficients of the current block can be derived.
[0164] In an example, if the size of the block to which a transform is applied is 8×8 and R=16 (i.e., R / N=16 / 64=1 / 4), the RST according to (a) of FIG. 1 can be expressed as a matrix operation shown in Equation 5 below. In this case, the storage and multiplication calculation can be reduced to about 1 / 4 by the reduction factor. Figure 7
[0165] In the disclosure, a matrix operation can be understood as an operation of obtaining a column vector by multiplying the column vector with a matrix disposed on the left side of the column vector.
[0166] [Equation 5]
[0167]
[0168] In Equation 5, r1to r 64 may represent residual samples of a target block, and specifically can be transform coefficients generated by applying a primary transform. As a result of the calculation of Equation 5, transform coefficients c1to c i of the target block can be derived, and the process of deriving c i may be shown in Equation 6.
[0169] [Equation 6]
[0170]
[0171] As a result of the calculation of Equation 6, transform coefficients c1to c R of the target block can be derived. That is, when R = 16, transform coefficients c1to c 16 If a regular transform is applied instead of RST, and a transform matrix of 64 x 64 (N x N) size is multiplied by residual samples of 64 x 1 (N x 1) size, 16 (R) transform coefficients are derived for the target block only because RST is applied, although 64 (N) transform coefficients are derived for the target block. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data transmitted by the encoding device 200 to the decoding device 300 is reduced, and thus the transmission efficiency between the encoding device 200 and the decoding device 300 can be improved.
[0172] When considered from the perspective of the size of the transform matrix, the size of the regular transform matrix is 64 x 64 (N x N), but the size of the reduced transform matrix is reduced to 16 x 64 (R x N), and thus the storage usage rate can be reduced by a ratio of R / N in the case of performing RST compared to the case of performing a regular transform. In addition, when compared to the number of multiplication calculations N x N in the case of using a regular transform matrix, the number of multiplication calculations (R x N) can be reduced by a ratio of R / N using a reduced transform matrix.
[0173] In an example, the transformer 232 of the encoding device 200 can derive transform coefficients of a target block by performing a primary transform and a secondary transform based on RST on residual samples of the target block. These transform coefficients can be communicated to an inverse transformer of the decoding device 300, and the inverse transformer 322 of the decoding device 300 can derive modified transform coefficients based on inverse reduced secondary transform (RST) for the transform coefficients, and can derive residual samples of the target block based on inverse primary transform for the modified transform coefficients.
[0174] The inverse RST matrix T N×RThe size of the reduced inverse transform block (b) is N x R, which is smaller than the size of the regular inverse transform matrix N x N, and the reduced transform matrix T R×N has a transposed relationship.
[0175] Figure 7 The matrix T t in the reduced inverse transform block (b) can be referred to as an inverse RST matrix T N×R T (superscript T means transposition). As Figure 7 indicated in (b), when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block or the residual samples of the target block can be derived. The inverse RST matrix T R×N T can be represented as (T R×N ) T N×R .
[0176] More specifically, when the inverse RST is used as a secondary inverse transform, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. In addition, the inverse RST can be used as an inverse primary transform, and in this case, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, the residual samples of the target block can be derived.
[0177] In an example, if the size of the block to which the inverse transform is applied is 8 x 8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), the RST according to (b) of Figure 7 can be represented as a matrix operation as shown in the following equation 7.
[0178] [Equation 7]
[0179]
[0180] In equation 7, c1 to c 16 may represent the transform coefficients of the target block. As a result of the calculation of equation 7, r j representing the modified transform coefficients of the target block or the residual samples of the target block can be derived, and the process of deriving r j may be as shown in equation 8.
[0181] [Equation 8]
[0182]
[0183] As a result of the calculation of Equation 8, r1 to r N From the perspective of the size of the inverse transform matrix, the size of the regular inverse transform matrix is 64x64 (NxN), but the size of the inverse reduced transform matrix is reduced to 64x16 (RxN), and thus the storage usage rate can be reduced by an R / N ratio in the case of performing the inverse RST compared to the case of performing the regular inverse transform. In addition, the number of multiplication calculations can be reduced by an R / N ratio (NxR) using the inverse reduced transform matrix when compared to the number of multiplication calculations NxN in the case of using the regular inverse transform matrix.
[0184] The transform set configuration shown in Table 2 can also be applied to 8x8 RST. That is, 8x8 RST can be applied according to the transform set in Table 2. Since one transform set includes two or three transforms (kernels) according to the intra prediction mode, it can be configured to select one of at most four transforms included in the case where the secondary transform is not applied. In the transform in which the secondary transform is not applied, it can be considered that an identity matrix is applied. Assuming that indices 0, 1, 2, and 3 are respectively assigned to the four transforms (for example, the index 0 can be assigned to the case where the identity matrix is applied, i.e., the case where the secondary transform is not applied), the transform index or lfnst index as a syntax element can be signaled for each transform coefficient block, thereby designating the transform to be applied. That is, for the top-left 8x8 block, through the transform index, 8x8 NSST in the RST configuration can be designated, or 8x8 lfnst when the LFNST is applied can be designated. The 8x8 lfnst and 8x8 RST refer to a transform that can be applied to an 8x8 region included in a transform coefficient block when W and H of a target block to be transformed are equal to or greater than 8, and the 8x8 region can be a top-left 8x8 region in the transform coefficient block. Similarly, the 4x4 lfnst and 4x4 RST refer to a transform that can be applied to a 4x4 region included in a transform coefficient block when W and H of a target block are equal to or greater than 4, and the 4x4 region can be a top-left 4x4 region in the transform coefficient block.
[0185] According to embodiments of the disclosure, for a transform in an encoding process, only 48 pieces of data can be selected, and a maximum 16x48 transform kernel matrix can be applied thereto, instead of applying a 16x64 transform kernel matrix to 64 pieces of data forming an 8x8 region. Here, "maximum" means that m has a maximum value of 16 in an m x 48 transform kernel matrix for generating m coefficients. That is, when RST is performed by applying an m x 48 transform kernel matrix (m ≤ 16) to an 8x8 region, 48 pieces of data are input, and m coefficients are generated. When m is 16, 48 pieces of data are input and 16 coefficients are generated. That is, assuming that the 48 pieces of data form a 48x1 vector, the 16x48 matrix and the 48x1 vector are sequentially multiplied, thereby generating a 16x1 vector. Here, the 48 pieces of data forming the 8x8 region can be appropriately arranged, thereby forming a 48x1 vector. For example, the 48x1 vector can be configured based on the 48 pieces of data constituting a region other than the right-bottom 4x4 region among the 8x8 region. Here, when matrix operation is performed by applying the maximum 16x48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients can be arranged in the top-left 4x4 region according to a scan order, and the top-right 4x4 region and the bottom-left 4x4 region can be padded with zeros.
[0186] For inverse transform in a decoding process, a transpose matrix of the aforementioned transform kernel matrix can be used. That is, when inverse RST or LFNST is performed in an inverse transform process performed by a decoding device, input coefficient data to which inverse RST is applied is configured in a one-dimensional vector according to a predetermined arrangement order, and a modified coefficient vector obtained by multiplying the one-dimensional vector with a corresponding inverse RST matrix on the left side of the one-dimensional vector can be arranged into a two-dimensional block according to the predetermined arrangement order.
[0187] In summary, in a transform process, when RST or LFNST is applied to an 8x8 region, 48 transform coefficients in the top-left region, the top-right region, and the bottom-left region of the 8x8 region other than the right-bottom region are subjected to matrix operation with a 16x48 transform kernel matrix. For the matrix operation, the 48 transform coefficients are input in a one-dimensional array. When the matrix operation is performed, 16 modified transform coefficients are derived, and the modified transform coefficients can be arranged in the top-left region of the 8x8 region.
[0188] In contrast, in the inverse transform process, when inverse RST or LFNST is applied to the 8x8 region, 16 transform coefficients corresponding to the upper left region of the 8x8 region among the transform coefficients in the 8x8 region can be input in a one-dimensional array according to the scan order, and can undergo a matrix operation with a 48x16 transform kernel matrix. That is, the matrix operation can be expressed as (48x16 matrix)*(16x1 transform coefficient vector)=(48x1 modified transform coefficient vector). Here, an nx1 vector can be interpreted to have the same meaning as an nx1 matrix, and thus can be expressed as an nx1 column vector. Also, * denotes matrix multiplication. When the matrix operation is performed, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper left region, the upper right region, and the lower left region in the 8x8 region except for the lower right region.
[0189] When the secondary inverse transform is based on RST, the inverse transformer 235 of the encoding apparatus 200 and the inverse transformer 322 of the decoding apparatus 300 can include an inverse reduced secondary transformer for deriving modified transform coefficients based on inverse RST on transform coefficients and an inverse primary transformer for deriving residual samples of a target block based on inverse primary transform on the modified transform coefficients. The inverse primary transform refers to an inverse transform of a primary transform applied to a residual. In the present disclosure, deriving transform coefficients based on a transform can refer to deriving transform coefficients by applying a transform.
[0190] The non-separable transform (LFNST) described above will be described in detail as follows. The LFNST can include a forward transform by an encoding apparatus and an inverse transform by a decoding apparatus.
[0191] The encoding apparatus receives a result (or a part of the result) derived after applying a primary (core) transform as input, and applies a forward secondary transform (secondary transform).
[0192] [Equation 9]
[0193] y = G T x
[0194] In Equation 9, x and y are an input and an output of the secondary transform, respectively, G is a matrix representing the secondary transform, and a transform basis vector is composed of a column vector. In the case of inverse LFNST, when the dimension of the transform matrix G is expressed as [number of rows x number of columns], in the case of forward LFNST, the transpose of the matrix G becomes G T .
[0195] For inverse LFNST, the dimensions of the matrix G are [48x16], [48x8], [16x16], [16x8], and the [48x8] matrix and the [16x8] matrix are partial matrices of 8 transform basis vectors sampled from the left side of the [48x16] matrix and the [16x16] matrix, respectively.
[0196] On the other hand, for forward LFNST, the dimensions of the matrix G T are [16x48], [8x48], [16x16], [8x16], and the [8x48] matrix and the [8x16] matrix are partial matrices obtained by sampling 8 transform basis vectors from the upper part of the [16x48] matrix and the [16x16] matrix, respectively.
[0197] Accordingly, in the case of forward LFNST, a [48x1] vector or a [16x1] vector can be inputted as x, and a [16x1] vector or an [8x1] vector can be outputted as y. In video encoding and decoding, the output of forward primary transform is two-dimensional (2D) data, and thus in order to construct a [48x1] vector or a [16x1] vector as input x, it is necessary to construct a one-dimensional vector by appropriately arranging the 2D data which is the output of forward transform.
[0198] Figure 8 is a diagram illustrating an order of arranging output data of forward primary transform into a one-dimensional vector according to an example. Figure 8 The left diagram of (a) and (b) of FIG. 1 illustrates an order for constructing a [48x1] vector, and Figure 8 The right diagram of (a) and (b) of FIG. 1 illustrates an order for constructing a [16x1] vector. In the case of LFNST, a one-dimensional vector x can be obtained by sequentially arranging 2D data in the same order as Figure 8 (a) and (b) of FIG. 1.
[0199] The arrangement direction of output data of forward primary transform can be determined according to the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is in a horizontal direction with respect to a diagonal direction, the output data of forward primary transform can be arranged in the order of (a) of FIG. 1, and when the intra prediction mode of the current block is in a vertical direction with respect to a diagonal direction, the output data of forward primary transform can be arranged in the order of (b) of FIG. 1. Figure 8 Figure 8 The arrangement direction of output data of forward primary transform can be determined according to the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is in a horizontal direction with respect to a diagonal direction, the output data of forward primary transform can be arranged in the order of (a) of FIG. 1, and when the intra prediction mode of the current block is in a vertical direction with respect to a diagonal direction, the output data of forward primary transform can be arranged in the order of (b) of FIG. 1.
[0200] According to an example, an arrangement order different from the arrangement order of (a) and (b) of FIG. 1 can be applied, and in order to derive an arrangement order to be applied, a value of a parameter can be derived. Figure 8 Figure 8 The arrangement order of (a) and (b) of FIG. 9 is the same result (y vector), and the column vectors of the matrix G can be rearranged according to the arrangement order. That is, the column vectors of G can be rearranged such that each element constituting the x vector is always multiplied by the same transform basis vector.
[0201] Since the output y derived through Equation 9 is a one-dimensional vector, when two-dimensional data is required as input data in a process of using the result of the forward secondary transform as input (for example, in a process of performing quantization or residual coding), the output y vector of Equation 9 needs to be arranged as 2D data again as appropriate.
[0202] Figure 9 is a diagram illustrating an order of arranging output data of the forward secondary transform into a two-dimensional block according to an example.
[0203] In the case of LFNST, the output values can be arranged in a 2D block according to a predetermined scan order. Figure 9 (a) of FIG. 10 shows that when the output y is a [16x1] vector, the output values are arranged at 16 positions of a 2D block according to a diagonal scan order. Figure 9 (b) of FIG. 10 shows that when the output y is an [8x1] vector, the output values are arranged at 8 positions of a 2D block according to a diagonal scan order, and the remaining 8 positions are padded with zeros. Figure 9 X in (b) of FIG. 10 indicates that it is padded with zeros.
[0204] According to another example, since the order of processing the output vector y when performing quantization or residual coding can be preset, the output vector y can not be arranged in a 2D block as shown in Figure 9 However, in the case of residual coding, data coding can be performed in a 2D block (for example, 4x4) unit (for example, CG (coefficient group)), and in this case, the data is arranged according to a specific order in the diagonal scan order as Figure 9 of FIG. 11.
[0205] Further, a decoding device can configure a one-dimensional input vector y by arranging two-dimensional data output through a dequantization process according to a preset scan order for inverse transform. The input vector y can be output as an output vector x through the following equation.
[0206] [Equation 10]
[0207] X = Gy
[0208] In the case of inverse LFNST, an output vector x can be derived by multiplying an input vector y, which is a [16x1] vector or a [8x1] vector, by a G matrix. For inverse LFNST, the output vector x can be a [48x1] vector or a [16x1] vector.
[0209] The output vector x is arranged in a two-dimensional block according to the order shown in Figure 8 and is arranged as two-dimensional data, and this two-dimensional data becomes input data (or a part of input data) of the inverse primary transform.
[0210] Accordingly, the inverse secondary transform is the inverse of the forward secondary transform process as a whole, and in the case of the inverse transform, unlike in the forward direction, the inverse secondary transform is applied first, and then the inverse primary transform is applied.
[0211] In the inverse LFNST, one of 8 [48x16] matrices and 8 [16x16] matrices can be selected as the transform matrix G. Whether to apply the [48x16] matrix or the [16x16] matrix depends on the size and shape of the block.
[0212] In addition, the 8 matrices can be derived from four transform sets as shown in Table 2 above, and each transform set can consist of two matrices. Which transform set is used among the 4 transform sets is determined according to the intra prediction mode, and more specifically, the transform set is determined based on the value of the intra prediction mode extended by considering wide-angle intra prediction (WAIP). Which matrix is selected from among the two matrices constituting the selected transform set is derived by index signaling. More specifically, 0, 1, and 2 can be transmitted as index values, 0 can indicate that the LFNST is not applied, and 1 and 2 can indicate any one of the two transform matrices constituting the transform set selected based on the intra prediction mode value.
[0213] Figure 10 is a diagram illustrating a wide-angle intra prediction mode according to an embodiment of the present document.
[0214] The general intra prediction mode value can have values from 0 to 66 and from 81 to 83, and the intra prediction mode value extended due to WAIP can have values from -14 to 83 as shown. The values from 81 to 83 indicate CCLM (cross-component linear model) modes, and the values from -14 to -1 and from 67 to 80 indicate intra prediction modes extended due to WAIP application.
[0215] When the width of the current prediction block is greater than the height, the upper reference pixels are generally closer to the position inside the block to be predicted. Therefore, prediction in the lower-left direction can be more accurate than in the upper-right direction. Conversely, when the height of the block is greater than the width, the left reference pixels are generally closer to the position inside the block to be predicted. Therefore, prediction in the upper-right direction can be more accurate than in the lower-left direction. Accordingly, it can be advantageous to apply remapping (i.e., mode index modification) to the index of the wide-angle intra prediction mode.
[0216] When the wide-angle intra prediction is applied, information on the existing intra prediction can be signaled, and after the information is parsed, the information can be remapped to the index of the wide-angle intra prediction mode. Thus, the total number of the intra prediction modes for a specific block (e.g., a non-square block of a specific size) can not be changed, that is, the total number of the intra prediction modes is 67, and the intra prediction mode coding for the specific block can not be changed.
[0217] Table 3 below shows a process of deriving a modified intra mode by remapping the intra prediction mode to the wide-angle intra prediction mode.
[0218] [Table 3]
[0219]
[0220] In Table 3, the extended intra prediction mode value is finally stored in the predModeIntra variable, and ISP_NO_SPLIT indicates that the CU block is not divided into sub-blocks by the intra sub-block partitioning (ISP) technique currently adopted in the VVC standard, and the cIdx variable values of 0, 1, and 2 indicate the cases of the luma component, the Cb component, and the Cr component, respectively. The log2 function shown in Table 3 returns a log value with a base of 2, and the Abs function returns an absolute value.
[0221] The variable predModeIntra indicating the intra prediction mode, and the height and width of the transform block, etc. are used as input values for the wide-angle intra prediction mode mapping process, and the output value is the modified intra prediction mode predModeIntra. The height and width of the transform block or the coding block can be the height and width of the current block for which the remapping of the intra prediction mode is performed. At this time, the variable whRatio reflecting the ratio of the width to the height can be set to Abs(Log2(nW / nH)).
[0222] For a non-square block, the intra prediction mode can be divided into two cases and modified.
[0223] First, if all of conditions (1) to (3) are satisfied, (1) the width of the current block is greater than the height, (2) the intra prediction mode before the modification is equal to or greater than 2, and (3) the intra prediction mode is less than a value derived as (8+2*whRatio) when the variable whRatio is greater than 1 and is less than 8 when the variable whRatio is less than or equal to 1 (predModeIntra is less than (whRatio>1)? (8+2*whRatio) : 8), the intra prediction mode is set to a value greater than predModeIntra by 65 [predModeIntra is set to equal to (predModeIntra+65)].
[0224] If different from the above, i.e., if conditions (1) to (3) are met, (1) the height of the current block is greater than the width, (2) the intra prediction mode before the modification is less than or equal to 66, and (3) the intra prediction mode is greater than a value derived as (60 - 2*whRatio) when whRatio is greater than 1 and is greater than 60 when whRatio is less than or equal to 1 (predModeIntra is greater than (whRatio>1)?(60-2*whRatio):60), the intra prediction mode is set to a value 67 less than predModeIntra [predModeIntra is set to equal to (predModeIntra-67)].
[0225] Table 2 above shows how the transform set is selected in LFNST based on the intra prediction mode value extended by WAIP. As shown, modes 14 to 33 and modes 35 to 80 are symmetric about the prediction direction around mode 34. For example, mode 14 and mode 54 are symmetric about the direction corresponding to mode 34. Therefore, the same transform set is applied to the modes located in the directions symmetric to each other, and this symmetry is also reflected in Table 2. Figure 10
[0226] In addition, it is assumed that the forward LFNST input data for mode 54 is symmetric to the forward LFNST input data for mode 14. For example, for mode 14 and mode 54, the two-dimensional data is rearranged to one-dimensional data according to the arrangement order shown in (a) and (b) of Figure 8 Figure 8 In addition, it can be seen that the pattern of the order shown in (a) and (b) of Figure 8 Figure 8 is symmetric about the direction (diagonal direction) indicated by mode 34.
[0227] In addition, as described above, which of the [48x16] matrix and the [16x16] matrix is applied to LFNST is determined by the size and shape of the transform target block.
[0228] Figure 11 is a diagram illustrating the block shape to which LFNST is applied. Figure 11 (a) of FIG. 1 shows a 4x4 block, Figure 11 (b) of FIG. 1 shows a 4x8 block and an 8x4 block, Figure 11 (c) of FIG. 1 shows a 4xN block or an Nx4 block, where N is 16 or more, Figure 11 (d) of FIG. 1 shows an 8x8 block, Figure 11 (e) of FIG. 1 shows an MxN block, where M≥8, N≥8, and N>8 or M>8.
[0229] In Figure 11 In the diagram, blocks with thick boundaries indicate the area where LFNST is applied. For Figure 11 For blocks (a) and (b), LFNST is applied to the top-left 4×4 region, and for Figure 11 Block (c) is individually applied to two consecutively arranged top-left 4×4 regions. Figure 11 In (a), (b), and (c), since the LFNST is applied in units of 4×4 regions, this LFNST will be referred to as "4×4 LFNST" in the following text. Based on the matrix dimension of G, a [16×16] or [16×8] matrix can be applied.
[0230] More specifically, a [16×8] matrix is applied to Figure 11 (a) consists of 4×4 blocks (4×4TU or 4×4CU), and a [16×16] matrix is applied to... Figure 11 The blocks in (b) and (c) are used to adjust the worst-case computational complexity to 8 multiplications per sample.
[0231] about Figure 8 In (d) and (e), LFNST is applied to the top-left 8×8 region, and this LFNST is referred to as "8×8 LFNST" below. As the corresponding transformation matrix, a [48×16] matrix or a [48×8] matrix can be applied. In the case of the forward LFNST, since the [48×1] vector (the X vector in Equation 9) is input as input data, not all sample values from the top-left 8×8 region are used as input values for the forward LFNST. That is, if... Figure 8 The left-hand order of (a) or Figure 11 As can be seen from the left-hand order of (b), the [48×1] vector can be constructed based on the samples belonging to the other three 4×4 blocks while leaving the bottom right 4×4 block as is.
[0232] A [48×8] matrix can be applied to Figure 11 The 8×8 blocks (8×8TU or 8×8CU) in (d) and the [48×16] matrix can be applied Figure 12 The 8×8 blocks in (e). This is also to adjust the worst-case computational complexity to 8 multiplications per sample.
[0233] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data (the Y vector in Equation 9, [8×1] or [16×1] vectors) are generated. In the forward LFNST, due to matrix G... T Due to its characteristic, the amount of output data is equal to or less than the amount of input data.
[0234] Figure 12 is a diagram illustrating an arrangement of output data of a forward LFNST according to an example, and shows a block in which output data of a forward LFNST is arranged according to a block shape.
[0235] In Figure 12 the shaded area at the top left of the block corresponds to an area in which output data of a forward LFNST is located, the positions marked with 0 indicate samples filled with the value 0, and the remaining area represents an area that is not changed by the forward LFNST. In the area that is not changed by the LFNST, the output data of the forward primary transform remains unchanged.
[0236] As described above, since the size of the transform matrix applied varies according to the shape of the block, the number of output data also varies. As Figure 12 , the output data of a forward LFNST can not completely fill the top left 4x4 block. In the case of (a) and (d) of Figure 9 , a [16x8] matrix and a [48x8] matrix are applied to the block or the partial area inside the block indicated by the thick line, respectively, and an [8x1] vector is generated as the output of the forward LFNST. That is, according to the scan order shown in (b) of Figure 12 , only 8 output data can be filled, as shown in (a) and (d) of Figure 11 , and 0 can be filled in the remaining 8 positions. In the case of the block to which the LFNST of (d) of Figure 12 is applied, as shown in (d) of Figure 12 , the two 4x4 blocks adjacent to the top left 4x4 block, the top right and the bottom left, are also filled with the value 0.
[0237] As described above, basically, by signaling the LFNST index, it is prescribed whether to apply the LFNST and the transform matrix to be applied. As Figure 12 shown, when the LFNST is applied, since the number of output data of the forward LFNST can be equal to or less than the number of input data, areas filled with zero values occur.
[0238] 1) As shown in (a) of Figure 12 , samples from the eighth position and subsequent positions in the scan order in the top left 4x4 block, i.e., from the ninth to the sixteenth.
[0239] 2) As shown in (d) and (e) of Figure 11 , when a [48x16] matrix or a [48x8] matrix is applied, the two 4x4 blocks adjacent to the top left 4x4 block or the second and third 4x4 blocks in the scan order.
[0240] Thus, if there is non-zero data by checking the regions 1) and 2), it is determined that LFNST is not applied, so that signaling of the corresponding LFNST index can be omitted.
[0241] According to an example, in the case of LFNST adopted in the VVC standard, since signaling of the LFNST index is performed after residual coding, the encoding device can know whether non-zero data (significant coefficients) exist at all positions within a TU or CU block through residual coding. Thus, the encoding device can determine whether to perform signaling of the LFNST index based on the presence of non-zero data, and the decoding device can determine whether to parse the LFNST index. When non-zero data does not exist in the regions specified in the above 1) and 2), signaling of the LFNST index is performed.
[0242] Further, for the adopted LFNST, the following simplified method can be applied.
[0243] (i) According to an example, the number of output data of the forward LFNST can be limited to a maximum of 16.
[0244] In the case of (c) of Figure 11 , 4x4 LFNST can be applied to two 4x4 regions adjacent to the upper left, respectively, and in this case, a maximum of 32 LFNST output data can be generated. When the number of output data of the forward LFNST is limited to a maximum of 16, in the case of a 4xN / Nx4 (N≥16) block (TU or CU), 4x4 LFNST is applied to only one 4x4 region of the upper left, and LFNST can be applied to Figure 13 all blocks at a time. By this, the implementation of image coding can be simplified.
[0245] (ii) According to an example, zeroing can be additionally applied to a region to which LFNST is not applied. In this document, zeroing can mean filling all positions belonging to a certain region with a value of 0. That is, zeroing can be applied to a region that is not changed due to LFNST, and the result of the forward one-time transform is maintained. As described above, since LFNST is divided into 4x4 LFNST and 8x8 LFNST, zeroing can be divided into two types ((ii)-(A) and (ii)-(B)) as follows.
[0246] (ii)-(A) When 4x4 LFNST is applied, a region to which 4x4 LFNST is not applied can be zeroed. Figure 13 is a diagram illustrating zeroing in a block to which 4x4 LFNST is applied according to an example.
[0247] As Figure 12 indicated, for a block to which 4x4 LFNST is applied, that is, for a block to which 4x4 LFNST is not applied, zeroing can be applied to a region adjacent to the upper left. Figure 13All blocks in (a), (b) and (c) can be zero-filled even if the entire region where LFNST is not applied.
[0248] on the other hand, Figure 14 (d) shows, according to an example, the remaining blocks for which 4×4 LFNST is not applied are zeroed when the maximum number of output data for the positive LFNST is limited to 16.
[0249] (ii)-(B) When 8×8LFNST is applied, areas where 8×8LFNST is not applied can be zeroed. Figure 14 This is a diagram illustrating the zeroing process in a block of 8×8 LFNST, based on the example.
[0250] like Figure 12 As shown, regarding the application of 8×8LFNST, that is, for Figure 12 In all blocks in (d) and (e), the entire region where LFNST is not applied can be filled with zeros.
[0251] (iii) Due to the zeroing proposed in (ii) above, the zero-filled regions can be different when LFNST is applied. Therefore, according to the zeroing proposed in (ii), and Figure 12 Compared to the case of LFNST, it can check for the presence of non-zero data over a wider area.
[0252] For example, when applying (ii)-(B), except Figure 14 Outside of the zero-filled regions in (d) and (e), Figure 12 After checking for non-zero data in the additional zero-filled region, the LFNST index can be signaled only if no non-zero data exists.
[0253] Of course, even with the zeroing proposed in application (ii), the presence of non-zero data can be checked in the same way as existing LFNST index signaling. That is, when checking for non-zero data... Figure 12 After confirming the presence of non-zero data within the zero-padded block, LFNST index signaling can be applied. In this case, the encoding device only performs a zeroing operation, and the decoding device does not assume this zeroing; that is, it only checks whether non-zero data exists within the zero-padded block. Figure 13 In regions explicitly marked as 0, LFNST index resolution can be performed.
[0254] Various implementations of the simplified methods for applying LFNST (combinations of (i), (ii)-(A), (ii)-(B), (iii)) can be derived. Of course, the combinations of the above simplified methods are not limited to the following implementations, and any combination can be applied to LFNST.
[0255] Embodiments
[0256] - Limit the number of output data of the forward LFNST to a maximum of 16 → (i)
[0257] - When the 4x4 LFNST is applied, all regions where the 4x4 LFNST is not applied are zeroed out → (ii)-(A)
[0258] - When the 8x8 LFNST is applied, all regions where the 8x8 LFNST is not applied are zeroed out → (ii)-(B)
[0259] - After checking whether non-zero data also exists in the existing region filled with zero values and the region filled with zero due to the additional zeroing ((ii)-(A), (ii)-(B)), the LFNST index is signaled only when non-zero data does not exist → (iii).
[0260] In the case of the embodiments, when the LFNST is applied, the region in which non-zero output data can exist is limited to the inside of the top-left 4x4 region. More specifically, in (a) of Figure 14 Figure 13 In (a) of Figure 14 Figure 15 In (b) and (d) of
[0261] Therefore, after the LFNST is applied, after checking whether non-zero data exists in a position that the residual coding process does not allow (at a position beyond the last position), it can be determined whether to signal the LFNST index.
[0262] In the case of the zeroing method proposed in (ii), due to the number of data finally generated when both the primary transform and the LFNST are applied, the amount of calculation required to perform the entire transform process can be reduced. That is, when the LFNST is applied, since zeroing is applied to the forward primary transform output data existing in a region where the LFNST is not applied, there is no need to generate data for a region that becomes zeroed during the execution of the forward primary transform. Therefore, the amount of calculation required to generate the corresponding data can be reduced. The additional effects of the zeroing method proposed in (ii) are summarized as follows.
[0263] First, as described above, the amount of calculation required to perform the entire transform process is reduced.
[0264] In particular, when applying (ii)-(B), the worst case calculation amount is reduced, so that the transform process can be made lighter. In other words, generally, a large amount of calculation is required to perform a primary transform of a large size. By applying (ii)-(B), the amount of data derived as a result of performing the forward LFNST can be reduced to 16 or less. In addition, as the size of the entire block (TU or CU) increases, the effect of reducing the amount of transform operation further increases.
[0265] Second, the amount of calculation required for the entire transform process can be reduced, thereby reducing the power consumption required to perform the transform.
[0266] Third, the delay involved in the transform process is reduced.
[0267] A secondary transform such as LFNST adds the amount of calculation to the existing primary transform, thus increasing the overall delay time involved in performing the transform. In particular, in the case of intra prediction, since the reconstructed data of the neighboring block is used in the prediction process, the increase in the delay due to the secondary transform during encoding results in an increase in the delay until the reconstruction. This can result in an increase in the overall delay of the intra prediction encoding.
[0268] However, if the zeroing out proposed in (ii) is applied, the delay time of performing the primary transform can be greatly reduced when applying the LFNST, the delay time of the entire transform is maintained or reduced, so that the encoding apparatus can be more simply implemented.
[0269] In addition, in the conventional intra prediction, the encoding target block is regarded as one encoding unit, and encoding is performed without being divided. However, intra sub partition (ISP) encoding means that the intra prediction encoding is performed by dividing the block to be currently encoded in a horizontal direction or a vertical direction. In this case, the reconstructed block can be generated by performing encoding / decoding in units of the divided block, and the reconstructed block can be used as a reference block for the next divided block. According to the embodiment, in the ISP encoding, one encoding block can be divided into two or four sub blocks and encoded, and in the ISP, intra prediction is performed in one sub block by referring to the reconstructed pixel values of the sub block located in the adjacent left side or the adjacent upper side. Hereinafter, "encoding" can be used as a concept including both encoding performed by an encoding apparatus and decoding performed by a decoding apparatus.
[0270] The ISP is a technique of dividing a block for which prediction is performed into two or four sub partitions in a vertical direction or a horizontal direction according to the size of the block. For example, the minimum block size to which the ISP can be applied is 4x8 or 8x4. When the block size is greater than 4x8 or 8x4, the block is divided into 4 sub partitions.
[0271] When the ISP is applied, the sub-blocks are sequentially encoded from left to right or from top to bottom according to the division type (e.g., horizontally or vertically), and after reconstruction processing is performed via inverse transformation and intra prediction for one sub-block, encoding of the next sub-block can be performed. For the leftmost or topmost sub-block, reconstructed pixels of an already encoded coding block are referred to as in the conventional intra prediction method. Also, when each side of a subsequent internal sub-block is not adjacent to a previous sub-block, reconstructed pixels of an already encoded neighboring coding block are referred to in order to derive reference pixels adjacent to the corresponding side as in the conventional intra prediction method.
[0272] In the ISP encoding mode, all of the sub-blocks can be encoded with the same intra prediction mode, and a flag indicating whether ISP encoding is used and a flag indicating a direction (horizontal or vertical) in which division is performed can be signaled. At this time, the number of sub-blocks can be adjusted to 2 or 4 depending on the block shape, and when the size (width x height) of one sub-block is less than 16, division can not be allowed for the corresponding sub-block, or application of ISP encoding itself can be set to be limited.
[0273] Also, in the case of the ISP prediction mode, one coding unit is divided into two or four sub-blocks and is predicted, and the same intra prediction mode is applied to the two or four sub-blocks that are divided.
[0274] As described above, in the division direction, both a horizontal direction (when an MxN coding unit having a horizontal length and a vertical length of M and N, respectively, is divided in the horizontal direction, if the MxN coding unit is divided into two, the MxN coding unit is divided into Mx(N / 2) blocks, and if the MxN coding unit is divided into four, the MxN coding unit is divided into Mx(N / 4) blocks) and a vertical direction (when the MxN coding unit is divided in the vertical direction, if the MxN coding unit is divided into two, the MxN coding unit is divided into (M / 2)xN blocks, and if the MxN coding unit is divided into four, the MxN coding unit is divided into (M / 4)xN blocks) are possible. When the MxN coding unit is divided in the horizontal direction, the sub-blocks are encoded in an up-down order, and when the MxN coding unit is divided in the vertical direction, the sub-blocks are encoded in a left-right order. In the case of horizontal (vertical) direction division, the reconstructed pixel values of the upper (left) sub-block can be referred to in order to predict the currently encoded sub-block.
[0275] A transform can be applied to a residual signal generated in a unit of a partition block through an ISP prediction method. A multi-transform selection (MTS) technique based on a DST-7 / DCT-8 combination and an existing DCT-2 can be applied to a forward one-time transform (core transform) and a forward low-frequency non-separable transform (LFNST) can be applied to a transform coefficient generated according to the one-time transform to generate a final modified transform coefficient.
[0276] That is, the LFNST can be applied to the partition blocks divided by applying the ISP prediction mode and the same intra prediction mode is applied to the divided partition blocks as described above. Accordingly, when the LFNST set derived based on the intra prediction mode is selected, the derived LFNST set can be applied to all of the partition blocks. That is, because the same intra prediction mode is applied to all of the partition blocks, the same LFNST set can be applied to all of the partition blocks.
[0277] Further, according to an example, the LFNST can be applied only to a transform block having both a horizontal length and a vertical length of 4 or more. Accordingly, when the horizontal length or the vertical length of the partition block divided according to the ISP prediction method is less than 4, the LFNST is not applied and the LFNST index is not signaled. Further, when the LFNST is applied to each partition block, the corresponding partition block can be considered as one transform block. Of course, when the ISP prediction method is not applied, the LFNST can be applied to the coding block.
[0278] The application of the LFNST to each partition block will be described in detail as follows.
[0279] According to an example, after the forward LFNST is applied to each partition block, only up to 16 (8 or 16) coefficients are left in the top-left 4x4 region in the scan order of the transform coefficients, and then a zero-out can be applied in which the remaining positions and regions are all filled with 0.
[0280] Alternatively, according to an example, when the length of one side of the partition block is 4, the LFNST is applied only to the top-left 4x4 region, and when the length of all sides (i.e., width and height) of the partition block is 8 or more, the LFNST can be applied to the remaining 48 coefficients within the top-left 8x8 region except for the bottom-right 4x4 region.
[0281] Alternatively, according to an example, in order to adjust the worst-case computational complexity to 8 multiplications per sample, when each partition block is 4x4 or 8x8, only 8 transform coefficients can be output after the forward LFNST is applied. That is, when the partition block is 4x4, an 8x16 matrix can be applied as a transform matrix, and when the partition block is 8x8, an 8x48 matrix can be applied as a transform matrix.
[0282] Also, in the current VVC standard, LFNST index signaling is performed in units of coding units. Therefore, in the ISP prediction mode and when LFNST is applied to all partition blocks, the same LFNST index value can be applied to the corresponding partition blocks. That is, when an LFNST index value is transmitted once at the coding unit level, the corresponding LFNST index can be applied to all partition blocks in the coding unit. As described above, the LFNST index value can have values of 0, 1, and 2, in which 0 indicates a case where LFNST is not applied, and 1 and 2 indicate two transform matrices existing in one LFNST set when LFNST is applied.
[0283] As described above, the LFNST set is determined by the intra prediction mode, and in the case of the ISP prediction mode, because all partition blocks in the coding unit are predicted in the same intra prediction mode, the partition blocks can refer to the same LFNST set.
[0284] As another example, LFNST index signaling is still performed in units of coding units, but in the case of the ISP prediction mode, it is not determined whether to uniformly apply LFNST to all partition blocks, and for each partition block, it can be determined whether to apply the LFNST index value signaled at the coding unit level and whether to apply LFNST through a separate condition. Here, the separate condition can be signaled in the form of a flag for each partition block through a bitstream, and when the flag value is 1, the LFNST index value signaled at the coding unit level is applied, and when the flag value is 0, LFNST can not be applied.
[0285] Hereinafter, a method of maintaining the worst case computational complexity when applying LFNST to the ISP mode will be described.
[0286] In the case of the ISP mode, when LFNST is applied, in order to maintain the number of multiplications per sample (or per coefficient, per position) to a certain value or less, the application of LFNST can be limited. According to the size of the partition block, by applying LFNST as follows, the number of multiplications per sample (or per coefficient, per position) can be maintained to 8 or less.
[0287] 1. When both the horizontal length and the vertical length of the partition block are 4 or more, the same method as the worst case computational complexity control method for LFNST in the current VVC standard can be applied.
[0288] That is, when the partition block is a 4x4 block, an 8x16 matrix obtained by sampling the upper 8 rows from a 16x16 matrix can be applied instead of the 16x16 matrix in the forward direction, and a 16x8 matrix obtained by sampling the left 8 columns from the 16x16 matrix can be applied in the inverse direction. Also, when the partition block is an 8x8 block, in the forward direction, instead of a 16x48 matrix, an 8x48 matrix obtained by sampling the upper 8 rows from the 16x48 matrix is applied, and in the inverse direction, instead of a 48x16 matrix, a 48x8 matrix obtained by sampling the left 8 columns from the 48x16 matrix can be applied.
[0289] In the case of a 4xN or Nx4 (N>4) block, when the forward transform is performed, 16 coefficients generated after the 16x16 matrix is applied only to the upper left 4x4 block can be disposed in the upper left 4x4 region, and the other regions can be filled with a value of 0. Also, when the inverse transform is performed, the 16 coefficients located in the upper left 4x4 block are disposed in a scan order to form an input vector, and then 16 output data can be generated by multiplication by the 16x16 matrix. The generated output data can be disposed in the upper left 4x4 region, and the remaining regions except for the upper left 4x4 region can be filled with a value of 0.
[0290] In the case of an 8xN or Nx8 (N>8) block, when the forward transform is performed, 16 coefficients generated after the 16x48 matrix is applied to an ROI region within only the upper left 8x8 block (the remaining regions except for the right lower 4x4 block from the upper left 8x8 block) can be disposed in the upper left 4x4 region, and all the other regions can be filled with a value of 0. Also, when the inverse transform is performed, the 16 coefficients located in the upper left 4x4 region are disposed in a scan order to form an input vector, and then 48 output data can be generated by multiplication by the 48x16 matrix. The generated output data can be filled in the ROI region, and all the other regions can be filled with a value of 0.
[0291] As another example, to keep the number of multiplications per sample (or per coefficient, per position) to a certain value or less, the number of multiplications per sample (or per coefficient, per position) based on the ISP coding unit size, not the size of the ISP partition block, can be kept to 8 or less. When only one block among the ISP partition blocks satisfies the condition for applying LFNST, the worst case complexity calculation of applying LFNST can be based on the corresponding coding unit size, not the size of the partition block. For example, if a luma coding block of a particular coding unit is coded with ISP by being split into four partition blocks having a 4x4 size, and there are no non-zero transform coefficients for two of them, it can be configured such that instead of generating 8, 16 transform coefficients are generated for each of the other two partition blocks.
[0292] Hereinafter, a method of signaling the LFNST index in the case of the ISP mode will be described.
[0293] As described above, the LFNST index can have values 0, 1, and 2, in which 0 indicates that the LFNST is not applied, and 1 and 2 indicate one of the two LFNST kernel matrices included in the selected LFNST set. The method of transmitting the LFNST index in the current VVC standard will be described as follows.
[0294] 1. The LFNST index can be transmitted once for each coding unit (CU). In the case of dual tree, separate LFNST indices can be signaled for each of the luma block and the chroma block.
[0295] 2. When the LFNST index is not signaled, the LFNST index is inferred to be 0 as a default value. The LFNST index value is inferred to be 0 in the following cases.
[0296] A. When in a mode in which a transform is not applied (e.g., transform skip, BDPCM, lossless coding, etc.)
[0297] B. When the primary transform is not DCT-2 (when it is DST7 or DCT8), that is, when the horizontal transform or the vertical transform is not DCT-2
[0298] C. When the horizontal length or the vertical length of the luma block of the coding unit exceeds the size of the transformable maximum luma transform, for example, when the size of the luma block of the coding block is 128x16 in the case where the size of the transformable maximum luma transform is 64, the LFNST is not applied.
[0299] In the case of dual tree, it is determined whether each of the coding unit of the luma component and the coding unit of the chroma component exceeds the size of the maximum luma transform. That is, it is checked whether the luma block exceeds the size of the transformable maximum luma transform, and it is checked whether the chroma block exceeds the horizontal / vertical length of the corresponding luma block for the color format and the size of the maximum transformable luma transform. For example, when the color format is 4:2:0, each of the horizontal / vertical length of the corresponding luma block is twice the horizontal / vertical length of the chroma block, and the transform size of the corresponding luma block is twice the transform size of the chroma block. As another example, when the color format is 4:4:4, the horizontal / vertical length of the corresponding luma block is the same as the horizontal / vertical length of the chroma block.
[0300] 64-length transform or 32-length transform can mean a transform applied horizontally or vertically with a length of 64 or 32, respectively, and the "transform size" can mean the corresponding length of 64 or 32.
[0301] In the single tree case, it is checked whether the horizontal or vertical length of the luma block exceeds the size of the transformable maximum luma transform block, and if it exceeds the size, the LFNST index signaling can be skipped.
[0302] D. The LFNST index can be transmitted only when both the horizontal length and the vertical length of the coding unit are greater than or equal to 4.
[0303] In the dual tree case, the LFNST index can be signaled only for the case where both the horizontal length and the vertical length of the corresponding component (i.e., luma or chroma component) are greater than or equal to 4.
[0304] In the single tree case, the LFNST index can be signaled only for the case where both the horizontal length and the vertical length of the luma component are greater than or equal to 4.
[0305] E. In the case where the last non-zero coefficient position is not the DC position (the top-left position of the block), when it is a dual tree type luma block, the LFNST index is transmitted if the last non-zero coefficient position is not the DC position. When it is a dual tree type chroma block, the LFNST index is transmitted if even one of the last non-zero coefficient position of Cb and the last non-zero coefficient position of Cr is not the DC position.
[0306] In the single tree type case, the LFNST index is transmitted if even one of the last non-zero coefficient positions in the luma component, the Cb component, and the Cr component is not the DC position.
[0307] Here, if a coded block flag (CBF) value indicating whether there is a transform coefficient for one transform block is 0, the last non-zero coefficient position of the transform block is not checked to determine whether to signal the LFNST index. That is, if the CBF value is 0, since the transform is not applied to the block, the last non-zero coefficient position can not be considered when checking the condition for the LFNST index signaling.
[0308] For example, 1) in the case of the dual tree type and the luma component, if the corresponding CBF value is 0, the LFNST index is not signaled, 2) in the case of the dual tree type and the chroma component, if the CBF value of Cb is 0 and the CBF value of Cr is 1, the LFNST index is transmitted by checking only the last non-zero coefficient position of Cr, and 3) in the case of the single tree type, for all of luma, Cb, and Cr, the last non-zero coefficient position is checked only for the component for which the corresponding CBF value is 1.
[0309] F.When it is confirmed that there is a transform coefficient at a position other than a position where an LFNSF transform coefficient can exist, LFNST index signaling can be skipped. In the case of 4x4 transform blocks and 8x8 transform blocks, LFNST transform coefficients can exist at 8 positions from the DC position according to the transform coefficient scan order in the VVC standard, and all the remaining positions are padded with 0. In addition, when the transform block is not a 4x4 transform block and an 8x8 transform block, LFNST transform coefficients can exist at 16 positions from the DC position according to the transform coefficient scan order in the VVC standard, and all the remaining positions are padded with 0.
[0310] Therefore, if a non-zero transform coefficient exists in a region to be padded with 0 after performing residual coding, LFNST index signaling can be omitted.
[0311] In addition, the ISP mode can be applied only to a luma block or can be applied to both a luma block and a chroma block. As described above, when ISP prediction is applied, prediction is performed by splitting a corresponding coding unit into 2 or 4 sub-blocks, and a transform can be applied to each sub-block. Therefore, when a condition in which an LFNST index is signaled is determined on a coding unit basis, the fact that LFNST is applicable to each sub-block should also be considered. In addition, when an ISP prediction mode is applied only to a specific component (e.g., a luma block), the LFNST index should be signaled by considering the fact that splitting into sub-blocks is achieved only for the component. The possible LFNST index signaling schemes when in the ISP mode are summarized as follows.
[0312] 1. The LFNST index can be signaled once per coding unit (CU). In the case of dual tree, separate LFNST indices can be signaled for each of the luma and chroma blocks.
[0313] 2. When the LFNST index is not signaled, the LFNST index is inferred to be 0 as a default value. The LFNST index value is inferred to be 0 in the following cases.
[0314] A. When in a mode in which no transform is applied (e.g., transform skip, BDPCM, lossless coding, etc.)
[0315] B. When the horizontal length or the vertical length of the luma block of the coding unit exceeds the size of the transformable maximum luma transform, for example, when the size of the luma block of the coding block is 128x16 in the case where the size of the transformable maximum luma transform is 64, LFNST is not applicable.
[0316] It is also possible to determine whether to signal the LFNST index based on the size of the partition block rather than the coding unit. That is, when the horizontal length or the vertical length of the corresponding luma block of the partition block exceeds the size of the transformable maximum luma transform, the LFNST index signaling can be skipped, and the LFNST index can be inferred to be 0.
[0317] In the dual tree case, it is determined whether each of the coding unit or the partition block of the luma component and the coding unit or the partition block of the chroma component exceeds the maximum transform block size. That is, each of the horizontal and vertical lengths of the coding unit or the partition block of the luma is compared with the maximum luma transform size, and if any of them is greater than the maximum luma transform size, the LFNST is not applied, and in the case of the coding unit or the partition block of the chroma, the horizontal / vertical length of the corresponding luma block for the color format is compared with the transformable maximum luma transform size. For example, when the color format is 4:2:0, each of the horizontal / vertical length of the corresponding luma block is twice the horizontal / vertical length of the chroma block, and the transform size of the corresponding luma block is twice the transform size of the chroma block. As another example, when the color format is 4:4:4, the horizontal / vertical length of the corresponding luma block is the same as the horizontal / vertical length of the chroma block.
[0318] In the single tree case, it is checked whether the luma block (coding unit or partition block) exceeds the size of the transformable maximum luma transform block, and if it exceeds the size, the LFNST index signaling can be skipped.
[0319] C. If the LFNST included in the current VVC standard is applied, the LFNST index can be transmitted only when both the horizontal length and the vertical length of the partition block are greater than or equal to 4.
[0320] If up to LFNST for 2xM (1xM) or Mx2 (Mx1) blocks in addition to the LFNST included in the current VVC standard is applied, the LFNST index is transmitted only when the partition block has a size greater than or equal to the size of the 2xM (1xM) or Mx2 (Mx1) block. Here, PxQ block is greater than or equal to RxS block means P≥R and Q≥S.
[0321] In summary, the LFNST index can be transmitted only when the partition block has a size greater than or equal to the minimum size applicable to the LFNST. In the dual tree case, the LFNST index can be signaled only when the partition block of the luma or chroma component has a size greater than or equal to the minimum size applicable to the LFNST. In the single tree case, the LFNST index can be signaled only when the partition block of the luma component has a size greater than or equal to the minimum size applicable to the LFNST.
[0322] In this document, an MxN block equal to or greater than a KxL block means that M is equal to or greater than K and N is equal to or greater than L. An MxN block greater than a KxL block means that M is equal to or greater than K and N is equal to or greater than L, in which M is greater than K or N is greater than L. An MxN block less than or equal to a KxL block means that M is less than or equal to K and N is less than or equal to L. An MxN block less than a KxL block means that M is less than or equal to K and N is less than or equal to L, in which M is less than K or N is less than L.
[0323] D. In the case where the last non-zero coefficient position is not the DC position (the top-left position of the block), when it is a dual tree type luma block, if the corresponding last non-zero coefficient position is not the DC position in even any one of all the sub-blocks, the LFNST signaling can be performed. When it is a dual tree type and chroma block, if even any one of the last non-zero coefficient position of all the sub-blocks of Cb (here, when the ISP mode is not applied to the chroma component, the number of sub-blocks is considered to be 1) and the last non-zero coefficient position of all the sub-blocks of Cr (here, when the ISP mode is not applied to the chroma component, the number of sub-blocks is considered to be 1) is not the DC position, the corresponding LNFST index can be signaled.
[0324] In the case of a single tree type, if the last non-zero position is not the DC position even in any one of all the sub-blocks of the luma component, the Cb component, and the Cr component, the corresponding LNFST index can be signaled.
[0325] Here, if the coded block flag (CBF) value indicating whether there is a transform coefficient for each sub-block is 0, the last non-zero coefficient position of the corresponding sub-block is not checked to determine whether to signal the LFNST index. That is, if the CBF value is 0, since a transform is not applied to the block, the last non-zero coefficient position is not considered when checking the condition for LFNST index signaling.
[0326] For example, 1) in the case of a dual tree type and a luma component, when determining whether to signal the LFNST index for each sub-block, if the corresponding CBF value is 0, 2) in the case of a dual tree type and a chroma component, if the CBF value of Cb is 0 and the CBF value of Cr is 1, whether to signal the LFNST index can be determined by checking only the non-zero coefficient position of Cr of each sub-block, and 3) in the case of a single tree type, whether to signal the LFNST index can be determined by checking the last non-zero coefficient position only for blocks in which the CBF value is 1 among all the sub-blocks of the luma component, the Cb component, and the Cr component.
[0327] In the case of the ISP mode, the image information can be configured not to check the last non-zero coefficient position, and an implementation thereof can be as follows.
[0328] i. In the case of ISP mode, LFNST index signaling can be allowed by skipping the check of the non-zero coefficient position for both luma and chroma blocks. That is, even if the last non-zero coefficient position is a DC position or the corresponding CBF value is 0 for all partition blocks, LFNST index signaling can be allowed.
[0329] ii. In the case of ISP mode, based on the foregoing scheme, the check of the non-zero coefficient position can be skipped only for luma blocks, and the non-zero coefficient position can be checked for chroma blocks. For example, in the case of dual tree type and luma blocks, LFNST index signaling is allowed without checking the non-zero coefficient position, and in the case of dual tree type and chroma blocks, it can be determined whether to signal the LFNST index by checking whether there is a DC position for the last non-zero coefficient position through the foregoing scheme.
[0330] iii. In the case of ISP mode and single tree type, scheme i or ii can be applied. That is, in the case of ISP mode and single tree type, when scheme i is applied, LFNST index signaling can be allowed by skipping the check of the last non-zero coefficient position for both luma and chroma blocks. Alternatively, scheme ii can be applied to determine to signal the LFNST index in such a way that the check for the non-zero coefficient position is skipped for the partition blocks of the luma component and performed for the partition blocks of the chroma component according to the foregoing scheme (when ISP is not applied to the chroma component, the number of partition blocks can be considered as 1).
[0331] E. When it is confirmed that transform coefficients exist in positions other than positions in which LFNST transform coefficients can exist even for one partition block among all partition blocks, LFNST index signaling can be omitted.
[0332] For example, in the case of 4x4 partition blocks and 8x8 partition blocks, LFNST transform coefficients can exist at 8 positions from the DC position according to the transform coefficient scan order in the VVC standard, and all the remaining positions are padded with 0. In addition, when the transform block is greater than or equal to the 4x4 size and is not a 4x4 partition block and an 8x8 partition block, LFNST transform coefficients can exist at 16 positions from the DC position according to the transform coefficient scan order in the VVC standard, and all the remaining positions are padded with 0.
[0333] Therefore, if non-zero transform coefficients exist in a region to be padded with 0 after performing residual encoding, LFNST index signaling can be omitted.
[0334] Also, in the case of the ISP mode, instead of DCT-2, DST-7 is applied without signaling for the MTS index in the current VVC standard by independently considering the length condition for each of the horizontal direction and the vertical direction. It is determined whether the horizontal or vertical length is greater than or equal to 4 or greater than or equal to 16, and a primary transform kernel is determined according to the determined result. Accordingly, when in the ISP mode and when LFNST is applicable, the transform combination can be configured as follows.
[0335] 1. In the case of the LFNST index being 0 (including the case where the LFNST index is inferred to be 0), when ISP is included in the current VVC standard, it can comply with the condition for determining a primary transform. That is, it is checked whether the length condition (the condition that the length is greater than or equal to 4 and less than or equal to 16) is independently satisfied for each of the horizontal direction and the vertical direction, and if the condition is satisfied, DST-7 is applied for the primary transform instead of DCT-2, and if the condition is not satisfied, DCT-2 can be applied.
[0336] 2. In the case of the LFNST index being greater than 0, the following two configurations are possible for the primary transform.
[0337] A. DCT-2 is applied to both the horizontal direction and the vertical direction.
[0338] B. When ISP is included in the current VVC standard, it can comply with the condition for determining a primary transform. That is, it is checked whether the length condition (the condition that the length is greater than or equal to 4 and less than or equal to 16) is independently satisfied for each of the horizontal direction and the vertical direction, and if the condition is satisfied, DST-7 is applied instead of DCT-2, and if the condition is not satisfied, DCT-2 can be applied.
[0339] In the ISP mode, the image information can be configured such that the LFNST index is not transmitted per coding unit but per partition block. In this case, it can be determined whether to signal the LFNST index by considering that only one partition block exists in the unit in which the LFNST index is transmitted in the aforementioned LFNST index signaling scheme.
[0340] Also, the signaling order of the LFNST index and the MTS index will be described below.
[0341] According to an example, the LFNST index signaled in residual coding can be coded after the coded position of the last non-zero coefficient position, and the MTS index can be coded immediately after the LFNST index. In case of this configuration, the LFNST index can be signaled for each transform unit. Alternatively, the LFNST index can be coded after the coding of the last significant coefficient position even if not signaled in residual coding, and the MTS index can be coded after the LFNST index.
[0342] The syntax of residual coding according to an example is as follows.
[0343] [Table 4]
[0344]
[0345]
[0346] The meanings of the main variables shown in Table 4 are as follows.
[0347] 1. cbWidth, cbHeight: width and height of the current coding block
[0348] 2. log2TbWidth, log2TbHeight: log2 value of the width and height of the current transform block, which can be reduced by reflecting zero-out to the top-left region where non-zero coefficients can exist.
[0349] 3. sps_lfnst_enabled_flag: flag indicating whether LFNST is enabled, if the flag value is 0, it indicates that LFNST is not enabled, and if the flag value is 1, it indicates that LFNST is enabled. It is defined in the sequence parameter set (SPS).
[0350] 4. CuPredMode[chType][x0][y0]: prediction mode corresponding to the variable chType and (x0, y0) position, chType can have values of 0 and 1, where 0 indicates a luma component and 1 indicates a chroma component. The (x0, y0) position indicates a position on a picture, and MODE_INTRA (intra prediction) and MODE_INTER (inter prediction) can be values of CuPredMode[chType][x0][y0].
[0351] 5. IntraSubPartitionsSplit[x0][y0]: the content of (x0, y0) position is the same as in clause 4. It indicates which ISP partition is applied at the (x0, y0) position, ISP_NO_SPLIT indicates that the coding unit corresponding to the (x0, y0) position is not divided into partition blocks.
[0352] 6. intra_mip_flag[ x0 ][ y0 ] : The content of (x0, y0) is the same as in clause 4 above. The intra_mip_flag is a flag indicating whether the matrix-based intra prediction (MIP) prediction mode is applied or not. If the flag value is 0, it indicates that MIP is not enabled, and if the flag value is 1, it indicates that MIP is enabled.
[0353] 7. cIdx : The value 0 indicates luma, and the values 1 and 2 indicate Cb and Cr, respectively, which are chroma components.
[0354] 8. treeType : Indicates single tree and dual tree, etc. (SINGLE_TREE : single tree, DUAL_TREE_LUMA : dual tree for luma component, DUAL_TREE_CHROMA : dual tree for chroma component)
[0355] 9. tu_cbf_cb[ x0 ][ y0 ] : The content of (x0, y0) is the same as in clause 4. It indicates a coding block flag (CBF) of the Cb component. If its value is 0, it means that there is no non-zero coefficient in the corresponding transform unit of the Cb component, and if its value is 1, it indicates that there is a non-zero coefficient in the corresponding transform unit of the Cb component.
[0356] 10. lastSubBlock : It indicates the position of the subblock (coefficient group (CG)) in which the last non-zero coefficient is located in the scan order. 0 indicates a subblock containing a DC component, and in the case of a value greater than 0, it is not a subblock containing a DC component.
[0357] 11. lastScanPos : It indicates the position of the last significant coefficient within a subblock in the scan order. If a subblock includes 16 positions, there can be values from 0 to 15.
[0358] 12. lfnst_idx[ x0 ][ y0 ] : LFNST index syntax element to be parsed. If not parsed, it is inferred to have a value of 0. That is, the default value is set to 0, indicating that LFNST is not applied.
[0359] 13. LastSignificantCoeffX, LastSignificantCoeffY : They indicate the x and y coordinates in which the last significant coefficient is located in the transform block. The x coordinate starts from 0 and increases from left to right, and the y coordinate starts from 0 and increases from top to bottom. If the values of both variables are 0, it means that the last significant coefficient is located at the DC.
[0360] 14. cu_sbt_flag: a flag indicating whether sub-block transform (SBT) included in the current VVC standard is enabled. If the flag value is 0, it indicates that SBT is not enabled, and if the flag value is 1, it indicates that SBT is enabled.
[0361] 15. sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: flags indicating whether explicit MTS is applied to inter CUs and intra CUs, respectively. If the corresponding flag value is 0, it indicates that MTS is not enabled for inter CUs or intra CUs, and if the corresponding flag value is 1, it indicates that MTS is enabled.
[0362] 16. tu_mts_idx[x0][y0]: an MTS index syntax element to be parsed. If not parsed, it is inferred to have a value of 0. That is, the default value is set to 0, indicating that DCT-2 is enabled in both the horizontal direction and the vertical direction.
[0363] As shown in Table 4, in the case of a single tree, it can be determined whether to signal the LFNST index using only the last significant coefficient position condition for luminance. That is, if the position of the last significant coefficient is not DC and the last significant coefficient exists in the top-left sub-block (CG) (e.g., a 4x4 block), the LFNST index is signaled. In this case, in the case of a 4x4 transform block and an 8x8 transform block, the LFNST index is signaled only when the last significant coefficient exists in positions 0 to 7 in the top-left sub-block.
[0364] In the case of a dual tree, the LFNST index is signaled independently for each of luminance and chrominance, and in the case of chrominance, the LFNST index can be signaled by applying the last significant coefficient position condition only to the Cb component. For the Cr component, the corresponding condition can not be checked, and if the Cb CBF value is 0, the LFNST index can be signaled by applying the last significant coefficient position condition to the Cr component.
[0365] "Min(log2TbWidth, log2TbHeight) >= 2" of Table 4 can be expressed as "Min(tbWidth, tbHeight) >= 4", and "Min(log2TbWidth, log2TbHeight) >= 4" can be expressed as "Min(tbWidth, tbHeight) >= 16".
[0366] In Table 4, log2ZoTbWidth and log2ZoTbHeight mean the log value of the width and height of the top-left region in which the last significant coefficient can exist by zero-out, base-2.
[0367] As shown in Table 4, the log2ZoTbWidth and log2ZoTbHeight values can be updated in two places. The first is before parsing the MTS index or LFNST index value, and the second is after parsing the MTS index.
[0368] The first update is before parsing the MTS index (tu_mts_idx[x0][y0]) value, so log2ZoTbWidth and log2ZoTbHeight can be set regardless of the MTS index value.
[0369] After parsing the MTS index, log2ZoTbWidth and log2ZoTbHeigh are set for the MTS index greater than 0 (DST-7 / DCT-8 combination). When DST-7 / DCT-8 is independently applied to each of the horizontal direction and the vertical direction in one transform, there can be up to 16 significant coefficients per row or column in each direction. That is, up to 16 transform coefficients can be derived for each row or column starting from the left or top after applying DST-7 / DCT-8 of length 32 or more. Accordingly, in a 2D block, when DST-7 / DCT-8 is applied to both the horizontal direction and the vertical direction, significant coefficients can exist in only up to a 16x16 top-left region.
[0370] In addition, when DCT-2 is independently applied to each of the horizontal direction and the vertical direction in the current one transform, there can be up to 32 significant coefficients per row or column in each direction. That is, up to 32 transform coefficients can be derived for each row or column starting from the left or top when DCT-2 of length 64 or more is applied. Accordingly, in a 2D block, when DCT-2 is applied to both the horizontal direction and the vertical direction, significant coefficients can exist in only up to a 32x32 top-left region.
[0371] In addition, when DST-7 / DCT-8 is applied on one side and DCT-2 is applied on the other side for the horizontal direction and the vertical direction, there can be 16 significant coefficients in the former direction and 32 significant coefficients in the latter direction. For example, in the case of a 64x8 transform block, if DCT-2 is applied in the horizontal direction and DST-7 is applied in the vertical direction (which can occur in the case of applying implicit MTS), significant coefficients can exist in up to a 32x8 top-left region.
[0372] If log2ZoTbWidth and log2ZoTbHeight are updated in two places as shown in Table 4, that is, before parsing the MTS index, the range of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be determined by log2ZoTbWidth and log2ZoTbHeight as shown in the following table.
[0373] [Table 5]
[0374]
[0375] Additionally, in this case, the maximum value of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be set by reflecting the log2ZoTbWidth and log2ZoTbHeight values in the binarization process of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix.
[0376] [Table 6]
[0377]
[0378] According to an example, when the signaling of Table 4 is applied in the case of applying ISP mode and LFNST, the specification text can be configured as shown in Table 7. Compared with Table 4, the condition of signaling the LFNST index only in the case of not including the ISP mode (IntraSubPartitionsSplit[x0][y0]==ISP_NO_SPLIT in Table 4) is deleted.
[0379] In single tree, when the LFNST index transmitted for the luma component (cIdx = 0) is reused for the chroma component, the LFNST index transmitted for the first ISP sub-block in which there is a significant coefficient can be applied to the chroma transform block. Alternatively, even in single tree, the LFNST index for the chroma component can be signaled separately from the LFNST index for the luma component. The description of the variables in Table 7 is the same as in Table 4.
[0380] [Table 7]
[0381]
[0382] According to an example, the LFNST index and / or the MTS index can be signaled at the coding unit level. As described above, the LFNST index can have three values 0, 1, and 2, where 0 indicates that no LFNST is applied, and 1 and 2 indicate that the first and second candidates of two LFNST core candidates are included in the selected LFNST set, respectively. The LFNST index is coded by truncated unary binarization, and the values 0, 1, 2 can be coded as bin strings of 0, 10, 11, respectively.
[0383] According to an example, the LFNST can be applied only when DCT-2 is applied in one transform in both the horizontal direction and the vertical direction. Thus, if the MTS index is signaled after the LFNST index is signaled, the MTS index can be signaled only when the LFNST index is 0, and one transform can be performed by applying DCT-2 in both the horizontal direction and the vertical direction without signaling the MTS index when the LFNST index is not 0.
[0384] The MTS index can have values 0, 1, 2, 3, and 4, where 0, 1, 2, 3, and 4 can indicate DCT-2 / DCT-2, DST-7 / DST-7, DCT-8 / DST-7, DST-7 / DCT-8, DCT-8 / DCT-8 applied to the horizontal direction and the vertical direction, respectively. In addition, the MTS index can be coded by truncated unary binarization, and the values 0, 1, 2, 3, 4 can be coded as bin strings of 0, 10, 110, 1110, 1111, respectively.
[0385] The LFNST index and the MTS index can be signaled at the coding unit level, and the MTS index can be sequentially coded after the LFNST index at the coding unit level. The coding unit syntax table for this is as follows.
[0386] [Table 8]
[0387]
[0388] The variable LfnstDcOnly and the variable LfnstZeroOutSigCoeffFlag of Table 8 can be set as shown in Table 9 below.
[0389] The variable LfnstDcOnly is equal to 1 when for a transform block with a coding block flag (CBF) equal to 1 (equal to 0 if there is at least one significant coefficient in the corresponding block, otherwise equal to 0) all the last significant coefficients are located at the DC position (top-left position), otherwise it is equal to 0. Specifically, in the case of dual tree luma, the position of the last significant coefficient is checked for a luma transform block and in the case of dual tree chroma, the position of the last significant coefficient is checked for both the transform block of Cb and the transform block of Cr. In the case of single tree, the position of the last significant coefficient can be checked for the transform blocks of luma, Cb and Cr.
[0390] The variable LfnstZeroOutSigCoeffFlag is equal to 0 if there is a significant coefficient at the zero-out position when LFNST is applied, otherwise it is equal to 1.
[0391] The lfnst_idx[x0][y0] included in Table 8 and subsequent tables indicates the LFNST index of the corresponding coding unit, while the tu_mts_idx[x0][y0] indicates the MTS index of the corresponding coding unit.
[0392] As shown in Table 8, the conditions for signaling lfnst_idx[x0][y0] can include a condition for checking whether the value of transform_skip_flag[x0][y0] is 0 (!transform_skip_flag[x0][y0]). In this case, the condition for checking whether the value of the existing tu_mts_idx[x0][y0] is 0 (i.e., checking whether DCT-2 is performed in both the horizontal and vertical directions) can be omitted.
[0393] The transform_skip_flag[x0][y0] indicates whether the coding unit is coded in a transform skip mode in which the transform is skipped, and this flag is signaled before the MTS index and the LFNST index. That is, since lfnst_idx[x0][y0] is signaled before the value of tu_mtx_idx[x0][y0] is signaled, it is possible to check only the condition regarding the value of transform_skip_flag[x0][y0].
[0394] As shown in Table 8, a plurality of conditions are checked when coding tu_mts_idx[x0][y0], and as described above, tu_mts_idx[x0][y0] is signaled only when the value of lfnst_idx[x0][y0] is 0.
[0395] tu_cbf_luma[x0][y0] is a flag indicating whether there is a valid coefficient for the luma component, and cbWidth and cbHeight indicate the width and height of the coding unit of the luma component, respectively.
[0396] According to Table 8, when both the width and height of the coding unit of the luma component are 32 or less, tu_mts_idx[x0][y0] is signaled, that is, whether to apply MTS is determined by the width and height of the coding unit of the luma component.
[0397] According to another example, when transform block (TU) tiling occurs (for example, when the maximum transform size is set to 32, a coding unit of 64x64 is divided into 4 transform blocks of 32x32 and coded), the MTS index can be signaled based on the size of each transform block. For example, when both the width and height of the transform block are 32 or less, the same MTS index value can be applied to all transform blocks in the coding unit, thereby applying the same primary transform. In addition, when transform block tiling occurs, the value of tu_cbf_luma[x0][y0] in Table 8 can be the CBF value of the top-left transform block, or can be set to 1 even when the CBF value of one of all transform blocks is 1.
[0398] As shown in Table 8, even in the ISP mode (IntraSubPartitionsSplitType!= ISP_NO_SPLIT), lfnst_idx[x0][y0] can be configured to be signaled, and the same LFNST index value can be applied to all ISP partition blocks.
[0399] In addition, tu_mts_idx[x0][y0] can be signaled only in the mode other than the ISP mode (IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT).
[0400] As shown in Table 8, when the MTS index is signaled immediately after the LFNST index, information about the primary transform cannot be known when residual coding is performed. That is, the MTS index is signaled after residual coding. Therefore, in the residual coding section, the section in which zeroing is performed while only 16 coefficients are reserved for DST-7 or DCT-8 of length 32 can be changed as shown in Table 9 below.
[0401] [Table 9]
[0402]
[0403]
[0404] As shown in Table 9, in determining log2ZoTbWidth and log2ZoTbHeight (where log2ZoTbWidth and log2ZoTbHeight denote the values of the base-2 logarithm of the width and height of the top-left region left after performing zero-out, respectively), the checking of the value of tu_mts_idx[x0][y0] can be omitted.
[0405] The binarization of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix in Table 9 can be determined based on log2ZoTbWidth and log2ZoTbHeight as shown in Table 6.
[0406] In addition, as shown in Table 9, when log2ZoTbWidth and log2ZoTbHeight are determined in residual coding, a condition of checking sps_mts_enable_flag can be added.
[0407] According to an example, when position information about the last significant coefficient of a luma transform block is recorded in a residual coding process, the MTS index can be signaled as shown in Table 10.
[0408] [Table 10]
[0409]
[0410] In Table 10, LumaLastSignificantCoeffX and LumaLastSignificantCoeffY indicate the X coordinate and Y coordinate of the last significant coefficient position of a luma transform block, respectively. In Table 10, a condition that both LumaLastSignificantCoeffX and LumaLastSignificantCoeffY must be less than 16 is added. When any one of them is 16 or more, DCT-2 is applied in both the horizontal direction and the vertical direction, it can be inferred that the signaling of tu_mts_idx[x0][y0] is omitted, and DCT-2 is applied in both the horizontal direction and the vertical direction.
[0411] When both LumaLastSignificantCoeffX and LumaLastSignificantCoeffY are smaller than 16, it means that the last significant coefficient exists in the top-left 16x16 region. In the current VVC standard, when DST-7 or DCT-8 with length 32 is applied, it indicates the possibility that the zero-out which keeps only 16 transform coefficients from left or above has been applied. Therefore, the transform kernel for one transform can be specified by signaling tu_mts_idx[x0][y0].
[0412] In addition, according to another example, the coding unit syntax table, the transform unit syntax table, and the residual coding syntax table are as follows. According to Table 11, the MTS index is moved from the transform unit level to the coding unit level syntax and signaled after the LFNST index signaling. In addition, the following constraint has been removed: LFNST is not allowed when ISP is applied to a coding unit. When ISP is applied to a coding unit, the constraint that LFNST is not allowed is removed, so that LFNST can be applied to all intra prediction blocks. In addition, both the MTS index and the LFNST index are conditionally signaled in the last part of the coding unit level.
[0413] [Table 11]
[0414]
[0415] [Table 12]
[0416]
[0417] [Table 13]
[0418]
[0419] In Table 11, MtsZeroOutSigCoeffFlag is initially set to 1, and the value can be changed in residual coding in Table 13. When there is a significant coefficient in the region to be padded with 0 by zero-out (LastSignificantCoeffX>15||LastSignificantCoeffY>15), the value of the variable MtsZeroOutSigCoeffFlag changes from 1 to 0, in which case the MTS index is not signaled, as shown in Table 11.
[0420] In addition, as shown in Table 11, when tu_cbf_luma[x0][y0] is 0, mts_idx[x0][y0] coding can be omitted. That is, when the CBF value of the luma component is 0, since no transform is applied, there is no need to signal the MTS index, so the MTS index coding can be omitted.
[0421] According to an example, the above technical features can be implemented with another conditional syntax. For example, after performing the MTS, a variable indicating whether there is a significant coefficient in a region other than the DC region of the current block can be derived, and when the variable indicates that there is a significant coefficient in the region other than the DC region, the MTS index can be signaled. That is, the presence of a significant coefficient in the region other than the DC region of the current block indicates that the value of tu_cbf_luma[x0][y0] is 1, and in this case, the MTS index can be signaled.
[0422] The variable can be denoted as MtsDcOnly, and after the variable MtsDcOnly is initially set to 1 at the coding unit level, the value is changed to 0 when it is determined that there is a significant coefficient in the region other than the DC region of the current block in the residual coding level. When the variable MtsDcOnly is 0, the picture information can be configured such that the MTS index is signaled.
[0423] When tu_cbf_luma[x0][y0] is 0, since the residual coding syntax is not invoked at the transform unit level of Table 12, the initial value 1 of the variable MtsDcOnly is maintained. In this case, since the variable MtsDcOnly is not changed to 0, the picture information can be configured not to signal the MTS index. That is, the MTS index is not parsed and signaled.
[0424] In addition, the decoding device can determine the color index cIdx of the transform coefficient to derive the variable MtsZeroOutSigCoeffFlag of Table 13. The color index cIdx of 0 indicates the luminance component.
[0425] According to an example, since the MTS can be applied only to the luminance component of the current block, the decoding device can determine whether the color index is luminance when deriving the variable MtsZeroOutSigCoeffFlag for determining whether to parse the MTS index.
[0426] The variable MtsZeroOutSigCoeffFlag is a variable indicating whether zero-out is performed when MTS is applied. It indicates whether there are transform coefficients in a region other than the top-left region of the last significant coefficient due to zero-out after MTS is performed (i.e., a region other than the top-left 16x16 region). The variable MtsZeroOutSigCoeffFlag is initially set to 1 at the coding unit level, as shown in Table 11 (MtsZeroOutSigCoeffFlag = 1), and its value can be changed from 1 to 0 at the residual coding level when there are transform coefficients in a region other than the 16x16 region, as shown in Table 13 (MtsZeroOutSigCoeffFlag = 0). When the value of the variable MtsZeroOutSigCoeffFlag is 0, the MTS index is not signaled.
[0427] As shown in Table 13, at the residual coding level, the non-zero-out region in which non-zero transform coefficients can exist can be set depending on whether zero-out accompanying MTS is performed, and even in this case, the color index (cIdx) is 0, the non-zero-out region can be set to the top-left 16x16 region of the current block.
[0428] As such, in deriving the variable determining whether to parse the MTS index, it is determined whether the color component is a luma or a chroma. However, since LFNST can be applied to both the luma component and the chroma component of the current block, the color component is not determined in deriving the variable for determining whether to parse the LFNST index.
[0429] For example, Table 11 shows the variable LfnstZeroOutSigCoeffFlag, which can indicate whether zero-out is performed when LFNST is applied. The variable LfnstZeroOutSigCoeffFlag indicates whether there are significant coefficients in a second region of the current block other than a first region at the top left. The value is initially set to 1, and when there are significant coefficients in the second region, the value can be changed to 0. The LFNST index can be parsed only when the value of the variable LfnstZeroOutSigCoeffFlag initially set remains 1. In determining and deriving whether the value of the variable LfnstZeroOutSigCoeffFlag is 1, since LFNST can be applied to both the luma component and the chroma component of the current block, the color index of the current block is not determined.
[0430] Figure 15 CCLM applicable when deriving an intra prediction mode of a chroma block according to an embodiment is exemplified.
[0431] In the present specification, a "reference sample template" can refer to a set of neighboring reference samples of a current chroma block used for predicting the current chroma block. The reference sample template can be predefined, and information about the reference sample template can be signaled from the encoding apparatus 200 to the decoding apparatus 300.
[0432] Referring to Figure 15 A set of shaded samples in a single line adjacent to a 4x4 block as the current chroma block refers to a reference sample template. The reference sample template is configured as a reference sample of a single line, and a reference sample region in a luminance region corresponding to the reference sample template is configured as two lines, as shown in Figure 15
[0433] In an embodiment, when intra coding of a chroma image is performed in a joint exploration test model (JEM) used in a joint video exploration team (JVET), a cross-component linear model (CCLM) can be used. The CCLM is a method of predicting a pixel value of a chroma image from a pixel value of a reconstructed luminance image, and is based on a high correlation between a luminance image and a chroma image.
[0434] The CCLM prediction of Cb and Cr chroma images can be performed based on the following equation.
[0435] [Equation 11]
[0436] Pred C (i, j) = a · Rec′ L (i, j) + β
[0437] Here, Pred c (i, j) denotes a Cb or Cr chroma image to be predicted, Rec L ′(i, j) denotes a reconstructed luminance image adjusted to a chroma block size, and (i, j) denotes coordinates of a pixel. In a 4:2:0 color format, since the size of a luminance image is twice that of a chroma image, it is necessary to generate Rec L ′ having a chroma block size by downsampling, and thus a pixel of a luminance image to be used for the chroma image Pred L (i, j) can be considered taking into account Rec c (2i, 2j) and neighboring pixels. Rec L ′(i, j) can be referred to as a downsampled luminance sample.
[0438] For example, Rec L ′(i, j) can be derived using six neighboring pixels as shown in the following equation.
[0439] [Equation 12]
[0440] Rec′ L (x, y) = (2 x Rec L (x, y) = (2 x Rec L (x, y) = (2 x Rec L (x, y) = (2 x Rec L (x, y) = (2 x Rec L (x, y) = (2 x Rec L (x, y) = (2 x Rec
[0441] and β represent the cross-correlation and average difference between the neighboring templates of a Cb or Cr chroma block and predModeIntra For example, α and β are represented by Equation 13.
[0442] [Equation 13]
[0443]
[0444]
[0445] L(n) represents neighboring reference samples and / or left neighboring samples of a luma block corresponding to a current chroma image, C(n) represents neighboring reference samples and / or left neighboring samples of a current chroma block to which encoding is currently applied, and (i,j) represents a pixel position. In addition, L(n) can represent upper neighboring samples and / or left neighboring samples of a down-sampled current luma block. N can represent the total number of pixel pair (luma and chroma) values used to calculate CCLM parameters, and can indicate a value that is twice the smaller value of the width and height of the current chroma block.
[0446] A picture can be divided into a sequence of coding tree units (CTU). The CTU can correspond to a coding tree block (CTB). Alternatively, the CTU can include a coding tree block of luma samples and a corresponding coding tree block of chroma samples. Depending on whether the luma block and the corresponding chroma block have separate partitioning structures, the tree type can be classified as SINGLE_TREE or DUAL_TREE. SINGLE_TREE can indicate that the chroma block has the same partitioning structure as the luma block, while DUAL_TREE can indicate that the chroma component block has a different partitioning structure from the partitioning structure of the luma block.
[0447] When applying LFNST to a chroma transform block according to an example, information about a collocated luma transform block is needed.
[0448] The existing specification text regarding the concerned section is shown in the following table.
[0449] [Table 14]
[0450]
[0451]
[0452] As shown in Table 14, when the current intra prediction mode is the CCLM mode, the value of the variable predModeIntra for the chroma transform block is determined by taking the intra prediction mode value of the co-located chroma transform block (the portion indicated in italics). The intra prediction mode value (the predModeIntra value) of the luma transform block can then be used to determine the LFNST set.
[0453] However, the variables nTbW and nTbH that are input as input values of this transform process represent the width and height of the current transform block. Thus, when the current block is a luma transform block, the variables nTbW and nTbH can represent the width and height of the luma transform block, while when the current block is a chroma transform block, the variables nTbW and nTbH represent the width and height of the chroma transform block.
[0454] Here, the variables nTbW and nTbH of the italicized portion of Table 14 represent the width and height of the chroma transform block that do not reflect the color format and thus do not accurately indicate the reference position of the luma transform block corresponding to the chroma transform block. Thus, the italicized portion of Table 14 can be modified as shown in the following table.
[0455] [Table 15]
[0456]
[0457]
[0458] As shown in Table 15, nTbW and nTbH are changed to (nTbW*SubWidthC) / 2 and (nTbH*SubHeightC) / 2, respectively. xTbY and yTbY can represent the luma position in the current picture (the top-left sample of the current luma transform block relative to the top-left luma sample of the current picture), while nTbW and nTbH can represent the width and height of the currently coded transform block (the variable nTbW specifies the width of the current transform block, and the variable nTbH specifies the height of the current transform block).
[0459] When the currently coded transform block is a chroma (Cb or Cr) transform block, nTbW and nTbH are the width and height of the chroma transform block, respectively. Thus, when the currently coded transform block is a chroma transform block (cldx > 0), the width and height of the luma transform block need to be used to obtain the reference location of the collocated luma transform block when obtaining the reference location. In Table 15, SubWidthC and SubHeightC are values set according to the color format (e.g., 4:2:0, 4:2:2 or 4:4:4), and in particular, the width and height ratios between the luma component and the chroma components (see Table 16 below). Thus, in the case of a chroma transform block, (nTbW * SubWidthC) and (nTbH * SubHeightC) can be the width and height of the collocated luma transform block, respectively.
[0460] Thus, xTbY + (nTbW * SubWidthC) / 2 and yTbY + (nTbH * SubHeightC) / 2 represent values of the central position in the collocated luma transform block based on the top-left position of the current picture, and thus accurately indicate the collocated luma transform block.
[0461] [Table 16]
[0462]
[0463] In Table 15, the variable predModeIntra represents the intra prediction mode value, and the value of the variable predModeIntra equal to INTRA LT CCLM, INTRA L CCLM or INTRA T CCLM indicates that the current transform block is a chroma transform block. According to an example, in the current VVC standard, INTRA LT CCLM, INTRA L CCLM and INTRA T CCLM correspond to mode values 81, 82 and 83 among the intra prediction mode values, respectively. Thus, as shown in Table 15, the values of xTbY + (nTbW * SubWidthC) / 2 and yTbY + (nTbH * SubHeightC) / 2 need to be used to obtain the reference location of the collocated luma transform block.
[0464] As shown in Table 15, the value of predModeIntra is updated in view of both the variable intra mip flag [xTbY + (nTbW * SubWidthC) / 2] [yTbY + (nTbH * SubHeightC) / 2] and the variable CuPredMode [0] [xTbY + (nTbW * SubWidthC) / 2] [yTbY + (nTbH * SubHeightC) / 2].
[0465] intra_mip_flag is a variable indicating whether the current transform block (or coding unit) is coded by a matrix-based intra prediction (MIP) method, and intra_mip_flag[ x ][ y ] is a flag value indicating whether MIP is applied to a position corresponding to coordinates (x, y) based on a luma component when a top-left position of a current picture is defined as (0, 0). The x and y coordinates are increased from left to right and from top to bottom, respectively, and when the flag indicating whether MIP is applied is 1, the flag indicates that MIP is applied. When the flag indicating whether MIP is applied is 0, the flag indicates that MIP is not applied. MIP can be applied only to a luma block.
[0466] According to a modified portion of Table 15, when the value of intra_mip_flag[ xTbY + (nTbW * SubWidthC) / 2 ][ yTbY + (nTbH * SubHeightC) / 2 ] in the collocated luma transform block is 1, the value of predModeIntra is set to the planar mode (INTRA_PLANAR).
[0467] The value of the variable CuPredMode[ 0 ][ xTbY + (nTbW * SubWidthC) / 2 ][ yTbY + (nTbH * SubHeightC) / 2 ] represents a prediction mode value corresponding to coordinates (xTbY + (nTbW * SubWidthC) / 2, yTbY + (nTbH * SubHeightC) / 2) when a top-left position of a current picture of a luma component is defined as (0, 0). The prediction mode value can have MODE_INTRA, MODE_IBC, MODE_PLT, and MODE_INTER values representing an intra prediction mode, an intra block copy (IBC) prediction mode, a palette (PLT) coding mode, and an inter prediction mode, respectively. According to Table 15, when the value of the variable CuPredMode[ 0 ][ xTbY + (nTbW * SubWidthC) / 2 ][ yTbY + (nTbH * SubHeightC) / 2 ] is MODE_IBC or MODE_PLT, the value of the variable predModeIntra is set to the DC mode. In other cases than these two cases, the value of the variable predModeIntra is set to IntraPredModeY[ xTbY + (nTbW * SubWidthC) / 2 ][ yTbY + (nTbH * SubHeightC) / 2 ] (an intra prediction mode value corresponding to a central position in the collocated luma transform block).
[0468] According to an example, the value of the variable predModeIntra can be updated one more time based on the updated predModeIntra value of Table 15, as shown in the following table, considering whether wide-angle intra prediction is performed.
[0469] [Table 17]
[0470]
[0471] The input values of predModeIntra, nTbW and nTbH in the mapping process shown in Table 17 are the same as the updated variable predModeIntra of Table 15 and the values of nTbW and nTbH referenced in Table 15, respectively.
[0472] In Table 17, nCbW and nCbH represent the width and height of a coding block corresponding to a transform block, respectively, and the variable IntraSubPartitionsSplitType indicates whether an ISP mode is applied, where IntraSubPartitionsSplitType equal to ISP_NO_SPLIT indicates that the coding unit is not partitioned by ISP (i.e., the ISP mode is not applied). The variable IntraSubPartitionsSplitType not equal to ISP_NO_SPLIT indicates that the ISP mode is applied, and thus the coding unit is partitioned into two or four sub-partition blocks. In Table 17, cIdx is an index indicating a color component. The cIdx value equal to 0 indicates a luma block, and the cIdx value not equal to 0 indicates a chroma block. The predModeIntra value output by the mapping process of Table 17 is an updated value considering whether a wide-angle intra prediction (WAIP) mode is applied.
[0473] For the predModeIntra value updated by Table 17, the LFNST set can be determined by the mapping relationship shown in the following table.
[0474] [Table 18]
[0475] lfnstTrSetIdx Figure 16 predModeIntra<0 1 0 <= predModeIntra <= 1 0 2 <= predModeIntra <= 12 1 13 <= predModeIntra <= 23 2 24 <= predModeIntra <= 44 3 45 <= predModeIntra <= 55 2 56 <= predModeIntra <= 80 1
[0476] In the above table, lfnstTrSetldx denotes an index indicating an LFNST set and has a value from 0 to 3, which indicates that a total of four LFNST sets are configured. Each LFNST set can include two transform kernels, i.e., an LFNST kernel (the transform kernel can be a 16x16 matrix or a 16x48 matrix depending on the area to which the LFNST is applied based on a forward direction), and the transform kernel to be applied among the two transform kernels can be specified through signaling of an LFNST index. In addition, whether to apply the LFNST can also be specified through the LFNST index. In the current VVC standard, the LFNST index can have values 0, 1, and 2, 0 indicating that the LFNST is not applied, and 1 and 2 indicating the two transform kernels, respectively.
[0477] The following drawings are provided to describe specific examples of the present disclosure. Since specific names of apparatuses or names of specific signals / messages / fields exemplified in the drawings are provided for examples, technical features of the present disclosure are not limited to the specific names used in the following drawings.
[0478] Figure 16 is a flowchart illustrating an operation of a video decoding apparatus according to an embodiment of the present disclosure.
[0479] Figure 4 to Figure 15 Each process disclosed in the Figure 3 to Figure 15 description. Accordingly, a description of specific details will be omitted or will be schematically made with respect to details of the Figure 17 description. Thus, a detailed description will not be made with respect to details overlapping with the
[0480] The decoding apparatus 300 according to the embodiment can obtain intra prediction mode information and an LFNST index from a bitstream (S1610).
[0481] The intra prediction mode information can include an intra prediction mode of a neighboring block (e.g., a left neighboring block and / or an above neighboring block) of the current block, and an MPM index indicating one of MPM candidates in an MPM list derived based on additional candidate modes or residual intra prediction mode information indicating one of residual intra prediction modes not included in the MPM candidates.
[0482] In addition, the intra mode information can include flag information sps_cclm_enabled_flag indicating whether CCLM is applied to the current block and information intra_chroma_pred_mode about an intra prediction mode of a chroma component.
[0483] The LFNST index information is received as syntax information, and the syntax information is received as a bin string of binarization containing 0 and 1.
[0484] The syntax element of the LFNST index according to the present embodiment can indicate whether to apply inverse LFNST or inverse non-separable transform and any one of the transform kernel matrices included in the transform set, and when the transform set includes two transform kernel matrices, the syntax element of the transform index can have three values.
[0485] That is, according to the embodiment, the value of the syntax element of the LFNST index can include 0 indicating that inverse LFNST is not applied to the target block, 1 indicating a first transform kernel matrix among the transform kernel matrices, and 2 indicating a second transform kernel matrix among the transform kernel matrices.
[0486] The decoding device 300 can decode information on quantized transform coefficients of the current block from the bitstream, and can derive quantized transform coefficients of the target block based on the information on the quantized transform coefficients of the current block. The information on the quantized transform coefficients of the target block can be included in a sequence parameter set (SPS) or a slice header, and can include at least one of information on whether to apply RST, information on a reduction factor, information on a minimum transform size for applying RST, information on a maximum transform size for applying RST, inverse RST size, and information on a transform index indicating any one of the transform kernel matrices included in the transform set.
[0487] The decoding device 300 can derive transform coefficients by dequantizing the residual information (i.e., quantized transform coefficients) on the current block, and can arrange the derived transform coefficients in a predetermined scan order.
[0488] Specifically, the derived transform coefficients can be arranged in units of 4x4 blocks according to the inverse diagonal scan order, and the transform coefficients in the 4x4 blocks can also be arranged according to the inverse diagonal scan order. That is, the dequantized transform coefficients can be arranged according to the inverse scan order applied in a video codec such as in VVC or HEVC.
[0489] The transform coefficients derived based on the residual information can be dequantized transform coefficients as described above, or can be quantized transform coefficients. That is, the transform coefficients can be any data for checking whether there is non-zero data in the current block, regardless of quantization.
[0490] The decoding device can derive the intra prediction mode of the chroma block to be the CCLM mode based on the intra prediction mode information (S1620).
[0491] For example, the decoding device can receive information on the intra prediction mode of the current chroma block through the bitstream, and can derive the intra prediction mode of the current chroma block to be the CCLM mode based on the intra prediction mode information.
[0492] The CCLM mode can include an upper-left CCLM mode, an upper CCLM mode, or a left CCLM mode.
[0493] As described above, the decoding device can derive the residual samples by applying the LFNST as a non-separable transform or the MTS as a separable transform, and can perform these transforms based on the LFNST index indicating the LFNST kernel (i.e., the LFNST matrix) and the MTS index indicating the MTS kernel, respectively.
[0494] For the LFNST, it is necessary to determine the LFNST set, and the LFNST set has a mapping relationship with the intra prediction mode of the current block.
[0495] The decoding device can update the intra prediction mode of the chroma block based on the intra prediction mode of the luma block corresponding to the chroma block, for the inverse LFNST of the chroma block (S1630).
[0496] According to an example, the updated intra prediction mode can be derived as the intra prediction mode corresponding to a specific position in the luma block, and the specific position can be set based on the color format of the chroma block.
[0497] The specific position can be a central position of the luma block, and can be represented as ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)).
[0498] In the central position, xTbY and yTbY represent the top-left coordinates of the luma block, i.e., the top-left position in the luma sample reference of the current transform block, nTbW and nTbH represent the width and height of the chroma block, and SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)) represents the central position of the luma transform block, and IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the intra prediction mode of the luma block at this position.
[0499] SubWidthC and SubHeightC can be derived as shown in Table 16. That is, when the color format is 4:2:0, SubWidthC and SubHeighC are 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0500] As shown in Table 15, in order to specify a specific position of a luma block corresponding to a chroma block regardless of a color format, the color format is reflected in a variable indicating the specific position.
[0501] According to an example, when an intra prediction mode of a luma block corresponding to the specific position is a matrix-based intra prediction (hereinafter, "MIP") mode, the decoding device can set an updated intra prediction mode to an intra planar mode.
[0502] The MIP mode can be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). When the MIP is applied to a current block, prediction samples of the current block can be derived i) using neighboring reference samples that have undergone an averaging process, ii) by performing a matrix vector multiplication process, and iii) by further performing a horizontal / vertical interpolation process.
[0503] Alternatively, according to an example, when the intra prediction mode corresponding to the specific position is an intra block copy (IBC) mode or a palette mode, the decoding device can set the updated intra prediction mode to an intra DC mode.
[0504] The IBC prediction mode or the palette mode can be used to encode a content image / video including a game, such as screen content coding (SCC). The IBC basically performs prediction within a current picture, but can be performed similarly to inter prediction, with the difference being that a reference block is derived within the current picture. That is, the IBC can use at least one of the inter prediction techniques described in the present disclosure. The palette mode can be considered as an example of intra coding or intra prediction. When the palette mode is applied, values of samples in a picture can be signaled based on information about a palette table and a palette index.
[0505] As described above, when the intra prediction mode of the central position is the MIP mode, the IBC mode, or the palette mode, the intra prediction mode of the chroma block can be updated to a specific mode such as the intra planar mode or the intra DC mode.
[0506] When the intra prediction mode of the central position is not the MIP mode, the IBC mode, and the palette mode, the intra prediction mode of the chroma block can be updated to the intra prediction mode of the luma block for the central position in order to reflect the association between the chroma block and the luma block.
[0507] When the chroma block is not square, the decoding device can remap the updated intra prediction mode to a wide-angle intra prediction mode (S1640).
[0508] In order to determine the LFNST set, the intra prediction mode can be remapped (that is, updated or modified) by reflecting the wide-angle intra mode, as shown in Table 17.
[0509] The input values of predModeIntra, nTbW and nTbH of the mapping process presented in Table 17 are the updated predModeIntra in Table 15 and the values of nTbW and nTbH corresponding to the width and height of the chroma block referred in Table 15, respectively.
[0510] A variable whRatio representing the ratio of the width and height of the block can be set as Abs(Log2(nW / nH)). Alternatively, it can be expressed as Abs(Log2(nW)-Log2(nH)).
[0511] For example, when the width of the chroma block is greater than the height, the updated intra mode is 2 or greater, and the updated intra mode is a variable (whRatio>1)?(8+2*whRatio):8 [where whRatio is Abs(Log2(nW / nH))], the updated intra mode can be remapped to “updated intra mode + 65”.
[0512] In this document, the “x?y:z” operator indicates that if x is TRUE, x becomes y, and if x is otherwise, x equals z (if x is TRUE, the value of y is evaluated; otherwise, the value of z is evaluated).
[0513] Alternatively, when the height of the chroma block is greater than the width, the updated intra mode is 66 or less, and the updated intra mode is a variable (whRatio>1)?(60-2*whRatio):60 [where whRatio is Abs(Log2(nW / nH))], the updated intra mode can be remapped to “updated intra mode - 67”.
[0514] That is, the decoding device can re-update the intra prediction mode for determining the LFNST set by reflecting the non-square shape of the chroma block.
[0515] The decoding device can determine the LFNST set including the LFNST matrix based on the remapped intra prediction mode (S1650), and can derive the transform coefficient of the chroma block based on the LFNST matrix derived from the LFNST set (S1660).
[0516] Any one of a plurality of LFNST matrices can be selected based on the LFNST set and the LFNST index.
[0517] As shown in Table 18, the LFNST transform set is derived according to the intra prediction mode, and 81 to 83 indicating the CCLM mode among the intra prediction modes are omitted because the LFNST transform set is derived in the CCLM mode using the intra mode value of the corresponding luma block or the re-mapped wide-angle intra mode value.
[0518] According to an example, as shown in Table 18, any one of the four LFNST sets can be determined according to the intra prediction mode of the current block, and the LFNST set to be applied to the current chroma block can also be determined.
[0519] The decoding device can perform inverse RST (e.g., inverse LFNST) by applying the LFNST matrix to the dequantized transform coefficients, thereby deriving modified transform coefficients of the current chroma block.
[0520] The decoding device can derive residual samples from the transform coefficients by one inverse transform (S1670). The MTS can be used for the one inverse transform.
[0521] In addition, the decoding device can generate reconstructed samples based on the residual samples of the current block and the prediction samples of the current block. The current block can be the current luma block or the current chroma block.
[0522] The following drawings are provided to describe specific examples of the present disclosure. Since specific names of apparatuses exemplified in the drawings or names of specific signals / messages / fields are provided for illustration, the technical features of the present disclosure are not limited to the specific names used in the following drawings.
[0523] Figure 17 is a flowchart illustrating an operation of a video encoding apparatus according to an embodiment of the present disclosure.
[0524] Figure 4 to Figure 15 Each process disclosed in Figure 2 described above. Accordingly, the description of specific details overlapping with the details described above will be omitted or will be schematically made. Figure 4 to Figure 15 described above. Accordingly, the description of specific details overlapping with the details described above will be omitted or will be schematically made.
[0525] The encoding apparatus 200 according to an embodiment can derive the intra prediction mode of the chroma block as the CCLM mode (S1710).
[0526] For example, the encoding apparatus can determine the intra prediction mode of the current chroma block based on a rate-distortion (RD) cost (or RDO). Here, the RD cost can be derived based on a sum of absolute differences (SAD). The encoding apparatus can determine the CCLM mode as the intra prediction mode of the current chroma block based on the RD cost.
[0527] The CCLM mode can include an upper-left CCLM mode, an upper CCLM mode, or a left CCLM mode.
[0528] The encoding device can encode information on an intra prediction mode for the current chroma block, and can signal the information on the intra prediction mode through a bitstream. The prediction-related information for the current chroma block can include the information on the intra prediction mode.
[0529] The encoding device can derive prediction samples of the chroma block based on the CCLM mode (S1720).
[0530] According to an embodiment, the encoding device can derive residual samples of the chroma block based on the prediction samples (S1730).
[0531] According to an embodiment, the encoding device can derive transform coefficients of the chroma block based on a primary transform on the residual samples.
[0532] The primary transform can be performed through a plurality of transform kernels, in which case, the transform kernels can be selected based on the intra prediction mode.
[0533] The encoding device can update the intra prediction mode of the chroma block based on an intra prediction mode of a luma block corresponding to the chroma block for LFNST of the chroma block (S1740).
[0534] As shown in Table 15, the encoding device can update the CCLM mode of the chroma block based on the intra prediction mode of the luma block corresponding to the chroma block (when predModeIntra is equal to INTRA_LT_CCLM, INTRA_L_CCLM, or INTRA_T_CCLM, predModeIntra is derived as follows).
[0535] According to an example, the updated intra prediction mode can be derived from an intra prediction mode corresponding to a specific position in the luma block, and the specific position can be set based on a color format of the chroma block.
[0536] The specific position can be a central position of the luma block, and can be expressed as ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)).
[0537] In the central position, xTbY and yTbY represent the top-left coordinates of the luma block, i.e., the top-left position in the luma sample reference of the current transform block, nTbW and nTbH represent the width and height of the chroma block, and SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY + (nTbW * SubWidthC) / 2), (yTbY + (nTbH * SubHeightC) / 2)) represents the central position of the luma transform block, and IntraPredModeY[xTbY + (nTbW * SubWidthC) / 2][yTbY + (nTbH * SubHeightC) / 2] represents the intra prediction mode of the luma block at this position.
[0538] SubWidthC and SubHeightC can be derived as shown in Table 16. That is, SubWidthC and SubHeightC are 2 when the color format is 4:2:0, and SubWidthC is 2 and SubHeightC is 1 when the color format is 4:2:2.
[0539] As shown in Table 15, in order to specify a specific position of a luma block corresponding to a chroma block regardless of a color format, the color format is reflected in a variable indicating the specific position.
[0540] According to an example, when an intra prediction mode of a luma block corresponding to a specific position is a matrix-based intra prediction (hereinafter, "MIP") mode, an encoding device can set an updated intra prediction mode to an intra planar mode.
[0541] The MIP mode can be referred to as affine linear weighted intra prediction (ALWIP) or matrix weighted intra prediction (MWIP). When the MIP is applied to a current block, a prediction sample of the current block can be derived i) using neighboring reference samples that have undergone an averaging process, ii) by performing a matrix vector multiplication process, and iii) by further performing a horizontal / vertical interpolation process.
[0542] Alternatively, according to an example, when the intra prediction mode corresponding to a specific position is an intra block copy (IBC) mode or a palette mode, an encoding device can set an updated intra prediction mode to an intra DC mode.
[0543] The IBC prediction mode or the palette mode can be used for encoding a content image / video including a game, e.g., screen content coding (SCC). The IBC basically performs prediction within a current picture, but can be performed similarly to inter prediction, with the difference that the reference block is derived within the current picture. That is, the IBC can use at least one of the inter prediction techniques described in the disclosure. The palette mode can be considered as an example of intra coding or intra prediction. When the palette mode is applied, values of samples in a picture can be signaled based on information about a palette table and palette indices.
[0544] In summary, when the intra prediction mode of the central position is the MIP mode, the IBC mode, or the palette mode, the intra prediction mode of the chroma block can be updated to a specific mode such as the intra planar mode or the intra DC mode.
[0545] When the intra prediction mode of the central position is not the MIP mode, the IBC mode, and the palette mode, the intra prediction mode of the chroma block can be updated to the intra prediction mode of the luma block for the central position to reflect the correlation between the chroma block and the luma block.
[0546] When the chroma block is not square, the encoding device can remap the updated intra prediction mode to the wide-angle intra prediction mode (S1750).
[0547] To determine the LFNST set, the intra prediction mode can be remapped (i.e., updated or modified) by reflecting the wide-angle intra mode, as shown in Table 17.
[0548] The input values of predModeIntra, nTbW, and nTbH of the mapping process presented in Table 17 are the updated predModeIntra in Table 15 and the values of nTbW and nTbH corresponding to the width and height of the chroma block referred to in Table 15, respectively.
[0549] A variable whRatio representing the ratio of the width and the height of the block can be set to Abs(Log2(nW / nH)).
[0550] For example, when the width of the chroma block is greater than the height, the updated intra mode is 2 or greater, and the updated intra mode is the variable (whRatio>1)?(8+2*whRatio):8 [where whRatio is Abs(Log2(nW / nH))], the updated intra mode can be remapped to "updated intra mode + 65".
[0551] Or, when the height of the chroma block is greater than the width, the updated intra mode is 66 or less, and the updated intra mode is a variable (whRatio>1)?(60-2*whRatio):60 [where whRatio is Abs(Log2(nW / nH))], the updated intra mode can be remapped to "updated intra mode-67".
[0552] That is, the encoding device can re-update the intra prediction mode for determining the LFNST set by reflecting the non-square of the chroma block.
[0553] The encoding device can determine the LFNST set including the LFNST matrix based on the remapped intra prediction mode (S1760), and can derive modified transform coefficients of the chroma block based on the residual samples and the LFNST matrix (S1770).
[0554] The encoding device can determine a transform set based on the mapping relationship according to the intra prediction mode applied to the current block, and can perform LFNST, i.e., a non-separable transform, based on any one of the two LFNST matrices included in the transform set.
[0555] As described above, a plurality of transform sets can be determined according to the intra prediction mode of the transform block to be transformed. The matrix applied to the LFNST is a transpose of the matrix used in the inverse LFNST.
[0556] In one example, the LFNST matrix can be a non-square matrix having a smaller number of rows than a number of columns.
[0557] The encoding device can derive quantized transform coefficients by performing quantization based on the modified transform coefficients of the current chroma block, and can encode and output image information including information on the quantized transform coefficients, information on the intra prediction mode, and an LFNST index indicating the LFNST matrix (S1780).
[0558] Specifically, the encoding device 200 can generate information on the quantized transform coefficients and can encode the generated information on the quantized transform coefficients.
[0559] In one example, the information on the quantized transform coefficients can include at least one of information on whether to apply the LFNST, information on a reduction factor, information on a minimum transform size for applying the LFNST, and information on a maximum transform size for applying the LFNST.
[0560] The encoding apparatus can encode flag information indicating whether CCLM is applied to the current block as sps_cclm_enabled_flag and information about an intra prediction mode of a chroma component as intra_chroma_pred_mode as information about an intra mode.
[0561] The information about the CCLM mode as intra_chroma_pred_mode can indicate a top-left CCLM mode, a top CCLM mode, or a left CCLM mode.
[0562] In the disclosure, at least one of quantization / dequantization and / or transform / inverse transform can be omitted. When quantization / dequantization is omitted, a quantized transform coefficient can be referred to as a transform coefficient. When transform / inverse transform is omitted, a transform coefficient can be referred to as a coefficient or a residual coefficient, or can still be referred to as a transform coefficient for consistency of expression.
[0563] In addition, in the disclosure, a quantized transform coefficient and a transform coefficient can be referred to as a transform coefficient and a scaled transform coefficient, respectively. In this case, residual information can include information about a transform coefficient, and the information about the transform coefficient can be signaled through a residual coding syntax. A transform coefficient can be derived based on the residual information (or the information about the transform coefficient), and a scaled transform coefficient can be derived through inverse transform (scaling) of the transform coefficient. A residual sample can be derived based on inverse transform (transform) of the scaled transform coefficient. These details can also be applied / expressed in other parts of the disclosure.
[0564] In the above-described embodiments, a method is explained based on a flowchart by means of a series of steps or blocks, but the disclosure is not limited to the order of the steps, and a certain step can be performed in a different order or step from the above-described order or step, or concurrently with other steps. In addition, it can be understood by one of ordinary skill in the art that the steps shown in the flowchart are not exclusive, and one or more steps in the flowchart can be incorporated or one or more steps in the flowchart can be deleted without affecting the scope of the disclosure.
[0565] The above-described method according to the disclosure can be implemented in the form of software, and an encoding apparatus and / or a decoding apparatus according to the disclosure can be included in an apparatus for image processing such as a television, a computer, a smart phone, a set-top box, and a display device.
[0566] When the embodiments of the disclosure are implemented by software, the above-described methods can be implemented as modules (steps, functions, etc.) for performing the above-described functions. The modules can be stored in the memory and can be executed by the processor. The memory can be inside or outside the processor, and can be connected to the processor in various well-known ways. The processor can include an application-specific integrated circuit (ASIC), other chipsets, logic circuit, and / or data processing device. The memory can include read-only memory (ROM), random access memory (RAM), flash memory, memory card, storage medium, and / or other storage device. That is, the embodiments described in the disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each of the drawings can be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0567] In addition, the decoding apparatus and the encoding apparatus according to the disclosure can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and can be used to process a video signal or a data signal. For example, the over-the-top (OTT) video device can include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0568] In addition, the processing method according to the disclosure can be produced in the form of a program executed by a computer, and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes various storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (for example, transmission over the Internet). In addition, the bitstream generated by the encoding method can be stored in a computer-readable recording medium or transmitted through a wired or wireless communication network. In addition, the embodiments of the disclosure can be implemented as a computer program product by program codes, and the program codes can be executed on a computer according to the embodiments of the disclosure. The program codes can be stored on a computer-readable carrier.
[0569] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims of the disclosure can be combined to be implemented or performed in an apparatus, and the technical features of the apparatus claims can be combined to be implemented or performed in a method. Also, the technical features of the method claims and the technical features of the apparatus claims can be combined to be implemented or performed in an apparatus, and the technical features of the method claims and the technical features of the apparatus claims can be combined to be implemented or performed in a method.
Claims
1. A decoding device for image decoding, the decoding device comprising: a memory; and at least one processor connected to the memory, the at least one processor configured to: obtain, from a bitstream, intra prediction mode information and a low-frequency non-separable transform (LFNST) index; derive, based on the intra prediction mode information, an intra prediction mode of a chroma block to a cross-component linear model (CCLM) mode; update the intra prediction mode of the chroma block based on an intra prediction mode of a luma block corresponding to the chroma block; modify the updated intra prediction mode to a wide-angle intra prediction mode based on the chroma block not being square; determine, based on the modified intra prediction mode, a LFNST set including a LFNST matrix; derive, based on a LFNST matrix derived from the LFNST index and the LFNST set, transform coefficients of the chroma block; and derive, based on the transform coefficients, residual samples of the chroma block, wherein the updated intra prediction mode is derived to be an intra prediction mode corresponding to a particular position in the luma block, wherein the particular position is set based on a color format of the chroma block, wherein the particular position is set to ((xTbY + (nTbW * SubWidthC) / 2), (yTbY + (nTbH * SubHeightC) / 2)), wherein xTbY and yTbY represent top-left coordinates of the luma block, wherein nTbW and nTbH represent a width and a height of the chroma block, wherein SubWidthC and SubHeightC represent variables corresponding to the color format, and wherein based on the width of the chroma block being greater than the height, the updated intra mode being 2 or greater and the updated intra mode being less than a variable (whRatio > 1)? (8 + 2 * whRatio) : 8, the updated intra mode is modified to "updated intra mode + 65", where whRatio is Abs(Log2(nW / nH)) and here nW = nTbW and nH = nTbH. based on the color format being 4:2:0, SubWidthC and SubHeightC being 2, and 2. The decoding device of claim 1, wherein, wherein based on the color format being 4:2:2, SubWidthC being 2 and SubHeightC being 1. based on the intra prediction mode corresponding to the particular position being a palette mode or an intra block copy (IBC) mode, the intra prediction mode of the chroma block is updated to an intra direct current (DC) mode.
3. The decoding device of claim 1, wherein, based on the intra prediction mode corresponding to the particular position being a matrix-based intra prediction (MIP) mode, the intra prediction mode of the chroma block is updated to an intra planar mode.
4. The decoding device of claim 1, wherein, 5. An encoding device for image encoding, the encoding device comprising: a memory; and at least one processor connected to the memory, the at least one processor configured to: deriving an intra prediction mode of a chroma block as a cross-component linear model, CCLM, mode; deriving prediction samples of the chroma block based on the CCLM mode; deriving residual samples of the chroma block based on the prediction samples; updating the intra prediction mode of the chroma block based on an intra prediction mode of a luma block corresponding to the chroma block; modifying the updated intra prediction mode as a wide-angle intra prediction mode based on the chroma block not being square; determining a low-frequency non-separable transform, LFNST, set including a LFNST matrix based on the modified intra prediction mode; and deriving modified transform coefficients of the chroma block based on the residual samples and the LFNST matrix, wherein the updated intra prediction mode is derived as an intra prediction mode corresponding to a particular position in the luma block, wherein the particular position is set based on a color format of the chroma block, wherein the particular position is set as ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)), wherein xTbY and yTbY represent top-left coordinates of the luma block, wherein nTbW and nTbH represent a width and a height of the chroma block, wherein SubWidthC and SubHeightC represent variables corresponding to the color format, and wherein the updated intra mode is modified as "updated intra mode + 65" based on the width of the chroma block being greater than the height, the updated intra mode being 2 or greater and the updated intra mode being less than a variable (whRatio>1)? (8+2*whRatio):8, where whRatio is Abs(Log2(nW / nH)), here, nW=nTbW and nH=nTbH.
6. The encoding device of claim 5, wherein, based on the color format being 4:2:0, SubWidthC and SubHeightC being 2, and wherein the updated intra mode is modified as "updated intra mode + 65" based on the color format being 4:2:2, SubWidthC being 2 and SubHeightC being 1.
7. The encoding device of claim 5, wherein, based on the intra prediction mode corresponding to the particular position being a palette mode or an intra block copy, IBC, mode, the intra prediction mode of the chroma block is updated as an intra direct current, DC, mode.
8. The encoding device of claim 5, wherein, based on the intra prediction mode corresponding to the particular position being a matrix-based intra prediction, MIP, mode, the intra prediction mode of the chroma block is updated as an intra planar mode.
9. An apparatus for transmitting data of an image, the apparatus comprising: at least one processor configured to obtain a bitstream for an image, wherein the bitstream is generated based on processing of deriving an intra prediction mode of a chroma block as a cross-component linear model (CCLM) mode, deriving a prediction sample of the chroma block based on the CCLM mode, deriving a residual sample of the chroma block based on the prediction sample, updating the intra prediction mode of the chroma block based on an intra prediction mode of a luma block corresponding to the chroma block, modifying the updated intra prediction mode as a wide-angle intra prediction mode based on the chroma block not being square, determining a low-frequency non-separable transform (LFNST) set including a LFNST matrix based on the modified intra prediction mode, deriving modified transform coefficients of the chroma block based on the residual sample and the LFNST matrix, and encoding image information related to the modified transform coefficients, and a transmitter configured to transmit the data of the bitstream including the image information, wherein the updated intra prediction mode is derived as an intra prediction mode corresponding to a particular position in the luma block, wherein the particular position is set based on a color format of the chroma block, wherein the particular position is set as ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)), wherein xTbY and yTbY represent top-left coordinates of the luma block, wherein nTbW and nTbH represent a width and a height of the chroma block, wherein SubWidthC and SubHeightC represent variables corresponding to the color format, and wherein the updated intra mode is modified as "updated intra mode+65" based on the width of the chroma block being greater than the height, the updated intra mode being 2 or greater and the updated intra mode being less than a variable (whRatio>1)? (8+2*whRatio):8, where whRatio is Abs(Log2(nW / nH)), here, nW=nTbW and nH=nTbH.
Citation Information
Patent Citations
Coding video data using derived chroma mode
CN110100436A
Intra-prediction-based image coding method and apparatus thereof
WO2019198997A1