Image decoding apparatus, image encoding apparatus, and image data transmission apparatus
By employing CCLM mode to update the intra-frame prediction mode and LFNST transform in image coding, the problem of efficient compression of high-resolution images/videos is solved, improving coding efficiency and device performance.
Patent Information
- Application Number
- CN202511558541.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-29
- Filing Date
- 2020-10-29
- Publication Date
- 2026-01-02
AI Technical Summary
Existing technologies suffer from high costs due to increased information volume when transmitting and storing high-resolution, high-quality images/videos, and there is a lack of efficient compression technologies to address the growing demand for immersive media and image/video broadcasting of real images.
An image coding method is adopted, which updates the intra-frame prediction mode of the chroma block to the intra-frame DC mode through the cross-component linear model (CCLM) mode, and uses the LFNST matrix to perform LFNST transformation to derive the LFNST set, thereby improving coding efficiency.
It increases image/video compression efficiency, improves the efficiency of encoding LFNST index and secondary transformation, and optimizes the overall performance of image encoding devices.
Smart Images

Figure CN121262385A_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application No. 202080090728.1 (International Application No.: PCT / KR2020 / 014924, Application Date: October 29, 2020, Invention Title: Transformation-Based Image Coding Method and Apparatus). Technical Field
[0002] This disclosure relates to image coding techniques, and more specifically, to methods and apparatus for transform-based image coding systems. Background Technology
[0003] Today, the demand for high-resolution and high-quality images / videos, such as 4K, 8K, or even higher Ultra High Definition (UHD) images / videos, is constantly growing across various fields. As image / video data becomes higher resolution and higher quality, the amount of information or bits transmitted increases compared to traditional image data. Therefore, transmission and storage costs increase when using media such as traditional wired / wireless broadband lines to transmit image data or when using existing storage media to store image / video data.
[0004] In addition, there is increasing interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms, and broadcasting of images / videos with image characteristics that differ from real images such as game images is on the rise.
[0005] Therefore, there is a need for efficient image / video compression techniques to effectively compress, transmit, store, and reproduce information with high resolution and high quality images / videos that have the various characteristics described above. Summary of the Invention
[0006] Technical Purpose
[0007] One aspect of this disclosure is to provide a method and apparatus for increasing image coding efficiency.
[0008] Another technical aspect of this disclosure is to provide a method and apparatus for increasing the efficiency of encoding LFNST indexes.
[0009] Another technical aspect of this disclosure is to provide a method and apparatus for improving the efficiency of a secondary transformation by encoding an LFNST index.
[0010] Another aspect of this disclosure is to provide an image coding method and an image coding device for deriving LFNST transform sets using intra-frame modes with luma blocks in CCLM mode.
[0011] Technical solution
[0012] According to embodiments of this disclosure, an image decoding method performed by a decoding device is provided. The method may include the following steps: updating the intra-prediction mode of the chroma block to an intra-DC mode based on the fact that the intra-prediction mode of the chroma block is a cross-component linear model (CCLM) mode and the intra-prediction mode of the luma block corresponding to the chroma block is a palette mode; determining an LFNST set including an LFNST matrix based on the updated intra-prediction mode; and performing LFNST on the chroma block based on the LFNST matrix derived from the LFNST set, wherein the intra-DC mode is an intra-prediction mode corresponding to a specific position in the luma block.
[0013] The specific location is set based on the color format of the chroma block.
[0014] The specific location is the center of the brightness block.
[0015] The specific position is set to ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)), where xTbY and yTbY represent the top-left coordinates of the luma block, nTbW and nTbH represent the width and height of the chroma block, and SubWidthC and SubHeightC represent variables corresponding to the color format.
[0016] When the color format is 4:2:0, SubWidthC and SubHeightC are both 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0017] When the intra-prediction mode corresponding to a specific location is MIP mode, the intra-prediction mode of the chroma block is updated to intra-plane mode.
[0018] When the intra-prediction mode corresponding to a specific location is IBC mode, the intra-prediction mode of the chroma block is updated to intra-DC mode.
[0019] According to another embodiment of this disclosure, an image coding method performed by an encoding device is provided. The method may include: updating the intra-prediction mode of the chroma block to an intra-DC mode based on the fact that the intra-prediction mode of the chroma block is a cross-component linear model (CCLM) mode and the intra-prediction mode of the luma block corresponding to the chroma block is a palette mode; determining an LFNST set including an LFNST matrix based on the updated intra-prediction mode; and performing LFNST on the chroma block based on residual samples and the LFNST matrix, wherein the intra-DC mode is an intra-prediction mode corresponding to a specific position in the luma block.
[0020] According to another embodiment of the present disclosure, a digital storage medium may be provided that stores image data including a bitstream and encoded image information generated according to an image encoding method performed by an encoding device.
[0021] According to another embodiment of the present disclosure, a digital storage medium can be provided that stores image data including encoded image information and bitstreams to enable a decoding device to perform an image decoding method.
[0022] Technical effect
[0023] According to this disclosure, the overall image / video compression efficiency can be increased.
[0024] According to this disclosure, the efficiency of encoding LFNST indexes can be increased.
[0025] According to this disclosure, the efficiency of the second transformation can be increased by encoding the LFNST index.
[0026] According to this disclosure, an image coding method and image coding device for deriving the LFNST transform set using an intra-frame mode with luma blocks in CCLM mode can be provided.
[0027] The effects achievable through the specific examples of this disclosure are not limited to those listed above. For example, various technical effects may exist that can be understood or derived from this disclosure by one of ordinary skill in the art. Therefore, the specific effects of this disclosure are not limited to those expressly described herein, but may include various effects that can be understood or derived from the technical features of this disclosure. Attached Figure Description
[0028] Figure 1 Examples of video / image coding systems to which this disclosure can be applied are illustrated schematically.
[0029] Figure 2 This is a diagram that schematically illustrates the configuration of a video / image encoding device to which this disclosure can be applied.
[0030] Figure 3 This is a diagram that schematically illustrates the configuration of a video / image decoding device to which this disclosure can be applied.
[0031] Figure 4 An example of an intra-frame orientation mode with 65 predicted directions is shown.
[0032] Figure 5 This is a diagram illustrating a wide-angle intra-frame prediction mode according to an implementation method described in this document.
[0033] Figure 6This is a diagram used to describe the CCLM that can be applied when deriving the intra-prediction mode of the chroma block according to the implementation method.
[0034] Figure 7 Multiple variations of the implementation scheme according to this document are illustrated schematically.
[0035] Figure 8 This is a diagram used to illustrate the implementation of the RST according to this document.
[0036] Figure 9 This is a diagram illustrating the order in which the output data of a forward first transformation is arranged into a one-dimensional vector, based on the example.
[0037] Figure 10 This is a diagram illustrating the order in which the output data of the forward quadratic transform is arranged into a two-dimensional vector, based on the example.
[0038] Figure 11 This is a diagram illustrating the block shape to which LFNST is applied.
[0039] Figure 12 This is a diagram illustrating the arrangement of the output data of the forward LFNST according to an embodiment.
[0040] Figure 13 This is a diagram illustrating that the amount of output data for the positive LFNST, as shown in the example, is limited to a maximum of 16.
[0041] Figure 14 This is a diagram illustrating the zeroing process in a block of 4×4 LFNST, based on the example.
[0042] Figure 15 This is a diagram illustrating the zeroing process in a block of 8×8 LFNST, based on the example.
[0043] Figure 16 This is a diagram illustrating the zeroing of a block that has been applied using 8×8 LFNST, based on another example.
[0044] Figure 17 This is a diagram illustrating examples of horizontal and vertical traversal scanning methods used to encode a palette index map.
[0045] Figure 18 This is a flowchart used to illustrate the image decoding method based on the example.
[0046] Figure 19 This is a flowchart used to illustrate the image encoding method based on the example.
[0047] Figure 20 The structure of a content streaming system to which this disclosure is applied is illustrated. Detailed Implementation
[0048] While this disclosure may be readily modified and includes various embodiments, specific embodiments thereof have been illustrated by way of example in the accompanying drawings and will now be described in detail. However, this is not intended to limit this disclosure to the specific embodiments disclosed herein. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the technical concept of this disclosure. The singular form may include the plural form unless the context clearly indicates otherwise. Terms such as “comprising” and “having” are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and should therefore not be construed as pre-excluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0049] Furthermore, for ease of description of their different features and functions, the components in the accompanying drawings described herein are illustrated independently; however, this does not imply that each component is implemented by a separate piece of hardware or software. For example, any two or more of these components may be combined to form a single component, and any single component may be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of this disclosure, provided they do not depart from the spirit of this disclosure.
[0050] In the following description, preferred embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Furthermore, in the drawings, the same reference numerals are used for the same components, and repeated descriptions of the same components will be omitted.
[0051] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Video Coding Universal) standard (ITU-T Rec. H.266), the next generation video / image coding standard after VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), EVC (Essential Video Coding) standard, AVS2 standard, etc.).
[0052] This document provides various implementations related to video / image encoding, and these implementations may be combined and performed in combination with each other unless otherwise specified.
[0053] In this document, video can refer to a collection of images over a period of time. Typically, an image is a unit representing a specific time region, while a strip / patch is a unit that constitutes a part of an image. A strip / patch can include one or more coding tree units (CTUs). An image can consist of one or more strips / patches. An image can consist of one or more patch groups. A patch group can include one or more patches.
[0054] A pixel or primitive (pel) can refer to the smallest unit that makes up a picture (or image). Alternatively, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. Alternatively, a sample can refer to a pixel value in the spatial domain, or, when the pixel value is transformed to the frequency domain, it can refer to the transform coefficients in the frequency domain.
[0055] A unit can represent the basic unit of image processing. A unit may include a specific region and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the context, units and terms such as blocks and regions may be used interchangeably. Typically, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0056] In this document, the terms “ / ” and “,” should be interpreted as indicating “and / or”. For example, the expression “A / B” can mean “A and / or B”. Additionally, “A, B” can mean “A and / or B”. Furthermore, “A / B / C” can mean “at least one of A, B, and / or C”.
[0057] Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" could include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".
[0058] In this disclosure, "at least one of A and B" can mean "only A", "only B" or "both A and B". Furthermore, in this disclosure, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".
[0059] Furthermore, in this disclosure, "at least one of A, B, and C" may mean "A only", "B only", "C only" or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0060] Additionally, the parentheses used in this disclosure can indicate "for example". Specifically, when indicated as "prediction (intra-frame prediction)", it can mean that "intra-frame prediction" is proposed as an example of "prediction". In other words, "prediction" in this disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" is proposed as an example of "prediction". Furthermore, when indicated as "prediction (i.e., intra-frame prediction)", this can also mean that "intra-frame prediction" is proposed as an example of "prediction".
[0061] The technical features described individually in one of the accompanying drawings of this disclosure may be implemented individually or simultaneously.
[0062] Figure 1 Examples of video / image coding systems to which this disclosure can be applied are illustrated schematically.
[0063] Reference Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.
[0064] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0065] Video sources can be obtained through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, video / image archives including previously captured video / images, etc. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated by computers, etc. In this case, the video / image capture process can be replaced by a process that generates related data.
[0066] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0067] A transmitter can send encoded video / image information or data, output in bitstream form, to a receiver in a receiving device via a digital storage medium or network, either as a file or a stream. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to a decoding device.
[0068] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0069] The renderer can render decoded video / images. The rendered video / images can then be displayed on a monitor.
[0070] Figure 2 This diagram schematically illustrates the configuration of a video / image encoding apparatus to which this disclosure may be applied. In the following, the term "video encoding apparatus" may include an image encoding apparatus.
[0071] Reference Figure 2 The encoding device 200 may include an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 described above may be constituted by one or more hardware components (e.g., an encoder chipset or processor). Furthermore, the memory 270 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0072] Image partitioner 210 can divide an input image (or picture or frame) input to encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a maximum coding unit (LCU), the coding units can be recursively partitioned according to a quadtree-binary-tritree (QTBTTT) structure. For example, based on a quadtree structure, a binary tree structure, and / or a ternary tree structure, a coding unit can be partitioned into multiple coding units of varying depths. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary tree structure can be applied later. Alternatively, a binary tree structure can be applied first. The encoding process according to this disclosure can be performed based on the final coding units without further partitioning. In this case, the maximum coding unit can be directly used as the final coding unit based on the encoding efficiency according to the image characteristics. Alternatively, the coding units can be recursively partitioned into deeper coding units as needed, thereby allowing the optimally sized coding unit to be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be separate or distinct from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transformation unit may be a unit for deriving the transform coefficients and / or a unit for deriving the residual signal from the transform coefficients.
[0073] Depending on the context, units and terms such as blocks and regions can be used to represent each other. Typically, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Samples can be used as a term corresponding to pixels or primitives (pellets) in a picture (or image).
[0074] Subtractor 231 subtracts the prediction signal (prediction block, prediction sample array) output from predictor 220 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to converter 232. Predictor 220 can perform prediction on the processing target block (hereinafter referred to as "current block") and can generate a prediction block that includes prediction samples of the current block. Predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various prediction-related information such as prediction mode information and send the generated information to entropy encoder 240. The prediction information can be encoded in entropy encoder 240 and output as a bitstream.
[0075] Intra-predictor 222 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the reference samples can be located near or separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, the directional modes can include, for example, 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes can be used depending on the settings. Intra-predictor 222 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.
[0076] Inter-frame predictor 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on the correlation between motion information of neighboring blocks and the current block, using blocks, sub-blocks, or samples as a basis. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same as or different from each other. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in jump mode and merge mode, the inter-frame predictor 221 can use motion information of neighboring blocks as motion information of the current block. In jump mode, unlike merge mode, residual signals cannot be sent. In motion information prediction (motion vector prediction, MVP) mode, motion vectors of neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0077] Predictor 220 can generate prediction signals based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to the prediction of a block, and can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform prediction on blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as in games with screen content coding (SCC). Although IBC essentially performs prediction within the current block, its execution is similar to inter-frame prediction in that it derives a reference block within the current block. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0078] The predicted signals generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate the reconstructed signal or the residual signal. The transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, transform techniques may include Discrete Cosine Transform (DCT) and Discrete Sine Transform (DST). The transformation can be at least one of the following: graph-based transformation (KLT), graph-based transformation (GBT), or conditional nonlinear transformation (CNT). Here, GBT refers to the transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to the transformation obtained based on a prediction signal generated using all previously reconstructed pixels. Furthermore, the transformation process can be applied to square pixel blocks of the same size, or to blocks of variable size that are not square.
[0079] Quantizer 233 quantizes the transform coefficients and sends them to entropy encoder 240, which encodes the quantized signal (information about the quantized transform coefficients) and outputs the encoded signal in a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector based on the coefficient scan order and generate information about the quantized transform coefficients based on this one-dimensional vector form. Entropy encoder 240 can perform various encoding methods such as exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 240 can encode information required for video / image reconstruction, other than the quantized transform coefficients (e.g., values of syntax elements), either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the unit level of the Network Abstraction Layer (NAL). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. In this disclosure, information and / or syntax elements sent from the encoding device to / signaled to the decoding device may be included in the video / image information. The video / image information can be encoded using the encoding process described above and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks, communication networks, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that sends the signal output from the entropy encoder 240 or a memory (not shown) that stores it may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoder 240.
[0080] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transform using vectorized transform coefficients via dequantizer 234 and inverse transformer 235, the residual signal (residual block or residual sample) can be reconstructed. Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222, thereby generating a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). When there is no residual for the processing target block, as in the case of applying a jump mode, the prediction block can be used as the reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the target image, and, as described later, for inter-frame prediction of the next image by filtering.
[0081] In addition, luminance mapping with chroma scaling (LMCS) can be applied in image encoding and / or reconstruction processing.
[0082] Filter 260 can improve subjective / objective video quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270, specifically in the DPB of memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive ring filter, bilateral filter, etc. As discussed later in the description of each filtering method, filter 260 can generate various filtering-related information and send the generated information to entropy encoder 240. The filtering information can be encoded in entropy encoder 240 and output as a bitstream.
[0083] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. Accordingly, the encoding device can avoid prediction mismatch between the encoding device 100 and the decoding device when applying inter-frame prediction, and can also improve encoding efficiency.
[0084] Memory 270 DPB can store modified reconstructed images for use as reference images in inter-frame predictor 221. Memory 270 can store motion information of blocks in the current image from which motion information has been derived (or encoded) and / or motion information of blocks in reconstructed images. The stored motion information can be sent to inter-frame predictor 221 to be used as motion information of neighboring blocks or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 222.
[0085] Figure 3This is a diagram that schematically illustrates the configuration of a video / image decoding device to which this disclosure can be applied.
[0086] Reference Figure 3 The video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra-frame predictor 331 and an inter-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above may be constituted by one or more hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0087] When the input includes a bitstream containing video / image information, the decoding device 300 can interact with data already prepared therein. Figure 2 The processing of video / image information in the encoding device correspondingly reconstructs the image. For example, the decoding device 300 can deduce units / blocks based on information related to block segmentation obtained from the bitstream. The decoding device 300 can perform decoding by using processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, an encoding unit, which can be segmented along a quadtree structure, binary tree structure, and / or ternary tree structure using encoding tree units or maximum encoding units. One or more transform units can be derived using encoding units. And, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproducer.
[0088] Decoding device 300 can receive data from... in the form of a bitstream. Figure 2The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc. In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or general constraint information. In this disclosure, the signaling / receiving information and / or syntax elements, which will be described subsequently, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information about the target syntax element and the decoding information of neighboring and target blocks, or information about symbols / bins decoded in previous steps, predict the bin generation probability based on the determined context model, and perform arithmetic decoding on the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model after determining it using information about symbols / bins decoded for the next symbol / bin. Prediction information from the information decoded in the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantization transform coefficients) and associated parameter information that have undergone entropy decoding in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, filtering information from the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) that receives the signal output from the encoding device can also configure the decoding device 300 as an internal / external component, and the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to this disclosure can be referred to as a video / image / picture encoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0089] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement can be performed based on the order of coefficient scans already performed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0090] The inverse converter 322 obtains the residual signal (residual block, residual sample array) by performing an inverse transformation on the transformation coefficients.
[0091] The predictor can perform predictions on the current block and generate a prediction block that includes prediction samples for the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from the entropy decoder 310, and specifically, can determine the intra-frame / inter-frame prediction mode.
[0092] The predictor can generate a predicted signal based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to the prediction of a block, and can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform intra-block copying (IBC) for the prediction of a block. Intra-frame block copying can be used for content image / video coding such as in games with screen content coding (SCC). Although IBC essentially performs prediction within the current block, its execution is similar to inter-frame prediction in that it derives a reference block within the current block. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0093] The intra-predictor 331 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the reference samples can be located near or separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-predictor 331 can determine the prediction mode applied to the current block by using the prediction modes applied to neighboring blocks.
[0094] Inter-frame predictor 332 can deduce the predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on the correlation between motion information of neighboring blocks and the current block, on a block, sub-block, or sample basis. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference image index of the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the mode of inter-frame prediction for the current block.
[0095] Adder 340 can generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from predictor 330. When there is no residual for processing the target block, as in the case of applying a jump mode, the prediction block can be used as the reconstruction block.
[0096] Adder 340 can be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current block, and as described later, it can be output by filtering or used for inter-frame prediction of the next image.
[0097] In addition, luminance mapping with chroma scaling (LMCS) can be applied in image decoding processing.
[0098] Filter 350 can improve subjective / objective video quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be sent to memory 360, specifically to the DPB of memory 360. Various filtering methods can include, for example, deblocking filtering, adaptive sample shifting, adaptive ring filtering, bilateral filtering, etc.
[0099] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks in the current image from which motion information has been derived (or decoded) and / or motion information of blocks in a reconstructed image. The stored motion information can be sent to inter-frame predictor 332 to be used as motion information of neighboring blocks or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 331.
[0100] The examples described in this specification in the predictor 330, dequantizer 321, inverse transformer 322 and filter 350 of the decoding device 300 can be similarly or correspondingly applied to the predictor 220, dequantizer 234, inverse transformer 235 and filter 260 of the encoding device 200, respectively.
[0101] As described above, prediction is performed to improve compression efficiency during video encoding. Accordingly, a prediction block can be generated that includes prediction samples for the current block, which is the target block for encoding. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived identically in both the encoding and decoding devices, and the encoding device can improve image encoding efficiency by signaling to the decoding device information about the residual between the original block and the prediction block (residual information), not the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed image including the reconstructed block.
[0102] Residual information can be generated through transformation and quantization processes. For example, an encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients. This allows it to signal the associated residual information to the decoding device (via a bitstream). Here, the residual information can include the value information, position information, transform technique, transform kernel, quantization parameters, etc., of the quantized transform coefficients. The decoding device can perform quantization / dequantization processes based on the residual information and derive residual samples (or residual sample blocks). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can derive the residual block by performing dequantization / inverse transform on the quantized transform coefficients to serve as a reference for inter-frame prediction of the next image, and can generate a reconstructed image based on this.
[0103] Figure 4 The intra-frame orientation patterns for 65 predicted directions are schematically shown.
[0104] In intra-frame prediction according to the implementation method of this document, it can be as follows: Figure 4 The diagram shows the use of 67 intra-frame prediction modes.
[0105] This expands the existing 35 orientation modes to 67 orientation modes for intra-frame coding of high-resolution images and more accurate prediction. Figure 4 The arrow indicated by the dashed line points to the 32 new orientation modes added from the original 35.
[0106] The Intra-Plane (INTRA_PLANAR) mode and Intra-DC (INTRA_DC) mode are identical to the existing Intra-Plane and Intra-DC modes. The 32 added orientation modes can be applied to all block sizes and can be applied to both intra-encoding and decoding of the luma and chroma components.
[0107] Reference Figure 4 Based on the intra-prediction mode 34 with a left-top diagonal prediction direction, intra-prediction modes can be divided into intra-prediction modes with horizontal directionality and intra-prediction modes with vertical directionality. Figure 4 In the diagram, H and V denote horizontal and vertical orientation, respectively, and the numbers -32 to 32 indicate a displacement of 1 / 32 unit at the sample grid position. These numbers can represent the offset for the mode index value. Intra-prediction modes 2 to 33 are horizontally oriented, and intra-prediction modes 34 to 66 are vertically oriented. Strictly speaking, intra-prediction mode 34 can be considered neither horizontal nor vertical, but it can be classified as horizontally oriented when determining the transform set of the quadratic transform. This is because the input data is transposed for a vertical orientation mode symmetric to intra-prediction mode 34, and the input data alignment method for the horizontal mode is used for intra-prediction mode 34. Transposing the input data means switching the rows and columns of the two-dimensional M×N block data to N×M data. Intra-prediction modes 18 and 50 can represent the horizontal and vertical intra-prediction modes, respectively, and intra-prediction mode 2 can be called the upper-right diagonal intra-prediction mode because it has a left reference pixel and performs prediction in the upper-right direction. Similarly, intra-prediction mode 34 can be referred to as the bottom-right diagonal intra-prediction mode, while intra-prediction mode 66 can be referred to as the bottom-left diagonal intra-prediction mode.
[0108] Figure 5 This is a diagram illustrating a wide-angle intra-frame prediction mode according to an implementation method described in this document.
[0109] Typical intra-prediction mode values can have values from 0 to 66 and from 81 to 83, and intra-prediction mode values extended due to WAIP can have values from -14 to 83 as shown. Values from 81 to 83 indicate CCLM (Cross-Component Linear Model) mode, and values from -14 to -1 and from 67 to 80 indicate intra-prediction mode extended due to WAIP application.
[0110] When the width of the current prediction block is greater than its height, the top reference pixel is typically closer to the interior of the block to be predicted. Therefore, prediction in the lower left direction is more accurate than prediction in the upper right direction. Conversely, when the height of the block is greater than its width, the left reference pixel is typically closer to the interior of the block to be predicted. Therefore, prediction in the upper right direction is more accurate than prediction in the lower left direction. Thus, applying remapping (i.e., mode index modification) to the index of the wide-angle intra-frame prediction mode can be advantageous.
[0111] When wide-angle intra-prediction is applied, information about existing intra-prediction patterns can be signaled, and after the information is parsed, it can be remapped to the index of the wide-angle intra-prediction pattern. Therefore, the total number of intra-prediction patterns used for a specific block (e.g., a non-square block of a specific size) can remain unchanged; that is, the total number of intra-prediction patterns is 67, and the encoding of the intra-prediction patterns used for a specific block can remain unchanged.
[0112] Table 1 illustrates the process of deriving the modified intra-frame mode by remapping the intra-frame prediction mode to the wide-angle intra-frame prediction mode.
[0113] [Table 1]
[0114]
[0115] In Table 1, the extended intra-prediction mode values are ultimately stored in the `predModeIntra` variable, and `ISP_NO_SPLIT` indicates that the CU block is not divided into sub-partitions using the intra-segmentation (ISP) technique currently used in the VVC standard. The `cIdx` variable values of 0, 1, and 2 indicate the cases for the luma, Cb, and Cr components, respectively. The `log2` function shown in Table 1 returns a log value with a base of 2, and the `Abs` function returns the absolute value.
[0116] The variable `predModeIntra`, which indicates the intra-prediction mode, along with the height and width of the transform block, serves as the input value for the wide-angle intra-prediction mode mapping process, and the output value is the modified intra-prediction mode `predModeIntra`. The height and width of the transform block or coded block can be the height and width of the current block used for intra-prediction mode remapping. In this case, the variable `whRatio`, which reflects the width-to-width ratio, can be set to `Abs(Log2(nW / nH))`.
[0117] For non-square blocks, the intra-prediction mode can be divided into two cases and modified accordingly.
[0118] First, if all conditions (1) to (3) are met, (1) the width of the current block is greater than its height, (2) the intra-prediction mode before modification is equal to or greater than 2, and (3) the intra-prediction mode is less than the value derived as (8+2*whRatio) when the variable whRatio is greater than 1 and less than 8 when the variable whRatio is less than or equal to 1 (predModeIntra is less than (whRatio > 1) ? (8 + 2 * whRatio) : 8), then the intra-prediction mode is set to a value 65 greater than predModeIntra [predModeIntra is set to equal to (predModeIntra+65)].
[0119] If the above is different, i.e., if conditions (1) to (3) are satisfied, (1) the height of the current block is greater than the width, (2) the intra-prediction mode before modification is less than or equal to 66, and (3) the intra-prediction mode is greater than the value derived as (60-2*whRatio) when whRatio is greater than 1 and greater than 60 when whRatio is less than or equal to 1 (predModeIntra is greater than (whRatio > 1) ? (60-2 * whRatio) : 60), then the intra-prediction mode is set to a value 67 smaller than predModeIntra [predModeIntra is set to equal to (predModeIntra-67)].
[0120] When performing intra-prediction on the current block, prediction of the luma component block (luma block) and its chroma component block (chroma block) of the current block can be performed. In this case, the intra-prediction mode of the chroma component (chroma block) can be set separately from the intra-prediction mode of the luma component (luma block).
[0121] In this specification, "chroma block" and "chroma image" can refer to the same meaning as "chrominance block" and "chrominance image," and therefore, "chroma" and "chrominance" can be used interchangeably. Similarly, "luma block" and "luma image" can refer to the same meaning as "luminance block" and "luminance image," and therefore, "luma" and "luminance" can be used interchangeably.
[0122] In this specification, "current chroma block" can refer to the chroma component block of the current block as the current coding unit, and "current luma block" can refer to the luma component block of the current block as the current coding unit. Therefore, the current luma block and the current chroma block correspond to each other. However, the current luma block and the current chroma block do not always have the same block shape and the same number of blocks, but may have different block shapes and different numbers of blocks depending on the circumstances. In some cases, the current chroma block may correspond to the current luma region, in which case the current luma region may include at least one luma block.
[0123] The intra-frame prediction mode for chroma components can be indicated based on intra-frame chroma prediction mode information, and this information can be signaled using the `intra_chroma_pred_mode` syntax element. For example, the intra-frame chroma prediction mode information can indicate one of the following: planar mode, DC mode, vertical mode, horizontal mode, derived mode (DM), and CCLM mode. Here, when using... Figure 4 Of the 67 intra-prediction modes illustrated, the planar mode can indicate intra-prediction mode 0, the DC mode can indicate intra-prediction mode 1, the vertical mode can indicate intra-prediction mode 50, and the horizontal mode can indicate intra-prediction mode 18. DM can also be referred to as direct mode. CCLM can be referred to as LM.
[0124] DM and CCLM are correlated intra-prediction modes that use information about the luma block to predict the chroma block. DM can refer to a mode in which the same intra-prediction mode as the luma component is applied as the intra-prediction mode for the chroma component. CCLM can refer to an intra-prediction mode in which, during the generation of the predicted chroma block, reconstructed samples of the luma block are subsampled, and then samples derived by applying CCLM parameters α and β to the subsampled samples are used as the predicted samples for the chroma block.
[0125] Figure 6 This is a diagram illustrating the CCLM that can be applied when deriving the intra-frame prediction mode of chroma blocks according to an implementation method.
[0126] In this specification, "reference sample template" can refer to the set of neighboring reference samples of the current chroma block used to predict the current chroma block. Reference sample templates can be predefined, and information about the reference sample templates can be signaled from the encoding device 200 to the decoding device 300.
[0127] Reference Figure 6 The set of shaded samples in a single row adjacent to the current 4×4 chroma block indicates the reference sample template. The reference sample template is configured as a single row of reference samples, while the reference sample areas in the corresponding luminance regions are configured as two rows, such as... Figure 6 As shown.
[0128] In this implementation, when intra-frame coding of the chroma image is performed in the Joint Exploration Test Model (JEM) used in the Joint Video Exploration Group (JVET), a cross-component linear model (CCLM) can be used. CCLM is a method for predicting the pixel values of the chroma image based on the pixel values of the reconstructed luminance image, and is based on the high correlation between the luminance and chroma images.
[0129] CCLM prediction of Cb and Cr chromaticity images can be performed based on the following formula.
[0130] [Formula 1]
[0131]
[0132] Here, Pred c (i,j) represents the Cb or Cr chromaticity image to be predicted, Rec L (i,j) represents the reconstructed luminance image adjusted to the chroma block size, and (i,j) represents the pixel coordinates. In the 4:2:0 color format, since the luminance image is twice the size of the chroma image, it is necessary to generate a Rec image with the chroma block size through downsampling. L Therefore, Rec can be considered. L (2i,2j) and neighboring pixels are used to take the Pred colorimetric image. c The pixels of the brightness image at (i,j). Rec L (i,j) can be referred to as the downsampled brightness sample.
[0133] For example, as shown in the following formula, Rec can be derived using six adjacent pixels. L '(i,j).
[0134] [Equation 2]
[0135]
[0136] α and β represent the adjacent templates of the Cb or Cr chromaticity blocks. Figure 6 The cross-correlation and average difference between adjacent templates of the brightness blocks in the mid-shade region. For example, α and β are represented by Equation 3.
[0137] [Formula 3]
[0138]
[0139] Here, L(n) represents the neighboring reference samples and / or left neighbor samples of the luma block corresponding to the current chroma image, C(n) represents the neighboring reference samples and / or left neighbor samples of the current chroma block currently being encoded, and (i,j) represents the pixel position. Alternatively, L(n) can represent the downsampled upper neighbor samples and / or left neighbor samples of the current luma block. N can represent the total number of pixel pairs (luminance and chroma) values used to calculate the CCLM parameters, and can indicate a value that is twice the smaller of the width and height of the current chroma block.
[0140] Furthermore, an image can be divided into a sequence of Coded Tree Units (CTUs). A CTU can correspond to a Coded Tree Block (CTB). Alternatively, a CTU may include a coded tree block for luma samples and two coded tree blocks for corresponding chroma samples. The tree type can be classified as single-tree or dual-tree based on whether the luma block and its corresponding chroma block have separate partitioning structures. A single-tree indicates that the chroma block has the same partitioning structure as the luma block, while a dual-tree indicates that the chroma component block has a different partitioning structure than the luma block.
[0141] The following will describe the transformation process required for the image encoding or decoding process disclosed herein.
[0142] Figure 7 Multiple transformation techniques according to embodiments of the present disclosure are illustrated schematically.
[0143] Reference Figure 7 The converter can correspond to the aforementioned Figure 2 The converter in the encoding device, and the inverse converter can correspond to the aforementioned Figure 2 Inverse converter in encoding devices, or Figure 3 The inverse converter in the decoding device.
[0144] The transformer can derive (first) transform coefficients (S710) by performing a first transform based on residual samples (residual sample array) in the residual block. This first transform can be referred to as the core transform. In this paper, the first transform can be based on multiple transform selection (MTS), and when multiple transforms are used as a first transform, it can be referred to as a multi-core transform.
[0145] Multi-core transform can represent a method of performing transforms by additionally using Discrete Cosine Transform (DCT) Type 2 and Discrete Sine Transform (DST) Type 7, DCT Type 8, and / or DST Type 1. In other words, multi-core transform can represent a method of transforming a spatial domain residual signal (or residual block) into frequency domain transform coefficients (or primary transform coefficients) based on multiple transform kernels selected from DCT Type 2, DST Type 7, DCT Type 8, and DST Type 1. In this paper, from the perspective of the transformer, primary transform coefficients can be referred to as temporary transform coefficients.
[0146] In other words, when applying conventional transform methods, transform coefficients can be generated by applying a spatial-to-frequency domain transform to the residual signal (or residual block) based on DCT type 2. In contrast, when applying multi-core transforms, transform coefficients (or single-stage transform coefficients) can be generated by applying a spatial-to-frequency domain transform to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. In this paper, DCT type 2, DST type 7, DCT type 8, and DST type 1 can be referred to as transform types, transform kernels, or transform cores. These DCT / DST transform types can be defined based on basis functions.
[0147] When performing a multi-core transform, a vertical transform kernel and a horizontal transform kernel can be selected from the transform kernels for the target block. A vertical transform can be performed on the target block based on the vertical transform kernel, and a horizontal transform can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform can indicate the transform of the horizontal components of the target block, and the vertical transform can indicate the transform of the vertical components of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target (CU or sub-block), including the residual block.
[0148] Furthermore, according to the example, if a transformation is performed by applying an MTS, the mapping relationship of the transformation kernels can be set by setting specific basis functions to predetermined values and combining the basis functions to be applied in the vertical or horizontal transformation. For example, when the horizontal transformation kernel is denoted as trTypeHor and the vertical transformation kernel is denoted as trTypeVer, a value of 0 for trTypeHor or trTypeVer can be set to DCT2, a value of 1 for trTypeHor or trTypeVer can be set to DST7, and a value of 2 for trTypeHor or trTypeVer can be set to DCT8.
[0149] In this scenario, the MTS index information can be encoded and signaled to the decoding device to indicate any one of the multiple transform cores. For example, MTS index 0 can indicate that both trTypeHor and trTypeVer values are 0, MTS index 1 can indicate that both trTypeHor and trTypeVer values are 1, MTS index 2 can indicate that trTypeHor is 2 and trTypeVer is 1, MTS index 3 can indicate that trTypeHor is 1 and trTypeVer is 2, and MTS index 4 can indicate that both trTypeHor and trTypeVer values are 2.
[0150] In one example, the transformation kernel set based on MTS index information is shown in the table below.
[0151] [Table 2]
[0152]
[0153] The transformer can perform a quadratic transformation based on the (first) transform coefficients to derive modified (second) transform coefficients (S720). A first transform is a transformation from the spatial domain to the frequency domain, while a quadratic transform refers to using the correlation between the (first) transform coefficients to transform into a more compact representation. A quadratic transform can include an inseparable transform. In this case, the quadratic transform can be called an inseparable quadratic transform (NSST) or a mode-dependent inseparable quadratic transform (MDNSST). An inseparable quadratic transform can represent a transformation based on an inseparable transform matrix, performing a quadratic transform on the (first) transform coefficients derived from the first transform to generate modified transform coefficients (or quadratic transform coefficients) for the residual signal. In this case, the vertical and horizontal transforms can be applied non-separately (or not independently) to the (first) transform coefficients, but rather applied one-time based on the inseparable transform matrix. In other words, an inseparable quadratic transform can represent a transform method where the vertical and horizontal components of the (first) transform coefficients are not separated, and for example, a two-dimensional signal (transform coefficients) is rearranged into a one-dimensional signal through a specific defined direction (e.g., row-first or column-first direction), and then modified transform coefficients (or quadratic transform coefficients) are generated based on the inseparable transform matrix. For example, M×N blocks are arranged in rows according to row-first order, in the order of first row, second row, ..., and Nth row. M×N blocks are arranged in columns according to column-first order, in the order of first column, second column, ..., and Mth column. An inseparable quadratic transform can be applied to the upper left region of a block containing (first) transform coefficients (hereinafter referred to as a transform coefficient block). For example, when both the width W and height H of the transform coefficient block are 8 or greater, an 8×8 inseparable quadratic transform can be applied to the upper left 8×8 region of the transform coefficient block. Furthermore, if both the width (W) and height (H) of the transform coefficient block are 4 or greater, and the width (W) or height (H) of the transform coefficient block is less than 8, then a 4×4 inseparable quadratic transform can be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block. However, the implementation is not limited to this; for example, even if only the condition that the width W or height H of the transform coefficient block is 4 or greater is met, a 4×4 inseparable quadratic transform can be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block.
[0154] Specifically, for example, if a 4×4 input block is used, the inseparable quadratic transformation can be performed as follows.
[0155] A 4×4 input block X can be represented as follows.
[0156] [Formula 4]
[0157]
[0158] If X is represented as a vector, then the vector It can be represented as follows.
[0159] [Formula 5]
[0160]
[0161] In Equation 5, the vector It is a one-dimensional vector obtained by rearranging the two-dimensional block X of Equation 4 according to the row priority order.
[0162] In this case, the inseparable quadratic transformation can be calculated as follows.
[0163] [Formula 6]
[0164]
[0165] In this formula, represents the transformation coefficient vector, while T represents the 16×16 (inseparable) transformation matrix.
[0166] Using Equation 6 above, the 16×1 transformation coefficient vector can be derived. And the vector can be scanned in order (horizontal, vertical, and diagonal, etc.). Reorganize into 4×4 blocks. However, the above calculation is an example, and the hypercube-Givens transform (HyGT) and similar methods can also be used to calculate inseparable quadratic transformations in order to reduce the computational complexity of inseparable quadratic transformations.
[0167] Furthermore, in inseparable quadratic transforms, the transform kernel (or transform type) can be selected as mode-dependent. In this case, the mode can include intra-frame prediction mode and / or inter-frame prediction mode.
[0168] As described above, an inseparable quadratic transformation can be performed based on an 8×8 transformation or a 4×4 transformation determined by the width (W) and height (H) of the transform coefficient block. An 8×8 transformation is a transformation applicable to an 8×8 region contained within the transform coefficient block when both W and H are equal to or greater than 8, and this 8×8 region can be the top-left 8×8 region within the transform coefficient block. Similarly, a 4×4 transformation is a transformation applicable to a 4×4 region contained within the transform coefficient block when both W and H are equal to or greater than 4, and this 4×4 region can be the top-left 4×4 region within the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, while the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.
[0169] Here, to select mode-dependent transform kernels, two inseparable quadratic transform kernels can be configured for each transform set of inseparable quadratic transforms for both 8×8 and 4×4 transforms, and there can be four transform sets. That is, four transform sets can be configured for 8×8 transforms, and four transform sets can be configured for 4×4 transforms. In this case, each transform set in the four transform sets for 8×8 transforms can include two 8×8 transform kernels, and each transform set in the four transform sets for 4×4 transforms can include two 4×4 transform kernels.
[0170] However, as the size of the transformation (i.e., the size of the region to which the transformation is applied) can be, for example, a size other than 8×8 or 4×4, the number of sets can be n, and the number of transformation kernels in each set can be k.
[0171] The transform set can be referred to as the NSST set or the LFNST set. A specific set within the transform set can be selected, for example, based on the intra-prediction mode of the current block (CU or sub-block). The Low-Frequency Inseparable Transform (LFNST) can be an example of a reduced inseparable transform, which will be described later, and represents an inseparable transform for low-frequency components.
[0172] For reference, for example, intra-prediction modes may include two non-directional (or non-angular) intra-prediction modes and 65 directional (or angular) intra-prediction modes. Non-directional intra-prediction modes may include planar intra-prediction mode number 0 and DC intra-prediction mode number 1, and directional intra-prediction modes may include 65 intra-prediction modes numbered 2 through 66. However, this is an example, and this document can be applied even if the number of intra-prediction modes differs. Furthermore, in some cases, intra-prediction mode number 67 may be used, and intra-prediction mode number 67 may represent a linear model (LM) mode.
[0173] Based on the example, it can be mapped according to Figure 4 and Figure 5 The four transform sets of the intra-prediction mode are shown in the table below, for example.
[0174] [Table 3]
[0175]
[0176] As shown in Table 3, any one of the four transform sets, i.e., lfnstTrSetIdx, can be mapped to any one of the four indices (i.e., 0 to 3) according to the intra-frame prediction mode.
[0177] When a specific set is determined to be used for an inseparable quadratic transform, one of the k transform kernels in that set can be selected using the inseparable quadratic transform index. The encoding device can derive the inseparable quadratic transform index indicating the specific transform kernel based on rate-distortion (RD) check and can signal the inseparable quadratic transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the inseparable quadratic transform index. For example, lfnst index 0 can refer to the first inseparable quadratic transform kernel, lfnst index 1 can refer to the second inseparable quadratic transform kernel, and lfnst index 2 can refer to the third inseparable quadratic transform kernel. Alternatively, lfnst index 0 can indicate that the first inseparable quadratic transform is not applied to the target block, and lfnst indexes 1 through 3 can indicate three transform kernels.
[0178] The converter can perform an inseparable quadratic transform based on the selected transform core and obtain modified (quadratic) transform coefficients. As mentioned above, the modified transform coefficients can be derived as transform coefficients quantized by a quantizer and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse converter in the encoding device.
[0179] Furthermore, as mentioned above, if the second transformation is omitted, the (first) transformation coefficients, which are the output of the first (separable) transformation, can be derived as the transformation coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.
[0180] The inverse transformer can perform a series of processes in the reverse order of those already executed in the aforementioned transformers. The inverse transformer can receive (dequantized) transform coefficients and derive (first) transform coefficients by performing a second (inverse) transform (S750), and obtain residual blocks (residual samples) by performing a first (inverse) transform on the (first) transform coefficients (S760). In this regard, from the perspective of the inverse transformer, the first transform coefficients can be referred to as modified transform coefficients. As described above, the encoding and decoding devices can generate reconstructed blocks based on the residual blocks and the prediction blocks, and can generate reconstructed images based on the reconstructed blocks.
[0181] The decoding device may also include a second-order inverse transform application determiner (or a component for determining whether to apply the second-order inverse transform) and a second-order inverse transform determiner (or a component for determining the second-order inverse transform). The second-order inverse transform application determiner can determine whether to apply the second-order inverse transform. For example, the second-order inverse transform can be NSST, RST, or LFNST, and the second-order inverse transform application determiner can determine whether to apply the second-order inverse transform based on a second-order transform flag obtained by parsing the bitstream. In another example, the second-order inverse transform application determiner can determine whether to apply the second-order inverse transform based on the transform coefficients of the residual block.
[0182] A secondary inverse transform determiner can determine the secondary inverse transform. In this case, the secondary inverse transform determiner can determine the secondary inverse transform applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra-prediction mode. In implementations, the secondary transform determination method can be determined depending on the primary transform determination method. Various combinations of primary and secondary transforms can be determined based on the intra-prediction mode. Furthermore, in the example, the secondary inverse transform determiner can determine the region where the secondary inverse transform is applied based on the size of the current block.
[0183] Furthermore, as mentioned above, if the second (inverse) transform is omitted, the (dequantized) transform coefficients can be received, a first (separable) inverse transform can be performed, and a residual block (residual sample) can be obtained. As mentioned above, the encoding and decoding devices can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed image based on the reconstructed block.
[0184] Furthermore, in this disclosure, a reduced quadratic transformation (RST) in which the size of the transformation matrix (kernel) is reduced can be applied to the concept of NSST in order to reduce the computational and storage requirements of the inseparable quadratic transformation.
[0185] Furthermore, the transform kernel, transform matrix, and coefficients constituting the transform kernel matrix described in this disclosure, i.e., kernel coefficients or matrix coefficients, can be represented in 8 bits. This is feasible in decoding and encoding devices, and compared to existing 9-bit or 10-bit representations, it reduces the amount of storage required to store the transform kernel and can reasonably accommodate performance degradation. Additionally, representing the kernel matrix in 8 bits allows for the use of smaller multipliers and is more suitable for Single Instruction Multiple Data (SIMD) instructions for optimal software implementation.
[0186] In this specification, the term "RST" can refer to a transformation performed on the residual samples of a target block based on a transformation matrix whose size is reduced according to a reduction factor. When performing a reduction transformation, the computational cost required for the transformation can be reduced due to the smaller size of the transformation matrix. In other words, RST can be used to address computational complexity issues that arise when transforming large blocks or when transforming indivisible blocks.
[0187] RST can be referred to by various terms such as reduced transform, reduced quadratic transform, reduced transform, simplified transform, and simple transform, and the names that RST can be called are not limited to the examples listed. Alternatively, since RST is performed primarily in the low-frequency region of the transform block that includes non-zero coefficients, it can be called low-frequency inseparable transform (LFNST). The transform index can be called the LFNST index.
[0188] Furthermore, when performing a second inverse transform based on RST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include: an inverse reduced second transformer that derives modified transform coefficients based on the inverse RST of the transform coefficients; and an inverse first transformer that derives the residual samples of the target block based on the inverse first transform of the modified transform coefficients. An inverse first transform refers to the inverse transform of a first transform applied to the residuals. In this disclosure, deriving transform coefficients based on a transform can mean deriving the transform coefficients by applying a transform.
[0189] Figure 8 This is a diagram illustrating an embodiment of the RST according to the present disclosure.
[0190] In this disclosure, "target block" may refer to the current block, residual block, or transform block to be encoded.
[0191] In the example RST, an N-dimensional vector can be mapped to an R-dimensional vector in another space, thus determining the reduced transformation matrix, where R is less than N. N can refer to the square of the length of the side of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the reduction factor can refer to the R / N value. The reduction factor can be called a reduction factor, shrinkage factor, simplification factor, or other various terms. Furthermore, R can be called a reduction coefficient, but depending on the situation, the reduction factor can refer to R. Additionally, depending on the situation, the reduction factor can refer to the N / R value.
[0192] In this example, the reduction factor or reduction coefficient can be signaled via a bitstream, but the example is not limited to this. For instance, a predetermined value for the reduction factor or reduction coefficient can be stored in each of the encoding device 200 and the decoding device 300, and in this case, the reduction factor or reduction coefficient does not need to be signaled separately.
[0193] The size of the reduced transformation matrix, as shown in the example, can be less than N×N (the size of the regular transformation matrix) and can be limited as shown in Equation 4 below.
[0194] [Formula 7]
[0195]
[0196] Figure 8 The matrix T in the reduced transformation block shown in (a) can refer to the matrix T in Equation 7. R×N .like Figure 8 As shown in (a), when the reduced transformation matrix T R×N When multiplied by the residual sample of the target block, the transformation coefficients of the current block can be derived.
[0197] In the example, if the size of the block to which the transformation is applied is 8×8 and R=16 (i.e., R / N = 16 / 64 = 1 / 4), then according to Figure 8 The RST of (a) can be represented as the matrix operation shown in Equation 8. In this case, the storage and multiplication computations can be reduced to approximately 1 / 4 by a reduction factor.
[0198] In this disclosure, matrix operations can be understood as operations on column vectors obtained by multiplying a column vector by a matrix placed to the left of the column vector.
[0199] [Formula 8]
[0200]
[0201] In Equation 8, r1 to r 64 The residual samples of the target block can be represented, and specifically, they can be the transformation coefficients generated by applying a single transformation. As a result of the calculation in Equation 8, the transformation coefficients c of the target block can be derived. i And derive c i The process can be shown in Equation 9.
[0202] [Formula 9]
[0203]
[0204] As a result of Equation 9, the transformation coefficients c1 to c of the target block can be derived. R In other words, when R = 16, the transformation coefficients c1 to c of the target block can be derived. 16If a conventional transform is applied instead of an RST, and a 64×64 (N×N) transform matrix is multiplied by a 64×1 (N×1) residual sample, only 16(R) transform coefficients are derived for the target block because of the application of the RST, even though 64(N) transform coefficients are derived for the target block. Since the total number of transform coefficients used for the target block is reduced from N to R, the amount of data sent from the encoding device 200 to the decoding device 300 is reduced, thus improving the transmission efficiency between the encoding device 200 and the decoding device 300.
[0205] When considering the size of the transformation matrix, the size of a regular transformation matrix is 64×64 (N×N), but the size of a reduced transformation matrix is reduced to 16×64 (R×N). Therefore, compared to performing a regular transformation, the storage usage ratio of performing an RST can be reduced. Furthermore, compared to the number of multiplications (N×N) when using a regular transformation matrix, using a reduced transformation matrix can reduce the number of multiplications (R×N) by the R / N ratio.
[0206] In the example, the transformer 232 of the encoding device 200 can derive the transform coefficients of the target block by performing a first transform and an RST-based second transform on the residual samples of the target block. These transform coefficients can be passed to the inverse transformer of the decoding device 300, and the inverse transformer 322 of the decoding device 300 can derive the modified transform coefficients based on the inverse reduced second transform (RST) for the transform coefficients, and can derive the residual samples of the target block based on the inverse first transform for the modified transform coefficients.
[0207] Based on the example inverse RST matrix T N×R Its size is N×R, which is larger than the size of the conventional inverse transformation matrix N×N, and is the same as the reduced transformation matrix T shown in Equation 7. R×N It has a transpose relationship.
[0208] Figure 8 The matrix T in the reduced inverse transform block shown in (b) t It can refer to the inverse RST matrix T N×R T (The superscript T indicates transpose). For example... Figure 8 As shown in (b), when the inverse RST matrix T N×R T Multiplying by the transform coefficients of the target block allows for the derivation of the modified transform coefficients of the target block or the residual samples of the target block. The inverse RST matrix T R×N T It can be represented as (T) R×N ) T N×R .
[0209] More specifically, when the inverse RST is used as a second inverse transformation, when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. Furthermore, the inverse RST can be used as the inverse first-order transform, and in this case, when the inverse RST matrix T... N×R T When multiplied by the transformation coefficients of the target block, the residual sample of the target block can be derived.
[0210] In the example, if the size of the block to which the inverse transform is applied is 8×8 and R=16 (i.e., R / N = 16 / 64 = 1 / 4), then according to Figure 8 The RST of (b) can be represented as a matrix operation as shown in Equation 10.
[0211] [Formula 10]
[0212]
[0213] In Equation 10, c1 to c 16 This can represent the transformation coefficients of the target block. As a result of Equation 10, the transformation coefficients representing the modifications to the target block or the r of the residual samples of the target block can be derived. j And derive r j The process can be shown in Equation 11.
[0214] [Equation 11]
[0215]
[0216] As a result of the calculation in Equation 11, the transformation coefficients representing the modification of the target block or the residual samples of the target block, r1 to r2, can be derived. N From the perspective of the size of the inverse transformation matrix, the size of the regular inverse transformation matrix is 64×64 (N×N), but the size of the inverse reduced transformation matrix is reduced to 64×16 (R×N). Therefore, compared with performing the regular inverse transformation, the storage utilization rate of performing the inverse RST can be reduced by the R / N ratio. In addition, when comparing the number of multiplications N×N when using the regular inverse transformation matrix, using the inverse reduced transformation matrix can reduce the number of multiplications (N×R) by the R / N ratio.
[0217] The transform set configuration shown in Table 3 can also be applied to 8×8 RST. That is, 8×8 RST can be applied based on the transform sets in Table 3. Since a transform set includes two or three transforms (kernels) depending on the intra-prediction mode, it can be configured to select one of up to four transforms, including those without applying a secondary transform. In the transforms without applying a secondary transform, the application of an identity matrix can be considered. Assuming indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case where the identity matrix is applied, i.e., without applying a secondary transform), the transform index or lfnst index, which is used as a syntax element, can be signaled for each transform coefficient block, thereby specifying the transform to be applied. That is, for the top-left 8×8 block, the 8×8 NSST in the RST configuration can be specified via the transform index, or the 8×8 lfnst can be specified when applying LFNST. 8×8 lfnst and 8×8 RST refer to transformations of 8×8 regions within a transform coefficient block when both W and H of the target block are equal to or greater than 8, and the 8×8 region can be the top-left 8×8 region within the transform coefficient block. Similarly, 4×4 lfnst and 4×4 RST refer to transformations of 4×4 regions within a transform coefficient block when both W and H of the target block are equal to or greater than 4, and the 4×4 region can be the top-left 4×4 region within the transform coefficient block.
[0218] According to embodiments of this disclosure, for the transformation during the encoding process, only 48 data points can be selected, and a maximum 16×48 transformation kernel matrix can be applied to them, instead of applying a 16×64 transformation kernel matrix to the 64 data points forming an 8×8 region. Here, "maximum" means that m has a maximum value of 16 in the m×48 transformation kernel matrix to generate m coefficients. That is, when performing RST by applying an m×48 transformation kernel matrix (m≤16) to an 8×8 region, 48 data points are input, and m coefficients are generated. When m is 16, 48 data points are input, and 16 coefficients are generated. That is, assuming 48 data points form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied sequentially, thereby generating a 16×1 vector. Here, the 48 data points forming the 8×8 region can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on 48 data points constituting the region other than the lower right 4×4 region within the 8×8 region. Here, when matrix operations are performed by applying a maximum 16×48 transformation kernel matrix, 16 modified transformation coefficients are generated. These 16 modified transformation coefficients can be arranged in the upper left 4×4 region according to the scan order, and the upper right 4×4 region and the lower left 4×4 region can be filled with zeros.
[0219] For the inverse transform in the decoding process, the transpose of the aforementioned transform kernel matrix can be used. That is, when performing inverse RST or LFNST during the inverse transform performed by the decoding device, the input coefficient data for applying inverse RST is arranged in a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the one-dimensional vector with the corresponding inverse RST matrix to the left of the one-dimensional vector is arranged in a two-dimensional block according to a predetermined arrangement order.
[0220] In summary, during the transformation process, when RST or LFNST is applied to an 8×8 region, matrix operations are performed on the 48 transformation coefficients in the upper left, upper right, and lower left regions of the 8×8 region (excluding the lower right region) with a 16×48 transformation kernel matrix. For matrix operations, the 48 transformation coefficients are input as a one-dimensional array. When performing matrix operations, 16 modified transformation coefficients are derived, and these modified coefficients can be arranged in the upper left region of the 8×8 region.
[0221] Conversely, in the inverse transform process, when the inverse RST or LFNST is applied to an 8×8 region, the 16 transform coefficients corresponding to the upper left region of the 8×8 region can be input as a one-dimensional array according to the scan order, and matrix operations can be performed with a 48×16 transform kernel matrix. That is, the matrix operation can be expressed as (48×16 matrix) * (16×1 transform coefficient vector) = (48×1 modified transform coefficient vector). Here, an n×1 vector can be interpreted as having the same meaning as an n×1 matrix, and therefore can be represented as an n×1 column vector. Furthermore, * denotes matrix multiplication. When performing matrix operations, 48 modified transform coefficients can be derived, and these 48 modified transform coefficients can be arranged in the upper left, upper right, and lower left regions of the 8×8 region, excluding the lower right region.
[0222] When the inverse quadratic transform is based on the Regression-Simplified Transform (RST), the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include an inverse reduced quadratic transformer for deriving modified transform coefficients based on the inverse RST of the transform coefficients, and an inverse first-order transformer for deriving residual samples of the target block based on the inverse first-order transform of the modified transform coefficients. The inverse first-order transform refers to the inverse transform applied to the first-order transform of the residuals. In this disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying a transform.
[0223] The Non-Separate Transform (LFNST) described above will be described in detail below. LFNST may include a forward transform performed by the encoding device and an inverse transform performed by the decoding device.
[0224] The encoding device receives the result (or part of the result) derived after applying a first (core) transform as input and applies a forward second transform (second transform).
[0225] [Equation 12]
[0226]
[0227] In Equation 12, x and y are the input and output of the quadratic transformation, respectively, and G is the matrix representing the quadratic transformation, with the transformation basis vectors consisting of column vectors. In the case of inverse LFNST, when the dimension of the transformation matrix G is expressed as [number of rows × number of columns], in the case of forward LFNST, the transpose of matrix G becomes Ginverted. T Dimensions.
[0228] For the inverse LFNST, the dimensions of matrix G are [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices of the eight transformed basis vectors sampled from the left side of the [48×16] matrix and the [16×16] matrix, respectively.
[0229] On the other hand, for a positive LFNST, matrix G T The dimensions are [16×48], [8×48], [16×16], and [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transformation basis vectors from the upper part of the [16×48] matrix and the [16×16] matrix, respectively.
[0230] Therefore, in the case of forward LFNST, a [48×1] vector or a [16×1] vector can be used as input x, and a [16×1] vector or an [8×1] vector can be used as output y. In video encoding and decoding, the output of the forward first transform is two-dimensional (2D) data, so in order to construct a [48×1] vector or a [16×1] vector as input x, it is necessary to construct a one-dimensional vector by properly arranging the 2D data as the output of the forward transform.
[0231] Figure 9 This is a diagram illustrating the order in which the output data of a forward first transformation is arranged into a one-dimensional vector, based on the example. Figure 9 The left figures of (a) and (b) show the order used to construct the [48×1] vector, and Figure 9 The right figures (a) and (b) illustrate the order used to construct the [16×1] vector. In the case of LFNST, this can be achieved by combining 2D data with... Figure 9 Arrange the same order in (a) and (b) to obtain a one-dimensional vector x.
[0232] The orientation of the output data for the forward first transform can be determined based on the intra-prediction mode of the current block. For example, when the intra-prediction mode of the current block is horizontal relative to the diagonal direction, the orientation can be determined by... Figure 9 The output data of the forward first transform are arranged in the order of (a), and when the intra-prediction mode of the current block is perpendicular to the diagonal direction, it can be arranged according to... Figure 9 The output data of the first forward transformation are arranged in the order of (b).
[0233] Based on the example, different methods can be applied. Figure 9 The arrangement order of (a) and (b), and for derivation and application Figure 9 The arrangement order of (a) and (b) results in the same outcome (y vector), and the column vectors of matrix G can be rearranged according to the arrangement order. That is, the column vectors of G can be rearranged such that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0234] Since the output y derived by Equation 12 is a one-dimensional vector, when two-dimensional data is required as input data in the process of using the result of the forward quadratic transform as input (e.g., in the process of performing quantization or residual coding), the output y vector of Equation 12 needs to be properly arranged as 2D data again.
[0235] Figure 10 This is a diagram illustrating the order in which the output data of the forward quadratic transform is arranged into a two-dimensional vector, based on the example.
[0236] In the case of LFNST, the output values can be arranged in 2D blocks according to a predetermined scan order. Figure 10 (a) shows how the output values are arranged at 16 positions in a 2D block according to the diagonal scan order when the output y is a [16×1] vector. Figure 10 (b) shows that when the output y is an [8×1] vector, the output values are arranged in 8 positions of the 2D block according to the diagonal scan order, and the remaining 8 positions are filled with zeros. Figure 10 In (b), X indicates that it is filled with zeros.
[0237] According to another example, since the order in which the output vector y is processed during quantization or residual coding can be preset, the output vector y does not need to be arranged as shown in the example. Figure 10 In the 2D block shown. However, in the case of residual coding, data encoding can be performed in 2D block (e.g., 4×4) cells (e.g., CG (coefficient group)), and in this case, according to as Figure 10 The data is arranged in a specific order within the diagonal scanning sequence.
[0238] Furthermore, the decoding device can configure the one-dimensional input vector y by arranging the two-dimensional data output from the dequantization process according to a preset scan order used for the inverse transform. The input vector y can be output as the output vector x using the following formula.
[0239] [Equation 13]
[0240]
[0241] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or an [8×1] vector, by the G matrix. For inverse LFNST, the output vector x can be a [48×1] vector or a [16×1] vector.
[0242] The output vector x is based on Figure 9 The sequence shown is arranged in a two-dimensional block and is arranged as two-dimensional data, which becomes the input data (or part of the input data) for the inverse first transformation.
[0243] Therefore, the inverse quadratic transform is the opposite of the forward quadratic transform process in general, and in the case of the inverse transform, unlike in the forward direction, the inverse quadratic transform is applied first, followed by the inverse first transform.
[0244] In the inverse LFNST, one of eight [48×16] matrices and eight [16×16] matrices can be chosen as the transformation matrix G. Whether to apply the [48×16] matrix or the [16×16] matrix depends on the size and shape of the block.
[0245] Additionally, eight matrices can be derived from the four transform sets shown in Table 2 above, and each transform set can consist of two matrices. The choice of which of the four transform sets to use is determined based on the intra-prediction mode, and more specifically, based on the values of the intra-prediction mode extended by taking into account wide-angle intra-prediction (WAIP). The selection of which matrix from the two matrices constituting the chosen transform set is derived via index signaling. More specifically, 0, 1, and 2 can be used as transmit index values; 0 can indicate that LFNST is not applied, and 1 and 2 can indicate either of the two transform matrices constituting the transform set selected based on the intra-prediction mode values.
[0246] Furthermore, Table 3 above illustrates how to select the transform set in LFNST based on the intra-prediction mode values extended by WAIP. For example... Figure 5 As shown, modes 14 to 33 and modes 35 to 80 are symmetrical about the prediction directions around mode 34. For example, modes 14 and 54 are symmetrical about the direction corresponding to mode 34. Therefore, the same set of transformations is applied to modes located in mutually symmetrical directions, and this symmetry is also reflected in Table 3.
[0247] Furthermore, it is assumed that the positive LFNST input data of mode 54 is symmetrical to the positive LFNST input data of mode 14. For example, for modes 14 and 54, according to Figure 9 (a) and Figure 9 The arrangement shown in (b) rearranges the two-dimensional data into one-dimensional data. Furthermore, it can be seen that... Figure 9 (a) and Figure 9 The pattern in the sequence shown in (b) is symmetrical about the direction indicated by pattern 34 (diagonal direction).
[0248] Furthermore, as mentioned above, the size and shape of the target block determine which transformation matrix, either the [48×16] matrix or the [16×16] matrix, will be applied to the LFNST.
[0249] Figure 11 This is a diagram illustrating the block shape to which LFNST is applied. Figure 11 (a) shows a 4×4 block. Figure 11 (b) shows 4×8 blocks and 8×4 blocks. Figure 11 (c) shows a 4×N block or an N×4 block, where N is 16 or greater. Figure 11 (d) shows an 8×8 block. Figure 11 (e) shows an M×N block where M≥8, N≥8 and N>8 or M>8.
[0250] exist Figure 11 In the diagram, blocks with thick boundaries indicate the area where LFNST is applied. For Figure 11 For blocks (a) and (b), LFNST is applied to the top-left 4×4 region, and for Figure 11 Block (c) is individually applied to two consecutively arranged top-left 4×4 regions. Figure 11 In (a), (b), and (c), since the LFNST is applied in units of 4×4 regions, this LFNST will be referred to as "4×4 LFNST" in the following text. As the corresponding transformation matrix, a [16×16] or [16×8] matrix can be applied to Equations 12 and 13 based on the matrix dimension of G.
[0251] More specifically, a [16×8] matrix is applied to Figure 11 (a) 4×4 blocks (4×4 TU or 4×4 CU), and a [16×16] matrix is applied to Figure 11 The blocks in (b) and (c) are used to adjust the worst-case computational complexity to 8 multiplications per sample.
[0252] about Figure 11In (d) and (e), LFNST is applied to the top-left 8×8 region, and this LFNST is referred to as "8×8 LFNST" below. As the corresponding transformation matrix, a [48×16] matrix or a [48×8] matrix can be applied. In the case of the forward LFNST, since the [48×1] vector (the X vector in Equation 2) is input as input data, not all sample values from the top-left 8×8 region are used as input values for the forward LFNST. That is, if... Figure 9 The left-hand order of (a) or Figure 9 As can be seen from the left-hand order of (b), the [48×1] vector can be constructed based on the samples belonging to the other three 4×4 blocks while leaving the bottom right 4×4 block as is.
[0253] A [48×8] matrix can be applied to Figure 11 The 8×8 blocks (8×8 TU or 8×8 CU) in (d) and the [48×16] matrix can be applied Figure 11 The 8×8 blocks in (e). This is also to adjust the worst-case computational complexity to 8 multiplications per sample.
[0254] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data (the Y vector in Equation 12, [8×1] or [16×1] vectors) are generated. In the forward LFNST, due to matrix G... T Due to its characteristic, the amount of output data is equal to or less than the amount of input data.
[0255] Figure 12 This is a diagram illustrating the arrangement of the output data of the forward LFNST according to an example, and showing the blocks in which the output data of the forward LFNST is arranged according to the block shape.
[0256] exist Figure 12 The shaded area in the upper left corner of the block shown corresponds to the region where the output data of the forward LFNST is located. The positions marked with 0 indicate samples filled with a value of 0, and the remaining areas represent regions that were not altered by the forward LFNST. In regions not altered by LFNST, the output data of the first forward transform remains unchanged.
[0257] As mentioned above, since the size of the applied transformation matrix varies depending on the shape of the block, the amount of output data also varies. Figure 12 The output data of a forward LFNST may not completely fill the top-left 4×4 block. Figure 12In cases (a) and (d), the [16×8] matrix and the A[48×8] matrix are applied to the block indicated by the thick line or a portion of the area inside the block, respectively, and an [8×1] vector is generated as the output of the positive LFNST. That is, according to Figure 10 The scan order shown in (b) can fill only 8 output data, such as Figure 12 As shown in (a) and (d), zeros can be filled in the remaining 8 positions. Figure 11 In the case of the LFNST application block of (d), such as Figure 12 As shown in (d), the two 4×4 blocks adjacent to the top-left 4×4 block, the top-right and bottom-left blocks, are also filled with the value 0.
[0258] As described above, essentially, by signaling the LFNST index, it is specified whether to apply LFNST and the transformation matrix to be applied. Figure 12 As shown, when LFNST is applied, since the number of output data of the positive LFNST can be equal to or less than the number of input data, the following area filled with zero values appears.
[0259] 1) such as Figure 12 As shown in (a), the samples are from the eighth position and the subsequent positions in the scanning order of the top left 4×4 block, that is, from the ninth to the sixteenth position.
[0260] 2) such as Figure 12 As shown in (d) and (e), when applying a [48×16] matrix or a [48×8] matrix, the two 4×4 blocks adjacent to the top left 4×4 block or the second and third 4×4 blocks in the scan order.
[0261] Therefore, if non-zero data is found in regions 1) and 2), it is determined that LFNST has not been applied, so the signaling for the corresponding LFNST index can be omitted.
[0262] Based on the example, such as in the case of LFNST used in the VVC standard, since the signaling for the LFNST index is executed after residual coding, the encoding device can know from the residual coding whether non-zero data (valid coefficients) exists at all locations within the TU or CU block. Therefore, the encoding device can determine whether to execute signaling regarding the LFNST index based on the presence of non-zero data, and the decoding device can determine whether to parse the LFNST index. The signaling for the LFNST index is executed when non-zero data does not exist in the areas specified in 1) and 2) above.
[0263] Because truncated unary codes are applied as the binarization method for the LFNST index, the LFNST index consists of up to two bins, and 0, 10, and 11 are assigned as binary codes for the possible LFNST index values 0, 1, and 2, respectively. According to the example, context-based CABAC encoding can be applied to the first bin (regular encoding), and context-based CABAC encoding can also be applied to the second bin. The encoding of the LFNST index is shown in the table below.
[0264] [Table 4]
[0265]
[0266] As shown in Table 4, for the first bin (binIdx = 0), context 0 is applied in the case of a single tree, while context 1 can be applied in the case of a non-single tree. Furthermore, as shown in Table 4, context 2 can be applied to the second bin (binIdx = 1). That is, two contexts can be assigned to the first bin, and one context can be assigned to the second bin, and each context can be distinguished by the ctxInc value (0, 1, 2).
[0267] Here, a single tree means that the luma and chroma components are encoded using the same coding structure. When coding units are partitioned while having the same coding structure, and the size of the coding unit becomes less than or equal to a certain threshold, and the luma and chroma components are encoded with separate tree structures, the corresponding coding units are considered as dual trees, and therefore, the context of the first bin can be determined. That is, as shown in Table 4, context 1 can be assigned.
[0268] Alternatively, context 0 can be used when the value of the variable treeType is assigned to SINGLE_TREE of the first bin; otherwise, context 1 can be used.
[0269] In addition, the following simplification method can be applied to the LFNST used.
[0270] (i) As shown in the example, the number of output data for a positive LFNST can be limited to a maximum of 16.
[0271] exist Figure 11 In case (c), the 4×4 LFNST can be applied to two adjacent 4×4 regions to the upper left, and in this case, a maximum of 32 LFNST output data can be generated. When the number of output data for the forward LFNST is limited to a maximum of 16, in the case of 4×N / N×4 (N≥16) blocks (TU or CU), the 4×4 LFNST is applied only to one 4×4 region to the upper left, and the LFNST can be applied only to... Figure 11 All blocks are processed at once. This simplifies the implementation of image encoding.
[0272] Figure 13 The example shows that the amount of output data for a positive LFNST is limited to a maximum of 16. (See example.) Figure 13 When LFNST is applied to the top left 4×4 region of a 4×N or N×4 block (where N is 16 or greater), the output data of the forward LFNST becomes 16.
[0273] (ii) As in the example, zeroing can be additionally applied to regions where LFNST has not been applied. In this document, zeroing can mean filling all positions belonging to a particular region with a value of 0. That is, zeroing can be applied to regions that have not changed due to LFNST and maintain the result of a positive first transformation. As mentioned above, since LFNST is divided into 4×4 LFNST and 8×8 LFNST, zeroing can be divided into two types as follows ((ii)-(A) and (ii)-(B)).
[0274] (ii)-(A) When 4×4 LFNST is applied, the area where 4×4 LFNST is not applied can be zeroed. Figure 14 This is a diagram illustrating the zeroing process in a block of 4×4 LFNST, based on the example.
[0275] like Figure 14 As shown, regarding the block that applied 4×4 LFNST, that is, for Figure 12 All blocks in (a), (b) and (c) where LFNST is not applied can be filled with zeros.
[0276] on the other hand, Figure 14 (d) shows that when the maximum number of output data for the positive LFNST is limited to 16 (e.g. Figure 13 When (as shown), zeroing is performed on the remaining blocks that have not applied 4×4 LFNST.
[0277] (ii)-(B) When 8×8 LFNST is applied, areas where 8×8 LFNST is not applied can be zeroed. Figure 15 This is a diagram illustrating the zeroing process in a block of 8×8 LFNST, based on the example.
[0278] like Figure 15 As shown, regarding the application of 8×8 LFNST to the block, that is, for Figure 12 In all blocks in (d) and (e), the entire area where LFNST is not applied can be filled with zeros.
[0279] (iii) Due to the zeroing presented in (ii) above, the zero-filled area may not be the same as when LFNST was applied. Therefore, it can be determined by comparison. Figure 12 In the case of LFNST, a wider area is used to perform zeroing as proposed in (ii) to check for the presence of non-zero data.
[0280] For example, when (ii)-(B) is applied, in the examination Figure 12 After checking whether there is non-zero data in the zero-filled regions in (d) and (e), additional checks are performed. Figure 15 The presence of non-zero data in the region filled with 0s can be checked, and signaling for the LFNST index can be executed only if no non-zero data exists.
[0281] Of course, even with the zeroing proposed in application (ii), the existence of non-zero data can be checked in the same way as existing LFNST index signaling. That is, when checking... Figure 12 After confirming the presence of non-zero data within the zero-padded block, LFNST index signaling can be applied. In this case, the encoding device only performs zeroing and the decoding device does not assume zeroing; that is, it only checks whether non-zero data exists within the zero-padded block. Figure 12 In regions explicitly marked as 0, LFNST index resolution can be performed.
[0282] Alternatively, according to another example, it can be as follows: Figure 16 The zeroing process is executed as shown. Figure 16 This is a diagram illustrating the zeroing of a block in an 8×8 LFNST application, based on another example.
[0283] like Figure 14 and Figure 15 As shown, zeroing can be applied to all areas except those where LFNST is applied, or it can be applied only to areas such as... Figure 16 The shown area. Zeroing is only applied to... Figure 16 Except for the top-left 8x8 area, the zeroing can be omitted from the bottom-right 4x4 block within the top-left 8x8 area.
[0284] Various implementations of the simplified methods for applying LFNST (combinations of (i), (ii)-(A), (ii)-(B), (iii)) can be derived. Of course, the combinations of the above simplified methods are not limited to the following implementations, and any combination can be applied to LFNST.
[0285] Implementation
[0286] - Limit the number of output data for the forward LFNST to a maximum of 16. (i)
[0287] - When 4×4 LFNST is applied, all areas where 4×4 LFNST is not applied are zeroed out. (II)-(A)
[0288] - When 8×8 LFNST is applied, all areas where 8×8 LFNST is not applied are zeroed out. (II)-(B)
[0289] - After checking whether non-zero data also exists in existing areas filled with zero values and areas filled with zeros due to additional zeroing ((ii)-(A), (ii)-(B)), the LFNST index is signaled only if no non-zero data exists. (iii).
[0290] In the implementation scenario, when LFNST is applied, the area containing non-zeroed data is limited to the upper left 4×4 area. More specifically, in Figure 14 (a) and Figure 15 In case (a), the eighth position in the scan order is the last position where non-zero data can exist. Figure 14 (b) and (c) and Figure 15 In case (b), the sixteenth position in the scan order (i.e., the position at the bottom right edge of the top left 4×4 block) is the last position in which data other than 0 can exist.
[0291] Therefore, after applying LFNST, and after checking whether non-zero data exists at a position that is not allowed in the residual encoding process (at a position beyond the last position), it can be determined whether to signal the LFNST index.
[0292] In the case of the zeroing method proposed in (ii), the computational cost required to perform the entire transformation process can be reduced because of the amount of data ultimately generated when both the first transformation and LFNST are applied. That is, when LFNST is applied, since zeroing is applied to regions where the output data of the forward first transformation exists without LFNST, it is not necessary to generate data for regions that are zeroed during the forward first transformation. Therefore, the computational cost required to generate the corresponding data can be reduced. The additional effects of the zeroing method proposed in (ii) are summarized below.
[0293] First, as mentioned above, reduce the amount of computation required to perform the entire transformation process.
[0294] Specifically, when (ii)-(B) is applied, the worst-case computational cost is reduced, making the transformation process lighter. In other words, generally, a large amount of computation is required to perform a single transformation of a large size. By applying (ii)-(B), the amount of data derived as a result of performing a forward LFNST can be reduced to 16 or less. Furthermore, as the size of the entire block (TU or CU) increases, the effect of reducing the number of transformation operations further increases.
[0295] Secondly, it can reduce the amount of computation required for the entire transformation process, thereby reducing the power consumption required to perform the transformation.
[0296] Third, it reduces the delay involved in the transformation process.
[0297] Quadratic transforms such as LFNST add computational complexity to existing single transforms, thus increasing the overall latency involved in performing the transform. Specifically, in the case of intra-frame prediction, the increased latency due to the quadratic transform during encoding leads to an increase in latency until reconstruction because reconstructed data from adjacent blocks is used during prediction. This can result in an increase in the overall latency of intra-frame predictive coding.
[0298] However, if the zeroing proposed in application (ii) is applied, the delay time for performing a single transformation can be greatly reduced when LFNST is applied, maintaining or reducing the delay time of the entire transformation, making it easier to implement the encoding device.
[0299] In addition, the signaling for the LFNST index and the MTS index will be described below.
[0300] According to another example, the coding unit syntax table, transform unit syntax table, and residual coding syntax table are as follows. According to Table 5, the MTS index is moved from the transform unit level to the coding unit level syntax and is signaled after the LFNST index signaling. Additionally, the following constraint has been removed: LFNST is not allowed when ISP is applied to a coding unit. The constraint that LFNST is not allowed when ISP is applied to a coding unit has been removed, allowing LFNST to be applied to all intra-prediction blocks. Furthermore, both the MTS index and the LFNST index are conditionally signaled at the end of the coding unit level.
[0301] [Table 5]
[0302]
[0303] [Table 6]
[0304]
[0305] [Table 7]
[0306]
[0307] The meanings of the main variables shown in the table are as follows.
[0308] 1. cbWidth, cbHeight: The width and height of the current encoding block.
[0309] 2. log2TbWidth, log2TbHeight: The base-2 logarithmic values of the width and height of the current transform block, which can be reduced to the upper left region where non-zero coefficients can exist by reflecting zeroing.
[0310] 3. sps_lfnst_enabled_flag: A flag indicating whether LFNST is enabled. If the flag value is 0, it indicates that LFNST is not enabled, and if the flag value is 1, it indicates that LFNST is enabled. It is defined in the Sequence Parameter Set (SPS).
[0311] 4. CuPredMode[chType][x0][y0]: The prediction mode corresponding to the variable chType and the position (x0, y0). chType can have values of 0 and 1, where 0 indicates the luma component and 1 indicates the chroma component. The position (x0, y0) indicates the location on the image, and MODE_INTRA (intra-frame prediction) and MODE_INTER (inter-frame prediction) can be used as the values of CuPredMode[chType][x0][y0].
[0312] 5. IntraSubPartitionsSplit[x0][y0]: The content at position (x0, y0) is the same as in item 4. It indicates which ISP partition was applied at position (x0, y0), and ISP_NO_SPLIT indicates that the coding unit corresponding to position (x0, y0) was not divided into a partition block.
[0313] 6. intra_mip_flag[x0][y0]: The content at position (x0, y0) is the same as in point 4 above. intra_mip_flag is a flag indicating whether matrix-based intra-frame prediction (MIP) prediction mode is applied. If the flag value is 0, it indicates that MIP is not enabled; if the flag value is 1, it indicates that MIP is enabled.
[0314] 7. cIdx: A value of 0 indicates luminance, and values of 1 and 2 indicate the Cb and Cr of the chromaticity components, respectively.
[0315] 8. treeType: Indicates single-tree and dual-tree types (SINGLE_TREE: single-tree, DUAL_TREE_LUMA: dual-tree for luminance component, DUAL_TREE_CHROMA: dual-tree for chrominance component).
[0316] 9. lastSubBlock: This indicates the position of the subblock (coefficient group (CG)) containing the last non-zero coefficient in the scan order. 0 indicates a subblock containing the DC component, and a value greater than 0 indicates a subblock that does not contain the DC component.
[0317] 10. lastScanPos: This indicates the position of the last valid coefficient within a sub-block in scan order. If a sub-block contains 16 positions, it can have values from 0 to 15.
[0318] 11. lfnst_idx[x0][y0]: The LFNST index syntax element to be parsed. If not parsed, it is inferred to be 0. That is, the default value is set to 0, indicating that LFNST is not applied.
[0319] 12. LastSignificantCoeffX, LastSignificantCoeffY: These indicate the x and y coordinates of the last significant coefficient in the transform block. The x-coordinate starts at 0 and increases from left to right, and the y-coordinate starts at 0 and increases from top to bottom. If both variables are 0, it means the last significant coefficient is located at DC.
[0320] 13. cu_sbt_flag: A flag indicating whether Subblock Transformation (SBT) included in the current VVC standard is enabled. If the flag value is 0, it indicates that SBT is not enabled, and if the flag value is 1, it indicates that SBT is enabled.
[0321] 14. sps_explicit_mts_inter_enabled_flag, sps_explicit_mts_intra_enabled_flag: These flags indicate whether explicit MTS is applied to inter-frame CUs and intra-frame CUs, respectively. If the corresponding flag value is 0, it indicates that MTS is not enabled for inter-frame or intra-frame CUs; if the corresponding flag value is 1, it indicates that MTS is enabled.
[0322] 15. tu_mts_idx[x0][y0]: The MTS index syntax element to be parsed. If not parsed, it is inferred to be 0. That is, the default value is set to 0, indicating that DCT-2 is enabled in both the horizontal and vertical directions.
[0323] As shown in Table 5, when encoding mts_idx[x0][y0], multiple conditions are checked, and tu_mts_idx[x0][y0] is signaled only when lfnst_idx[x0][y0] is zero.
[0324] Additionally, tu_cbf_luma[x0][y0] is a flag indicating whether there are valid coefficients for the luminance component.
[0325] According to Table 5, when both the width and height of the coding unit of the luminance component are 32 or less, a signal is sent to mts_idx[x0][y0] (Max(cbWidth, cbHeight) <= 32). In other words, whether MTS is applied is determined by the width and height of the coding unit of the luminance component.
[0326] Furthermore, according to Table 5, even in ISP mode (IntraSubPartitionsSplitType != ISP_NO_SPLIT), lfnst_idx[x0][y0] can be configured to be signaled, and the same LFNST index value can be applied to all ISP partition blocks.
[0327] On the other hand, mts_idx[x0][y0] can only be signaled when not in ISP mode (IntraSubPartitionsSplit[x0][y0] == ISP_NO_SPLIT).
[0328] In Table 7, log2ZoTbWidth and log2ZoTbHeight represent the logarithmic values of the width and height of the top-left region, which can contain the last valid coefficients, with a base of 2 (base-2), respectively. The check on the value of mts_idx[x0][y0] can be omitted.
[0329] Additionally, as shown in the example, when determining log2ZoTbWidth and log2ZoTbHeight in the residual encoding, a condition for checking sps_mts_enable_flag can be added.
[0330] If a valid coefficient exists at the zeroing position when applying LFNST, the variable LfnstZeroOutSigCoeffFlag in Table 5 is 0; otherwise, the variable is 1. The variable LfnstZeroOutSigCoeffFlag can be set according to several conditions shown in Table 7.
[0331] According to the example, when all the last valid coefficients are located at the DC position of a transform block with a coded block flag of 1 (CBF equals 0 if at least one valid coefficient exists in the corresponding block, otherwise CBF equals 0), the variable LfnstDcOnly in Table 5 is equal to 1; otherwise, it is equal to 0. Specifically, in the case of dual-tree luma, the position of the last valid coefficient is checked relative to one luma transform block, while in the case of dual-tree chroma, the position of the last valid coefficient is checked relative to both the Cb transform block and the Cr transform block. In the case of single-tree, the position of the last valid coefficient can be checked relative to the luma, Cb, and Cr transform blocks.
[0332] In Table 5, MtsZeroOutSigCoeffFlag is initially set to 1, and this value can be changed in the residual coding in Table 7. When a valid coefficient exists in the region to be filled with 0 by zeroing (LastSignificantCoeffX>15| |LastSignificantCoeffY>15), the value of the variable MtsZeroOutSigCoeffFlag changes from 1 to 0. In this case, no signal is sent to the MTS index, as shown in Table 5.
[0333] Furthermore, as shown in Table 5, when tu_cbf_luma[x0][y0] is 0, the encoding of mts_idx[x0][y0] can be omitted. That is, when the CBF value of the luminance component is 0, since no transformation is applied, there is no need to signal the MTS index, so the MTS index encoding can be omitted.
[0334] Based on the example, the above technical features can be implemented using another conditional syntax. For example, after executing MTS, a variable indicating whether a valid coefficient exists in the region outside the DC region of the current block can be derived, and when this variable indicates the existence of a valid coefficient in the region outside the DC region, the MTS index can be signaled. That is, the existence of a valid coefficient in the region outside the DC region of the current block indicates that the value of tu_cbf_luma[x0][y0] is 1, and in this case, the MTS index can be signaled.
[0335] This variable can be represented as MtsDcOnly, and after MtsDcOnly is initially set to 1 at the coding unit level, its value is changed to 0 at the residual coding level when it is determined that there are valid coefficients in the current block outside the DC region. When MtsDcOnly is 0, the image information can be configured to signal the MTS index.
[0336] When tu_cbf_luma[x0][y0] is 0, the initial value of the variable MtsDcOnly is maintained at 1 because the residual coding syntax is not invoked at the transform unit level in Table 6. In this case, since the variable MtsDcOnly is not changed to 0, the image information can be configured not to signal the MTS index. That is, neither parsing nor signaling the MTS index is performed.
[0337] Furthermore, the decoding device can determine the color index cIdx of the transform coefficients to derive the variable MtsZeroOutSigCoeffFlag in Table 7. A color index cIdx of 0 indicates the luminance component.
[0338] According to the example, since MTS can be applied only to the luminance component of the current block, the decoding device can determine whether the color index is luminance when deriving the variable MtsZeroOutSigCoeffFlag used to determine whether to parse the MTS index.
[0339] The variable `MtsZeroOutSigCoeffFlag` indicates whether zeroing is performed when MTS is applied. It indicates whether transform coefficients exist in the region outside the top-left region of the last valid coefficients (i.e., the region outside the top-left 16×16 region) after MTS is performed due to zeroing. The variable `MtsZeroOutSigCoeffFlag` is initially set to 1 at the coding unit level, as shown in Table 5 (`MtsZeroOutSigCoeffFlag=1`), and its value can be changed from 1 to 0 at the residual coding level when transform coefficients exist in the region outside the 16×16 region, as shown in Table 7 (`MtsZeroOutSigCoeffFlag=0`). When the value of the variable `MtsZeroOutSigCoeffFlag` is 0, no signal is sent to the MTS index.
[0340] As shown in Table 7, at the residual coding level, the non-zero region, in which non-zero transform coefficients may exist, can be set depending on whether the zeroing of the accompanying MTS is performed. Even in this case, when the color index (cIdx) is 0, the non-zero region can be set to the upper left 16×16 region of the current block.
[0341] Thus, when deriving the variable used to determine whether to resolve the MTS index, it is necessary to determine whether the color component is luminance or chrominance. However, since LFNST can be applied to both the luminance and chrominance components of the current block, the color component is uncertain when deriving the variable used to determine whether to resolve the LFNST index.
[0342] For example, Table 5 shows the variable LfnstZeroOutSigCoeffFlag, which can indicate whether zeroing is performed when LFNST is applied. The variable LfnstZeroOutSigCoeffFlag indicates whether a valid coefficient exists in a second region of the current block, excluding the first region in the upper left. This value is initially set to 1 and can be changed to 0 when a valid coefficient exists in the second region. The LFNST index can only be resolved if the initially set value of the variable LfnstZeroOutSigCoeffFlag remains 1. When determining and deriving whether the value of the variable LfnstZeroOutSigCoeffFlag is 1, the color index of the current block is uncertain because LFNST can be applied to both the luma and chroma components of the current block.
[0343] As shown in the example, prediction can be based on palette encoding. Palette encoding is a useful technique for representing blocks that consist of a small number of unique color values. Instead of applying prediction and transformation to the block, in palette mode, an index representing the value of each sample is signaled. To decode a block encoded using palette mode, the decoder needs to decode both the palette entry and the index. Palette entries can be represented by a palette table and can be encoded using a palette table encoding tool.
[0344] For example, when selecting a palette mode, information about the palette table can be signaled. The palette table can include an index corresponding to each pixel. The palette table can be constructed from pixel values used in previous blocks to create a palette prediction table. For example, previously used pixel values can be stored in a specific buffer (palette predictor), and palette predictor information (palette_predictor_run) for constructing the current palette can be received from the buffer. That is, the palette predictor can include data indicating at least a portion of the palette index map used for the current block. When the number of palette entries used to represent the current block is insufficient to construct a palette prediction entry from the palette predictor, pixel information for the current palette entry can be sent separately.
[0345] Palette mode is signaled at the CU level and is typically used when most pixels in the CU can be represented by a representative set of pixel values. That is, in palette mode, samples in the CU can be represented as a representative set of pixel values. Such a set can be called a palette. When a sample has a value close to a pixel value in the palette, a palette index (palette_idx_idc) or information (run_copy_flag, copy_above_palette_indices_flag) indicating the index corresponding to the pixel value in the palette can be signaled. For samples with pixel values other than those in the palette, the samples can be marked with an escape symbol, and the quantization of the sample value can be signaled directly. In this document, pixels or pixel values can be referred to as samples or sample values.
[0346] To decode a block encoded in palette mode, the decoder needs palette entry information and palette index information. When a palette index corresponds to an escape symbol, the (quantized) escape value can be signaled as an additional component. Additionally, the encoder must derive the appropriate palette for the CU and deliver it to the decoder.
[0347] For efficient encoding of palette items, a palette predictor can be maintained. The maximum size of the palette predictor and the palette can be signaled in the SPS. Alternatively, the palette predictor and the maximum palette size can be predefined. For example, depending on whether the current block is single-tree or dual-tree, the palette predictor and the maximum palette size can be defined as 31 and 15 respectively. In the VVC standard, an `sps_palette_enabled_flag` indicating whether palette mode is enabled can be sent. Then, a `pred_mode_plt_coding` flag indicating whether the current coding unit is encoded in palette mode can be sent. The palette predictor can be initialized at the beginning of each brick or stripe.
[0348] For each item in the palette predictor, a reuse flag can be signaled to indicate whether it is part of the current palette. The reuse flag can be sent using a run-length encoding of 0. Then, the number of new palette items can be signaled using zero-order exponent Golomb encoding. Finally, the component values of the new palette items can be signaled. After encoding the current CU, the palette predictor can be updated using the current palette, and items from the old palette predictor that are not reused in the current palette can be added to the end of the new palette predictor until the maximum allowed size (palette stuffing) is reached.
[0349] The index can be encoded using both horizontal and vertical traversal scans to encode a palette index map. The scan order can be explicitly signaled from the bitstream using flag information such as `palette_transpose_flag`. In the following text, for simplicity, this document will primarily describe horizontal scans. However, this can also be applied to vertical scans.
[0350] Figure 17 Examples are shown to illustrate the horizontal and vertical traversal scanning methods used to encode palette index maps.
[0351] Figure 17 (a) shows an example of using a horizontal traversal scan to encode a palette index map, and Figure 17 (b) shows an example of using vertical traversal scans to encode a palette index map.
[0352] like Figure 17 As shown in (a), when using a horizontal scan, the palette index can be encoded by performing a horizontal scan from the samples in the first row (top row) of the current block (i.e., the current CU) to the samples in the last row (bottom row).
[0353] like Figure 17 As shown in (b), when using a vertical scan, the palette index can be encoded by performing a vertical scan from the samples in the first column (leftmost column) of the current block (i.e., the current CU) to the samples in the last column (rightmost column).
[0354] When applying LFNST to a chroma transform block according to the example, you need to refer to information about juxtaposed luminance transform blocks.
[0355] The existing standard texts for the relevant sections are shown in the table below.
[0356] [Table 8]
[0357]
[0358] As shown in Table 8, when the current intra-prediction mode is CCLM mode, the value of the variable predModeIntra of the chroma transform block is determined by taking the intra-prediction mode value of the co-bit chroma transform block (the part indicated in italics). The intra-prediction mode value (predModeIntra value) of the luma transform block can then be used to determine the LFNST set.
[0359] However, the variables nTbW and nTbH, which are input as input values to this transformation process, represent the width and height of the current transform block. Therefore, when the current block is a luminance transform block, the variables nTbW and nTbH can represent the width and height of the luminance transform block, while when the current block is a chrominance transform block, the variables nTbW and nTbH represent the width and height of the chrominance transform block.
[0360] Here, the variables nTbW and nTbH in the italicized portion of Table 8 represent the width and height of the chroma transform block, which do not reflect the color format and therefore do not accurately indicate the reference position of the luminance transform block corresponding to the chroma transform block. Therefore, the italicized portion of Table 8 can be modified as shown in the table below.
[0361] [Table 9]
[0362]
[0363] As shown in Table 9, nTbW and nTbH are changed to (nTbW*SubWidthC) / 2 and (nTbH*SubHeightC) / 2, respectively. xTbY and yTbY can represent the brightness position in the current image (relative to the top-left brightness sample of the current image, the top-left sample of the current brightness transform block), while nTbW and nTbH can represent the width and height of the currently encoded transform block (the variable nTbW specifies the width of the current transform block, and the variable nTbH specifies the height of the current transform block).
[0364] When the currently encoded transform block is a chroma (Cb or Cr) transform block, nTbW and nTbH are the width and height of the chroma transform block, respectively. Therefore, when the currently encoded transform block is a chroma transform block (cIdx>0), the width and height of the luma transform block are needed to obtain the reference position of the juxtaposed luma transform block. In Table 9, SubWidthC and SubHeightC are values set according to the color format (e.g., 4:2:0, 4:2:2, or 4:4:4), specifically, the width ratio and height ratio between the luma and chroma components, respectively (see Table 10 below). Therefore, in the case of a chroma transform block, (nTbW*SubWidthC) and (nTbH*SubHeightC) can be the width and height relative to the juxtaposed luma transform block, respectively.
[0365] Therefore, xTbY+(nTbW*SubWidthC) / 2 and yTbY+(nTbH*SubHeightC) / 2 represent the values of the center position in the juxtaposed brightness transformation blocks based on the top-left position of the current image, and thus accurately indicate the juxtaposed brightness transformation blocks.
[0366] [Table 10]
[0367]
[0368] In Table 9, the variable `predModeIntra` represents the intra-prediction mode value. A value of `predModeIntra` equal to `INTRA_LT_CCLM`, `INTRA_L_CCLM`, or `INTRA_T_CCLM` indicates that the current transform block is a chroma transform block. According to the example, in the current VVC standard, `INTRA_LT_CCLM`, `INTRA_L_CCLM`, and `INTRA_T_CCLM` correspond to mode values 81, 82, and 83 in the intra-prediction mode values, respectively. Therefore, as shown in Table 9, the values of `xTbY+(nTbW*SubWidthC) / 2` and `yTbY+(nTbH*SubHeightC) / 2` are needed to obtain the reference position of the juxtaposed luma transform blocks.
[0369] As shown in Table 9, the value of predModeIntra is updated based on the variables intra_mip_flag[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] and CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2].
[0370] `intra_mip_flag` is a variable indicating whether the current transform block (or coding unit) is encoded using a matrix-based intra-prediction (MIP) method, and `intra_mip_flag[x][y]` is a flag value indicating whether MIP is applied to the position corresponding to the coordinates (x,y) based on the luma component when the top-left position of the current image is defined as (0,0). The x and y coordinates increase from left to right and from top to bottom, respectively, and when the flag indicating whether MIP is applied is 1, it indicates that MIP is applied. When the flag indicating whether MIP is applied is 0, it indicates that MIP is not applied. MIP can be applied only to luma blocks.
[0371] According to the modified part of Table 9, when the value of intra_mip_flag[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] in the juxtaposed brightness transformation block is 1, the value of predModeIntra is set to planar mode (INTRA_PLANAR).
[0372] The value of the variable CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the prediction mode value corresponding to the coordinates (xTbY+(nTbW*SubWidthC) / 2, yTbY+(nTbH*SubHeightC) / 2) when the top-left position of the current image for the luminance component is defined as (0,0). The prediction mode value can have MODE_INTRA, MODE_IBC, MODE_PLT, and MODE_INTER values representing intra-frame prediction mode, intra-block copy (IBC) prediction mode, palette (PLT) coding mode, and inter-frame prediction mode, respectively. According to Table 17, when the value of variable CuPredMode[0][xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] is MODE_IBC or MODE_PLT, the value of variable predModeIntra is set to DC mode. In all other cases, the value of variable predModeIntra is set to IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] (the intra-prediction mode value corresponding to the center position in the juxtaposed luma transform block).
[0373] As shown in the example table below, considering whether to perform wide-angle intra-frame prediction, the value of the variable predModeIntra can be updated again based on the updated predModeIntra value in Table 9.
[0374] [Table 11]
[0375]
[0376] In the mapping process shown in Table 11, the input values of predModeIntra, nTbW, and nTbH are the same as the updated variables predModeIntra and the values of nTbW and nTbH referenced in Table 9.
[0377] In Table 11, nCbW and nCbH represent the width and height of the coding block corresponding to the transform block, respectively. The variable IntraSubPartitionsSplitType indicates whether ISP mode is applied. IntraSubPartitionsSplitType equal to ISP_NO_SPLIT indicates that the coding unit is not partitioned by ISP (i.e., ISP mode is not applied). IntraSubPartitionsSplitType not equal to ISP_NO_SPLIT indicates that ISP mode is applied, and therefore the coding unit is partitioned into two or four partition blocks. In Table 11, cIdx is the index indicating the color component. A cIdx value of 0 indicates a luma block, while a cIdx value not equal to 0 indicates a chroma block. The predModeIntra value output by the mapping process in Table 11 is an updated value considering whether Wide-Angle Intra Prediction (WAIP) mode is applied.
[0378] For the predModeIntra value updated via Table 11, the LFNST set can be determined using the mapping relationship shown in the table below.
[0379] [Table 12]
[0380]
[0381] In the table above, `lfnstTrSetIdx` indicates the index of the LFNST set and has a value from 0 to 3, indicating that a total of four LFNST sets are configured. Each LFNST set can include two transform kernels: an LFNST kernel (based on the forward direction depending on the region where the LFNST is applied; the transform kernel can be a 16×16 matrix or a 16×48 matrix), and the transform kernel to be applied can be specified by the signaling of the LFNST index. Additionally, whether to apply LFNST can be specified via the LFNST index. In the current VVC standard, the LFNST index can have values of 0, 1, and 2, where 0 indicates that LFNST is not applied, and 1 and 2 indicate two transform kernels respectively.
[0382] The following figures are provided to illustrate specific examples of this disclosure. Since the specific names of the devices or signals / messages / fields illustrated in the figures are provided for illustrative purposes only, the technical features of this disclosure are not limited to the specific names used in the following figures.
[0383] Figure 18 This is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.
[0384] Figure 18 Each process disclosed in the document is based on a reference. Figures 4 to 17Some details described. Therefore, they will be omitted or will be illustrated with references. Figures 3 to 17 The description of overlapping details is a specific detail.
[0385] According to the implementation method, the decoding device 300 can obtain intra-frame prediction mode information and LFNST index from the bit stream (S1810).
[0386] Intra-prediction mode information may include the intra-prediction modes of the current block’s neighboring blocks (e.g., the left neighboring block and / or the upper neighboring block), as well as the MPM index indicating one of the MPM candidates in the list of most probable modes (MPM) derived based on additional candidate modes, or the remaining intra-prediction mode information indicating one of the remaining intra-prediction modes not included in the MPM candidates.
[0387] Additionally, intra-frame mode information may include the flag information sps_cclm_enabled_flag indicating whether CCLM is applied to the current block and the information intra_chroma_pred_mode regarding the intra-frame prediction mode of the chroma components.
[0388] The LFNST index information is received as syntax information, and the syntax information is received as a binary bin string containing 0s and 1s.
[0389] According to this embodiment, the syntax elements of the LFNST index can indicate whether the inverse LFNST or the inverse inseparable transformation is applied, as well as any one of the transformation kernel matrices included in the transformation set. When the transformation set includes two transformation kernel matrices, the syntax elements of the transformation index can have three values.
[0390] In other words, according to the implementation, the values of the syntax elements of the LFNST index can include: 0, which indicates that the inverse LFNST is not applied to the target block; 1, which indicates the first transformation kernel matrix in the transformation kernel matrix; and 2, which indicates the second transformation kernel matrix in the transformation kernel matrix.
[0391] Decoding device 300 can decode information about the quantization transform coefficients of the current block from the bitstream, and can deduce the quantization transform coefficients of the target block based on the information about the quantization transform coefficients of the current block. Information about the quantization transform coefficients of the target block can be included in the Sequence Parameter Set (SPS) or stripe header, and may include at least one of the following: information about whether RST is applied, information about the reduction factor, information about the minimum transform size for applying RST, information about the maximum transform size for applying RST, the inverse RST size, and information about the transform index indicating any one of the transform kernel matrices included in the transform set.
[0392] The decoding device 300 can derive the transform coefficients by dequantizing the residual information about the current block (i.e., the quantized transform coefficients), and can arrange the derived transform coefficients in a predetermined scan order.
[0393] Specifically, the derived transform coefficients can be arranged in 4×4 blocks according to the inverse diagonal scan order, and the transform coefficients in the 4×4 blocks can also be arranged according to the inverse diagonal scan order. In other words, the dequantized transform coefficients can be arranged according to the inverse scan order applied in video codecs (such as in VVC or HEVC).
[0394] The transform coefficients derived from the residual information can be dequantized transform coefficients as described above, or they can be quantized transform coefficients. In other words, the transform coefficients can be any data used to check whether there is non-zero data in the current block, and are unrelated to quantization.
[0395] The decoding device can update the intra-prediction mode of the chroma block to intra-DC mode based on the fact that the intra-prediction mode of the chroma block is CCLM mode and the intra-prediction mode of the corresponding luma block is palette mode (S1820).
[0396] The decoding device can deduce the intra-prediction mode of a chroma block as CCLM mode based on intra-prediction mode information. For example, the decoding device can receive information about the intra-prediction mode of the current chroma block through a bitstream, and can deduce the intra-prediction mode of the current chroma block as CCLM mode based on the intra-prediction mode information.
[0397] CCLM modes can include top-left CCLM mode, top CCLM mode, or left CCLM mode.
[0398] As described above, the decoding device can derive residual samples by applying LFNST as an inseparable transform or MTS as a separable transform, and these transforms can be performed based on the LFNST index indicating the LFNST kernel (i.e., the LFNST matrix) and the MTS index indicating the MTS kernel, respectively.
[0399] For LFNST, it is necessary to determine the LFNST set, and this LFNST set needs to have a mapping relationship with the intra-prediction mode of the current block.
[0400] The decoding device can update the intra-prediction mode of the chroma block based on the intra-prediction mode of the luma block corresponding to the chroma block, for use in the inverse LFNST of the chroma block.
[0401] Based on the example, the updated intra-prediction mode can be derived as an intra-prediction mode corresponding to a specific position in the luma block, and that specific position can be set based on the color format of the chroma block.
[0402] A specific position can be the center of the luminance block and can be represented by ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)).
[0403] In the central position, xTbY and yTbY represent the top-left coordinates of the luma block, i.e., the top-left position in the luma sample reference of the current transform block. nTbW and nTbH represent the width and height of the chroma block, while SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)) represents the central position of the luma transform block, while IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the intra-prediction mode of the luma block at that position.
[0404] SubWidthC and SubHeightC can be derived as shown in Table 10. That is, when the color format is 4:2:0, SubWidthC and SubHeightC are both 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0405] As shown in Table 9, in order to specify the specific position of the luminance block corresponding to the chroma block independently of the color format, the color format is reflected in the variable indicating the specific position.
[0406] As described above, when the intra-prediction mode corresponding to a specific position of the luma block is the palette mode, the decoding device can update the intra-prediction mode to the intra-DC mode.
[0407] Palette encoding is a useful technique for representing blocks that comprise a small number of unique color values. Instead of applying prediction and transformation to the blocks, in palette mode, an index representing the value of each sample is signaled. To decode a block encoded using palette mode, the decoder needs to decode both the palette entries and their indices. Palette entries can be represented by a palette table and can be encoded using a palette table encoding tool.
[0408] Alternatively, according to the example, when the intra-prediction mode corresponding to a particular location is the intra-block copy (IBC) mode, the decoding device may set the updated intra-prediction mode to the intra-DC mode.
[0409] IBC essentially performs predictions within the current image, but it can be performed similarly to inter-frame prediction, except that the reference block is derived within the current image. In other words, IBC can use at least one of the inter-frame prediction techniques described in this document.
[0410] IBC prediction mode, or palette mode, can be used to encode content images / videos, including those from games, such as Screen Content Coding (SCC). IBC essentially performs prediction within the current frame, but can be performed similarly to inter-frame prediction, except that a reference block is derived within the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure. Palette mode can be considered an example of intra-frame coding or intra-frame prediction. When applying palette mode, the values of samples in the image can be signaled based on information about the palette table and palette indices.
[0411] Alternatively, according to an example, when the intra-prediction mode of the luma block corresponding to a particular location is matrix-based intra-prediction (hereinafter, "MIP") mode, the decoding device may set the updated intra-prediction mode to intra-plane mode.
[0412] The MIP mode can be called affine linear weighted intra-frame prediction (ALWIP) or matrix weighted intra-frame prediction (MWIP). When MIP is applied to the current block, the predicted samples for the current block can be derived by i) using neighboring reference samples that have already undergone averaging, ii) performing matrix-vector multiplication, and iii) performing further horizontal / vertical interpolation.
[0413] In summary, when the intra-prediction mode of the central position is MIP mode, IBC mode, or palette mode, the intra-prediction mode of the chroma block can be updated to a specific mode, such as intra-plane mode or intra-DC mode.
[0414] When the intra-prediction mode of the central position is not MIP mode, IBC mode, or palette mode, the intra-prediction mode of the chroma block can be updated to the intra-prediction mode of the luma block at the central position in order to reflect the correlation between the chroma block and the luma block.
[0415] The decoding device can determine the LFNST set including the LFNST matrix based on the updated intra-frame prediction mode (S1830), and can derive the transform coefficients by performing LFNST on the chroma block based on the LFNST matrix derived from the LFNST set (S1840).
[0416] You can select any one of multiple LFNST matrices based on the LFNST set and LFNST index.
[0417] As shown in Table 12, the LFNST transform set is derived based on the intra-prediction mode, and 81 to 83, which indicate the CCLM mode in the intra-prediction mode, are omitted because the LFNST transform set is derived using the intra-mode values of the corresponding luma blocks in the CCLM mode.
[0418] As shown in Table 12, as an example, any one of the four LFNST sets can be determined based on the intra-prediction mode of the current block, and the LFNST set to be applied to the current chroma block can also be determined.
[0419] The decoding device can derive the modified transform coefficients of the current chroma block by performing an inverse RST (e.g., inverse LFNST) by applying the LFNST matrix to the dequantized transform coefficients.
[0420] The decoding device can derive the residual samples from the transform coefficients through the inverse first transform (S1850), and when the current block is a chroma block, it can derive the residual samples of the chroma block based on the transform coefficients. MTS can be used for the inverse first transform.
[0421] In addition, the decoding device can generate reconstructed samples based on the residual samples of the current block and the predicted samples of the current block (S1860).
[0422] The following figures are provided to illustrate specific examples of this disclosure. Since the specific names of the devices or signals / messages / fields illustrated in the figures are for illustrative purposes only, the technical features of this disclosure are not limited to the specific names used in the following figures.
[0423] Figure 19 This is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present disclosure.
[0424] Figure 19 Each process disclosed in the document is based on a reference. Figures 4 to 17 Some details described. Therefore, they will be omitted or will be illustrated with references. Figure 2 and Figures 4 to 17 The description of overlapping details is a specific detail.
[0425] According to the implementation method, the encoding device 200 can derive the prediction sample of the chroma block based on the fact that the intra-frame prediction mode of the chroma block is CCLM mode (S1910).
[0426] The encoding device can first deduce the intra-prediction mode of the chroma block to CCLM mode.
[0427] For example, the coding device can determine the intra-prediction mode for the current chroma block based on the rate-distortion (RD) cost (or RDO). Here, the RD cost can be derived based on the sum of absolute differences (SAD). The coding device can then determine the CCLM mode as the intra-prediction mode for the current chroma block based on the RD cost.
[0428] CCLM modes can include top-left CCLM mode, top CCLM mode, or left CCLM mode.
[0429] The encoding device can encode information about the intra-prediction mode of the current chroma block and can signal this information via a bitstream. Prediction-related information about the current chroma block can include information about the intra-prediction mode.
[0430] According to the implementation method, the encoding device can derive the residual samples of the chroma block based on the predicted samples (S1920).
[0431] According to the implementation method, the encoding device can derive the transformation coefficients of the chroma block based on a single transformation of the residual sample (S1930).
[0432] A single transformation can be performed using multiple transformation kernels. In this case, the transformation kernel can be selected based on the intra-frame prediction mode.
[0433] The encoding device can update the intra-prediction mode of the chroma block to the intra-DC mode of the LFNST for the chroma block (S1940) based on the fact that the intra-prediction mode of the chroma block is CCLM mode and the intra-prediction mode of the corresponding luma block is palette mode.
[0434] As shown in Table 9, the encoding device can update the CCLM mode of the chroma block based on the intra-prediction mode of the luma block corresponding to the chroma block (when predModeIntra is equal to INTRA_LT_CCLM, INTRA_L_CCLM or INTRA_T_CCLM, the predModeIntra is derived as follows:).
[0435] According to the example, the updated intra-prediction mode can be derived to correspond to an intra-prediction mode for a specific location in the luma block, and that specific location can be set based on the color format of the chroma block.
[0436] A specific position can be the center of the luminance block and can be represented by ((xTbY+(nTbW*SubWidthC) / 2),(yTbY+(nTbH*SubHeightC) / 2)).
[0437] In the central position, xTbY and yTbY represent the top-left coordinates of the luma block, i.e., the top-left position in the luma sample reference of the current transform block. nTbW and nTbH represent the width and height of the chroma block, while SubWidthC and SubHeightC correspond to variables corresponding to the color format. ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)) represents the central position of the luma transform block, while IntraPredModeY[xTbY+(nTbW*SubWidthC) / 2][yTbY+(nTbH*SubHeightC) / 2] represents the intra-prediction mode of the luma block at that position.
[0438] SubWidthC and SubHeightC can be derived as shown in Table 10. That is, when the color format is 4:2:0, SubWidthC and SubHeightC are both 2, and when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
[0439] As shown in Table 9, in order to specify the specific position of the luminance block corresponding to the chroma block independently of the color format, the color format is reflected in the variable indicating the specific position.
[0440] As described above, when the intra-prediction mode of the luma block corresponding to a specific location is palette mode, the encoding device can update the intra-prediction mode to intra-DC mode. Palette encoding is a useful technique for representing blocks that comprise a small number of unique color values. Instead of applying prediction and transformation to the block, in palette mode, an index representing the value of each sample is signaled. To decode a block encoded using palette mode, the decoder needs to decode the palette entries and indices. Palette entries can be represented by a palette table and can be encoded using a palette table encoding tool.
[0441] Alternatively, according to the example, when the intra prediction mode corresponding to a particular location is the intra block copy (IBC) mode, the encoding device may set the updated intra prediction mode to the intra DC mode.
[0442] IBC essentially performs predictions within the current image, but it can be performed similarly to inter-frame prediction, except that it derives a reference block within the current image. In other words, IBC can use at least one of the inter-frame prediction techniques described in this document.
[0443] IBC prediction mode, or palette mode, can be used to encode content images / videos, including those from games, such as Screen Content Coding (SCC). IBC essentially performs prediction within the current frame, but can be performed similarly to inter-frame prediction, except that a reference block is derived within the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure. Palette mode can be considered an example of intra-frame coding or intra-frame prediction. When applying palette mode, the values of samples in the image can be signaled based on information about the palette table and palette indices.
[0444] Alternatively, according to an example, when the intra-prediction mode of the luma block corresponding to a particular location is matrix-based intra-prediction (hereinafter, "MIP") mode, the encoding device may set the updated intra-prediction mode to intra-plane mode.
[0445] The MIP mode can be called affine linear weighted intra-frame prediction (ALWIP) or matrix weighted intra-frame prediction (MWIP). When MIP is applied to the current block, the predicted samples for the current block can be derived by i) using neighboring reference samples that have already undergone averaging, ii) performing matrix-vector multiplication, and iii) performing further horizontal / vertical interpolation.
[0446] In summary, when the intra-prediction mode of the central position is MIP mode, IBC mode, or palette mode, the intra-prediction mode of the chroma block can be updated to a specific mode, such as intra-plane mode or intra-DC mode.
[0447] When the intra-prediction mode of the central position is not MIP mode, IBC mode, or palette mode, the intra-prediction mode of the chroma block can be updated to the intra-prediction mode of the luma block at the central position to reflect the correlation between the chroma block and the luma block.
[0448] The coding device can determine the LFNST set including the LFNST matrix based on the updated intra-prediction mode (S1950), and can derive the modified transform coefficients by performing LFNST on the chroma block based on the residual samples and the LFNST matrix (S1960).
[0449] The coding device can determine the transform set based on the mapping relationship according to the intra-prediction mode applied to the current block, and can perform LFNST, i.e., inseparable transform, based on either of the two LFNST matrices contained in the transform set.
[0450] As described above, multiple transform sets can be determined based on the intra-prediction mode of the transform block to be transformed. The matrix applied to LFNST is the transpose of the matrix used in inverse LFNST.
[0451] In one example, the LFNST matrix can be a non-square matrix with fewer rows than columns.
[0452] The encoding device can derive the quantized transform coefficients by performing quantization based on the modified transform coefficients of the current chroma block, and can encode and output image information including information about the quantized transform coefficients, information about the intra-frame prediction mode, and LFNST index indicating the LFNST matrix (S1970).
[0453] Specifically, the encoding device 200 can generate information about the quantization transform coefficients and can encode the generated information about the quantization transform coefficients.
[0454] In one example, information about the quantization transform coefficients may include at least one of the following: information about whether LFNST is applied, information about the reduction factor, information about the minimum transform size for applying LFNST, and information about the maximum transform size for applying LFNST.
[0455] The encoding device can encode the flag information indicating whether CCLM is applied to the current block as sps_cclm_enabled_flag and the information about the intra-prediction mode of the chroma components as intra_chroma_pred_mode into information about the intra-mode.
[0456] Information about the CCLM mode as intra_chroma_pred_mode can indicate top-left CCLM mode, top CCLM mode, or left CCLM mode.
[0457] In this disclosure, at least one of quantization / dequantization and / or transformation / inverse transformation may be omitted. When quantization / dequantization is omitted, the quantization transformation coefficients may be referred to as transformation coefficients. When transformation / inverse transformation is omitted, the transformation coefficients may be referred to as coefficients or residual coefficients, or, for the sake of consistency, may still be referred to as transformation coefficients.
[0458] Furthermore, in this disclosure, quantization transform coefficients and transform coefficients can be referred to as transform coefficients and scaling transform coefficients, respectively. In this case, residual information can include information about the transform coefficients, and this information can be signaled via residual coding syntax. Transform coefficients can be derived based on residual information (or information about transform coefficients), and scaling transform coefficients can be derived through the inverse transform (scaling) of the transform coefficients. Residual samples can be derived based on the inverse transform (scaling) of the scaling transform coefficients. These details can also be applied / expressed in other parts of this disclosure.
[0459] In the above embodiments, the method is explained based on a flowchart using a series of steps or blocks. However, this disclosure is not limited to the order of the steps, and a step may be performed in a different order or sequence than described above, or a step may be performed concurrently with other steps. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps in the flowchart may be deleted without affecting the scope of this disclosure.
[0460] The methods described above according to this disclosure can be implemented in software form, and the encoding and / or decoding devices according to this disclosure can be included in devices for image processing such as televisions, computers, smartphones, set-top boxes, and display devices.
[0461] When the embodiments of this disclosure are implemented by software, the above methods can be implemented as modules (steps, functions, etc.) for performing the above functions. These modules can be stored in memory and can be executed by a processor. The memory can be internal or external to the processor and can be connected to the processor in various well-known ways. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0462] Furthermore, the decoding and encoding devices using this disclosure can include multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices (such as video communication), mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, and medical video devices, and can be used to process video signals or data signals. For example, over-the-top (OTT) video devices can include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0463] Furthermore, the processing methods applied to this disclosure can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having the data structure according to this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include various storage devices and distributed storage devices for storing computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of a carrier wave (e.g., transmission over the Internet). Furthermore, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks. Additionally, embodiments of this disclosure can be implemented as computer program products by program code, and the program code can be executed on a computer according to embodiments of this disclosure. The program code can be stored on a computer-readable carrier.
[0464] Figure 20 The structure of a content streaming system applying this disclosure is illustrated.
[0465] Furthermore, the content streaming system using this disclosure can generally include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0466] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then sends it to a streaming server. As another example, in cases where the multimedia input device, such as a smartphone, camera, or camcorder, directly generates the bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method disclosed herein. Furthermore, the streaming server can temporarily store the bitstream during the sending or receiving process.
[0467] The streaming server sends multimedia data to the user's device via a web server based on the user's request. The web server acts as a tool to notify the user of available services. When a user requests a desired service, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this context, the content streaming system may include a separate control server, which in this case controls the commands / responses between the corresponding devices within the content streaming system.
[0468] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0469] For example, user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, board-type PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. The servers in the content streaming system can operate as distributed servers, and in this case, data received by each server can be processed in a distributed manner.
[0470] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims can be combined to be implemented or performed in a device, and the technical features of the device claims can be combined to be implemented or performed in a method. Furthermore, the technical features of the method claims and the device claims can be combined to be implemented or performed in a device, and the technical features of the method claims and the device claims can be combined to be implemented or performed in a method.
Claims
1. An image decoding device, the image decoding device comprising: Memory; as well as At least one processor connected to the memory, wherein the at least one processor is configured to: Obtain intra-frame prediction mode information and LFNST index from the bitstream; The intra-prediction mode based on chroma blocks is the cross-component linear model (CCLM) mode, and the prediction mode of the luma block corresponding to the chroma block is the palette mode, so as to update the intra-prediction mode of the chroma block to the intra-DC mode. The LFNST set, including the LFNST matrix, is determined based on the updated intra-frame prediction mode; and LFNST is performed on the chroma block based on the LFNST matrix derived from the LFNST set. Wherein, the color palette mode is the prediction mode corresponding to a specific position in the brightness block, and The specific position is set to ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)). Where xTbY and yTbY represent the top-left coordinates of the brightness block. Wherein, nTbW and nTbH represent the width and height of the chromaticity block, and Here, SubWidthC and SubHeightC represent variables corresponding to the color format.
2. The image decoding device according to claim 1, wherein, The specific location is set based on the color format of the chroma block.
3. The image decoding device according to claim 2, wherein, The specific location is the center of the brightness block.
4. The image decoding device according to claim 1, wherein, When the color format is 4:2:0, SubWidthC and SubHeightC are both 2, and Wherein, when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
5. The image decoding device according to claim 1, wherein, When the intra-prediction type corresponding to the specific position is MIP mode, the intra-prediction mode of the chroma block is updated to intra-plane mode.
6. The image decoding device according to claim 1, wherein, When the prediction mode corresponding to the specific location is IBC mode, the intra-prediction mode of the chroma block is updated to intra-DC mode.
7. An image encoding device, the image encoding device comprising: Memory; as well as At least one processor connected to the memory, wherein the at least one processor is configured to: The predicted samples of the chroma block are derived based on the fact that the intra-frame prediction mode for the chroma block is the cross-component linear model (CCLM) mode. Based on the predicted samples, residual samples for the chroma block are derived; The intra-prediction mode of the chroma block is updated to intra-DC mode based on the fact that the intra-prediction mode of the chroma block is the cross-component linear model (CCLM) mode and the prediction mode of the luma block corresponding to the chroma block is the palette mode. The LFNST set, including the LFNST matrix, is determined based on the updated intra-frame prediction mode; and LFNST is performed on the chroma block based on the residual samples and the LFNST matrix. Wherein, the color palette mode is the prediction mode corresponding to a specific position in the brightness block, and The specific position is set to ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)). Where xTbY and yTbY represent the top-left coordinates of the brightness block. Wherein, nTbW and nTbH represent the width and height of the chromaticity block, and Here, SubWidthC and SubHeightC represent variables corresponding to the color format.
8. The image encoding device according to claim 7, wherein, The specific location is set based on the color format of the chroma block.
9. The image encoding device according to claim 8, wherein, The specific location is the center of the brightness block.
10. The image encoding device according to claim 7, wherein, When the color format is 4:2:0, SubWidthC and SubHeightC are both 2, and Wherein, when the color format is 4:2:2, SubWidthC is 2 and SubHeightC is 1.
11. The image encoding device according to claim 7, wherein, When the intra-prediction type corresponding to the specific position is MIP mode, the intra-prediction mode of the chroma block is updated to intra-plane mode.
12. The image encoding device according to claim 7, wherein, When the prediction mode corresponding to the specific location is IBC mode, the intra-prediction mode of the chroma block is updated to intra-DC mode.
13. An apparatus for transmitting data relating to image information, the apparatus comprising: At least one processor is configured to generate a bitstream for the image information, wherein the bitstream is generated based on the following operations: deriving prediction samples for the chroma block based on the fact that the intra-prediction mode for the chroma block is a cross-component linear model (CCLM) mode; deriving residual samples for the chroma block based on the prediction samples; updating the intra-prediction mode of the chroma block to an intra-DC mode based on the fact that the intra-prediction mode of the chroma block is a CCLM mode and the prediction mode of the corresponding luma block is a palette mode; determining an LFNST set including an LFNST matrix based on the updated intra-prediction mode; performing LFNST on the chroma block based on the residual samples and the LFNST matrix; and encoding the residual information to generate the bitstream; and A transmitter configured to transmit the data comprising the bit stream. Wherein, the color palette mode is the prediction mode corresponding to a specific position in the brightness block, and The specific position is set to ((xTbY+(nTbW*SubWidthC) / 2), (yTbY+(nTbH*SubHeightC) / 2)). Where xTbY and yTbY represent the top-left coordinates of the brightness block. Wherein, nTbW and nTbH represent the width and height of the chromaticity block, and Here, SubWidthC and SubHeightC represent variables corresponding to the color format.