Decoding device, encoding device, and data transmitting device
Patent Information
- Application Number
- CN202310770648.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-06
- Filing Date
- 2019-12-05
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2039-12-05
AI Technical Summary
因此,当使用诸如传统有线/无线宽带线这样的介质来发送图像数据或者使用现有存储介质来存储图像/视频数据时,其传输成本和存储成本增加
[0019] According to this disclosure, the overall image/video compression efficiency can be increased.
Smart Images

Figure CN116684643B_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application No. 201980075783.0 (International Application No.: PCT / KR2019 / 017104, Application Date: December 5, 2019, Invention Title: Image Coding Method and Apparatus Based on Quadratic Transformation). Technical Field
[0002] This disclosure generally relates to image coding techniques, and more specifically, to a transform-based image coding method and apparatus in an image coding system. Background Technology
[0003] Today, the demand for high-resolution and high-quality images / videos, such as 4K, 8K, or even higher Ultra High Definition (UHD) images / videos, is constantly growing across various fields. As image / video data becomes higher resolution and higher quality, the amount of information or bits transmitted increases compared to traditional image data. Therefore, transmission and storage costs increase when using media such as traditional wired / wireless broadband lines to transmit image data or when using existing storage media to store image / video data.
[0004] In addition, there is increasing interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms, and broadcasting of images / videos with image characteristics that differ from real images such as game images is on the rise.
[0005] Therefore, there is a need for efficient image / video compression techniques to effectively compress, transmit, store, and reproduce information with high resolution and high quality images / videos that have the various characteristics described above. Summary of the Invention
[0006] Technical issues
[0007] The technical aspect of this disclosure is to provide methods and apparatus for increasing image coding efficiency.
[0008] Another aspect of this disclosure is to provide methods and apparatus for increasing conversion efficiency.
[0009] Another technical aspect of this disclosure is to provide a method and apparatus for increasing the efficiency of a quadratic transformation by encoding the transformation index.
[0010] Another technical aspect of this disclosure is to provide an image coding method and an image coding device based on Reduced Quadratic Transform (RST).
[0011] Another technical aspect of this disclosure is to provide an image coding method and image coding device based on transform sets and capable of increasing coding efficiency.
[0012] Technical solution
[0013] According to embodiments of this disclosure, an image decoding method performed by a decoding device is provided. The method may include: deriving quantized transform coefficients of a target block from a bitstream; deriving transform coefficients by dequantization based on the quantized transform coefficients of the target block; deriving modified transform coefficients based on an inverse reduced quadratic transform (RST) of the transform coefficients; deriving residual samples of the target block based on an inverse first transform of the modified transform coefficients; and generating a reconstructed image based on the residual samples of the target block, wherein the inverse RST may be performed based on a transform set and a transform kernel matrix selected from two transform kernel matrices included in each transform set, the transform set being determined based on a mapping relationship according to an intra-prediction mode applied to the target block, and the inverse RST may be performed based on a transform index relating to whether the inverse RST is applied and one of the transform kernel matrices included in the transform set.
[0014] According to another embodiment of this disclosure, a decoding apparatus for performing image decoding is provided. The decoding apparatus may include: an entropy decoder that derives information about the prediction and quantized transform coefficients of a target block from a bitstream; a predictor that generates prediction samples of the target block based on the information about the prediction; a dequantizer that derives transform coefficients by dequantization based on the quantized transform coefficients of the target block; an inverse transformer that includes an inverse reduced quadratic transform (RST) that derives modified transform coefficients based on an inverse RST of the transform coefficients, and an inverse first transform that derives residual samples of the target block based on an inverse first transform of the modified transform coefficients; and an adder that generates reconstructed samples based on the residual samples and the prediction samples, wherein the inverse RST can be performed based on a transform set and a transform kernel matrix selected from two transform kernel matrices included in each transform set, the transform set being determined based on a mapping relationship according to an intra-prediction mode applied to the target block, and the inverse RST can be performed based on a transform index relating to an indication of whether the inverse RST is applied and one of the transform kernel matrices included in the transform set.
[0015] According to another embodiment of this disclosure, an image coding method performed by an encoding device is provided. The method may include: deriving prediction samples based on an intra-prediction mode applied to a target block; deriving residual samples of the target block based on the prediction samples; deriving transform coefficients of the target block based on a first-order transform of the residual samples; deriving modified transform coefficients based on a reduced quadratic transform (RST) of the transform coefficients, the inverse RST being performed based on a transform set and a transform kernel matrix selected from two transform kernel matrices included in each transform set, the transform set being determined based on a mapping relationship according to the intra-prediction mode applied to the target block; deriving quantized transform coefficients by performing quantization based on the modified transform coefficients; and generating a transform index indicating whether the inverse RST is applied and a transform kernel matrix included in the transform set.
[0016] According to another embodiment of the present disclosure, a digital storage medium may be provided that stores image data including encoded image information generated according to an image encoding method performed by an encoding device.
[0017] According to another embodiment of this disclosure, a digital storage medium can be provided that stores image data including encoded image information to enable a decoding device to perform an image decoding method.
[0018] Technical effect
[0019] According to this disclosure, the overall image / video compression efficiency can be increased.
[0020] According to this disclosure, the efficiency of the secondary transformation can be increased by changing the encoding of the index.
[0021] According to this disclosure, image coding efficiency can be increased by performing image coding based on a transform set. Attached Figure Description
[0022] Figure 1 Examples of video / image coding systems to which this disclosure can be applied are illustrated schematically.
[0023] Figure 2 This is a diagram that schematically illustrates the configuration of a video / image encoding device to which this disclosure can be applied.
[0024] Figure 3 This is a diagram that schematically illustrates the configuration of a video / image decoding device to which this disclosure can be applied.
[0025] Figure 4 Multiple transformation techniques according to embodiments of the present disclosure are illustrated schematically.
[0026] Figure 5 Examples of directional intra-frame modes in 65 prediction directions are shown.
[0027] Figure 6 This is a diagram illustrating an embodiment of the RST according to the present disclosure.
[0028] Figure 7 This is a diagram illustrating the scanning order of transformation coefficients according to an embodiment of the present disclosure.
[0029] Figure 8 This is a flowchart illustrating the reverse RST process according to an embodiment of the present disclosure.
[0030] Figure 9 This is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.
[0031] Figure 10 This is a control flow diagram illustrating the reverse RST according to an embodiment of the present disclosure.
[0032] Figure 11 This is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present disclosure.
[0033] Figure 12 This is a control flow diagram illustrating an embodiment of RST according to the present disclosure.
[0034] Figure 13 The structure of a content streaming system applying this disclosure is illustrated. Detailed Implementation
[0035] While this disclosure may be readily modified and includes various embodiments, specific embodiments thereof have been illustrated by way of example in the accompanying drawings and will now be described in detail. However, this is not intended to limit this disclosure to the specific embodiments disclosed herein. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the technical concept of this disclosure. The singular form may include the plural form unless the context clearly indicates otherwise. Terms such as “comprising” and “having” are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and should therefore not be construed as pre-excluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0036] Furthermore, for ease of description of their different features and functions, the components in the accompanying drawings described herein are illustrated independently; however, this does not imply that each component is implemented by a separate piece of hardware or software. For example, any two or more of these components may be combined to form a single component, and any single component may be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of this disclosure, provided they do not depart from the spirit of this disclosure.
[0037] In the following description, preferred embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Furthermore, in the drawings, the same reference numerals are used for the same components, and repeated descriptions of the same components will be omitted.
[0038] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Video Coding Universal) standard (ITU-T Rec.H.266), next-generation video / image coding standards after VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), EVC (Essential Video Coding) standard, AVS2 standard, etc.).
[0039] This document provides various implementations related to video / image encoding, and these implementations may be combined and performed in combination with each other unless otherwise specified.
[0040] In this document, video can refer to a collection of images over a period of time. Typically, an image is a unit representing a specific time region, while a strip / patch is a unit that constitutes a part of an image. A strip / patch may include one or more coding tree units (CTUs). An image may consist of one or more strips / patches. An image may consist of one or more patch groups. A patch group may include one or more patches.
[0041] A pixel or primitive (pel) can refer to the smallest unit that makes up a picture (or image). Alternatively, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. Alternatively, a sample can refer to a pixel value in the spatial domain, or, when the pixel value is transformed to the frequency domain, it can refer to the transform coefficients in the frequency domain.
[0042] A unit can represent the basic unit of image processing. A unit may include a specific region and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the context, units and terms such as blocks and regions may be used interchangeably. Typically, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0043] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" can mean "A and / or B". Additionally, "A, B" can mean "A and / or B". Furthermore, "A / B / C" can mean "at least one of A, B, and / or C". Additionally, "A / B / C" can mean "at least one of A, B, and / or C".
[0044] Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" could include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".
[0045] Figure 1 Examples of video / image coding systems to which this disclosure can be applied are illustrated schematically.
[0046] Reference Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.
[0047] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0048] Video sources can be obtained through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, video / image archives including previously captured video / images, etc. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated by computers, etc. In this case, the video / image capture process can be replaced by a process that generates related data.
[0049] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0050] A transmitter can send encoded video / image information or data, output in bitstream form, to a receiver in a receiving device via a digital storage medium or network, either as a file or a stream. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to a decoding device.
[0051] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0052] The renderer can render decoded video / images. The rendered video / images can then be displayed on a monitor.
[0053] Figure 2 This diagram schematically illustrates the configuration of a video / image encoding apparatus to which this disclosure may be applied. In the following, the term "video encoding apparatus" may include an image encoding apparatus.
[0054] Reference Figure 2 The encoding device 200 may include an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 described above may be constituted by one or more hardware components (e.g., an encoder chipset or processor). Furthermore, the memory 270 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0055] Image partitioner 210 can divide an input image (or picture or frame) input to encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a maximum coding unit (LCU), the coding units can be recursively partitioned according to a quadtree-binary-tritree (QTBTTT) structure. For example, based on a quadtree structure, a binary tree structure, and / or a ternary tree structure, a coding unit can be partitioned into multiple coding units of varying depths. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary tree structure can be applied later. Alternatively, a binary tree structure can be applied first. The encoding process according to this disclosure can be performed based on the final coding units without further partitioning. In this case, the maximum coding unit can be directly used as the final coding unit based on the encoding efficiency according to the image characteristics. Alternatively, the coding units can be recursively partitioned into deeper coding units as needed, thereby allowing the optimally sized coding unit to be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be separate from or distinct from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transformation unit may be a unit for deriving the transform coefficients and / or a unit for deriving the residual signal from the transform coefficients.
[0056] Depending on the context, units and terms such as blocks and regions can be used to represent each other. Typically, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Samples can be used as a term corresponding to pixels or primitives (pellets) in a picture (or image).
[0057] Subtractor 231 subtracts the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 221 or intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to converter 232. In this case, as shown, the unit in encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as subtractor 231. The predictor can perform prediction on the processing target block (hereinafter referred to as "current block") and can generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various information related to the prediction (e.g., prediction mode information) and send the generated information to entropy encoder 240. The information about the prediction can be encoded in entropy encoder 240 and output as a bitstream.
[0058] Intra-predictor 222 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the reference samples can be located near or separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, the directional modes can include, for example, 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes can be used depending on the settings. Intra-predictor 222 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.
[0059] Inter-frame predictor 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on the correlation between motion information of neighboring blocks and the current block, on a block, sub-block, or sample basis. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same as or different from each other. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in jump mode and merge mode, the inter-frame predictor 221 can use motion information of neighboring blocks as motion information of the current block. In jump mode, unlike merge mode, residual signals cannot be sent. In motion information prediction (motion vector prediction, MVP) mode, motion vectors of neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0060] Predictor 220 can generate prediction signals based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to the prediction of a block, and can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform prediction on a block based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video encoding such as games, etc. Although IBC essentially performs prediction within the current block, its execution is similar to inter-frame prediction in that it derives a reference block within the current block. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0061] The predicted signals generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate the reconstructed signal or the residual signal. The transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques can include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transform obtained based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size, or to blocks of variable size that are not square.
[0062] Quantizer 233 quantizes the transform coefficients and sends them to entropy encoder 240, which encodes the quantized signal (information about the quantized transform coefficients) and outputs the encoded signal in a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector based on the coefficient scan order and generate information about the quantized transform coefficients based on this one-dimensional vector form. Entropy encoder 240 can perform various encoding methods such as exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 240 can encode information required for video / image reconstruction, other than the quantized transform coefficients (e.g., values of syntax elements), either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form on a unit-by-unit basis in the Network Abstraction Layer (NAL). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. In this disclosure, information and / or syntax elements sent from the encoding device to / signaled to the decoding device may be included in the video / image information. The video / image information can be encoded using the encoding process described above and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks, communication networks, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that sends the signal output from the entropy encoder 240 or a memory (not shown) that stores it may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoder 240.
[0063] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transform using vectorized transform coefficients via dequantizer 234 and inverse transformer 235, the residual signal (residual block or residual sample) can be reconstructed. Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222, thereby generating a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). When there is no residual for the processing target block, as in the case of applying a jump mode, the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current block, and, as described later, for inter-frame prediction of the next image through filtering.
[0064] In addition, luminance mapping with chroma scaling (LMCS) can be applied in image encoding and / or reconstruction processing.
[0065] Filter 260 can improve subjective / objective video quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270, specifically in the DPB of memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive ring filter, bilateral filter, etc. As discussed later in the description of each filtering method, filter 260 can generate various filtering-related information and send the generated information to entropy encoder 240. The filtering information can be encoded in entropy encoder 240 and output as a bitstream.
[0066] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. Accordingly, the encoding device can avoid prediction mismatch in the encoding device 100 and the decoding device when applying inter-frame prediction, and can also improve encoding efficiency.
[0067] Memory 270DPB can store modified reconstructed images for use as reference images in inter-frame predictor 221. Memory 270 can store motion information of blocks in the current image from which motion information has been derived (or encoded) and / or motion information of blocks in reconstructed images. The stored motion information can be sent to inter-frame predictor 221 to be used as motion information of neighboring blocks or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 222.
[0068] Figure 3This is a diagram that schematically illustrates the configuration of a video / image decoding device to which this disclosure can be applied.
[0069] Reference Figure 3 The video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above may be constituted by one or more hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0070] When the input includes a bitstream containing video / image information, the decoding device 300 can interact with data already prepared therein. Figure 2 The processing of video / image information in the encoding device correspondingly reconstructs the image. For example, the decoding device 300 can deduce units / blocks based on information related to block segmentation obtained from the bitstream. The decoding device 300 can perform decoding by using processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, an encoding unit, which can be segmented along a quadtree structure, binary tree structure, and / or ternary tree structure using encoding tree units or maximum encoding units. One or more transform units can be derived using encoding units. And, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproducer.
[0071] Decoding device 300 can receive data from... in the form of a bitstream. Figure 2The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc. In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or general constraint information. In this disclosure, the signaling / receiving information and / or syntax elements, which will be described subsequently, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information about the target syntax element and the decoding information of neighboring and target blocks, or information about symbols / bins decoded in previous steps, predict the bin generation probability based on the determined context model, and perform arithmetic decoding on the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model after determining it using information about symbols / bins decoded for the context model of the next symbol / bin. Prediction information from the information decoded in the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantization transform coefficients) and associated parameter information that have undergone entropy decoding in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, filtering information from the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) that receives the signal output from the encoding device can also configure the decoding device 300 as an internal / external component, and the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to this disclosure can be referred to as a video / image / picture encoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0072] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement can be performed based on the order of coefficient scans already performed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0073] The dequantizer 321 obtains the residual signal (residual block, residual sample array) by performing an inverse transform on the transform coefficients.
[0074] The predictor can perform predictions on the current block and generate a prediction block that includes prediction samples for the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from the entropy decoder 310, and specifically, can determine the intra-frame / inter-frame prediction mode.
[0075] The predictor can generate a predicted signal based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to the prediction of a block, and can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform intra-block copying (IBC) for the prediction of a block. Intra-block copying can be used for content image / video encoding such as in games with screen content coding (SCC). Although IBC essentially performs prediction within the current block, its execution is similar to inter-frame prediction in that it derives a reference block within the current block. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0076] The intra-predictor 331 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the reference samples can be located near or separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-predictor 331 can determine the prediction mode applied to the current block by using the prediction modes applied to neighboring blocks.
[0077] Inter-frame predictor 332 can deduce the predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on the correlation between motion information of neighboring blocks and the current block, on a block, sub-block, or sample basis. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference image index of the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the mode of inter-frame prediction for the current block.
[0078] Adder 340 can generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from predictor 330. When there is no residual for processing the target block, as in the case of applying the jump mode, the prediction block can be used as the reconstruction block.
[0079] Adder 340 can be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current block, and as described later, it can be output by filtering or used for inter-frame prediction of the next image.
[0080] In addition, luminance mapping with chroma scaling (LMCS) can be applied in image decoding processing.
[0081] Filter 350 can improve subjective / objective video quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be sent to memory 360, specifically to the DPB of memory 360. Various filtering methods can include, for example, deblocking filtering, adaptive sample shifting, adaptive ring filtering, bilateral filtering, etc.
[0082] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks in the current image from which motion information has been derived (or decoded) and / or motion information of blocks in a reconstructed image. The stored motion information can be sent to inter-frame predictor 332 to be used as motion information of neighboring blocks or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 331.
[0083] The examples described in this specification in the predictor 330, dequantizer 321, inverse transformer 322 and filter 350 of the decoding device 300 can be similarly or correspondingly applied to the predictor 220, dequantizer 234, inverse transformer 235 and filter 260 of the encoding device 200, respectively.
[0084] As described above, prediction is performed to improve compression efficiency during video encoding. Accordingly, a prediction block can be generated that includes prediction samples for the current block, which is the target block for encoding. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived identically in both the encoding and decoding devices, and the encoding device can improve image encoding efficiency by signaling to the decoding device information about the residual between the original block and the prediction block (residual information), not the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed image including the reconstructed block.
[0085] Residual information can be generated through transformation and quantization processes. For example, an encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients. This allows it to signal the associated residual information to the decoding device (via a bitstream). Here, the residual information can include the value information, position information, transform technique, transform kernel, quantization parameters, etc., of the quantized transform coefficients. The decoding device can perform quantization / dequantization processes based on the residual information and derive residual samples (or residual sample blocks). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can derive the residual block by performing dequantization / inverse transform on the quantized transform coefficients to serve as a reference for inter-frame prediction of the next image, and can generate a reconstructed image based on this.
[0086] Figure 4 Multiple transformation techniques according to embodiments of the present disclosure are illustrated schematically.
[0087] Reference Figure 4 The converter can correspond to the aforementioned Figure 2 The converter in the encoding device, and the inverse converter can correspond to the aforementioned Figure 2 Inverse converter in encoding devices, or Figure 3 The inverse converter in the decoding device.
[0088] The transformer can derive (first) transform coefficients (S410) by performing a first transform based on residual samples (residual sample array) in the residual block. This first transform can be referred to as the core transform. In this paper, the first transform can be based on multiple transform selection (MTS), and when multiple transforms are used as a first transform, it can be referred to as a multi-core transform.
[0089] Multi-core transform can refer to a method that additionally uses Discrete Cosine Transform (DCT) Type 2 and Discrete Sine Transform (DST) Type 7, DCT Type 8, and / or DST Type 1 for transformation. In other words, multi-core transform can represent a method that transforms a spatial domain residual signal (or residual block) into frequency domain transform coefficients (or primary transform coefficients) based on multiple transform kernels selected from DCT Type 2, DST Type 7, DCT Type 8, and DST Type 1. In this paper, from the perspective of the transformer, primary transform coefficients can be referred to as temporary transform coefficients.
[0090] In other words, when applying conventional transform methods, transform coefficients can be generated by applying a spatial-to-frequency domain transform to the residual signal (or residual block) based on DCT type 2. In contrast, when applying multi-core transforms, transform coefficients (or single-stage transform coefficients) can be generated by applying a spatial-to-frequency domain transform to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. In this paper, DCT type 2, DST type 7, DCT type 8, and DST type 1 can be referred to as transform types, transform kernels, or transform cores.
[0091] For reference, the DCT / DST transform type can be defined based on basis functions, and the basis functions can be shown in the table below.
[0092] [Table 1]
[0093]
[0094] If a multi-core transform is performed, a vertical transform core and a horizontal transform core can be selected from the transform cores for the target block. A vertical transform can be performed on the target block based on the vertical transform core, and a horizontal transform can be performed on the target block based on the horizontal transform core. Here, the horizontal transform can represent the transform of the horizontal components of the target block, and the vertical transform can represent the transform of the vertical components of the target block. The vertical transform core / horizontal transform core can be adaptively determined based on the prediction mode and / or transform index of the target block (CU or sub-block), including the residual block.
[0095] Furthermore, according to the example, if a transformation is performed by applying an MTS, the mapping relationship of the transformation kernels can be set by setting specific basis functions to predetermined values and combining the basis functions to be applied in the vertical or horizontal transformation. For example, when the horizontal transformation kernel is denoted as trTypeHor and the vertical transformation kernel is denoted as trTypeVer, a value of 0 for trTypeHor or trTypeVer can be set to DCT2, a value of 1 for trTypeHor or trTypeVer can be set to DST7, and a value of 2 for trTypeHor or trTypeVer can be set to DCT8.
[0096] In this scenario, the MTS index information can be encoded and signaled to the decoding device to indicate any one of the multiple transform cores. For example, MTS index 0 can indicate that both trTypeHor and trTypeVer values are 0, MTS index 1 can indicate that both trTypeHor and trTypeVer values are 1, MTS index 2 can indicate that trTypeHor is 2 and trTypeVer is 1, MTS index 3 can indicate that trTypeHor is 1 and trTypeVer is 2, and MTS index 4 can indicate that both trTypeHor and trTypeVer values are 2.
[0097] The transformer can derive modified (secondary) transform coefficients by performing a secondary transform based on the (first) transform coefficients (S420). A primary transform is a transform from the spatial domain to the frequency domain, while a secondary transform refers to a transformation to a more compressed representation by utilizing the correlation between the (first) transform coefficients. Secondary transforms can include inseparable transforms. In this case, the secondary transform can be called an inseparable secondary transform (NSST) or a mode-dependent inseparable secondary transform (MDNSST). An inseparable secondary transform can represent a transform that generates modified transform coefficients (or secondary transform coefficients) for the residual signal by performing a secondary transform on the (first) transform coefficients derived from the primary transform based on an inseparable transform matrix. In this case, the vertical and horizontal transforms may not be applied separately to the (first) transform coefficients (or the horizontal and vertical transforms may not be applied independently), but the transform matrix can be applied once based on the inseparable transform. In other words, an inseparable quadratic transform can represent a transformation method where the vertical and horizontal components of the (first) transform coefficients are not separated, and, for example, a two-dimensional signal (transform coefficients) is rearranged into a one-dimensional signal through a defined direction (e.g., a first row direction or a first column direction), and then modified transform coefficients (or quadratic transform coefficients) are generated based on the inseparable transform matrix. For example, M×N blocks are arranged in rows according to row priority, in the order of first row, second row, ..., and Nth row. M×N blocks are arranged in columns according to column priority, in the order of first column, second column, ..., and Nth column. An inseparable quadratic transform can be applied to the upper left region of a block containing (first) transform coefficients (hereinafter referred to as a transform coefficient block). For example, if the width (W) and height (H) of the transform coefficient block are both equal to or greater than 8, an 8×8 inseparable quadratic transform can be applied to the upper left 8×8 region of the transform coefficient block. Furthermore, if the width (W) and height (H) of the transform coefficient block are both equal to or greater than 4, and the width (W) or height (H) of the transform coefficient block is less than 8, then a 4×4 inseparable quadratic transform can be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block. However, the embodiments are not limited to this, and for example, even if only the condition that the width (W) or height (H) of the transform coefficient block is equal to or greater than 4 is met, a 4×4 inseparable quadratic transform can also be applied to the upper left min(8,W)×min(8,H) region of the transform coefficient block.
[0098] Specifically, for example, if a 4×4 input block is used, the inseparable quadratic transformation can be performed as follows.
[0099] A 4×4 input block X can be represented as follows.
[0100] [Formula 1]
[0101]
[0102] If X is represented as a vector, then the vector It can be represented as follows.
[0103] [Equation 2]
[0104]
[0105] In Equation 2, the vector It is a one-dimensional vector obtained by rearranging the two-dimensional block X of Equation 1 according to the row priority order.
[0106] In this case, the inseparable quadratic transformation can be calculated as follows.
[0107] [Formula 3]
[0108]
[0109] In this formula, represents the transformation coefficient vector, while T represents the 16×16 (inseparable) transformation matrix.
[0110] Using Equation 3 above, the 16×1 transformation coefficient vector can be derived. Furthermore, the vector can be scanned in sequence (horizontal, vertical, and diagonal, etc.). Reorganize into 4×4 blocks. However, the above calculation is an example, and the hypercube-Givens transform (HyGT) and similar methods can also be used to calculate inseparable quadratic transformations in order to reduce the computational complexity of inseparable quadratic transformations.
[0111] Furthermore, in inseparable quadratic transforms, the transform kernel (or transform type) can be selected as mode-dependent. In this case, the mode can include intra-frame prediction mode and / or inter-frame prediction mode.
[0112] As described above, an inseparable quadratic transformation can be performed based on an 8×8 transformation or a 4×4 transformation determined by the width (W) and height (H) of the transform coefficient block. An 8×8 transformation is a transformation applicable to an 8×8 region contained within the transform coefficient block when both W and H are equal to or greater than 8, and this 8×8 region can be the top-left 8×8 region within the transform coefficient block. Similarly, a 4×4 transformation is a transformation applicable to a 4×4 region contained within the transform coefficient block when both W and H are equal to or greater than 4, and this 4×4 region can be the top-left 4×4 region within the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, while the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.
[0113] Here, to select pattern-based transform kernels, for inseparable quadratic transforms used for both 8×8 and 4×4 transforms, three inseparable quadratic transform kernels can be configured for each transform set, and there can be 35 transform sets. That is, 35 transform sets can be configured for 8×8 transforms, and 35 transform sets can be configured for 4×4 transforms. In this case, each of the 35 transform sets for 8×8 transforms can contain three 8×8 transform kernels, and each of the 35 transform sets for 4×4 transforms can contain three 4×4 transform kernels. The transform size, number of sets, and number of transform kernels in each set mentioned above are for illustrative purposes only. Alternatively, sizes other than 8×8 or 4×4 can be used, n sets can be configured, and each set can include k transform kernels.
[0114] The transform set can be called the NSST set, and the transform kernel in the NSST set can be called the NSST kernel. For example, a specific set from the transform set can be selected based on the intra-prediction mode of the target block (CU or sub-block).
[0115] For reference, as an example, intra-prediction modes may include two non-directional (or non-angular) intra-prediction modes and 65 directional (or angular) intra-prediction modes. The non-directional intra-prediction modes may include a 0-plane intra-prediction mode and a 1-DC intra-prediction mode, and the directional intra-prediction modes may include 65 intra-prediction modes between intra-prediction mode 2 and intra-prediction mode 66. However, this is an example, and this disclosure can be applied to cases where the number of intra-prediction modes differs. Furthermore, depending on the circumstances, an intra-prediction mode 67 may be used, and intra-prediction mode 67 may represent a linear model (LM) mode.
[0116] Figure 5 Intra-frame orientation patterns for 65 prediction directions are illustrated illustratively.
[0117] Reference Figure 5 Based on intra-prediction mode 34 with a left-top diagonal prediction direction, intra-prediction modes with horizontal directionality and intra-prediction modes with vertical directionality can be classified. Figure 5H and V refer to the horizontal and vertical orientations, respectively, and the numbers -32 to 32 indicate displacements in units of 1 / 32 at the sample grid positions. This can represent the offset of the mode index value. Intra-prediction modes 2 to 33 are horizontally oriented, while intra-prediction modes 34 to 66 are vertically oriented. Furthermore, strictly speaking, intra-prediction mode 34 can be considered neither horizontal nor vertical, but in terms of the transform set used to determine the quadratic transform, it can be classified as horizontally oriented. This is because the input data is transposed for a vertically oriented mode symmetric to intra-prediction mode 34, and the input data alignment method used for the horizontal mode is applied to intra-prediction mode 34. Transposing the input data means transforming the rows and columns of the two-dimensional block data M×N into N×M data. Intra-prediction modes 18 and 50 can represent the horizontal and vertical intra-prediction modes, respectively, and intra-prediction mode 2 can be called the upper-right diagonal intra-prediction mode because it has a left reference pixel and predicts in the upper-right direction. In the same way, intra prediction mode 34 can be called the bottom right diagonal intra prediction mode, and intra prediction mode 66 can be called the bottom left diagonal intra prediction mode.
[0118] In this case, for example, the mapping between the 35 transform sets and the intra-prediction mode can be shown in the table below. For reference, if the LM mode is applied to the target block, the quadratic transform need not be applied to the target block.
[0119] [Table 2]
[0120] In-frame mode 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 set 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 In-frame mode 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67(LM) set 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 null
[0121] Furthermore, if a specific set is determined to be used, one of the k transform kernels in that specific set can be selected using an inseparable quadratic transform index. The encoding device can derive the inseparable quadratic transform index indicating the specific transform kernel based on rate-distortion (RD) check and can signal the inseparable quadratic transform index to the decoding device. The decoding device can then select one of the k transform kernels in the specific set based on the inseparable quadratic transform index. For example, NSST index value 0 can indicate the first inseparable quadratic transform kernel, NSST index value 1 can indicate the second inseparable quadratic transform kernel, and NSST index value 2 can indicate the third inseparable quadratic transform kernel. Alternatively, NSST index value 0 can indicate that the first inseparable quadratic transform is not applied to the target block, and NSST index values 1 through 3 can indicate three transform kernels.
[0122] Return to reference Figure 4The converter can perform an inseparable quadratic transform based on the selected transform core and obtain modified (quadratic) transform coefficients. As mentioned above, the modified transform coefficients can be derived as transform coefficients quantized by a quantizer and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse converter in the encoding device.
[0123] Furthermore, as mentioned above, if the second transformation is omitted, the (first) transformation coefficients, which are the output of the first (separable) transformation, can be derived as the transformation coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.
[0124] The inverse transformer can perform a series of processes in the reverse order of those already performed in the aforementioned transformers. The inverse transformer can receive (dequantized) transform coefficients and derive (first) transform coefficients by performing a second (inverse) transform (S450), and obtain residual blocks (residual samples) by performing a first (inverse) transform on the (first) transform coefficients (S460). In this regard, from the perspective of the inverse transformer, the first transform coefficients can be referred to as modified transform coefficients. As described above, the encoding and decoding devices can generate reconstructed blocks based on the residual blocks and the prediction blocks, and can generate reconstructed images based on the reconstructed blocks.
[0125] The decoding device may also include a second-order inverse transform application determiner (or a component for determining whether to apply the second-order inverse transform) and a second-order inverse transform determiner (or a component for determining the second-order inverse transform). The second-order inverse transform application determiner can determine whether to apply the second-order inverse transform. For example, the second-order inverse transform can be NSST or RST, and the second-order inverse transform application determiner can determine whether to apply the second-order inverse transform based on a second-order transform flag obtained by parsing the bitstream. In another example, the second-order inverse transform application determiner can determine whether to apply the second-order inverse transform based on the transform coefficients of the residual block.
[0126] A second-order inverse transform determiner can determine the second-order inverse transform. In this case, the second-order inverse transform determiner can determine the second-order inverse transform applied to the current block based on the NSST (or RST) transform set specified according to the intra-prediction mode. In implementations, the second-order transform determination method can be determined depending on the first-order transform determination method. Various combinations of the first and second-order transforms can be determined based on the intra-prediction mode. Furthermore, in the example, the second-order inverse transform determiner can determine the region where the second-order inverse transform is applied based on the size of the current block.
[0127] Furthermore, as mentioned above, if the second (inverse) transform is omitted, the (dequantized) transform coefficients can be received, a first (separable) inverse transform can be performed, and a residual block (residual sample) can be obtained. As mentioned above, the encoding and decoding devices can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed image based on the reconstructed block.
[0128] Furthermore, in this disclosure, a reduced quadratic transformation (RST) in which the size of the transformation matrix (kernel) is reduced can be applied to the concept of NSST in order to reduce the computational and storage requirements of the inseparable quadratic transformation.
[0129] Furthermore, the transform kernel, transform matrix, and coefficients constituting the transform kernel matrix described in this disclosure, i.e., kernel coefficients or matrix coefficients, can be represented in 8 bits. This is feasible for implementation in decoding and encoding devices, and compared to existing 9-bit or 10-bit representations, it reduces the amount of storage required to store the transform kernel and can reasonably accommodate performance degradation. Additionally, representing the kernel matrix in 8 bits allows for the use of smaller multipliers and is more suitable for Single Instruction Multiple Data (SIMD) instructions for optimal software implementation.
[0130] In this specification, the term "RST" can refer to a transformation performed on the residual samples of a target block based on a transformation matrix whose size is reduced according to a reduction factor. When performing a reduction transformation, the computational cost required for the transformation can be reduced due to the smaller size of the transformation matrix. In other words, RST can be used to address computational complexity issues that arise when transforming large blocks or when transforming indivisible blocks.
[0131] RST can be referred to by various terms such as reduced transform, reduced quadratic transform, reduced transform, simplified transform, and simple transform, and the names that RST can be called are not limited to the examples listed. Alternatively, since RST is performed primarily in the low-frequency region of the transform block that includes non-zero coefficients, it can be called low-frequency inseparable transform (LFNST).
[0132] Furthermore, when performing a second inverse transform based on RST, the inverse transformer 235 of the encoding device 200 and the inverse transformer 322 of the decoding device 300 may include: an inverse reduced second transformer that derives modified transform coefficients based on the inverse RST of the transform coefficients; and an inverse first transformer that derives the residual samples of the target block based on the inverse first transform of the modified transform coefficients. An inverse first transform refers to the inverse transform of a first transform applied to the residuals. In this disclosure, deriving transform coefficients based on a transform can mean deriving the transform coefficients by applying a transform.
[0133] Figure 6 This is a diagram illustrating an embodiment of the RST according to the present disclosure.
[0134] In this specification, the term "target block" may refer to the current block or residual block on which coding is performed.
[0135] In the example RST, an N-dimensional vector can be mapped to an R-dimensional vector in another space, thus determining the reduced transformation matrix, where R is less than N. N can refer to the square of the length of the side of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the reduction factor can refer to the R / N value. The reduction factor can be called a reduction factor, shrinkage factor, simplification factor, or other various terms. Furthermore, R can be called a reduction coefficient, but depending on the situation, the reduction factor can refer to R. Additionally, depending on the situation, the reduction factor can refer to the N / R value.
[0136] In this example, the reduction factor or reduction coefficient can be signaled via a bitstream, but the example is not limited to this. For instance, a predetermined value for the reduction factor or reduction coefficient can be stored in each of the encoding device 200 and the decoding device 300, and in this case, the reduction factor or reduction coefficient does not need to be signaled separately.
[0137] The size of the reduced transformation matrix, as shown in the example, can be less than N×N (the size of the regular transformation matrix) and can be limited as shown in Equation 4 below.
[0138] [Formula 4]
[0139]
[0140] Figure 6 The matrix T in the reduced transformation block shown in (a) can refer to the matrix T in Equation 4. R×N .like Figure 6 As shown in (a), when the reduced transformation matrix T R×N The transformation coefficients of the target block can be derived by multiplying by the residual samples of the target block.
[0141] In the example, if the size of the block to which the transformation is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then according to Figure 6 The RST of (a) can be represented as the matrix operation shown in Equation 5. In this case, the storage and multiplication computations can be reduced to approximately 1 / 4 by a reduction factor.
[0142] [Formula 5]
[0143]
[0144] In Equation 5, r1 to r 64The residual samples of the target block can be represented, and specifically, they can be the transformation coefficients generated by applying a single transformation. As a result of the calculation in Equation 5, the transformation coefficients c of the target block can be derived. i And derive c i The process can be shown in Equation 6.
[0145] [Formula 6]
[0146]
[0147] As a result of Equation 6, the transformation coefficients c1 to c of the target block can be derived. R In other words, when R = 16, the transformation coefficients c1 to c of the target block can be derived. 16 If a conventional transform is applied instead of an RST, and a 64×64 (N×N) transform matrix is multiplied by a 64×1 (N×1) residual sample, only 16(R) transform coefficients are derived for the target block because of the application of the RST, even though 64(N) transform coefficients are derived for the target block. Since the total number of transform coefficients used for the target block is reduced from N to R, the amount of data sent from the encoding device 200 to the decoding device 300 is reduced, thus improving the transmission efficiency between the encoding device 200 and the decoding device 300.
[0148] When considering the size of the transformation matrix, the size of a regular transformation matrix is 64×64 (N×N), but the size of a reduced transformation matrix is reduced to 16×64 (R×N). Therefore, compared to performing a regular transformation, the storage utilization rate of performing an RST can be reduced by the R / N ratio. Furthermore, compared to the number of multiplications (N×N) when using a regular transformation matrix, using a reduced transformation matrix can reduce the number of multiplications (R×N) by the R / N ratio.
[0149] In the example, the transformer 232 of the encoding device 200 can derive the transform coefficients of the target block by performing a first transform and an RST-based second transform on the residual samples of the target block. These transform coefficients can be passed to the inverse transformer of the decoding device 300, and the inverse transformer 322 of the decoding device 300 can derive the modified transform coefficients based on the inverse reduced second transform (RST) for the transform coefficients, and can derive the residual samples of the target block based on the inverse first transform for the modified transform coefficients.
[0150] Based on the example inverse RST matrix T N×R Its size is N×R, which is larger than the size of the conventional inverse transformation matrix N×N, and is the same as the reduced transformation matrix T shown in Equation 4. R×N It has a transpose relationship.
[0151] Figure 6The matrix T in the reduced inverse transform block shown in (b) t It can refer to the inverse RST matrix T N×R T (The superscript T indicates transpose). For example... Figure 6 As shown in (b), when the inverse RST matrix T N×R T Multiplying by the transform coefficients of the target block allows us to derive the modified transform coefficients of the target block or the residual samples of the target block. The inverse RST matrix T R×N T It can be represented as (T) R×N T ) NxR .
[0152] More specifically, when the inverse RST is used as a second inverse transformation, when the inverse RST matrix T N×R T When multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. Furthermore, the inverse RST can be used as the inverse first-order transform, and in this case, when the inverse RST matrix T... N×R T When multiplied by the transformation coefficients of the target block, the residual sample of the target block can be derived.
[0153] In the example, if the size of the block to which the inverse transform is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then according to Figure 6 The RST of (b) can be represented as the matrix operation shown in Equation 7.
[0154] [Formula 7]
[0155]
[0156] In Equation 7, c1 to c 16 This can represent the transformation coefficients of the target block. As a result of Equation 7, the transformation coefficients representing the modifications to the target block or the r values of the residual samples of the target block can be derived. j And derive r j The process can be shown in Equation 8.
[0157] [Formula 8]
[0158]
[0159] As a result of Equation 8, the transformation coefficients representing the modification of the target block or the residual samples of the target block, r1 to r2, can be derived. NFrom the perspective of the size of the inverse transformation matrix, the size of the regular inverse transformation matrix is 64×64 (N×N), but the size of the inverse reduced transformation matrix is reduced to 64×16 (R×N). Therefore, compared with performing the regular inverse transformation, the storage utilization rate of performing the inverse RST can be reduced by the R / N ratio. In addition, when comparing the number of multiplications N×N when using the regular inverse transformation matrix, using the inverse reduced transformation matrix can reduce the number of multiplications (N×R) by the R / N ratio.
[0160] The transform set configuration shown in Table 2 can also be applied to 8×8 RST. That is, 8×8 RST can be applied based on the transform sets in Table 2. Since a transform set includes two or three transforms (kernels) depending on the intra-prediction mode, it can be configured to select one of up to four transforms, including those without applying a secondary transform. In the transform without applying a secondary transform, the application of an identity matrix can be considered. Assuming indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case where the identity matrix is applied, i.e., without applying a secondary transform), the NSST index, which is a syntax element, can be signaled for each transform coefficient block, thereby specifying the transform to be applied. That is, an 8×8 NSST can be specified for the top-left 8×8 block via the NSST index, and an 8×8 RST can be specified in the RST configuration. 8×8 NSST and 8×8 RST refer to transforms that can be applied to the 8×8 region included in the transform coefficient block when the W and H of the target block to be transformed are both equal to or greater than 8, and the 8×8 region can be the top-left 8×8 region in the transform coefficient block. Similarly, 4×4NSST and 4×4RST refer to transformations that can be applied to the 4×4 region included in the transform coefficient block when both W and H of the target block are equal to or greater than 4, and the 4×4 region can be the upper left 4×4 region in the transform coefficient block.
[0161] If the (forward) 8×8 RST shown in Equation 4 is applied, 16 effective transform coefficients are generated. Therefore, considering that the 64 input data points forming the 8×8 region are reduced to 16 output data points, and from the perspective of the two-dimensional region, only 1 / 4 of this region is filled with effective transform coefficients, the 16 output data points obtained by applying the forward 8×8 RST can fill the upper left region of this block (transform coefficients 1 to 16), as shown below. Figure 7 As shown.
[0162] Figure 7 This is a diagram illustrating the transformation coefficient scanning order according to an embodiment of the present disclosure. As described above, when the forward scanning order starts from the first transformation coefficient, transformation coefficients from the 64th to the 17th can be scanned in the forward scanning order in the following order: Figure 7The arrows indicate the direction and order in which the reverse scan is performed.
[0163] exist Figure 7 In the diagram, the top-left 4×4 region is the Region of Interest (ROI) filled with valid transform coefficients, while the remaining regions are empty. By default, empty regions can be filled with 0.
[0164] In other words, when an 8×8 RST with a 16×64 positive transformation matrix is applied to an 8×8 region, the output transformation coefficients can be arranged in the upper left 4×4 region, and according to... Figure 7 The scanning order can fill regions where no output transform coefficients exist with 0 (from transform coefficient 64 to 17).
[0165] If in Figure 7 If non-zero valid transform coefficients are found outside the ROI, it is determined that 8×8 RST has not yet been applied, and therefore NSST index encoding can be omitted. Conversely, if in Figure 7 If no non-zero transform coefficients are found outside the ROI (e.g., if transform coefficients are set to 0 in regions outside the ROI when applying 8×8 RST), then 8×8 RST has likely been applied, and NSST index coding can be performed. Since it is necessary to check for the presence of non-zero transform coefficients, this conditional NSST index coding can be performed after residual coding.
[0166] This disclosure discloses a method for optimizing the design and association of an RST, which can be applied from the RST structure described in this embodiment to a 4×4 block. Some concepts can be applied not only to 4×4 RSTs, but also to 8×8 RSTs or other types of transformations.
[0167] Figure 8 This is a flowchart illustrating the reverse RST process according to an embodiment of the present disclosure.
[0168] Figure 8 Each operation disclosed in the document can be performed by Figure 3 The decoding device 300 shown is used to perform this operation. Specifically, S800 can be performed by... Figure 3 The dequantizer 321 shown is executed, and S810 and S820 can be performed by... Figure 3 The inverse transformer 322 shown is used to perform this. Therefore, for the above reference... Figure 3 Specific details regarding the overlap of content will be omitted or summarized. In this disclosure, RST can be applied to a forward transform, and inverse RST can refer to being applied to a backward transform.
[0169] In implementation, the specific operation according to the reverse RST may differ from the specific operation according to RST only in that their operation order is reversed, and the specific operation according to the reverse RST may be substantially similar to the specific operation according to RST. Therefore, those skilled in the art will readily understand that the description of S800 to S820 for the reverse RST described below can be applied to RST in the same or similar manner.
[0170] According to the embodiment, the decoding device 300 can derive the transform coefficients by performing dequantization on the quantization transform coefficients of the target block (S800).
[0171] Decoding device 300 can determine whether to apply an inverse quadratic transform after the inverse first transform and before the inverse quadratic transform. For example, the inverse quadratic transform can be NSST or RST. For instance, the decoding device can determine whether to apply the inverse quadratic transform based on the quadratic transform flag parsed from the bitstream. In another example, the decoding device can determine whether to apply the inverse quadratic transform based on the transform coefficients of the residual block.
[0172] Decoding device 300 can determine the inverse quadratic transform. In this case, decoding device 300 can determine the inverse quadratic transform applied to the current block based on the NSST (or RST) transform set specified according to the intra-prediction mode. In implementations, the quadratic transform determination method can be determined depending on the primary transform determination method. For example, it can be determined that RST or LFNST is applied only when DCT-2 is used as the transform kernel in the primary transform. Alternatively, various combinations of primary and secondary transforms can be determined according to the intra-prediction mode.
[0173] Furthermore, in the example, the decoding device 300 can determine the region to which the inverse quadratic transform is applied based on the size of the current block before determining the inverse quadratic transform.
[0174] According to the embodiment, the decoding device 300 can select a transform kernel (S810). More specifically, the decoding device 300 can select a transform kernel based on at least one of the following: transform index, the width and height of the region to which the transform is applied, the intra-prediction mode used in image decoding, and the color components of the target block. However, the example is not limited to this; for example, the transform kernel can be predefined, and separate information for selecting the transform kernel can be provided without signaling it.
[0175] In one example, CIdx can indicate information about the color components of a target block. If the target block is a luma block, CIdx can indicate 0, and if the target block is a chroma block (e.g., a Cb block or a Cr block), CIdx can indicate a non-zero value (e.g., 1).
[0176] According to the implementation, the decoding device 300 can apply the inverse RST to the transform coefficients based on the selected transform kernel and reduction factor (S820).
[0177] Below, a method for determining a quadratic NSST set (i.e., a quadratic transform set or transform set) considering intra-frame prediction mode and block size according to embodiments of the present disclosure is proposed.
[0178] In the implementation, the set for the current transform block can be configured based on the intra-prediction mode described above, thereby applying a transform set including transform kernels of various sizes to the transform block. The transform sets in Table 3 are represented by 0 to 3 in Table 4.
[0179] [Table 3]
[0180] In-frame mode 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 NSST Collection 0 0 2 2 2 2 2 2 2 2 2 2 2 18 18 18 18 18 18 18 18 18 18 18 34 34 34 34 34 34 34 34 34 34 In-frame mode 34 35 36 37 38 39 40 41 42 43 44 45 46 47 46 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 NSST Collection 34 34 34 34 34 34 34 34 34 34 34 18 18 18 18 18 18 18 18 18 18 18 2 2 2 2 2 2 2 2 2 2 2
[0181] [Table 4]
[0182] In-frame mode 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 NSST Collection 0 0 1 1 1 1 1 1 1 1 1 1 1 2 2 2 2 2 2 2 2 2 2 2 3 3 3 3 3 3 3 3 3 3 In-frame mode 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 NSST Collection 3 3 3 3 3 3 3 3 3 3 3 2 2 2 2 2 2 2 2 2 2 2 1 1 1 1 1 1 1 1 1 1 1 1
[0183] The indices 0, 2, 18, and 34 shown in Table 3 correspond to 0, 1, 2, and 3 in Table 4, respectively. In Tables 3 and 4, only four transform sets are used instead of 35, thus significantly reducing memory space.
[0184] The various numbers of transformation kernel matrices that can be included in each transformation set can be set as shown in the table below.
[0185] [Table 5]
[0186]
[0187] [Table 6]
[0188]
[0189] [Table 7]
[0190]
[0191] According to Table 5, two transformation kernels are available for each transformation set, so the range of the transformation index is 0 to 2.
[0192] According to Table 6, two transform kernels are available for transform set 0 (i.e., the transform set based on DC mode and planar mode in intra-prediction mode), and one transform kernel is used for each of the remaining transform sets. Here, the range of available transform indices for transform set 1 is 0 to 2, while the range of transform indices for the remaining transform sets 1 to 3 is 0 to 1.
[0193] According to Table 7, one transformation kernel is available for each transformation set, so the range of the transformation index is 0 to 1.
[0194] In the transform set mapping of Table 3, a total of four transform sets can be used, and these four transform sets can be rearranged to be distinguished by indices 0, 1, 2, and 3, as shown in Table 4. Tables 8 and 9 illustrate the four transform sets that can be used for secondary transforms, with Table 8 presenting transform kernel matrices applicable to 8×8 blocks and Table 9 presenting transform kernel matrices applicable to 4×4 blocks. Each transform set in Tables 8 and 9 includes two transform kernel matrices, and these two transform kernel matrices can be applied to all intra-prediction modes shown in Table 5.
[0195] [Table 8]
[0196]
[0197]
[0198]
[0199]
[0200]
[0201]
[0202]
[0203]
[0204] [Table 9]
[0205]
[0206]
[0207]
[0208]
[0209]
[0210]
[0211]
[0212]
[0213] All the exemplary transformation kernel matrices shown in Table 8 are transformation kernel matrices multiplied by 128 as a scaling factor. In the g_aiNsst8×8[N1][N2]
[16]
[64] array present in the matrix array of Table 8, N1 represents the number of transformation sets (N1 is 4 or 35, distinguished by indices 0, 1, ... and N1-1), N2 represents the number of transformation kernel matrices included in each transformation set (1 or 2), and
[16]
[64] represents a 16×64 reduced quadratic transformation (RST).
[0214] As shown in Tables 3 and 4, when the transformation set includes a transformation kernel matrix, either the first transformation kernel matrix or the second transformation kernel matrix can be used for the transformation set in Table 8.
[0215] Although 16 transformation coefficients are output when applying RST, only m transformation coefficients can be output when only the m×64 portion of the 16×64 matrix is applied. For example, by setting m=8 and multiplying only by the 8×64 matrix from above, only eight transformation coefficients are output, halving the computational cost. To further reduce computation in the worst case, the 8×64 matrix can be applied to an 8×8 transformation unit (TU).
[0216] All the exemplary transformation kernel matrices applicable to the 4×4 region shown in Table 9 are transformation kernel matrices multiplied by 128 as a scaling value. In the matrix array g_aiNsst4×4[N1][N2]
[16]
[64] present in Table 9, N1 represents the number of transformation sets (N1 is 4 or 35, distinguished by indices 0, 1, ... and N1-1), N2 represents the number of transformation kernel matrices included in each transformation set (1 or 2), and
[16]
[16] represents a 16×16 transformation.
[0217] As shown in Tables 3 and 4, when the transformation set includes a transformation kernel matrix, either the first transformation kernel matrix or the second transformation kernel matrix can be used for the transformation set in Table 9.
[0218] As in 8×8 RST, only m×16 parts of the 16×16 matrix can be applied, resulting in the output of only m transformation coefficients. For example, by setting m=8 and multiplying only by the 8×16 matrix from above, only eight transformation coefficients can be output, halving the computational cost. To further reduce computation in the worst-case scenario, the 8×16 matrix can be applied to a 4×4 transformation unit (TU).
[0219] Basically, the transformation kernel matrices listed in Table 9 that can be applied to 4×4 regions can be applied to 4×4TU, 4×MTU, and M×4TU (M>4, 4×MTU and M×4TU can be divided into 4×4 regions, and each specified transformation kernel matrix can be applied to that region, or the transformation kernel matrix can be applied only to the largest top-left 4×8 or 8×4 region), or only to the top-left 4×4 region. If the quadratic transformation is configured to be applied only to the top-left 4×4 region, the transformation kernel matrices shown in Table 8 that can be applied to 8×8 regions may not be necessary.
[0220] To reduce the computational load in the worst-case scenario, the following implementation method can be proposed. Hereinafter, a matrix comprising M rows and N columns is represented as an M×N matrix, and the M×N matrix refers to the transformation matrix applied in the forward transform (i.e., when the encoding device performs the transform (RST)). Therefore, in the inverse transform (inverse RST) performed by the decoding device, an N×M matrix obtained by transposing the M×N matrix can be used.
[0221] 1) In the case of a block (e.g., a transformation unit) with a width W and a height H (where W ≥ 8 and H ≥ 8), a transformation kernel matrix applicable to an 8×8 region is applied to the top-left 8×8 region of the block. In the case of W = 8 and H = 8, only the 8×64 portion of a 16×64 matrix can be applied. That is, eight transformation coefficients can be generated.
[0222] 2) In the case of a block (e.g., a transformation unit) with a width of W and a height of H (where one of W and H is less than 8, i.e., one of W and H is 4), a transformation kernel matrix applicable to a 4×4 region is applied to the upper left region of the block. In the case of W=4 and H=4, only the 8×16 portion of a 16×16 matrix can be applied, in which case eight transformation coefficients are generated.
[0223] If (W, H) = (4, 8) or (8, 4), then the quadratic transformation is applied only to the top-left 4×4 region. If W or H is greater than 8 (i.e., if one of W and H is equal to or greater than 16 and the other is 4), then the quadratic transformation is applied only to the two top-left 4×4 blocks. In other words, only the top-left 4×8 or 8×4 region can be divided into two 4×4 blocks, and the specified transformation kernel matrix can be applied to them.
[0224] 3) In the case of a block (e.g., a transformation unit) with a width of W and a height of H (where both W and H are 4), a quadratic transformation may not be required.
[0225] 4) In the case of a block with width W and height H (e.g., a transform unit), the number of coefficients generated by applying a quadratic transform can be kept to 1 / 4 or less of the area of the transform unit (i.e., the total number of pixels included in the transform unit = W × H). For example, when both W and H are 4, the upper 4 × 16 matrix of a 16 × 16 matrix can be applied to generate four transform coefficients.
[0226] Assuming the quadratic transformation is applied only to the top-left 8×8 region of the entire transform unit (TU), eight or fewer coefficients are needed for a 4×8 or 8×4 transform unit. Therefore, the upper 8×16 matrix of a 16×16 matrix can be applied to the top-left 4×4 region. A maximum of 16×64 matrices can be applied to an 8×8 transform unit (generating a maximum of 16 coefficients). In a 4×N or N×4 (N≥16) transform unit, a 16×16 matrix can be applied to the top-left 4×4 block, or the upper 8×16 matrix of a 16×16 matrix can be applied to two top-left 4×4 blocks. Similarly, in a 4×8 or 8×4 transform unit, eight transform coefficients can be generated by applying the upper 4×16 matrix of a 16×16 matrix to two top-left 4×4 blocks.
[0227] 5) The maximum size of the quadratic transformation applied to a 4×4 region can be limited to 8×16. In this case, the amount of memory required to store the transformation kernel matrix applied to the 4×4 region can be reduced by half compared to a 16×16 matrix.
[0228] For example, among all the transform kernel matrices shown in Table 9, the maximum size can be limited to 8×16 by extracting only the upper 8×16 matrix of each 16×16 matrix, and the actual image coding system can be implemented as an 8×16 matrix that only stores the transform kernel matrix.
[0229] If the maximum applicable transformation size is 8×16, and the maximum number of multiplications required to generate a coefficient is limited to 8, then a maximum of 8×16 matrices can be applied to 4×4 blocks, and a maximum of 8×16 matrices can be applied to each of the two top-left 4×4 blocks included in a 4×N block or an N×4 block (N≥8, N=2n, n≥3). For example, an 8×16 matrix can be stored for one of the top-left 4×4 blocks in a 4×N block or an N×4 block (N≥8, N=2n, n≥3).
[0230] According to the implementation method, when the encoding specifies the index to be applied to the secondary transform of the luminance component, specifically when a transform set includes two transform kernel matrices, it is necessary to specify whether to apply the secondary transform and which transform kernel matrix is applied to the secondary transform. For example, when no secondary transform is applied, the transform index can be encoded as 0, while when the secondary transform is applied, the transform indices for the two transform sets can be encoded as 1 and 2 respectively.
[0231] In this case, truncated unary codes can be used when encoding the transform index. For example, the binary codes for 0, 10, and 11 can be assigned to transform indices 0, 1, and 2 respectively, thereby encoding the transform index.
[0232] Additionally, when encoding transform indices using truncated unary codes, different CABAC contexts can be assigned to each bin. In the example above, two CABAC contexts can be used to encode transform indices 0, 10, and 11.
[0233] When encoding specifies the transform index to be applied to the quadratic transform of the chroma component, specifically when a transform set includes two transform kernel matrices, similar to encoding the transform index for the luminance component's quadratic transform, it is necessary to specify whether to apply the quadratic transform and which transform kernel matrix to apply in the quadratic transform. For example, when no quadratic transform is applied, the transform index can be encoded as 0, while when a quadratic transform is applied, the transform indices of the two transform sets can be encoded as 1 and 2 respectively.
[0234] In this case, truncated unary codes can be used when encoding the transform index. For example, the binary codes for 0, 10, and 11 can be assigned to transform indices 0, 1, and 2 respectively, thereby encoding the transform index.
[0235] Additionally, when encoding transform indices using truncated unary codes, different CABAC contexts can be assigned to each bin. In the example above, two CABAC contexts can be used to encode transform indices 0, 10, and 11.
[0236] According to the implementation, different CABAC context sets can be allocated based on the chroma intra-prediction mode. For example, when the chroma intra-prediction mode is divided into non-directional modes (e.g., planar mode or DC mode) and other directional modes (i.e., divided into two groups), when encoding 0, 10, and 11 in the above example, a corresponding CABAC context set (including two contexts) can be allocated to each group.
[0237] When chroma intra-prediction modes are divided into multiple groups and assigned corresponding CABAC context sets, it is necessary to find the chroma intra-prediction mode values before encoding the transform index of the secondary transform. However, in chroma direct mode (DM), since the luma intra-prediction mode values are used as is, it is also necessary to find the luma component intra-prediction mode values. Therefore, when encoding information about the chroma components, data dependencies on luma component information may occur. Therefore, in chroma DM, when encoding the transform index of the secondary transform without information about the intra-prediction mode, data dependencies can be removed by mapping to specific groups. For example, if the chroma intra-prediction mode is chroma DM, the transform index can be encoded using the corresponding CABAC context set assuming a planar mode or DC mode, or other directional modes can be assumed to apply the corresponding CABAC context set.
[0238] Figure 9 This is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.
[0239] Figure 9 Each operation shown can be performed by Figure 3 The decoding device 300 shown performs this operation. Specifically, S910 can be performed by... Figure 3 The entropy decoder 310 shown is executed, and S920 can be performed by... Figure 3 The dequantizer 321 shown is executed, and S930 and S940 can be performed by... Figure 3 The inverse converter 322 shown is executed, and S950 can be performed by... Figure 3 The adder 340 shown is executed. The operation according to S910 to S950 is based on reference... Figures 4 to 8 Some of the details mentioned above. Therefore, regarding the references above... Figures 3 to 8 Descriptions of specific details that overlap will be omitted or simplified.
[0240] According to the embodiment, the decoding device 300 can derive the quantization transform coefficients of the target block from the bitstream (S910). Specifically, the decoding device 300 can decode information about the quantization transform coefficients of the target block from the bitstream, and can derive the quantization transform coefficients of the target block based on the information about the quantization transform coefficients of the target block. The information about the quantization transform coefficients of the target block can be included in the sequence parameter set (SPS) or the stripe header, and can include information about whether a reduction transform (RST) is applied, information about the reduction factor, information about the minimum transform size with RST applied, information about the maximum transform size with RST applied, information about the size of the reduction inverse transform, and information about the transform index indicating any transform kernel matrix included in the transform set.
[0241] According to the embodiment, the decoding device 300 can derive the transform coefficients by dequantizing the quantization transform coefficients of the target block (S920).
[0242] According to the embodiment, the decoding device 300 can derive the modified transform coefficients based on the inverse reduced quadratic transform (RST) of the transform coefficients (S930).
[0243] In the example, the inverse RST can be performed based on the inverse RST transformation matrix, and the inverse RST transformation matrix can be a non-square matrix in which the number of columns is less than the number of rows.
[0244] In an implementation, S930 may include: decoding the transform index; determining, based on the transform index, whether the conditions for applying the inverse RST are met; selecting a transform kernel matrix; and, when the conditions for applying the inverse RST are met, applying the inverse RST to the transform coefficients based on the selected transform kernel matrix and / or a reduction factor. In this case, the size of the reduced inverse transform matrix may be determined based on the reduction factor.
[0245] According to the implementation method, the decoding device 300 can derive the residual sample of the target block based on the inverse transform of the modified transform coefficients (S940).
[0246] The decoding device 300 can perform an inverse first-order transform on the modified transform coefficients of the target block. In this case, a reduced inverse transform can be applied, or a regular separable transform can be used as an inverse first-order transform.
[0247] According to the implementation method, the decoding device 300 can generate a reconstructed sample based on the residual sample of the target block and the predicted sample of the target block (S950).
[0248] Referring to S930, it can be identified that the residual sample of the target block is derived based on the inverse RST of the transform coefficients of the target block. From the perspective of the size of the inverse transform matrix, since the size of the regular inverse transform matrix is N×N, while the size of the inverse RST matrix is reduced to N×R, the memory usage for performing the inverse RST can be reduced by an R / N ratio compared to performing the regular transform. Furthermore, compared to the number of multiplications (N×N) when using the regular inverse transform matrix, using the inverse RST matrix can reduce the number of multiplications (N×R) by an R / N ratio. Additionally, since only R transform coefficients need to be decoded when applying the inverse RST, the total number of transform coefficients of the target block can be reduced from N to R compared to the number of transform coefficients needed to be decoded when applying the regular inverse transform, thus increasing decoding efficiency. In other words, according to S930, the (inverse) transform efficiency and decoding efficiency of the decoding device 300 can be increased through the inverse RST.
[0249] Figure 10This is a control flow diagram illustrating the reverse RST according to an embodiment of the present disclosure.
[0250] Decoding device 300 receives information about the transform index and intra-frame prediction mode from the bit stream (S1000).
[0251] This information is received as syntax information, and the syntax information is received as a bin string containing 0s and 1s.
[0252] The entropy decoder 310 can deduce binarization information about the syntax elements of the transform index.
[0253] This operation generates a candidate set of binary values that the syntax elements for the received transform index can have. According to this implementation, the syntax elements used for the transform index can be binarized by truncating unary codes.
[0254] The syntax elements of the transform index according to this embodiment can indicate whether the inverse RST has been applied and one of the transform kernel matrices included in the transform set. When the transform set includes two transform kernel matrices, the syntax elements of the transform index can have three values.
[0255] In other words, according to the implementation method, the values of the syntax elements of the transform index can include: 0, which indicates that the inverse RST is not applied to the target block; 1, which indicates the first transform kernel matrix of the transform kernel matrix; and 2, which indicates the second transform kernel matrix of the transform kernel matrix.
[0256] In this case, the three values of the syntax element of the transform index can be encoded as 0, 10, and 11 respectively according to the truncated unary encoding. That is, the value 0 of the syntax element can be binarized to "0", the value 1 of the syntax element can be binarized to "10", and the value 2 of the syntax element can be binarized to "11".
[0257] The entropy decoder 310 can deduce context information (i.e., context model) about the bin string of the transformation index (S1010), and can decode the bin of the bin string of the syntax element based on the context information (S1020).
[0258] In other words, the entropy decoder 310 receives the bin string that has been binarized by truncating the unary encoding, and decodes the syntax elements of the transform index using the candidate set of binary values.
[0259] According to this implementation, context information for different entries (i.e., different probability models) can be applied to the two bins of the transform index respectively. That is, all two bins of the transform index can be decoded using a contextual approach rather than a bypass approach, wherein the first bin of the transform index's syntax element bins can be decoded based on first context information, and the second bin of the transform index's syntax element bins can be decoded based on second context information.
[0260] The value of the syntax element applied to the target block can be derived from the binary values that the syntax element of the transform index can have through the decoding based on context information (S1030).
[0261] In other words, it can be deduced which of the transformation indices 0, 1, and 2 applies to the current target block.
[0262] The inverse transformer 322 of the decoding device 300 determines the transform set based on the mapping relationship according to the intra-prediction mode applied to the target block (S1040), and can perform inverse RST based on the values of the syntax elements of the transform set and the transform index (S1050).
[0263] As described above, multiple transform sets can be determined based on the intra-prediction mode of the transform block to be transformed, and RST can be performed based on any of the transform kernel matrices included in the transform set indicated by the transform index.
[0264] Figure 11 This is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present disclosure.
[0265] Figure 11 Each operation shown can be performed by Figure 2 The encoding device 200 shown is used to perform this. Specifically, S1110 can be performed by... Figure 2 The predictor 220 shown is used to perform this, and S1120 can be executed by... Figure 2 The subtractor 231 shown is used to perform the operation, and S1130 and S1140 can be performed by... Figure 2 The converter 232 shown is used to perform this, and S1150 can be performed by... Figure 2 The quantizer 233 and entropy encoder 240 shown are used for execution. The operations according to S1110 to S1150 are based on… Figures 4 to 8 Some of the content described above. Therefore, regarding the reference above... Figures 4 to 8 Descriptions of specific details that overlap will be omitted or simplified.
[0266] According to the implementation method, the encoding device 200 can derive prediction samples based on the intra-prediction mode applied to the target block (S1110).
[0267] The encoding device 200 according to the embodiment can derive the residual sample of the target block (S1120).
[0268] The encoding device 200 according to the embodiment can derive the transform coefficients of the target block based on a single transform of the residual samples (S1130). A single transform can be performed using multiple transform kernels, and the transform kernel can be selected based on the intra-frame prediction mode.
[0269] The decoding device 300 can perform a secondary transformation on the transform coefficients of the target block, specifically NSST, in which case NSST can be performed based on reduced transform (RST) or not based on RST. When NSST is performed based on reduced transform, the operation according to S1140 can be performed.
[0270] According to the implementation, the encoding device 200 can derive the modified transform coefficients of the target block based on the RST of the transform coefficients (S1140). In the example, the RST can be performed based on the reduced transform matrix or the transform kernel matrix, and the reduced transform matrix can be a non-square matrix in which the number of rows is less than the number of columns.
[0271] In an implementation, S1140 may include: determining whether the conditions for applying RST are met; generating and encoding a transform index based on the determination; selecting a transform kernel; and when the conditions for applying RST are met, applying RST to the residual samples based on the selected transform kernel matrix and / or a reduction factor. In this case, the size of the reduced transform kernel matrix may be determined based on the reduction factor.
[0272] The encoding device 200 according to the embodiment can derive quantized transform coefficients by performing quantization based on the modified transform coefficients of the target block, and can encode information about the quantized transform coefficients (S1150).
[0273] Specifically, the encoding device 200 can generate information about the quantization transform coefficients and can encode the generated information about the quantization transform coefficients.
[0274] In the example, information about the quantization transform coefficients may include at least one of the following: information about whether RST was applied, information about the reduction factor, information about the minimum transform size with RST applied, and information about the maximum transform size with RST applied.
[0275] Referring to S1140, the transform coefficients of the target block can be derived using an RST based on residual samples. From the perspective of the transform kernel matrix size, since the size of the regular transform kernel matrix is N×N while the size of the reduced transform matrix is reduced to R×N, the memory usage can be reduced by an R / N ratio compared to performing a regular transform. Furthermore, compared to the number of multiplications (N×N) when using a regular transform kernel matrix, using a reduced transform kernel matrix can reduce the number of multiplications (R×N) by an R / N ratio. Additionally, since only R transform coefficients need to be derived when applying RST, the total number of transform coefficients of the target block can be reduced from N to R compared to deriving N transform coefficients when applying a regular transform, thus reducing the amount of data sent from the encoding device 200 to the decoding device 300. In other words, according to S1140, the transform efficiency and encoding efficiency of the encoding device 200 can be increased through RST.
[0276] Figure 12 This is a control flow diagram illustrating an embodiment of RST according to the present disclosure.
[0277] First, the encoding device 200 can determine the transform set based on the mapping relationship according to the intra-prediction mode applied to the target block (S1200).
[0278] Transformer 232 can derive the transformation coefficients by performing an RST based on any of the transformation kernel matrices included in the transformation set (S1210).
[0279] In this embodiment, the transform coefficients can be modified transform coefficients obtained by a first transform followed by a second transform, and each transform set may include two transform kernel matrices.
[0280] When RST is executed, information about RST can be encoded by the entropy encoder 240.
[0281] The entropy encoder 240 can derive the value of the syntax element that indicates the transform index of any of the transform kernel matrices included in the transform set (S1220).
[0282] The syntax elements of the transform index according to this embodiment can indicate whether (inverse) RST has been applied and any of the transform kernel matrices included in the transform set. When the transform set includes two transform kernel matrices, the syntax elements of the transform index can have three values.
[0283] According to the implementation, the value of the syntax element of the transform index can be deduced as: 0, which indicates that (inverse) RST is not applied to the target block; 1, which indicates the first transform kernel matrix of the transform kernel matrix; or 2, which indicates the second transform kernel matrix of the transform kernel matrix.
[0284] The entropy encoder 240 can binarize the derivation value of the syntax element of the transformation index (S1230).
[0285] The entropy encoder 240 can binarize the three values of the syntax element of the transform index into 0, 10, and 11 according to the truncated unary code. That is, the value 0 of the syntax element can be binarized into "0", the value 1 of the syntax element can be binarized into "10", and the value 2 of the syntax element can be binarized into "11", and the entropy encoder 240 can binarize the syntax element of the derived transform index into one of "0", "10", and "11".
[0286] The entropy encoder 240 can deduce context information (i.e., context model) about the bin string of the transformation index (S1240), and can encode the bin of the bin string of the syntax element based on the context information (S1250).
[0287] According to this embodiment, the context information of different entries can be applied to the two bins of the transform index respectively. That is, all two bins of the transform index can be encoded using a context method instead of a bypass method, wherein the first bin of the transform index's syntax element bins can be decoded based on the first context information, and the second bin of the transform index's syntax element bins can be decoded based on the second context information.
[0288] The encoded bin string of the syntax elements can be output as a bitstream to the decoding device 300 or an external device.
[0289] In the above embodiments, the method is explained based on a flowchart using a series of steps or blocks. However, this disclosure is not limited to the order of the steps, and a step may be performed in a different order or sequence than described above, or a step may be performed concurrently with other steps. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps in the flowchart may be deleted without affecting the scope of this disclosure.
[0290] The methods described above according to this disclosure can be implemented in software form, and the encoding and / or decoding devices according to this disclosure can be included in devices for image processing such as televisions, computers, smartphones, set-top boxes, and display devices.
[0291] When the embodiments of this disclosure are implemented by software, the above methods can be implemented as modules (steps, functions, etc.) for performing the above functions. These modules can be stored in memory and can be executed by a processor. The memory can be internal or external to the processor and can be connected to the processor in various well-known ways. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0292] Furthermore, the decoding and encoding devices using this disclosure can include multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices (such as video communication), mobile streaming devices, storage media, portable video cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, and medical video devices, and can be used to process video signals or data signals. For example, over-the-top (OTT) video devices can include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0293] Furthermore, the processing methods of this disclosure can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having the data structure according to this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include various storage devices and distributed storage devices for storing computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of a carrier wave (e.g., transmission over the Internet). Furthermore, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks. Additionally, embodiments of this disclosure can be implemented as computer program products by program code, and the program code can be executed on a computer according to embodiments of this disclosure. The program code can be stored on a computer-readable carrier.
[0294] Figure 13The structure of a content streaming system applying this disclosure is illustrated.
[0295] Furthermore, the content streaming system using this disclosure can generally include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0296] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then sends it to a streaming server. As another example, in cases where the multimedia input device, such as a smartphone, camera, or camcorder, directly generates the bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method disclosed herein. Furthermore, the streaming server can temporarily store the bitstream during the sending or receiving process.
[0297] The streaming server sends multimedia data to the user's device via a web server based on the user's request. The web server acts as a tool to notify the user of available services. When a user requests a desired service, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this context, the content streaming system may include a separate control server, which in this case controls the commands / responses between the corresponding devices within the content streaming system.
[0298] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0299] For example, user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, board-type PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. The servers in the content streaming system can operate as distributed servers, and in this case, data received by each server can be processed in a distributed manner.
Claims
1. A decoding device for image decoding, the decoding device comprising: Memory; as well as At least one processor connected to the memory, the at least one processor being configured to: Derive the quantization transform coefficients for the target block from the bitstream; The transformation coefficients are derived by dequantization based on the quantization transformation coefficients used for the target block; The modified transformation coefficients are derived based on the inverse quadratic transformation used for the transformation coefficients; The residual samples for the target block are derived based on the inverse first-order transform of the transformed coefficients used for the modification; and A reconstructed image is generated based on the residual samples used for the target block. The inverse quadratic transformation is performed based on the transformation kernel matrix. The transformation kernel matrix is obtained based on the transformation index and the transformation set. Wherein, the transformation index represents at least one of first index information, second index information, and third index information; the first index information indicates that the inverse quadratic transform is not applied to the target block; the second index information indicates that the first transform kernel matrix is the transform kernel matrix used for the inverse quadratic transform; and the third index information indicates that the second transform kernel matrix is the transform kernel matrix used for the inverse quadratic transform. The transform set is determined based on the intra-prediction mode of the target block. Wherein, based on the intra-frame prediction mode being equal to 0 or 1, the transform set is determined as the first transform set index value. Specifically, based on the intra-frame prediction mode being greater than or equal to 2 and less than or equal to 12, the transform set is determined as the second transform set index value. Specifically, based on the intra-frame prediction mode being greater than or equal to 13 and less than or equal to 23, the transform set is determined as the third transform set index value. Specifically, based on the intra-frame prediction mode being greater than or equal to 24 and less than or equal to 44, the transform set is determined as the fourth transform set index value. Wherein, based on the intra-frame prediction mode being greater than or equal to 45 and less than or equal to 55, the transform set is determined as the index value of the third transform set, and Wherein, based on the intra-frame prediction mode being greater than or equal to 56, the transform set is determined as the index value of the second transform set.
2. An encoding device for image encoding, the encoding device comprising: Memory; as well as At least one processor connected to the memory, the at least one processor being configured to: Predicted samples are derived based on the intra-prediction mode applied to the target block; Based on the predicted samples, residual samples for the target block are derived; The transformation coefficients for the target block are derived based on a first transformation applied to the residual sample. The modified transformation coefficients are derived based on the quadratic transformation of the transformation coefficients. The quantized transform coefficients are derived by performing quantization based on the modified transform coefficients. Determine the transformation set and the transformation kernel matrix; as well as A transformation index is generated representing at least one of a first index, a second index, and a third index, wherein the first index indicates that the inverse quadratic transform is not applied to the target block, the second index indicates that the first transform kernel matrix is the transform kernel matrix used for the inverse quadratic transform, and the third index indicates that the second transform kernel matrix is the transform kernel matrix used for the inverse quadratic transform. The transform set is determined based on the intra-prediction mode of the target block. Wherein, based on the intra-frame prediction mode being equal to 0 or 1, the transform set is determined as the first transform set index value. Specifically, based on the intra-frame prediction mode being greater than or equal to 2 and less than or equal to 12, the transform set is determined as the second transform set index value. Specifically, based on the intra-frame prediction mode being greater than or equal to 13 and less than or equal to 23, the transform set is determined as the third transform set index value. Specifically, based on the intra-frame prediction mode being greater than or equal to 24 and less than or equal to 44, the transform set is determined as the fourth transform set index value. Wherein, based on the intra-frame prediction mode being greater than or equal to 45 and less than or equal to 55, the transform set is determined as the index value of the third transform set, and Wherein, based on the intra-frame prediction mode being greater than or equal to 56, the transform set is determined as the index value of the second transform set.
3. An apparatus for transmitting data for an image, the apparatus comprising: At least one processor is configured to perform the following steps to obtain a bitstream for the image: deriving prediction samples based on an intra-prediction mode applied to a target block; deriving residual samples of the target block based on the prediction samples; deriving transform coefficients for the target block based on a first transform for the residual samples; deriving modified transform coefficients based on a second transform of the transform coefficients; deriving quantized transform coefficients by performing quantization based on the modified transform coefficients; determining a transform set and a transform kernel matrix; and generating a transform index representing at least one of first index information, second index information, and third index information, wherein the first index information indicates that the inverse second transform is not applied to the target block, the second index information indicates the first transform kernel matrix as the transform kernel matrix for the inverse second transform, and the third index information indicates the second transform kernel matrix as the transform kernel matrix for the inverse second transform; and A transmitter configured to send the data including the bit stream. The transform set is determined based on the intra-prediction mode of the target block. Wherein, based on the intra-frame prediction mode being equal to 0 or 1, the transform set is determined as the first transform set index value. Specifically, based on the intra-frame prediction mode being greater than or equal to 2 and less than or equal to 12, the transform set is determined as the second transform set index value. Specifically, based on the intra-frame prediction mode being greater than or equal to 13 and less than or equal to 23, the transform set is determined as the third transform set index value. Specifically, based on the intra-frame prediction mode being greater than or equal to 24 and less than or equal to 44, the transform set is determined as the fourth transform set index value. Wherein, based on the intra-frame prediction mode being greater than or equal to 45 and less than or equal to 55, the transform set is determined as the index value of the third transform set, and Wherein, based on the intra-frame prediction mode being greater than or equal to 56, the transform set is determined as the index value of the second transform set.
Citation Information
Patent Citations
Multi-core parallel video decoding method for allocating tasks and data by row in staggered manner
CN105376583A
Template matching-based encoding and decoding method and device
WO2018119609A1