Image decoding apparatus, image encoding apparatus, and data transmission apparatus
By dividing the image block into sub-partition transformation blocks and applying the LFNST matrix and index selection transformation coefficients, the efficient encoding problem of high-resolution images/videos is solved, and encoding efficiency and storage costs are improved.
Patent Information
- Application Number
- CN202510679975.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-20
- Filing Date
- 2020-09-15
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art When transmitting and storing high resolution, high-quality images/videos, the increase in the amount of information leads to high costs, and the image/video broadcasting requirements for immersive media and real images are increasing, requiring efficient compression and encoding technologies.
By using the image encoding method, by dividing the current block into sub-partition transformation blocks, selecting the transformation coefficients using the LFNST matrix and index, deriving the residual sample, and increasing the encoding efficiency.
Improves image/video compression efficiency and transform index coding efficiency, suitable for high resolution and high-quality image/video encoding.
Smart Images

Figure CN120263979A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the original application number 202080065580.6 (International Application Number: PCT / KR2020 / 012405, Application Date: September 15, 2020, Invention Title: Transform-Based Image Coding Method and Apparatus). Technical Field
[0002] The present disclosure relates to an image coding technology, and more particularly, to a method and apparatus for coding an image based on a transform in an image coding system. Background Art
[0003] Nowadays, the demand for high-resolution and high-quality images / videos such as 4K, 8K, or higher ultra-high definition (UHD) images / videos has been continuously increasing in various fields. As image / video data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases compared to traditional image data. Therefore, when transmitting image data using a medium such as a traditional wired / wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.
[0004] In addition, nowadays, the interest in and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing, and the broadcasting of images / videos having image characteristics different from those of real images such as game images is increasing.
[0005] Therefore, there is a need for an efficient image / video compression technology that can effectively compress, transmit, store, and reproduce information of high-resolution and high-quality images / videos having various characteristics as described above. Summary of the Invention
[0006] Technical Problem
[0007] One technical aspect of the present disclosure is to provide a method and apparatus for increasing image coding efficiency.
[0008] Another technical aspect of the present disclosure is to provide a method and apparatus for increasing the efficiency of transform index coding.
[0009] Yet another technical aspect of the present disclosure is to provide an image coding method and apparatus using LFNST.
[0010] Yet another technical aspect of the present disclosure is to provide a method and apparatus for coding an image for applying LFNST to a sub-partition transform block.
[0011] Technical Solution
[0012] In one aspect, a method for decoding an image performed by a decoding device is provided. The method includes: when dividing a current block into sub-partition transform blocks, deriving prediction samples of the current block based on intra prediction mode information; determining an LFNST set including an LFNST matrix based on the intra prediction mode derived from the intra prediction mode information; selecting one of the LFNST matrices based on the LFNST set and an LFNST index; deriving transform coefficients of the sub-partition transform blocks based on the selected LFNST matrix; and deriving residual samples of the current block based on the transform coefficients.
[0013] The same LFNST set and the same LFNST index can be applied to the sub-partition transform blocks divided from the current block.
[0014] The same intra prediction mode can be applied to the sub-partition transform blocks divided from the current block, and prediction samples can be derived for each sub-partition transform block.
[0015] When the width and height of the sub-partition transform block are 4 or greater, the transform coefficients can be derived based on the selected LFNST matrix.
[0016] When the size (width × height) of the current block is 8 × 4, the current block can be vertically divided, and when the size (width × height) of the current block is 4 × 8, the current block can be horizontally divided.
[0017] When the size (width × height) of the current block is greater than 4 × 8 or 8 × 4, the current block can be divided into four sub-partition transform blocks in the horizontal or vertical direction.
[0018] According to an embodiment of this document, a method for encoding an image performed by an encoding device is provided. The method includes: when dividing a current block into sub-partition transform blocks, deriving prediction samples of each sub-partition transform block based on the intra prediction mode applied to the current block; deriving residual samples of the current block based on the prediction samples; applying a transform once to the residual samples to derive transform coefficients; and deriving modified transform coefficients of the sub-partition transform blocks based on the LFNST set mapped to the intra prediction mode and the LFNST matrix included in the LFNST set.
[0019] According to another embodiment of the present disclosure, a digital storage medium can be provided, which stores image data including a bitstream and encoded image information generated according to an image encoding method performed by an encoding device.
[0020] According to another embodiment of the present disclosure, a digital storage medium can be provided, which stores image data including encoded image information and a bitstream to enable a decoding device to perform an image decoding method.
[0021] Technical Effects
[0022] According to the present disclosure, the overall image / video compression efficiency can be increased.
[0023] According to the present disclosure, the efficiency of transform index coding can be increased.
[0024] The technical aspects of the present disclosure can provide an image coding method and apparatus using LFNST.
[0025] The technical aspects of the present disclosure can provide a method and apparatus for coding an image for applying LFNST to a sub-partition transform block.
[0026] The effects that can be obtained through specific examples of the present disclosure are not limited to the effects listed above. For example, there may be various technical effects that can be understood by those of ordinary skill in the relevant art or derived from the present disclosure. Therefore, the specific effects of the present disclosure are not limited to those explicitly described in the present disclosure, and may include various effects that can be understood or derived based on the technical features of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 An example of a video / image coding system to which the present disclosure can be applied is schematically illustrated.
[0028] Figure 2 is a diagram schematically illustrating the configuration of a video / image coding apparatus to which the present disclosure can be applied.
[0029] Figure 3 is a diagram schematically illustrating the configuration of a video / image decoding apparatus to which the present disclosure can be applied.
[0030] Figure 4 A multi-transform scheme according to an embodiment of the present document is schematically illustrated.
[0031] Figure 5 An intra-frame directional mode of 65 prediction directions is schematically shown.
[0032] Figure 6 is a diagram for explaining RST according to an embodiment of the present document.
[0033] Figure 7 is a diagram illustrating the order of arranging the output data of a forward single transform as a one-dimensional vector according to an example.
[0034] Figure 8 is a diagram illustrating the order of arranging the output data of a forward double transform as a two-dimensional vector according to an example.
[0035] Figure 9 is a diagram illustrating a wide-angle intra-frame prediction mode according to an embodiment of the present document.
[0036] Figure 10 It is a diagram illustrating the block shapes to which the LFNST is applied.
[0037] Figure 11 It is a diagram illustrating the arrangement of the output data of the forward LFNST according to an embodiment.
[0038] Figure 12 It is a diagram illustrating that the number of the output data of the forward LFNST according to an example is limited to a maximum value of 16.
[0039] Figure 13 It is a diagram illustrating the zeroing in the block where the 4×4 LFNST is applied according to an example.
[0040] Figure 14 It is a diagram illustrating the zeroing in the block where the 8×8 LFNST is applied according to an example.
[0041] Figure 15 It is a diagram illustrating the zeroing in the block where the 8×8 LFNST is applied according to another example.
[0042] Figure 16 It is a diagram illustrating an example of the sub - blocks into which one coded block is divided.
[0043] Figure 17 It is a diagram illustrating another example of the sub - blocks into which one coded block is divided.
[0044] Figure 18 It is a diagram illustrating the symmetry between the M×2 (M×1) block and the 2×M (1×M) block according to an embodiment.
[0045] Figure 19 It is a diagram illustrating an example of transposing the 2×M block according to an embodiment.
[0046] Figure 20 It illustrates the scanning order of the 8×2 or 2×8 region according to an embodiment.
[0047] Figure 21 It is a flowchart illustrating a method for decoding an image according to an embodiment.
[0048] Figure 22 It is a flowchart illustrating a method for encoding an image according to an embodiment.
[0049] Figure 23 It is a diagram illustrating the structure of the content stream system to which the content of this document is applied. Detailed implementation manners
[0050] Although the present disclosure may be susceptible to various modifications and include various embodiments, specific embodiments thereof have been shown by way of example in the drawings and will now be described in detail. However, this is not intended to limit the present disclosure to the specific embodiments disclosed herein. The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the technical concept of the present disclosure. Unless the context clearly indicates otherwise, the singular forms may include the plural forms. Terms such as "including" and "having" are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and thus should not be construed as precluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0051] In addition, for the convenience of describing different characteristic functions from each other, each component in the drawings described herein is illustrated independently. However, it is not meant that each component is implemented by separate hardware or software. For example, any two or more of these components may be combined to form a single component, and any single component may be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of the patent right of the present disclosure as long as they do not depart from the essence of the present disclosure.
[0052] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the drawings. In addition, in the drawings, the same reference numerals are used for the same components, and repeated descriptions of the same components will be omitted.
[0053] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), the EVC (Essential Video Coding) standard, the AVS2 standard, etc.).
[0054] In this document, various embodiments related to video / image coding may be provided, and unless otherwise specified, these embodiments may be combined with each other and executed.
[0055] In this document, video may refer to a collection of a series of images over a period of time. Generally, a picture refers to a unit representing an image of a specific time region, and a slice / tile is a unit that forms part of a picture. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / titles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles.
[0056] A pixel or picture element (pel) may refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. Alternatively, a sample can mean a pixel value in the spatial domain, or when the pixel value is transformed into the frequency domain, it can mean a transform coefficient in the frequency domain.
[0057] A unit can represent the basic unit of image processing. A unit can include at least one of a specific region and information related to the region. A unit can include one luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the situation, terms such as unit and terms like block, region, etc. can be used interchangeably. Usually, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients composed of M columns and N rows.
[0058] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" can mean "A and / or B". Additionally, "A, B" can mean "A and / or B". Additionally, "A / B / C" can mean "at least one of A, B, and / or C". Additionally, "A / B / C" can mean "at least one of A, B, and / or C".
[0059] Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" can include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".
[0060] In this disclosure, "at least one of A and B" can mean "only A", "only B", or "both A and B". Additionally, in this disclosure, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".
[0061] Furthermore, in this disclosure, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0062] In addition, the parentheses used in the present disclosure may indicate "for example". Specifically, when it is indicated as "prediction (intra prediction)", it may mean that "intra prediction" is presented as an example of "prediction". In other words, "prediction" in the present disclosure is not limited to "intra prediction", and "intra prediction" is presented as an example of "prediction". In addition, when it is indicated as "prediction (i.e., intra prediction)", this may also mean that "intra prediction" is presented as an example of "prediction".
[0063] The technical features separately described in one of the drawings in the present disclosure may be implemented separately or may be implemented simultaneously.
[0064] Figure 1 Examples of a video / image encoding system to which the present disclosure can be applied are schematically illustrated.
[0065] Referring to Figure 1 , the video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transfer the encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.
[0066] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0067] The video source may obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source may include a video / image capturing device and / or a video / image generating device. The video / image capturing device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generating device may include, for example, a computer, a tablet computer, and a smart phone, and may (electronically) generate video / images. For example, virtual video / images may be generated by a computer or the like. In this case, the video / image capturing process may be replaced by a process of generating relevant data.
[0068] The encoding device may encode the input video / images. The encoding device may perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.
[0069] The transmitter can send the encoded video / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating a media file in a predetermined file format and can include elements for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to the decoding device.
[0070] The decoding device can decode the video / image by performing a series of processes such as dequantization, inverse transformation, prediction, etc. corresponding to the operations of the encoding device.
[0071] The renderer can render the decoded video / image. The rendered video / image can be displayed via a display.
[0072] Figure 2 FIG. is a diagram schematically illustrating the configuration of a video / image encoding device to which the present disclosure can be applied. Hereinafter, the so-called video encoding device may include an image encoding device.
[0073] Referring to Figure 2 , the encoding device 200 can include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, an inverse transformer 235. The residual processor 230 can further include a subtractor 231. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be constituted by one or more hardware components (e.g., an encoder chipset or a processor). In addition, the memory 270 can include a decoded picture buffer (DPB) and can be constituted by a digital storage medium. The hardware components can further include the memory 270 as an internal / external component.
[0074] The image divider 210 may divide an input image (or picture or frame) input to the encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a largest coding unit (LCU), the coding unit may be recursively divided according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, based on a quadtree structure, a binary tree structure, and / or a ternary tree structure, one coding unit may be divided into a plurality of coding units with a deeper depth. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or the ternary tree structure may be applied later. Alternatively, the binary tree structure may be applied first. The encoding process according to the present disclosure may be performed based on the final coding unit that is not further divided. In this case, based on the encoding efficiency according to the image characteristics, the largest coding unit may be directly used as the final coding unit. Alternatively, the coding unit may be recursively divided into coding units with a deeper depth as needed, whereby the coding unit with the optimal size may be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, a processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be separate from or divided from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.
[0075] Depending on the situation, the terms unit and terms such as block, region, etc. may be used in place of each other. In general, an M×N block may represent a set of samples or transformation coefficients composed of M columns and N rows. Samples generally may represent pixels or pixel values, and may represent only the pixels / pixel values of the luminance component, or only the pixels / pixel values of the chrominance component. Samples may be used as a term corresponding to the pixels or pels of a picture (or image).
[0076] The subtractor 231 subtracts the prediction signal (prediction block, prediction sample array) output from the predictor 220 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. The predictor 220 can perform prediction on the block to be processed (hereinafter referred to as "current block") and can generate a prediction block including the prediction samples of the current block. The predictor 220 can determine whether to apply intra prediction or inter prediction based on the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various information related to prediction such as prediction mode information and send the generated information to the entropy encoder 240. The information about prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0077] The intra predictor 222 can predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the reference samples can be located near the current block or separated from the current block. In intra prediction, the prediction mode can include multiple non - directional modes and multiple directional modes. The non - directional modes can include, for example, the DC mode and the planar mode. Depending on the level of detail of the prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used depending on the settings. The intra predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0078] The inter - frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted on a block, sub - block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, the neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same as each other or different from each other. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter - frame predictor 221 can configure a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter - frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, the residual signal cannot be transmitted. In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of a neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0079] The predictor 220 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra - frame prediction or inter - frame prediction to the prediction of a block, and can also apply intra - frame prediction and inter - frame prediction simultaneously. This can be referred to as combined inter - frame and intra - frame prediction (CIIP). Additionally, the predictor can be based on the intra - block copy (IBC) prediction mode or the palette mode in order to perform prediction on the block. The IBC prediction mode or the palette mode can be used for content image / video coding such as in games like screen content coding (SCC). Although IBC basically performs prediction in the current block, the way it is performed is similar to inter - frame prediction in that it derives a reference block in the current block. That is, IBC can use at least one of the inter - frame prediction techniques described in this disclosure.
[0080] The prediction signal generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT means a transform obtained from a graph when relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform processing can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size rather than square blocks.
[0081] Quantizer 233 may quantize the transform coefficients and send them to entropy encoder 240, and entropy encoder 240 may encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream. The information about the quantized transform coefficients may be referred to as residual information. Quantizer 233 may rearrange the quantized transform coefficients of the block type into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Entropy encoder 240 may perform various coding methods such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. Entropy encoder 240 may encode, together or separately, information required for video / image reconstruction other than the quantized transform coefficients (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) may be sent or stored in the form of a bitstream on a unit basis of a network abstraction layer (NAL). The video / image information may also include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc. Additionally, the video / image information may also include general constraint information. In the present disclosure, the information and / or syntax elements sent from the encoding device to / signaled to the decoding device may be included in the video / image information. The video / image information may be encoded through the above encoding process and included in the bitstream. The bitstream may be transmitted through a network or stored in a digital storage medium. Here, the network may include a broadcast network, a communication network, and / or the like, and the digital storage medium may include various storage media such as a USB, an SD, a CD, a DVD, a Blu-ray, an HDD, an SSD, etc. A transmitter (not shown) that sends the signal output from entropy encoder 240 or a memory (not shown) that stores it may be configured as an internal / external element of encoding device 200, or the transmitter may be included in entropy encoder 240.
[0082] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transformation to the quantized transform coefficients using the dequantizer 234 and the inverse transformer 235, a residual signal (residual block or residual samples) can be reconstructed. The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222, enabling the generation of a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the processing target block as in the case of applying the skip mode, the prediction block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the target picture and, as described subsequently, can be used for inter-frame prediction of the next picture through filtering.
[0083] In addition, in the picture encoding and / or reconstruction process, a luminance mapping with chroma scaling (LMCS) can be applied.
[0084] The filter 260 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and can store the modified reconstructed picture in the memory 270, particularly in the DPB of the memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. As discussed subsequently in the description of each filtering method, the filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240. The information about the filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0085] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-frame predictor 221. Accordingly, the encoding device can avoid prediction mismatches in the encoding device 100 and the decoding device when applying inter-frame prediction and can also improve the encoding efficiency.
[0086] The memory 270 DPB can store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the blocks in the current picture from which the motion information has been derived (or encoded) and / or the motion information of the blocks in the already reconstructed pictures. The stored motion information can be sent to the inter-frame predictor 221 to be used as the motion information of neighboring blocks or temporally neighboring blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 222.
[0087] Figure 3It is a diagram schematically illustrating the configuration of a video / image decoding device to which the present disclosure can be applied.
[0088] Referring to Figure 3 , the video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra predictor 331 and an inter predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be constituted by one or more hardware components (e.g., a decoder chipset or a processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0089] When receiving a bitstream including video / image information, the decoding device 300 may reconstruct an image corresponding to the processing of the video / image information that has been processed in the Figure 2 encoding device accordingly. For example, the decoding device 300 may derive units / blocks based on information related to block segmentation obtained from the bitstream. The decoding device 300 may perform decoding by using the processing units applied in the encoding device. Thus, the decoded processing unit may be, for example, an encoding unit, which may be divided along a quadtree structure, a binary tree structure, and / or a ternary tree structure with a coding tree unit or a largest coding unit. One or more transform units may be derived with the encoding unit. And, the reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproducer.
[0090] The decoding device 300 may receive in the form of a bitstream from Figure 2The signal output by the encoding device, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), video parameter sets (VPS), etc. Additionally, the video / image information can also include conventional constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the conventional constraint information. In the present disclosure, the signaled / received information and / or syntax elements described subsequently can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the decoding target syntax element information and the decoding information of the neighboring and decoding target blocks or the information of the symbols / bins decoded in the previous step to determine the context model, predict the bin generation probability according to the determined context model, and perform arithmetic decoding on the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model using the information of the symbols / bins decoded by the context model for the next symbol / bin after determining the context model. Among the information decoded in the entropy decoder 310, the information about prediction can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantized transform coefficients) for which entropy decoding has been performed in the entropy decoder 310 and the associated parameter information can be input to the residual processor 320. The residual processor 320 can derive the residual signal (residual block, residual sample, residual sample array). Additionally, the information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) that receives the signal output by the encoding device can also configure the decoding device 300 as an internal / external component, and the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to the present disclosure can be referred to as a video / image / picture encoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0091] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into the form of a two-dimensional block. In this case, the rearrangement can be performed based on the order of coefficient scanning that has been performed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.
[0092] The inverse transformer 322 obtains a residual signal (residual block, residual sample array) by performing an inverse transform on the transform coefficients.
[0093] The predictor can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and specifically can determine the intra / inter prediction mode.
[0094] The predictor can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction to the prediction of a block, and can also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). Additionally, the predictor can perform intra block copy (IBC) for the prediction of a block. Intra block copy can be used for content image / video coding such as games like screen content coding (SCC). Although IBC basically performs prediction in the current block, the way it is performed is similar to inter prediction in that it derives a reference block in the current block. That is, IBC can use at least one of the inter prediction techniques described in this disclosure.
[0095] The intra predictor 331 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the reference samples can be located near the current block or separated from the current block. In intra prediction, the prediction mode can include multiple non - directional modes and multiple directional modes. The intra predictor 331 can determine the prediction mode to be applied to the current block by using the prediction mode applied to neighboring blocks.
[0096] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information can be predicted on a block, sub-block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction can include information indicating the mode of inter-frame prediction for the current block.
[0097] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor 330. When there is no residual for the target block to be processed, as in the case of applying the skip mode, the prediction block can be used as the reconstructed block.
[0098] The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block in the current block, and as described later, can be output through filtering or used for inter-frame prediction of the next picture.
[0099] In addition, in the picture decoding process, a luminance mapping with chroma scaling (LMCS) can be applied.
[0100] The filter 350 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can send the modified reconstructed picture to the memory 360, especially to the DPB of the memory 360. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0101] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter - predictor 332. The memory 360 can store the motion information of blocks in the current picture from which the motion information has been derived (or decoded) and / or the motion information of blocks in the reconstructed pictures. The stored motion information can be sent to the inter - predictor 332 to be used as the motion information of neighboring blocks or temporally neighboring blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra - predictor 331.
[0102] In this specification, the examples described in the predictor 330, de - quantizer 321, inverse transformer 322, and filter 350 of the decoding device 300 can be similarly or correspondingly applied to the predictor 220, de - quantizer 234, inverse transformer 235, and filter 260 of the encoding device 200, respectively.
[0103] As described above, prediction is performed to improve the compression efficiency when performing video encoding. Accordingly, a prediction block including prediction samples for the current block, which is the block to be encoded, can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived identically in the encoding device and the decoding device, and the encoding device can improve the image encoding efficiency by signaling to the decoding device not the original sample values of the original block itself but the information about the residual between the original block and the prediction block (residual information). The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and generate a reconstructed picture including the reconstructed block.
[0104] The residual information can be generated through a transformation process and a quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transformation process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients, so that it can signal the associated residual information to the decoding device (through a bitstream). Here, the residual information can include value information, position information, transformation technique, transformation kernel, quantization parameter, etc. of the quantized transform coefficients. The decoding device can perform a quantization / de - quantization process based on the residual information and derive residual samples (or a residual sample block). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can de - quantize / inverse - transform the quantized transform coefficients to derive a residual block for use as a reference for inter - prediction of the next picture, and can generate a reconstructed picture based on this.
[0105] Figure 4 A multi - transformation technique according to an embodiment of the present disclosure is schematically illustrated.
[0106] Reference Figure 4 , the transformer can correspond to the transformer in the aforementioned Figure 2 encoding device, and the inverse transformer can correspond to the inverse transformer in the aforementioned Figure 2 encoding device, or Figure 3 the inverse transformer in the decoding device.
[0107] The transformer can derive (primary) transform coefficients (S410) by performing a single transformation based on the residual samples (residual sample array) in the residual block. This single transformation can be referred to as the core transformation. In this document, the single transformation can be based on Multiple Transform Selection (MTS), and when multiple transforms are used as the single transformation, it can be referred to as a multi-core transformation.
[0108] The multi-core transformation can represent a method of additionally using Discrete Cosine Transform (DCT) type 2 and Discrete Sine Transform (DST) type 7, DCT type 8, and / or DST type 1 for transformation. That is, the multi-core transformation can represent a transformation method of transforming a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. In this document, from the perspective of the transformer, the primary transform coefficients can be referred to as temporary transform coefficients.
[0109] In other words, when applying a conventional transformation method, transform coefficients can be generated by applying a transformation from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2. In contrast, when applying a multi-core transformation, transform coefficients (or primary transform coefficients) can be generated by applying a transformation from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. In this document, DCT type 2, DST type 7, DCT type 8, and DST type 1 can be referred to as transformation types, transform kernels, or transform cores. These DCT / DST transformation types can be defined based on basis functions.
[0110] When performing a multi-core transformation, a vertical transform kernel and a horizontal transform kernel for the target block can be selected from the transform kernels, a vertical transformation can be performed on the target block based on the vertical transform kernel, and a horizontal transformation can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transformation can indicate the transformation of the horizontal component of the target block, and the vertical transformation can indicate the transformation of the vertical component of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target (CU or sub-block) including the residual block.
[0111] In addition, according to the example, if a transformation is performed by applying MTS, the mapping relationship of the transformation kernel can be set by setting a specific basis function to a predetermined value and combining the basis functions to be applied in the vertical transformation or the horizontal transformation. For example, when the horizontal transformation kernel is represented as trTypeHor and the vertical transformation kernel is represented as trTypeVer, trTypeHor or trTypeVer with a value of 0 can be set to DCT2, trTypeHor or trTypeVer with a value of 1 can be set to DST7, and trTypeHor or trTypeVer with a value of 2 can be set to DCT8.
[0112] In this case, the MTS index information can be encoded and signaled to the decoding device to indicate any one of the multiple transformation kernel sets. For example, MTS index 0 can indicate that both the trTypeHor and trTypeVer values are 0, MTS index 1 can indicate that both the trTypeHor and trTypeVer values are 1, MTS index 2 can indicate that the trTypeHor value is 2 and the trTypeVer value is 1, MTS index 3 can indicate that the trTypeHor value is 1 and the trTypeVer value is 2, and MTS index 4 can indicate that both the trTypeHor and trTypeVer values are 2.
[0113] In one example, the transformation kernel sets according to the MTS index information are shown in the following table.
[0114] [Table 1]
[0115] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2
[0116] The transformer can perform a secondary transform based on the (primary) transform coefficients to derive modified (secondary) transform coefficients (S420). The primary transform is a transform from the spatial domain to the frequency domain, and the secondary transform refers to a transform to a more compact representation using the correlation existing between the (primary) transform coefficients. The secondary transform can include an inseparable transform. In this case, the secondary transform can be referred to as an inseparable secondary transform (NSST) or a mode-dependent inseparable secondary transform (MDNSST). The NSST can represent a transform that performs a secondary transform on the (primary) transform coefficients derived through the primary transform based on an inseparable transform matrix to generate modified transform coefficients (or secondary transform coefficients) for the residual signal. Here, based on the inseparable transform matrix, the transform can be applied once to the (primary) transform coefficients without separating the vertical transform and the horizontal transform (or applying the horizontal / vertical transform independently). In other words, the NSST does not separately apply to the (primary) transform coefficients in the vertical and horizontal directions, and can represent, for example, a transform method of rearranging a two-dimensional signal (transform coefficients) into a one-dimensional signal in a specific predetermined direction (e.g., row-major order or column-major order) and then generating modified transform coefficients (or secondary transform coefficients) based on the inseparable transform matrix. For example, the row-major order is set in rows in the order of the first row, the second row, …, and the Nth row for M×N blocks, and the column-major order is set in rows in the order of the first column, the second column, …, and the Mth column for M×N blocks. The NSST can be applied to the upper left region of a block configured with (primary) transform coefficients (hereinafter referred to as a transform coefficient block). For example, when both the width W and the height H of the transform coefficient block are 8 or more, an 8×8 NSST can be applied to the upper left 8×8 region of the transform coefficient block. In addition, while both the width (W) and the height (H) of the transform coefficient block are 4 or more, when the width (W) or the height (H) of the transform coefficient block is less than 8, a 4×4 NSST can be applied to the upper left min(8, W)×min(8, H) region of the transform coefficient block. However, the embodiments are not limited thereto. For example, even if only the condition that the width W or the height H of the transform coefficient block is 4 or more is satisfied, a 4×4 NSST can be applied to the upper left end min(8, W)×min(8, H) region of the transform coefficient block.
[0117] Specifically, for example, if a 4×4 input block is used, the inseparable secondary transform can be performed as follows.
[0118] The 4×4 input block X can be represented as follows.
[0119] [Equation 1]
[0120]
[0121] If X is represented in the form of a vector, the vector It can be expressed as follows.
[0122] [Equation 2]
[0123]
[0124] In Equation 2, the vector is a one-dimensional vector obtained by rearranging the two-dimensional block X of Equation 1 in row-major order.
[0125] In this case, the inseparable quadratic transform can be calculated as follows.
[0126] [Equation 3]
[0127]
[0128] In this equation, represents the transform coefficient vector, and T represents a 16×16 (inseparable) transform matrix.
[0129] Through the aforementioned Equation 3, the 16×1 transform coefficient vector can be derived, and the vector can be reorganized into 4×4 blocks in a scanning order (such as horizontal, vertical, and diagonal). However, the above calculation is an example, and the hypercube-Givens transform (HyGT), etc., can also be used for the calculation of the inseparable quadratic transform to reduce the computational complexity of the inseparable quadratic transform.
[0130] In addition, in the inseparable quadratic transform, the transform kernel (or transform core, transform type) can be selected to be mode-dependent. In this case, the mode can include an intra prediction mode and / or an inter prediction mode.
[0131] As described above, the inseparable quadratic transform can be performed based on an 8×8 transform or a 4×4 transform determined based on the width (W) and height (H) of the transform coefficient block. The 8×8 transform refers to a transform that can be applied to an 8×8 region included in the transform coefficient block when both W and H are equal to or greater than 8, and the 8×8 region can be the upper left 8×8 region in the transform coefficient block. Similarly, the 4×4 transform refers to a transform that can be applied to a 4×4 region included in the transform coefficient block when both W and H are equal to or greater than 4, and the 4×4 region can be the upper left 4×4 region in the transform coefficient block. For example, the 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and the 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.
[0132] Here, in order to select mode-related transform kernels, two non-separable quadratic transform kernels can be configured for each transform set for non-separable quadratic transforms for both 8×8 transforms and 4×4 transforms, and there can be four transform sets. That is, four transform sets can be configured for 8×8 transforms, and four transform sets can be configured for 4×4 transforms. In this case, each of the four transform sets for 8×8 transforms can include two 8×8 transform kernels, and each of the four transform sets for 4×4 transforms can include two 4×4 transform kernels.
[0133] However, as the size of the transform (i.e., the size of the region to which the transform is applied) can be, for example, a size other than 8×8 or 4×4, the number of sets can be n, and the number of transform kernels in each set can be k.
[0134] The transform sets can be referred to as NSST sets or LFNST sets. A specific set among the transform sets can be selected, for example, based on the intra prediction mode of the current block (CU or sub-block). Low-frequency non-separable transform (LFNST) can be an example of a reduced non-separable transform, which will be described later, and represents a non-separable transform for low-frequency components.
[0135] As a reference, for example, the intra prediction mode can include two non-directional (or non-angle) intra prediction modes and 65 directional (or angle) intra prediction modes. The non-directional intra prediction modes can include the planar intra prediction mode numbered 0 and the DC intra prediction mode numbered 1, and the directional intra prediction modes can include 65 intra prediction modes numbered from 2 to 66. However, this is an example, and this document can be applied even if the number of intra prediction modes is different. In addition, in some cases, the intra prediction mode numbered 67 can also be used, and the intra prediction mode numbered 67 can represent a linear model (LM) mode.
[0136] Figure 5 The intra-directional mode with 65 prediction directions is schematically shown.
[0137] Referring to Figure 5 , based on the intra prediction mode 34 with the upper-left diagonal prediction direction, the intra prediction mode can be divided into an intra prediction mode with horizontal directivity and an intra prediction mode with vertical directivity. In Figure 5In this case, H and V respectively indicate horizontal directionality and vertical directionality, and the numbers -32 to 32 indicate displacements of 1 / 32 units at the sample grid positions. These numbers can represent offsets for the mode index values. Intra prediction modes 2 to 33 have horizontal directionality, and intra prediction modes 34 to 66 have vertical directionality. Strictly speaking, intra prediction mode 34 can be regarded as neither horizontal nor vertical, but can be classified as belonging to the horizontal directionality when determining the transform set of the second transform. This is because the input data is transposed for the vertically oriented mode symmetric to intra prediction mode 34, and the input data alignment method for the horizontal mode is used for intra prediction mode 34. Transposing the input data means switching the rows and columns of the two-dimensional M×N block data into N×M data. Intra prediction mode 18 and intra prediction mode 50 can respectively represent the horizontal intra prediction mode and the vertical intra prediction mode, and intra prediction mode 2 can be called the upper right diagonal intra prediction mode because intra prediction mode 2 has a left reference pixel and performs prediction in the upper right direction. Similarly, intra prediction mode 34 can be called the lower right diagonal intra prediction mode, and intra prediction mode 66 can be called the lower left diagonal intra prediction mode.
[0138] According to an example, four transform sets can be mapped according to the intra prediction mode, for example, as shown in the following table.
[0139] [Table 2]
[0140] lfnstPredModeIntra fnstIrSetIdx lfnstPredModeIntra<0 1 0 <= lfnstPredModeIntra <= 1 0 2 <= lfnstPredModeIntra <= 12 1 13 <= lfnstPredModeIntra <= 23 2 24 <= lfnstPredModeIntra <= 44 3 45 <= lfnstPredModeIntra <= 55 2 56 <= lfnstPredModeIntra <= 80 1 81 <= lfnstPredModeIntra <= 83 0
[0141] As shown in Table 2, any one of the four transform sets, i.e., lfnstTrSetIdx, can be mapped to any one of the four indices (i.e., 0 to 3) according to the intra prediction mode.
[0142] When determining a specific set for the non-separable transform, one of the k transform kernels in the specific set can be selected by the non-separable second transform index. The encoding device can derive the non-separable second transform index indicating the specific transform kernel based on rate distortion (RD) checking, and can signal the non-separable second transform index to the decoding device. The decoding device can select one of the k transform kernels in the specific set based on the non-separable second transform index. For example, the lfnst index value 0 can refer to the first non-separable second transform kernel, the lfnst index value 1 can refer to the second non-separable second transform kernel, and the lfnst index value 2 can refer to the third non-separable second transform kernel. Alternatively, the lfnst index value 0 can indicate that the first non-separable second transform is not applied to the target block, and the lfnst index values 1 to 3 can indicate three transform kernels.
[0143] The transformer can perform an inseparable quadratic transform based on the selected transform kernel and can obtain modified (quadratic) transform coefficients. As described above, the modified transform coefficients can be derived as the transform coefficients quantized by the quantizer, and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.
[0144] In addition, as described above, if the quadratic transform is omitted, the (primary) transform coefficients, which are the output of the primary (separable) transform, can be derived as the transform coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.
[0145] The inverse transformer can perform a series of processes in an order opposite to the order already performed in the above-mentioned transformer. The inverse transformer can receive the (dequantized) transform coefficients and derive the (primary) transform coefficients by performing a quadratic (inverse) transform (S450), and can obtain the residual block (residual samples) by performing a primary (inverse) transform on the (primary) transform coefficients (S460). In this regard, from the perspective of the inverse transformer, the primary transform coefficients can be referred to as modified transform coefficients. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0146] The decoding device may further include a quadratic inverse transform application determiner (or an element for determining whether to apply the quadratic inverse transform) and a quadratic inverse transform determiner (or an element for determining the quadratic inverse transform). The quadratic inverse transform application determiner can determine whether to apply the quadratic inverse transform. For example, the quadratic inverse transform can be NSST, RST, or LFNST, and the quadratic inverse transform application determiner can determine whether to apply the quadratic inverse transform based on the quadratic transform flag obtained by parsing the bitstream. In another example, the quadratic inverse transform application determiner can determine whether to apply the quadratic inverse transform based on the transform coefficients of the residual block.
[0147] The quadratic inverse transform determiner can determine the quadratic inverse transform. In this case, the quadratic inverse transform determiner can determine the quadratic inverse transform applied to the current block based on the LFNST (NSST or RST) transform set specified according to the intra prediction mode. In an embodiment, the quadratic transform determination method can be determined depending on the primary transform determination method. Various combinations of the primary transform and the quadratic transform can be determined according to the intra prediction mode. In addition, in an example, the quadratic inverse transform determiner can determine the region to which the quadratic inverse transform is applied based on the size of the current block.
[0148] In addition, as described above, if the secondary (inverse) transformation is omitted, the (dequantized) transform coefficients can be received, a single (separable) inverse transformation can be performed, and a residual block (residual samples) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.
[0149] In addition, in the present disclosure, a reduced secondary transform (RST) in which the size of the transform matrix (kernel) is reduced can be applied in the concept of NSST, so as to reduce the computational amount and storage amount required for the non-separable secondary transform.
[0150] In addition, the transform kernel, transform matrix, and coefficients constituting the transform kernel matrix described in the present disclosure, that is, kernel coefficients or matrix coefficients, can be represented by 8 bits. This can be a condition implemented in the decoding device and the encoding device, and compared with the existing 9 bits or 10 bits, it can reduce the storage amount required to store the transform kernel, and can reasonably adapt to performance degradation. In addition, representing the kernel matrix by 8 bits can allow the use of small multipliers, and can be more suitable for single instruction multiple data (SIMD) instructions for optimal software implementation.
[0151] In this specification, the term "RST" may refer to a transform performed on the residual samples of a target block based on a transform matrix whose size is reduced according to a reduction factor. In the case of performing a reduction transform, due to the reduction in the size of the transform matrix, the computational amount required for the transform can be reduced. That is, RST can be used to solve the computational complexity problem that occurs in the transform of a large-sized block or a non-separable transform.
[0152] RST can be referred to by various terms such as reduction transform, reduced secondary transform, scaled transform, simplified transform, and simple transform, and the names that RST can be referred to are not limited to the listed examples. Alternatively, since RST is mainly performed in the low-frequency region including non-zero coefficients in the transform block, it can be referred to as a low-frequency non-separable transform (LFNST). The transform index can be referred to as the LFNST index.
[0153] In addition, when performing a secondary inverse transform based on RST, the inverse transformers 235 of the encoding device 200 and 322 of the decoding device 300 can include: an inverse reduced secondary transformer that derives modified transform coefficients based on the inverse RST of the transform coefficients; and an inverse primary transformer that derives the residual samples of the target block based on the inverse primary transform of the modified transform coefficients. The inverse primary transform refers to the inverse transform of the primary transform applied to the residual. In the present disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying the transform.
[0154] Figure 6FIG. 0 is a diagram illustrating an RST according to an embodiment of the present disclosure.
[0155] In the present disclosure, a "target block" may refer to a current block to be encoded, a residual block, or a transform block.
[0156] In an RST according to an example, an N-dimensional vector may be mapped to an R-dimensional vector located in another space, so that a reduction transform matrix may be determined, where R is less than N. N may refer to the square of the length of a side of a block to which a transform is applied, or the total number of transform coefficients corresponding to a block to which a transform is applied, and a reduction factor may refer to the R / N value. The reduction factor may be referred to as a reduction factor, a shrinking factor, a simplification factor, a simplifying factor, or various other terms. In addition, R may be referred to as a reduction coefficient, but depending on the situation, the reduction factor may refer to R. In addition, depending on the situation, the reduction factor may refer to the N / R value.
[0157] In an example, the reduction factor or the reduction coefficient may be signaled through a bitstream, but the example is not limited thereto. For example, a predetermined value for the reduction factor or the reduction coefficient may be stored in each of the encoding device 200 and the decoding device 300, and in this case, the reduction factor or the reduction coefficient may not be signaled separately.
[0158] The size of the reduction transform matrix according to an example may be R×N, which is less than N×N (the size of a conventional transform matrix), and may be defined as in Equation 4 below.
[0159] [Equation 4]
[0160]
[0161] Figure 6 The matrix T in the reduction transform block shown in (a) of may refer to the matrix T of Equation 4 R×N . As Figure 6 shown in (a) of, when the reduction transform matrix T R×N is multiplied by the residual samples of the target block, the transform coefficients of the current block may be derived.
[0162] In an example, if the size of a block to which a transform is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then the RST according to Figure 6 (a) of may be expressed as a matrix operation shown in Equation 5 below. In this case, the storage and multiplication calculations may be reduced to about 1 / 4 by the reduction factor.
[0163] In the present disclosure, a matrix operation may be understood as an operation of obtaining a column vector by multiplying a column vector by a matrix set on the left side of the column vector.
[0164] [Equation 5]
[0165]
[0166] In Equation 6, r1 to r 64 can represent the residual samples of the target block, and specifically can be the transform coefficients generated by applying a single transform. As a result of the calculation of Equation 5, the transform coefficients c i of the target block can be derived, and the process of deriving c i can be as shown in Equation 6.
[0167] [Equation 6]
[0168]
[0169] As a result of the calculation of Equation 6, the transform coefficients c1 to c R of the target block can be derived. That is, when R = 16, the transform coefficients c1 to c 16 of the target block can be derived. If a conventional transform is applied instead of RST, and a transform matrix of size 64×64 (N×N) is multiplied by a residual sample of size 64×1 (N×1), then only 16 (R) transform coefficients are derived for the target block because RST is applied, although 64 (N) transform coefficients are derived for the target block. Since the total number of transform coefficients for the target block is reduced from N to R, the amount of data sent from the encoding device 200 to the decoding device 300 is reduced, and thus the transmission efficiency between the encoding device 200 and the decoding device 300 can be improved.
[0170] When considered from the perspective of the size of the transform matrix, the size of the conventional transform matrix is 64×64 (N×N), but the size of the reduced transform matrix is reduced to 16×64 (R×N). Therefore, the storage utilization rate in the case of performing RST can be reduced by the ratio of R / N compared to the case of performing a conventional transform. Additionally, when compared with the number of multiplication calculations N×N in the case of using a conventional transform matrix, using the reduced transform matrix can reduce the number of multiplication calculations (R×N) by the ratio of R / N.
[0171] In the example, the transformer 232 of the encoding device 200 can derive the transform coefficients of the target block by performing a single transform on the residual samples of the target block and a secondary transform based on RST. These transform coefficients can be transmitted to the inverse transformer of the decoding device 300, and the inverse transformer 322 of the decoding device 300 can derive the modified transform coefficients based on the inverse reduced secondary transform (RST) for the transform coefficients, and can derive the residual samples of the target block based on the inverse single transform for the modified transform coefficients.
[0172] According to the inverse RST matrix T N×Ris of size N×R, which is smaller than the size of the conventional inverse transform matrix N×N, and is related to the reduced transform matrix T shown in Equation 4 R×N has a transpose relationship.
[0173] Figure 6 The matrix T in the reduced inverse transform block shown in (b) of t may refer to the inverse RST matrix T N×R T (the superscript T refers to transpose). As Figure 6 shown in (b) of N×R T When the inverse RST matrix T R×N T is multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block or the residual samples of the target block can be derived. The inverse RST matrix T R×N can be expressed as (T T N×R ).
[0174] More specifically, when the inverse RST is used as a secondary inverse transform, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, the modified transform coefficients of the target block can be derived. In addition, the inverse RST can be used as an inverse primary transform, and in this case, when the inverse RST matrix T N×R T is multiplied by the transform coefficients of the target block, the residual samples of the target block can be derived.
[0175] In the example, if the size of the block to which the inverse transform is applied is 8×8 and R = 16 (i.e., R / N = 16 / 64 = 1 / 4), then the RST in (b) of Figure 6 can be represented by the matrix operation shown in Equation 7 below.
[0176] [Equation 7]
[0177]
[0178] In Equation 7, c1 to c 16 can represent the transform coefficients of the target block. As a result of the calculation of Equation 7, r j , which represents the modified transform coefficients of the target block or the residual samples of the target block, can be derived, and the process of deriving r j can be as shown in Equation 8.
[0179] [Equation 8]
[0180]
[0181] As a result of the calculation of Equation 8, r1 to r representing the modified transform coefficients of the target block or the residual samples of the target block can be derived N From the perspective of the size of the inverse transform matrix, the size of the conventional inverse transform matrix is 64×64 (N×N), but the size of the inverse reduction transform matrix is reduced to 64×16 (R×N). Therefore, compared with the case of performing a conventional inverse transform, the storage utilization rate in the case of performing an inverse RST can be reduced by the ratio of R / N. In addition, when compared with the number of multiplication calculations N×N in the case of using a conventional inverse transform matrix, using the inverse reduction transform matrix can reduce the number of multiplication calculations (N×R) by the ratio of R / N.
[0182] The transform set configuration shown in Table 2 can also be applied to 8×8 RST. That is, 8×8 RST can be applied according to the transform set in Table 2. Since, according to the intra prediction mode, one transform set includes two or three transforms (kernels), it can be configured to select one of up to four transforms including the case where no secondary transform is applied. Among the transforms without applying the secondary transform, applying an identity matrix can be considered. Assuming that indices 0, 1, 2, and 3 are assigned to the four transforms respectively (for example, index 0 can be assigned to the case of applying the identity matrix, that is, the case where no secondary transform is applied), the transform index or lfnst index as a syntax element can be signaled for each transform coefficient block, thereby specifying the transform to be applied. That is, for the upper left 8×8 block, through the transform index, 8×8 NSST in the RST configuration can be specified, or 8×8 lfnst can be specified when LFNST is applied. 8×8 lfnst and 8×8 RST refer to the transforms that can be applied to the 8×8 region included in the transform coefficient block when both the W and H of the target block to be transformed are equal to or greater than 8, and the 8×8 region can be the upper left 8×8 region in the transform coefficient block. Similarly, 4×4 lfnst and 4×4 RST refer to the transforms that can be applied to the 4×4 region included in the transform coefficient block when both the W and H of the target block are equal to or greater than 4, and the 4×4 region can be the upper left 4×4 region in the transform coefficient block.
[0183] According to an embodiment of the present disclosure, for the transform in the encoding process, only 48 pieces of data can be selected, and a maximum 16×48 transform kernel matrix can be applied thereto, instead of applying a 16×64 transform kernel matrix to 64 pieces of data forming an 8×8 region. Here, "maximum" means that m has a maximum value of 16 in the m×48 transform kernel matrix for generating m coefficients. That is, when performing RST by applying an m×48 transform kernel matrix (m≤16) to an 8×8 region, 48 pieces of data are input, and m coefficients are generated. When m is 16, 48 pieces of data are input and 16 coefficients are generated. That is, assuming that 48 pieces of data form a 48×1 vector, the 16×48 matrix and the 48×1 vector are multiplied in sequence, thereby generating a 16×1 vector. Here, the 48 pieces of data forming the 8×8 region can be appropriately arranged to form a 48×1 vector. For example, a 48×1 vector can be constructed based on 48 pieces of data constituting a region other than the lower right 4×4 region among the 8×8 region. Here, when performing matrix operations by applying a maximum 16×48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients can be arranged in the upper left 4×4 region according to the scanning order, and the upper right 4×4 region and the lower left 4×4 region can be filled with zeros.
[0184] For the inverse transform in the decoding process, the transpose matrix of the aforementioned transform kernel matrix can be used. That is, when performing inverse RST or LFNST in the inverse transform process performed by the decoding device, the input coefficient data for applying inverse RST is configured in a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the one-dimensional vector by the corresponding inverse RST matrix on the left side of the one-dimensional vector can be arranged in a two-dimensional block according to a predetermined arrangement order.
[0185] In summary, in the transform process, when RST or LFNST is applied to an 8×8 region, matrix operations are performed on 48 transform coefficients in the upper left region, the upper right region, and the lower left region of the 8×8 region except for the lower right region with a 16×48 transform kernel matrix. For the matrix operations, 48 transform coefficients are input in a one-dimensional array. When performing the matrix operations, 16 modified transform coefficients are derived, and the modified transform coefficients can be arranged in the upper left region of the 8×8 region.
[0186] Conversely, in the inverse transformation process, when applying the inverse RST or LFNST to an 8×8 region, the corresponding 16 transform coefficients in the 8×8 region corresponding to the upper-left region of the 8×8 region among the transform coefficients in the 8×8 region can be input as a one-dimensional array according to the scanning order, and can undergo matrix operations with a 48×16 transform kernel matrix. That is, the matrix operation can be expressed as (48×16 matrix) * (16×1 transform coefficient vector) = (48×1 modified transform coefficient vector). Here, an n×1 vector can be interpreted as having the same meaning as an n×1 matrix and can thus be represented as an n×1 column vector. In addition, * represents matrix multiplication. When performing the matrix operation, 48 modified transform coefficients can be derived, and the 48 modified transform coefficients can be arranged in the upper-left region, upper-right region, and lower-left region of the 8×8 region except for the lower-right region.
[0187] When the inverse quadratic transformation is based on RST, the inverse transformers 235 of the encoding device 200 and the inverse transformers 322 of the decoding device 300 can include an inverse reduced quadratic transformer for deriving modified transform coefficients based on the inverse RST of the transform coefficients and an inverse primary transformer for deriving the residual samples of the target block based on the inverse primary transform of the modified transform coefficients. The inverse primary transform refers to the inverse transform of the primary transform applied to the residuals. In the present disclosure, deriving transform coefficients based on a transform can refer to deriving transform coefficients by applying the transform.
[0188] The non-separable transform (LFNST) described above will be described in detail below. The LFNST can include a forward transform performed by the encoding device and an inverse transform performed by the decoding device.
[0189] The encoding device receives the result (or a part of the result) derived after applying the primary (core) transform as input and applies a forward quadratic transform (quadratic transform).
[0190] [Equation 9]
[0191] y = G T x
[0192] In Equation 9, x and y are the input and output of the quadratic transform respectively, G is a matrix representing the quadratic transform, and the transform basis vectors are composed of column vectors. In the case of the inverse LFNST, when the dimension of the transform matrix G is represented as [number of rows × number of columns], in the case of the forward LFNST, the transpose of the matrix G becomes the dimension of G T of.
[0193] For the inverse LFNST, the dimensions of matrix G are [48×16], [48×8], [16×16], [16×8], and the [48×8] matrix and the [16×8] matrix are partial matrices of 8 transform basis vectors sampled from the left side of the [48×16] matrix and the [16×16] matrix, respectively.
[0194] On the other hand, for the forward LFNST, matrix G T has dimensions of [16×48], [8×48], [16×16], [8×16], and the [8×48] matrix and the [8×16] matrix are partial matrices obtained by sampling 8 transform basis vectors from the upper part of the [16×48] matrix and the [16×16] matrix, respectively.
[0195] Therefore, in the case of the forward LFNST, a [48×1] vector or a [16×1] vector can be used as the input x, and a [16×1] vector or an [8×1] vector can be used as the output y. In video coding and decoding, the output of the forward single transform is two-dimensional (2D) data. Therefore, in order to construct a [48×1] vector or a [16×1] vector as the input x, it is necessary to construct a one-dimensional vector by appropriately arranging the 2D data that is the output of the forward transform.
[0196] Figure 7 is a diagram illustrating the order of arranging the output data of the forward single transform into a one-dimensional vector according to the example. Figure 7 The left diagrams of (a) and (b) of illustrate the order for constructing a [48×1] vector, and Figure 7 the right diagrams of (a) and (b) of illustrate the order for constructing a [16×1] vector. In the case of LFNST, a one-dimensional vector x can be obtained by arranging the 2D data in the same order as in Figure 7 the (a) and (b) of.
[0197] The arrangement direction of the output data of the forward single transform can be determined according to the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is in the horizontal direction with respect to the diagonal direction, the output data of the forward single transform can be arranged in the order of Figure 7 the (a) of, and when the intra prediction mode of the current block is in the vertical direction with respect to the diagonal direction, the output data of the forward single transform can be arranged in the order of Figure 7 the (b) of.
[0198] According to the example, an arrangement order different from Figure 7 the (a) and (b) of can be applied, and in order to derive and apply Figure 7The same result (y vector) for the arrangement orders of (a) and (b) can rearrange the column vectors of matrix G according to the arrangement order. That is, the column vectors of G can be rearranged such that each element constituting the x vector is always multiplied by the same transformation basis vector.
[0199] Since the output y derived by Equation 9 is a one-dimensional vector, when two-dimensional data is required as input data during the process of using the result of the forward quadratic transform as input (e.g., during quantization or residual coding), the output y vector of Equation 9 needs to be appropriately rearranged into 2D data again.
[0200] Figure 8 is a diagram illustrating the order of arranging the output data of the forward quadratic transform into a two-dimensional vector according to an example.
[0201] In the case of LFNST, the output values can be arranged in a 2D block according to a predetermined scan order. Figure 8 (a) of shows that when the output y is a [16×1] vector, the output values are arranged at 16 positions in a 2D block according to the diagonal scan order. Figure 8 (b) of shows that when the output y is an [8×1] vector, the output values are arranged at 8 positions in a 2D block according to the diagonal scan order, and the remaining 8 positions are filled with zeros. Figure 8 X in (b) indicates that it is filled with zeros.
[0202] According to another example, since the order of processing the output vector y during quantization or residual coding can be preset, the output vector y may not be arranged in a 2D block as shown in Figure 8 . However, in the case of residual coding, data coding can be performed in units of 2D blocks (e.g., 4×4) (e.g., CG (coefficient group)), and in this case, the data is arranged according to a specific order in the diagonal scan order as shown in Figure 8 .
[0203] In addition, the decoding device can configure the one-dimensional input vector y by arranging the two-dimensional data output by the dequantization process according to a preset scan order for inverse transformation. The input vector y can be output as an output vector x by the following equation.
[0204] [Equation 10]
[0205] x = Gy
[0206] In the case of inverse LFNST, the output vector x can be derived by multiplying the input vector y, which is a [16×1] vector or an [8×1] vector, by the G matrix. For inverse LFNST, the output vector x can be a [48×1] vector or a [16×1] vector.
[0207] The output vector x is arranged in a two-dimensional block according to the order shown in Figure 7 and is arranged as two-dimensional data, and this two-dimensional data becomes the input data (or a part of the input data) of the inverse first transformation.
[0208] Therefore, the inverse second transformation as a whole is the reverse of the forward second transformation process, and in the case of the inverse transformation, different from the forward direction, the inverse second transformation is first applied, and then the inverse first transformation is applied.
[0209] In the inverse LFNST, one of eight [48×16] matrices and eight [16×16] matrices can be selected as the transformation matrix G. Whether to apply the [48×16] matrix or the [16×16] matrix depends on the size and shape of the block.
[0210] In addition, eight matrices can be derived from the four transformation sets shown in Table 2 above, and each transformation set can consist of two matrices. Which transformation set to use among the four transformation sets is determined according to the intra-frame prediction mode, and more specifically, based on the value of the intra-frame prediction mode extended by considering wide-angle intra-frame prediction (WAIP). Which matrix to select from the two matrices constituting the selected transformation set is derived by index signaling. More specifically, 0, 1, and 2 can be used as the transmitted index values, 0 can indicate that the LFNST is not applied, and 1 and 2 can indicate any one of the two transformation matrices constituting the transformation set selected based on the intra-frame prediction mode value.
[0211] Figure 9 FIG. is a diagram illustrating a wide-angle intra-frame prediction mode according to an embodiment of this document.
[0212] The general intra-frame prediction mode values can have values from 0 to 66 and from 81 to 83, and the intra-frame prediction mode values extended due to WAIP can have values from -14 to 83 as shown. The values from 81 to 83 indicate the CCLM (Cross-Component Linear Model) mode, and the values from -14 to -1 and from 67 to 80 indicate the intra-frame prediction mode extended due to the application of WAIP.
[0213] When the width of the current prediction block is greater than the height, the upper reference pixel is generally closer to the position inside the block to be predicted. Therefore, predicting in the lower left direction can be more accurate than in the upper right direction. On the contrary, when the height of the block is greater than the width, the left reference pixel is generally closer to the position inside the block to be predicted. Therefore, predicting in the upper right direction can be more accurate than in the lower left direction. Therefore, it may be advantageous to apply remapping (i.e., mode index modification) to the index of the wide-angle intra-frame prediction mode.
[0214] When applying wide-angle intra prediction, information about existing intra prediction can be signaled, and after the information is parsed, the information can be remapped to an index of the wide-angle intra prediction mode. Therefore, the total number of intra prediction modes for a specific block (e.g., a non-square block of a specific size) can be unchanged, that is, the total number of intra prediction modes is 67, and the intra prediction mode coding for a specific block can be unchanged.
[0215] Table 3 below shows the process of deriving a modified intra mode by remapping the intra prediction mode to a wide-angle intra prediction mode.
[0216] [Table 3]
[0217]
[0218] In Table 3, the extended intra prediction mode value is finally stored in the predModeIntra variable, and ISP_NO_SPLIT indicates that the CU block is not partitioned into sub-partitions by the intra sub-partition (ISP) technique currently adopted in the VVC standard, and the cIdx variable values of 0, 1, and 2 indicate the cases of the luminance component, the Cb component, and the Cr component, respectively. The log2 function shown in Table 3 returns the log value with base 2, and the Abs function returns the absolute value.
[0219] Variables such as the predModeIntra indicating the intra prediction mode and the height and width of the transform block are used as input values for the wide-angle intra prediction mode mapping process, and the output value is the modified intra prediction mode predModeIntra. The height and width of the transform block or the coding block can be the height and width of the current block for the remapping of the intra prediction mode. At this time, the variable whRatio reflecting the ratio of width to width can be set to Abs(Log2(nW / nH)).
[0220] For non-square blocks, the intra prediction mode can be divided into two cases and modified.
[0221] First, if all of the conditions (1) to (3) are satisfied, (1) the width of the current block is greater than the height, (2) the intra prediction mode before modification is equal to or greater than 2, and (3) the intra prediction mode is less than the value derived as (8 + 2*whRatio) when the variable whRatio is greater than 1 and less than 8 when the variable whRatio is less than or equal to 1 (predModeIntra < (whRatio > 1)? (8 + 2*whRatio) : 8), then the intra prediction mode is set to a value 65 greater than predModeIntra [predModeIntra is set to be equal to (predModeIntra + 65)].
[0222] If different from the above, i.e., if conditions (1) to (3) are satisfied, (1) the height of the current block is greater than the width, (2) the intra prediction mode before modification is less than or equal to 66, and (3) the intra prediction mode is greater than the value derived as (60 - 2 * whRatio) when whRatio is greater than 1 and greater than 60 when whRatio is less than or equal to 1 (predModeIntra > (whRatio > 1)? (60 - 2 * whRatio) : 60), then the intra prediction mode is set to a value 67 less than predModeIntra [predModeIntra is set to be equal to (predModeIntra - 67)].
[0223] Table 2 above shows how to select a transform set in LFNST based on the intra prediction mode values extended by WAIP. As Figure 9 shown, modes 14 to 33 and modes 35 to 80 are symmetric about the prediction direction around mode 34. For example, mode 14 and mode 54 are symmetric about the direction corresponding to mode 34. Therefore, the same transform set is applied to modes located in symmetric directions, and this symmetry is also reflected in Table 2.
[0224] In addition, it is assumed that the forward LFNST input data of mode 54 is symmetric with the forward LFNST input data of mode 14. For example, for mode 14 and mode 54, according to Figure 7 the arrangement order shown in (a) of Figure 7 and the arrangement order shown in (b) of Figure 7 rearrange the two-dimensional data into one-dimensional data. Additionally, it can be seen that the pattern of the order shown in (a) of Figure 7 and (b) of
[0225] is symmetric about the direction (diagonal direction) indicated by mode 34.
[0226] Figure 10 is a diagram illustrating the block shapes to which LFNST is applied. Figure 10 (a) of Figure 10 shows a 4×4 block, Figure 10 shows a 4×8 block and an 8×4 block, Figure 10 shows a 4×N block or an N×4 block, where N is 16 or greater, Figure 10 shows an 8×8 block,
[0227] In Figure 10In, the blocks with thick boundaries indicate the regions to which the LFNST is applied. For Figure 10 the blocks of (a) and (b), the LFNST is applied to the upper left 4×4 region, and for Figure 10 the block of (c), the LFNST is separately applied to two continuously arranged upper left 4×4 regions. In Figure 10 (a), (b), and (c), since the LFNST is applied in units of 4×4 regions, this LFNST will be referred to as "4×4 LFNST" hereinafter. Based on the matrix dimension of G, a [16×16] or [16×8] matrix can be applied.
[0228] More specifically, a [16×8] matrix is applied to the 4×4 block (4×4 TU or 4×4 CU) of Figure 10 (a), and a [16×16] matrix is applied to the blocks in Figure 10 (b) and (c). This is to adjust the worst-case computational complexity to 8 multiplications per sample.
[0229] Regarding Figure 10 (d) and (e), the LFNST is applied to the upper left 8×8 region, and this LFNST is referred to as "8×8 LFNST" hereinafter. As the corresponding transformation matrix, a [48×16] matrix or a [48×8] matrix can be applied. In the case of the forward LFNST, since a [48×1] vector (the X vector in Equation 9) is input as the input data, not all the sample values in the upper left 8×8 region are used as the input values of the forward LFNST. That is, as can be seen from the left order of Figure 7 (a) or the left order of Figure 7 (b), a [48×1] vector can be constructed based on the samples belonging to the remaining 3 4×4 blocks while leaving the lower right 4×4 block unchanged.
[0230] A [48×8] matrix can be applied to the 8×8 block (8×8 TU or 8×8 CU) in Figure 10 (d), and a [48×16] matrix can be applied to the 8×8 block in Figure 10 (e). This is also to adjust the worst-case computational complexity to 8 multiplications per sample.
[0231] Depending on the block shape, when the corresponding forward LFNST (4×4 or 8×8 LFNST) is applied, 8 or 16 output data (the Y vector in Equation 9, [8×1] or [16×1] vector) are generated. In the forward LFNST, due to the characteristics of the matrix G T , the number of output data is equal to or less than the number of input data.
[0232] Figure 11 FIG. is a diagram illustrating the arrangement of the output data of the forward LFNST according to the example, and shows the blocks in which the output data of the forward LFNST is arranged according to the block shape.
[0233] In Figure 11 the shaded area in the upper left of the shown block corresponds to the area where the output data of the forward LFNST is located, the positions marked with 0 indicate the samples filled with the value 0, and the remaining area represents the area not changed by the forward LFNST. In the area not changed by the LFNST, the output data of the forward transform remains unchanged.
[0234] As described above, since the size of the applied transformation matrix varies according to the block shape, the number of output data also varies. As Figure 11 , the output data of the forward LFNST may not completely fill the upper left 4×4 block. In Figure 11 cases (a) and (d), a [16×8] matrix and an A[48×8] matrix are respectively applied to the blocks indicated by the thick lines or partial areas inside the blocks, and an [8×1] vector as the output of the forward LFNST is generated. That is, according to Figure 8 the scanning order shown in (b), only 8 output data can be filled, as shown in Figure 11 cases (a) and (d), and 0 can be filled in the remaining 8 positions. In Figure 10 the case of the block to which the LFNST is applied in (d), as shown in Figure 11 (d), the two 4×4 blocks in the upper right and lower left adjacent to the upper left 4×4 block are also filled with the value 0.
[0235] As described above, basically, by signaling the LFNST index, it is specified whether the LFNST is applied and the transformation matrix to be applied. As Figure 11 shown, when the LFNST is applied, since the number of output data of the forward LFNST can be equal to or less than the number of input data, there are areas filled with zero values as follows.
[0236] 1) As Figure 11 shown in (a), the samples from the eighth position and subsequent positions in the scanning order in the upper left 4×4 block, that is, the samples from the ninth to the sixteenth.
[0237] 2) As Figure 11 shown in (d) and (e), when a [48×16] matrix or a [48×8] matrix is applied, the two 4×4 blocks adjacent to the upper left 4×4 block or the second and third 4×4 blocks in the scanning order.
[0238] Therefore, if there is non-zero data in the check regions 1) and 2), it is determined that LFNST is not applied, so that signaling of the corresponding LFNST index can be omitted.
[0239] According to the example, for instance, in the case of LFNST adopted in the VVC standard, since signaling of the LFNST index is performed after residual coding, the encoding device can know whether there is non-zero data (valid coefficients) at all positions within a TU or CU block through residual coding. Therefore, the encoding device can determine whether to perform signaling regarding the LFNST index based on the presence of non-zero data, and the decoding device can determine whether to parse the LFNST index. When there is no non-zero data in the regions specified in 1) and 2) above, signaling of the LFNST index is performed.
[0240] Since the truncated unary code is applied as the binarization method of the LFNST index, the LFNST index consists of up to two bins, and 0, 10, and 11 are respectively assigned as the binary codes for the possible LFNST index values 0, 1, and 2. In the case of the currently used LFNST for VVC, context-based CABAC coding is applied to the first bin (conventional coding), and bypass coding is applied to the second bin. The total number of contexts for the first bin is 2. When (DCT-2, DCT-2) is applied as the one-time transform pair for the horizontal and vertical directions and the luminance component and chrominance component are encoded in a dual-tree type, one context is assigned and the other context is applied to the remaining cases. The coding of the LFNST index is shown in the following table.
[0241] [Table 4]
[0242]
[0243] In addition, for the adopted LFNST, the following simplified method can be applied.
[0244] (i) According to the example, the number of output data of the forward LFNST can be limited to a maximum of 16.
[0245] In Figure 10 case (c), 4×4 LFNST can be respectively applied to two 4×4 regions adjacent to the upper left, and in this case, a maximum of 32 LFNST output data can be generated. When the number of output data of the forward LFNST is limited to a maximum of 16, in the case of a 4×N / N×4 (N≥16) block (TU or CU), 4×4 LFNST is only applied to one 4×4 region in the upper left, and LFNST can be applied to Figure 10 all blocks only once. Through this, the implementation manner of image coding can be simplified.
[0246] Figure 12 It is shown that the number of output data of the forward LFNST according to the example is limited to a maximum value of 16. In Figure 12 , when the LFNST is applied to the uppermost left 4×4 region in a 4×N or N×4 block (where N is 16 or greater), the output data of the forward LFNST becomes 16.
[0247] (ii) According to the example, the regions to which the LFNST is not applied can be additionally cleared. In this document, clearing can mean filling all positions belonging to a specific region with a value of 0. That is, clearing can be applied to the regions that are not changed due to the LFNST, and the result of the forward one-time transformation is maintained. As described above, since the LFNST is divided into 4×4 LFNST and 8×8 LFNST, the clearing can be divided into two types ((ii)-(A) and (ii)-(B)) as follows.
[0248] (ii)-(A) When the 4×4 LFNST is applied, the regions to which the 4×4 LFNST is not applied can be cleared. Figure 13 FIG. is an illustration of the clearing in a block to which the 4×4 LFNST according to the example is applied.
[0249] As Figure 13 shown, regarding the block to which the 4×4 LFNST is applied, that is, for all the blocks in Figure 11 (a), (b), and (c), the entire region to which the LFNST is not applied can be filled with zeros.
[0250] On the other hand, Figure 13 (d) of FIG. shows that when the maximum value of the number of output data of the forward LFNST is limited to 16 (as Figure 12 shown), clearing is performed on the remaining blocks to which the 4×4 LFNST is not applied.
[0251] (ii)-(B) When the 8×8 LFNST is applied, the regions to which the 8×8 LFNST is not applied can be cleared. Figure 14 FIG. is an illustration of the clearing in a block to which the 8×8 LFNST according to the example is applied.
[0252] As Figure 14 shown, regarding the block to which the 8×8 LFNST is applied, that is, for all the blocks in Figure 11 (d) and (e), the entire region to which the LFNST is not applied can be filled with zeros.
[0253] (iii) Due to the clearing presented in (ii) above, the regions filled with zeros may not be the same as those when the LFNST is applied. Therefore, according to the comparison Figure 11For the wider area of the LFNST, perform the clearing proposed in (ii) to check for the presence of non-zero data.
[0254] For example, when (ii)-(B) is applied, after checking for the presence of non-zero data in the zero-filled areas in (d) and (e) of Figure 11 , additionally check for the presence of non-zero data in the area filled with 0 in Figure 14 . Signaling for the LFNST index can be performed only when there is no non-zero data.
[0255] Of course, even when the clearing proposed in (ii) is applied, the presence of non-zero data can be checked in the same way as the existing LFNST index signaling. That is, after checking for the presence of non-zero data in the blocks filled with zero in Figure 11 , the LFNST index signaling can be applied. In this case, the encoding device only performs the clearing and the decoding device does not assume the clearing, that is, only checks whether non-zero data exists only in the area clearly marked as 0 in Figure 11 . LFNST index parsing can be performed.
[0256] Alternatively, according to another example, the clearing as shown in Figure 15 can be performed. Figure 15 is a diagram illustrating the clearing in the block applying the 8×8 LFNST according to another example.
[0257] As shown in Figure 13 and Figure 14 , the clearing can be applied to all areas except the area where the LFNST is applied, or the clearing can be applied only to a local area, as shown in Figure 15 . The clearing is applied only to the area except the upper left 8×8 area of Figure 15 . The clearing may not be applied to the lower right 4×4 block within the upper left ×8 area.
[0258] Various embodiments of the combination of the simplified methods ((i), (ii)-(A), (ii)-(B), (iii)) for applying the LFNST can be derived. Of course, the combination of the above simplified methods is not limited to the following embodiments, and any combination can be applied to the LFNST.
[0259] Embodiment
[0260] - Limit the number of output data of the forward LFNST to a maximum of 16 → (i)
[0261] - When applying the 4×4 LFNST, all areas where the 4×4 LFNST is not applied are cleared → (II)-(A)
[0262] - When applying 8×8 LFNST, all areas where 8×8 LFNST is not applied are cleared → (II)-(B)
[0263] - After checking whether non-zero data also exists in the existing areas filled with zero values and the areas filled with zero due to additional clearing ((ii)-(A), (ii)-(B)), signal the LFNST index only when non-zero data does not exist → (iii).
[0264] In the case of the embodiment, when applying LFNST, the areas where non-zero output data can exist are limited to the inside of the upper left 4×4 area. More specifically, in Figure 13 of (a) and Figure 14 of (a), the eighth position in the scanning order is the last position where non-zero data can exist. In Figure 13 of (b) and (c) and Figure 14 of (b), the sixteenth position in the scanning order (i.e., the position of the lower right edge of the upper left 4×4 block) is the last position where data other than 0 can exist.
[0265] Therefore, after applying LFNST, after checking whether non-zero data exists at positions where the residual coding process does not allow (at positions beyond the last position), it can be determined whether to signal the LFNST index.
[0266] In the case of the clearing method proposed in (ii), due to the amount of data finally generated when both a transform and LFNST are applied once, the computational amount required to perform the entire transform process can be reduced. That is, when LFNST is applied, since clearing is applied to the areas where the forward one-time transform output data exists and LFNST is not applied, it is not necessary to generate data for the areas that become cleared during the execution of the forward one-time transform. Therefore, the computational amount required to generate the corresponding data can be reduced. The additional effects of the clearing method proposed in (ii) are summarized as follows.
[0267] First, as described above, reduce the computational amount required to perform the entire transform process.
[0268] In particular, when applying (ii)-(B), the worst-case computational amount is reduced, making the transform process lighter. In other words, generally, a large amount of computation is required to perform a large-size one-time transform. By applying (ii)-(B), the amount of data derived as a result of performing the forward LFNST can be reduced to 16 or less. Additionally, as the size of the entire block (TU or CU) increases, the effect of reducing the amount of transform operations further increases.
[0269] Second, the computational amount required for the entire transformation process can be reduced, thereby reducing the power consumption required to perform the transformation.
[0270] Third, the latency involved in the transformation process is reduced.
[0271] Secondary transformations such as LFNST add computational amount to the existing primary transformation, thus increasing the overall latency time involved when performing the transformation. In particular, in the case of intra prediction, since the reconstructed data of adjacent blocks is used during the prediction process, during encoding, the increase in latency due to the secondary transformation leads to an increase in the latency until reconstruction. This can result in an increase in the overall latency of intra prediction coding.
[0272] However, if the zeroing proposed in application (ii) is applied, the latency time for performing the primary transformation can be greatly reduced when applying LFNST, maintaining or reducing the latency time of the entire transformation, such that the encoding device can be implemented more simply.
[0273] In traditional intra prediction, the block to be currently encoded is regarded as one coding unit, and encoding is performed without division. However, intra-subpartition (ISP) coding means performing intra prediction coding by dividing the block to be currently encoded in the horizontal direction or the vertical direction. In this case, reconstructed blocks can be generated by performing encoding / decoding in units of the divided blocks, and the reconstructed blocks can be used as reference blocks for the next divided blocks. According to an embodiment, in ISP coding, one coding block can be divided into two or four sub-blocks for encoding, and in ISP, within one sub-block, intra prediction is performed by referring to the reconstructed pixel values of the sub-blocks adjacent on the left or adjacent on the upper side. Hereinafter, "encoding" can be used as a concept including both the encoding performed by the encoding device and the decoding performed by the decoding device.
[0274] Table 5 shows the number of sub-blocks divided according to the block size when applying ISP, and the sub-partitions divided according to ISP can be referred to as transform units (TUs).
[0275] [Table 5]
[0276] Block size (CU) Number of partitions 4×4 Not available 4×8、8×4 2 All other cases 4
[0277] ISP divides a block predicted to be intra in the luminance into two or four sub-partitions in the vertical direction or the horizontal direction according to the size of the block. For example, the minimum block size for which ISP can be applied is 4×8 or 8×4. When the block size is larger than 4×8 or 8×4, the block is divided into 4 sub-partitions.
[0278] Figure 16 and 17 illustrates an example of the sub-blocks into which one coding block is divided, and more specifically, Figure 16An example of a division where the coding block (width (W) × height (H)) is a 4×8 block or an 8×4 block is illustrated, and Figure 17 An example of a division in a case where the coding block is not a 4×8 block, an 8×4 block, or a 4×4 block is illustrated.
[0279] When applying ISP, the sub - blocks are encoded in order from left to right or from top to bottom according to the division type (e.g., horizontally or vertically), and after performing the reconstruction process via inverse transform and intra - prediction for one sub - block, the encoding of the next sub - block can be performed. For the left - most or top - most sub - block, the reconstructed pixels of the already - encoded coding block are referred to, as in the conventional intra - prediction method. In addition, when each side of a subsequent internal sub - block is not adjacent to the previous sub - block, in order to derive the reference pixels adjacent to the corresponding side, the reconstructed pixels of the already - encoded adjacent coding block are referred to, as in the conventional intra - prediction method.
[0280] In the ISP coding mode, all sub - blocks can be encoded with the same intra - prediction mode, and a flag indicating whether to use ISP coding and a flag indicating whether to divide in which direction (horizontal or vertical) can be signaled. As Figure 16 and Figure 17 shown, the number of sub - blocks can be adjusted to 2 or 4 according to the shape of the block, and when the size (width × height) of a sub - block is less than 16, it can be restricted so that division into the corresponding sub - blocks is not allowed or ISP coding itself is not applied.
[0281] In the case of the ISP prediction mode, one coding unit is divided into two or four partitioned blocks (i.e., sub - blocks) and prediction is performed, and the same intra - prediction mode is applied to the two or four partitioned blocks.
[0282] As described above, in the division direction, both the horizontal direction (when an M×N coding unit with horizontal length M and vertical length N is divided in the horizontal direction, if the M×N coding unit is divided into two, the M×N coding unit is divided into M×(N / 2) blocks, and if the M×N coding unit is divided into four blocks, the M×N coding unit is divided into M×(N / 4) blocks)) and the vertical direction (when the M×N coding unit is divided in the vertical direction, if the M×N coding unit is divided into two, the M×N coding unit is divided into (M / 2)×N blocks, and if the M×N coding unit is divided into four, the M×N coding unit is divided into (M / 4)×N blocks) are possible. When dividing the M×N coding unit in the horizontal direction, the partitioned blocks are encoded in the up - down order, and when dividing the M×N coding unit in the vertical direction, the partitioned blocks are encoded in the left - right order. In the case of horizontal (vertical) division, the reconstructed pixel values of the upper (left) partitioned block can be referred to for predicting the currently - encoded partitioned block.
[0283] A transform can be applied to the residual signal generated in block units by the ISP prediction method. The multi-transform selection (MTS) technique based on the DST-7 / DCT-8 combination and the existing DCT-2 can be applied to the forward primary transform (core transform), and the forward low-frequency non-separable transform (LFNST) can be applied to the transform coefficients generated by the primary transform to generate the final modified transform coefficients.
[0284] That is, the LFNST can be applied to the blocks divided by applying the ISP prediction mode, and the same intra-frame prediction mode is applied to the divided blocks, as described above. Therefore, when selecting the set of LFNSTs derived based on the intra-frame prediction mode, the derived set of LFNSTs can be applied to all the blocks. That is, since the same intra-frame prediction mode is applied to all the blocks, the same set of LFNSTs can be applied to all the blocks.
[0285] According to an embodiment, the LFNST can be applied only to transform blocks having both a horizontal length and a vertical length of 4 or greater. Therefore, when the horizontal length or the vertical length of the divided block according to the ISP prediction method is less than 4, the LFNST is not applied and the LFNST index is not signaled. Further, when the LFNST is applied to each block, the corresponding block can be regarded as one transform block. When the ISP prediction method is not applied, the LFNST can be applied to the coding block.
[0286] A method of applying the LFNST to each block will be described in detail.
[0287] According to an embodiment, after applying the forward LFNST to each block, only up to 16 (8 or 16) coefficients are left in the upper left 4×4 region in the transform coefficient scan order, and then zeroing can be applied, where the remaining positions and regions are all filled with 0.
[0288] Alternatively, according to an embodiment, when the length of one side of the block is 4, the LFNST is applied only to the upper left 4×4 region, and when the length of all sides of the block (i.e., width and height) is 8 or greater, the LFNST can be applied to the remaining 48 coefficients in the upper left 8×8 region except for the lower right 4×4 region.
[0289] Alternatively, according to an embodiment, in order to adjust the worst-case computational complexity to 8 multiplications per sample, when each block is 4×4 or 8×8, only 8 transform coefficients can be output after applying the forward LFNST. That is, when the block is 4×4, an 8×16 matrix can be used as the transform matrix, and when the block is 8×8, an 8×48 matrix can be used as the transform matrix.
[0290] In the current VVC standard, the LFNST index signaling is performed on a coding unit basis. Therefore, in the ISP prediction mode and when LFNST is applied to all partition blocks, the same LFNST index value can be applied to the corresponding partition blocks. That is to say, when the LFNST index value is sent once at the coding unit level, the corresponding LFNST index can be applied to all partition blocks in the coding unit. As described above, the LFNST index value can have values of 0, 1, and 2, where 0 indicates the case where LFNST is not applied, and 1 and 2 indicate two transform matrices existing in an LFNST set when LFNST is applied.
[0291] As described above, the LFNST set is determined by the intra prediction mode. And in the case of the ISP prediction mode, since all partition blocks in the coding unit are predicted under the same intra prediction mode, the partition blocks can refer to the same LFNST set.
[0292] As another example, the LFNST index signaling is still performed on a coding unit basis, but in the case of the ISP prediction mode, it is not determined whether LFNST is uniformly applied to all partition blocks, and for each partition block, it can be determined whether to apply the LFNST index value signaled at the coding unit level and whether to apply LFNST through separate conditions. Here, the separate conditions can be signaled in the bitstream in the form of a flag for each partition block, and when the flag value is 1, the LFNST index value signaled at the coding unit level is applied, and when the flag value is 0, LFNST may not be applied.
[0293] In a coding unit where the ISP mode is applied, an example of applying LFNST when the length of one side of a partition block is less than 4 is described as follows.
[0294] First, when the size of the partition block is N×2 (2×N), LFNST can be applied to the upper-left M×2 (2×M) region (where M≤N). For example, when M = 8, the upper-left region becomes 8×2 (2×8), so the region where there are 16 residual signals can be the input of the forward LFNST, and a forward transform matrix of R×16 (R≤16) can be applied.
[0295] Here, the forward LFNST matrix can be a separate additional matrix other than the matrices included in the current VVC standard. In addition, for worst-case complexity control, an 8×16 matrix that samples only the upper 8 row vectors of a 16×16 matrix can be used for transformation. The complexity control method will be described in detail later.
[0296] Secondly, when the size of the partition block is N×1 (1×N), the LFNST can be applied to the upper-left M×1 (1×M) region (where M≤N). For example, when M = 16, the upper-left region becomes 16×1 (1×16), so there are 16 regions of residual signals that can be the input of the forward LFNST, and a forward transform matrix of R×16 (R≤16) can be applied.
[0297] Here, the corresponding forward LFNST matrix can be a separate additional matrix other than the matrices included in the current VVC standard. In addition, to control the worst-case complexity, an 8×16 matrix in which only the upper 8 row vectors of the 16×16 matrix are sampled can be used for the transform. The complexity control method will be described in detail later.
[0298] The first embodiment and the second embodiment can be applied simultaneously, or either of the two embodiments can be applied. In particular, in the case of the second embodiment, since one transform is considered in the LFNST, it has been experimentally observed that the improvement in compression performance obtained in the existing LFNST is relatively small compared to the LFNST index signaling cost. However, in the case of the first embodiment, an improvement in compression performance similar to that obtained from the conventional LFNST is observed. That is, in the case of ISP, the contribution of the application of the 2×N and N×2 LFNSTs to the actual compression performance can be experimentally examined.
[0299] In the current VVC's LFNST, symmetry between intra prediction modes is applied. The same set of LFNSTs is applied to two directional modes set around mode 34 (prediction in the 45-degree diagonal direction in the lower right corner). For example, the same set of LFNSTs is applied to mode 18 (horizontal direction prediction mode) and mode 50 (vertical direction prediction mode). However, in modes 35 to 66, when the forward LFNST is applied, the input data is transposed and then the LFNST is applied.
[0300] VVC supports the wide-angle intra prediction (WAIP) mode. Considering the WAIP mode, a set of LFNSTs is derived based on the modified intra prediction mode. For the modes extended by WAIP, the set of LFNSTs is determined by using symmetry, just like in the general intra prediction direction modes. For example, since mode -1 is symmetric with mode 67, the same set of LFNSTs is applied, and since mode -14 is symmetric with mode 80, the same set of LFNSTs is applied. Modes 67 to 80 apply the LFNST transform after transposing the input data before applying the forward LFNST.
[0301] When applying LFNST to the upper-left M×2 (M×1) block, since the block to which LFNST is applied is non-square, the symmetry of LFNST cannot be applied. Therefore, instead of applying the symmetry based on the intra prediction mode as in the LFNST of Table 2, the symmetry between the M×2 (M×1) block and the 2×M (1×M) block can be applied.
[0302] Figure 18 FIG. is a diagram illustrating the symmetry between the M×2 (M×1) block and the 2×M (1×M) block according to an embodiment.
[0303] As Figure 18 shown, since it can be considered that mode 2 in the M×2 (M×1) block is symmetric to mode 66 in the 2×M (1×M) block, the same set of LFNSTs can be applied to the 2×M (1×M) block and the M×2 (M×1) block.
[0304] In this case, in order to apply the set of LFNSTs applied to the M×2 (M×1) block to the 2×M (1×M) block, the set of LFNSTs is selected based on mode 2 instead of mode 66. That is, before applying the forward LFNST, after transposing the input data of the 2×M (1×M) block, the LFNST can be applied.
[0305] Figure 19 FIG. is a diagram illustrating an example of transposing a 2×M block according to an embodiment.
[0306] Figure 19 (a) of FIG. is a diagram illustrating that the LFNST can be applied by reading the input data of the 2×M block in column-major order, Figure 19 (b) of FIG. is a diagram illustrating that the LFNST can be applied by reading the input data of the M×2 (M×1) block in row-major order. The method of applying the LFNST to the upper-left M×2 (M×1) or 2×M (M×1) block is described below.
[0307] 1. First, as Figure 19 (a) and (b) of FIG. shown, the input data is arranged to form the input vector of the forward LFNST. For example, referring to Figure 18 , for the M×2 block predicted in mode 2, following the order in Figure 19 (b), for the 2×M block predicted in mode 66, the input data is arranged in the order of Figure 19 (a), and then the LFNST set set for mode 2 can be applied.
[0308] 2. For the M×2 (M×1) block, considering WAIP, the set of LFNSTs is determined based on the modified intra prediction mode. As described above, a preset mapping relationship is established between the intra prediction mode and the set of LFNSTs, which can be represented by a mapping table.
[0309] For a 2×M (1×M) block, a symmetric pattern around the prediction pattern (pattern 34 in the case of the VVC standard) in the downward 45-degree diagonal direction from the modified intra prediction mode can be obtained considering the WAIP, and then the LFNST set can be determined based on the corresponding symmetric pattern and the mapping table. The symmetric pattern (y) around pattern 34 can be derived by the following formula. The mapping table will be described in more detail below.
[0310] [Equation 11]
[0311] If 2 ≤ x ≤ 66, then y = 68 - x,
[0312] otherwise (x ≤ -1 or x ≥ 67), y = 66 - x
[0313] 3. When applying the forward LFNST, the transform coefficients can be derived by multiplying the input data prepared in process 1 by the LFNST kernel. The LFNST kernel can be selected according to the LFNST set determined in process 2 and a predetermined LFNST index.
[0314] For example, when M = 8 and a 16×16 matrix is applied as the LFNST kernel, 16 transform coefficients can be generated by multiplying the matrix by 16 input data. The generated transform coefficients can be arranged in the upper left 8×2 or 2×8 region in the scan order used in the VVC standard.
[0315] Figure 20 The scan order of the 8×2 or 2×8 region according to the embodiment is illustrated.
[0316] All regions except the upper left 8×2 or 2×8 region can be filled with zero values (cleared), or the existing transform coefficients after applying the transform once can be left as they are. The predetermined LFNST index can be one of the LFNST index values (0, 1, 2) tried when calculating the RD cost while changing the LFNST index value in the programming process.
[0317] In the case of a configuration where the worst-case computational complexity is adjusted to a certain level or lower (e.g., 8 multiplications / sample), for example, after generating only 8 transform coefficients by multiplying an 8×16 matrix that only takes the upper 8 rows of the 16×16 matrix, the transform coefficients can be set in Figure 20 the scan order, and clearing can be applied to the remaining coefficient regions. The worst-case complexity control will be described later.
[0318] 4. When inverse LFNST is applied, a preset number (e.g., 16) of transform coefficients are set as an input vector, and the LFNST set obtained from process 2 and an LFNST kernel (e.g., a 16×16 matrix) derived from the selected parsed LFNST index are selected, and then an output vector can be derived by multiplying the LFNST kernel with the corresponding input vector.
[0319] In the case of M×2(M×1) blocks, the output vector can be expressed as Figure 19 In the case of 2×M (1×M) blocks, the output vector can be Figure 19 (a) Column priority setting.
[0320] The remaining areas except the area where the corresponding output vector is set within the upper left M×2 (M×1) or 2×M (M×2) area and the area except the upper left M×2 (M×1) or 2×M (M×2) area in the partition block (M×2 in the partition block) can all be cleared to have zero values, or can be configured to keep the reconstructed transform coefficients as is through residual encoding and inverse quantization processing.
[0321] When constructing the input vector, as in point 3, the input data can be Figure 20 The scanning order arrangement is carried out, and in order to control the computational complexity of the worst case to a certain level or lower, the input vector can be constructed by reducing the number of input data (for example, 8 instead of 16).
[0322] For example, when M=8, if 8 input data are used, only the left 16×8 matrix can be taken from the corresponding 16×16 matrix and multiplied to obtain 16 output data. Worst case complexity control will be described later.
[0323] In the above embodiment, when LFNST is applied, the case where symmetry is applied between M×2 (M×1) blocks and 2×M (1×M) blocks is shown, but according to another example, a different LFNST set may be applied to each of the two block shapes.
[0324] Hereinafter, various examples of a mapping method using an intra prediction mode and a LFNST set configuration of an ISP mode will be described.
[0325] In the case of ISP mode, the LFNST set configuration may be different from the existing LFNST set. In other words, a core different from the existing LFNST core may be applied, and a mapping table different from the mapping table between the intra prediction mode index and the LFNST set applied to the current VVC standard may be applied. The mapping table applied to the current VVC standard may be the same as the mapping table of Table 2.
[0326] In Table 2, the preModeIntra value represents the intra prediction mode value changed considering WAIP, and the lfnstTrSetIdx value is an index value indicating a specific LFNST set. Each LFNST set is configured with two LFNST kernels.
[0327] When applying the ISP prediction mode, if both the horizontal length and the vertical length of each sub-block are equal to or greater than 4, the same kernel as the LFNST kernel applied in the current VVC standard can be applied, and the mapping table can be applied as it is. A mapping table and LFNST kernel different from the current VVC standard can be applied.
[0328] When applying the ISP prediction mode, when the horizontal length or the vertical length of each sub-block is less than 4, a mapping table and LFNST kernel different from the current VVC standard can be applied. Hereinafter, Tables 6 to 8 represent the mapping tables between the intra prediction mode values (intra prediction mode values changed considering WAIP) and the LFNST sets, which can be applied to M×2 (M×1) blocks or 2×M (1×M) blocks.
[0329] [Table 6]
[0330] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0 <= predModeIntra <= 1 0 2 <= predModeIntra <= 12 1 13 <= predModeIntra <= 23 2 24 <= predModeIntra <= 34 3 35 <= predModeIntra <= 44 4 45 <= predModeIntra <= 55 5 56 <= predModeIntra <= 66 6 67 <= predModeIntra <= 80 6 81 <= predModeIntra <= 83 0
[0331] [Table 7]
[0332] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0 <= predModeIntra <= 1 0 2 <= predModeIntra <= 23 1 24 <= predModeIntra <= 44 2 45 <= predModeIntra <= 66 3 67 <= predModeIntra <= 80 3 81 <= predModeIntra <= 83 0
[0333] [Table 8]
[0334] predModeIntra lfnstTrSetIdx predModeIntra<0 1 0 <= predModeIntra <= 1 0 2 <= predModeIntra <= 80 1 81 <= predModeIntra <= 83 0
[0335] The first mapping table of Table 6 is configured with seven LFNST sets, the mapping table of Table 7 is configured with four LFNST sets, and the mapping table of Table 8 is configured with two LFNST sets. As another example, when it is configured with one LFNST set, the lfnstTrSetIdx value can be fixed to 0 with respect to the preModeIntra value.
[0336] Hereinafter, a method for maintaining the worst-case computational complexity when applying LFNST to the ISP mode will be described.
[0337] In the case of the ISP mode, when applying LFNST, in order to keep the number of multiplications per sample (or per coefficient, per position) at a certain value or less, the application of LFNST may be restricted. Depending on the size of the sub-block, by applying LFNST as follows, the number of multiplications per sample (or per coefficient, per position) can be kept at 8 or less.
[0338] 1. When both the horizontal length and the vertical length of the partition block are 4 or greater, the same method as the worst-case computational complexity control method for LFNST in the current VVC standard can be applied.
[0339] That is to say, when the partition block is a 4×4 block, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix instead of the 16×16 matrix can be applied in the forward direction, and a 16×8 matrix obtained by sampling the left 8 columns from a 16×16 matrix can be applied in the inverse direction. In addition, when the partition block is an 8×8 block, in the forward direction, instead of a 16×48 matrix, an 8×48 matrix obtained by sampling the upper 8 rows from a 16×48 matrix is applied, and in the inverse direction, instead of a 48×16 matrix, a 48×8 matrix obtained by sampling the left 8 columns from a 48×16 matrix can be applied.
[0340] In the case of a 4×N or N×4 (N>4) block, when performing the forward transform, the 16 coefficients generated after applying the 16×16 matrix only to the upper-left 4×4 block can be set in the upper-left 4×4 region, and the other regions can be filled with the value 0. Additionally, when performing the inverse transform, the 16 coefficients located in the upper-left 4×4 block are set in scan order to form an input vector, and then 16 output data can be generated by multiplying by the 16×16 matrix. The generated output data can be set in the upper-left 4×4 region, and the remaining regions except the upper-left 4×4 region can be filled with the value 0.
[0341] In the case of an 8×N or N×8 (N>8) block, when performing the forward transform, the 16 coefficients generated after applying the 16×48 matrix to the ROI region (the remaining region except the lower-right 4×4 block in the upper-left 8×8 block) within only the upper-left 8×8 block can be set in the upper-left 4×4 region, and all other regions can be filled with the value 0. Furthermore, when performing the inverse transform, the 16 coefficients located in the upper-left 4×4 region are set in scan order to form an input vector, and then 48 output data can be generated by multiplying by the 48×16 matrix. The generated output data can be filled in the ROI region, and all other regions can be filled with the value 0.
[0342] 2. When the size of the partition block is N×2 or 2×N and LFNST is applied to the upper-left M×2 or 2×M region (M≤N), a matrix sampled according to the N value can be applied.
[0343] In the case of M = 8, for the partition block of N = 8, that is, an 8×2 or 2×8 block, in the case of the forward transform, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix can be applied instead of the 16×16 matrix, and in the case of the inverse transform, a 16×8 matrix obtained by sampling the left 8 columns from a 16×16 matrix can be applied instead of the 16×16 matrix.
[0344] When N is greater than 8, in the case of the forward transform, the 16 output data generated after applying the 16×16 matrix to the upper left 8×2 or 2×8 block are set in the upper left 8×2 or 2×8 block, and the remaining area can be filled with the value 0. In the case of the inverse transform, the 16 coefficients located in the upper left 8×2 or 2×8 block are set in scan order to form an input vector, and then 16 output data can be generated by multiplying by the 16×16 matrix. The generated output data can be placed in the upper left 8×2 or 2×8 block, and all the remaining areas can be filled with the value 0.
[0345] 3. When the size of the partition block is N×1 or 1×N and LFNST is applied to the upper left M×1 or 1×M area (M ≤ N), a matrix sampled according to the N value can be applied.
[0346] In the case of M = 16, for the partition block of N = 16, that is, a 16×1 or 1×16 block, in the case of the forward transform, an 8×16 matrix obtained by sampling the upper 8 rows from a 16×16 matrix can be applied instead of the 16×16 matrix, and in the case of the inverse transform, a 16×8 matrix obtained by sampling the left 8 columns from a 16×16 matrix can be applied instead of the 16×16 matrix.
[0347] When N is greater than 16, in the case of the forward transform, the 16 output data generated after applying the 16×16 matrix to the upper left 16×1 or 1×16 block can be set in the upper left 16×1 or 1×16 block, and the remaining area can be filled with the value 0. In the case of the inverse transform, the 16 coefficients located in the upper left 16×1 or 1×16 block can be set in scan order to form an input vector, and then 16 output data can be generated by multiplying by the 16×16 matrix. The generated output data can be set in the upper left 16×1 or 1×16 block, and all the remaining areas can be filled with the value 0.
[0348] As another example, in order to keep the number of multiplications per sample (or per coefficient, per position) at a certain value or less, the number of multiplications per sample (or per coefficient, per position) based on the ISP coding unit size rather than the size of the ISP partition block can be kept at 8 or less. When only one block among the ISP partition blocks satisfies the condition for applying LFNST, the worst-case complexity calculation of applying LFNST can be based on the corresponding coding unit size rather than the size of the partition block.
[0349] The following drawings are provided to describe specific examples of the present disclosure. Since the specific names of the devices shown in the drawings or the names of specific signals / messages / fields are provided for illustration purposes, the technical features of the present disclosure are not limited to the specific names used in the following drawings.
[0350] Figure 21 is a flowchart illustrating the operation of a video decoding device according to an embodiment of this document.
[0351] Figure 21 Each step disclosed in Figures 4 to 20 is based on some of the content described above in Figures 3 to 20 Therefore, detailed descriptions that are repetitive with those described above in
[0352] The decoding device 300 according to the embodiment can obtain intra prediction mode information and LFNST index information from the bitstream (S2110).
[0353] LFNST is an inseparable transform that applies the transform without separating the coefficients in a specific direction, which is different from a single transform that separates and transforms the transform target coefficients in the vertical or horizontal direction. The inseparable transform can be a low-frequency inseparable transform that applies the transform only to the low-frequency region rather than the entire target block to be transformed.
[0354] The LFNST index information can be received as syntax information, and the syntax information can be received as a binarized empty string including 0 and 1.
[0355] The syntax element of the LFNST index according to the present embodiment can indicate whether to apply the inverse LFNST or the inverse inseparable transform and any one of the transform kernel matrices included in the transform set, and when the transform set includes two transform kernel matrices, the syntax element of the transform index can have three values.
[0356] That is, according to the embodiment, the syntax element values of the LFNST index can include 0, 1, and 2. 0 indicates the case where the inverse LFNST is not applied to the target block, 1 indicates the first transform kernel matrix among the transform kernel matrices, and 2 indicates the second transform kernel matrix among the transform kernel matrices.
[0357] Such intra prediction mode information and LFNST index information can be signaled at the coding unit level.
[0358] In addition, the decoding device can derive residual information from the bitstream, e.g., the quantized transform coefficients for the current block (i.e., the transform block to be transformed).
[0359] More specifically, the decoding device 300 can decode information on the quantized transform coefficients of the target block from the bitstream, and can derive the quantized transform coefficients of the current block based on the information on the quantized transform coefficients of the current block. The information on the quantized transform coefficients of the target block can be included in the sequence parameter set (SPS) or the slice header, and can include at least one of information on whether to apply a reduced transform (RST), information on a reduction factor, information on the minimum transform size to which the reduced transform is applied, information on the maximum transform size to which the reduced transform is applied, the reduced inverse transform size, and information on a transform index indicating any one of the transform kernel matrices included in the transform set.
[0360] The decoding device 300 can perform dequantization on the quantized transform coefficients of the current block to derive the transform coefficients.
[0361] The derived transform coefficients can be arranged in a reverse diagonal scan order in units of 4×4 blocks, and the transform coefficients within a 4×4 block can be arranged in a reverse diagonal scan order. That is, the transform coefficients for which inverse quantization has been performed can be set in the reverse scan order applied in a video codec such as in VCC or HEVC.
[0362] When the current block is divided into sub-partition transform blocks, the decoding device can derive the prediction samples of the current block based on the intra prediction mode information, i.e., perform intra prediction (S2120).
[0363] By receiving and parsing flag information indicating whether ISP coding or an ISP mode is applied, the decoding device can derive whether the current block is divided into a predetermined number of sub-partition transform blocks. Here, the current block can be a coding block. In addition, the decoding device can derive the size and number of the divided sub-blocks from the flag information indicating in which direction the current block is divided.
[0364] For example, as Figure 16 shown, when the size (width × height) of the current block is 8×4, the current block can be vertically divided and divided into two sub-blocks, and when the size (width × height) of the current block is 4×8, the current block can be horizontally divided and divided into two sub-blocks. Alternatively, as Figure 17As shown, when the size (width × height) of the current block is greater than 4×8 or 8×4, that is, when the size of the current block is: 1) 4×N or N×4 (N≥16) or 2) M×N (M≥8, N≥8), the current block can be divided into 4 sub-blocks in the horizontal or vertical direction.
[0365] In this document, the fact that an M×N block is greater than or equal to a K×L block means that M is greater than or equal to K and N is greater than or equal to L. Additionally, the fact that an M×N block is greater than a K×L block can mean that M is greater than or equal to K while N is greater than or equal to L and M is greater than K or N is greater than L. The fact that an M×N block is less than or equal to a K×L block means that M is less than or equal to K and N is less than or equal to L, and the fact that an M×N block is less than a K×L block means that M is less than or equal to K while N is less than or equal to L and M is less than K or N is less than L.
[0366] The same intra prediction mode can be applied to the sub-partition transform blocks divided from the current block, and the decoding device can derive the prediction samples of each sub-partition transform block. That is, the decoding device performs intra prediction sequentially from left to right or from top to bottom according to the division form of the sub-partition transform blocks, for example, horizontally or vertically. For the leftmost or topmost sub-block, the reconstructed pixels of the already encoded coding block are referred to, as in the traditional intra prediction method. Additionally, for each side of the subsequent internal sub-partition transform blocks, when it is not adjacent to the previous sub-partition transform block, in order to derive the reference pixels adjacent to the corresponding side, the reconstructed pixels of the already encoded adjacent coding block are referred to, as in the traditional intra prediction method.
[0367] The decoding device can determine an LFNST set including the LFNST matrix based on the intra prediction mode derived from the intra prediction mode information (S2130), and select any one of the multiple LFNST matrices based on the LFNST set and the LFNST index (S2140).
[0368] In this case, the same LFNST set and the same LFNST index can be applied to the sub-partition transform blocks divided from the current block. That is, since the same intra prediction mode is applied to the sub-partition transform blocks, the LFNST set determined based on the intra prediction mode can be equally applied to all sub-partition transform blocks. Additionally, since the LFNST index is signaled at the coding unit level, the same LFNST matrix can be applied to the sub-partition transform blocks divided from the current block.
[0369] Furthermore, since LFNST can be applied when the height and width of the transform block are 4 or greater, in order to apply LFNST to the sub-partition transform blocks, the width and height of the divided sub-partition transform blocks should also be 4 or greater.
[0370] As described above, a transform set can be determined based on the intra prediction mode of a transform block to be transformed, and an inverse LFNST can be performed based on any one of the transform kernel matrices (i.e., LFNST matrices) included in the transform set indicated by the LFNST index. The matrix applied to the inverse LFNST can be referred to as an inverse LFNST matrix or an LFNST matrix, and such a matrix can have any name as long as it has a transpose relationship with the matrix used for the forward LFNST.
[0371] In one example, the inverse LFNST matrix can be a non-square matrix in which the number of columns is less than the number of rows.
[0372] The decoding device can derive the transform coefficients of the sub-partitioned transform block based on the LFNST matrix (S2150).
[0373] A predetermined number of transform coefficients as the output data of the LFNST can be derived based on the size of the sub-partitioned transform block. For example, when the height and width of the sub-partitioned transform block are 8 or greater, 48 transform coefficients can be derived, as shown on the left side of Figure 7 and when the width and height of the sub-partitioned transform block are not 8 or greater, that is, when the width or height of the sub-partitioned transform block is 4 or greater and less than 8, 16 transform coefficients can be derived, as shown on the right side of Figure 7
[0374] As shown in Figure 7 48 transform coefficients can be arranged in the upper left, upper right, and lower left 4×4 regions of the upper left 8×8 region of the sub-partitioned transform block, and 16 transform coefficients can be arranged in the upper left 4×4 region of the sub-partitioned transform block.
[0375] The 48 transform coefficients and the 16 transform coefficients can be arranged in the vertical direction or the horizontal direction according to the intra prediction mode of the sub-partitioned transform block. For example, when the intra prediction mode is the horizontal direction (modes 2 to 34 in Figure 9 ) based on the diagonal direction (mode 34 in Figure 9 ), the transform coefficients can be arranged in the horizontal direction, that is, in the row-major order, as shown in Figure 7 (a) of Figure 9 and when the intra prediction mode is the vertical direction (modes 35 to 66 in Figure 7 ) based on the diagonal direction, the transform coefficients can be arranged in the horizontal direction, that is, in the column-major order, as shown in (b) of
[0376]
[0377] In this case, a simplified inverse transform can be applied to the inverse first transform, or a conventional separable transform can be used. Alternatively, the MTS can be used as the inverse first transform.
[0378] Subsequently, the decoding device 300 can generate reconstructed samples based on the residual samples of the current block and the predicted samples of the current block.
[0379] The following drawings are provided to describe specific examples of the present disclosure. Since the specific names of the devices shown in the drawings or the names of specific signals / messages / fields are provided for illustration, the technical features of the present disclosure are not limited to the specific names used in the following drawings.
[0380] Figure 22 is a flowchart illustrating the operation of a video encoding device according to an embodiment of this document.
[0381] Figure 22 Each step disclosed in is based on some of the content described above in Figures 4 to 20 Therefore, detailed descriptions that are repetitive with those described above in Figure 2 and Figures 4 to 20 will be omitted or simplified.
[0382] According to an embodiment, the encoding device 200 can derive predicted samples for each sub-partition transform block based on the intra prediction mode applied to the current block (S2210).
[0383] The encoding device can determine whether ISP encoding or an ISP mode is applied to the current block (i.e., the encoding block), determine the direction in which the current block will be divided according to the determination result, and derive the size and number of the divided sub-blocks.
[0384] For example, when the size (width × height) of the current block is 8 × 4, as Figure 16 shown, the current block can be vertically divided into two sub-blocks. When the size (width × height) of the current block is 4 × 8, the current block can be horizontally divided into two sub-blocks. Alternatively, as Figure 17 shown, when the size (width × height) of the current block is greater than 4 × 8 or 8 × 4, that is, when the size of the current block is: 1) 4 × N or N × 4 (N ≥ 16) or 2) M × N (M ≥ 8, N ≥ 8), the current block can be divided into 4 sub-blocks in the horizontal or vertical direction.
[0385] The same intra prediction mode can be applied to the sub - partition transform blocks divided from the current block, and the encoding device can derive the prediction samples for each sub - partition transform block. That is, the encoding device performs intra prediction sequentially from left to right or from top to bottom according to the division form of the sub - partition transform blocks, for example, horizontally or vertically. For the left - most or top - most sub - blocks, the reconstructed pixels of the already - encoded coding blocks are referred to, as in the traditional intra prediction method. Additionally, for each side of the subsequent internal sub - partition transform blocks, when it is not adjacent to the previous sub - partition transform block, in order to derive the reference pixels adjacent to the corresponding side, the reconstructed pixels of the already - encoded adjacent coding blocks are referred to, as in the traditional intra prediction method.
[0386] The encoding device 200 can derive the residual samples of the current block based on the prediction samples (S2220).
[0387] Furthermore, the encoding device 200 can derive the transform coefficients of the current block based on a single transformation of the residual samples (S2230).
[0388] The single transformation can be performed by multiple transform kernels, and in this case, the transform kernel can be selected based on the intra prediction mode.
[0389] The encoding device 200 can determine whether to perform a secondary transformation or a non - separable transformation, especially LFNST, on the transform coefficients of the current block.
[0390] When it is determined to perform LFNST, the encoding device 200 can derive the modified transform coefficients of the sub - partition transform blocks based on the LFNST set mapped to the intra prediction mode and the LFNST matrix included in the LFNST set (S2240).
[0391] The encoding device 200 can determine the LFNST set based on the mapping relationship according to the intra prediction mode applied to the current block, and perform LFNST, that is, non - separable transformation, based on one of the two LFNST matrices included in the LFNST set.
[0392] In this case, the same LFNST set and the same LFNST index can be applied to the sub - partition transform blocks divided from the current block. That is, because the same intra prediction mode is applied to the sub - partition transform blocks, the LFNST set determined based on the intra prediction mode can also be equally applied to all sub - partition transform blocks. Additionally, since the LFNST index is encoded in units of coding units, the same LFNST matrix can be applied to the sub - partition transform blocks divided from the current block.
[0393] In addition, since LFNST can be applied when the height and width of a transform block are 4 or greater, in order to apply LFNST to a sub-partitioned transform block, the width and height of the partitioned sub-partitioned transform block should also be 4 or greater.
[0394] As described above, a transform set can be determined according to the intra prediction mode of the transform block to be transformed. The matrix applied to LFNST and the matrix for inverse LFNST have a transpose relationship.
[0395] In one example, the LFNST matrix can be a non-square matrix in which the number of rows is less than the number of columns.
[0396] The region where the transform coefficients serving as the input data for LFNST are located can be derived based on the size of the sub-partitioned transform block. For example, when the height and width of the sub-partitioned transform block are 8 or greater, the region is the upper-left, upper-right, and lower-left 4×4 regions of the upper-left 8×8 region of the sub-partitioned transform block, as shown on the left side of Figure 7 and when the height and width of the sub-partitioned transform block are not equal to or greater than 8, the region can be the upper-left 4×4 region of the current block, as shown on the right side of Figure 7 .
[0397] The transform coefficients of the region can be read in column-major order or row-major order according to the intra prediction mode of the sub-partitioned transform block for multiplication with the LFNST matrix and arranged in one dimension.
[0398] 48 modified transform coefficients or 16 modified transform coefficients can be read in the vertical direction or the horizontal direction according to the intra prediction mode of the sub-partitioned transform block and arranged in one dimension. For example, when the intra prediction mode is in the horizontal direction based on the diagonal direction (mode 34 in Figure 9 and modes 2 to 34 in Figure 9 ), the transform coefficients can be arranged in the horizontal direction, that is, arranged in row-major order, as shown in (a) of Figure 7 and when the intra prediction mode is in the vertical direction based on the diagonal direction (modes 35 to 66 in Figure 9 ), the transform coefficients can be arranged in the horizontal direction, that is, arranged in column-major order, as shown in (b) of Figure 7 .
[0399] In one embodiment, an encoding device may include the following steps: determining whether the encoding device is in a condition for applying LFNST, generating and encoding an LFNST index based on the determination, selecting a transform kernel matrix, and applying LFNST to the residual samples based on the selected transform kernel matrix and / or a reduction factor when the encoding device is in a condition for applying LFNST. In this case, the size of the reduced transform kernel matrix can be determined based on the reduction factor.
[0400] The encoding device may perform quantization based on the modified transform coefficients of the current block to derive quantized transform coefficients, and encode information about the quantized transform coefficients and the LFNST index indicating the LFNST matrix (S2250).
[0401] That is, the encoding device may generate residual information including information about the quantized transform coefficients. The residual information may include the above-mentioned transform-related information / syntax elements. The encoding device may encode the image / video information including the residual information and output the encoded image / video information in the form of a bitstream.
[0402] More specifically, the encoding device 200 may generate information about the quantized transform coefficients and encode the information about the generated quantized transform coefficients.
[0403] The syntax element of the LFNST index according to the present embodiment may indicate whether to apply (inverse) LFNST and any one of the LFNST matrices included in the LFNST set, and when the LFNST set includes two transform kernel matrices, the syntax element of the LFNST index may have three values.
[0404] According to an embodiment, when the partition tree structure of the current block is a double-tree type, the LFNST index may be encoded for each of the luminance block and the chrominance block.
[0405] According to an embodiment, the syntax element value of the transform index may be derived as 0, 1, and 2. 0 indicates the case where (inverse) LFNST is not applied to the current block, 1 indicates the first LFNST matrix in the LFNST matrix, and 2 indicates the second LFNST matrix in the LFNST matrix.
[0406] In the present disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for the sake of consistency in expression.
[0407] In addition, in the present disclosure, the quantized transform coefficients and the transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled through the residual coding syntax. The transform coefficients may be derived based on the residual information (or the information about the transform coefficients), and the scaled transform coefficients may be derived through the inverse transform (scaling) of the transform coefficients. The residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. These details may also be applied / expressed in other parts of the present disclosure.
[0408] In the above-described embodiments, the method is explained based on a flowchart by means of a series of steps or blocks. However, the present disclosure is not limited to the order of the steps, and a certain step may be performed in an order or steps different from the above order or steps, or a certain step may be performed concurrently with other steps. In addition, those of ordinary skill in the art will understand that the steps shown in the flowchart are not exclusive, and one or more steps in the flowchart may be incorporated or deleted without affecting the scope of the present disclosure.
[0409] The above method according to the present disclosure may be implemented in software form, and the encoding device and / or decoding device according to the present disclosure may be included in devices for image processing such as a television, a computer, a smart phone, a set-top box, and a display device.
[0410] When the embodiments in the present disclosure are implemented by software, the above method may be implemented as modules (steps, functions, etc.) for performing the above functions. These modules may be stored in a memory and may be executed by a processor. The memory may be inside or outside the processor and may be connected to the processor in various well-known ways. The processor may include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present disclosure may be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing may be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip.
[0411] In addition, the decoding device and the encoding device applying the present disclosure may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device (such as video communication), a mobile streaming device, a storage medium, a camera, a video on demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and may be used to process video signals or data signals. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.
[0412] In addition, the processing method of the present disclosure can be produced in the form of a program executable by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes various storage devices and distributed storage devices that store computer-readable data. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission on the Internet). Additionally, the bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted through a wired or wireless communication network. Further, embodiments of the present disclosure can be implemented as a computer program product by program code, and the program code can be executed on a computer according to the embodiments of the present disclosure. The program code can be stored on a computer-readable carrier.
[0413] Figure 23 Illustrated is the structure of a content stream system applying the present disclosure.
[0414] In addition, the content stream system applying the present disclosure can generally include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0415] The encoding server is used to compress the content input from a multimedia input device such as a smart phone, a camera, a video camera, etc. into digital data to generate a bitstream and send it to the streaming server. As another example, in the case where a multimedia input device such as a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method of the present disclosure. And the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0416] The streaming server sends multimedia data to the user device through the web server based on the user's request, and the web server serves as a tool for notifying the user of what services exist. When the user requests a service that the user wants, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this regard, the content stream system can include a separate control server, and in this case, the control server is used to control commands / responses between corresponding devices in the content stream system.
[0417] The streaming server can receive content from a media storage device and / or an encoding server. For example, in the case of receiving content from the encoding server, the content can be received in real time. In this case, in order to provide the streaming service smoothly, the streaming server can store the bitstream for a predetermined time.
[0418] For example, the user device can include a mobile phone, a smartphone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigator, a slate PC, a tablet PC, a superbook, a wearable device (e.g., a watch-type terminal (smartwatch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, a digital signage, etc. Each server in the content streaming system can operate as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.
[0419] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined to be implemented or executed in a device, and the technical features of the device claims can be combined to be implemented or executed in a method. In addition, the technical features of the method claims and the device claims can be combined to be implemented or executed in a device, and the technical features of the method claims and the device claims can be combined to be implemented or executed in a method.
Claims
1. An image decoding device, the image decoding device comprising: a memory; and at least one processor, the at least one processor being connected to the memory and configured to: obtain, from a bitstream, intra prediction mode information and low-frequency non-separable transform (LFNST) index information for a current block, wherein the current block is divided into a plurality of sub-partition transform blocks based on an intra sub-partition mode, and intra prediction is performed for each of the sub-partition transform blocks in the current block in the intra sub-partition mode, and wherein, based on the width and height of the sub-partition transform blocks obtained by dividing the current block being greater than or equal to 4, obtain the LFNST index information for the current block from the bitstream; derive prediction samples for the sub-partition transform blocks based on an intra prediction mode derived from the intra prediction mode information for the current block; determine an LFNST set including an LFNST matrix based on the intra prediction mode; select one of the LFNST matrices based on the LFNST set and the LFNST index information; derive transform coefficients for the sub-partition transform blocks based on the one of the LFNST matrices; derive residual samples for the sub-partition transform blocks based on the transform coefficients; and generate reconstructed samples for the sub-partition transform blocks based on the prediction samples and the residual samples.
2. The image decoding device according to claim 1, wherein The same LFNST set and the same LFNST index information are applied to the sub-partition transform blocks divided from the current block.
3. The image decoding device according to claim 1, wherein, The same intra prediction mode is applied to the sub-partition transform blocks divided from the current block.
4. The image decoding device according to claim 1, wherein In response to the width and height of each sub-partition transform block being greater than or equal to 4, derive the transform coefficients based on the one of the LFNST matrices.
5. The image decoding device according to claim 1, wherein, Based on the size (width × height) of the current block being 8 × 4, the current block is vertically divided, and wherein, based on the size (width × height) of the current block being 4 × 8, the current block is horizontally divided.
6. The image decoding device according to claim 5, wherein, Based on the size (width × height) of the current block being greater than 4 × 8 or 8 × 4, the current block is divided into four sub-partition transform blocks in the horizontal direction or the vertical direction.
7. An image encoding device, the image encoding device comprising: a memory; and at least one processor, the at least one processor being connected to the memory and configured to: derive prediction samples for sub-partition transform blocks obtained by dividing a current block based on an intra prediction mode of the current block, wherein the current block is divided into a plurality of sub-partition transform blocks based on an intra sub-partition mode, and intra prediction is performed for each of the sub-partition transform blocks in the current block in the intra sub-partition mode; derive residual samples for the sub-partition transform blocks based on the prediction samples; apply a transform once to the residual samples to derive transform coefficients; Derive modified transform coefficients for the sub-partition transform block based on a set of low-frequency non-separable transforms (LFNSTs) mapped to the intra prediction mode and an LFNST matrix included in the set of LFNSTs; and Encode image information including LFNST index information related to the LFNST matrix and intra prediction mode information related to the intra prediction mode, wherein the LFNST index information is encoded into the bitstream based on the width and height of the sub-partition transform block being greater than or equal to 4.
8. The image encoding device according to claim 7, wherein, The same set of LFNSTs and the same LFNST matrix are applied to the sub-partition transform block.
9. The image encoding apparatus according to claim 7, wherein, In response to the width and height of each sub-partition transform block being greater than or equal to 4, derive the modified transform coefficients based on the LFNST matrix.
10. The image encoding device according to claim 7, wherein, Based on the size (width × height) of the current block being 8 × 4, the current block is vertically divided, and wherein, based on the size (width × height) of the current block being 4 × 8, the current block is horizontally divided.
11. The image encoding device according to claim 10, wherein, Based on the size (width × height) of the current block being greater than 4 × 8 or 8 × 4, the current block is divided into four sub-partition transform blocks in the horizontal or vertical direction.
12. A transmitting device for image data, the transmitting device comprising: At least one processor configured to obtain a bitstream for the image, wherein the bitstream is generated based on the following operations: Derive prediction samples for sub-partition transform blocks obtained by partitioning a current block based on the intra prediction mode of the current block, wherein the current block is divided into a plurality of sub-partition transform blocks based on an intra sub-partition mode, and intra prediction is performed for each of the sub-partition transform blocks in the current block in the intra sub-partition mode; Derive residual samples for the sub-partition transform blocks based on the prediction samples; Apply a transform once to the residual samples to derive transform coefficients; Derive modified transform coefficients for the sub-partition transform blocks based on a set of low-frequency non-separable transforms (LFNSTs) mapped to the intra prediction mode and an LFNST matrix included in the set of LFNSTs; and Encode image information including LFNST index information related to the LFNST matrix and intra prediction mode information related to the intra prediction mode; and A transmitter configured to transmit the data including the bitstream, wherein the LFNST index information is encoded into the bitstream based on the width and height of the sub-partition transform block being greater than or equal to 4.