Transform-based Video Compilation Method and Its Device

The method addresses the high costs of transmitting and storing high-resolution images by using high-frequency zeroing and multiple transform selection to enhance coding efficiency and reduce data loss, specifically for high-quality images and videos.

CN114342386BActive Publication Date: 2025-07-15LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080058618.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-08
Filing Date
2020-08-07
Publication Date
2025-07-15
Estimated Expiration
2040-08-07

AI Technical Summary

Technical Problem

The prior art has problems with increasing transmission and storage costs when transmitting and storing high resolution, high-quality images and videos, especially when dealing with immersive media such as virtual reality and holograms, image compilation efficiency is low and residual compilation efficiency is insufficient.

Method used

The image compilation method based on high-frequency zeroing and multiple transformation selection is adopted. By derive the transformation coefficients of the current block and applying zeroing blocks and multiple transformation selection, the image compilation efficiency is improved and data loss is reduced.

Benefits of technology

It improves the compression efficiency of images and videos, reduces transmission and storage costs, and enhances the efficiency of image compilation and the effect of residual compilation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114342386B_ABST
    Figure CN114342386B_ABST
Patent Text Reader

Abstract

A video decoding method according to this document includes a step of deriving transform coefficients for a current block based on residual information, wherein the step of deriving the transform coefficients includes a step of deriving a zeroing block indicating a region in the current block where valid transform coefficients may exist, and wherein the zeroing block is derived based on flag information indicating whether multiple transform selection (MTS) using multiple transform kernels can be applied to the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to image compilation technology, and more particularly, to a transform-based image compilation method and apparatus in an image compilation system. Background Art

[0002] Recently, the demand for high-resolution and high-quality images and videos such as ultra-high definition (HUD) images and 4K or 8K or larger videos has been increasing in various fields. As image and video data become high-resolution and high-quality, the amount of information or the number of bits transmitted relatively increases compared to existing image and video data. Therefore, if media such as existing wired or wireless broadband lines are used to transmit image data or existing storage media are used to store image and video data, the transmission cost and storage cost increase.

[0003] In addition, recently, the interest and demand for immersive media such as virtual reality (VR), augmented reality (AR) content, or holograms have been increasing. The broadcast of images and videos with image characteristics different from those of real images, such as game images, has been increasing.

[0004] Therefore, in order to effectively compress and transmit or store and play back information of high-resolution and high-quality images and videos with such various characteristics, efficient image and video compression technologies are required. Summary of the Invention

[0005] Technical Problem

[0006] A technical aspect of the present disclosure is to provide a method and apparatus for improving image compilation efficiency.

[0007] Another technical aspect of the present disclosure is to provide a method and apparatus for improving residual compilation efficiency.

[0008] Yet another technical aspect of the present disclosure is to provide a method and apparatus for improving residual compilation efficiency by compiling transform coefficients based on high-frequency zeroing.

[0009] Still another technical aspect of the present disclosure is to provide a method and apparatus for improving the efficiency of image compilation in which high-frequency zeroing is performed based on multiple transform selections.

[0010] Yet another technical aspect of the present disclosure is to provide a method and apparatus for compiling an image, which can reduce data loss when performing high-frequency zeroing.

[0011] Another technical aspect of the present disclosure is to provide a method and apparatus for deriving a context model for last valid transform coefficient position information based on a current block size when encoding transform coefficients for a current block (or current transform block) based on high-frequency zeroing.

[0012] Technical solution

[0013] According to an embodiment of the present disclosure, an image decoding method performed by a decoding device is provided. The method includes: deriving transform coefficients for a current block based on residual information, where the derivation of the transform coefficients includes deriving a zero-out block indicating a region where valid transform coefficients may exist in the current block, and deriving the zero-out block based on flag information indicating whether multiple transform selection (MTS) using multiple transform kernels is applicable to the current block.

[0014] When MTS is applied, the width or height of the zero-out block can be set to 16, and when MTS is not applied, the width or height of the zero-out block can be set to 32 or less.

[0015] When the flag information indicating whether sub-block transform for performing a transform on a partitioned compilation unit is applied to the current block is 1, the width or height of the zero-out block can be 16.

[0016] When the height of the partitioned sub-block is less than 64 and the width of the sub-block is 32, the width of the zero-out block can be set to 16.

[0017] When the width of the partitioned sub-block is less than 64 and the height of the sub-block is 32, the height of the zero-out block can be set to 16.

[0018] The transform kernel can be derived based on the partition direction of the current block and the position of the sub-block to which the transform is applied.

[0019] The residual information may include last valid coefficient prefix information, and the maximum value of the last valid coefficient prefix information can be derived based on the size of the zero-out block.

[0020] According to another embodiment of the present disclosure, an image encoding method performed by an encoding device is provided. The method includes: deriving residual samples for a current block and deriving transform coefficients based on the residual samples for the current block, where the derivation of the transform coefficients may include deriving a zero-out block indicating a region where valid transform coefficients may exist in the current block based on whether multiple transform selection (MTS) using multiple transform kernels is applied to the current block.

[0021] According to still another embodiment of the present disclosure, a digital storage medium can be provided that stores image data including encoded image information and a bitstream generated according to the image encoding method performed by the encoding device.

[0022] According to another embodiment of the present disclosure, a digital storage medium can be provided that stores image data including encoded image information and a bitstream to enable a decoding device to perform an image decoding method.

[0023] Beneficial effects

[0024] According to the present disclosure, the overall image / video compression efficiency can be improved.

[0025] According to the present disclosure, the efficiency of residual encoding can be improved.

[0026] According to the present disclosure, the efficiency of residual encoding can be improved by encoding transform coefficients based on high-frequency zeroing.

[0027] According to the present disclosure, the efficiency of image encoding in which high-frequency zeroing is performed based on multiple transform selection can be improved.

[0028] According to the present disclosure, data loss during the execution of high-frequency zeroing can be reduced, thereby improving the efficiency of image encoding.

[0029] The effects achievable through specific examples of the present disclosure are not limited to the effects listed above. For example, there can be various technical effects that can be understood by those of ordinary skill in the relevant art or derived from the present disclosure. Therefore, the specific effects of the present disclosure are not limited to the effects explicitly described in the present disclosure and can include various effects that can be understood or derived from the technical features of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Schematically illustrates an example of a video / image encoding system to which the present disclosure is applicable.

[0031] Figure 2 Is a diagram schematically illustrating the configuration of a video / image encoding device to which the present disclosure is applicable.

[0032] Figure 3 Is a diagram schematically illustrating the configuration of a video / image decoding device to which the present disclosure is applicable.

[0033] Figure 4 Schematically illustrates a multiple transform technique according to an embodiment of the present disclosure.

[0034] Figure 5 Illustrates an MTS applied to sub-block transform according to an example of the present disclosure.

[0035] Figure 6 Illustrates 32-point zeroing applied to sub-block transform according to an example of the present disclosure.

[0036] Figure 7It is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.

[0037] Figure 8 It is a flowchart illustrating the process for a video decoding device to derive transform coefficients according to an embodiment of the present disclosure.

[0038] Figure 9 It is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present disclosure.

[0039] Figure 10 It is a flowchart illustrating the process for a video encoding device to encode transform coefficients and information according to an embodiment of the present disclosure.

[0040] Figure 11 Illustrates the structure of a content streaming system to which the present disclosure is applied. Detailed Embodiments

[0041] This document can be modified in various ways and can have various embodiments, and specific embodiments will be illustrated and described in detail in the accompanying drawings. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are used to describe specific embodiments, rather than to limit the technical spirit of this document. Unless clearly indicated otherwise in the context, singular expressions include plural expressions. Terms such as "including" or "having" in this specification should be understood as indicating the presence of the features, numbers, steps, operations, elements, components, or combinations thereof described in this specification, without excluding the possibility of the presence or addition of one or more features, numbers, steps, operations, elements, components, or combinations thereof.

[0042] In addition, for the convenience of describing features and functions related to different aspects, the elements in the accompanying drawings described in this document are illustrated independently. This does not mean that each element is implemented as a separate piece of hardware or separate software. For example, at least two elements can be combined to form a single element, or a single element can be divided into multiple elements. Embodiments in which elements are combined and / or separated are also included within the scope of the rights of this document, unless it deviates from the essence of this document.

[0043] Hereinafter, the preferred embodiments of this document will be described more specifically with reference to the accompanying drawings. Hereinafter, in the accompanying drawings, the same reference numerals are used for the same elements, and redundant descriptions of the same elements may be omitted.

[0044] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Recommendation H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Recommendation H.265), the EVC (essential video coding) standard, the AVS2 standard, etc.).

[0045] In this document, various embodiments related to video / image coding can be provided, and, unless otherwise specified, the embodiments can be combined with and executed by each other.

[0046] In this document, video can mean a collection of a series of images over time. Generally, a picture means a unit representing an image of a specific time region, and a slice / tile is a unit that forms part of a picture. A slice / tile can include one or more coding tree units (CTUs). A picture can be composed of one or more slices / tiles. A picture can be composed of one or more tile groups. A tile group can include one or more tiles.

[0047] A pixel or pel can mean the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or the value of a pixel, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.

[0048] A unit can represent a basic unit of image processing. A unit can include at least one of a specific region and information related to that region. A unit can include one luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the situation, the terms unit and terms such as block, region, etc. can be used interchangeably. Generally, an MxN block can include a set (or array) of samples (or sample arrays) or transform coefficients composed of M columns and N rows.

[0049] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" can mean "A and / or B". Additionally, "A, B" can mean "A and / or B". Additionally, "A / B / C" can mean "at least one of A, B, and / or C". Additionally, "A / B / C" can mean "at least one of A, B, and / or C".

[0050] In addition, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" can include 1) "only A", 2) "only B", and / or 3) both "A and B". In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".

[0051] In the present disclosure, "at least one of A and B" can mean "only A", "only B", or both "A and B". Additionally, in the present disclosure, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".

[0052] In addition, in the present disclosure, "at least one of A, B, and C" can mean "only A", "only B", "only C", or any combination of "A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".

[0053] In addition, the parentheses used in the present disclosure can mean "for example". Specifically, when indicated as "prediction (intra prediction)", this may mean that "intra prediction" is presented as an example of "prediction". That is, "prediction" in the present disclosure is not limited to "intra prediction", and "intra prediction" can be presented as an example of "prediction". Additionally, when indicated as "prediction (i.e., intra prediction)", this may also mean that "intra prediction" is presented as an example of "prediction".

[0054] The technical features separately described in one of the drawings in the present disclosure can be implemented separately or can be implemented simultaneously.

[0055] Figure 1 An example of a video / image coding system to which the embodiments of this document can be applied is schematically illustrated.

[0056] Referring to Figure 1 , the video / image coding system can include a first device (source device) and a second device (receiving device). The source device can transfer the encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.

[0057] The source device can include a video source, an encoding device, and a transmitter. The receiving device can include a receiver, a decoding device, and a renderer. The encoding device can be referred to as a video / image encoding device, and the decoding device can be referred to as a video / image decoding device. The transmitter can be included in the encoding device. The receiver can be included in the decoding device. The renderer can include a display, and the display can be configured as a separate device or an external component.

[0058] A video source can obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet computer, and a smart phone, and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer or the like. In this case, the video / image capture process can be replaced by a process of generating relevant data.

[0059] An encoding device can encode the input video / images. The encoding device can perform a series of processes such as prediction, transformation, and quantization for compression and compilation efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0060] A transmitter can send the encoded video / image information or data output in the form of a bitstream to the receiver of a receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating a media file in a predetermined file format and can include elements for sending via a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to a decoding device.

[0061] A decoding device can decode the video / images by performing a series of processes such as dequantization, inverse transformation, prediction, etc., corresponding to the operations of the encoding device.

[0062] A renderer can render the decoded video / images. The rendered video / images can be displayed on a display.

[0063] Figure 2 is a diagram schematically depicting the configuration of a video / image encoding device to which this document can be applied. Hereinafter, the so-called video encoding device can include an image encoding device.

[0064] Refer to Figure 2, the encoding device 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-predictor 221 and an intra-predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a de-quantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 described above may be constituted by one or more hardware components (e.g., an encoder chipset or a processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.

[0065] The image partitioner 210 partitions an input image (or picture or frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a largest coding unit (LCU), the coding unit may be recursively partitioned according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, a coding unit may be divided into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure may be applied first, and then the binary tree structure and / or the ternary tree structure may be applied. Alternatively, the binary tree structure may be applied first. The coding process according to this document may be performed based on the final coding unit that is not further partitioned. In this case, based on the coding efficiency according to the image characteristics, the largest coding unit may be directly used as the final coding unit. Alternatively, the coding unit may be recursively partitioned into coding units with an even deeper depth as needed, such that the coding unit with the optimal size may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be divided or split from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal based on the transformation coefficients.

[0066] Depending on the context, terms such as unit and terms like block, region, etc. can be used interchangeably. Under normal circumstances, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can generally represent pixels or pixel values, and can represent only the pixels / pixel values of the luminance component, or only the pixels / pixel values of the chrominance component. Samples can be used as a term corresponding to the pixels or pels of a picture (or image).

[0067] The subtractor 231 subtracts the prediction signal (prediction block, prediction sample, or prediction sample array) output from the predictor 220 from the input image signal (original block, original sample, or original sample array) to generate a residual signal (residual block, residual sample, or residual sample array), and the generated residual signal is sent to the transformer 232. The predictor 220 can perform prediction on the block to be processed (hereinafter referred to as the "current block") and can generate a prediction block including the prediction samples of the current block. The predictor 220 can determine whether to apply intra prediction or inter prediction on the basis of the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various information related to prediction, such as prediction mode information, and send the generated information to the entropy encoder 240. Information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.

[0068] The intra predictor 222 can predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the reference samples can be located near the current block or separated from the current block. In intra prediction, the prediction mode can include a variety of non - directional modes and a variety of directional modes. The non - directional modes can include, for example, the DC mode and the planar mode. Depending on the level of detail of the prediction direction, the directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used depending on the settings. The intra predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0069] The inter - frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted on a block, sub - block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same as or different from each other. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter - frame predictor 221 can configure a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes. For example, in the skip mode and the merge mode, the inter - frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, a residual signal cannot be transmitted. In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of neighboring blocks can be used as a motion vector prediction term, and the motion vector of the current block can be indicated by signaling a motion vector difference.

[0070] The predictor 220 can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra - frame prediction or inter - frame prediction to predict a block, and can also apply intra - frame prediction and inter - frame prediction simultaneously. This can be referred to as combined intra - inter prediction (CIIP). In addition, the predictor can perform prediction on a block based on the intra - block copy (IBC) prediction mode or the palette mode. The IBC prediction mode or the palette mode can be used for content image / video coding of games, etc., such as screen content coding (SCC). Although IBC basically performs prediction in the current block, the aspect of deriving a reference block in the current block can be performed similarly to inter - frame prediction. That is, IBC can use at least one of the inter - frame prediction techniques described in this disclosure.

[0071] The prediction signal generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate a reconstructed signal or a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include a discrete cosine transform (DCT), a discrete sine transform (DST), a graph-based transform (GBT), or a conditional non-linear transform (CNT), etc. Here, GBT means a transform obtained from a graph when the relationship information between pixels is represented by the graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform process can be applied to square pixel blocks of the same size or can be applied to blocks of variable size instead of square blocks.

[0072] Quantizer 233 may quantize the transform coefficients and send them to entropy encoder 240, and entropy encoder 240 may encode the quantized signal (information regarding the quantized transform coefficients) and output the encoded signal in a bitstream. The information regarding the quantized transform coefficients may be referred to as residual information. Quantizer 233 may rearrange the block-based quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Entropy encoder 240 may perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. Entropy encoder 240 may encode, together or separately, information necessary for video / image reconstruction other than the quantized transform coefficients (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) may be sent or stored in the form of a bitstream on a unit basis of a network abstraction layer (NAL). The video / image information may also include information regarding various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc. In addition, the video / image information may also include general constraint information. In the present disclosure, the information and / or syntax elements sent / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above encoding process and included in the bitstream. The bitstream may be sent through a network or stored in a digital storage medium. Here, the network may include a broadcast network, a communication network, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that sends the signal output from entropy encoder 240 and / or a storage device (not shown) that stores it may be configured as an internal / external element of encoding device 200, or the transmitter may be included in entropy encoder 240.

[0073] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transformation to the quantized transform coefficients via the dequantizer 234 and the inverse transform unit 235, a residual signal (residual block or residual sample) can be reconstructed. The adder 250 adds the reconstructed residual signal to the prediction signal output from the predictor 220, so that a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample, or reconstructed sample array) can be generated. When the target block has no residual as in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra prediction of the next processing target block in the current block, and as will be described later, can be used for inter prediction of the next picture through filtering.

[0074] Meanwhile, during picture encoding and / or reconstruction, luminance mapping and chrominance scaling (LMCS) can be applied.

[0075] The filter 260 can improve the subjective / objective video quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and can store the modified reconstructed picture in the memory 270, specifically in the DPB of the memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. As will be discussed later in the description of each filtering method, the filter 260 can generate various information related to filtering, and send the generated information to the entropy encoder 290. The information about filtering can be encoded in the entropy encoder 290 and output in the form of a bitstream.

[0076] The modified reconstructed picture that has been sent to the memory 270 can be used as a reference picture in the inter predictor 280. By this, the encoding device can avoid prediction mismatches in the encoding device 200 and the decoding device when applying inter prediction, and can also improve the encoding efficiency.

[0077] The memory 270 DPB can store the modified reconstructed picture so as to use it as a reference picture in the inter predictor 221. The memory 270 can store the motion information of the blocks in the current picture from which the motion information (or encoded for it) has been derived, and / or the motion information of the blocks in the pictures that have been reconstructed. The stored motion information can be sent to the inter predictor 221 to be used as the motion information of neighboring blocks or temporally neighboring blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture, and send them to the intra predictor 222.

[0078] Figure 3 is a diagram schematically depicting the configuration of a video / image decoding device to which this document can be applied.

[0079] Reference Figure 3 , the video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be constituted by one or more hardware components (e.g., a decoder chipset or a processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0080] When receiving a bitstream including video / image information, the decoding device 300 may reconstruct an image corresponding to the processing of the video / image information in the Figure 2 encoding device accordingly. For example, the decoding device 300 may derive units / blocks based on information related to block segmentation obtained from the bitstream. The decoding device 300 may perform decoding by using the processing units applied in the encoding device. Thus, the decoded processing units may be, for example, coding units, which may be segmented along a quadtree structure, a binary tree structure, and / or a ternary tree structure with coding tree units or maximum coding units. One or more transform units may be derived with coding units. And, the reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproducer.

[0081] The decoding device 300 is capable of receiving, in the form of a bitstream, from Figure 2a signal output by an encoding device, and the received signal can be decoded by an entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can also include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), video parameter sets (VPS), etc. In addition, the video / image information can also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. In the present disclosure, the signaled / received information and / or syntax elements described later can be decoded through the decoding process and are obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on coding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements necessary for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the decoding target syntax element information and decoding information of the neighboring and decoding target blocks or the information of the symbols / bins decoded in the previous step to determine the context model, predict the bin generation probability according to the determined context model, and perform arithmetic decoding on the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model using the information of the symbols / bins decoded by the context model for the next symbol / bin after determining the context model. Among the information decoded in the entropy decoder 310, the information about prediction can be provided to the predictor 330, and the information about the residuals that has been entropy decoded in the entropy decoder 310, i.e., the quantized transform coefficients and the associated parameter information, can be input to the dequantizer 321. In addition, among the information decoded in the entropy decoder 310, the information about filtering can be provided to the filter 350. Meanwhile, a receiver (not shown) that receives the signal output by the encoding device can further configure the decoding device 300 as internal / external components, and the receiver can be a component of the entropy decoder 310. Meanwhile, the decoding device according to the present disclosure can be referred to as a video / image / picture encoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, while the sample decoder can include at least one of the dequantizer 321, the inverse transformator 322, the predictor 330, the adder 340, the filter 350, and the memory 360.

[0082] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the order of coefficient scanning that has been performed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.

[0083] The inverse transformer 322 obtains a residual signal (residual block, residual sample array) by performing an inverse transform on the transform coefficients.

[0084] The predictor can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and specifically can determine the intra / inter prediction mode.

[0085] The predictor can generate a prediction signal based on various prediction methods. For example, the predictor can apply intra prediction or inter prediction to predict a block, and can also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). Additionally, the predictor can perform intra block copy (IBC) to predict a block. Intra block copy can be used for content image / video compilation such as games, such as screen content compilation (SCC). Although IBC basically performs prediction in the current block, the aspect of deriving a reference block in the current block can be performed similarly to inter prediction. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.

[0086] The intra predictor 331 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the samples referred to can be located in the neighborhood of the current block or far from the current block. In intra prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to neighboring blocks.

[0087] The inter-frame predictor 332 may derive a prediction block of a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information may be predicted on a block, sub-block, or sample basis based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter-frame predictor 332 may configure a motion information candidate list based on neighboring blocks, and derive a motion vector and / or a reference picture index of the current block based on the received candidate selection information. The inter-frame prediction may be performed based on various prediction modes, and information about the prediction may include information indicating an inter-frame prediction mode for the current block.

[0088] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (prediction block, prediction sample array) output from the predictor 330. When there is no residual for a target block as in the case of applying the skip mode, the prediction block may be used as the reconstructed block.

[0089] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, may be output through filtering as described below, or may be used for inter-frame prediction of the next picture.

[0090] In addition, luminance mapping and chrominance scaling (LMCS) may be applied to the picture decoding process.

[0091] The filter 350 may improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 may generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 360, specifically, in the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering processing, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.

[0092] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter - predictor 332. The memory 360 can store the motion information of the blocks in the current picture from which the motion information has been derived (or decoded), and / or the motion information of the blocks in the reconstructed picture. The stored motion information can be sent to the inter - predictor 260 to be used as the motion information of neighboring blocks or temporally neighboring blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and send them to the intra - predictor 331.

[0093] In this specification, the examples described in the predictor 330, de - quantizer 321, inverse transformer 322, and filter 350 of the decoding device 300 can be similarly or correspondingly applied to the predictor 220, de - quantizer 234, inverse transformer 235, and filter 260 of the encoding device 200, respectively.

[0094] As described above, when performing video coding, prediction is performed to improve the compression efficiency. A prediction block including prediction samples of the current block, that is, the target coding block, can be generated through prediction. In this case, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived in the same way in both the encoding device and the decoding device. The encoding device can improve the image coding efficiency by signaling information about the residual between the original block and the prediction block rather than the original sample values of the original block itself (residual information) to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, can generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and can generate a reconstructed picture including the reconstructed block.

[0095] Residual information can be generated through a transform and quantization process. For example, the encoding device can derive a residual block between the original block and the prediction block, can derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, can derive quantized transform coefficients by performing a quantization process on the transform coefficients, and can signal (through the bitstream) the relevant residual information to the decoding device. In this case, the residual information can include information such as value information, position information, transform scheme, transform kernel, and quantization parameters of the quantized transform coefficients. The decoding device can perform an inverse quantization / inverse transform process based on the residual information and can derive residual samples (or a residual block). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. In addition, the encoding device can perform inverse quantization / inverse transform on the quantized transform coefficients to derive a residual block for reference in the inter - prediction of subsequent pictures and can generate a reconstructed picture.

[0096] Figure 4 A multi - transform technique according to an embodiment of the present disclosure is schematically illustrated.

[0097] With reference to Figure 4 , the transformer can correspond to the transformer in the encoding device as described above Figure 2 , and the inverse transformer can correspond to the inverse transformer in the encoding device as described above Figure 2 , or can correspond to the inverse transformer in the decoding device as described above Figure 3 .

[0098] The transformer can derive (primary) transform coefficients (S410) by performing a primary transform based on the residual samples (residual sample array) in the residual block. This primary transform can be referred to as the core transform. In this document, the primary transform can be based on multi-transform selection (MTS), and when multi-transform is applied as the primary transform, it can be referred to as multi-core transform.

[0099] The multi-core transform can represent a method of additionally using Discrete Cosine Transform (DCT) type 2 and Discrete Sine Transform (DST) type 7, DCT type 8, and / or DST type 1 for transformation. That is, the multi-core transform can represent a transform method of transforming a residual signal (or residual block) in the spatial domain into transform coefficients (or primary transform coefficients) in the frequency domain based on multiple transform kernels selected from DCT type 2, DST type 7, DCT type 8, and DST type 1. In this document, the primary transform coefficients can be referred to as temporal transform coefficients from the perspective of the transformer.

[0100] That is, when applying a conventional transform method, transform coefficients can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2. However, when applying the multi-core transform, transform coefficients (or primary transform coefficients) can be generated by applying a transform from the spatial domain to the frequency domain to the residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1. Here, DCT type 2, DST type 7, DCT type 8, and DST type 1 can be referred to as transform types, transform kernels, or transform cores. These DCT / DST types can be defined based on basis functions.

[0101] If the multi-core transform is performed, a vertical transform kernel and a horizontal transform kernel for the target block can be selected from among the transform kernels, the vertical transform can be performed on the target block based on the vertical transform kernel, and the horizontal transform can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform can represent the transform for the horizontal component of the target block, and the vertical transform can represent the transform for the vertical component of the target block. The vertical transform kernel / horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target block (CU or sub-block) including the residual block.

[0102] In addition, according to an example, if the primary transformation is performed by applying MTS, the mapping relationship of the transform kernel can be set by setting a specific basis function to a predetermined value and combining the basis functions to be applied in the vertical transformation or the horizontal transformation. For example, when the horizontal transform kernel is expressed as trTypeHor and the vertical transform kernel is expressed as trTypeVer, the value 0 of trTypeHor or trTypeVer can be set to DCT2, the value 1 of trTypeHor or trTypeVer can be set to DST-7, and the value 2 of trTypeHor or trTypeVer can be set to DCT-8.

[0103] In this case, the MTS index information can be encoded and signaled to the decoding device to indicate any one of a plurality of transform kernel sets. For example, MTS index 0 can indicate that both the trTypeHor and trTypeVer values are 0, MTS index 1 can indicate that both the trTypeHor and trTypeVer values are 1, MTS index 2 can indicate that the trTypeHor value is 2 and the trTypeVer value is 1, MTS index 3 can indicate that the trTypeHor value is 1 and the trTypeVer value is 2, and MTS index 4 can indicate that both the trTypeHor and trTypeVer values are 2.

[0104] In one example, the transform kernel sets according to the MTS index information are illustrated in the following table.

[0105] [Table 1]

[0106] tu_mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTvpeVer 0 1 1 2 2

[0107] The transformer can derive a modified (secondary) transform coefficient by performing a secondary transform based on a (primary) transform coefficient (S420). The primary transform is a transform from the spatial domain to the frequency domain, and the secondary transform refers to a transform that uses the correlation existing between the (primary) transform coefficients for a more compressed representation. The secondary transform may include a non-separable transform. In this case, the secondary transform may be referred to as a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform can represent generating a modified transform coefficient (or secondary transform coefficient) for a residual signal by performing a secondary transform on the (primary) transform coefficients derived by the primary transform based on a non-separable transform matrix. At this time, the vertical transform and the horizontal transform may not be separately applied to the (primary) transform coefficients (or may not be independently applied to the horizontal transform and the vertical transform), but the transform may be applied at once based on the non-separable transform matrix. In other words, the non-separable secondary transform can represent a transform method in which the vertical component and the horizontal component of the (primary) transform coefficients are not separated, and for example, a two-dimensional signal (transform coefficient) is rearranged into a one-dimensional signal in a certain determined direction (e.g., row-major order or column-major order), and then a modified transform coefficient (or secondary transform coefficient) is generated based on the non-separable transform matrix. For example, according to the row-major order, M×N blocks are arranged in a row in the order of the first row, the second row, …, and the Nth row. According to the column-major order, M×N blocks are arranged in a row in the order of the first column, the second column, …, and the Nth column. The non-separable secondary transform can be applied to the upper left region of a block configured with (primary) transform coefficients (hereinafter, may be referred to as a transform coefficient block). For example, if both the width (W) and the height (H) of the transform coefficient block are equal to or greater than 8, an 8×8 non-separable secondary transform can be applied to the upper left 8×8 region of the transform coefficient block. In addition, if both the width (W) and the height (H) of the transform coefficient block are equal to or greater than 4, and the width (W) or the height (H) of the transform coefficient block is less than 8, a 4×4 non-separable secondary transform can be applied to the upper left min(8, W)×min(8, H) region of the transform coefficient block. However, this embodiment is not limited thereto, and for example, even if only the width (W) or the height (H) of the transform coefficient block satisfies the condition of being equal to or greater than 4, a 4×4 non-separable secondary transform can be applied to the upper left min(8, W)×min(8, H) region of the transform coefficient block.

[0108] The transformer can perform a non-separable secondary transform based on a selected transform kernel and can obtain a modified (secondary) transform coefficient. As described above, the modified transform coefficient can be derived as a transform coefficient quantized by a quantizer, can be encoded and signaled to a decoding device, and is transmitted to a dequantizer / inverse transformer in an encoding device.

[0109] Meanwhile, as described above, if the secondary transform is omitted, the (primary) transform coefficients that are the output of the primary (separable) transform can be derived as the transform coefficients quantized by the quantizer as described above, and can be encoded and signaled to the decoding device, and transmitted to the dequantizer / inverse transformer in the encoding device.

[0110] The inverse transformer can perform the series of processes in an order opposite to the order in which the series of processes have been performed in the above transformer. The inverse transformer can receive the (dequantized) transform coefficients, and derive the (primary) transform coefficients by performing the secondary (inverse) transform (S450), and can obtain the residual block (residual samples) by performing the primary (inverse) transform on the (primary) transform coefficients (S460). In this regard, from the perspective of the inverse transformer, the primary transform coefficients can be referred to as modified transform coefficients. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.

[0111] The decoding device may further include a secondary inverse transform application determiner (or an element for determining whether to apply the secondary inverse transform) and a secondary inverse transform determiner (or an element for determining the secondary inverse transform). The secondary inverse transform application determiner can determine whether to apply the secondary inverse transform. For example, the secondary inverse transform can be NSST or RST, and the secondary inverse transform application determiner can determine whether to apply the secondary inverse transform based on the secondary transform flag obtained by parsing the bitstream. In another example, the secondary inverse transform application determiner can determine whether to apply the secondary inverse transform based on the transform coefficients of the residual block.

[0112] The secondary inverse transform determiner can determine the secondary inverse transform. In this case, the secondary inverse transform determiner can determine the secondary inverse transform to be applied to the current block based on the NSST (or RST) transform set specified according to the intra prediction mode. In an embodiment, the secondary transform determination method can be determined depending on the primary transform determination method. Various combinations of the primary transform and the secondary transform can be determined according to the intra prediction mode. In addition, in an example, the secondary inverse transform determiner can determine the region to which the secondary inverse transform is applied based on the size of the current block.

[0113] Meanwhile, as described above, if the secondary (inverse) transform is omitted, the (dequantized) transform coefficients can be received, the primary (separable) inverse transform can be performed, and the residual block (residual samples) can be obtained. As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and the prediction block, and can generate a reconstructed picture based on the reconstructed block.

[0114] Meanwhile, in the present disclosure, a reduced secondary transform (RST) with a reduced size of a transform matrix (kernel) can be applied in the context of NSST so as to reduce the amount of computation and memory required for the non-separable secondary transform.

[0115] Meanwhile, the transform kernel, transform matrix, and coefficients constituting the transform kernel matrix, i.e., kernel coefficients or matrix coefficients, described in the present disclosure can be expressed in 8 bits. This can be a condition implemented in a decoding device and an encoding device, and can reduce the amount of memory required to store the transform kernel due to a reasonable performance degradation compared to existing 9 bits or 10 bits. In addition, expressing the kernel matrix in 8 bits can allow the use of a small multiplier and may be more suitable for single instruction multiple data (SIMD) instructions for optimal software implementation.

[0116] In this specification, the term "RST" may refer to a transform performed on residual samples of a target block based on a transform matrix whose size is reduced according to a reduction factor. In the case of performing a reduction transform, since the size of the transform matrix is reduced, the amount of computation required for the transform can be reduced. That is, RST can be used to solve the computational complexity problem that occurs during non-separable transforms or transforms of large-sized blocks.

[0117] RST can be referred to by various terms such as reduction transform, reduced secondary transform, reduced transform, simplified transform, simple transform, etc., and the names by which RST can be referred to are not limited to the examples listed. Alternatively, since RST is mainly performed in the low-frequency region including non-zero coefficients in a transform block, it can be referred to as a low-frequency non-separable transform (LFNST).

[0118] Meanwhile, when performing an inverse secondary transform based on RST, the inverse transformers 235 of the encoding device 200 and 322 of the decoding device 300 can include an inverse reduced secondary transformer that derives modified transform coefficients based on an inverse RST of the transform coefficients, and an inverse primary transformer that derives residual samples of the target block based on an inverse primary transform of the modified transform coefficients. The inverse primary transform refers to the inverse transform of the primary transform applied to the residuals. In the present disclosure, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying the transform.

[0119] Hereinafter, a reduced multi-transform technique (reduced adaptive multi-selection (or set) (RMTS)) is described.

[0120] As described above, when a combination of multiple transforms (DCT-2, DST-7, DCT-8, DST-1, DCT-5, etc.) is selectively used in a multi-transform technique (multi-transform set or adaptive multi-transform) for the primary transform, the transform can be applied only to predefined regions to reduce complexity, rather than performing the transform in all cases, thereby significantly reducing complexity in the worst case.

[0121] For example, when applying the primary transform to an M×M pixel block based on the previous reduction transform (RT) method, calculations can be performed only on the transform blocks of an R×R block (M >= R), rather than obtaining an M×M transform block. As a result, non-zero valid coefficients exist only in the R×R region, and the transform coefficients in other regions can be considered zero without being calculated. The following table illustrates three examples of reduced adaptive multi-transform (RAMT) using predefined reduction transform factor (R) values for the size of the block used to apply the primary transform.

[0122] [Table 2]

[0123] Transform size Reduction transform 1 Reduction transform 2 Reduction transform 3 8x8 4x4 6x6 6x6 16x16 8x8 12x12 8x8 32x32 16x16 16x16 16x16 64x64 32x32 16x16 16x16 128x128 32x32 16x16 16x16

[0124] According to one example, when applying the reduced multi-transform illustrated above, the reduction transform factor can be determined based on the primary transform. For example, when the primary transform is DCT2, it is computationally simple compared to other primary transforms, and thus the reduction transform may not be used for small blocks or a relatively large R value can be used for small blocks, thereby minimizing the reduction in compilation performance. For example, different reduction transform factors can be used for DCT2 and other transforms as follows.

[0125] [Table 3]

[0126] Transform size Reduction transform for DCT2 Reduction transform other than DCT2 8x8 8x8 4x4 16x16 16x16 8x8 32x32 32x32 16x16 64x64 32x32 32x32 128x128 32x32 32x32

[0127] As shown in Table 3, when the primary transform is DCT2, the size of the transform remains unchanged when the size of the block to be transformed is 8×8 or 16×16, and the reduced size of the transform is limited to 32×32 when the size of the block is 32×32 or larger.

[0128] Alternatively, according to the example, when the flag value indicating whether to apply MTS is 0 (i.e., when DCT2 is applied for both the horizontal and vertical directions), for both (horizontal and vertical) directions, only 32 coefficients from the left and top can be left and the high-frequency components can be zeroed, that is, set to 0 (zeroing implementation method 1).

[0129] For example, in a 64×64 transform unit (TU), the transform coefficients are only left in the upper-left 32×32 region; in a 64×16 TU, the transform coefficients are only left in the upper-left 32×16 region; and in an 8×64 TU, the transform coefficients are only left in the upper-left 8×32 region. That is to say, there are transform coefficients corresponding to a maximum length of only up to 32 in both width and height.

[0130] This zeroing method can be applied only to the residual signal to which intra prediction is applied or can be applied only to the residual signal to which inter prediction is applied. Alternatively, the zeroing method can be applied to both the residual signal to which intra prediction is applied and the residual signal to which inter prediction is applied.

[0131] The change in the transform block size that can be expressed as the previous zeroing or high-frequency zeroing is a process of zeroing (determining to be 0) the transform coefficients related to a certain value or higher frequency in a (transform) block having a first width (or length) of W1 and a first height (or length) of H1. When high-frequency zeroing is applied, the transform coefficient values of all transform coefficients outside the low-frequency transform coefficient region configured based on a second width of W2 and a second height of H2 in the (transform) block can be determined (set) to 0. The outside of the low-frequency transform coefficient region can be referred to as the high-frequency transform coefficient region. In the example, the low-frequency transform coefficient region can be a rectangular region starting from the upper-left of the (transform) block.

[0132] That is to say, high-frequency zeroing can be defined as setting all transform coefficients at positions defined by an x coordinate of w or greater and a y coordinate of h or greater to 0, where the horizontal x coordinate value of the upper-left position of the current transform block (TB) is set to 0 and its vertical y coordinate value is set to 0 (and where the x coordinate increases from left to right and the y coordinate increases downward).

[0133] In the present disclosure, specific terms or expressions are used to define specific information or concepts. For example, as described above, in this specification, the process of zeroing the transform coefficients corresponding to a frequency of a specific value or greater value in a (transform) block having a first width (or length) of W1 and a first height (or length) of H1 is defined as "high-frequency zeroing", the region that has undergone zeroing through high-frequency zeroing is defined as the "high-frequency transform coefficient region", and the region that has not undergone zeroing is defined as the "low-frequency transform coefficient region". To indicate the size of the low-frequency transform coefficient region, a second width (or length) of W2 and a second height (or length) of H2 are used.

[0134] However, "high frequency zeroing" can be replaced with various terms such as high frequency zeroing, high frequency zeroing-out, high-frequency zeroing-out, high-frequency zero-out, and zero-out, and "high frequency transform coefficient region" can be replaced with various terms such as high frequency zeroing application region, high frequency zeroing region, high frequency region, high frequency coefficient region, high frequency zeroing region, and zeroing region, and "low frequency transform coefficient region" can be replaced with various terms such as high frequency zeroing non-application region, low frequency region, low frequency coefficient region, and restricted region. Therefore, specific terms or expressions used in this document to define specific information or concepts need to be interpreted throughout the specification according to the content indicated by the term, in view of various operations, functions, and effects, and are not limited to the designation.

[0135] Alternatively, according to an example, the low frequency transform coefficient region refers to the region remaining after performing high frequency zeroing or the region in which valid transform coefficients are left, that is, the region where non-zero transform coefficients can exist, and can be referred to as the zeroing region or zeroing block.

[0136] According to an example, when the flag value indicating whether to apply MTS is 1, that is, when different transforms (DST-7 or DCT-8) other than DCT2 are applicable to the horizontal and vertical directions, the transform coefficients can be left only in the upper left region, and the remaining region can be zeroed as follows (zeroing implementation 2).

[0137] - When the width (w) is equal to or greater than 2 n only the transform coefficients corresponding to the length of w / 2 starting from the left p can be left, and the remaining transform coefficients can be fixed to 0 (zeroed).

[0138] - When the height (h) is equal to or greater than 2 m only the transform coefficients corresponding to the length of h / 2 starting from the top q can be left, and the remaining transform coefficients can be fixed to 0 (zeroed).

[0139] Here, m, n, p, and q can be integers equal to or greater than 0, and can be specifically as follows.

[0140] 1) (m,n,p,q) = (5,5,1,1)

[0141] 2) (m,n,p,q) = (4,4,1,1)

[0142] In configuration 1), the transform coefficients are only left in the upper-left 16×16 region of the 32×16 TU, and the transform coefficients are only left in the upper-left 8×16 region of the 8×32 TU.

[0143] This zeroing method can be applied only to the residual signal to which intra prediction is applied or can be applied only to the residual signal to which inter prediction is applied. Alternatively, the zeroing method can be applied to both the residual signal to which intra prediction is applied and the residual signal to which inter prediction is applied.

[0144] Alternatively, according to another example, when the flag value indicating whether to apply MTS is 1, that is, when different transforms (DST-7 or DCT-8) other than DCT2 are applicable to the horizontal and vertical directions, the transform coefficients can be left only in the upper-left region, and the remaining regions can be zeroed as follows (zeroing implementation 3).

[0145] - When the height (h) is equal to or greater than the width (w) and equal to or greater than 2 n only the transform coefficients in the upper-left w×(h / 2 p ) region can be left, and the remaining transform coefficients can be fixed to 0 (zeroed).

[0146] - When the width (w) is greater than the height (h) and equal to or greater than 2 m only the transform coefficients in the upper-left (w / 2 q )×h region can be left, and the remaining transform coefficients can be fixed to 0 (zeroed).

[0147] Under the above conditions, when the height (h) and the width (w) are the same, the vertical length is reduced (h / 2p), but the horizontal length can also be reduced (w / 2q).

[0148] Here, m, n, p, and q can be integers equal to or greater than 0, and can be specifically as follows.

[0149] 1) (m, n, p, q) = (4, 4, 1, 1)

[0150] 2) (m, n, p, q) = (5, 5, 1, 1)

[0151] In configuration 1), the transform coefficients are only left in the upper-left 16×16 region of the 32×16 TU, and the transform coefficients are only left in the upper-left 8×8 region of the 8×16 TU.

[0152] This zeroing method can be applied only to the residual signal to which intra prediction is applied or can be applied only to the residual signal to which inter prediction is applied. Alternatively, the zeroing method can be applied to both the residual signal to which intra prediction is applied and the residual signal to which inter prediction is applied.

[0153] In the foregoing embodiments, the transform coefficient region is restricted depending on whether the flag value indicating whether to apply MTS is 0 or the flag value indicating whether to apply MTS is 1. According to one example, a combination of these embodiments is possible.

[0154] 1) Zeroing Embodiment 1 + Zeroing Embodiment 2

[0155] 2) Zeroing Embodiment 1 + Zeroing Embodiment 3

[0156] As mentioned in Zeroing Embodiment 2 and Zeroing Embodiment 3, the zeroing method can be applied only to the residual signal to which intra prediction is applied or can be applied only to the residual signal to which inter prediction is applied. Alternatively, the zeroing method can be applied to both the residual signal to which intra prediction is applied and the residual signal to which inter prediction is applied. Therefore, when the MTS flag is 1, the following table can be configured (when the MTS flag is 1, Zeroing Embodiment 1 can be applied). Here, the MTS flag can also be configured to indicate the MTS index for the transform kernel for MTS. For example, an MTS index of 0 can indicate that Zeroing Embodiment 1 is applied.

[0157] [Table 4]

[0158] Configuration index Intra prediction residual signal Inter prediction residual signal 1 Do not apply zeroing Do not apply zeroing 2 Do not apply zeroing Zeroing implementation 2 3 Do not apply zeroing Zeroing implementation 3 4 Zeroing implementation 2 Do not apply zeroing 5 Zeroing implementation 2 Zeroing implementation 2 6 Zeroing implementation 2 Zeroing implementation 3 7 Zeroing implementation 3 Do not apply zeroing 8 Zeroing implementation 3 Zeroing implementation 2 9 Zeroing implementation 3 Zeroing implementation 3

[0159] In Zeroing Embodiment 1, Zeroing Embodiment 2, and Zeroing Embodiment 3, the regions in the TU that inevitably include the value 0 are clearly defined. That is, the regions other than the upper left region where transform coefficients are allowed to exist are zeroed. Therefore, according to the embodiment, it can be configured to bypass the regions where the transform coefficients clearly have the value 0 as a result of the entropy coding of the residual signal, rather than performing residual coding on them. For example, the following configurations are possible.

[0160] 1) In HEVC or VVC, a flag (subblock_flag) indicating whether there are non-zero transform coefficients in a coefficient group (CG, which can be a 4x4 or 2x2 block depending on the shape of the subblock and the TU block and the luminance component / chrominance component) is coded. Only when subblock_flag is 1, the inside of the CG is scanned and the coefficient level values are coded. Therefore, for the CG belonging to the region where zeroing is performed, subblock_flag can be default set to the value 0 instead of being coded.

[0161] 2) In HEVC or VVC, the positions of the last coefficients in the forward scan order (last_coefficient_position_x in the X direction and last_coefficient_position_y in the Y direction) are compiled first. Generally, last_coefficient_position_x and last_coefficient_position_y can respectively have a maximum value of (width of the TU - 1) and a maximum value of (height of the TU - 1). However, when the region where non-zero coefficients can exist is restricted due to zeroing, the maximum values of last_coefficient_position_x and last_coefficient_position_y are also restricted. Therefore, the maximum values of last_coefficient_position_x and last_coefficient_position_y can be restricted in view of zeroing and then can be compiled. For example, when the binarization method applied to last_coefficient_position_x and last_coefficient_position_y is truncated unary binarization, the maximum length of the truncated unary code (the codeword length that last_coefficient_position_x and last_coefficient_position_y can have) can be reduced based on the adjusted maximum value.

[0162] As described above, when zeroing is applied, especially in the case where the top-left 16x16 region is a low-frequency transform coefficient region (which can be hereinafter referred to as 32-point reduced MTS or RMT 32), zeroing can be applied both when the MTS technique is applied and when 32-point DST-7 or 32-point DCT-8 is applied.

[0163] Figure 5 The figure shows MTS applied to sub-block transform according to an example of the present disclosure.

[0164] According to the example, sub-block transform (SBT) can be applied in which a coding unit is divided into sub-blocks and a transform process is performed on the sub-blocks. The sub-block transform is applied to the residual signal generated by inter-frame prediction, and the residual signal block is partitioned into two partitioned sub-blocks, and a separable transform is applied to only one of the sub-blocks according to the sub-block transform. The sub-blocks can be divided in the horizontal direction or the vertical direction, and the width or height of the partitioned sub-blocks can be 1 / 2 or 1 / 4 of the coding unit. When the sub-block transform is applied, since only one of the two partitioned sub-blocks is transformed, there is only residual data in the transformed sub-block and no residual data in the remaining sub-block.

[0165] When the width or height of the sub-block to which the transform is applied is 64 or greater, DCT-2 can be applied in both the horizontal and vertical directions, while when the width and height of the sub-block to which the transform is applied are 32 or less, DST-7 or DCT-8 can be applied. Thus, when applying SBT, zeroing can be performed by applying RMTS32 only when all sides of the sub-block to which the transform is applied have a length of 32 or less. That is, DST-7 or DCT-8 with a length of 32 or less can be applied in each direction (horizontal and vertical direction), leaving at most 16 transform coefficients for each row or column.

[0166] As Figure 5 shown, when a block is partitioned and the transform is applied to region A, DST-7 or DCT-8 can be applied to each side and the transforms applied in the horizontal and vertical directions are not limited to Figure 5 the example illustrated in Figure 5 In

[0167] As Figure 5 shown, the block to which the transform is applied can be located on the left or right side within the entire block or on the top or bottom within the entire block. Additionally, Figure 5 the block can be a residual signal generated by inter-frame prediction. A flag indicating whether the transform is applied to only one sub-block of the residual signal partitioned as Figure 5 shown can be signaled, and when the flag is 1, a flag indicating whether the block is vertically partitioned or horizontally partitioned as Figure 5 shown can also be set by signaling.

[0168] A flag indicating whether the block A to which the transform is actually applied is located on the left or right side within the entire block or a flag indicating whether the block A is located on the top or bottom can also be signaled.

[0169] As Figure 5 illustrated, when determining the horizontal and vertical transforms for a specific block instead of specifying the horizontal and vertical transforms by MTS signaling, RMTS32 proposed above can be applied to each side when the corresponding sides in the horizontal and vertical directions have a length of 32. In RMTS32, residual compilation can be omitted for the zeroing region, or residual compilation can be performed by scanning only the non-zeroing region.

[0170] Figure 6 Illustrated is the 32-point zeroing applied to sub-block transform according to an example of the present disclosure.

[0171] When RMTS32 is applied to a sub-block of a partitioned block as Figure 5 shown, residual data may exist after transformation as Figure 6 shown. That is, Figure 6 shows RMTS32 being applied to a sub-block among blocks partitioned according to the application of sub-block transformation, to which transformation is performed.

[0172] With respect to the width w and height h of the original transform block, the width and height of block A to which transformation is actually applied may be w / 2 and h / 2 or w / 4 and h / 4, respectively.

[0173] In summary, if DST-7 or DCT-8 of length 32 can be applied in each of the horizontal and vertical directions, RMTS32 can be applied to any block to which transformation is applied. Whether to apply DST-7 or DCT-8 of length 32 can be determined by preset signaling or can be determined without signaling according to predetermined compilation conditions.

[0174] In the case where MTS is disabled (e.g., in VVC, MTS can be disabled when "sps_mts_enable_flag" is set to 0 in the sequence parameter set), when applying SBT, DCT-2 is applied in both the horizontal and vertical directions, instead of the combination of DST-7 and DCT-8 presented in Table 3.

[0175] Therefore, when MTS is disabled, even if the block to be transformed is partitioned into sub-blocks, it is necessary to prevent the application of RMTS32. As described above, when MTS is not applied, DCT-2 can be applied in the primary transformation instead of DST-7 or DCT-8, and the top-left block in which non-zero transform coefficients to which DCT-2 is applied can exist can undergo high-frequency zeroing that reduces the width and height to 32. That is, the top-left block in which non-zero transform coefficients to which inverse DCT-2 is applied can have its width and height reduced to 32, but is not zeroed to a width and height less than 32. This is to prevent data loss due to zeroing, and the width or height of the top-left block to which inverse DCT-2 is applied is not reduced to 16.

[0176] If MTS is disabled, it is necessary to explicitly check the width or height of the partitioned sub-block to which DCT-2 is applied so that the width or height of the partitioned sub-block is not reduced to 16.

[0177] According to an example, by determining "sps_mts_enabled_flag" in the residual compilation syntax, it can be configured not to perform zeroing due to RMTS32 when applying SBT.

[0178] The normative texts reflecting the foregoing embodiments may be shown in the following tables (Tables 5 to 15).

[0179] [Table 5]

[0180]

[0181]

[0182] [Table 6]

[0183]

[0184]

[0185] Tables 5 and 6 include the picture information signaled in the sequence parameter set for picture compilation and include several flag information related to transformation.

[0186] The sps_transform_skip_enabled_flag specifies whether transform skip is applied, that is, whether the transform_skip_flag can be present in the transform unit syntax.

[0187] The sps_mts_enabled_flag is the flag information that specifies whether MTS, i.e., the multiple transform selection technique, can be explicitly used. The sps_mts_enabled_flag being equal to 1 specifies that the sps_explicit_mts_intra_enabled_flag and the sps_explicit_mts_inter_enabled_flag are present in the sequence parameter set syntax.

[0188] When any one of the sps_explicit_mts_intra_enabled_flag and the sps_explicit_mts_inter_enabled_flag is 1, it is specified that MTS can be applied to the transform unit, which means that the tu_mts_idx can be present in the transform unit syntax. For example, when the sps_mts_enabled_flag is equal to 1 and the sps_explicit_mts_intra_enabled_flag is equal to 0, implicit MTS can be applied to the intra-compiled unit.

[0189] The sps_sbt_enabled_flag is the flag information that specifies whether the previous sub-block transform can be applied to the compiled unit for inter prediction.

[0190] When sps_sbt_enabled_flag is equal to 1, the sps_sbt_max_size_64_flag that signals the maximum width and height of the coding unit for enabling sub-block transform can be used.

[0191] When sps_sbt_max_size_64_flag is equal to 0, the maximum width and height of the coding unit for enabling block transform are 32, and when sps_sbt_max_size_64_flag is equal to 1, the maximum width and height of the coding unit for enabling sub-block transform are derived to be equal to 64.

[0192] [Table 7]

[0193]

[0194]

[0195] Table 7 shows the transform-related information signaled in the picture parameter set. log2_transform_skip_max_size_minus2 is the information for deriving the maximum block size for transform skip, and the maximum block size for transform skip is derived as the value of 2 to the power of (log2_transform_skip_max_size_minus2 + 2) (1<<(log2_transform_skip_max_size_minus2 + 2)).

[0196] [Table 8]

[0197]

[0198]

[0199] [Table 9]

[0200]

[0201]

[0202]

[0203] Table 8 and Table 9 show the syntax and semantics of the coding units to which inter prediction is applied, and the partition shape to which SBT is applied can be determined by the four syntax elements in Table 8.

[0204] The cu_sbt_flag specifies whether the SBT is applied to the compilation unit, and the cu_sbt_quad_flag is flag information that specifies whether the block to which the transformation is applied is 1 / 4 of the entire block when a compilation unit is partitioned into two sub-blocks. When cu_sbt_quad_flag is equal to 0, the partitioned sub-block has a size of 1 / 2 of the width or height of the compilation unit, and when cu_sbt_quad_flag is equal to 1, the partitioned sub-block has a size of 1 / 4 of the width or height of the compilation unit. When the width of the compilation unit is w and its height is h, the height of the partitioned block can be h1 = (1 / 4) x h or its width can be w1 = (1 / 4) x w.

[0205] cu_sbt_horizontal_flag being equal to 1 specifies that the compilation unit is partitioned horizontally, i.e., in the horizontal direction, and cu_sbt_horizontal_flag being equal to 0 specifies that the compilation unit is partitioned vertically, i.e., in the vertical direction.

[0206] When the cu_sbt_pos_flag value is equal to 0, the transformation is applied to the upper sub-block among the sub-blocks partitioned horizontally, and the transformation is applied to the left sub-block among the sub-blocks partitioned vertically. When the cu_sbt_pos_flag value is equal to 1, the transformation is applied to the lower sub-block among the sub-blocks partitioned horizontally, and the transformation is applied to the right sub-block among the sub-blocks partitioned vertically.

[0207] The following table shows trTypeHor and trTypeVer according to cu_sbt_horizontal_flag and cu_sbt_pos_flag.

[0208] [Table 10]

[0209] cu_sbt_horizontal_flag cu_sbt_pos_flag trTypeHor trTypeVer 0 0 2 1 0 1 1 1 1 0 1 2 1 1 1 1

[0210] As described above, when trTypeHor represents the horizontal transformation kernel and trTypeVer represents the vertical transformation kernel, a trTypeHor or trTypeVer value of 0 can be set for DCT-2, and a trTypeHor or trTypeVer value of 1 can be set for DST-7, and a trTypeHor or trTypeVer value of 2 can be set for DCT-8. Therefore, when the length of at least one side of the partitioned block to which the transformation is applied is 64 or greater, DCT-2 can be applied in both the horizontal and vertical directions; otherwise, DST-7 or DCT-8 can be applied. When partitioning the current block and performing a transformation on the sub-blocks, MTS can be implicitly applied as shown in Table 10.

[0211] When the width and height of a sub-block are both 32 or less and thus DST-7 or DCT-8 is applied, that is, when MTS is applied to the sub-block, the previous RMTS can be applied. For example, when the length of the sub-block in each direction is 32, only 16 transform coefficients can be left by applying DST-7 or DCT-8 of length 32.

[0212] However, when the width or height of a sub-block is 64 or greater, DCT-2 can be applied in both the horizontal and vertical directions, and only 32 transform coefficients can be left according to high-frequency zeroing, instead of applying RMTS in each direction.

[0213] [Table 11]

[0214]

[0215]

[0216] Table 11 shows parts of the syntax and semantics of the transform unit according to the example. tu_mts_idx[x0][y0] specifies the MTS index applied to the transform block, and trTypeHor and trTypeVer can be determined according to the MTS index as shown in Table 1.

[0217] According to another example, mts_idx can be signaled at the compilation unit level instead of at the transform unit level.

[0218] [Table 12]

[0219]

[0220]

[0221]

[0222] Table 12 shows parts of the syntax and semantics of the residual compilation according to the example.

[0223] The syntax elements last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix in Table 12 specify the (x, y) position information regarding the last non-zero transform coefficient in a transform block. Specifically, last_sig_coeff_x_prefix specifies the prefix of the column position of the last significant coefficient in the transform block in scan order, last_sig_coeff_y_prefix specifies the prefix of the row position of the last significant coefficient in the transform block in scan order, last_sig_coeff_x_suffix specifies the suffix of the column position of the last significant coefficient in the transform block in scan order, and last_sig_coeff_y_suffix specifies the suffix of the row position of the last significant coefficient in the transform block in scan order. Here, the significant coefficient can refer to a non-zero coefficient. Additionally, the scan order can be a top-right diagonal scan order. Alternatively, the scan order can be a horizontal scan order or a vertical scan order. The scan order can be determined based on whether intra / inter prediction is applied to the target block (CB or CB including TB) and / or a specific intra / inter prediction mode.

[0224] The zeroing region can be configured in the residual compilation of Table 12 based on tu_mts_idx[x0][y0] of Table 11.

[0225] Furthermore, when cu_sbt_flag is equal to 1, the height of the block to which the transform is applied is 32 or less (log2TbHeight < 6), its width is 32 (log2TbWidth < 6 && log2TbWidth > 4), and when sps_mts_enabled_flag is equal to 1, the width of the top-left region where non-zero transform coefficients can exist is set to 16 (log2ZoTbWidth = 4). Similarly, when cu_sbt_flag is equal to 1, the width of the block to which the transform is applied is 32 or less (log2TbWidth < 6), its height is 32 (log2TbHeight < 6 && log2TbHeight > 4), and when sps_mts_enabled_flag is equal to 1, the height of the top-left region where non-zero transform coefficients can exist is set to 16 (log2ZoTbHeight = 4).

[0226] The block to which the transform is applied can be an original transform block or a partitioned sub-block.

[0227] Here, log2ZoTbWidth and log2ZoTbHeight represent the maximum width and maximum height of the top-left region in a transform block where non-zero coefficients can exist, and the top-left region can be referred to as the "reduced transform block". That is, log2ZoTbWidth and log2ZoTbHeight can be defined as variables indicating the width and height of the reduced transform block.

[0228] That is, according to the example, when sub-block transform is applied to a coding unit, MTS needs to be applied so as to apply reduction to zero (RMTS), where the width of the top-left block where non-zero coefficients can exist is reduced to 16 and the remaining region is 0.

[0229] When MTS is not applied, even if it is indicated that SBT is applied to a coding unit (when cu_sbt_flag is 1), DCT-2 rather than MTS needs to be applied to the sub-block to which the transform is applied. For example, when sps_mts_enabled_flag is equal to 0, even if SBT is applied to a coding block, DCT-2 is applied instead of DST-7 or DCT-8.

[0230] In addition, when SBT is applied to a coding block, DCT-2 is applied only when the width or height of the sub-block is 64 or greater, and DCT-2 is not applied in other cases.

[0231] That is, in order to ensure that DCT-2 is applied to a sub-block and to prevent RMTS that can cause data loss from being applied to the sub-block to which DCT-2 is applied, the encoding device configures the image information to check the value of sps_mts_enabled_flag, and the decoding device checks the value of sps_mts_enabled_flag in residual coding according to the configured image information.

[0232] In summary, when only cu_sbt_flag is indicated as 1 during image decoding, the value of sps_mts_enabled_flag signaled in the sequence parameter set can be checked when setting the transform block size according to RMTS, so as to prevent RMTS from being applied to the sub-block to which DCT-2 is applied. When cu_sbt_flag is equal to 1 and sps_mts_enabled_flag is equal to 1, RMTS can be applied, so the width and height of the top-left block where non-zero transform coefficients can exist can be set to 16, and when cu_sbt_flag is equal to 1 and sps_mts_enabled_flag is equal to 0, MTS is not applied because DCT-2 is used for the transform, so reduction to zero according to RMTS is not applied either. In this case, the width and height of the top-left block where non-zero transform coefficients can exist can be set to 32 or less.

[0233] For example, in the case of a 32x32 transform block to which a sub-block transform is applied, when sps_mts_enabled_flag is equal to 0, DCT-2 is used as the transform kernel, so the width or height of length 32 is not reduced to 16.

[0234] As described above, when MTS is disabled, it is possible to prevent RMTS of length zeroed to 16 by checking sps_mts_enabled_flag.

[0235] In other cases, when cu_sbt_flag is not 1, the height of the transform block is greater than 32, the width of the transform block is not 32, or MTS is not applied, the width of the transform block can be set to the smaller value between the width of the transform block and 32. That is, the maximum width of the transform block can be limited to 32 by high-frequency zeroing. In addition, when cu_sbt_flag is not 1, the width of the transform block is greater than 32, the height of the transform block is not 32, or MTS is not applied, the height of the transform block can be set to the smaller value between the height of the transform block and 32. That is, the maximum height of the transform block can be limited to 32 by zeroing.

[0236] When SBT is applied, if the length of at least one side of the partitioned block is 64 or greater, DCT-2 can be applied in both the horizontal and vertical directions. Otherwise, DST-7 or DCT-8 can be applied as shown in Table 10. Therefore, when SBT is applied, zeroing is performed by applying RMTS32 only when all two sides of the partitioned block to which the transform is applied have lengths of 32 or less. That is, when the length of the block in each direction is 32, DST-7 or DCT-8 of length 32 can be applied, leaving only 16 transform coefficients.

[0237] As shown in Table 12, when RMTS32 is applied, the width and height of the remaining non-zeroed area (low-frequency transform coefficient area) can be considered as the width and height of the actual transform block instead of using the width and height of the original transform block used for compilation (log2ZoTbWidth = 4 or log2ZoTbHeight = 4) to perform compilation.

[0238] For example, in the case where the width x height of the original transform block is 32x16, when RMTS32 is applied, non-zero coefficients exist only in the upper left 16x16 area due to zeroing. Therefore, the width and height of the upper left area where non-zero transform coefficients can exist are set to 16 and 16 respectively, and then compilation of syntax elements (such as last_sig_coeff_x_prefix and last_sig_coeff_y_prefix) can be performed.

[0239] In summary, according to the residual compilation of Table 12, log2ZoTbWidth and log2ZoTbHeight, which indicate the maximum width and maximum height where non-zero coefficients can exist, are set before compiling last_sig_coeff_x_prefix. The position of the last non-zero coefficient is compiled by applying log2ZoTbWidth and log2ZoTbHeight, and then the width and height of the actual transform block are respectively changed to log2ZoTbWidth and log2ZoTbHeight (log2TbWidth = log2ZoTbWidth and log2TbHeight = log2ZoTbHeight). Subsequently, the syntax elements can be encoded according to the changed values. Therefore, the remaining top-left region excluding the zeroing area in the residual compilation can be configured as a new transform block, and then the residual samples can be derived.

[0240] When reducing the size of the transform block to the low-frequency transform coefficient region by zeroing the high-frequency transform coefficients, the values of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be restricted to the range from 0 to a value between (log2ZoTbWidth << 1) – 1 and (log2ZoTbHeight << 1) – 1 as shown in the semantics of Table 12.

[0241] [Table 13]

[0242]

[0243]

[0244]

[0245]

[0246]

[0247]

[0248]

[0249]

[0250] Table 13 illustrates the transform process when SBT is applied and MTS is implicitly applied to the compilation unit (cu_sbt_flag equals 1 and Max(nTbW,nTbH) is less than or equal to 32, and implicitMtsEnabled is set to equal 1).

[0251] The variables trTypeHor representing the horizontal transform kernel and trTypeVer representing the vertical transform kernel can be derived based on Table 8-14, and Table 8-14 in Table 13 can correspond to Table 1 of this document.

[0252] In addition, when SBT is applied to a compilation unit, the variables trTypeHor and trTypeVer can be derived based on Table 8-15, and Table 8-15 in Table 13 can correspond to Table 10 of this document.

[0253] The size of the block to which zeroing is applied, i.e., the zeroing block illustrated in Table 12, is expressed as nonZeroW and nonZeroH in Table 13. nonZeroW and nonZeroH can be defined as variables representing the width and height of the top-left block where non-zero transform coefficients can exist.

[0254] When LFNST is not applied, nonZeroW can be set to the smaller value between the value based on whether trTypeHor is greater than 0 ((trTypeHor > 0)? 16:32) and the width of the transform block (nTbW) (nonZeroW = Min(nTbW, (trTypeHor > 0)? 16:32)). When trTypeHor is greater than 0, since MTS is applied, "(trTypeHor > 0)? 16:32" is set to 16, and thus nonZeroW is set to the smaller value between the width of the transform block (nTbW) and 16. However, when trTypeHor is not greater than 0, since MTS is not applied, "(trTypeHor > 0)? 16:32" is set to 32, and thus nonZeroW is set to the smaller value between the width of the transform block (nTbW) and 32.

[0255] Similarly, when LFNST is not applied, nonZeroH can be set to the smaller value between the value based on whether trTypeVer is greater than 0 ((trTypeVer > 0)? 16:32) and the height of the transform block (nTbH) (nonZeroH = Min(nTbH, (trTypeVer > 0)? 16:32)). When trTypeVer is greater than 0, since MTS is applied, "(trTypeVer > 0)? 16:32" is set to 16, and thus nonZeroH is set to the smaller value between the height of the transform block (nTbH) and 16. However, when trTypeVer is not greater than 0, since MTS is not applied, "(trTypeVer > 0)? 16:32" is set to 32, and thus nonZeroH is set to the smaller value between the height of the transform block (nTbH) and 32.

[0256] That is, depending on whether MTS is applied, i.e., whether trTypeHor and trTypeVer can have a value of 0 or greater, the size of the zeroing block is set to 16 or 32.

[0257] Residual sample values can be derived based on nonZeroW and nonZeroH that take into account the zeroing setting (when nTbH is greater than 1, each (vertical) column of scaling transform coefficients d[x][y] with x = 0..nonZeroW-1, y = 0..nonZeroH-1 is transformed into e[x][y] with x = 0..nonZeroW-1, y = 0..nTbH-1 by calling the one-dimensional transform process specified in Clause 8.7.4.4 for each column x = 0..nonZeroW-1 with the height of the transform block nTbH, the non-zero height of the scaling transform coefficients nonZeroH, the list d[x][y] with y = 0..nonZeroH-1, and the transform type variable trType set equal to trTypeVer as inputs, and the output is the list e[x][y] with y = 0..nTbH-1).

[0258] When changing the size of the transform block by applying zeroing, the size of the transform block for context selection for last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can also be changed. Table 14 shows the binarization of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix considering the reduced transform block, and Table 15 shows the process of deriving ctxInc (context increment) for deriving last_sig_coeff_x_prefix and lastsig_coeff_y_prefix. Since contexts can be selected and distinguished by the context increment, the context model can be derived based on the context increment.

[0259] [Table 14]

[0260]

[0261]

[0262] [Table 15]

[0263]

[0264]

[0265] As illustrated in Table 14, the maximum value (cMax) of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix is set based on log2ZoTbWidth and log2ZoTbHeight corresponding to the width and height of the reduced transform block (cMax = (cMax = (log2ZoTbWidth << 1) – 1, cMax = (log2ZoTbHeight << 1) – 1)). When the truncated unary is used for the binarization of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix, the maximum value (cMax) of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix can be set to be equal to the maximum value of the codewords used for the binarization of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix. Therefore, the maximum length of the prefix codeword representing the last significant coefficient prefix information can be derived based on the size of the zeroed block.

[0266] As illustrated in Table 15, for the CABAC context of two syntax elements, namely last_sig_coeff_x_prefix and last_sig_coeff_y_prefix, the size of the original transform block (TU) in the low-frequency transform coefficient region rather than the reduced transform block is applied (log2TbSize is set to be equal to log2TbWidth, log2TbSize is set to be equal to log2TbHeight).

[0267] In summary, according to the example, the residual samples can be derived based on the last significant coefficient position information, where the context model can be derived based on the size of the original transform block whose size has not changed, and the last significant coefficient position can be derived based on the size of the transform block to which zeroing has been applied. Here, the size of the transform block to which zeroing has been applied, i.e., the zeroed block, is specifically the width or height that is less than the size of the original transform block, width or height.

[0268] The following drawings are provided to describe specific examples of the present disclosure. Since the specific names of the devices illustrated in the drawings or the names of specific signals / messages / fields are provided for illustration purposes, the technical features of the present disclosure are not limited to the specific names used in the following drawings.

[0269] Figure 7 is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.

[0270] Figure 7 Each operation illustrated inFigure 3 is performed by the decoding device 300 illustrated in. Specifically, S700 and S710 can be performed by Figure 3 the entropy decoder 310 illustrated in, S720 can be performed by Figure 3 the dequantizer 321 illustrated in, S730 can be performed by Figure 3 the inverse transformer 322 illustrated in, and S740 can be performed by Figure 3 the adder 340 illustrated in. The operations according to S700 to S740 are based on some of the foregoing details described with reference to Figures 4 to 6 Accordingly, descriptions of specific details that overlap with those described above with reference to Figures 3 to 6 will be omitted or will be briefly described.

[0271] According to an embodiment, the decoding device may receive a bitstream including residual information (S700). Specifically, the entropy decoder 310 of the decoding device may receive a bitstream including residual information.

[0272] According to an embodiment, the decoding device may derive quantization transform coefficients for a current block based on the residual information included in the bitstream (S710). Specifically, the entropy decoder 310 of the decoding device may quantize the transform coefficients for the current block based on the residual information included in the bitstream.

[0273] According to an embodiment, the decoding device may derive transform coefficients from the quantization transform coefficients based on a dequantization process (S720). Specifically, the dequantizer 321 of the decoding device may derive transform coefficients from the quantization transform coefficients based on the dequantization process.

[0274] According to an embodiment, the decoding device may derive residual samples for a current block by applying an inverse transform to the derived transform coefficients (S730). Specifically, the inverse transformer 322 of the decoding device may derive residual samples for the current block by applying an inverse transform to the derived transform coefficients.

[0275] According to an embodiment, the decoding device may generate a reconstructed picture based on the residual samples for the current block (S740). Specifically, the adder 340 of the decoding device may generate a reconstructed picture based on the residual samples for the current block.

[0276] In an embodiment, the unit of the current block may be a transform block (TB).

[0277] In an embodiment, each transform coefficient for the current block may be associated with a high-frequency transform coefficient region including transform coefficients of 0 or a low-frequency transform coefficient region including at least one valid transform coefficient.

[0278] In an embodiment, the residual information may include last significant coefficient prefix information and last significant coefficient suffix information regarding the position of the last significant transform coefficient among the transform coefficients for the current block.

[0279] In one example, the last significant coefficient prefix information may have a maximum value determined based on the size of the zeroed block.

[0280] In an embodiment, the position of the last significant transform coefficient may be determined based on a prefix codeword indicating the last significant coefficient prefix information and the last significant coefficient suffix information.

[0281] In an embodiment, the maximum length of the prefix codeword may be determined based on the low-frequency transform coefficient region, i.e., the size of the zeroed block.

[0282] In an embodiment, the size of the zeroed block may be determined based on the width and height of the current block.

[0283] In an embodiment, the last significant coefficient prefix information may include x-axis prefix information and y-axis prefix information, and the prefix codeword may be a codeword for the x-axis prefix information and a codeword for the y-axis prefix information.

[0284] In one example, the x-axis prefix information may be expressed as last_sig_coeff_x_prefix, the y-axis prefix information may be expressed as last_sig_coeff_y_prefix, and the position of the last significant transform coefficient may be expressed as (LastSignificantCoeffX, LastSignificantCoeffY).

[0285] In an embodiment, the residual information may include information regarding the size of the zeroed block.

[0286] Figure 8 FIG. is a flowchart illustrating a process for a video decoding device to derive transform coefficients according to an embodiment of the present disclosure.

[0287] Figure 8 Each operation illustrated in Figure 3 may be performed by the decoding device 300 illustrated in Figure 3 Specifically, S800 to S840 may be performed by the entropy decoder 310 illustrated in

[0288] First, as illustrated, a zeroed block for the current block may be derived (S800). As described above, the zeroed block refers to a low-frequency transform coefficient region including non-zero significant transform coefficients, and the width or height of the zeroed block may be derived based on whether the MTS using multiple transform kernels is applicable to the current block, whether sub-block transform is applied, and the width or height of the current block.

[0289] According to the example, when applying MTS, the decoding device can set the width or height of the zeroing block to 16, and when not applying MTS, the width or height of the zeroing block can be set to 32 or less. In this case, it can be determined whether MTS is applicable based on the sps_mts_enabled_flag indicating whether MTS is applicable.

[0290] Specifically, when applying DST-7 or DCT-8 instead of DCT-2 as the transform kernel for the inverse primary transform, the width of the current block is 32, and the height of the current block is 32 or less, the width of the zeroing block can be set to 16. When the above conditions are not met, that is, when the transform kernel is DCT-2, the width of the current block is not 32, or the height of the current block is 64 or greater, the width of the zeroing block can be set to the smaller value between the width of the current block and 32.

[0291] Similarly, when applying DST-7 or DCT-8 instead of DCT-2 as the transform kernel for the inverse primary transform, the height of the current block is 32, and the width of the current block is 32 or less, the height of the zeroing block can be set to 16. When the above conditions are not met, that is, when the transform kernel is DCT-2, the height of the current block is not 32, or the width of the current block is 64 or greater, the height of the zeroing block can be set to the smaller value between the height of the current block and 32.

[0292] In addition, according to the example, the width or height of the zeroing block can be derived based on the flag information (cu_sbt_flag) indicating whether the current block is partitioned into sub-blocks and transformed. For example, when the value of the flag indicating whether the current block is partitioned into sub-blocks and transformed is 1, the width of the partitioned sub-block is 32, and the height of the sub-block is less than 64, the width of the top-left region where non-zero transform coefficients can exist in the sub-block can be set to 16. Alternatively, when the value of the flag indicating whether the current block is partitioned into sub-blocks and transformed is 1, the height of the partitioned sub-block is 32, and the width of the sub-block is less than 64, the height of the top-left region where non-zero transform coefficients can exist in the sub-block can be set to 16.

[0293] The transform kernel can be derived based on the partition direction of the current block and the position of the sub-block to which the transform is applied as shown in Table 10.

[0294] The width or height of the zeroing block can be derived based on the MTS index of the current block or the flag information indicating whether MTS is applied to the transform of the current block.

[0295] The size of the zeroing block can be smaller than the size of the current block. Specifically, the width of the zeroing block can be smaller than the width of the current block, and the height of the zeroing block can be smaller than the height of the current block.

[0296] In an embodiment, the size of the zeroing block can be one of 32x16, 16x32, 16x16, or 32x32.

[0297] In an embodiment, the size of the current block can be 64x64, and the size of the zeroing block can be 32x32.

[0298] The decoding device can derive a context model for the last significant coefficient position information based on the width or height of the current block (S810).

[0299] According to an example, the context model can be derived based on the size of the original transform block rather than the size of the zeroing block. Specifically, context deltas for the x-axis prefix information and y-axis prefix information corresponding to the last significant coefficient prefix information can be derived based on the size of the original transform block.

[0300] The decoding device can derive the value of the last significant coefficient position information based on the derived context model (S820).

[0301] As described above, the last significant coefficient position information can include last significant coefficient prefix information and last significant coefficient suffix information, and the value of the last significant coefficient position can be derived based on the context model.

[0302] The decoding device can derive the last significant coefficient position based on the derived value of the last significant coefficient position information and the width or height of the zeroing block (S830).

[0303] In one example, the decoding device can derive the last significant coefficient position within a range of the size of the zeroing block that is smaller than the size of the current block rather than the original current block. That is, the transform coefficients to which the transform is applied can be derived within a range of the size of the zeroing block rather than the size of the current block.

[0304] In one example, the last significant coefficient prefix information can have a maximum value determined based on the size of the zeroing block.

[0305] In one example, the last significant coefficient position can be derived based on a prefix codeword indicating the last significant coefficient prefix information and the last significant coefficient suffix information, and the maximum length of the prefix codeword can be determined based on the size of the zeroing block.

[0306] The decoding device can derive the transform coefficients according to the last significant coefficient position derived based on the width or height of the zeroing block (S840).

[0307] The transform coefficients can be derived through the residual compilation process of Table 13.

[0308] Subsequently, the decoding device may derive residual samples by performing at least one of the foregoing inseparable inverse secondary transform and inverse primary transform based on Table 1 and Table 10.

[0309] The following drawings are provided to describe specific examples of the present disclosure. Since the specific names of the devices illustrated in the drawings or the names of specific signals / messages / fields are provided for illustration purposes, the technical features of the present disclosure are not limited to the specific names used in the following drawings.

[0310] Figure 9 is a flowchart illustrating the operation of a video coding device according to an embodiment of the present disclosure.

[0311] Figure 9 Each operation illustrated in Figure 2 can be performed by the encoding device 200 illustrated in Figure 2 Specifically, S900 can be performed by the subtractor 231 illustrated in Figure 2 S910 can be performed by the transformer 232 illustrated in Figure 2 S920 can be performed by the quantizer 233 illustrated in Figure 2 and S930 can be performed by the entropy encoder 240 illustrated in Figures 4 to 6 The operations according to S900 to S930 are based on some foregoing details described with reference to Figure 2 and Figures 4 to 6 Therefore, the description of specific details overlapping with those described above with reference to

[0312] The encoding device according to an embodiment may derive residual samples for a current block (S900). Specifically, the subtractor 231 of the encoding device may derive residual samples for the current block.

[0313] The encoding device according to an embodiment may transform the residual samples for the current block to derive transform coefficients for the current block (S910). Specifically, the transformer 232 of the encoding device may transform the residual samples of the current block to derive transform coefficients for the current block.

[0314] The encoding device according to an embodiment may derive quantized transform coefficients from the transform coefficients based on quantization (S920). Specifically, the quantizer 233 of the encoding device may derive quantized transform coefficients from the transform coefficients based on quantization.

[0315] The encoding device according to an embodiment may encode residual information including information about the quantized transform coefficients (S930). Specifically, the entropy encoder 240 of the encoding device may encode residual information including information about the quantized transform coefficients.

[0316] In an embodiment, each transform coefficient for a current block may be associated with a high-frequency transform coefficient region including transform coefficients of 0 or a zeroed block which is a low-frequency transform coefficient region including at least one valid transform coefficient.

[0317] In an embodiment, the residual information may include last significant coefficient prefix information and last significant coefficient suffix information regarding the position of the last valid transform coefficient among the transform coefficients for the current block.

[0318] In an embodiment, the position of the last valid transform coefficient may be determined based on a prefix codeword indicating the last significant coefficient prefix information and the last significant coefficient suffix information.

[0319] In one example, the last significant coefficient prefix information may have a maximum value determined based on the size of the zeroed block.

[0320] In an embodiment, the maximum length of the prefix codeword may be determined based on the size of the zeroed block.

[0321] In an embodiment, the size of the zeroed block may be determined based on the width and height of the current block.

[0322] In an embodiment, the last significant coefficient prefix information may include x-axis prefix information and y-axis prefix information, and the prefix codeword may be a codeword regarding the x-axis prefix information and a codeword for the y-axis prefix information.

[0323] In one example, the x-axis prefix information may be expressed as last_sig_coeff_x_prefix, the y-axis prefix information may be expressed as last_sig_coeff_y_prefix, and the position of the last valid transform coefficient may be expressed as (LastSignificantCoeffX,LastSignificantCoeffY).

[0324] In an embodiment, the residual information may include information regarding the size of the zeroed block.

[0325] Figure 10 FIG. is a flowchart illustrating a process for encoding transform coefficients and information according to an embodiment of the present disclosure.

[0326] Figure 10 Each operation disclosed in Figure 2 may be performed by the encoding device 200 disclosed in Figure 2 . Specifically, S1000 and S1010 may be performed by the transformer 232, and S1020 to S1040 may be performed by the

[0327] First, as shown in the figure, a zeroing block for the current block can be derived (S1000). As described above, the zeroing block refers to a low-frequency transform coefficient region including non-zero valid transform coefficients, and the width or height of the zeroing block can be derived based on whether the MTS using multiple transform kernels is applicable to the current block, whether sub-block transformation is applied, and the width or height of the current block.

[0328] According to an example, the encoding device can set the width or height of the zeroing block to 16 when applying MTS (when DST-7 / DCT-8 is applicable), and can set the width or height of the zeroing block to 32 or less when not applying MTS.

[0329] Specifically, when applying DST-7 or DCT-8 instead of DCT-2 as the transform kernel for primary transformation, the width of the current block is 32, and the height of the current block is 32 or less, the width of the zeroing block can be set to 16. When the above conditions are not met, that is, when the transform kernel is DCT-2, the width of the current block is not 32, or the height of the current block is 64 or greater, the width of the zeroing block can be set to the smaller value between the width of the current block and 32.

[0330] Similarly, when applying DST-7 or DCT-8 instead of DCT-2 as the transform kernel for primary transformation, the height of the current block is 32, and the width of the current block is 32 or less, the height of the zeroing block can be set to 16. When the above conditions are not met, that is, when the transform kernel is DCT-2, the height of the current block is not 32, or the width of the current block is 64 or greater, the height of the zeroing block can be set to the smaller value between the height of the current block and 32.

[0331] In addition, according to an example, the width or height of the zeroing block can be derived based on whether the current block is partitioned into sub-blocks and transformed. For example, when the current block is partitioned into sub-blocks and transformed, the width of the partitioned sub-block is 32, and the height of the sub-block is less than 64, the width of the sub-block can be set to 16. Alternatively, when the current block is partitioned into sub-blocks and transformed, the height of the partitioned sub-block is 32, and the width of the sub-block is less than 64, the height of the sub-block can be set to 16.

[0332] The transform kernel can be derived based on the partitioning direction of the current block and the position of the sub-block to which the transformation is applied as shown in Table 10.

[0333] In an embodiment, the size of the zeroing block can be one of 32x16, 16x32, 16x16, or 32x32.

[0334] In an embodiment, the size of the current block can be 64x64, and the size of the zeroing block can be 32x32.

[0335] The encoding device may derive transform coefficients based on a zeroing block (S1010).

[0336] The encoding device may derive transform coefficients from residual samples by performing at least one of the foregoing transformation processes, i.e., primary transformation and non-separable secondary inverse transformation, based on Table 1 and Table 10.

[0337] The encoding device may derive the last valid coefficient position based on the width or height of the derived zeroing block (S1020).

[0338] In one example, the encoding device may derive the last valid coefficient position within a range of the size of the zeroing block that is less than or equal to the size of the current block rather than the size of the original current block. That is, the transform coefficients to which the transformation is applied may be derived within the range of the size of the zeroing block rather than the size of the current block.

[0339] In one example, the last valid coefficient position may be derived based on a prefix codeword indicating last valid coefficient prefix information and last valid coefficient suffix information, and the maximum length of the prefix codeword may be determined based on the size of the zeroing block.

[0340] The encoding device may derive a context model for the last valid coefficient position information based on the width or height of the current block (S1030).

[0341] According to an embodiment, the context model may be derived based on the size of the original transform block rather than the size of the zeroing block. Specifically, context deltas of x-axis prefix information and y-axis prefix information corresponding to the last valid coefficient prefix information may be derived based on the size of the original transform block.

[0342] The encoding device may encode the position information of the value regarding the last valid coefficient position based on the derived context model (S1040).

[0343] As described above, the last valid coefficient position information may include last valid coefficient prefix information and last valid coefficient suffix information, and the value of the last valid coefficient position may be encoded based on the context model.

[0344] In the present disclosure, at least one of quantization / dequantization and / or transformation / inverse transformation may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transformation / inverse transformation is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for the sake of consistency in expression.

[0345] In addition, in the present disclosure, the quantized transform coefficients and the transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled through a residual coding syntax. The transform coefficients may be derived based on the residual information (or the information about the transform coefficients), and the scaled transform coefficients may be derived through an inverse transform (scaling) of the transform coefficients. The residual samples may be derived based on an inverse transform (transformation) of the scaled transform coefficients. These details may also be applied / expressed in other parts of the present disclosure.

[0346] In the above embodiments, the method is illustrated based on a flowchart by means of a series of steps or blocks, but the present disclosure is not limited to the order of the steps, and a certain step may be executed in an order or steps different from the above order or steps or simultaneously with another step. In addition, those of ordinary skill in the art can understand that the steps shown in the flowchart are not exclusive, and one or more steps of the flowchart may be incorporated or removed without affecting the scope of the present disclosure.

[0347] The above method according to the present disclosure may be implemented in software form, and the encoding device and / or decoding device according to the present disclosure may be included in an image processing device such as a TV, a computer, a smart phone, a set-top box, a display device, etc.

[0348] When the embodiments in the present disclosure are specifically implemented by software, the above method may be specifically embodied as a module (procedure, function, etc.) for performing the above functions. The module may be stored in a memory and may be executed by a processor. The memory may be inside or outside the processor and may be connected to the processor in various well-known ways. The processor may include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory may include a read only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in the present disclosure may be specifically implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing may be specifically implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip.

[0349] In addition, the decoding device and encoding device applying the present disclosure may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camera, a video-on-demand (VoD) service providing device, an over-the-top (OTT) video device, an Internet streaming service providing device, a three-dimensional (3D) video device, a video phone video device, and a medical video device, and may be used to process video signals or data signals. For example, an over-the-top (OTT) video device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smart phone, a tablet PC, a digital video recorder (DVR), etc.

[0350] In addition, the processing method applying the present disclosure may be generated in the form of a program executed by a computer and stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium embodied in the form of a carrier wave (e.g., transmission through the Internet). Additionally, a bitstream generated by an encoding method may be stored in a computer-readable recording medium or transmitted through a wired or wireless communication network. Additionally, embodiments of the present disclosure may be embodied as a computer program product by program code, and the program code may be executed on a computer by embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.

[0351] Figure 11 Illustrated is the structure of a content streaming system applying the present disclosure.

[0352] In addition, the content streaming system applying the present disclosure may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0353] The encoding server compresses the content input from multimedia input devices such as smart phones, cameras, video cameras, etc. into digital data to generate a bitstream, and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, camera, video camera, etc. directly generates a bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method of the embodiments of this document. And the streaming server can temporarily store the bitstream in the process of sending or receiving the bitstream.

[0354] The streaming server sends the multimedia data to the user device through the network server based on the user's request, and the network server serves as a medium for informing the user of the service. When the user requests a desired service from the network server, the network server transmits it to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between the devices in the content streaming system.

[0355] The streaming server can receive the content from the media memory and / or the encoding server. For example, when receiving the content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined time.

[0356] For example, the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a watch-type terminal (smart watch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, a digital signage, etc. Each of the servers in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.

[0357] The claims disclosed in this document can be combined in various ways. For example, the technical features of the method claims of this disclosure can be combined to be implemented or executed in a device, and the technical features of the device claims can be combined to be implemented or executed in a method. In addition, the technical features that can combine the method claims and the device claims can be combined to be implemented or executed in a device, and the technical features that can combine the method claims and the device claims can be combined to be implemented or executed in a method.

Claims

1. An image decoding method performed by a decoding device, comprising: Receiving a bitstream including residual information; Deriving transform coefficients for a current block based on the residual information; And Deriving residual samples for the current block based on the transform coefficients, Wherein the derivation of the transform coefficients includes deriving a zeroing block related to a region where valid transform coefficients may exist in the current block, Wherein the zeroing block is derived based on a first flag, the first flag indicating whether a second flag and a third flag exist in a sequence parameter set SPS of the bitstream, and Wherein the second flag indicates whether multi-transform selection MTS index information can exist for an intra-coded unit, and the third flag indicates whether the MTS index information can exist for an inter-coded unit.

2. The image decoding method according to claim 1, wherein, Based on the first flag indicating that the second flag and the third flag exist in the SPS of the bitstream, the width or height of the zeroing block is set to 16, and Wherein, based on the first flag indicating that the second flag and the third flag do not exist in the SPS of the bitstream, the width or height of the zeroing block is set to be less than or equal to 32.

3. The image decoding method according to claim 1, wherein, Based on a fourth flag related to whether sub-block transform applied to sub-blocks derived by partitioning a coded unit is applied to the current block being equal to 1, the width or height of the zeroing block is set to 16.

4. The image decoding method according to claim 3, wherein, Based on the height of the sub-block being less than 64 and the width of the sub-block being 32, the width of the zeroing block is set to 16.

5. The image decoding method according to claim 3, wherein, Based on the width of the sub-block being less than 64 and the height of the sub-block being 32, the height of the zeroing block is set to 16.

6. The image decoding method according to claim 3, wherein, Deriving a transform kernel for the current block based on the partitioning direction of the current block and the position of the sub-block to which the transform is applied.

7. The image decoding method according to claim 1, wherein, The residual information includes last significant coefficient prefix information, and Wherein, the maximum value of the last significant coefficient prefix information is derived based on the size of the zeroing block.

8. The image decoding method according to claim 1, wherein, Deriving the zeroing block for the luminance component of the current block.

9. An image encoding method performed by an image encoding device, comprising: Deriving residual samples for a current block; Deriving transform coefficients based on the residual samples for the current block; And Encoding residual information including information about the transform coefficients, Wherein the derivation of the transform coefficients includes deriving a zeroing block related to a region where valid transform coefficients may exist in the current block, Wherein the zeroing block is derived based on a first flag, the first flag indicating whether a second flag and a third flag exist in a sequence parameter set SPS of the bitstream, and Wherein the second flag indicates whether multi-transform selection MTS index information can exist for an intra-coded unit, and the third flag indicates whether the MTS index information can exist for an inter-coded unit.

10. The image encoding method according to claim 9, wherein, Based on the first flag indicating that the second flag and the third flag exist in the SPS of the bitstream, the width or height of the zeroing block is set to 16, and Among them, based on the first flag indicating that the second flag and the third flag do not exist in the SPS of the bitstream, the width or height of the zeroing block is set to be less than or equal to 32.

11. The image encoding method according to claim 9, wherein, Based on the fourth flag related to whether the sub-block transform for performing a transform on a sub-block derived by partitioning a coding unit is applied to the current block being equal to 1, the width or height of the zeroing block is set to 16.

12. The image encoding method according to claim 11, wherein, Based on the height of the sub-block being less than 64 and the width of the sub-block being 32, the width of the zeroing block is set to 16, and Among them, based on the width of the sub-block being less than 64 and the height of the sub-block being 32, the height of the zeroing block is set to 16.

13. The image encoding method according to claim 11, wherein, Determine the transform kernel for the current block based on the partitioning direction of the current block and the position of the sub-block to which the transform is applied.

14. The image encoding method according to claim 9, wherein, The residual information includes last significant coefficient prefix information, and Among them, the maximum value of the last significant coefficient prefix information is derived based on the size of the zeroing block.

15. A method for transmitting data of an image, comprising: Obtain a bitstream for an image, where the bitstream is generated by the following steps: derive residual samples for a current block, derive transform coefficients based on the residual samples for the current block, and encode residual information including information about the transform coefficients to generate the bitstream; and Transmit data including the bitstream, Among them, the derivation of the transform coefficients includes deriving a zeroing block related to a region where valid transform coefficients may exist in the current block, Among them, the zeroing block is derived based on a first flag, the first flag indicating whether a second flag and a third flag exist in the sequence parameter set SPS of the bitstream, and Among them, the second flag indicates whether multi-transform selection MTS index information can exist for an intra coding unit, and the third flag indicates whether the MTS index information can exist for an inter coding unit.