Video signal processing method and device therefor

The video signal processing method addresses the challenge of low coding efficiency by dynamically adjusting dequantized transform coefficients, resulting in improved compression and decoding performance.

WO2025127632A1PCT designated stage expired Publication Date: 2025-06-19WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/020005
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-24
Filing Date
2024-12-06
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing video signal processing methods struggle with achieving high coding efficiency due to limitations in removing redundant information, particularly in handling spatial, temporal, and probabilistic correlations in video signals.

Method used

A video signal processing method and device that includes a processor for decoding video signals by obtaining a quantization index, dequantizing transform coefficients, and adjusting them based on specific conditions such as prediction modes and block characteristics, to improve coding efficiency.

Benefits of technology

The proposed method enhances coding efficiency by dynamically adjusting dequantized transform coefficients, leading to improved compression and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024020005_19062025_PF_FP_ABST
    Figure KR2024020005_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A video signal decoding device is disclosed. The video signal decoding device comprises a processor. The processor: obtains a quantization index for a current block from a bitstream; obtains an inversely quantized transform coefficient by multiplying the quantization index by a scale factor; determines whether to adjust the inversely quantized transform coefficient; in the case of adjusting the inversely quantized transform coefficient, adjusts the inversely quantized transform coefficient, and generates an error block by inversely transforming a residual signal of the current block by using the adjusted inversely quantized transform coefficient; and in the case of not adjusting the inversely quantized transform coefficient, generates an error block by inversely transforming the residual signal of the current block by using the inversely quantized transform coefficient.
Need to check novelty before this filing date? Find Prior Art

Description

Video signal processing method and device therefor

[0001] The present invention relates to a method and device for processing a video signal, and more particularly, to a method and device for processing a video signal for encoding or decoding a video signal.

[0002] Compression coding refers to a series of signal processing techniques for transmitting digitized information over communication lines or storing it in a format suitable for storage media. Objects subject to compression coding include audio, video, and text, and the technology that performs compression coding on video in particular is called video compression. Compression coding of video signals is achieved by removing redundant information by considering spatial, temporal, and probabilistic correlations. However, with the recent advancements in various media and data transmission media, there is a growing demand for more efficient video signal processing methods and devices.

[0003] The purpose of this specification is to provide a video signal processing method and a device therefor to improve the coding efficiency of a video signal.

[0004] The present specification provides a video signal processing method and a device therefor. A video signal decoding device according to an embodiment of the present invention includes a processor. The processor obtains a quantization index for a current block from a bitstream, obtains a dequantized transform coefficient by multiplying the quantization index by a scale factor, determines whether to adjust the dequantized transform coefficient, adjusts the dequantized transform coefficient when the dequantized transform coefficient is adjusted, and inversely transforms a residual signal of the current block using the adjusted dequantized transform coefficient to generate an error block, and inversely transforms a residual signal of the current block using the inverse quantized transform coefficient when the dequantized transform coefficient is not adjusted, to generate an error block.

[0005] The processor may obtain a quantization index and a moving quantization index that moves the value of the quantization index by a predetermined offset from 0, obtain a first inverse quantized transform coefficient by multiplying the quantization index by the scale factor, obtain a second inverse quantized transform coefficient by multiplying the moving quantization index by the scale factor, and obtain the adjusted inverse quantized transform coefficient by weighting the first inverse quantized transform coefficient and the second inverse quantized transform coefficient.

[0006] The processor can determine the weight of the weighted average based on whether the current block is predicted in an intra prediction mode.

[0007] The above processor can determine whether to adjust the dequantized transform coefficients on a transform block basis.

[0008] The above processor can determine whether to adjust the dequantized transform coefficients of the current transform block based on whether the sum of the absolute values ​​of the non-zero coefficients of the current transform block is greater than a predetermined value.

[0009] The above processor can determine whether to adjust the dequantized transform coefficient based on the quantizer applied to the quantization index.

[0010] The processor may determine whether to adjust the value of the dequantized transform coefficient based on state information used in the transition procedure for the quantized index or updated state change.

[0011] The processor may obtain a value of the adjusted dequantized transform coefficient by adding a predefined offset to the dequantized transform coefficient according to a quantization scale factor and an absolute value of the quantization index. The quantization scale factor may be a coefficient multiplied by the quantization index to obtain the dequantized transform coefficient.

[0012] The value of the above offset can be determined based on whether the current block is predicted in intra prediction mode.

[0013] The processor can parse information signaling whether value adjustment of a dequantized transform coefficient is activated in a Sequence Parameter Set (SPS) included in the bitstream, and determine whether to adjust the value of the dequantized transform coefficient based on the information signaling whether value adjustment of the dequantized transform coefficient is activated.

[0014] The above processor can perform the inverse transformation using a transformation type derived from a transformation type list.

[0015] The above list of conversion types can be reordered based on template cost.

[0016] The above processor can configure the transformation type list based on whether the size of the luminance block of the current block is a pre-specified size.

[0017] The above processor can configure the above transformation type list as MTS, NSPT, DCT2, IDTR, and Transform SKIP.

[0018] A video signal encoding device according to an embodiment of the present invention includes a processor.

[0019] The processor obtains a dequantized transform coefficient by multiplying a quantization index by a scale factor, determines whether to adjust the dequantized transform coefficient, adjusts the dequantized transform coefficient if the dequantized transform coefficient is adjusted, inversely transforms a residual signal of the current block using the adjusted dequantized transform coefficient to generate an error block, and inversely transforms a residual signal of the current block using the inverse quantized transform coefficient if the dequantized transform coefficient is not adjusted, and includes a quantization index for the current block in a bitstream.

[0020] The processor obtains a shifted quantization index by shifting the quantization index and the value of the quantization index from 0 by a predetermined offset, and obtains a first inverse quantized transform coefficient by multiplying the quantization index by the scale factor.

[0021] The second dequantized transform coefficient can be obtained by multiplying the above-described moving quantization index by the above-described scale factor, and the adjusted dequantized transform coefficient can be obtained by weighting the first dequantized transform coefficient and the second dequantized transform coefficient.

[0022] The processor can determine the weight of the weighted average based on whether the current block is predicted in an intra prediction mode.

[0023] The above processor can determine whether to adjust the dequantized transform coefficients on a transform block basis.

[0024] The above processor can determine whether to adjust the dequantized transform coefficients of the current transform block based on whether the sum of the absolute values ​​of the non-zero coefficients of the current transform block is greater than a predetermined value.

[0025] In a computer-readable non-transitory storage medium storing a bitstream according to an embodiment of the present invention, the bitstream is decoded by a decoding method. The decoding method includes: a step of obtaining a dequantized transform coefficient by multiplying the quantization index by a scale factor; a step of determining whether to adjust the dequantized transform coefficient; a step of adjusting the dequantized transform coefficient when adjusting the dequantized transform coefficient, and generating an error block by inversely transforming a residual signal of the current block using the adjusted dequantized transform coefficient; and a step of not adjusting the dequantized transform coefficient when inversely transforming a residual signal of the current block using the dequantized transform coefficient to generate an error block.

[0026] This specification provides a method for efficiently processing a video signal.

[0027] The effects that can be obtained from this specification are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by those skilled in the art to which the present invention pertains from the description below.

[0028] FIG. 1 is a schematic block diagram of a video signal encoding device according to one embodiment of the present specification.

[0029] FIG. 2 is a schematic block diagram of a video signal decoding device according to one embodiment of the present specification.

[0030] Figure 3 illustrates an embodiment in which a coding tree unit within a picture is divided into coding units.

[0031] Figure 4 illustrates one embodiment of a method for signaling the splitting of a quad tree and a multi-type tree.

[0032] Figures 5 and 6 illustrate the intra prediction method according to an embodiment of the present invention in more detail.

[0033] Figure 7 is a diagram showing the locations of surrounding blocks used to construct a motion candidate list in inter prediction.

[0034] FIG. 8 illustrates a method for determining a reference pixel line based on a template according to one embodiment of the present specification.

[0035] FIG. 9 is a diagram illustrating a block vector related to an IBC encoding method according to one embodiment of the present specification.

[0036] Figure 10 illustrates a method for predicting the current block using RRIBC in the horizontal direction.

[0037] Figure 11 illustrates a method for predicting the current block using RRIBC in the vertical direction.

[0038] FIG. 12 illustrates a block vector of a block encoded in Intra TMP mode according to one embodiment of the present specification.

[0039] FIG. 13 illustrates a case where a current block is divided by GPM mode according to one embodiment of the present specification and the divided area is encoded by IBC mode.

[0040] FIG. 14 illustrates how a current block is encoded in IBC-CIIP mode according to one embodiment of the present specification.

[0041] FIG. 15 illustrates an example of a reference region and filter shape used to derive CCCM parameters according to one embodiment of the present specification.

[0042] FIG. 16 illustrates types of transform kernels that can be used in video coding according to one embodiment of the present specification.

[0043] FIG. 17 illustrates a transformation set table for LFNST and NSPT transformations according to one embodiment of the present specification.

[0044] Figure 18 shows an example of ROI after LFNST transformation.

[0045] FIG. 19 illustrates a method for deriving a multi-transform set and a LFNST / NSPT set according to one embodiment of the present specification.

[0046] Figure 20 illustrates a mapping table according to one embodiment of the present specification.

[0047] Figure 21 illustrates a conversion type set table according to one embodiment of the present specification.

[0048] Figure 22 illustrates a conversion type combination table according to one embodiment of the present specification.

[0049] FIG. 23 illustrates a threshold value table for an IDT conversion type according to one embodiment of the present specification.

[0050] FIG. 24 illustrates block boundaries and samples around the boundaries in a deblocking filtering process according to one embodiment of the present specification.

[0051] Figure 25 shows two independent scalar quantizers according to an embodiment of the present invention.

[0052] Figure 26 shows a state transition table according to an embodiment of the present invention.

[0053] Figure 27 shows a state transition graph according to an embodiment of the present invention.

[0054] Figure 28 illustrates a process of restoring a conversion coefficient according to an embodiment of the present invention.

[0055] Figure 29 shows an optimal trellis path according to one embodiment of the present invention.

[0056] Fig. 30 shows a restoration order of quantized transform coefficients according to an embodiment of the present invention.

[0057] FIG. 31 illustrates a method for updating the state of syntax elements and quantized transform coefficients according to one embodiment of the present invention.

[0058] Figure 32 shows a state transition graph according to one embodiment of the present invention.

[0059] Figure 33 shows a state transition graph according to one embodiment of the present invention.

[0060] Figure 34 shows a state transition graph according to one embodiment of the present invention.

[0061] Figure 35 illustrates a separate transition procedure according to an embodiment of the present invention.

[0062] FIG. 36 illustrates a method for setting the number of states according to a temporal layer according to an embodiment of the present invention.

[0063] FIG. 37 shows a quantization index for a block of size 8 x 8 according to an embodiment of the present invention.

[0064] Figure 38 shows a method for changing the number of states according to an embodiment of the present specification.

[0065] Figure 39 shows a method for changing the number of states according to an embodiment of the present specification.

[0066] Figure 40 shows a scan order for performing dependent quantization according to an embodiment of the present invention.

[0067] Figure 41 shows an example of changing the inverse quantized transform coefficients.

[0068] Figure 42 shows an example of an offset added to the value of the inverse quantized transform coefficient.

[0069] FIG. 43 shows signaling to SPS whether to enable value adjustment of inverse quantized transform coefficients according to an embodiment of the present invention.

[0070] FIG. 44 shows a constraint flag related to adjusting the value of a dequantized transform coefficient in the general_constraint_info() syntax structure according to an embodiment of the present invention.

[0071] Figure 45 shows a flowchart for parsing the transformation type of the current block according to an embodiment of the present invention.

[0072] FIG. 46 shows a flowchart for parsing the transformation type of the current block according to another embodiment of the present invention.

[0073] FIG. 47 shows a flowchart for parsing the transformation type of the current block according to another embodiment of the present invention.

[0074] FIG. 48 shows a method for calculating a template cost for each transformation type candidate in a transformation type list according to an embodiment of the present invention.

[0075] Figure 49 shows the coordinates of a luminance block corresponding to the coordinates of a chrominance block according to an embodiment of the present invention.

[0076] The terms used in this specification have been selected from widely used and current terms, taking into account the functions of the present invention. However, these terms may vary depending on the intentions of those skilled in the art, customs, or the emergence of new technologies. Furthermore, in certain cases, the applicant may arbitrarily select terms, in which case their meanings will be described in the description of the relevant invention. Therefore, it should be noted that the terms used in this specification should be interpreted based on their substantive meaning and the overall content of this specification, rather than simply their names.

[0077] In this specification, 'A and / or B' may be interpreted to mean 'comprising at least one of A or B'.

[0078] In this specification, some terms may be interpreted as follows. Coding may be interpreted as encoding or decoding, depending on the case. In this specification, a device that encodes a video signal to generate a video signal bitstream is referred to as an encoding device or encoder, and a device that decodes a video signal bitstream to restore a video signal is referred to as a decoding device or decoder. In addition, in this specification, a video signal processing device is used as a term that includes both an encoder and a decoder. Information is a term that includes values, parameters, coefficients, elements, etc., and since the meaning may be interpreted differently depending on the case, the present invention is not limited thereto. 'Unit' is used to mean a basic unit of image processing or a specific location of a picture, and refers to an image area that includes at least one of a luminance component and a chroma component. In addition, 'block' refers to an image area including specific components among luminance components and chrominance components (i.e., Cb and Cr). However, depending on the embodiment, terms such as 'unit', 'block', 'partition', 'signal', and 'region' may be used interchangeably. In addition, in this specification, 'current block' means a block that is currently scheduled to be encoded, and 'reference block' means a block that has already been encoded or decoded and is used as a reference in the current block. In addition, in this specification, terms such as 'luma', 'luminance', and 'Y' may be used interchangeably. In addition, in this specification, terms such as 'chroma', 'chroma', 'color difference', and 'Cb or Cr' may be used interchangeably, and since chrominance components are divided into two, Cb and Cr, each chrominance component may be used separately. In addition, in this specification, a unit may be used as a concept including all of a coding unit, a prediction unit, and a transformation unit.A picture refers to a field or a frame, and depending on the embodiment, the terms may be used interchangeably. Specifically, if the captured image is an interlace image, one frame is divided into an odd (or odd, top) field and an even (or even, bottom) field, and each field is configured as one picture unit and can be encoded or decoded. If the captured image is a progressive image, one frame is configured as a picture and can be encoded or decoded. In addition, in this specification, the terms 'error signal', 'residual signal', 'residual signal', 'residual signal', and 'differential signal' may be used interchangeably. In addition, in this specification, the terms 'intra prediction mode', 'intra prediction directional mode', 'intra-screen prediction mode', and 'intra-screen prediction directional mode' may be used interchangeably. In addition, in this specification, the terms 'motion', 'movement', and the like may be used interchangeably. In addition, in this specification, 'left', 'upper left', 'upper left', 'upper right', 'right', 'lower right', 'lower left', and 'lower left' can be used interchangeably with 'left', 'upper left', 'top', 'upper right', 'right', 'lower right', 'bottom', and 'lower left'. In addition, element and member can be used interchangeably with each other. POC (Picture Order Count) indicates the temporal position information of a picture (or frame), can be the playback order displayed on the screen, and can have a unique POC for each picture.

[0079] FIG. 1 is a schematic block diagram of a video signal encoding device (100) according to one embodiment of the present specification. Referring to FIG. 1, the encoding device (100) of the present invention includes a transform unit (110), a quantization unit (115), an inverse quantization unit (120), an inverse transform unit (125), a filtering unit (130), a prediction unit (150), and an entropy coding unit (160).

[0080] The transform unit (110) obtains a transform coefficient value by transforming the residual signal, which is the difference between the input video signal and the prediction signal generated by the prediction unit (150). For example, a discrete cosine transform (DCT), a discrete sine transform (DST), or a wavelet transform may be used. The discrete cosine transform and the discrete sine transform divide the input picture signal into blocks and perform the transformation. The coding efficiency may vary depending on the distribution and characteristics of the values ​​within the transform domain during the transformation. The transform kernel used for the transformation on the residual block may be a transform kernel having separable characteristics of vertical transformation and horizontal transformation. In this case, the transformation on the residual block may be performed separately as vertical transformation and horizontal transformation. For example, the encoder may perform vertical transformation by applying the transform kernel in the vertical direction of the residual block. Additionally, the encoder can perform horizontal transformation by applying a transformation kernel in the horizontal direction of the residual block. In the present disclosure, the transformation kernel may be used as a term referring to a set of parameters used for transformation of the residual signal, such as a transformation matrix, a transformation array, a transformation function, or a transformation. For example, the transformation kernel may be any one of a plurality of available kernels. Additionally, transformation kernels based on different transformation types may be used for each of the vertical transformation and the horizontal transformation.

[0081] The transformation coefficients are distributed in a way that increases toward the upper left corner of the block, and decreases toward 0 toward the lower right corner. As the current block size increases, there is a high probability of a significant number of 0 coefficients in the lower right corner. To reduce the transformation complexity of large blocks, the remaining regions can be reset to 0, leaving only the upper left corner as an arbitrary region.

[0082] Additionally, error signals may exist only in some regions of a coding block. In this case, the conversion process may be performed only on some arbitrary regions. For example, in a block of size 2Nx2N, an error signal may exist only in the first 2NxN block, and the conversion process may be performed only on the first 2NxN block, but the conversion process may not be performed on the second 2NxN block and may not be encoded or decoded. Here, N can be any positive integer.

[0083] The encoder may perform an additional transform before the transform coefficients are quantized. The aforementioned transform method may be referred to as a primary transform, and the additional transform may be referred to as a secondary transform. The secondary transform may be optional for each residual block. In one embodiment, the encoder may improve coding efficiency by performing the secondary transform on areas where it is difficult to concentrate energy in the low-frequency region using only the primary transform. For example, the secondary transform may be additionally performed on blocks where residual values ​​appear large in directions other than the horizontal or vertical direction of the residual block. Unlike the primary transform, the secondary transform may not be performed separately into a vertical transform and a horizontal transform. Such a secondary transform may be referred to as a Low Frequency Non-Separable Transform (LFNST).

[0084] The quantization unit (115) quantizes the transformation coefficient value output from the transformation unit (110).

[0085] In order to increase coding efficiency, rather than coding the picture signal as it is, a method is used to predict a picture using an already coded area through a prediction unit (150), and to obtain a restored picture by adding the residual value between the original picture and the predicted picture to the predicted picture. In order to prevent mismatches from occurring in the encoder and decoder, when the encoder performs prediction, information that is also available to the decoder must be used. To this end, the encoder performs a process of restoring the encoded current block. The inverse quantization unit (120) inversely quantizes the transform coefficient value, and the inverse transform unit (125) restores the residual value using the inverse quantized transform coefficient value. Meanwhile, the filtering unit (130) performs a filtering operation to improve the quality of the restored picture and enhance the coding efficiency. For example, a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter may be included. The filtered picture is stored in the Decoded Picture Buffer (DPB, 156) to be output or used as a reference picture.

[0086] A deblocking filter is a filter that removes distortion within blocks generated at the boundaries between blocks in a reconstructed picture. The encoder can determine whether to apply a deblocking filter to a certain boundary based on the distribution of pixels in several columns or rows based on an arbitrary boundary (edge) within the block. When applying a deblocking filter to a block, the encoder can apply a long filter, a strong filter, or a weak filter depending on the deblocking filtering strength. Additionally, horizontal and vertical filtering can be processed in parallel. Sample adaptive offset (SAO) can be used to correct the offset from the original image on a pixel-by-pixel basis for the residual block to which the deblocking filter has been applied. To correct the offset for a specific picture, the encoder can divide the pixels contained in the image into a certain number of regions, determine the regions to perform offset correction, and apply the offset to those regions (Band Offset). Alternatively, the encoder can use a method (Edge Offset) that applies an offset by considering the edge information of each pixel. An adaptive loop filter (ALF) is a method that divides pixels included in an image into a predetermined group, determines a filter to be applied to each group, and performs filtering differently for each group. Information regarding whether to apply ALF can be signaled on a coding unit basis, and the shape and filter coefficients of the ALF filter to be applied can vary depending on each block. In addition, an ALF filter of the same shape (fixed shape) can be applied regardless of the characteristics of the target block to which it is applied.

[0087] The prediction unit (150) includes an intra prediction unit (152) and an inter prediction unit (154). The intra prediction unit (152) performs intra prediction within a current picture, and the inter prediction unit (154) performs inter prediction to predict the current picture using a reference picture stored in a decoded picture buffer (156). The intra prediction unit (152) performs intra prediction from reconstructed regions within the current picture and transfers intra encoding information to the entropy coding unit (160). The intra encoding information may include at least one of an intra prediction mode, an MPM (Most Probable Mode) flag, an MPM index, and information about a reference sample. The inter prediction unit (154) may be configured to include a motion estimation unit (154a) and a motion compensation unit (154b). The motion estimation unit (154a) refers to a specific area of ​​the restored reference picture to find the part most similar to the current area and obtains a motion vector value, which is the distance between the areas. The motion information (reference direction indication information (L0 prediction, L1 prediction, bidirectional prediction), reference picture index, motion vector information, etc.) for the reference area obtained by the motion estimation unit (154a) is transferred to the entropy coding unit (160) so that it can be included in the bitstream. Using the motion information transferred from the motion estimation unit (154a), the motion compensation unit (154b) performs inter-motion compensation to generate a prediction block for the current block. The inter-prediction unit (154) transfers inter-encoding information including motion information for the reference area to the entropy coding unit (160).

[0088] According to an additional embodiment, the prediction unit (150) may include an intra block copy (IBC) prediction unit (not shown). The IBC prediction unit performs IBC prediction on reconstructed samples in the current picture and transfers IBC encoding information to the entropy coding unit (160). The IBC prediction unit obtains a block vector value indicating a reference region used for prediction of the current region by referring to a specific region in the current picture. The IBC prediction unit may perform IBC prediction using the obtained block vector value. The IBC prediction unit transfers the IBC encoding information to the entropy coding unit (160). The IBC encoding information may include at least one of size information of the reference region, block vector information (index information for block vector prediction of the current block within a motion candidate list, and block vector difference information).

[0089] When the picture prediction as above is performed, the transformation unit (110) obtains a transformation coefficient value by transforming the residual value between the original picture and the predicted picture. At this time, the transformation can be performed in units of specific blocks within the picture, and the size of the specific block can be varied within a preset range. The quantization unit (115) quantizes the transformation coefficient value generated by the transformation unit (110) and transfers the quantized transformation coefficient to the entropy coding unit (160).

[0090] The quantized transform coefficients in the form of a two-dimensional array can be rearranged into a one-dimensional array for entropy coding. The method of scanning the quantized transform coefficients can be determined by which scanning method is used depending on the size of the transform block and the prediction mode within the screen. For example, diagonal, vertical, and horizontal scanning can be applied. This scanning information can be signaled on a block-by-block basis and can be derived according to predetermined rules.

[0091] The entropy coding unit (160) entropy-codes information representing quantized transform coefficients, intra-coding information, inter-coding information, etc. to generate a video signal bitstream. The entropy coding unit (160) may use a variable length coding (VLC) method and an arithmetic coding method. The variable length coding (VLC) method converts input symbols into continuous codewords, and the length of the codewords may be variable. For example, frequently occurring symbols are expressed as short codewords, and infrequently occurring symbols are expressed as long codewords. A context-based adaptive variable length coding (CAVLC) method may be used as a variable length coding method. Arithmetic coding converts continuous data symbols into a single prime number by using the probability distribution of each data symbol, and arithmetic coding can obtain the optimal prime number bits required to express each symbol. Context-based Adaptive Binary Arithmetic Coding (CABAC) can be used as an arithmetic coding.

[0092] CABAC is a binary arithmetic coding method that uses multiple context models generated based on experimentally obtained probabilities. The context models can also be referred to as context models. First, if the symbols are not in binary form, the encoder binarizes each symbol using exp-Golomb, etc. The binarized 0 or 1 can be described as a bin. The CABAC initialization process is divided into context initialization and arithmetic coding initialization. Context initialization is the process of initializing the occurrence probability of each symbol, and is determined by the symbol type, quantization parameter (QP), and slice type (I, P, B). A context model with this initialization information can use probability-based values ​​obtained through experiments. The context model provides the occurrence probability of the Least Probable Symbol (LPS) or Most Probable Symbol (MPS) for the symbol to be currently encoded, as well as information (valMPS) on which bin value corresponds to the MPS between 0 and 1. One of several context models is selected through the context index (ctxIdx), and the context index can be derived from information about the current block to be encoded or information about the surrounding blocks. Initialization for binary arithmetic coding is performed based on the probability model selected from the context model. Binary arithmetic coding is performed by dividing the data into probability intervals based on the occurrence probabilities of 0 and 1, and then encoding is performed through a process in which the probability interval corresponding to the bin to be processed becomes the entire probability interval for the bin to be processed next. The location information within the probability interval in which the last bin has been processed is output. However, since the probability interval cannot be divided infinitely, if it is reduced to a certain size, a renormalization process is performed to expand the probability interval and output the corresponding location information. In addition, after each bin is processed, a probability update process can be performed in which the probability for the next bin to be processed is newly set based on the information of the processed bin.

[0093] A bitstream may consist of one or more coded video sequences (CVSs), and a CVS may be encoded independently of other CVSs. Each CVS may consist of one or more layers, and each layer may represent a specific quality level, a specific resolution, or a general image, a depth map, or a transparency map. Furthermore, a CLVS may mean a layer-wise CVS composed of consecutive (in decoding order) PUs within the same layer. For example, there may be a CLVS for a specific quality layer, and there may be a CLVS for a depth map.

[0094] The above generated bitstream is encapsulated into NAL (Network Abstraction Layer) units as basic units. NAL units are divided into VCL (Video Coding Layer) NAL units containing video data and non-VCL NAL units containing parameter information for decoding video data, and there are various types of VCL or non-VCL NAL units. A NAL unit consists of NAL header information and RBSP (Raw Byte Sequence Payload) data, and the NAL header information includes summary information about the RBSP. The RBSP of a VCL NAL unit includes an integer number of encoded coding tree units. In order to decode a bitstream in a video decoder, the bitstream must first be divided into NAL unit units, and then each divided NAL unit must be decoded. Meanwhile, information required for decoding a video signal bitstream can be transmitted as included in a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), an Adaptation Parameter Set (APS), etc. An RBSP of a VCL NAL unit can include an integer number of coding tree units. A VPS is a parameter set configured with a common syntax by extracting duplicate parameters from an SPS parameter set signaled for each layer in a bitstream that supports image quality, resolution, and frame rate scalability or a bitstream that supports multi-view.SPS is a parameter set that includes at least one of the following: Profile, which contains information about acceptable coding tools (or algorithms) and video formats; Level, which contains information about the decoder's processing capabilities, such as the resolution and frame rate of processable video and the allowable memory size; Tier, which contains information about the maximum bit rate that can be processed; and information about the resolution, bit depth, and whether or not a function can be enabled. PPS is a parameter set that includes at least one of the following: resolution of the video, tile division information, whether or not to enable weight prediction, quantization parameters, and filtering-related information. APS is a parameter set that includes one of the following: ALF filter coefficient information, LMCS-related parameters, and quantization scale parameters, depending on the APS type. APS is divided into prefix APS, which is signaled before the VCL NAL unit, and suffix APS, which is signaled after the VCL NAL unit. In the case of ALF APS, it is efficient to apply the ALF filter coefficients derived from the previous picture to the next picture, so it can be signaled as suffix APS.

[0095] Meanwhile, the block diagram of FIG. 1 illustrates an encoding device (100) according to one embodiment of the present specification, and the blocks shown separately logically distinguish elements of the encoding device (100). Accordingly, the elements of the encoding device (100) described above may be mounted as one chip or as multiple chips depending on the design of the device. According to one embodiment, the operations of each element of the encoding device (100) described above may be performed by a processor (not shown).

[0096] FIG. 2 is a schematic block diagram of a video signal decoding device (200) according to one embodiment of the present specification. Referring to FIG. 2, the decoding device (200) of the present invention includes an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (225), a filtering unit (230), and a prediction unit (250).

[0097] The entropy decoding unit (210) entropy decodes the video signal bitstream to extract transform coefficient information, intra-coding information, inter-coding information, etc. for each region. For example, the entropy decoding unit (210) can obtain a binarization code for transform coefficient information of a specific region from the video signal bitstream. In addition, the entropy decoding unit (210) inversely binarizes the binarization code to obtain a quantized transform coefficient. The inverse quantization unit (220) inversely quantizes the quantized transform coefficient, and the inverse transform unit (225) restores the residual value using the inverse quantized transform coefficient. The video signal processing device (200) restores the original pixel value by adding the residual value obtained by the inverse transform unit (225) and the prediction value obtained by the prediction unit (250).

[0098] Meanwhile, the filtering unit (230) performs filtering on the picture to improve the image quality. This may include a deblocking filter to reduce block distortion and / or an adaptive loop filter to remove distortion from the entire picture. The filtered picture is output or stored in the decoded picture buffer (DPB, 256) to be used as a reference picture for the next picture.

[0099] The prediction unit (250) includes an intra prediction unit (252) and an inter prediction unit (254). The prediction unit (250) generates a prediction picture by utilizing the encoding type decoded through the entropy decoding unit (210) described above, the transform coefficients for each region, intra / inter encoding information, etc. In order to restore the current block on which decoding is performed, the decoded region of the current picture or other pictures including the current block may be used. A picture (or tile / slice) that uses only the current picture for restoration, i.e., performs intra prediction or intra BC prediction, is called an intra picture or I picture (or tile / slice), and a picture (or tile / slice) that can perform all of intra prediction, inter prediction, and intra BC prediction is called an inter picture (or tile / slice). A picture (or tile / slice) that uses at most one motion vector and reference picture index to predict sample values ​​of each block among inter-pictures (or tiles / slices) is called a predictive picture or P-picture (or tile / slice), and a picture (or tile / slice) that uses at most two motion vectors and reference picture indices is called a bi-predictive picture or B-picture (or tile / slice). In other words, a P-picture (or tile / slice) uses at most one motion information set to predict each block, and a B-picture (or tile / slice) uses at most two motion information sets to predict each block. Here, a motion information set includes one or more motion vectors and one reference picture index.

[0100] The intra prediction unit (252) generates a prediction block using intra encoding information and reconstructed samples within the current picture. As described above, the intra encoding information may include at least one of an intra prediction mode, a Most Probable Mode (MPM) flag, and an MPM index. The intra prediction unit (252) predicts sample values ​​of the current block using reconstructed samples located on the left and / or above the current block as reference samples. In the present disclosure, the reconstructed samples, the reference samples, and the samples of the current block may represent pixels. Additionally, the sample values ​​may represent pixel values.

[0101] According to one embodiment, the reference samples may be samples included in a neighboring block of the current block. For example, the reference samples may be samples adjacent to the left boundary of the current block and / or samples adjacent to the upper boundary. In addition, the reference samples may be samples located on a line within a preset distance from the left boundary of the current block and / or samples located on a line within a preset distance from the upper boundary of the current block among the samples of the neighboring blocks of the current block. In this case, the neighboring blocks of the current block may include at least one of a left (L) block, an upper (A) block, a lower left (BL) block, an above right (AR) block, or an above left (AL) block adjacent to the current block. The neighboring blocks of the current block may be reference blocks for predicting the current block.

[0102] The inter prediction unit (254) generates a prediction block using the reference picture and inter encoding information stored in the decoded picture buffer (256). The inter encoding information may include a set of motion information (reference picture index, motion vector information, etc.) of the current block with respect to the reference block. Inter prediction may include L0 prediction, L1 prediction, and bi-prediction. L0 prediction is prediction using one reference picture included in the L0 picture list, and L1 prediction means prediction using one reference picture included in the L1 picture list. For this, one set of motion information (e.g., motion vector and reference picture index) may be required. In the bi-prediction method, up to two reference areas can be used, and these two reference areas may exist in the same reference picture or may exist in different pictures, respectively. That is, in the bi-prediction method, up to two sets of motion information (e.g., motion vectors and reference picture indices) can be used, and the two motion vectors may correspond to the same reference picture index or may correspond to different reference picture indices. At this time, the reference pictures are pictures that are located temporally before or after the current picture, and may be completed pictures that have already been restored. According to one embodiment, the two reference areas used in the bi-prediction method may be areas selected from each of the L0 picture list and the L1 picture list. In addition, a prediction method that uses only reference pictures having a POC smaller than the POC of the current picture or uses only reference pictures having a POC larger than the POC of the current picture based on the POC (picture order count) indicating the display order of the current picture can be called uni-directional prediction.Also, a prediction method that uses both a reference picture with a POC (picture order count) smaller than that of the current picture and a reference picture with a POC larger than that of the current picture based on the POC indicating the display order of the current picture can be called bi-directional prediction. A prediction method that uses only one reference picture in uni-directional prediction can be called uni-prediction, and a prediction method that uses two reference pictures in uni-directional prediction can be called bi-prediction or bi-prediction.

[0103] The inter prediction unit (254) can obtain a reference block of the current block using a motion vector and a reference picture index. The reference block exists in a reference picture corresponding to the reference picture index. In addition, a sample value of a block specified by the motion vector or an interpolated value thereof can be used as a predictor of the current block. For motion prediction with sub-pel unit pixel accuracy, for example, an 8-tap interpolation filter can be used for a luminance signal and a 4-tap interpolation filter can be used for a chrominance signal. However, the interpolation filter for sub-pel unit motion prediction is not limited thereto. In this way, the inter prediction unit (254) performs motion compensation to predict the texture of the current unit from a previously restored picture. At this time, the inter prediction unit can use a motion information set.

[0104] According to an additional embodiment, the prediction unit (250) may include an IBC prediction unit (not shown). The IBC prediction unit may reconstruct the current region by referring to a specific region including reconstructed samples within the current picture. The IBC prediction unit may perform IBC prediction using IBC encoding information obtained from the entropy decoding unit (210). The IBC encoding information may include block vector information.

[0105] A restored video picture is generated by adding the predicted value output from the intra prediction unit (252) or inter prediction unit (254) and the residual value output from the inverse transformation unit (225). That is, the video signal decoding device (200) restores the current block using the predicted block generated from the prediction unit (250) and the residual obtained from the inverse transformation unit (225).

[0106] Meanwhile, the block diagram of FIG. 2 illustrates a decoding device (200) according to one embodiment of the present specification, and the blocks shown separately illustrate logically distinguishing elements of the decoding device (200). Accordingly, the elements of the aforementioned decoding device (200) may be mounted as one chip or as multiple chips depending on the design of the device. According to one embodiment, the operations of each element of the aforementioned decoding device (200) may be performed by a processor (not shown).

[0107] Meanwhile, the technology proposed in this specification is applicable to both the methods and devices of the encoder and decoder, and the parts described as signaling and parsing may be described for convenience of explanation. In general, signaling can be described as encoding each syntax from the encoder's perspective, and parsing can be described as interpreting each syntax from the decoder's perspective. That is, each syntax can be included in the bitstream from the encoder and signaled, and the decoder can parse the syntax and use it in the restoration process. At this time, the sequence of bits for each syntax listed in the prescribed hierarchical structure can be referred to as a bitstream.

[0108] A picture can be encoded by dividing it into sub-pictures, slices, tiles, etc. A sub-picture can include one or more slices or tiles. When a picture is encoded by dividing it into multiple slices or tiles, all slices or tiles within the picture must be decoded before it can be displayed on the screen. On the other hand, when a picture is encoded into multiple sub-pictures, only any sub-picture can be decoded and displayed on the screen. A slice can include multiple tiles or sub-pictures, or a tile can include multiple sub-pictures or slices. Sub-pictures, slices, and tiles can be encoded or decoded independently, which is effective for parallel processing and improving processing speed. However, there is a disadvantage in that the amount of bits increases because the encoded information of adjacent sub-pictures, slices, and tiles cannot be used. Sub-pictures, slices, and tiles can be encoded by dividing them into multiple coding tree units (CTUs).

[0109] FIG. 3 illustrates an embodiment in which a Coding Tree Unit (CTU) within a picture is divided into Coding Units (CUs). In the process of coding a video signal, a picture may be divided into a sequence of Coding Tree Units (CTUs). A Coding Tree Unit may be composed of a luminance (luma) Coding Tree Block (CTB), two chroma (chroma) Coding Tree Blocks, and their encoded syntax information. One Coding Tree Unit may be composed of one Coding Unit, or one Coding Tree Unit may be split into multiple Coding Units. One Coding Unit may be composed of a luminance Coding Block (CB), two chroma Coding Blocks, and their encoded syntax information. One Coding Block may be split into multiple Sub-Coding Blocks. One Coding Unit may be composed of one Transform Unit (TU), or one Coding Unit may be split into multiple Transform Units. A single transform unit may consist of a luminance transform block (TB), two chrominance transform blocks, and their encoded syntax information. A coding tree unit may be divided into multiple coding units. A coding tree unit may not be divided and may also be a leaf node. In this case, the coding tree unit itself may be a coding unit.

[0110] A coding unit refers to a basic unit for processing a picture in the video signal processing described above, such as intra / inter prediction, transformation, quantization, and / or entropy coding. The size and shape of a coding unit within a picture may not be constant. A coding unit may have a square or rectangular shape. A rectangular coding unit (or rectangular block) includes a vertical coding unit (or vertical block) and a horizontal coding unit (or horizontal block). In the present specification, a vertical block is a block whose height is greater than its width, and a horizontal block is a block whose width is greater than its height. In addition, a non-square block in the present specification may refer to a rectangular block, but the present invention is not limited thereto.

[0111] Referring to Fig. 3, the coding tree unit is first partitioned into a Quad Tree (QT) structure. That is, in the Quad Tree structure, one node with a size of 2NX2N can be partitioned into four nodes with a size of NXN. In this specification, the Quad Tree may also be referred to as a quaternary tree. The Quad Tree partitioning can be performed recursively, and not all nodes need to be partitioned to the same depth.

[0112] Meanwhile, the leaf node of the aforementioned quad tree can be further split into a multi-type tree (MTT) structure. According to an embodiment of the present invention, in the multi-type tree structure, one node can be split into a binary or ternary tree structure of horizontal or vertical split. That is, the multi-type tree structure has four split structures: vertical binary split, horizontal binary split, vertical ternary split, and horizontal ternary split. According to an embodiment of the present invention, in each of the above tree structures, the width and height of the node can both have a power of 2 value. For example, in the binary tree (BT) structure, a node of size 2NX2N can be split into two NX2N nodes by vertical binary split, and can be split into two 2NXN nodes by horizontal binary split. Also, in a Ternary Tree (TT) structure, a node of size 2NX2N can be split into nodes of size (N / 2)X2N, NX2N, and (N / 2)X2N by vertical ternary splitting, and into nodes of size 2NX(N / 2), 2NXN, and 2NX(N / 2) by horizontal ternary splitting. This multi-type tree splitting can be performed recursively.

[0113] A leaf node of a multi-type tree can be a coding unit. If the coding unit is not larger than the maximum transformation length, the coding unit can be used as a unit of prediction and / or transformation without further splitting. In one embodiment, if the width or height of the current coding unit is larger than the maximum transformation length, the current coding unit can be split into multiple transformation units without explicit signaling regarding the splitting. Meanwhile, in the quad tree and multi-type tree described above, at least one of the following parameters can be predefined or transmitted through an RBSP of a higher-level set, such as a PPS, an SPS, or a VPS. 1) CTU size: The size of the root node of the quad tree, 2) MinQtSize: The minimum allowed QT leaf node size, 3) MaxBtSize: The maximum allowed BT root node size, 4) MaxTT size (MaxTtSize): The maximum allowed TT root node size, 5) MaxMttDepth: The maximum allowed depth of an MTT split from a leaf node of a QT, 6) MinBT size (MinBtSize): The minimum allowed BT leaf node size, 7) MinTT size (MinTtSize): The minimum allowed TT leaf node size.

[0114] Fig. 4 illustrates one embodiment of a method for signaling splitting of a quad tree and a multi-type tree. Pre-configured flags may be used to signal splitting of the quad tree and the multi-type tree described above. Referring to Fig. 4, at least one of a flag 'split_cu_flag' indicating whether a node is split, a flag 'split_qt_flag' indicating whether a quad tree node is split, a flag 'mtt_split_cu_vertical_flag' indicating a splitting direction of a multi-type tree node, or a flag 'mtt_split_cu_binary_flag' indicating a splitting shape of a multi-type tree node may be used.

[0115] According to an embodiment of the present invention, a flag 'split_cu_flag' indicating whether a current node is split may be signaled first. If the value of 'split_cu_flag' is 0, it indicates that the current node is not split, and the current node becomes a coding unit. If the current node is a coding tree unit, the coding tree unit includes one coding unit that is not split. If the current node is a quad tree node 'QT node', the current node is a leaf node 'QT leaf node' of the quad tree and becomes a coding unit. If the current node is a multi-type tree node 'MTT node', the current node is a leaf node 'MTT leaf node' of the multi-type tree and becomes a coding unit.

[0116] When the value of 'split_cu_flag' is 1, the current node can be split into nodes of a quad tree or a multi-type tree depending on the value of 'split_qt_flag'. The coding tree unit is the root node of the quad tree and can be first split into a quad tree structure. In the quad tree structure, 'split_qt_flag' is signaled for each node 'QT node'. When the value of 'split_qt_flag' is 1, the node is split into four square nodes, and when the value of 'split_qt_flag' is 0, the node becomes a leaf node 'QT leaf node' of the quad tree, and the node is split into multi-type nodes. According to an embodiment of the present invention, quad tree splitting can be limited depending on the type of the current node. Quad tree splitting may be allowed if the current node is a coding tree unit (root node of a quad tree) or a quad tree node, and quad tree splitting may not be allowed if the current node is a multi-type tree node. Each quad tree leaf node 'QT leaf node' may be further split into a multi-type tree structure. As described above, if 'split_qt_flag' is 0, the current node may be split into multi-type nodes. To indicate the splitting direction and splitting shape, 'mtt_split_cu_vertical_flag' and 'mtt_split_cu_binary_flag' may be signaled. If the value of 'mtt_split_cu_vertical_flag' is 1, a vertical split of the node 'MTT node' is indicated, and if the value of 'mtt_split_cu_vertical_flag' is 0, a horizontal split of the node 'MTT node' is indicated.Additionally, if the value of 'mtt_split_cu_binary_flag' is 1, the node 'MTT node' is split into two rectangular nodes, and if the value of 'mtt_split_cu_binary_flag' is 0, the node 'MTT node' is split into three rectangular nodes.

[0117] The tree partitioning structure allows luminance blocks and chrominance blocks to be partitioned in the same manner. That is, the chrominance blocks can be partitioned by referring to the partitioning form of the luminance blocks. If the current chrominance block is smaller than a predetermined size, the chrominance block may not be partitioned even if the luminance block is partitioned.

[0118] The luminance block and the chrominance block may have the same tree partitioning structure, which may be referred to as a single tree. If the current block is encoded with a single tree, the partitioning structure, encoding mode information, motion information, etc. of the luminance block and the chrominance block may be the same, and information related to other error signals may be different between the luminance block and the chrominance block. In addition, the luminance block and the chrominance block may have different tree partitioning structures, which may be referred to as a dual tree. If the current block is encoded with a dual tree, the partitioning structure, encoding mode information, motion information, etc. of the luminance block and the chrominance block may be different in at least one or more.

[0119] There may be a close correlation between a luminance block and its corresponding chrominance block. Therefore, when the current block is encoded and decoded using a dual tree, the encoder and decoder can use the segmentation information, encoding mode information, and motion information of the luminance block to encode the chrominance block.

[0120] A node to be divided into the smallest unit can be processed as a single coding block. If the current block is a coding block, the coding block can be divided into multiple sub-blocks (sub-coding blocks), and the prediction information of each sub-block can be the same or different. For example, if the coding unit is an intra mode, the intra prediction modes of each sub-block can be the same or different. Furthermore, if the coding unit is an inter mode, the motion information of each sub-block can be the same or different. Furthermore, each sub-block can be encoded or decoded independently. Each sub-block can be distinguished by a sub-block index (sbIdx). Furthermore, when a coding unit is divided into sub-blocks, it can be divided horizontally, vertically, or diagonally. In intra mode, the mode that divides the current coding unit into two or four sub-blocks horizontally or vertically is called ISP (Intra Sub-Partitions). In inter mode, the mode that divides the current coding block diagonally is called GPM (Geometric partitioning mode). In GPM mode, the position and direction of the diagonal line are derived using a predefined angle table, and the index information of the angle table is signaled.

[0121] The motion information may include one or more of reference direction indication information, reference picture information, motion vector, motion resolution, affine model, CPMV (control point motion vector), block vector, block vector resolution, MHP information, LIC information, filtering information, BCW information, and RRIBC information.

[0122] The reference direction indication information is composed of L0 prediction, L1 prediction, L0 and L1 prediction, and L0 prediction and L1 prediction are uni-prediction and unidirectional prediction, and L0 and L1 prediction are bi-prediction. And L0 and L1 prediction can be uni-prediction or bi-directional prediction. Here, L0 prediction is predicted using reference pictures in the L0 reference picture list, and L1 prediction can be predicted using reference pictures in the L1 reference picture list. In the L0 reference picture list, a reference picture with a POC smaller than the POC of the current picture can be added to the reference picture list based on the POC of the current picture. In addition, the L0 reference picture list can be composed in order from a reference picture with a POC closer to the POC of the current picture to a reference picture with a POC farther from the POC. In the L1 reference picture list, a reference picture with a POC larger than the POC of the current picture can be added to the reference picture list based on the POC of the current picture. Additionally, the L1 reference picture list can be organized in the order of reference pictures with a POC closer to the POC of the current picture to reference pictures with a POC farther away. The L0 and L1 reference picture lists can vary for each slice, subpicture, and picture. Additionally, the L0 reference picture list can include reference pictures of the L1 reference picture list. And the L1 reference picture list can include reference pictures of the L0 reference picture list.

[0123] Reference picture information may vary for each block and may be index information indicating which reference picture in the L0 reference picture list and / or the L1 reference picture list is used to predict the current block. The reference picture information may include one or more of L0 reference picture information and L1 reference picture information.

[0124] A motion vector is information that indicates the block that best matches the current block in a reference picture. It is a value that represents the distance to the reference block in horizontal and vertical coordinates based on the upper left position of the current block in the picture.

[0125] Motion resolution represents the resolution of the motion vector, and motion resolution can be expressed in units of 4 pixels, 1 pixel (integer pixel), 1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel, 1 / 16 pixel, etc.

[0126] A block vector is information indicating the block that best matches the current block in the already restored area within the current picture. It is a value that represents the distance to the reference block in horizontal and vertical coordinates based on the upper left position of the current block within the picture.

[0127] Block vector resolution can be expressed in units of 4 pixels, 1 pixel (integer pixel), 1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel, 1 / 16 pixel, etc.

[0128] MHP information may include whether additional motion information is applied and additional motion information.

[0129] LIC information may include whether LIC applies to the current block.

[0130] Filtering information may include filtering type and coefficient information applied according to the motion resolution of the current block.

[0131] BCW information may include whether BCW is applied to the current block.

[0132] RRIBC information may include information about whether RRIBC is applied and the RRIBC type if the current block is encoded in IBC mode.

[0133] Picture prediction (motion compensation) for coding is performed on coding units that are no longer divisible (i.e., leaf nodes of a coding tree unit). The basic unit performing this prediction is referred to below as a prediction unit or prediction block.

[0134] Hereinafter, the term "unit" used in this specification may be used as a replacement for the prediction unit, which is the basic unit for performing prediction. However, the present invention is not limited thereto, and can be understood more broadly as a concept that includes the coding unit.

[0135] Figures 5 and 6 illustrate an intra prediction method according to an embodiment of the present invention in more detail. As described above, the intra prediction unit predicts sample values ​​of the current block using restored samples located to the left and / or above the current block as reference samples.

[0136] First, FIG. 5 illustrates an embodiment of reference samples used for prediction of a current block in an intra prediction mode. According to an embodiment, the reference samples may be samples adjacent to a left boundary of the current block and / or samples adjacent to an upper boundary. As illustrated in FIG. 5, when the size of the current block is WXH and samples of a single reference line adjacent to the current block are used for intra prediction, reference samples may be set using at most 2W+2H+1 surrounding samples located on the left and / or upper sides of the current block.

[0137] Meanwhile, pixels of multiple reference lines may be used for intra prediction of the current block. The multiple reference lines may be composed of n lines located within a preset range from the current block. According to one embodiment, when pixels of multiple reference lines are used for intra prediction, separate index information indicating the lines to be set as reference pixels may be signaled, and this may be referred to as a reference line index.

[0138] In addition, if at least some of the samples to be used as reference samples have not yet been restored, the intra prediction unit may perform a reference sample padding process to obtain a reference sample. In addition, the intra prediction unit may perform a reference sample filtering process to reduce an error in intra prediction. That is, filtering may be performed on the surrounding samples and / or the reference samples obtained by the reference sample padding process to obtain filtered reference samples. The intra prediction unit predicts samples of the current block using the reference samples obtained in this manner. The intra prediction unit predicts samples of the current block using unfiltered reference samples or filtered reference samples. In the present disclosure, the surrounding samples may include samples on at least one reference line. For example, the surrounding samples may include adjacent samples on a line adjacent to a boundary of the current block.

[0139] Next, FIG. 6 illustrates an embodiment of prediction modes used for intra prediction. For intra prediction, intra prediction mode information indicating an intra prediction direction may be signaled. The intra prediction mode information indicates any one of a plurality of intra prediction modes constituting an intra prediction mode set. If the current block is an intra prediction block, the decoder receives intra prediction mode information of the current block from the bitstream. The intra prediction unit of the decoder performs intra prediction on the current block based on the extracted intra prediction mode information.

[0140] According to an embodiment of the present invention, an intra prediction mode set may include all intra prediction modes used for intra prediction (e.g., a total of 67 intra prediction modes). More specifically, the intra prediction mode set may include a planar mode, a DC mode, and a plurality of (e.g., 65) angular modes (i.e., directional modes). Each intra prediction mode may be indicated by a preset index (i.e., an intra prediction mode index). For example, as illustrated in FIG. 6, an intra prediction mode index 0 indicates a planar mode, and an intra prediction mode index 1 indicates a DC mode. In addition, intra prediction mode indices 2 to 66 may each indicate different angular modes. The angular modes each indicate different angles within a preset angular range. For example, an angular mode may indicate an angle within an angular range from 45 degrees to -135 degrees clockwise (i.e., a first angular range). The angular mode may be defined based on the 12 o'clock direction. At this time, intra prediction mode index 2 indicates horizontal diagonal (HDIA) mode, intra prediction mode index 18 indicates horizontal (HOR) mode, intra prediction mode index 34 indicates diagonal (DIA) mode, intra prediction mode index 50 indicates vertical (VER) mode, and intra prediction mode index 66 indicates vertical diagonal (VDIA) mode.

[0141] Meanwhile, the preset angle range may be set differently depending on the shape of the current block. For example, if the current block is a rectangular block, a wide-angle mode indicating an angle exceeding 45 degrees or less than -135 degrees in a clockwise direction may be additionally used. If the current block is a horizontal block, the angle mode may indicate an angle within an angle range (i.e., a second angle range) between (45+offset1) degrees and (-135+offset1) degrees in a clockwise direction. At this time, angle modes 67 to 76 that are outside the first angle range may be additionally used. In addition, if the current block is a vertical block, the angle mode may indicate an angle within an angle range (i.e., a third angle range) between (45-offset2) degrees and (-135-offset2) degrees in a clockwise direction. At this time, angle modes -10 to -1 that are outside the first angle range may be additionally used. According to an embodiment of the present invention, the values ​​of offset1 and offset2 may be determined differently depending on the ratio between the width and height of the rectangular block. Also, offset1 and offset2 can be positive.

[0142] According to an additional embodiment of the present invention, the plurality of angular modes constituting the intra prediction mode set may include a basic angular mode and an extended angular mode. In this case, the extended angular mode may be determined based on the basic angular mode.

[0143] According to one embodiment, the basic angle mode may be a mode corresponding to an angle used in intra prediction of the existing HEVC (High Efficiency Video Coding) standard, and the extended angle mode may be a mode corresponding to an angle newly added in intra prediction of the next-generation video codec standard. More specifically, the basic angle mode may be an angle mode corresponding to any one of the intra prediction modes {2, 4, 6, … , 66}, and the extended angle mode may be an angle mode corresponding to any one of the intra prediction modes {3, 5, 7, … , 65}. That is, the extended angle mode may be an angle mode between the basic angle modes within the first angle range. Therefore, the angle indicated by the extended angle mode may be determined based on the angle indicated by the basic angle mode.

[0144] According to another embodiment, the basic angle mode may be a mode corresponding to an angle within a preset first angle range, and the extended angle mode may be a wide-angle mode outside the first angle range. That is, the basic angle mode may be an angle mode corresponding to any one of the intra prediction modes {2, 3, 4, … , 66}, and the extended angle mode may be an angle mode corresponding to any one of the intra prediction modes {-14, –13, –12, … , –1} and {67, 68, … , 80}. The angle indicated by the extended angle mode may be determined as an opposite angle to the angle indicated by the corresponding basic angle mode. Accordingly, the angle indicated by the extended angle mode may be determined based on the angle indicated by the basic angle mode. Meanwhile, the number of extended angle modes is not limited thereto, and additional extended angles may be defined according to the size and / or shape of the current block. Meanwhile, the total number of intra prediction modes included in the intra prediction mode set may vary depending on the configuration of the basic angle mode and the extended angle mode described above.

[0145] In the above embodiment, the spacing between the extended angular modes may be set based on the spacing between the corresponding basic angular modes. For example, the spacing between the extended angular modes {3, 5, 7, … , 65} may be determined based on the spacing between the corresponding basic angular modes {2, 4, 6, … , 66}. In addition, the spacing between the extended angular modes {-14, -13, … , -1} may be determined based on the spacing between the corresponding opposite basic angular modes {53, 53, … , 66}, and the spacing between the extended angular modes {67, 68, … , 80} may be determined based on the spacing between the corresponding opposite basic angular modes {2, 3, 4, … , 15}. The angular spacing between the extended angular modes may be set to be equal to the angular spacing between the corresponding basic angular modes. In addition, the number of the extended angular modes in the intra prediction mode set may be set to be less than or equal to the number of the basic angular modes.

[0146] According to an embodiment of the present invention, an extended angular mode may be signaled based on a base angular mode. For example, a wide-angle mode (i.e., an extended angular mode) may replace at least one angular mode (i.e., a base angular mode) within a first angular range. The replaced base angular mode may be an angular mode corresponding to an opposite side of the wide-angle mode. That is, the replaced base angular mode is an angular mode corresponding to an angle in an opposite direction to an angle indicated by the wide-angle mode or an angle that is different from the angle in the opposite direction by a preset offset index. According to an embodiment of the present invention, the preset offset index is 1. An intra prediction mode index corresponding to the replaced base angular mode may be remapped to the wide-angle mode to signal the corresponding wide-angle mode. For example, the wide-angle modes {-14, -13, … , -1} may be signaled by intra prediction mode indices {52, 53, … , 66}, and the wide-angle modes {67, 68, … , 69} may be signaled by intra prediction mode indices {52, 53, … , 66}, respectively. , 80} can be signaled by intra prediction mode indices {2, 3, … , 15}, respectively. By having the intra prediction mode index for the basic angular mode signal the extended angular mode in this way, even if the configurations of the angular modes used for intra prediction of each block are different, the same set of intra prediction mode indices can be used to signal the intra prediction mode. Therefore, the signaling overhead due to changes in the intra prediction mode configuration can be minimized.

[0147] Meanwhile, whether to use the extended angle mode may be determined based on at least one of the shape and size of the current block. In one embodiment, if the size of the current block is larger than a preset size, the extended angle mode may be used for intra prediction of the current block, and otherwise, only the default angle mode may be used for intra prediction of the current block. In another embodiment, if the current block is a non-square block, the extended angle mode may be used for intra prediction of the current block, and if the current block is a square block, only the default angle mode may be used for intra prediction of the current block.

[0148] The intra prediction unit determines reference samples and / or interpolated reference samples to be used for intra prediction of the current block based on intra prediction mode information of the current block. If the intra prediction mode index indicates a specific angular mode, the reference sample or interpolated reference sample corresponding to the specific angle from the current sample of the current block is used for prediction of the current pixel. Therefore, different sets of reference samples and / or interpolated reference samples can be used for intra prediction depending on the intra prediction mode. After intra prediction of the current block is performed using the reference samples and intra prediction mode information, the decoder restores the sample values ​​of the current block by adding the residual signal of the current block obtained from the inverse transform unit to the intra prediction value of the current block.

[0149] Motion information used for inter prediction may include reference direction indication information (inter_pred_idc), reference picture indexes (ref_idx_l0, ref_idx_l1), and motion vectors (mvL0, mvL1). Reference picture list utilization information (predFlagL0, predFlagL1) may be set according to the reference direction indication information. As an example, in case of unidirectional prediction using an L0 reference picture, predFlagL0=1, predFlagL1=0 may be set. In case of unidirectional prediction using an L1 reference picture, predFlagL0=0, predFlagL1=1 may be set. In case of bidirectional prediction using both L0 and L1 reference pictures, predFlagL0=1, predFlagL1=1 may be set.

[0150] If the current block is a coding unit, the coding unit can be divided into multiple sub-blocks, and the prediction information of each sub-block can be the same or different. For example, if the coding unit is an intra mode, the intra prediction modes of each sub-block can be the same or different. In addition, if the coding unit is an inter mode, the motion information of each sub-block can be the same or different. In addition, each sub-block can be encoded or decoded independently. Each sub-block can be distinguished through a sub-block index (sbIdx).

[0151] The motion vector of the current block is likely to be similar to that of the surrounding blocks. Therefore, the motion vectors of the surrounding blocks can be used as motion vector predictors (mvp), and the motion vector of the current block can be derived using the motion vectors of the surrounding blocks. Furthermore, to improve the accuracy of the motion vector, the difference in motion vectors (mvd) between the optimal motion vector of the current block, found from the original image by the encoder, and the motion predictor can be signaled.

[0152] Motion vectors can have various resolutions, and the resolution of motion vectors can vary on a block-by-block basis. Motion vector resolution can be expressed in integer units, half-pixel units, quarter-pixel units, sixteenth-pixel units, and integer-4 pixel units. Since images such as screen content are in simple graphical forms such as characters, interpolation filters do not need to be applied, integer units and integer-4 pixel units can be selectively applied on a block-by-block basis. Blocks encoded in affine mode, which can express rotation and scale, have significant shape changes, so integer units, quarter-pixel units, and sixteenth-pixel units can be selectively applied on a block-by-block basis. Information on whether to selectively apply motion vector resolution on a block-by-block basis is signaled with amvr_flag. If applicable, which motion vector resolution to apply to the current block is signaled with amvr_precision_idx.

[0153] For blocks to which bidirectional prediction is applied, the weights between the two prediction blocks can be applied equally or differently when applying weighted average, and information about the weights is signaled through bcw_idx.

[0154] To improve the accuracy of the motion prediction value, Merge or advanced motion vector prediction (AMVP) methods can be selectively used on a block-by-block basis. The Merge method configures the motion information of the current block to be identical to the motion information of the adjacent blocks to the current block, and has the advantage of increasing the encoding efficiency of the motion information by spatially propagating the motion information without change in the homogeneous motion region. On the other hand, the AMVP method predicts motion information in the L0 and L1 prediction directions respectively to express accurate motion information and signals the most optimal motion information. After the decoder derives the motion information for the current block through the AMVP or Merge method, it uses the reference block located in the motion information derived from the reference picture as the prediction block for the current block.

[0155] A method for deriving motion information in Merge or AMVP may be to construct a motion candidate list using motion prediction values ​​derived from neighboring blocks of the current block, and then signal the index information for the optimal motion candidate. In the case of AMVP, since motion candidate lists are derived for each of L0 and L1, the optimal motion candidate indices (mvp_l0_flag, mvp_l1_flag) for each of L0 and L1 are signaled. In the case of Merge, since one motion candidate list is derived, one merge index (merge_idx) is signaled. The motion candidate lists derived from one coding unit may vary, and a motion candidate index or merge index may be signaled for each motion candidate list. In this case, a mode in which there is no information about the remaining blocks in a block encoded in Merge mode can be referred to as MergeSkip mode.

[0156] Bidirectional motion information for the current block can be derived using a combination of AMVP and Merge modes. For example, motion information in the L0 direction can be derived using AMVP, while motion information in the L1 direction can be derived using Merge. Conversely, Merge can be applied to L0, while AMVP can be applied to L1. This encoding mode can be referred to as AMVP-merge mode.

[0157] The motion candidates and motion information candidates in this specification may have the same meaning. Furthermore, the motion candidate list and motion information candidate list in this specification may have the same meaning.

[0158] SMVD (Symmetric MVD) is a method for reducing the bit volume of transmitted motion information by ensuring that the MVD (Motion Vector Difference) values ​​in the L0 and L1 directions are symmetrical in the case of bidirectional prediction. The MVD information in the L1 direction, which is symmetrical to the L0 direction, is not transmitted, and reference picture information in the L0 and L1 directions is also not transmitted and can be derived during the decoding process.

[0159] OBMC (Overlapped Block Motion Compensation) is a method that generates prediction blocks for the current block using motion information from surrounding blocks when the motion information between blocks differs. It then weights and averages these prediction blocks to generate a final prediction block for the current block. This method effectively reduces blocking artifacts that occur at block boundaries in motion-compensated images.

[0160] Typically, merge motion candidates have low motion accuracy. To improve the accuracy of these merge motion candidates, the MMVD (Merge mode with MVD) method can be used. The MMVD method is a method of compensating motion information using one candidate selected from several motion difference value candidates. Information about the compensation value of the motion information obtained through the MMVD method (e.g., an index indicating one candidate selected from among the motion difference value candidates) can be included in the bitstream and transmitted to the decoder. Compared to the conventional method of including the motion information difference value in the bitstream, the inclusion of information about the compensation value of the motion information in the bitstream can save bits.

[0161] MBVD (merge mode with block vector differences) mode is a mode that encodes the differential value for a block vector, similar to MMVD, which is a mode that encodes the differential value for a motion vector in merge mode. The MBVD method is a method that determines a block vector by using one candidate selected from several block vector difference value candidates. Information about the differential value of the block vector obtained through the MBVD method (e.g., an index indicating one candidate selected from among the block vector difference value candidates) can be included in the bitstream and transmitted to the decoder. Compared to including the existing block vector difference value in the bitstream, the amount of bits can be saved by including only a part of the information about the differential value of the block vector in the bitstream.

[0162] The TM (Template Matching) method constructs a template using the surrounding pixels of the current block, finds the matching area with the highest similarity to the template, and then compensates for motion information. TM (Template Matching) is a method that performs motion prediction in the decoder without including motion information in the bitstream to reduce the size of the encoded bitstream. Since the decoder does not have the original image, it can roughly derive motion information for the current block using already reconstructed surrounding blocks.

[0163] The DMVR (Decoder-side Motion Vector Refinement) method is a method of compensating motion information through the correlation of already restored reference images to find slightly more accurate motion information. It is a method of using the bidirectional motion information of the current block to use the best matching point between reference blocks within an arbitrary set area of ​​two reference pictures as a new bidirectional motion. When this DMVR is performed, the encoder performs DMVR on a block basis to compensate for the motion information, and then divides the block into sub-blocks and performs DMVR on each sub-block basis to compensate for the motion information of the sub-block again. This can be called MP-DMVR (Multi-pass DMVR).

[0164] The LIC (Local Illumination Compensation) method is a method of compensating for luminance changes between blocks. It is a method of deriving a linear model using neighboring pixels adjacent to the current block, and then compensating the luminance information of the current block through the linear model.

[0165] Existing video encoding methods perform motion compensation that only considers parallel translation in all directions, which reduces encoding efficiency when encoding videos containing motions commonly encountered in the real world, such as zooming in, zooming out, and rotation. To express such motions, an affine model-based motion prediction technique utilizing a four-parameter (rotation) or six-parameter (zoom in, zoom out, rotation) model can be applied.

[0166] Bi-Directional Optical Flow (BDOF) is used to compensate for predicted blocks by estimating pixel changes based on optical flow from a reference block of a block composed of bidirectional motion. The motion information derived from BDOF in VVC can be used to compensate for the motion of the current block.

[0167] PROF (Prediction Refinement with Optical Flow) is a technique for improving the accuracy of sub-block-level affine motion prediction to be similar to that of pixel-level motion prediction. Similar to BDOF, PROF calculates a correction value on a pixel-by-pixel basis for affine motion-compensated pixel values ​​on a sub-block-by-subblock basis based on optical flow to obtain a final prediction signal.

[0168] The CIIP (Combined Inter- / Intra-picture Prediction) method is a method of generating a final prediction block by weighting and averaging prediction blocks generated by the intra-picture prediction method and prediction blocks generated by the inter-picture prediction method when generating a prediction block for the current block.

[0169] The Intra Block Copy (IBC) method finds the most similar part to the current block in a previously restored area within the current picture and uses that reference block as a prediction block for the current block. At this time, information related to the block vector, which represents the distance between the current block and the reference block, can be included in the bitstream. The decoder can parse the block vector information contained in the BeastStream to calculate or set the block vector for the current block.

[0170] The BCW (Bi-prediction with CU-level Weights) method is a method that performs a weighted average on two motion-compensated prediction blocks by adaptively applying weights on a block-by-block basis, rather than generating a prediction block by averaging two motion-compensated prediction blocks from different reference pictures.

[0171] The Intra TMP (Template Matching Prediction) method is a method in which a video signal processing device constructs a reference template using pixel values ​​of neighboring blocks adjacent to the current block, finds the part most similar to the constructed reference template in an already restored area within the current picture, and then uses the reference block (the part found in the already restored area) as a prediction block for the current block.

[0172] The MHP (Multi-hypothesis prediction) method is a method of performing weight prediction using various prediction signals by transmitting additional motion information to unidirectional and bidirectional motion information during inter-screen prediction.

[0173] CCLM (Cross-component linear model) is a method that constructs a linear model using the high correlation between a luminance signal and a chrominance signal located at the same location as the luminance signal, and then predicts the chrominance signal using the linear model. A template is constructed using blocks adjacent to the current block that have been reconstructed, and the parameters for the linear model are derived from the template. Next, the current luminance block, which has been reconstructed to fit the size of the chrominance block, is optionally downsampled depending on the image format. Finally, the chrominance block of the current block is predicted using the downsampled luminance block and the linear model. In this case, the method of using two or more linear models is called MMLM (Multi-model Linear mode).

[0174] GLM (Gradient Linear Model) is a method of predicting a color difference signal through a linear model such as CCLM, which constructs a model by additionally reflecting the gradient between the luminance sample corresponding to the color difference sample and the surrounding luminance samples adjacent to that luminance sample.

[0175] Prediction methods that utilize the correlation between different signals, such as CCLM, MMLM, CCCM, and GLM, can be called Cross-Component Prediction (CCP). In other words, a method of predicting another signal (a chrominance signal, such as Cb or Cr) from one signal (e.g., a luminance signal) can be called Cross-Component Prediction (CCP).

[0176] The CCP merge method is a method of predicting the chrominance block of the current block using the CCP model (models such as CCLM, MMLM, and CCCM) used in the surrounding blocks.

[0177] The encoding modes (prediction modes) described herein may be described without the term "mode" or with the term "method" instead of "mode." For example, the CCLM mode may be described as the CCLM method or CCLM.

[0178]

[0179] Among intra prediction encoding techniques, the MIP (Matrix Intra Prediction) method is a matrix-based intra prediction method. Unlike the prediction method that has directionality from the pixels of the surrounding blocks adjacent to the current block, it is a method of obtaining a prediction signal using a predefined matrix and offset values ​​for the pixels on the left and top of the surrounding blocks.

[0180] To derive the intra-prediction mode of the current block, a template, which is a random region adjacent to the current block and reconstructed, is used. The intra-prediction mode derived from the template's surrounding pixels can be used to reconstruct the current block. First, the decoder generates a prediction template for the template using the surrounding pixels (reference) adjacent to the template. The intra-prediction mode that generates the prediction template most similar to the already reconstructed template can then be used to reconstruct the current block. This method is called TIMD (Template intra-mode derivation).

[0181] Typically, an encoder can determine a prediction mode for generating a prediction block and generate a bitstream containing information about the determined prediction mode. A decoder can parse the received bitstream to set an intra-prediction mode. At this time, the bit amount of information about the prediction mode may be about 10% of the total bitstream size. To reduce the bit amount of information about the prediction mode, the encoder may not include information about the intra-prediction mode in the bitstream. Accordingly, the decoder can derive (determine) an intra-prediction mode for restoring the current block using the characteristics of the surrounding blocks, and can restore the current block using the derived intra-prediction mode. At this time, the decoder can use a method of applying a Sobel filter in the horizontal and vertical directions to each neighboring pixel (pixel) of the current block to infer directional information, and then mapping the directional information to the intra-prediction mode to derive the intra-prediction mode. The method by which the decoder derives the intra prediction mode using surrounding blocks can be described as Decoder side intra mode derivation (DIMD).

[0182] If the current block is a luminance block, the video signal processing device can derive an intra prediction mode through the DIMD and TIMD methods. If the current block is a chroma block, since there is an already reconstructed luminance block corresponding to the chroma block, the video signal processing device can derive an intra prediction mode by applying the DIMD and TIMD methods using the reconstructed luminance block, and then use the derived intra prediction mode as the intra prediction mode of the chroma block. That is, the video signal processing device can derive directional information if the current block is a luminance block, and can not derive directional information if the current block is a chroma block, and can apply the directional information found in the luminance block to the chroma block. This mode can be called the 'DIMD Chroma' mode or the 'TIMD Chroma' mode.

[0183] Figure 7 is a diagram showing the locations of surrounding blocks used to construct a motion candidate list in inter prediction.

[0184] The neighboring blocks can be blocks of spatial position or blocks of temporal position. The neighboring block spatially adjacent to the current block can be at least one of a left (A1) block, a left below (A0) block, an above (B1) block, an above right (B0) block, and an above left (B2) block. The neighboring block temporally adjacent to the current block can be a block that includes an upper left pixel position of a bottom right (BR) block of the current block in a corresponding picture (Collocated picture). If the neighboring block temporally adjacent to the current block is encoded in intra mode or the neighboring block temporally adjacent to the current block exists in an unusable position, a block that includes a horizontal and vertical center (Center, Ctr) pixel position of the current block in the picture corresponding to the current picture (Collocated picture) can be used as the temporal neighboring block. Motion candidate information derived from corresponding pictures can be referred to as a Temporal Motion Vector Predictor (TMVP). Only one TMVP can be derived from a single block, or a block can be divided into multiple sub-blocks, and then a separate TMVP candidate can be derived for each sub-block. The method for deriving TMVPs at the sub-block level can be referred to as a sub-block Temporal Motion Vector Predictor (sbTMVP).

[0185] Whether the methods described in this specification are applied can be determined based on at least one of information about the slice type (e.g., whether it is an I slice, a P slice, or a B slice), whether it is a tile, whether it is a subpicture, the size of the current block, the depth of the coding unit, whether the current block is a luminance block or a chrominance block, whether it is a reference frame or a non-reference frame, the temporal layer according to the reference order and layer, etc. The information used to determine whether the methods described in this specification are applied can be information agreed upon in advance between the decoder and the encoder. In addition, such information can be determined according to a profile and a level. Such information can be expressed as a variable value, and information about the variable value can be included in the bitstream. That is, the decoder can determine whether the above-described methods are applied by parsing information about the variable value included in the bitstream. For example, whether the above-described methods are applied can be determined based on the horizontal length or the vertical length of the coding unit. The above-described methods can be applied if the horizontal or vertical length is 32 or more (e.g., 32, 64, 128, etc.). In addition, the above-described methods can be applied if the horizontal or vertical length is less than 32 (e.g., 2, 4, 8, 16). In addition, the above-described methods can be applied if the horizontal or vertical length is 4 or 8.

[0186] FIG. 8 illustrates a method for determining a reference pixel line based on a template according to one embodiment of the present specification.

[0187] Below, we describe a method for determining the optimal reference pixel line for the current block (for restoring the current block) based on a template. Here, the reference pixel line may have the same meaning as the reference sample line.

[0188] Referring to FIG. 8, a video signal processing device can construct a reference template using reference pixel lines adjacent to a current block. The video signal processing device can generate prediction samples for the positions of the reference template using reference pixel lines 1, 2, 3, etc. The video signal processing device can calculate a cost between the generated prediction samples and samples of the reference template. At this time, the cost can be calculated using a method such as SAD (Sum of Absolute Differences) or MRSAD (Mean-Removed SAD). The reference pixel corresponding to the minimum cost may be the optimal reference pixel. In addition, the encoder can rearrange the calculated costs in ascending order, construct a list for reference pixel lines, and then generate and signal a bitstream including information on an index for the optimal reference pixel line. The decoder can construct a list for reference pixel lines using the above-described method, parse the index for the optimal reference pixel line included in the bitstream, and generate a prediction sample using the reference pixel line indicated by the index. As described herein, a method for a video signal processing device to determine a reference pixel line based on a template can be described as a Template-based Multiple Reference Line (TMRL) method or a TMRL intra prediction method.

[0189] FIG. 9 is a diagram illustrating a block vector related to an IBC encoding method according to one embodiment of the present specification.

[0190] The IBC coding method (IBC mode) is a method of finding the part (reference block) that is most similar to the current block in an already reconstructed area within the current picture and using the reference block as a prediction block for the current block. At this time, the encoder can generate a bitstream that includes information related to a block vector, which is the distance between the current block and the reference block. The decoder can parse information related to the block vector included in the BeastStream to calculate or set the block vector for the current block. The IBC coding method can be applied to chrominance blocks. In the chrominance block, the block vector of the luminance block corresponding to the chrominance block can be used as the block vector of the chrominance block without finding a new block vector, and this coding method can be called DBV (Direct Block Vector) mode.

[0191] The RRIBC (Reconstruction-Reordered IBC) encoding mode can be used in IBC blocks (blocks to which the IBC encoding method is applied). RRIBC can be composed of vertical flips and horizontal flips. The reconstructed block of a block to which RRIBC is applied is flipped according to the RRIBC type of the current block. The encoder can flip the original block to be encoded before finding the most similar part to the current block from the reference picture. That is, the most similar part from the reference picture is found using the flipped original block. Therefore, the prediction block uses a block that has not been flipped, and the residual block can also be a block that has not been flipped. The decoder can flip the final reconstructed block according to the RRIBC type of the current block.

[0192] Figure 10 illustrates a method for predicting the current block using RRIBC in the horizontal direction.

[0193] Figure 11 illustrates a method for predicting the current block using RRIBC in the vertical direction.

[0194] In Figures 10 and 11, (Xn, Yn) represents the central position of the surrounding blocks, and (Xc, Yc) represents the central position of the current block. BV n h , BV n v represents the horizontal block vector and vertical block vector of the surrounding blocks, respectively, and BV C h , BV C v represent the horizontal block vector and the vertical block vector of the current block, respectively. As shown in Fig. 10, if the RRIBC type of the current block is horizontal, BV C h is 2 * (Xn - Xc) + BV n h It can be calculated as, and as in Fig. 53, if the RRIBC type of the current block is in the vertical direction, BV C v is 2 * (y n - y c ) + BV n v can be calculated as . At this time, BV n h , BV n v Since it uses restored blocks, BV n h , BV n v The sign of can be negative.

[0195] When the current block is encoded with RRIBC, the video signal processing device can determine an optimal block vector by constructing a block vector candidate list. At this time, the video signal processing device can construct the block vector candidate list according to the RRIBC type of the current block. For example, when the RRIBC type of the current block is horizontal, the video signal processing device can construct the block vector candidate list using only the neighboring blocks of the current block encoded with RRIBC in the horizontal direction. In addition, the video signal processing device can construct the block vector candidate list regardless of the RRIBC type of the current block. For example, when the RRIBC type of the current block is horizontal, the block vector candidate list can be constructed using not only the neighboring blocks of the current block encoded with RRIBC in the horizontal direction, but also the neighboring blocks of the current block encoded with RRIBC in the vertical direction, and / or blocks encoded with general motion, and / or blocks encoded with block vectors.

[0196] When a video signal processing device predicts a current block using general motion, the video signal processing device can construct a motion candidate list depending on whether the neighboring blocks of the current block are encoded in the IBC mode or the RRIBC mode. At this time, if the neighboring blocks of the current block are encoded in the RRIBC mode, the video signal processing device can construct the motion candidate list by additionally considering the RRIBC type. For example, when the video signal processing device constructs a motion candidate list for the current block, if the encoding mode of the neighboring blocks of the current block is the RRIBC mode and the RRIBC type is the vertical direction or the horizontal direction, the block vector of the neighboring blocks may not be included in the motion candidate list. Alternatively, when the video signal processing device predicts the current block using general motion, the video signal processing device can construct a motion candidate list regardless of the encoding mode of the neighboring blocks of the current block. In other words, the video signal processing device can construct a motion candidate list regardless of whether the neighboring blocks of the current block are encoded in the IBC mode or the RRIBC mode. For example, when a video signal processing device constructs a motion candidate list for a current block, the video signal processing device may include block vectors of the surrounding blocks in the motion candidate list even if the encoding mode of the surrounding blocks of the current block is the IBC mode and the RRIBC type is the vertical direction or the horizontal direction.

[0197] The RRIBC encoding method is effective for images with symmetrical characteristics. Symmetry can mean perfect horizontal (or vertical) symmetry with equal distances between the current block and the reference block around a central axis. The vertical (or horizontal) direction (symmetry axis) of the current block can be colinear with the vertical (or horizontal) direction (symmetry axis) of the reference block. In this case, the block vector in the vertical (or horizontal) direction can be set to '0', and the block vector in the horizontal direction can be set to any negative value other than '0'. The current block and the reference block can be configured to be symmetric with different distances apart from one central axis between the current block and the reference block, and the current block and the reference block can be located on different vertical lines. In this case, the block vector in the vertical direction can be set to any negative value other than '0'. That is, the video signal processing device can encode or decode the current block using a symmetric block having a block vector value of an arbitrary negative value other than '0' in both the horizontal and vertical directions.

[0198] FIG. 12 illustrates a block vector of a block encoded in Intra TMP mode according to one embodiment of the present specification.

[0199] Referring to FIG. 12, the Intra TMP method (Intra TMP encoding mode) is a method in which a video signal processing device constructs a reference template using pixel values ​​of neighboring blocks adjacent to a current block, then finds the most similar part to the reference template in an already reconstructed area (reference block) within the current picture, and then uses the reference block (Ref. luma block of FIG. 12) as a prediction block for the current block. At this time, there may be more than one block vector used to generate a prediction block for the current luminance block, and FIG. 12 shows a case in which there are two block vectors. The video signal processing device can generate a prediction block for the luminance block by weighting and averaging the reference blocks predicted from the two block vectors of the current luminance block. In the case of a chrominance block, the video signal processing device can derive a block vector from a luminance block corresponding to the chrominance block, and then generate a chrominance prediction block using the block vector. There may be more than one block vector used to generate a prediction block for the current chrominance block, and Fig. 12 shows a case where there are two block vectors for the chrominance block. In this case, the block vector of the chrominance block may be the same as or different from the block vector derived from the luminance block.

[0200] FIG. 13 illustrates a case where a current block is divided by GPM mode according to one embodiment of the present specification and the divided area is encoded by IBC mode.

[0201] The current block can be encoded or decoded in IBC-GPM mode. IBC-GPM mode may mean a mode in which the current block is divided using GPM mode, and the divided region is encoded in intra prediction mode or IBC prediction mode. Referring to Fig. 13, the current block can be divided into two regions based on a dotted line. At this time, one of the divided regions can be encoded in intra prediction mode, and the other can be encoded in IBC prediction mode. For example, the left region among the divided regions can be encoded in intra prediction mode, and the right region can be encoded in IBC prediction mode. At this time, the region encoded in IBC prediction mode can be encoded in Intra TMP prediction mode, and this can be described as IntraTMP-GPM mode.

[0202] FIG. 14 illustrates how a current block is encoded in IBC-CIIP mode according to one embodiment of the present specification.

[0203] A video signal processing device can predict the current block using a weighted average between a block predicted in Intra mode and a block predicted in IBC mode, which can be referred to as IBC-CIIP mode. In this case, a block predicted in Intra TMP mode can be used instead of a block predicted in IBC mode, which can be referred to as Intra TMP-CIIP mode.

[0204] A video signal processing device can obtain a reconstructed luminance block by adding a residual signal for a luminance prediction block of a current block and a luminance block, and can configure a CCP model by using a correlation between the reconstructed luminance block and the luminance prediction block. At this time, the CCP model can be one of CCLM, MMLM, GLM, CCCM, MM-CCCM, GL-CCCM, CCCM-ND, and CCCM-MDF. The CCP model derived from the luminance block can be applied to a chrominance prediction block to generate a first chrominance prediction block to which the CCP model is applied. A final chrominance block can be generated by adding an error signal for the first chrominance prediction block and the chrominance block.

[0205] FIG. 15 illustrates an example of a reference region and filter shape used to derive CCCM parameters according to one embodiment of the present specification.

[0206] CCCM (Convolutional cross-component intra prediction model) is a method of predicting a chrominance signal using a nonlinear model that utilizes the high correlation between a luminance signal and a chrominance signal located at the same location as the luminance signal. Fig. 15(a) shows the positional relationship between reference samples (vertical hatching, 1520) for applying CCCM to the current prediction block (diagonal hatching, 1510) and side samples (horizontal hatching) required when applying a cross-shaped filter. The current prediction block (MxN) can be composed of reference samples in the upper 6 rows (2Mx6), reference samples in the left 6 rows (6x2N), and 6x6 reference samples in the upper left, where the ratio of the number of chroma to luma samples is 1:1. When a video signal processing device applies a cross-shaped sample (Fig. 15(b)) filter to a chroma sample prediction relationship (Fig. 15(c)) for CCCM, cases where the reference sample area is exceeded may occur. At this time, additionally required samples may be side samples. The chroma sample prediction relationship of Fig. 15(c) may be applied to each chroma component (i.e., Cb, Cr). The sample at the C (Center) position in Fig. 15(b) may be a luma sample corresponding to the Cb and Cr chroma samples, and N (North), E (East), S (South), and W (West) may be luma samples adjacent to the luma sample at the C position. Side samples may additionally require one sample for an area other than the reference sample depending on the position of the C sample. If the sample value existing at the position of the side sample is not available, the sample value at the unavailable position may be padded with the C sample value. The P value of Fig. 15(c) may be a nonlinear term. The P value may be calculated as P = ( C*C + midVal ) >> bitDepth , and in the case of 10-bit content, P = ( C*C + 512 ) >> 10 .In Fig. 15(c), the B value can be an integer offset value as a bias value. The B value can be an intermediate value of bitDepth (bit depth). In the case of 10-bit content, the B value can be 512.

[0207] MM-CCCM (multi-model CCCM) is a method that derives two CCCM parameters based on the average value of the reference area (or the restored current luminance block).

[0208] GL-CCCM (Gradient and location based convolutional cross-component model) is an additional CCCM mode that uses gradient and location information. In the case of the existing CCCM mode, the video signal processing device can derive the chrominance sample for the current block using the luminance sample at the corresponding position from the predicted chrominance sample position, four samples around the luminance sample, and coefficient information. At this time, in the case of the GL-CCCM mode, the video signal processing device can derive the chrominance sample for the current block by reflecting the vertical and horizontal differences for the luminance sample at the corresponding position from the predicted chrominance sample position and eight samples around the luminance sample, and also using the position value of the current luminance sample and its coefficient information.

[0209] When CCCM mode is applied, the video signal processing device can apply a downsampling filter to match the resolution difference between the luminance block and the chrominance block. This is to reduce the resolution of the luminance block to that of the chrominance block. The mode that applies the downsampling filter can be described as CCCM-MDF (CCCM with multiple downsampling filters).

[0210] When the inter encoding mode is applied to the current block, the video signal processing device can derive linear and nonlinear models between the luminance prediction block (Y') derived using the motion information of the current block and the first chrominance prediction block (Cb', Cr') derived using the motion information of the current block (Derive filter), and then generate a reconstructed luminance block of the current block using the luminance residual block of the current block, and apply the derived linear and nonlinear models to the reconstructed luminance block of the current block (Apply filter) to generate a second chrominance prediction block of the current block. Thereafter, the video signal processing device can add the chrominance residual block to the second chrominance prediction block of the current block predicted using the linear and nonlinear models to finally generate the chrominance block (Cb, Cr) of the current block. This method can be described as a cross-component residual model (CCRM). CCRM, Inter CCCM, and Inter CCP may have the same meaning.

[0211] To improve coding efficiency, rather than coding the aforementioned residual signal as is, a method may be used in which the transform coefficient values ​​obtained by transforming the residual signal are quantized and the quantized transform coefficients are coded. As described above, the transform unit may obtain the transform coefficient values ​​by transforming the residual signal. At this time, the residual signal of a specific block may be distributed across the entire region of the current block. Accordingly, by performing a frequency-domain transformation on the residual signal, energy can be concentrated in the low-frequency region, thereby improving coding efficiency.

[0212] The encoder can obtain at least one residual block containing a residual signal for the current block. The residual block may be the current block or one of the blocks split from the current block. In this specification, the residual block may be described as a residual array or a residual matrix containing residual samples of the current block. Furthermore, in this specification, the residual block may represent a block having the same size as the size of a transform unit or a transform block.

[0213] An encoder can transform a residual block using a transform kernel. The transform kernel used for transforming the residual block may be a transform kernel having separable vertical and horizontal transform properties. In this case, the transform for the residual block may be performed separately as a vertical transform and a horizontal transform. For example, the encoder may perform a vertical transform by applying a transform kernel in the vertical direction of the residual block. Additionally, the encoder may perform a horizontal transform by applying a transform kernel in the horizontal direction of the residual block. In this specification, a transform kernel may be used as a term referring to a set of parameters used for transforming a residual signal, such as a transform matrix, a transform array, a transform function, or a transform. In one embodiment, the transform kernel may be any one of a plurality of available kernels. Additionally, transform kernels based on different transform types may be used for each of the vertical and horizontal transforms. That is, before performing the first transformation, the transformation method for the vertical and horizontal directions can be derived using at least one or more of the intra prediction mode of the current block, the encoding mode, the transformation method parsed from the bitstream, and the size information of the current block. In addition, in order to reduce the computational complexity in the transformation process for large blocks, a process of processing the high-frequency region as '0' while leaving only the low-frequency region can be performed. This process is called high-frequency zeroing, and the transformation size at the time of the actual first transformation can be set for this zeroing. In the high-frequency zeroing process, the low-frequency region can be set to an arbitrary fixed size, and for example, the horizontal or vertical size can be a combination of 4, 8, 16, 32, etc.

[0214] The encoder can quantize the transform block transformed from the residual block by passing it to the quantization unit. At this time, the transform block can include multiple transform coefficients. Specifically, the transform block can be composed of multiple transform coefficients arranged in a two-dimensional array. The size of the transform block, like the residual block, can be the same as either the current block or a block split from the current block. The transform coefficients passed to the quantization unit can be expressed as quantized values.

[0215] Additionally, the encoder may perform an additional transform before the transform coefficients are quantized. The aforementioned transform method may be referred to as a primary transform, and the additional transform may be referred to as a secondary transform. The secondary transform may be optional for each residual block. In one embodiment, the encoder may perform the secondary transform for areas where it is difficult to focus energy in the low-frequency region using only the primary transform, thereby improving coding efficiency. For example, the secondary transform may be added for blocks in which the residual values ​​appear significantly in directions other than the horizontal or vertical directions of the residual block. The residual values ​​of an intra-predicted block may be more likely to change in directions other than the horizontal or vertical directions compared to the residual values ​​of an inter-predicted block. Accordingly, the encoder may additionally perform the secondary transform on the residual signal of the intra-predicted block. Additionally, the encoder may omit the secondary transform on the residual signal of an inter-predicted block. High-frequency zeroing in the first conversion can also be performed in the second conversion process.

[0216] As another example, whether to perform a secondary transform may be determined based on the size of the current block or the remaining block. Furthermore, transform kernels having different sizes may be used based on the size of the current block or the remaining block. For example, an 8X8 secondary transform may be applied to a block in which the length of a shorter side among the width or the height is greater than or equal to a first preset length. Furthermore, a 4X4 secondary transform may be applied to a block in which the length of a shorter side among the width or the height is greater than or equal to a second preset length and less than the first preset length. In this case, the first preset length may be a value greater than the second preset length, but the present disclosure is not limited thereto. Furthermore, unlike the first transform, the secondary transform may not be performed separately into a vertical transform and a horizontal transform. Such a secondary transform may be referred to as a low frequency non-separable transform (LFNST).

[0217] In addition, for video signals in a specific region, high-frequency band energy may not be reduced even if frequency transform is performed due to rapid brightness changes. Accordingly, compression performance due to quantization may deteriorate. In addition, when transform is performed on a region where residual values ​​rarely exist, encoding time and decoding time may unnecessarily increase. Accordingly, transform for the residual signal in a specific region may be omitted. Whether or not to perform transform for the residual signal in a specific region may be determined by a syntax element related to the transform of the specific region. For example, the syntax element may include transform skip information. The transform skip information may be a transform skip flag. If the transform skip information for a residual block indicates transform skip, transform for the corresponding residual block is not performed. In this case, the encoder can immediately quantize the residual signal in the region where the transform is not performed.

[0218] The aforementioned transformation-related syntax elements may be information parsed from a video signal bitstream. A decoder may entropy decode the video signal bitstream to obtain the transformation-related syntax elements. Additionally, an encoder may entropy code the transformation-related syntax elements to generate a video signal bitstream.

[0219] The decoder can obtain encoding information necessary for decoding by parsing the transmitted bitstream. At this time, information related to the transformation process includes index information for the first and second transformation types and quantized transformation coefficients. The inverse transformation unit can obtain a residual signal by inversely transforming the inverse quantized transformation coefficients. First, the inverse transformation unit can detect whether an inverse transformation is performed for a specific region from a transformation-related syntax element of the specific region. According to one embodiment, if a transformation-related syntax element for a specific transformation block indicates a transformation skip, the transformation for the corresponding transformation block may be skipped. In this case, both the first inverse transformation and the second inverse transformation for the transformation block may be skipped. In addition, the inverse quantized transformation coefficients can be used as a residual signal. For example, the decoder can use the inverse quantized transformation coefficients as a residual signal to reconstruct the current block. Alternatively, the second inverse transform may be performed and the first inverse transform may be omitted, and the second inverse transformed value may be used as the residual signal. The first inverse transform described above represents the inverse transform for the first transform and may be referred to as the inverse primary transform. The second inverse transform represents the inverse transform for the second transform and may be referred to as the inverse secondary transform or the inverse LFNST. In the present invention, the first (inverse) transform may be referred to as the first (inverse) transform, and the second (inverse) transform may be referred to as the second (inverse) transform.

[0220] FIG. 16 illustrates types of transform kernels that can be used in video coding according to one embodiment of the present specification.

[0221] Fig. 16 shows the formulas of DCT-II, DCT-V (discrete cosine transform type-V), DCT-VIII (discrete cosine transform type-VIII), DST-I (discrete sine transform type-I), and DST-VII kernels applied to MTS. DCT and DST can be expressed as functions of cosine and sine, respectively, and when the basis function of the transform kernel for the number of samples N is expressed as Ti(j), the index i represents the index in the frequency domain, and the index j represents the index within the basis function. That is, as i becomes smaller, it represents a low-frequency basis function, and as i becomes larger, it represents a high-frequency basis function. When the basis function Ti(j) is expressed as a two-dimensional matrix, it can represent the j-th element of the i-th row, and since the transform kernels illustrated in Fig. 16 all have a separable characteristic, they can perform transformations in the horizontal and vertical directions respectively for the residual signal X. That is, when the residual signal block is X and the transform kernel matrix is ​​T, the transform for the residual signal X can be expressed as TXT'. Here, T' means the transpose of the transform kernel matrix T. Since DCT and DST are in decimal form, not integer form, it is burdensome to implement them as they are in hardware encoders and decoders. Therefore, the decimal form transform kernel must be approximated to an integer form transform kernel through scaling and rounding. The integer precision of the transform kernel can be determined as 8-bit or 10-bit, but if the precision is low, the encoding efficiency may decrease. Depending on the approximation, the orthonormal property of DCT and DST may not be maintained, but the encoding efficiency loss due to this is not large, so approximating the transform kernel to an integer form is advantageous in terms of implementing a hardware encoder and decoder.IDTR (Identity Transform) is a transformation whose result is the same as the original transformation. This is called an identity transformation. Typically, an identity transformation constructs a transformation matrix by setting "1" in positions where rows and columns have the same value. However, here, the identity transformation uses an arbitrary fixed value, not "1," to uniformly increase or decrease the value of the input residual signal.

[0222] In the above-mentioned first-order transform, the MTS transform, the transform is calculated by applying the transform kernel to the vertical and horizontal directions of the error block respectively, so it can be said to be a separable transform method. On the other hand, in the above-mentioned second-order transform, the LFNST transform, the transform is calculated by applying the transform kernel only once without applying the transform kernel to the vertical and horizontal directions respectively, so it can be said to be a non-separable transform method. In addition, the above-mentioned second-order transform is additionally applied to the first-order transformed transform coefficients of the block to which the DCT-2 transform is applied, so it can be said to be a two-step transform technique. The above-mentioned second-order transform has high encoding efficiency, but it has the disadvantage of being complex because the transform kernel is applied a total of three times. To reduce this complexity, the NSPT (Non-separable primary transform) method, which is a method of applying the transform using only the second-order transform method, can be applied. The NSPT transform method is a non-separable transform method, and is calculated by applying the transform kernel only once to the vertical and horizontal directions of the error block, without applying the transform kernel separately. In a video signal processing device, the error block of the current block can be transformed or inversely transformed using one of the following transform methods: MTS, DCT2 + LFNST, or NSPT transform.

[0223] NSPT transform can be a transform method that replaces the existing DCT2 + LFNST transform method. For blocks whose size is equal to or smaller than 16x16, one of the following kernels can be applied depending on the size of the transform block: NSPT4x4 (16x16 kernel), NSPT4x8 (32x20 kernel), NSPT8x4 (32x20 kernel), NSPT8x8 (64x32 kernel), NSPT4x16 (64x24 kernel), NSPT16x4 (64x24 kernel), NSPT8x16 (128x40 kernel), NSPT16x8 (128x40 kernel), NSPT4x32 (128x20 kernel), NSPT32x4 (128x20 kernel), NSPT8x32 (256x24 kernel), NSPT32x8 (256x24 kernel). NSPT, similar to LFNST, consists of 35 sets of transform kernels, each of which can have three candidates. The encoder can derive a set of transform kernels according to the intra prediction mode, and then generate and signal a bitstream containing information about the index of the optimal candidate among the three candidates. The decoder can parse the information about the signaled index, and then use the transform kernel candidate indicated by the index information among the set of transform kernels derived using the intra prediction mode to inversely transform the current transform coefficients and obtain a residual block. Zero-out may not be performed on a 4x4 block to which NSPT is applied. In addition, the number of coefficients to be zeroed out may vary depending on the size of the NSPT kernel. For example, NSPT with a size of 32x20 can be applied to a 4x8 block or an 8x4 block. Therefore, out of the 32 transform coefficients, only 20 transform coefficients can be zeroed out, while the remaining 12 transform coefficients can be zeroed out.

[0224] FIG. 17 illustrates a transformation set table for LFNST and NSPT transformations according to one embodiment of the present specification.

[0225] There can be 35 transform sets used in LFNST and NSPT transforms, and they can vary depending on the intra prediction mode (see FIG. 6). That is, the video signal processing device can derive the transform set index of the LFNST and NSPT transforms corresponding to the intra prediction mode (see FIG. 6) by referring to the transform set table of FIG. 17. In addition, the LFNST and NSPT transform sets can vary depending on the intra prediction mode, information on whether the current block is a luminance block or a chrominance block, the horizontal and vertical sizes of the current block, and whether the intra prediction directional mode of the current block is an extended angle mode. There can be an arbitrary number of transform matrices for each transform set. Here, the arbitrary number can be an integer greater than or equal to 1, and can be 3. The encoder can signal the index information for the optimal transform matrix among multiple transform matrices in the transform set by including it in the bitstream. After the decoder parses the index for the optimal transformation matrix, it can apply the inverse transformation using the transformation matrix corresponding to the index in the transformation set.

[0226] The encoder and decoder can apply three types of transformation kernels, LFNST4, LFNST8, and LFNST16, depending on the size of the transformation block. If the width and height of the current transformation block are greater than or equal to 16, the encoder and decoder can apply the LFNST16 transformation kernel. If the width and height of the current transformation block are greater than or equal to 8, the encoder and decoder can apply the LFNST8 transformation kernel. If the width and height of the current transformation block are less than 8, the encoder and decoder can apply the LFNST4 transformation kernel.

[0227] Figure 18 shows an example of ROI after LFNST transformation.

[0228] The gray block area in Fig. 18, which is the low-frequency part output after the encoder and decoder LFNST-convert the residual block, is referred to as the Region of Interest (ROI). Areas other than the pre-designated ROI can be zero-out processed, which is set to 0. In the case of LFNST16, only 96 samples are required. Therefore, as shown in (a) of Fig. 18, the remaining white sub-blocks except for 6 NxN sub-blocks can be zero-out processed. In the case of LFNST8, only 64 samples are required. Therefore, as shown in (b) of Fig. 18, the remaining white sub-blocks except for 4 NxN sub-blocks can be zero-out processed. In the case of LFNST4, there may not be a zero-out area. In this case, N is a positive integer and may be 4.

[0229] FIG. 19 illustrates a method for deriving a multi-transform set and a LFNST / NSPT set according to one embodiment of the present specification.

[0230] Referring to Fig. 19(a), the encoder can select which one of MTS, DCT2 + LFNST, and NSPT to apply to the residual block, and can transform the residual block and obtain transform coefficients based on the selected transform method. At this time, the encoder can signal information on which transform method was applied by including it in the bitstream. If the MTS transform is applied to the residual block, the transform coefficients of the residual block can be obtained by applying the MTS transform, and LFNST and NSPT transforms may not be applied. If the DCT2 + LFNST transform is applied to the residual block, the encoder can apply the DCT2 transform to the residual block to obtain the first transform coefficients, and apply the LFNST transform to the first transform coefficients to obtain the second transform coefficients. At this time, the MTS and NSPT transforms may not be applied. When the NSPT transform is applied to the residual block, the video encoder can output transform coefficients by applying the NSPT transform to the residual block, and at this time, the MTS and DCT2 + LFNST transforms may not be applied to the residual block.

[0231] Referring to Fig. 19(b), the decoder can parse information about which transform method was applied from the bitstream, and determine whether to apply one of the inverse transform methods among MTS, DCT2 + LFNST, and NSPT to the current transform coefficient based on the parsed information. Then, the decoder can perform an inverse transform on the transform coefficient based on the determined transform method and obtain a residual block. If the MTS transform is applied to the current transform coefficient, the decoder can perform an inverse MTS transform on the transform coefficient to obtain a residual block, and at this time, the LFNST and NSPT inverse transforms may not be applied. If the DCT2 + LFNST transform is applied to the secondary transform coefficient, the decoder can perform an inverse LFNST transform on the secondary transform coefficient to output a primary transform coefficient, and perform an inverse DCT2 transform on the primary transform coefficient to obtain a residual block, and at this time, the MTS and NSPT transforms may not be applied. If the NSPT transform is applied to the current transform coefficients, the decoder can obtain a residual block by performing the NSPT inverse transform on the current transform coefficients, and at this time, the MTS and DCT2 + LFNST transforms may not be applied.

[0232] A video signal processing device can derive a transform kernel for each of MTS, LFNST, and NSPT transforms (or inverse transforms) using an intra prediction mode. In addition, the video signal processing device can determine which transform (or inverse transform) is applied among MTS, DCT2 + LFNST, and NSPT. At this time, in order to determine the transform, at least one or more of the horizontal and vertical sizes of the current block, whether the components of the current block are luminance components or chrominance components, whether the current block is a single tree or a dual tree, information on whether the current block is encoded in intra mode or inter mode, information on the encoding mode of the current block (e.g., IBC, Intra TMP, Merge, AMVP, GPM, SGPM, CCLM, CCCM), and information on the quantization parameters of the current block may be used.

[0233] Figure 20 illustrates a mapping table according to one embodiment of the present specification.

[0234] Figure 21 illustrates a conversion type set table according to one embodiment of the present specification.

[0235] Figure 22 illustrates a conversion type combination table according to one embodiment of the present specification.

[0236] FIG. 23 illustrates a threshold value table for an IDT conversion type according to one embodiment of the present specification.

[0237] A method for selecting a set of multiple transforms available to the current block in a video signal processing device is described.

[0238] 1) First, the video signal processing device can derive the nSzIdxW and nSzIdxH values ​​based on the size of the current block to map the horizontal and vertical sizes of the current block into a single variable. nSzIdxW can be the minimum value among the logarithm of 2 calculated for the width of the current block, the difference of 2 with the decimal places discarded, and the three. nSzIdxH can be the minimum value among the logarithm of 2 calculated for the height of the current block, the difference of 2 with the decimal places discarded, and the three.

[0239] 2) Next, the video signal processing device can derive the intra-directional mode (predMode) of the current block. In the case of TIMD mode, the number of intra-prediction modes expanded from the existing 67 to 131 can be used, reducing the precision of the existing 67 modes.

[0240] 3) Next, the video signal processing device can derive the ucMode, nMdIdx, and isTrTransposed values.

[0241] A. If the current block is encoded in MIP mode, ucMode can be set to '0', nMdIdx can be set to '35', and isTrTransposed can be set to a value derived from MIP.

[0242] B. If the current block is not encoded in MIP mode, ucMode can be set to the intra directional mode (predMode) of the current block. predMode can mean an index value of the intra directional mode. predMode can be determined through an extended angle mode according to the aspect ratio of the current block. The video signal processing device can clip predMode to a value between 2 and 66. If predMode is greater than the 34th angle mode, which is a diagonal mode, the isTrTransposed value can be set to 1, and if predMode is less than or equal to 34, the isTrTransposed value can be set to 0. If predMode is greater than 34, the video signal processing device resets the value of predMode to a value that is a difference between 67 (the maximum value of the intra directional mode index) plus 1. For example, if predMode is 35, the video signal processing device can reset the 35th angle mode to the 33rd angle mode (67+1-35(predMode)). If predMode is 66, the video signal processing device can reset the 66th angle mode to the 2nd angle mode (67+1-66(predMode)). That is, by making it symmetrical with respect to the 34th angle mode, which is a diagonal mode, there is an effect of reducing the size of the transformation mapping table of Fig. 35 by about half.

[0243] 4) The video signal processing device can derive the nSzIdx value through nSzIdxW, nSzIdxH, and isTrTransposed values. If the isTrTransposed value is '1', the value obtained by multiplying nSzIdxH by 4 and then adding nSzIdxW can be set to nSzIdx. If the isTrTransposed value is '0', the value obtained by multiplying nSzIdxW by 4 and then adding nSzIdxH can be set to nSzIdx.

[0244] 5) The video signal processing device can derive nTrSet, which is an index of an available transformation type set, according to the predefined table of FIG. 20 using nSzIdx, which is the size information of the current block, and nMdIdx, which is the intra-directional mode information of the current block. FIG. 20 defines an index of a transformation type set according to the intra-screen directional mode (0 to 34 and MIP) of the current block and the size index (0 to 15) of the current block. Referring to FIG. 20, nTrSet can be 80, and if the size of the current block is 4x8 and the intra-screen directional mode of the current block is 13, nTrSet can be '7'.

[0245] 6) The video signal processing device can derive a set of transformation types corresponding to nTrSet from the table of FIG. 21 by parsing mts_idx included in the bitstream. The transformation types in the vertical and horizontal directions are set differently depending on whether predMode is a value greater than the 34th angle mode, which is a diagonal mode. 0 to 79 in the vertical column of FIG. 21 may correspond to nTrSet, and 0 to 3 in the horizontal column may correspond to mts_idx. Referring to FIG. 21, if nTrSet is 7 and the value of mts_idx is 3, 22 may be selected from (2, 17, 18, 22). In addition, DST1 and DCT5 corresponding to index 22 of the transformation type combination table of FIG. 22 are selected, and the vertical transformation type of the current block may be set to DST1 and the horizontal transformation type may be set to DCT5. 0 to 24 in the horizontal column of FIG. 22 are indices selected through FIG. 21, and 0 to 1 in the vertical column may represent vertical and horizontal transformation types, respectively. If the predicted directionality mode within the screen of the current block is greater than 34, which is a diagonal mode, the vertical and horizontal transformation types may be interchanged.

[0246] If the mts_idx value is '3' and the width and height of the current block are both 16 or less, the vertical or horizontal transformation type can be reset to the IDT transformation type through the process described below.

[0247] If the absolute difference between the index of the on-screen prediction directional mode of the current block and the index of the horizontal mode, 18, is less than any predetermined value, the vertical transformation type may be reset to the IDT transformation type. If the absolute difference between the index of the on-screen prediction directional mode of the current block and the index of the horizontal mode, 50, is less than any predetermined value, the horizontal transformation type may be reset to the IDT transformation type. At this time, the predetermined value is an integer and may be determined based on the horizontal or vertical size of the current block. For example, the predetermined value may be determined through the table of Fig. 20. The table of Fig. 23 (a) shows a case where the threshold value is set differently whenever the horizontal or vertical size differs by 4, and the table of Fig. 23 (b) shows a case where the threshold value is set differently whenever the horizontal or vertical size differs by double. If the size of the current block is 16x16, the video signal processing device may not reset the vertical transformation type to the IDT transformation type, but may maintain the existing transformation type as is.

[0248] The encoder and decoder can reduce the amount of bits for encoding mts_idx by adaptively changing the number of transform type sets for each block. The encoder and decoder can determine the number of transform type sets to be one, four, or six, depending on the sum of the absolute values ​​of the transform coefficients of the current transform block. If the number of transform type sets changes, the maximum number of bins for signaling mts_idx may change. If the sum of the absolute values ​​of the transform coefficients is less than or equal to 6, the number of transform type sets is one. Therefore, the encoder may not signal mts_idx. In this case, the decoder may not parse mts_idx. If the sum of the absolute values ​​of the transform coefficients is greater than 6 and less than or equal to 32, the number of transform type sets may be four. If the sum of the absolute values ​​of the transform coefficients is greater than 32, the number of transform kernel candidates may be six. At this time, the encoder can signal mts_idx based on the maximum number of transformation type sets. Additionally, the decoder can parse mts_idx based on the maximum number of transformation type sets.

[0249] When the current block is predicted in inter coding mode, the encoder and decoder can use one of four transform kernel sets: {(DST7, DST7), (DST7, DCT8), (DCT8, DST7), (DCT8, DCT8)}. The encoder can signal an index of which transform kernel set it used. The decoder can parse the index to determine which transform kernel set to use for the current transform block. When the configured transform kernel set is (DST7, DCT8), the encoder and decoder transform or inversely transform the current block using DST7 in the horizontal direction, and transform or inversely transform the current block using DCT8 in the vertical direction. To optimize complexity, the encoder and decoder can set the maximum CU size to which Inter MTS can be applied to 32x32 for images larger than 1920x1080 resolution. The encoder and decoder set the maximum CU size to which Inter MTS can be applied to images with a resolution other than 1920x1080 to 16x16. Therefore, the encoder and decoder can apply DCT2 transform in the horizontal and vertical directions without applying Inter MTS to transform blocks larger than 16x16, for example, transform blocks of size 32x32. In addition, the encoder and decoder can use separable KLT instead of DST7 and DCT8 in transform and inverse transform of transform blocks of size equal to or smaller than 16x16.

[0250] FIG. 24 illustrates block boundaries and samples around the boundaries in a deblocking filtering process according to one embodiment of the present specification.

[0251] Referring to (a) of Fig. 24, the part indicated by the dotted line between the P block and the Q block may denote a block boundary. The block boundary may exist at any fixed size, and a block boundary may exist at every size of 4.

[0252] Figure 24 (b) shows samples for which filtering is performed based on block boundaries. The video signal processing device can perform a process of alleviating blocking artifacts occurring at block boundaries by performing deblocking filtering on the currently restored block. The deblocking filtering process can include a process of determining a transform block boundary, determining a sub-block boundary, determining a length of a filter to perform filtering, determining a filtering strength (bS), determining a filtering parameter, determining whether to perform filtering, and determining a filtering type.

[0253] In independent scalar quantization, the reconstructed coefficient t'k for an input coefficient tk depends only on the associated quantization index qk. That is, the quantization index for any reconstructed coefficient has a different value from the quantization indices for other reconstructed coefficients. Here, t'k may be a value of tk that includes quantization error, and may be different or the same depending on the quantization parameter. Here, t'k can be named a reconstructed transform coefficient or an inverse quantized transform coefficient, and the quantization index can also be named a quantized transform coefficient.

[0254] In Uniform Reconstruction Quantizers (URQ), the reconstructed coefficients are spaced at equal intervals. The distance between two adjacent reconstructed values ​​is referred to as the quantization step size. Reconstructed values ​​can include zeros, and the entire set of available reconstructed values ​​can be uniquely defined by the quantization step size. The quantization step size can vary depending on the quantization parameter.

[0255] In conventional methods, quantization reduces the set of acceptable reconstructed transform coefficients, and the number of elements in this set can be finite. This limits the ability to minimize the average error between the original and reconstructed images. Vector quantization can be used as a method to minimize this average error.

[0256] A simple form of vector quantization used in video encoding is sign data hiding. This is a method in which the encoder does not encode the sign of a non-zero coefficient, and the decoder determines the sign of the coefficient based on whether the sum of the absolute values ​​of all coefficients is even or odd. To achieve this, the encoder may increase or decrease the value of at least one coefficient, and at least one coefficient may be selected and adjusted to be optimal in terms of rate-distortion cost. In one implementation, a coefficient having a value close to the boundary of the quantization interval may be selected.

[0257] Another vector quantization method is trellis-coded quantization, which is used in video coding as an optimal path search technique to obtain optimized quantization values ​​in dependent quantization. Block-wise, quantization candidates for all coefficients in the block are arranged in a trellis graph, and the optimal trellis path between the optimized quantization candidates is searched for by considering the cost for rate-distortion. Specifically, dependent quantization applied to video coding can be designed so that the set of allowable reconstructed transform coefficients for a transform coefficient depends on the value of the transform coefficient preceding the current transform coefficient in the restoration order. At this time, by selectively using multiple quantizers according to the transform coefficients, the average error between the original image and the reconstructed image can be minimized, thereby increasing the encoding efficiency.

[0258] Figure 25 shows two independent scalar quantizers according to an embodiment of the present invention.

[0259] A video signal processing device may require two quantizers to perform a dependent quantization method. In addition, the dependent quantization method of the video signal processing device may include a switching procedure between the two quantizers. Referring to FIG. 25, two quantizers (Q0, Q1) may be defined to perform the dependent quantization method. In uniform restoration quantization, all restored transform coefficients t' k is the quantization step size (Δ k ) can be distributed at regular intervals according to integer multiples of t' k can be the restored transform coefficients or the inverse quantized transform coefficients. The quantizer Q0 is a quantization step size (Δ k ) of an even multiple (e.g., -4Δ k , -2Δ k , 0, 2Δk , 4Δ k ) can be included, and the quantizer Q1 has a quantization step size (Δ k ) odd multiples (e.g., -3Δ k , -Δ k , 0, Δ k , 3Δ k ) can include. Both Q0 and Q1 quantizers can include restored transform coefficients having a value of '0'. The values ​​shown on the circle, which is the vertical coordinate of Fig. 25, are the quantization indices (q) for the restored transform coefficients. k ) means one quantization index (q k ) can be mapped to the value of one of the two restored transform coefficients. For example, the quantization index (q k ) -1 is the restored transform coefficient of the quantizer Q0 -2Δ k or the restored transformation coefficient of Q1 -Δ k can be mapped to the quantization index (q k ) may be a quantized transform coefficient, an encoded transform coefficient, or a transform coefficient obtained by a video signal processing device parsing a bitstream.

[0260] Unlike the existing independent scalar quantization, the transform coefficients within a block can be restored using a preset scan order. The preset scan order can be the restoration order in which the transform coefficients within a block are restored. In this case, the restoration order is determined by the quantization index (q). k ) may be the same as the order in which they are encoded or decoded. Under a given restoration order, the transition procedure between the two quantizers is 2 N It can be described by a state (N>=2) transition model. The video signal processing device currently has a transform coefficient t to be restored. k Status s for k Depending on t, we can decide which quantizer to use for the dependent quantization method. k The next transformation coefficient to be restored is t k+1 Status s fork+1 is the current state s k and the current quantization index q k Parity p calculated by k can be determined by p k is q k & can be the result value of 1, i.e. p k = q k & can be 1. & in this specification represents the bitwise operator AND, and outputs 1 if both the left operand and the right operand are odd, and outputs 0 if either the left operand or the right operand is even. q k If p is 1 k is 1, and q k If is 0, pk can be 0.

[0261] Figure 26 shows a state transition table according to an embodiment of the present invention.

[0262] Fig. 26 shows a quantizer selected according to a current state and a next state transitioned from the current state. Referring to Fig. 26, a video signal processing device can select a quantizer according to the current state. In addition, the video signal processing device can determine the next state according to a quantization index qk. For example, if the current state is 0, the video signal processing device can select a quantizer Q0. At this time, if the quantization index qk is even, the next state may be 0, and if the quantization index qk is odd, the next state may be 2. If the current state is 1, the video signal processing device can select a quantizer Q0. At this time, if the quantization index qk is even, the next state may be 2, and if it is odd, the next state may be 0. If the current state is 2, the video signal processing device can select a quantizer Q1. At this time, if the quantization index qk is even, the next state may be 1, and if it is odd, the next state may be 3. When the current state is 3, the video signal processing device can select the quantizer Q1. At this time, if the quantization index qk is even, the next state can be 3, and if it is odd, the next state can be 1. When dependent quantization is first performed, the video signal processing device can set the initial value of the state to 0.

[0263] Figure 27 shows a state transition graph according to an embodiment of the present invention.

[0264] Figure 27 is a diagram showing the state transition table of Figure 26 in graph form. The numbers (0, 1, 2, 3) in the circles of Figure 27 can represent the current state. t k The next transformation coefficient to be restored is t k+1 Status s for k+1 is the current state s k and the current quantization index q k Parity p calculated by kcan be determined by. Referring to Fig. 27, when the current state is 0, 1, the video signal processing device can select the quantizer Q0. In addition, when the current state is 2, 3, the video signal processing device can select the quantizer Q1. s k+1 can be expressed as 'QStateTransTable[ ][ ] = { 0, 2}, { 2, 0}, { 1, 3}, { 3, 1}}'. {0, 2} is the current state s k can mean a possible next state candidate when {2, 0} is the current state s k can mean the possible next state candidates when {1, 3} is the current state s k can mean the possible next state candidates when {3, 1} is the current state s k It can mean a possible next state candidate when 3. The possible next state candidate {x, y} in Table 1 can be {when the quantization index is even, when the quantization index is odd}.

[0265] Figure 28 illustrates a process of restoring a conversion coefficient according to an embodiment of the present invention.

[0266] Below, the process of restoring the conversion coefficients is described with reference to Fig. 28.

[0267] The video signal processing device has a quantization index q k can be obtained. The video signal processing device is currently in state s k can be initialized to 0. The video signal processing device is currently in state s k The quantizer can be selected according to the following. For example, a video signal processing device may be s k ≫ If the result value of 1 is 0, select the quantizer Q0, and s k≫ If the result of 1 is 1, the quantizer Q1 can be selected. '≫' is a bit shift operator that outputs a value by repeatedly dividing by 2 as many times as the number on the right. The video signal processing device uses the selected quantizer to t' k can be calculated and can be performed through two steps. First, the video signal processing device calculates the quantization index q k q' for quantization step size Δk by multiplying by integer. k can be calculated. At this time, the integer is an integer greater than or equal to 2 and can be a multiple of 2. For example, q' k is q k The value multiplied by 2 (integer) and (s) k ≫ 1) q in the result value k It can be a value obtained by subtracting the value multiplied by the sign of . Next, the video signal processing device can use uniform restoration quantization, and q' k Multiply by Δk to get t' k can be calculated. The video signal processing device is s k Wow q k Using , the state (s) for the transformation coefficients to be restored next k+1 ) can be updated. At this time, the state (s) for the transformation coefficients to be restored next k+1 ) can be updated according to Fig. 26 and Fig. 27. The video signal processing device can check whether all coefficients in the block have been restored. If all coefficients in the block have not been restored, the video signal processing device performs the quantizer selection step again to select the next transform coefficient to be restored (t k+1 ) can be restored. The video signal processing device can restore all transform coefficients (t') within the restored block if all coefficients within the block are restored. k ) can be output. The video signal processing device can restore the transform coefficients corresponding to all quantization indices in the current block. At this time, the restoration order for restoring the transform coefficients may be a preset restoration order.

[0268] Figure 29 shows an optimal trellis path according to one embodiment of the present invention.

[0269] For quantization, quantization candidates for all coefficients in a block can be arranged in a trellis graph. An optimal trellis path may be required for optimized quantization. At this time, the Viterbi algorithm can be used as a method for finding the optimal trellis path. The Viterbi algorithm can be a technique for finding the optimal trellis path between optimized quantization candidates in the encoder by considering the cost for rate-distortion. Figure 29 shows the optimal trellis path found using the Viterbi algorithm. The encoder can add an uncoded state to the four states described in Figures 26 and 27. The uncoded state can be used for a transform coefficient that is restored before the first non-zero transform coefficient among the transform coefficients restored according to the restoration order. Referring to Figure 29, the first non-zero transform coefficient may be a transform coefficient with a scan index of 2. At this time, the transform coefficients (scan index 0, 1) before the first non-zero transform coefficient can be set to the uncoded state so as not to affect the state transition. The dotted line path connecting each node in Fig. 29 may be the optimal trellis path searched by the encoder. S in Fig. 29 may represent a state. Information about the first non-zero transform coefficient can be included in the bitstream generated by the encoder.

[0270] Fig. 30 shows a restoration order of quantized transform coefficients according to an embodiment of the present invention.

[0271] Referring to FIG. 30, the current block may be 16 x 16 in size, and the arrows in (a) and (b) of FIG. 30 may indicate the order (direction) in which the decoder restores the quantized transform coefficients for the current block. The restoration order of the quantized transform coefficients may be the order in which the encoder performs encoding. The restoration order described in FIG. 30 may be a preset restoration order described in the present specification. (a) of FIG. 30 shows the restoration order for each sub-block when the current block is divided into 4x4 sub-blocks. (b) of FIG. 30 shows the restoration order for each coefficient within one of the divided sub-blocks. The restoration order described in (a) and (b) of FIG. 30 may be referred to as a diagonal scan order. (c) of FIG. 30 shows the uncoded area and the position of the first non-zero quantized transform coefficient. The gray area in (c) of Fig. 30 may be an area where coefficients are not encoded. An entirely gray sub-block may indicate that not all coefficients in the corresponding sub-block are encoded. Additionally, a black sub-block may be the position of the first non-zero quantized transform coefficient. Referring to (c) of Fig. 30, coefficients located before the first non-zero quantized transform coefficient in the restoration order are not encoded (because the quantized transform coefficient is 0) and may therefore be indicated as a gray area.

[0272] FIG. 31 illustrates a method for updating the state of syntax elements and quantized transform coefficients according to one embodiment of the present invention.

[0273] Figure 31 illustrates the syntax elements related to each quantized transform coefficient within a sub-block parsed according to the restoration order and the order in which the status of each quantized transform coefficient is updated. The method by which a video signal processing device parses the absolute value of a coefficient may include two methods. In the first method, after the video signal processing device parses multiple syntax elements, the video signal processing device can obtain the final coefficient through a combination of the parsed syntaxes. In the second method, the video signal processing device can obtain the final coefficient through a single syntax element. The first method has high encoding efficiency, but the parsing processing speed may be slow due to the process of deriving and updating a context model for parsing. The second method has high processing speed, but the encoding efficiency may be low. Therefore, the video signal processing device may use a combination of the first and second methods. The maximum number of bins that can be allowed in the current block may be limited. For example, if the number of currently used bins is less than or equal to a predetermined threshold, the video signal processing device may use the first method. If the number of currently used bins is greater than the threshold, the video signal processing device may use the second method. Here, the threshold may be a value determined based on the horizontal and vertical sizes of the current block.

[0274] Below, a method for parsing syntax elements related to coefficients within a sub-block is described. Each syntax element related to coefficients within a sub-block can be parsed in the restoration order (scan order) described in Fig. 30. C of Fig. 31 kis an index of coefficients of a sub-block within a current block according to an embodiment of the present invention. Referring to FIG. 31, a sub-block may include 16 coefficients in a 4 x 4 size. At this time, each coefficient of the sub-block may be indexed. k is an index related to the restoration order. The index of the coefficient restored first according to the restoration order (scan order) is C 15 , the index of the coefficient to be restored next is C 14 , the index of the coefficient to be restored next is C 13 And the index of the coefficient to be restored last may be C0.

[0275] First, the video signal processing device can perform the 'Pass 1' process of Fig. 31 (a first method of parsing multiple syntax elements and then obtaining a final coefficient by combining the syntaxes). Specifically, the video signal processing device can parse the syntax elements (sig_coeff_flag[k], abs_level_gtx_flag[k][0], par_level_flag[k], abs_level_gtx_flag[k][1]) for the 'Pass 1' process and parse each coefficient within the sub-block according to the restoration order. The video signal processing device can obtain the absolute value of the quantization index (or the quantized transform coefficient) by parsing the syntax elements (sig_coeff_flag[k], abs_level_gtx_flag[k][0], par_level_flag[k], abs_level_gtx_flag[k][1], abs_remainder[k]) related to the coefficients in the sub-block. The video signal processing device can obtain the absolute value of the quantization index or the absolute value of the quantized transform coefficient by parsing dec_abs_level[k]. At this time, the current state Sk can be used to derive a context model for sig_coeff_flag[k]. The video signal processing device can check the parity for the current quantization index based on the result of parsing the syntax elements of Pass 1. The video signal processing device can update the state used to parse the syntax element sig_coeff_flag[k] of the coefficient to be restored next based on the parity for the current quantization index. The video signal processing device can determine whether the number of currently used bins is greater than a threshold value. At this time, if the number of currently used bins is greater than the threshold value, the video signal processing device can terminate the process of Pass 1 and perform the process of Pass 2.If the number of currently used bins is less than or equal to a threshold, a 'Pass 1' process for the coefficient to be restored next can be performed according to the restoration order. Referring to FIG. 31, when the syntax element for the Pass 1 process of C2 is parsed, the number of currently used bins can be greater than the threshold. The video signal processing device can use the second method, a method of parsing one syntax element for restoration of the coefficients (C1, C0) to be restored next in the restoration order. At this time, the video signal processing device can parse one syntax element for restoration of C1, C0 in a bypass mode without a context model.

[0276] The video signal processing device can perform the second method of obtaining the final coefficient through one syntax element in the Pass 2 process. The video signal processing device can parse the remaining syntax element abs_remainder[k] for each coefficient restored in the Pass 1 process. At this time, the video signal processing device can parse abs_remainder[k] up to the point where the number of bins less than or equal to the threshold value of the previous Pass 1 process is used. Referring to FIG. 31, the video signal processing device C 15 Only abs_remainder[k] can be parsed from C2 to C3. At this time, the video signal processing device can parse abs_remainder[k] in bypass mode without a context model. After parsing abs_remainder[k], the video signal processing device can parse syntax element dec_abs_level[k] related to coefficients for which the number of bins greater than the threshold is used. At this time, the video signal processing device can parse dec_abs_level[k] in bypass mode without a context model. For binarization of dec_abs_level[k], the video signal processing device can use the current state s kcan be used. The video signal processing device can determine the parity of the current coefficient based on the result of parsing dec_abs_level[k]. The video signal processing device can update the state used for the coefficient to be restored next based on the parity of the current coefficient. The video signal processing device can obtain information about the sign of the coefficient by parsing coeff_sign_flag. The video signal processing device can parse coeff_sign_flag associated with all coefficients whose value of sig_coeff_flag is 1 within the subblock.

[0277] sig_coeff_flag[k] in FIG. 31 may be a syntax element indicating whether the kth coefficient in the restoration order is a non-zero coefficient. For example, when the value of sig_coeff_flag[k] is 1, sig_coeff_flag[k] may indicate that the kth coefficient is a non-zero coefficient. abs_level_gtx_flag[k][0] may be a syntax element indicating whether the kth coefficient in the restoration order is greater than 1. When the value of abs_level_gtx_flag[k][0] is 1, abs_level_gtx_flag[k][0] may indicate that the kth coefficient is greater than 1. par_level_flag[k] may be a syntax element indicating the parity of the kth coefficient. When the value of par_level_flag[k] is 1, par_level_flag[k] may indicate that the k-th coefficient is odd. abs_level_gtx_flag[k][1] may be a syntax element indicating whether the k-th coefficient in the restoration order is greater than 3. When the value of abs_level_gtx_flag[k][1] is 1, abs_level_gtx_flag[k][1] may indicate that the k-th coefficient is greater than 3. abs_remainder[k] may be a syntax element indicating the magnitude of the remaining absolute value of the k-th coefficient in the restoration order. dec_abs_level[k] may be a syntax element indicating the magnitude of the absolute value of the k-th coefficient in the restoration order. coeff_sign_flag[k] may be a syntax element indicating the sign (+, -) of the k-th coefficient in the restoration order.

[0278] The video signal processing device can obtain a quantized transform coefficient using the parsed syntax element. If the dependent quantization method is not used (if independent quantization is used), the video signal processing device can obtain a quantized transform coefficient by applying a sign obtained by parsing coeff_sign_flag to a quantization index obtained using the parsed syntax element. If the dependent quantization method is used, the video signal processing device can obtain a quantized transform coefficient using Equation 1. In Equation 1, TransCoeffLevel is a quantized transform coefficient, xC and yC are horizontal and vertical positions within a sub-block, AbsLevel is a quantization index, QState is state information, and coeff_sig_flag is sign information for the currently quantized transform coefficient. In addition, in Equation 1, (condition 1: 2) is a formula in which, if the condition is true, 1 is performed, and if it is false, 2 is performed. In addition, the state information used when calculating the current quantized transform coefficient can be updated using the current quantization index and then used when calculating the next quantized transform coefficient.

[0279] Mathematical formula 1

[0280] TransCoeffLevel[xC][yC] = (2 * AbsLevel[xC][yC] - (QState > 1 1 : 0)) * (1 - 2 * coeff_sign_flag[xC][yC])

[0281] The video signal device can dequantize the quantized transform coefficient obtained according to the above embodiment and output the dequantized transform coefficient. The video signal processing device can obtain the dequantized transform coefficient d[x][y] through component-wise multiplication between dz[x][y] and ls[x][y] as in Equation 2. In Equation 2, dz[x][y] may be a quantization index as TransCoeffLevel, and ls[x][y] is a scale factor. The video signal processing device can obtain the bdOffset value through Equation 3. The video signal processing device can obtain the bdShift value through Equation 4. x and y are position values ​​of samples within a sub-block. The video signal processing device can obtain rectNonTsFlag in Equation 4 through Equation 5. In Equation 5, nTbW is the horizontal size of the current transform block, nTbH is the vertical size of the current transform block, and Log2(X) is the value of the logarithm function with base 2 for the X value. In Equation 4, BitDepth is the bit depth used to represent the current sample (for example, in the case of 10 bits, the range of representing the sample can be 0 to 511). sh_dep_quant_used_flag is 1 if the current slice uses the dependent quantization method, and 0 if the dependent quantization method is not used. The dequantized transform coefficients d[x][y] can be clipped to have a value between '-(1<<15)' and '(1<<15)-1'. That is, if the d[x][y] value is less than '-(1<<15)', it is changed to '-(1<<15)', and if the d[x][y] value is greater than '(1<<15)-1', it is changed to '(1<<15)-1'.

[0282] Mathematical formula 2

[0283] d[x][y] = (dz[x][y] * ls[x][y] + bdOffset) >> bdShift

[0284] Mathematical formula 3

[0285] bdOffset = (1<<bdShift) > > 1

[0286] Mathematical formula 4

[0287] bdShift = BitDepth + rectNonTsFlag + ( ( Log2( nTbW ) + Log2( nTbH ) ) / 2 ) - 5 + sh_dep_quant_used_flag

[0288] Mathematical Formula 5

[0289] rectNonTsFlag = ( ( ( Log2(nTbW) + Log2(nTbH) ) & 1 ) = = 1 ) 1:0

[0290] In mathematical expression 2, ls[x][y] is a scale factor. When the dependent quantization method is used, the video signal processing device can obtain the value of ls[x][y] using mathematical expression 6. When the dependent quantization method is not used, the video signal processing device can obtain the value of ls[x][y] using mathematical expression 7. In mathematical expressions 6 and 7, m[x][y] is a quantization matrix. The quantization matrix is ​​set to 16 when at least one of the following is satisfied: when transform skip is applied to the current transform block, when a user quantization matrix is ​​not used, and when LFNST is applied to the current block. Otherwise, the video signal processing device parses the quantization matrix included in the bitstream to set the value of the quantization matrix. At this time, levelScale[j][k] is { { 40, 45, 51, 57, 64, 72}, { 57, 64, 72, 80, 90, 102}}. When rectNonTsFlag is 0, levelScale[j][k] is { 40, 45, 51, 57, 64, 72}, and for k values ​​0, 1, 2, 3, 4, and 5, respectively, the values ​​40, 45, 51, 57, 64, and 72 are used as levelScale values. When rectNonTsFlag is 1, levelScale[j][k] uses { 57, 64, 72, 80, 90, 102}, and for k values ​​0, 1, 2, 3, 4, 5, the values ​​57, 64, 72, 80, 90, 102 are used as levelScale values, respectively. In addition, qP is a quantization parameter.

[0291] Mathematical formula 6

[0292] ls[ x ][ y ] = ( m[ x ][ y ] * levelScale[ rectNonTsFlag ][ ( qP + 1 ) % 6 ] ) << ( ( qP + 1 ) / 6 )

[0293] Mathematical formula 7

[0294] ls[ x ][ y ] = ( m[ x ][ y ] * levelScale[ rectNonTsFlag ][ qP % 6 ] ) << ( qP / 6 )

[0295] Figures 32 to 34 illustrate state transition graphs according to one embodiment of the present invention.

[0296] The dependent quantization method described in this specification may require a process of searching for an optimal trellis path between optimized quantization candidates through four states. To search for an optimal trellis path, the video signal processing device may expand the current state to eight to diversify the trellis path. Figure 32 is a diagram showing a transition procedure for eight states. In Figure 32, k represents a quantization index q. k It can be. The transition procedure of Fig. 32 is configured so that when transitioning from the current state to the next state, when the parity of the k value in the current state is 0, the quantizer for the next state and when the parity of the k value in the current state is 1 are both the same. The quantizer (Q0, Q1) used according to the parity of the k value in the current state is not selected, and the video signal processing device uses a predetermined quantizer. For example, when the current state is 0, the next state can be 0 or 2. At this time, the video signal processing device uses the same quantizer Q0 for the state of 0 or 2. Similarly, when the current state is 1, the next state can be 5 or 7. At this time, the video signal processing device uses the same quantizer Q1 for the state of 5 or 7. In other words, although transition to the next state can be made depending on the k value, the video signal processing device uses the same quantizer for the candidates that can be transitioned to the next state. The next state s for the state transition of Fig. 32 k+1can be expressed as 'QStateTransTable[ ][ ] = { 0, 2}, { 5, 7}, { 1, 3}, { 6, 4}, { 2, 0}, { 4, 6}, { 3, 1}, { 7, 5}}'. {0, 2} is the current state s k {5, 7} can be a possible next state candidate when is 0. The current state s k When {1, 3} is 1, it can be a possible next state candidate. {1, 3} is the current state s k {6, 4} can be a possible next state candidate when s is 2. k When {2, 0} is 3, it can be a possible next state candidate. {2, 0} is the current state s k When {4, 6} is 4, it can be a possible next state candidate. {4, 6} is the current state s k When {3, 1} is 5, it can be a possible next state candidate. {3, 1} is the current state s k {7, 5} can be a possible next state candidate when s is 6. k When the quantization index is 7, it can be a possible next state candidate. The possible next state candidate {x, y} of 'QStateTransTable' can represent {when the quantization index is even, when the quantization index is odd}. The video signal processing device can use the same quantizer for the possible next state candidate {x, y} of QStateTransTable.

[0297] A video signal processing device may require a transition procedure for dynamically selecting a quantizer. Referring to FIG. 33, the encoder may select a quantizer to be used in the next state based on the parity of the k value. In addition, the encoder may select a predetermined quantizer regardless of the k value. Referring to FIG. 33, when the current state is 0, 1, 6, 7 (states marked with an X), the video signal processing device may use the same predetermined quantizer for the next state regardless of the parity of the k value. On the other hand, when the current state is 2, 3, 4, 5 (states not marked with an X), the video signal processing device may selectively use quantizers Q0 and Q1 for the next state based on the parity of the k value. Specifically, the video signal processing device may determine whether a predetermined quantizer is used based on the current state information and whether a quantizer is selected and used based on the parity of the k value. The next state s for the state transition of FIG. 33 k+1 can be expressed as QStateTransTable[ ][ ] = { { 0, 2}, { 3, 7}, { 1, 4}, { 6, 5}, { 5, 6}, { 4, 1}, { 2, 0}, { 7, 5}}. {0, 2} is the current state s k There are possible next state candidates when {3, 7} is the current state s k {1, 4} can be a possible next state candidate when s is 1. {1, 4} is the current state s k {6, 5} can be a possible next state candidate when s is 2. {6, 5} is the current state s k {5, 6} can be a possible next state candidate when s is 3. {5, 6} is the current state s k {4, 1} can be a possible next state candidate when s is 4. {4, 1} is the current state s k {2, 0} can be a possible next state candidate when s is 5. k{7, 5} can be a possible next state candidate when s is 6. {7, 5} is the current state s k can be a possible next state candidate when the quantization index is 7. The next state {x, y} of QStateTransTable can be {when the quantization index is even, when the quantization index is odd}. Referring to FIG. 33 and QStateTransTable, when the current state is 0, 1, 6, or 7, the video signal processing device can use the same quantizer as the quantizer used for the current state for the next state {x, y}. In addition, when the current state is 2, 3, 4, or 5, the video signal processing device can use a different quantizer than the quantizer used for the current state for the next state {x, y}.

[0298] Expanding the current state to eight and diversifying the trellis paths can lead to increased complexity. Below, we describe a method to address this complexity.

[0299] As shown in Fig. 34, the complexity problem can be solved using only two states. The transition procedure can be changed in various ways depending on the parity of the k value. Referring to Fig. 34 (a), if the parity of the k value is 0, the video signal processing device can select the current state as the next state. At this time, if the parity of the k value is 1, the video signal processing device can select the next state as a state different from the current state. In Fig. 34 (b), if the current state is 0, the video signal processing device follows the same transition procedure as Fig. 34 (a). If the current state is 1, when the parity of the k value is 0, the video signal processing device selects the next state as a state different from the current state. In addition, if the parity of the k value is 1, the video signal processing device can select the next state as the current state. In this way, various transition procedures can be used depending on the parity of the k value.

[0300] Next state s for state transition of (a) of Fig. 34k+1 can be expressed as QStateTransTable[ ][ ] = { { 0, 1}, { 1, 0}}. {0, 1} is the current state s k {1, 0} can be a possible next state candidate when s is 0. {1, 0} is the current state s k It can be a possible next state candidate when 1.

[0301] Next state s for state transition of (b) of Fig. 34 k+1 can be expressed as QStateTransTable[ ][ ] = { { 0, 1}, { 0, 1}}. {0, 1} is the current state s k When {0, 1} is 0, it can be a possible next state candidate. The current state s k When is 1, it can be a possible next state candidate. In QStateTransTable, the next state {x, y} can be {when the quantization index is even, when the quantization index is odd}.

[0302] An encoder can obtain a bitstream including information indicating which transition procedure among various transition procedures was used. At this time, the information indicating which transition procedure was used can be signaled at at least one level among an SPS, a PPS, a Picture header, a slice header, a Tile, and a CU of the bitstream. A decoder can parse the information indicating which transition procedure was used included in the bitstream to set a transition procedure to be used at at least one unit among an SPS, a PPS, a Picture header, a slice header, a Tile, and a CU.

[0303] When the current block is encoded at a low bit rate, the quantized transform coefficients within the current block may have a high probability of being 0. When 0 coefficients are continuously repeated, the decoder only performs a transition procedure, so the decoder efficiency is not significantly affected. However, when 0 coefficients are continuously repeated, the encoder performs cost calculation and optimal path search for quantization to 0 coefficients, which increases the complexity of the encoder and may also increase the memory required by the encoder. To reduce this unnecessary encoder complexity, the video signal processing device may use a separate transition procedure for 0 coefficients. The separate transition procedure is described below.

[0304] Figure 35 illustrates a separate transition procedure according to an embodiment of the present invention.

[0305] The video signal processing device currently has a quantization index q k can be obtained. The video signal processing device can initialize the current state. At this time, the current state to be initialized may be 0. The video signal processing device may obtain the current quantization index q k You can determine if q is 0. k If is not 0, the video signal processing device can perform the transition procedure described above through Figs. 25 to 34. The video signal processing device is currently in state s k A quantizer can be selected according to the selected quantizer. The video signal processing device uses the selected quantizer to transform the transform coefficient t' k can calculate and update the current state. q k If is 0, the video signal processing device selects the quantizer and t' kThe calculation and current state update process may not be performed. At this time, the video signal processing device can only set the current state. The video signal processing device can set the current state to a pre-specified state. The pre-specified state may be 0. The video signal processing device can determine whether all coefficients in the current block have been restored. If all coefficients have been restored, the video signal processing device can determine whether all transform coefficients (t') in the restored current block have been restored. k ) can be output. If not all coefficients have been restored, the video signal processing device selects a quantizer for the next coefficient to be restored, t' k The calculation and current state update process can be performed again.

[0306] q k If is 0, the video signal processing device can maintain the previously set state without resetting the current state. The video signal processing device can use the state previously set by the non-zero quantization index as the state for the next non-zero quantization index.

[0307] When the quantization index is 0, the video signal processing device can restore the transform coefficients using the transition procedure described above through FIGS. 25 to 34. When there are multiple or more consecutive quantization indices that are 0, the video signal processing device can set the current state to a pre-specified state. As the number of states for the dependent quantization method increases, the encoding efficiency increases, but the complexity also increases. Therefore, there may be a trade-off relationship between the encoding efficiency and the complexity. Hereinafter, a method for setting the number of states in a video signal processing device will be described.

[0308] FIG. 36 illustrates a method for setting the number of states according to a temporal layer according to an embodiment of the present invention.

[0309] Temporal layers are used to hierarchically support images from low frame rates to high frame rates. As the temporal layer number increases, the temporal resolution of the image, i.e., the frame rate of the image, increases. A video signal processing device can set the number of states based on the number of the temporal layer. The higher the number of the temporal layer, the higher the quantization parameter is used, and thus the amount of residual signal may be reduced. In a temporal layer with a higher number, there may be more coefficients with values ​​close to 0 than in a temporal layer with a lower number. Additionally, in a temporal layer with a higher number, fewer coefficients may be encoded than in a temporal layer with a lower number. When the number of the first temporal layer is greater than that of the second temporal layer, the number of states of the first temporal layer may be less than or equal to the number of states of the second temporal layer. For example, when the number of the temporal layer is 0, the number of states may be 8. When the number of the temporal layer is 1 or 2, the number of states may be 4. When the number of the temporal layer is 3, the number of states may be 2. Conversely, the higher the number of time layers, the more states there can be.

[0310] The video signal processing device can set the number of states based on the quantization parameter. For example, if the quantization parameter is less than or equal to 22, the number of states can be 8. If the quantization parameter is greater than 22 and less than or equal to 32, the number of states can be 4. If the quantization parameter is greater than 32, the number of states can be 2.

[0311] The video signal processing device can set the number of states of the current block depending on whether the current block's components are luminance blocks or chrominance blocks. For example, if the current block is a luminance block, the video signal processing device can set the number of states of the current block to 8. If the current block is a chrominance block, the video signal processing device can set the number of states of the current block to 2 or 4.

[0312] A video signal processing device can set the number of states based on syntax elements. The video signal processing device can set the number of states when the video signal processing device parses the syntax elements for the Pass 2 process to be greater than the number of states when the video signal processing device parses the syntax elements for the Pass 1 process of FIG. 31. For example, the number of states used to derive context models for sig_coeff_flag[k], abs_remainder[k], and dec_abs_level[k] may be different. The number of states used to derive the context model for sig_coeff_flag[k] may be 4. The number of states used to derive the context models for abs_remainder[k] and dec_abs_level[k] may be 8. Alternatively, the number of states used to derive the context model for sig_coeff_flag[k] may be 8. The number of states used to derive the context model for abs_remainder[k] and dec_abs_level[k] may be four. The video signal processing device may perform dependent dequantization by setting four states to the quantization index output by parsing sig_coeff_flag[k], abs_level_gtx_flag[k][0], par_level_flag[k], abs_level_gtx_flag[k][1], and abs_remainder[k]. In addition, the video signal processing device may perform dependent dequantization by setting eight states to the quantization index output by parsing dec_abs_level[k]. The number of states may be set to be less when the video signal processing device parses the syntax element for the Pass 2 process than when the video signal processing device parses the syntax element for the Pass 1 process of FIG. 31.The number of states when a syntax element for Pass 1 is parsed can be 8, and the number of states when a syntax element for Pass 2 is parsed can be 4.

[0313] FIG. 37 shows a quantization index for a block of size 8 x 8 according to an embodiment of the present invention.

[0314] Referring to Fig. 37, the quantization index qk has a characteristic of increasing from the lower right to the upper left of the block. The video signal processing device can perform the process of setting a state and selecting a quantizer using the quantization index according to the restoration order, for example, the restoration order (scan order) of Fig. 30. Referring to Fig. 37, the current block can be divided into four 4 x 4 sub-blocks, and the video signal processing device can perform restoration for each sub-block. At this time, the state information between each sub-block is connected, and the video signal processing device can use the last state information of the previous sub-block for the next sub-block. In addition, the size range of the quantization index can vary depending on the position of the sub-block. The video signal processing device can perform dependent quantization by setting the number of states differently depending on the position of the sub-block.

[0315] When a video signal processing device performs dependent quantization on the upper left sub-block, the video signal processing device can use eight states. This is because, for the upper left sub-block, the size of the quantization index is relatively large and the range of the index is relatively large. When a video signal processing device performs dependent quantization on the lower right sub-block, the video signal processing device can use two states. This is because, for the lower right sub-block, the size of the quantization index is relatively small and the range of the index value is also relatively small. In addition, the video signal processing device can use four states when performing dependent quantization on the remaining sub-blocks excluding the upper left / lower right sub-blocks.

[0316] Figures 38 and 39 illustrate a method of changing the number of states according to an embodiment of the present specification.

[0317] Fig. 38 illustrates a method for expanding the number of states according to an embodiment of the present disclosure. Since the number of states varies for each sub-block, the number of states can be expanded when a state transitions from one sub-block to another. The state transition point where the number of states is expanded can be referred to as an expansion point. When a state transitions from one sub-block, for example, sub-block 0 of Fig. 38, to another sub-block, for example, sub-block 1 of Fig. 38, the number of states can be expanded from 2 to 4. The method for expanding the number of states can reduce the complexity of the encoder by adaptively expanding the number of states.

[0318] Figure 39 illustrates a method for reducing the number of states according to an embodiment of the present disclosure. When a state transitions from one sub-block to another, the number of states may be reduced. The state transition point at which the number of states is reduced may be referred to as a reduction point. When transitioning from an arbitrary sub-block, for example, sub-block 0 of Figure 39, to another sub-block, for example, sub-block 1 of Figure 39, the number of states may be reduced from 8 to 4. A video signal processing device may reduce the number of states using a predefined rule at the reduction point. The predefined rule may be a method of leaving only odd-numbered nodes or only even-numbered nodes. Alternatively, the predefined rule may be a method of leaving only predefined node positions. If the previous node positions are 0 to 5, the predefined node positions may be 0 to 3. The method of reducing the number of states can reduce complexity by preemptively eliminating unlikely-to-be-used paths.

[0319] In addition, the video signal processing device can set the number of states based on the row and column positions of the coefficients. The row and column positions of the coefficients can be expressed in the form of coordinates (row, column). When the position of the coefficient is (i, j), the number of states can be set to x. For example, when the positions of the coefficients are (0, 0), (1, 0), (0, 1), (1, 1), the number of states can be 8, and the number of states for the remaining positions can be 4.

[0320] The encoder can include information about the number of states applied to each sub-block in the bitstream. The decoder can parse the information about the number of states included in the bitstream and determine the number of states applied to the sub-block. The decoder can perform dependent quantization based on the determined number of states.

[0321] Dependent quantization offers high encoding efficiency, but can increase encoder complexity. Therefore, the blocks to which dependent quantization is applied and the locations where it is applied can be configured in various ways.

[0322] A video signal processing device can selectively apply dependent quantization and independent quantization. An encoder can include information in a bitstream indicating whether the quantization applied to a block is dependent quantization or independent quantization. A decoder can parse the information included in the bitstream and determine whether the quantization applied to the corresponding block is dependent quantization or independent quantization. The video signal processing device can perform quantization according to the determined quantization method. Independent quantization may be quantization in which the set of allowable restored transform coefficients for a transform coefficient does not depend on the value of the transform coefficient preceding the current transform coefficient in the restoration order.

[0323] Whether dependent quantization is applied can be determined for each sub-block. That is, the lower right sub-block may have a relatively small impact on image quality as it is a high-frequency portion. The video signal processing device may perform conventional independent quantization on the lower right sub-block. The upper left sub-block may have a relatively large impact on image quality as it is a low-frequency portion. The video signal processing device may perform quantization using dependent quantization on the upper left sub-block. The encoder may generate a bitstream including information indicating whether dependent quantization is used for each sub-block. The decoder may parse the information included in the bitstream to determine the quantization to be applied to the corresponding sub-block and perform quantization according to the determined quantization method.

[0324] A video signal processing device can set a start position or an end position to which dependent quantization is applied according to a predetermined rule. Hereinafter, in the description, the start position and the end position respectively refer to the start position and the end position to which dependent quantization is applied. The video signal processing device can use both dependent quantization and independent quantization in a single block. At this time, coefficients to which dependent quantization is applied among coefficients within a single block can be consecutive in the restoration order. The video signal processing device can set the start position and the end position among coefficients to which dependent quantization is applied in the restoration order. The video signal processing device can set the start position and the end position based on at least one of the horizontal or vertical size of the current block, the quantization parameter, and the temporal layer information. Information indicating the start position and the end position respectively can be included in at least one of the SPS, the PPS, the Picture header, and the slice header. The encoder can include information indicating the start position and the end position respectively in the bitstream. The decoder can parse the information indicating the start position and the end position respectively included in the bitstream and perform quantization based on the parsed information. Information indicating the start position and the end position respectively can be set for each block. For example, the starting position of dependent quantization may be a position that is a first predetermined number of positions away from the position of the first non-zero coefficient in the restoration order. In this case, the first predetermined number may be an integer greater than or equal to 1. The ending position of dependent quantization may be a position that is a second predetermined number of positions away from the starting position in the restoration order. In this case, the second predetermined number may be an integer greater than or equal to 1. The last position at which the ending position can be set may be the upper left corner of the last block in the restoration order, (0, 0). The starting position and the ending position may each be set for each sub-block.For example, the starting position of the upper left 4 x 4 sub-block of Fig. 37 can be set to (2, 2) with an index of -6. Also, the ending position of Fig. 37 can be set to (0, 0) with a quantization index of -18. The restoration order can be the same as the restoration order described in Fig. 30.

[0325] A video signal processing device can independently apply dependent quantization to each subblock. When coefficient restoration for a subblock begins, the video signal processing device can initialize its state. Furthermore, the video signal processing device can select the optimal path for each subblock. This enables parallel processing for each subblock.

[0326] Figure 40 shows a scan order for performing dependent quantization according to an embodiment of the present invention.

[0327] The scan order for performing dependent quantization may be the same as the restoration order described in Fig. 30. According to the restoration order described in Fig. 30, blocks are scanned in order from the high-frequency part to the low-frequency part.

[0328] When an encoder performs rate-distortion optimized quantization (RDOQ), the low-frequency portion has a significant impact on the error signal. Therefore, it may be more efficient to search for the optimal path starting from the low-frequency portion than to search for the optimal path starting from the high-frequency portion. In the embodiment of Fig. 40, a scanning order that is the reverse of the scanning order described in Fig. 30 is used. Therefore, in the embodiment of Fig. 40, the video signal processing device can perform scanning in a diagonal order from the upper left to the upper right of the block.

[0329] The video signal processing device can determine a scanning order by selecting one of horizontal scan, vertical scan, and diagonal scan using at least one of the horizontal or vertical size of the current block, the horizontal and vertical size ratio of the current block, and the intra prediction direction mode. For example, if the horizontal size of the current block is greater than the vertical size, the video signal processing device can select vertical scan.

[0330] Since the current block is likely to be similar to the surrounding blocks, the quantization characteristics of the transform coefficients of the current block are likely to be similar to those of the surrounding blocks. Therefore, the video signal processing device can use the reconstruction values ​​for the transform coefficients already determined in the surrounding blocks to determine the reconstruction values ​​of the transform coefficients of the current block. In order to respond sensitively to changes in the characteristics of the surrounding blocks and the current block, a history-based state selection method or a method for selecting the number of states can be applied on a block-by-block basis or a transform block-by-block basis. For example, if the number of states used in the surrounding blocks is 8, the encoder and decoder can set the number of states of the current block to 8. Alternatively, the encoder and decoder can encode or decode the transform coefficients of the current block using the transition procedure used in the surrounding blocks (the transition procedure described above with reference to FIGS. 25 to 34).

[0331] The dequantized transform coefficients can be distributed at regular intervals proportional to the quantization index (or quantized transform coefficients). By adjusting the dequantized transform coefficients distributed at regular intervals, quantization distortion can be minimized. Shifting quantization centers can minimize quantization distortion by adjusting the distribution of the dequantized transform coefficients. This can increase encoding efficiency.

[0332] Figure 41 shows an example of changing the inverse quantized transform coefficients.

[0333] In Fig. 41, the horizontal axis coordinates represent the inverse quantized transform coefficients, which may be the same as the horizontal axis coordinates of Fig. 25. The inverse quantized transform coefficients may be expressed as the product of the N value calculated from the quantization index and the scale factor Δ. In addition, they may be distributed left and right with 0 as the standard. Fig. 41 (a) shows inverse quantized transform coefficients distributed at regular intervals. Fig. 41 (b) shows adjusted inverse quantized transform coefficients.

[0334] An embodiment of a video signal processing device for obtaining an adjusted dequantized transform coefficient is described. The video signal processing device can adjust the interval between dequantization indices. Specifically, the video signal processing device can reduce the interval between dequantization indices that are relatively close to 0, and widen the interval between dequantization indices that are relatively far from 0. Specifically, the video signal processing device can obtain a moving quantization index by shifting the value of the quantization index from 0 by a predetermined offset. At this time, the video signal processing device can obtain a first dequantized transform coefficient and a second dequantized transform coefficient by multiplying each of the quantization index and the moving quantization index by a scale factor. The video signal processing device can obtain an adjusted dequantized transform coefficient by weighting the first dequantized transform coefficient and the second dequantized transform coefficient. Specifically, the video signal processing device can obtain the first quantization index using Equation 1. The video signal processing device can generate the first dequantized transform coefficient by applying the first quantization index to Equation 2. If the first quantization index value is positive, the video signal processing device may generate a second quantization index by adding 1 to the first quantization index value. If the first quantization index value is negative, the video signal processing device may generate a second quantization index by adding -1 to the first quantization index value. The video signal processing device may generate a second dequantized transform coefficient by applying mathematical expression 2 to the second quantization index. The video signal processing device may obtain a weighted average value by applying a first weight to the first dequantized transform coefficient and a second weight to the second dequantized transform coefficient. At this time, the obtained weighted average value may be a final dequantized transform coefficient. The first weight may be a value greater than the second weight, and the first weight and the second weight may be integers.In addition, the video signal processing device can perform the dequantized transform coefficient adjustment only when the absolute value of the first quantization index value is smaller than a pre-specified value. The pre-specified value may be a positive integer. For example, the pre-specified value may be 64. In addition, the mathematical formula for obtaining the weighted average may be as shown in mathematical formula 8. At this time, DequantTcoeff is the final dequantized transform coefficient, DequantTcoeff1 is the first dequantized transform coefficient, DequantTcoeff2 is the second dequantized transform coefficient, m is the first weight, n is the second weight, and shift may be a pre-specified number. At this time, when m is (1024 - T) and n is T, the value of the shift may be 10. T may be a pre-specified positive integer. In a specific embodiment, the values ​​of the quantization indices [0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,…,63] may each have [0,63,31,21,15,12,10,9,7,7,6,5,5,4,4,4,3,3,3,3,3,3,2,2,2,2,2,2,2,2,2,2,1,…,1] as the value of T. Weighted averaging may be performed in such a way that as the first weight increases, the second weight decreases.

[0335] Mathematical formula 8

[0336] DequantTcoeff = (m * DequantTcoeff1 + n * DequantTcoeff2) >> shift

[0337] In the method for obtaining the previously adjusted dequantized transform coefficients, the video signal processing device may determine the first weight and the second weight by using at least one or more of information on whether the current block is intra mode or inter mode, an encoding mode of the current block (e.g., DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, and DMVR), intra prediction directionality information of the current block, information on whether a component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, position information of the last coefficient of the current transform block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of the non-zero coefficients of the current transform block, quantization parameter information, transform type information of the current block (or the type of transform kernel), and the horizontal and vertical sizes of the current block. Specifically, the video signal processing device may determine the weight value of the weighted average based on whether the current block is predicted in an intra prediction mode. In a specific embodiment, if the current block is predicted in intra prediction mode, the video signal processing device may apply a new first weight by adding a third weight to the first weight. At this time, the video signal processing device may apply a new second weight by subtracting the third weight from the second weight. The third weight may be a negative or positive integer.

[0338] Referring to FIG. 41, the video signal processing device can change the value of the adjusted dequantized transform coefficient in a direction further from 0 than the value of the existing dequantized transform coefficient. In addition, when the difference value between the value of the adjusted dequantized transform coefficient and the value of the existing dequantized transform coefficient is defined as an offset, the video signal processing device can change the value of the adjusted dequantized transform coefficient so that the offset value becomes smaller as the quantization index increases.

[0339] A video signal processing device can adaptively apply value adjustment of dequantized transform coefficients on a block-by-block basis or a transform block-by-block basis. An encoder can include information in a bitstream indicating whether to apply value adjustment of dequantized transform coefficients to the current block on a coding block-by-coding block (or transform block-by-transform) basis. A decoder can parse information indicating whether to apply value adjustment of dequantized transform coefficients from the bitstream and then set whether to apply value adjustment of dequantized transform coefficients to the current coding block (or transform block).

[0340] The video signal processing device may determine whether to adjust the value of the dequantized transform coefficient by using at least one of information on whether the current block is intra mode or inter mode, an encoding mode of the current block (e.g., DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, and DMVR), intra prediction directionality information of the current block, information on whether the component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, position information of the last coefficient of the current transform block, position information of the coefficient of the current transform block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of the non-zero coefficients of the current transform block, position information of a sub-block, quantization parameter information, transform type information of the current block (or the type of transform kernel), and the horizontal and vertical sizes of the current block. In a specific embodiment, the video signal processing device may determine whether to adjust the dequantized transform coefficient of the current transform block based on whether the sum of the absolute values ​​of the non-zero coefficients of the current transform block is greater than a predetermined value. For example, if the sum of the absolute values ​​of the non-zero coefficients of the current transform block is greater than a predetermined value, the video signal processing device may not adjust the value of the dequantized transform coefficient. In another specific embodiment, if the sum of the absolute values ​​of the non-zero coefficients of the current transform block is less than a predetermined value, the video signal processing device may not adjust the value of the dequantized transform coefficient. The predetermined value may be a positive integer. In another specific embodiment, if the position information of the coefficients of the current transform block is a predetermined position, the video signal processing device may adjust the value of the dequantized transform coefficient. At this time, the predetermined position may be a low-frequency area or a high-frequency area.Specifically, if the pre-designated position is included in the low-frequency region, the pre-designated position may be a position between 0 and a pre-designated number according to the anti-diagonal scanning order, assuming that the upper left sample position in FIG. 30 is 0. If the pre-designated position is included in the high-frequency region, the pre-designated position may be a pre-designated number of coefficient positions according to the anti-diagonal scanning order based on the last non-zero coefficient position. The pre-designated number and the pre-designated number may be integers. For example, the pre-designated number may be 5. In another specific embodiment, if the transform skip is applied to the current block, the video signal processing device may not adjust the value of the inverse quantized transform coefficient. In another specific embodiment, the video signal processing device may adjust the value of the inverse quantized transform coefficient according to the restoration order of the coefficient of the current transform block in the restoration order. The video signal processing device may adjust the value of the inverse quantized transform coefficient, where the order of the coefficient of the current transform block corresponds to the coefficient from the order (M) to the (MN)-th order of the last coefficient in the restoration order. At this time, the video signal processing device may not adjust the value of the inverse quantized transform coefficient corresponding to the (MN)th order from the last coefficient order (M) in the restoration order in the current block. Here, N may be a positive integer. The value of N may vary depending on the size of the current transform block. The value of N may be 5.

[0341] Even when the dependent quantization method is not applied and the independent quantization method is applied, the video signal processing device can adjust the value of the dequantized transform coefficient. The video signal processing device can generate a first dequantized transform coefficient by applying the first quantized transform coefficient to Equation 2. When the value of the first quantized transform coefficient is positive, the video signal processing device can generate a second quantized transform coefficient by adding 1 to the value of the first quantized transform coefficient. When the value of the first quantized transform coefficient is negative, the video signal processing device can generate a second quantized transform coefficient by adding -1 to the value of the first quantized transform coefficient. The video signal processing device can generate a second dequantized transform coefficient by applying the second quantized transform coefficient to Equation 2. The video signal processing device can perform weighted averaging by applying a first weight to the first dequantized transform coefficient and applying a second weight to the second dequantized transform coefficient. The video signal processing device generates a value obtained by performing the weighted averaging as a final dequantized transform coefficient. The first weight may be a value greater than the second weight, and the first and second weights may be integers. Furthermore, the video signal processing device may adjust the dequantized transform coefficients only when the absolute value of the first quantization index value is less than a predetermined value. The predetermined value may be a positive integer. For example, the predetermined value may be 64.

[0342] When the dependent quantization method is applied, the video signal processing device can selectively adjust the value of the dequantized transform coefficient for each quantizer. In the dependent quantization method, a Q0 quantizer and a Q1 quantizer are defined. The video signal processing device can select which quantizer to use depending on the status. At this time, the video signal processing device can adjust the value of the dequantized transform coefficient only when a pre-specified quantizer is used. The pre-specified quantizer can be either a Q0 quantizer or a Q1 quantizer.

[0343] When the dependent quantization method is applied, the video signal processing device can selectively adjust the value of the dequantized transform coefficient according to the state information used in the transition procedure between quantizers or the updated state change. If the quantizer applied to the previously dequantized transform coefficient is a Q0 quantizer (or a Q1 quantizer) and the quantizer applied to the currently dequantized transform coefficient is a Q1 quantizer (or a Q0 quantizer), i.e., since the quantizer has changed due to the updated state, the video signal processing device can adjust the currently dequantized transform coefficient. In another specific embodiment, if the state information used when obtaining the currently dequantized transform coefficient is a predetermined state, the video signal processing device can adjust the currently dequantized transform coefficient.

[0344] To save memory when implemented in hardware, one dequantized transform coefficient can be stored in a 2-byte memory. The dequantized transform coefficient value can range from -32768 to 32767. In addition, the dequantized transform coefficient value can be clipped so that the dequantized transform coefficient value can have a value within the range. Since the adjustment of the value of the dequantized transform coefficient uses a weighted average of the first dequantized transform coefficient and the second dequantized transform coefficient, it may require a lot of memory when performing the calculation. Therefore, an overflow problem may occur when calculating within a limited memory. Therefore, when the video signal processing device adjusts the value of the dequantized transform coefficient, the video signal processing device may adjust the value of the dequantized transform coefficient using an offset without applying the weighted average. In addition, if the quantization index value is greater than a specific value, the value of the dequantized transform coefficient may not be adjusted. Therefore, in certain cases, the process of calculating the second dequantized transform coefficient may not be necessary. Improvements to these unnecessary processes may be necessary to reduce complexity.

[0345] Figure 42 shows an example of an offset added to the value of the inverse quantized transform coefficient.

[0346] In adjusting the value of the offset-based dequantized transform coefficient, the video signal processing device can obtain the value of the adjusted dequantized transform coefficient by adding a predefined offset to the dequantized transform coefficient according to the quantization scale factor and the absolute value of the quantization index. At this time, the offset may be as shown in FIG. 42. In addition, when the quantization index value is positive, the video signal processing device can obtain the value of the adjusted dequantized transform coefficient by adding the offset value to the dequantized transform coefficient. In addition, when the quantization index value is negative, the video signal processing device can obtain the value of the adjusted dequantized transform coefficient by subtracting the offset value from the dequantized transform coefficient. When the quantization index value is 0, the dequantized transform coefficient is 0. In addition, the video signal processing device can determine whether to apply the offset value by comparing the absolute value of the quantization index with a pre-specified value. In an embodiment, when the absolute value of the quantization index is less than the pre-specified value, the video signal processing device can apply the offset value to the dequantized transform coefficient calculated using the quantization index. The predefined value can be an integer greater than or equal to 1. For example, the predefined value can be one of 1, 2, 3, 4, 5, and 6. In addition, the range of the inverse quantized transform coefficient values ​​can be clipped to have a predefined range of values. The predefined range can be from -32768 to 32767. (a) of Fig. 42 is a predefined offset value for quantization index values ​​from 1 to 5. (b) of Fig. 42 is a predefined offset value for quantization index values ​​from 1 to 6. In addition, (b) of Fig. 42 adds an offset for a wider range of quantization indices than (a) of Fig. 42.The video signal processing device can determine which table to apply among two different predefined tables ((a) and (b) of FIG. 42) by using at least one of the quantization parameters, intra prediction direction information of the current block, information on whether the component of the current block is a luminance component or a chrominance component, information on whether the current block has a dual tree structure applied, transformation type information of the current block (or type of transformation kernel), whether a user quantization matrix is ​​used, whether the block is a transformation skip block, and the horizontal and vertical sizes of the current block.

[0347] The video signal processing device can set the offset described above inversely or proportionally according to the absolute value of the quantization index. The video signal processing device can change the value of the existing inversely quantized transform coefficient so that the offset of the inversely quantized transform coefficient becomes smaller (or larger) as the magnitude of the absolute value of the quantization index increases. In addition, the video signal processing device can determine the offset value using at least one or more of the quantization parameter, the intra prediction directionality information of the current block, information on whether the component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, the transform type information of the current block (or the type of transform kernel), whether a user quantization matrix is ​​used, whether it is a transform skip block, and the width and height of the current block. Here, the transform skip block is a block that does not apply a transform (or inverse transform) process to the current block. From the decoder's perspective, when inverse quantization is performed, an error block is generated without performing an inverse transform. The video signal processing device adds the prediction block to the generated error block to generate the final reconstructed block.

[0348] In the offset-based dequantized transform coefficient acquisition method described above, the video signal processing device may determine an offset using at least one or more of information on whether the current block is intra mode or inter mode, an encoding mode of the current block (e.g., DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, DMVR), intra prediction directionality information of the current block, information on whether the component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, position information of the last coefficient of the current transform block, position information of the coefficient of the current transform block, position information of the sub-block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of the non-zero coefficients of the current transform block, quantization parameter information, transform type information of the current block (or type of transform kernel), and the horizontal and vertical sizes of the current block. In a specific embodiment, the value of the offset may be determined based on whether the current block is predicted in an intra prediction mode. Specifically, if the current block is predicted in intra prediction mode, the video signal processing device can apply a new offset by adding an additional offset to the base offset. The additional offset can be calculated based on the base offset.

[0349] Since the degree of quantization distortion varies depending on the quantization method, the video signal processing device can generate a final reconstructed block using a weighted average between blocks reconstructed through various quantization methods. The video signal processing device can configure a first reconstructed block through a first quantization technique, configure a second reconstructed block through a second quantization technique, and then perform a weighted average by applying a first weight to the first reconstructed block and a second weight to the second reconstructed block. The video signal processing device can generate a block obtained through the weighted average as a final reconstructed block.

[0350] A video signal processing device can generate a final error block by using a weighted average of error blocks obtained through various quantization methods. The video signal processing device can form a first error block through a first quantization technique, form a second error block through a second quantization technique, and then apply a first weight to the first error block and a second weight to the second error block to perform a weighted average. The video signal processing device can use the block obtained through the weighted average as the final error block. The video signal processing device can generate a reconstructed block by adding the final error block to a prediction block.

[0351] An encoder can obtain a quantization index (or quantized transform coefficient) by dividing a transform coefficient by a scale factor. A transform coefficient close to 0 can be divided by the scale factor to make the quantization index 0. A decoder can obtain a quantization index from a bitstream and obtain a dequantized transform coefficient (encoder transform coefficient) by multiplying the quantization index by the scale factor. The decoder can obtain a valid dequantized transform coefficient if the value of the quantization index is not 0. If the quantization index value is 0, the decoder obtains 0 as the dequantized transform coefficient. A video signal processing device can selectively restore the dequantized transform coefficient even if the quantization index value is 0. This can minimize quantization distortion. When the quantization index (or quantization transform coefficient) value is 0, the video signal processing device can obtain a non-zero dequantized transform coefficient by using at least one or more of information on whether the current block is intra mode or inter mode, an encoding mode of the current block (e.g., DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, DMVR), intra prediction directionality information of the current block, information on whether the component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, position information of the last coefficient of the current transform block, position information of the coefficient of the current transform block, position information of the sub-block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of the non-zero coefficients of the current transform block, quantization parameter information, transform type information of the current block (or the type of transform kernel), and the horizontal and vertical sizes of the current block.In a specific embodiment, if the current quantization index value is 0, the position of the quantization index within the current transform block is a position within a predetermined range, and the sum of the absolute values ​​of non-zero coefficients of the current transform block is greater than a predetermined value, the video signal processing device may determine a predetermined inverse quantized transform coefficient as the inverse quantized transform coefficient. The predetermined inverse quantized transform coefficient may vary depending on a scale factor. The position within the predetermined range may be an NxN block range based on the upper left position of the current transform block. The predetermined value and N may be positive integers.

[0352] The video signal processing device can also apply the dequantized transform coefficient adjustment described above to a block to which the transform skip mode is applied. The transform skip mode is a mode in which a transform process is not performed, and the encoder quantizes an error signal without transforming and encodes the quantized coefficients. In the transform skip mode, the decoder can dequantize the quantized coefficients to generate an error signal. Since the transform process is not performed in the transform skip mode, all quantized and dequantized coefficients can have the same importance regardless of the position of the coefficients within the block. Therefore, in the transform skip mode, the video signal processing device can output the modified dequantized coefficients by applying a predetermined value to all coefficients equally. In addition, the video signal processing device can use the modified dequantized coefficients as an error signal. The predetermined value can be a positive integer. For example, it can be one of 1, 2, 3, 4, 5, and 6. If the value of the dequantized coefficient is positive, the video signal processing device can output a modified dequantized coefficient by adding a predetermined value to the value of the dequantized coefficient. If the value of the dequantized coefficient is negative, the video signal processing device can change the predetermined value to a negative number and add the changed value to the value of the dequantized coefficient. This allows the changed dequantized coefficient to move further away from 0, thereby increasing the average size of the error signal. The embodiments described above can be applied not only when the current block corresponds to skip mode but also when it is a general enriched block. Specifically, the embodiments described above can be applied to blocks to which independent quantization and dependent quantization are applied.A video signal processing device can determine whether to apply dequantized coefficient adjustment on a block-by-block basis by using at least one or more of information on whether the current block is intra mode or inter mode, an encoding mode of the current block (e.g., DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, and DMVR), intra prediction directionality information of the current block, information on whether the component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, position information of the last coefficient of the current transform block, position information of the coefficient of the current transform block, position information of a sub-block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of the non-zero coefficients of the current transform block, quantization parameter information, transform type information of the current block (or type of transform kernel), and the horizontal and vertical sizes of the current block. The encoder can signal information on whether to apply dequantized coefficient adjustment. In addition, the encoder can additionally signal information on which offset value to apply. The decoder can parse information on whether to adjust the dequantized coefficients in the current block to determine whether to adjust the dequantized coefficients in the current block. If the video signal processing device adjusts the dequantized coefficients in the current block, the video signal processing device can additionally parse information related to the offset value. The video signal processing device can determine the offset value through the parsed information. The decoder can add the offset value to the dequantized coefficients and then generate the modified dequantized coefficients. If the transform skip mode is applied to the current block, since the dequantized coefficients are residual signals, the decoder can output the final reconstructed block by adding the predicted block samples and the residual signal samples. At this time, the offset value can be signaled using a list. Specifically, the encoder can construct an offset candidate list and then signal index information for the optimal offset in the offset candidate list.The decoder can parse the offset index information to determine the optimal offset from the offset candidate list. The video signal processing device can reorder the offset candidate list based on the template cost. The offset candidate list can include at least one of the offsets used in the surrounding blocks and the preset offsets.

[0353] A video signal processing device can output a compensated residual signal by adding an offset to a residual signal of a current block. The video signal processing device can determine, on a block-by-block basis, whether to generate a compensated residual signal by adding an offset to the residual signal by using at least one or more of information on whether the current block is intra mode or inter mode, an encoding mode of the current block (DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, DMVR, etc.), intra prediction directionality information of the current block, information on whether a component of the current block is a luminance component or a chrominance component, information on whether a dual tree structure is applied to the current block, information on whether a transform skip mode is applied to the current block, position information of the last coefficient of the current transform block, position information of the coefficients of the current transform block, position information of a sub-block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of the non-zero coefficients of the current transform block, quantization parameter information, transform type information of the current block (or type of transform kernel), and the horizontal and vertical sizes of the current block. Generating a compensated residual signal by adding an offset to the residual signal can be referred to as an offset-based residual change (ORC). A video signal processing device can perform ORC when a transform skip mode is applied to the current block. The encoder can signal whether to apply ORC to the current block. Additionally, the encoder can signal additional information regarding which offset value to apply. The decoder can parse the information regarding whether to apply ORC to the current block to determine whether to apply ORC to the current block. If ORC is applied to the current block, the decoder can additionally parse information related to the offset value. The decoder can determine the offset value based on the parsed information. The decoder can generate a modified residual signal sample by adding the offset value to the residual signal sample.Additionally, the decoder can output the final reconstructed block by adding the predicted block samples and the residual signal samples. At this time, the offset value can be signaled using a list. Specifically, the encoder can construct an offset candidate list and then signal index information for the optimal offset in the offset candidate list. The decoder can parse the offset index information and determine the offset indicated by the offset index information in the offset candidate list as the optimal offset. The offset candidate list can be reordered based on the template cost. The offset candidate list can include at least one or more of the offsets used in the surrounding blocks and the preset offsets. The template cost can be applied by modifying the template cost-based method described in FIG. 8.

[0354] When a dependent quantization method is applied to the current block, the video signal processing device can selectively apply ORC depending on which quantizer was used for the current block, the state information used in the transition procedure, and the updated state information. Specifically, when Q0 and Q1 quantizers exist, the video signal processing device can apply ORC only to samples reconstructed using the Q0 quantizer. In another specific embodiment, the video signal processing device can apply ORC only to samples reconstructed using the Q1 quantizer.

[0355] If the value of the residual signal sample changed through ORC falls outside the range of allowable samples, the video signal processing device can clip it to fall within the range of allowable samples. If the range of allowable samples is 0 to 1023 (in the case of 10 bits), and the value of the changed residual signal sample is less than 0, it can be changed to 0. Additionally, if the value of the changed residual signal sample is greater than 1023, it can be changed to 1023. Clipping can be applied in the same way to the dequantized transform coefficient adjustment described above.

[0356] A bitstream consists of one or more coded video sequences (CVSs), each of which is encoded independently of the others. Each CVS consists of one or more layers. Each layer can represent a specific quality level and a specific resolution. Additionally, a layer can represent a normal image, a depth map, and a transparency map. Additionally, a coded layer video sequence (CLVS) refers to a layer-wise CVS composed of consecutive PUs in the same layer in decoding order. For example, a bitstream may have a CLVS for a specific quality layer and may include a CLVS for a depth map.

[0357] The video signal processing device can selectively apply value adjustment of a dequantized transform coefficient to at least one of a VPS / SPS / PPS, a Slice / Tile header, a CTU, a CU, a PU, a TU, and a sample unit.

[0358] FIG. 43 shows signaling to SPS whether to enable value adjustment of inverse quantized transform coefficients according to an embodiment of the present invention.

[0359] An encoder can include in the SPS information signaling whether value adjustment of dequantized transform coefficients is enabled. A decoder can parse the information signaling whether value adjustment of dequantized transform coefficients is enabled from the SPS. The decoder can adjust the values ​​of the dequantized transform coefficients based on the information signaling whether value adjustment of the parsed dequantized transform coefficients is enabled.

[0360] In the embodiment of FIG. 43, when sps_shift_quantizer_enabled_flag is True, 1, value adjustment of the dequantized transform coefficient is enabled. When sps_shift_quantizer_enabled_flag is False, 0, value adjustment of the dequantized transform coefficient is disabled. If sps_shift_quantizer_enabled_flag is not parsed, the video signal processing device can infer it as 0.

[0361] Whether to enable value adjustment of dequantized transform coefficients can be signaled in at least one of SPS, PPS / sub-picture / Slice / Tile. The decoder can determine whether to enable value adjustment of dequantized transform coefficients by parsing one or more flags in at least one of SPS, PPS, sub-picture, Slice, and Tile.

[0362] FIG. 44 shows a constraint flag related to adjusting the value of a dequantized transform coefficient in the general_constraint_info() syntax structure according to an embodiment of the present invention.

[0363] The general_constraint_info() syntax can be called from the profile_tier_level() syntax. The profile_tier_level() syntax can be called from the sequence parameter set RBSP syntax, the video parameter set RBSP syntax, and the decoding capability information RBSP syntax. Individual syntax elements of the general_constraint_info() syntax can correspond to a sequence parameter set RBSP, and the activation of the corresponding sequence parameter set RBSP syntax element can be constrained by the definition of the corresponding flag.

[0364] The constraint flag related to adjusting the values ​​of the inverse quantized transform coefficients can be gci_no_shift_quantizer_constraint_flag.

[0365] If the gci_no_shift_quantizer_constraint_flag value is equal to 1, the video signal processing device must treat the sps_shift_quantizer_enabled_flag value as 0 for all pictures existing in OlsScope. Therefore, if the gci_no_shift_quantizer_constraint_flag value is equal to 1, the video signal processing device cannot adjust the values ​​of the dequantized transform coefficients. If the gci_no_shift_quantizer_constraint_flag value is equal to 0, there is no constraint that forces the video signal processing device to treat the sps_shift_quantizer_enabled_flag value as 0.

[0366] Figure 45 shows a flowchart for parsing the transformation type of the current block according to an embodiment of the present invention.

[0367] The decoder can parse the LFNST index for the current block. The video signal processing device can determine whether to parse the LFNST index by using at least one or more of the encoding mode of the current block (DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, DMVR, etc.), intra prediction direction information of the current block, information on whether the component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, position information of the last coefficient of the current transform block, position information of the coefficient of the current transform block, position information of the sub-block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of the non-zero coefficients of the current transform block, quantization parameter information, transform type information of the current block (or type of transform kernel), and the horizontal and vertical sizes of the current block. In a specific embodiment, if the minimum value of the horizontal and vertical sizes of the current block is less than a predetermined size or the encoding mode of the current block is SGPM, the video signal processing device may not parse the LFNST index of the current block. At this time, the video signal processing device may set the LFNST index to 0. At this time, 0 may indicate that LFNST is not used in the current transformation block. Additionally, the predetermined size may be an integer. For example, the predetermined size may be 4.

[0368] When the LFNST index is 0, the video signal processing device can parse information on whether MTS is applied for the current block. The video signal processing device can determine whether to parse the information on whether MTS is applied by using at least one of the encoding mode of the current block (DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, DMVR, etc.), intra prediction direction information of the current block, information on whether the component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, position information of the last coefficient of the current transform block, position information of the coefficient of the current transform block, position information of the sub-block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of the non-zero coefficients of the current transform block, quantization parameter information, transform type information of the current block (or type of transform kernel), and horizontal and vertical sizes of the current block. In a specific embodiment, if the maximum values ​​of the horizontal and vertical sizes of the current block are greater than a predetermined size or the encoding mode of the current block is SGPM, the video signal processing device may not parse the information on whether or not MTS is applied to the current block. At this time, the video signal processing device may set the information on whether or not MTS is applied to the current transform block to 0. Here, 0 may indicate that MTS is not applied to the current transform block, but DCT2 transform is applied. In addition, the predetermined size may be the maximum size of the current CTU. For example, the predetermined size may be 256.

[0369] If MTS is not applied to the current block, the video signal processing device may perform an inverse transform process using the DCT2 kernel on the current block. At this time, the video signal processing device may generate an error block after the inverse transform.

[0370] When MTS is applied to the current block, the video signal processing device parses the MTS index for the current block. The video signal processing device may determine whether to parse the MTS index by using at least one of the encoding mode of the current block (e.g., DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, and DMVR), intra prediction direction information of the current block, information on whether the component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, position information of the last coefficient of the current transform block, position information of the coefficient of the current transform block, position information of the sub-block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of the non-zero coefficients of the current transform block, quantization parameter information, transform type information of the current block (or type of transform kernel), and horizontal and vertical sizes of the current block. In a specific embodiment, if the maximum values ​​of the horizontal and vertical sizes of the current block are greater than a predetermined size or the encoding mode of the current block is SGPM, the video signal processing device may not parse the MTS index of the current block. At this time, the video signal processing device may perform an inverse transform on the current block using the DCT2 kernel. At this time, the predetermined size may be the maximum size of the current CTU. For example, the predetermined size may be 256.

[0371] When the LFNST index is greater than 0, the video signal processing device can determine whether to apply NSPT or DCT2 + LFNST to the current block by using at least one of the encoding mode of the current block (e.g., DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, and DMVR), intra prediction direction information of the current block, information on whether the component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, position information of the last coefficient of the current transform block, position information of the coefficient of the current transform block, position information of the sub-block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of the non-zero coefficients of the current transform block, quantization parameter information, transform type information of the current block (or the type of transform kernel), and horizontal and vertical sizes of the current block. In a specific embodiment, if the size of the current block is a predetermined size and the current block applies a dual tree structure or the current block is not a chrominance component, the video signal processing device may perform an inverse transformation by applying an NSPT transformation to the current block. At this time, the predetermined size may be one of 4x4, 4x8, 8x4, 8x8, 4x16, 16x4, 8x16, 16x8, 4x32, 32x4, 8x32, and 32x8. Alternatively, the predetermined size may be a size equal to or smaller than a size of MxN. M and N are positive integers and may be 4. The video signal processing device may obtain an NSPT index using the parsed LFNST index. The video signal processing device may derive an NSPT kernel using the NSPT index. The video signal processing device may perform an inverse transformation using the NSPT. At this time, the inverse transformation process may be applied to the embodiments described above.

[0372] A video signal processing device can use one LFNST index for NSPT inverse transformation or LFNST inverse transformation. Specifically, a decoder can use one parsed information (or syntax) to derive different encoding information. In addition, during entropy coding, a video signal processing device can derive a context model using encoding information with conflicting characteristics. At this time, the video signal processing device can parse one piece of information (or syntax) using the derived context model. An encoder can derive a context model using encoding information with conflicting characteristics, and then entropy code one piece of information (or syntax). In a specific embodiment, the encoding information with conflicting characteristics may be an NSPT index and an LFNST index. In addition, one piece of information to be entropy coded (or decoded) may be index information. One piece of index information may also be used as an NSPT index or an LFNST index. Additionally, when the video signal processing device derives a context model, the video signal processing device can utilize both NSPT and LFNST encoding information of the surrounding blocks.

[0373] Specifically, the process of a video signal processing device parsing the transformation type of the current block from the bitstream may be as shown in FIG. 45.

[0374] FIG. 46 shows a flowchart for parsing the transformation type of the current block according to another embodiment of the present invention.

[0375] A video signal processing device can determine whether to apply NSPT or DCT2 + LFNST to the current block using the method described above. If NSPT is applied to the current block, the video signal processing device can derive an NSPT context model and then perform entropy decoding to obtain an NSPT index. The video signal processing device can derive an NSPT kernel using the obtained NSPT index. The video signal processing device can perform an inverse transform by applying the NSPT kernel to the currently dequantized transform coefficients. The video signal processing device can obtain an error block for the current block through the inverse transform. If LFNST is applied to the current block, the video signal processing device can derive an LFNST context model. At this time, the NSPT context model and the LFNST context model may be different. The video signal processing device can obtain an LFNST index by performing entropy decoding. The video signal processing device can derive an LFNST kernel using the obtained LFNST index and then perform an inverse transform by applying the LFNST kernel to the currently dequantized transform coefficients. A video signal processing device can obtain an error block for a current block using inverse quantization. When the video signal processing device parses the NSPT index and LFNST index through entropy decoding, a fixed-length method or a variable-length method can be used in the binarization process. In the fixed-length method, the length of the codeword for each symbol is configured to be the same during binarization. In the variable-length method, the length of the codeword for each symbol can be configured to be different during binarization. The fixed-length method can be effective when the probability of occurrence of symbols is similar for all symbols. The variable-length method can be effective when the probability of occurrence of symbols is high for several symbols.

[0376] The video signal processing device can determine whether to parse information on whether MTS is applied to the current block using the NSPT index and the LFNST index. If the NSPT index is greater than 0 or the LFNST index is greater than 0, the video signal processing device may not parse information on whether MTS is applied to the current block. If the NSPT index is 0 and the LFNST index is 0, the video signal processing device may parse information on whether MTS is applied to the current block. The video signal processing device may perform subsequent processes using the described embodiments described above.

[0377] Specifically, a flowchart of a video signal processing device parsing and using the NSPT index and LFNST index from a bitstream may be as shown in FIG. 46.

[0378] FIG. 47 shows a flowchart for parsing the transformation type of the current block according to another embodiment of the present invention.

[0379] The video signal processing device can vary the available transformation methods and the parsing order depending on the size of the current block. NSPT can have higher encoding efficiency for transform blocks smaller than LFNST. Therefore, the available transformation methods for the current block can vary depending on the size of the current block. If the block size is smaller than or equal to a predetermined size, the available transformation methods for the video signal processing device can be DCT2, MTS, and NSPT. If the block size is larger than the predetermined size, the available transformation methods for the video signal processing device can be DCT2, MTS, and DCT2+LFNST. The predetermined size can vary depending on the size of the transform kernel. In addition, the predetermined size can be one of 4x4, 4x8, 8x4, 8x8, 4x16, 16x4, 8x16, 16x8, 4x32, 32x4, 8x32, and 32x8. Figure 47 is a flowchart illustrating a process for parsing transformation methods available for a current block, which vary depending on the size of the current block, according to an embodiment of the present invention. The transformation type for the current block may be determined in the following order.

[0380] When the size of the current block is a predetermined size, the video signal processing device can parse information on whether DCT2 is applied for the current block. The video signal processing device can determine whether to parse the information on whether DCT2 is applied by using at least one of an encoding mode of the current block (e.g., DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, and DMVR), intra prediction directionality information of the current block, information on whether a component of the current block is a luminance component or a chrominance component, information on whether a dual tree structure is applied to the current block, position information of the last coefficient of the current transform block, position information of the coefficient of the current transform block, position information of a sub-block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of non-zero coefficients of the current transform block, quantization parameter information, transform type information (or type of transform kernel) of the current block, and horizontal and vertical sizes of the current block. If DCT2 is applied to the current block, the video signal processing device can perform an inverse transformation through DCT2 on the current block to generate an error block. If DCT2 is not applied to the current block, the video signal processing device can parse information on whether MTS is applied for the current block. At this time, the video signal processing device can determine whether to parse the information on whether MTS is applied according to the method described above. If MTS is used in the current block, the video signal processing device can parse the MTS index for the current block. The video signal processing device can derive a transform kernel using the parsed MTS index using the method described above. The video signal processing device can perform an inverse transformation using the derived transform kernel to generate an error block. If MTS is not used in the current block, the video signal processing device can parse the NSPT index for the current block.A video signal processing device can derive a transform kernel using the parsed NSPT index using the method described above. The video signal processing device can then perform an inverse transform using the derived transform kernel to generate an error block.

[0381] If the size of the current block is not a pre-specified size, the video signal processing device can parse the LFNST index for the current block. The video signal processing device can determine whether to parse the LFNST index according to the method described above. If the parsed LFNST index is greater than 0, the video signal processing device can derive a transform kernel using the parsed LFNST index using the method described above. The video signal processing device can perform an inverse transform using the derived transform kernel to generate an error block. If the parsed LFNST index is 0, the video signal processing device can parse information on whether to apply MTS for the current block. The embodiments described above can be applied to the subsequent processes.

[0382] The video signal processing device can determine a transform method usable for the current block by using at least one or more of the encoding mode of the current block (e.g., DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, and DMVR), intra prediction direction information of the current block, information on whether a component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, position information of the last coefficient of the current transform block, position information of the coefficient of the current transform block, position information of a sub-block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of non-zero coefficients of the current transform block, quantization parameter information, transform type information of the current block (or type of transform kernel), and horizontal and vertical sizes of the current block.

[0383] In addition, the video signal processing device can signal and parse syntax for the LFNST index, NSPT index, MTS index, whether DCT2 is applied, whether MTS is applied, whether NSPT is applied, and whether LFNST is applied for the transformation type of the current block in various orders. In a specific embodiment, information on whether MTS is applied may be signaled and parsed first. If MTS is applied, the MTS index may be signaled and parsed. If MTS is not applied, information on whether LFNST and NSPT are applied may be parsed. In another specific embodiment, information on whether DCT2 is applied may be signaled and parsed first. If DCT2 is applied, other transformation type information may not be signaled and parsed. If DCT2 is not applied, information on whether MTS is applied, LFNST, and whether NSPT are applied may be parsed. The process of signaling and parsing other syntaxes may be applied to the embodiments described above.

[0384] A video signal processing device can derive which transformation (or inverse transformation) method to apply to a current block from a transformation type list. The video signal processing device can configure the transformation type list using at least one or more of the encoding mode of the current block (e.g., DIMD, TIMD, Intra TMP, SGPM, IBC, Merge, AMVP, and DMVR), intra prediction directionality information of the current block, information on whether the component of the current block is a luminance component or a chrominance component, information on whether the current block applies a dual tree structure, position information of the last coefficient of the current transform block, position information of the coefficient of the current transform block, position information of the sub-block, the number of non-zero coefficients of the current transform block, the sum of the absolute values ​​of the non-zero coefficients of the current transform block, quantization parameter information, transformation type information (or the type of transformation kernel) of the current block, and the horizontal and vertical sizes of the current block. In a specific embodiment, the video signal processing device can configure the transformation type list based on whether the size of the current luminance block is a predetermined size. Specifically, when the size of the current luminance block is a predetermined size, the video signal processing device can configure the transform type list as MTS, NSPT, DCT2, IDTR, and Transform SKIP. The encoder can include an index indicating an optimal transform type for the current block in the transform type list in the bitstream. The decoder can parse the transform type index from the bitstream and then set the transform type indicated by the transform type index in the derived transform type list as the transform type for the current block. The decoder can derive a transform kernel through the determined transform type. The decoder can perform an inverse transform through the derived transform kernel to generate an error block. In addition, the video signal processing device can reorder the transform type list based on a template cost and then determine an optimal index using the reordered transform type list.Additionally, the video signal processing device can determine the transformation type with the minimum cost according to the template cost from the transformation type list as the transformation type for the current block. The encoder may not signal the index for the optimal transformation type. The decoder may not parse the index for the optimal transformation type. At this time, the template cost for each transformation type candidate in the transformation type list can be calculated through SAD or MR-SAD between the reconstructed neighboring samples adjacent to the current block and the current block reconstructed using each transformation type candidate.

[0385] FIG. 48 shows a method for calculating a template cost for each transformation type candidate in a transformation type list according to an embodiment of the present invention.

[0386] The video signal processing device can perform inverse transformation using each transformation type candidate from the transformation type list. The video signal processing device can restore the current block using the generated error block. The video signal processing device can extract the samples of the boundary blocks (P in (a) of Fig. 48) from the restored current block. x,1 and P 1,y ) and the restored sample adjacent to the current block (R in (a) of Fig. 48 x,-1 , R x,0 , R -1,y , R 0,y ) can be applied to (b) of Fig. 48 to calculate the template cost. Here, w is the width of the current block, and h is the height of the current block. There are various other methods for calculating the template cost.

[0387] Figure 49 shows the coordinates of a luminance block corresponding to the coordinates of a chrominance block according to an embodiment of the present invention.

[0388] Depending on the video format of the current picture, the size of the luminance block and the size of the chrominance block of the current block may be different. If the video format of the current picture is 4:2:0, the chrominance block may be 1 / 4 the size of the luminance block. If the luminance block is 16x16, the chrominance block may be 8x8. Since the block sizes between the luminance block and the chrominance block are different, the transform types applicable to the luminance block and the chrominance block may be different. Alternatively, even if the block sizes between the luminance block and the chrominance block are different, the transform types applicable to the luminance block and the chrominance block may be the same. If a dual tree structure is applied to the current transform block, the encoder can signal the transform type information for the current luminance transform block and the transform type information for the current chrominance transform block separately and include them in the bitstream. The decoder can parse the transform type information for the current luminance transform block and the transform type information for the current chrominance transform block separately from the bitstream. At this time, the decoder can independently perform the inverse transformation process for the current luminance transformation block and the current chrominance transformation block. When a dual tree structure is applied to the current transformation block, the video signal processing device can derive the transformation type information for the current chrominance transformation block from the transformation type information for the current luminance transformation block corresponding to the current chrominance transformation block. At this time, the position of the luminance block for deriving the transformation type information of the luminance transformation block corresponding to the current chrominance transformation block may be a pre-specified position. In a specific embodiment, the luminance block may include C, TL, TR, BL, and BR positions, as shown in FIG. 49. The upper left position coordinate of the luminance block corresponding to the upper left position coordinate (xCbC, yCbC) of the chrominance block is (xCbL, yCbL). At this time, C represents the central position of the luminance block corresponding to the central position of the current chrominance block, and TL, TR, BL, and BR represent the positions of the luminance blocks corresponding to the upper left, upper right, lower left, and lower right positions of the current chrominance block, respectively.If a dual tree structure is applied to a current transform block and a luminance block corresponding to a current chroma block has NSPT (or LFNST, MTS, DCT2) transform (or inverse transform) applied, a video signal processing device can apply NSPT (or LFNST, MTS, DCT2) transform (or inverse transform) as a transform for the current chroma block. At this time, if the size of the chroma block is a block size to which NSPT cannot be applied, the video signal processing device can reset the transform type for the chroma block to a pre-specified transform type. The pre-specified transform type may be one of DCT2, DCT5, DCT7, DCT8, and DST transform types. The method of deriving transform type information for the current chroma transform block from transform type information for the corresponding current luminance transform block can also be applied when a dual tree structure is not applied to the current block (when a single tree structure is applied).

[0389] A video signal processing device can predict, signal, and parse transformation type information for a current chrominance block from transformation type information for a current luminance block corresponding to a current chrominance block. When signaling transformation type information for the current chrominance block, the encoder can derive transformation type information for the current luminance block. At this time, the encoder can signal information about whether the transformation type information for the current luminance block is the same as the transformation type information for the current chrominance block. If the transformation type for the current luminance block is the same as the transformation type information for the current chrominance block, the encoder can signal only information about whether the transformation type information for the current luminance block is the same as the transformation type information for the current chrominance block. If the transformation type for the current luminance block is not the same as the transformation type for the current chrominance block, the encoder can signal transformation type information for the current chrominance block. At this time, the video signal processing device can set transformation type information that is the same as the transformation type information for the current luminance block to not be used. This embodiment can reduce the number of bits required for signaling because the types of transformation types of the current chrominance block are reduced. The decoder can parse information on whether the transformation type of the chrominance block is the same as the transformation type of the luminance block. If the transformation type of the chrominance block is the same as the transformation type of the luminance block, the decoder can set the transformation type of the current chrominance block to the transformation type of the luminance block corresponding to the chrominance block. If the transformation type of the chrominance block is not the same as the transformation type of the luminance block, the decoder can parse the transformation type information of the current chrominance block. The decoder can set the parsed transformation type information to the transformation type of the chrominance block. At this time, when the decoder derives the transformation type information from the luminance block corresponding to the current chrominance block, the decoder can determine the position of the luminance block corresponding to the chrominance block according to the embodiments described above.

[0390] The embodiments of the present invention described above can vary the applicability or scope of application by using one or more of slice type information (e.g., information on whether it is I, P, or B), information on whether it is a tile, information on whether it is a subpicture, block size, CU depth, information on whether it is a CU block or a subblock, information on whether it is a prediction block or a transform block, information on whether it is a luminance signal or a chrominance signal, temporal layer information according to reference order and layer, AMVR resolution information of the current block, and information on whether it is a reference frame or a non-reference frame. The variable that determines the applicability or scope of application in this way can be set so that the encoder and decoder use a predetermined value. This can also be configured to use a value determined according to a profile and level. In addition, the encoder can include a variable value indicating the applicability or scope of application in the bitstream. The decoder can parse the variable value to determine the applicability or scope of application. The applicability can be determined on a block-by-block basis, a sub-block-by-block basis, a block-by-block basis, a tile-by-tile basis, or a slice-by-slice basis. The scope of application can vary depending on the size of the current block. Embodiments of the present invention can be applied to blocks whose current block size is larger or smaller than a predetermined size. The embodiments described above may be performed only on luminance signals and may not be performed on chrominance signals. In addition, information obtained from the luminance signal may be used in the chrominance signal. Whether the embodiments described above are performed may be determined based on whether the temporal layer information is larger or smaller than a predetermined value. Whether the embodiments described above are performed may be determined based on information regarding whether the frame is a reference frame or a non-reference frame. The embodiments described above may be applied when the horizontal or vertical length is 32 or more (e.g., 32, 64, and 128).Additionally, the embodiments described above may be applied when the horizontal or vertical length is less than 32 (e.g., 2, 4, 8, and 16). Additionally, the embodiments described above may be applied when the horizontal or vertical length is 4 or 8.

[0391]

[0392] The methods described herein may be performed by a processor of a decoder or encoder. Furthermore, the encoder may generate a bitstream decoded by the methods described above. Furthermore, the bitstream generated by the encoder may be stored on a computer-readable, non-transitory storage medium (recording medium).

[0393] While this specification primarily addresses the decoder's perspective, the same principles apply to encoders. While the term "parsing" in this specification focuses on the process of acquiring information from a bitstream, it can also be interpreted as configuring that information into a bitstream from an encoder's perspective. Therefore, the term "parsing" is not limited to the decoder's operations; it can also be interpreted as the act of configuring a bitstream in an encoder. Furthermore, such a bitstream can be stored and configured on a computer-readable recording medium.

[0394] The embodiments of the present invention described above may be implemented through various means. For example, the embodiments of the present invention may be implemented through hardware, firmware, software, or a combination thereof.

[0395] In the case of hardware implementation, the method according to embodiments of the present invention may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), processors, controllers, microcontrollers, microprocessors, etc.

[0396] When implemented via firmware or software, the methods according to embodiments of the present invention may be implemented in the form of modules, procedures, or functions that perform the functions or operations described above. The software code may be stored in memory and executed by the processor. The memory may be located internally or externally to the processor and may exchange data with the processor via various known means.

[0397] Some embodiments may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media can be any available media that can be accessed by a computer, and includes both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media can include both computer storage media and communication media. Computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Communication media typically contains computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules, or other transport mechanisms, and includes any information delivery media.

[0398] The foregoing description of the present invention is for illustrative purposes only, and those skilled in the art will readily appreciate that the present invention can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present invention. Therefore, the embodiments described above should be construed as illustrative and not restrictive in all respects. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.

[0399] The scope of the present invention is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present invention.

Claims

1. In a video signal decoding device, Includes a processor, The above processor Obtain the quantization index for the current block from the bitstream, Multiply the above quantized index by a scale factor to obtain the inverse quantized transform coefficient, Determine whether to adjust the above inverse quantized transform coefficients, When adjusting the above inverse quantized transform coefficients, the inverse quantized transform coefficients are adjusted, and the residual signal of the current block is inversely transformed using the adjusted inverse quantized transform coefficients to generate an error block. If the above inverse quantized transform coefficients are not adjusted, the residual signal of the current block is inversely transformed using the above inverse quantized transform coefficients to generate an error block. A video signal decoding device.

2. In paragraph 1, The above processor Obtain the above quantization index and the shifted quantization index by shifting the value of the above quantization index from 0 by a predetermined offset, Multiplying the above quantized index by the above scale factor to obtain a first inverse quantized transform coefficient, The second inverse quantized transform coefficient is obtained by multiplying the above moving quantization index by the above scale factor, Obtaining the adjusted inverse quantized transform coefficient by weighting the first inverse quantized transform coefficient and the second inverse quantized transform coefficient. Decoding device.

3. In paragraph 2, The above processor Determine the weight of the weighted average based on whether the current block is predicted in intra prediction mode. Decoding device.

4. In paragraph 1, The above processor Determines whether to adjust the above inverse quantized transform coefficients by transform block unit. Decoding device.

5. In paragraph 1, The above processor Determine whether to adjust the dequantized transform coefficients of the current transform block based on whether the sum of the absolute values ​​of the non-zero coefficients of the current transform block is greater than a pre-specified value. Decoding device.

6. In paragraph 1, The above processor Determine whether to adjust the dequantized transform coefficient based on the quantizer applied to the above quantization index. Decoding device.

7. In paragraph 1, The above processor Determines whether to adjust the value of the dequantized transform coefficient according to the state information used in the transition procedure for the above quantized index or the updated state change. Decoding device.

8. In paragraph 1, The above processor Obtaining the value of the adjusted dequantized transform coefficient by adding a predefined offset to the dequantized transform coefficient according to the quantization scale factor and the absolute value of the quantization index, The above quantization scale factor is a coefficient multiplied by the quantization index to obtain the inverse quantized transform coefficient. Decoding device.

9. In paragraph 8, The value of the above offset is determined based on whether the current block is predicted in intra prediction mode. Decoding device.

10. In paragraph 1, The above processor Parse information signaling whether adjustment of values ​​of dequantized transform coefficients is activated in the SPS (Sequence Parameter Set) included in the above bitstream, Determine whether to adjust the value of the dequantized transform coefficient based on information signaling whether adjustment of the value of the dequantized transform coefficient is activated. Decoding device.

11. In paragraph 1, The above processor The above inverse transformation is performed using a transformation type derived from a transformation type list. Decoding device.

12. In Article 11, The above list of conversion types is reordered based on template cost. Decoding device.

13. In Article 11, The above processor The transformation type list is constructed based on whether the size of the luminance block of the current block is a pre-specified size. Decoding device.

14. In paragraph 13, The above processor The above transformation type list consists of MTS, NSPT, DCT2, IDTR, and Transform SKIP. Decoding device.

15. In a video signal encoding device, Includes a processor, The above processor Multiply the quantized index by a scale factor to obtain the inverse quantized transform coefficients, Determine whether to adjust the above inverse quantized transform coefficients, When adjusting the above inverse quantized transform coefficients, the inverse quantized transform coefficients are adjusted, and the residual signal of the current block is inversely transformed using the adjusted inverse quantized transform coefficients to generate an error block. If the above inverse quantized transform coefficients are not adjusted, the residual signal of the current block is inversely transformed using the above inverse quantized transform coefficients to generate an error block, Including the quantization index for the current block in the bitstream. Encoding device.

16. In paragraph 15, The above processor Obtain the above quantization index and the shifted quantization index by shifting the value of the above quantization index from 0 by a predetermined offset, Multiplying the above quantized index by the above scale factor to obtain a first inverse quantized transform coefficient, The second inverse quantized transform coefficient is obtained by multiplying the above moving quantization index by the above scale factor, Obtaining the adjusted inverse quantized transform coefficient by weighting the first inverse quantized transform coefficient and the second inverse quantized transform coefficient. Encoding device.

17. In Article 16, The above processor Determine the weight of the weighted average based on whether the current block is predicted in intra prediction mode. Encoding device.

18. In paragraph 15, The above processor Determines whether to adjust the above inverse quantized transform coefficients by transform block unit. Encoding device.

19. In Article 18, The above processor Determine whether to adjust the dequantized transform coefficients of the current transform block based on whether the sum of the absolute values ​​of the non-zero coefficients of the current transform block is greater than a pre-specified value. Encoding device.

20. In a computer-readable non-transitory storage medium storing a bitstream, the bitstream is decoded by a decoding method, the decoding method A step of obtaining an inverse quantized transform coefficient by multiplying the above quantized index by a scale factor; A step of determining whether to adjust the above inverse quantized transform coefficients; When adjusting the dequantized transform coefficients, a step of adjusting the dequantized transform coefficients and inversely transforming the residual signal of the current block using the adjusted dequantized transform coefficients to generate an error block; and If the dequantized transform coefficients are not adjusted, a step of generating an error block by inversely transforming the residual signal of the current block using the dequantized transform coefficients is included. Non-transitory storage media.

Citation Information

Patent Citations

  • Adaptive inverse-quantization method and apparatus in video coding

    KR101893049B1

  • Method and apparatus for processing a video signal

    KR102424419B1

  • Entropy coding of transform coefficients suitable for dependent scalar quantization

    KR102533654B1

  • Method and apparatus for image encoding, and method and apparatus for image decoding

    KR102546694B1

  • Video decoder with reduced dynamic range transform with inverse transform shifting memory

    US20230052841A1