Video signal processing method using chroma mapping, and device therefor

The method enhances video signal processing efficiency by using chroma mapping to leverage the correlation between luminance and chrominance components, addressing inefficiencies in existing methods.

WO2026101039A1PCT designated stage Publication Date: 2026-05-15WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC
Filing Date
2025-10-15
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing video signal processing methods are inefficient due to the lack of effective techniques for enhancing coding efficiency by leveraging the correlation between luminance and chrominance components.

Method used

A video signal processing method and apparatus that utilize chroma mapping by generating a restoration block, obtaining a relational expression between luminance and chrominance components, and performing signal processing to enhance coding efficiency.

Benefits of technology

Improves coding efficiency by effectively utilizing the correlation between luminance and chrominance components, resulting in more efficient video signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025016276_15052026_PF_FP_ABST
    Figure KR2025016276_15052026_PF_FP_ABST
Patent Text Reader

Abstract

A video signal decoding device is disclosed. The video signal decoding device comprises a processor. The video signal decoding device comprises the processor, which generates a reconstructed block for the current block of the current picture included in a video signal, acquires a relational expression representing a correlation between a luma component and a chroma component of the reconstructed block, performs signal processing on the luma component, acquires the chroma component on which signal processing has been performed by using the luma component on which the signal processing has been performed and the relational expression, and generates a final reconstructed block for the current block on the basis of the luma component on which the signal processing has been performed and the chroma component on which the signal processing has been performed.
Need to check novelty before this filing date? Find Prior Art

Description

Video signal processing method using chroma mapping and apparatus for the same

[0001] The present invention relates to a method and apparatus for processing a video signal, and more specifically, to a video signal processing method and apparatus for encoding or decoding a video signal.

[0002] Compression encoding refers to a series of signal processing techniques used to transmit digitized information over communication lines or store it in a form suitable for storage media. Targets of compression encoding include voice, video, and text; specifically, the technology of performing compression encoding on video is called video image compression. Compression encoding of video signals is achieved by removing redundant information by considering spatial correlation, temporal correlation, and probabilistic correlation. However, due to recent advancements in various media and data transmission media, there is a demand for more efficient video signal processing methods and devices.

[0003] The present specification aims to improve the coding efficiency of a video signal by providing a video signal processing method using chroma mapping and an apparatus for the same.

[0004] The present specification provides a video signal processing method and an apparatus for the same. The video signal decoding apparatus includes a processor. The processor generates a restoration block for a current block of a current picture included in the video signal, obtains a relational expression representing the correlation between a luminance component and a chrominance component of the restoration block, performs signal processing on the luminance component, obtains a chrominance component for which signal processing has been performed using the luminance component for which signal processing has been performed and the relational expression, and generates a final restoration block for the current block based on the luminance component for which signal processing has been performed and the chrominance component for which signal processing has been performed.

[0005] The processor may use a downsampled luminance component when obtaining a relationship equation representing the correlation between the luminance component and the color difference component of the restoration block.

[0006] The processor can obtain information indicating whether to use the downsampled luminance component from the video signal, and when obtaining a relationship expression representing the correlation between the luminance component and the chrominance component of the restoration block, it can determine whether to use the downsampled luminance component according to the information indicating whether to use the downsampled luminance component.

[0007] The processor can obtain information regarding the format of the relationship from the video signal and obtain the relationship according to the information regarding the format of the relationship.

[0008] The processor obtains information regarding the blending of the color difference component before signal processing and the color difference component after signal processing from the video signal, and can blend the color difference sample before signal processing and the color difference sample after signal processing according to the information regarding blending.

[0009] The information regarding the above blending may include information regarding whether the above blending is applied.

[0010] The information regarding the above blending may include information regarding the type of the above blending.

[0011] The processor can obtain information regarding a window that specifies the range of components used to obtain the relationship from the video signal or the range of components to which the relationship is applied, and can obtain the relationship according to the information regarding the window or determine the range of color difference components to which the relationship is applied according to the information regarding the window.

[0012] The color difference component on which the signal processing was performed is clipped to obtain the clipped color difference component, and the luminance component on which the signal processing was performed and the clipped color difference component can be used to generate a final restored block for the current block.

[0013] The range of the clipping above can be determined according to the range of color difference sample values ​​before the signal processing is performed.

[0014] The above signal processing may include an in-loop filter.

[0015] An encoding device for a video signal according to an embodiment of the present invention includes a processor. The processor generates a restoration block for a current block of a current picture included in the video signal, obtains a relational expression representing the correlation between a luminance component and a chrominance component of the restoration block, performs signal processing on the luminance component, obtains a chrominance component for which signal processing has been performed using the luminance component for which signal processing has been performed and the relational expression, and generates a final restoration block for the current block based on the luminance component for which signal processing has been performed and the chrominance component for which signal processing has been performed.

[0016] The processor may use a downsampled luminance component when obtaining a relationship equation representing the correlation between the luminance component and the color difference component of the restoration block.

[0017] The processor may include information indicating whether to use the downsampled luminance component in a bit stream containing the video signal.

[0018] The processor can include information regarding the format of the relationship in a bit stream containing the video signal.

[0019] The processor can include information regarding the blending of the color difference component before signal processing and the color difference component after signal processing from the video signal in a bit stream containing the video signal.

[0020] The information regarding the above blending may include information regarding whether the above blending is applied.

[0021] The information regarding the above blending may include information regarding the type of the above blending.

[0022] A method of operating a decoding device for a video signal according to an embodiment of the present invention comprises the step of generating a restoration block for a current block of a current picture included in the video signal;

[0023] The method may include the steps of: obtaining a relationship equation representing the correlation between the luminance component and the color difference component of the above-mentioned restoration block; performing signal processing on the luminance component; obtaining a color difference component that has undergone signal processing using the luminance component that has undergone signal processing and the relationship equation; and generating a final restoration block for the current block based on the luminance component that has undergone signal processing and the color difference component that has undergone signal processing.

[0024] According to an embodiment of the present invention, a method for generating a bit stream containing a video signal, wherein the bit stream is included in a storage medium, comprises: generating a restoration block for a current block of a current picture containing the video signal; obtaining a relationship expression representing a correlation between a luminance component and a chrominance component of the restoration block; performing signal processing on the luminance component; obtaining a chrominance component for which signal processing has been performed using the luminance component for which signal processing has been performed and the relationship expression; and generating a final restoration block for the current block based on the luminance component for which signal processing has been performed and the chrominance component for which signal processing has been performed.

[0025] The present specification provides a method for efficiently processing video signals.

[0026] The effects obtainable in this specification are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description below.

[0027] FIG. 1 is a schematic block diagram of a video signal encoding device according to one embodiment of the present specification.

[0028] FIG. 2 is a schematic block diagram of a video signal decoding device according to one embodiment of the present specification.

[0029] FIG. 3 illustrates an embodiment in which a coding tree unit is divided into coding units within a picture.

[0030] FIG. 4 illustrates an example of a method for signaling the splitting of a quad tree and a multi-type tree.

[0031] FIGS. 5 and 6 illustrate an intra-prediction method according to an embodiment of the present invention in more detail.

[0032] Figure 7 is a diagram showing the locations of surrounding blocks used to construct a list of motion candidates in inter prediction.

[0033] FIG. 8 illustrates a method for determining a reference pixel line based on a template according to one embodiment of the present specification.

[0034] FIG. 9 is a diagram showing block vectors related to an IBC encoding method according to one embodiment of the present specification.

[0035] Figure 10 shows a method for predicting the current block using RRIBC in the horizontal direction.

[0036] Figure 11 shows a method for predicting the current block using RRIBC in the vertical direction.

[0037] FIG. 12 shows a block vector of a block encoded in Intra TMP mode according to one embodiment of the present specification.

[0038] FIG. 13 illustrates a case where the current block is divided by a GPM mode according to one embodiment of the present specification, and the divided area is encoded in an IBC mode.

[0039] FIG. 14 illustrates a method in which a current block is encoded in IBC-CIIP mode according to one embodiment of the present specification.

[0040] FIG. 15 shows an example of a reference region and filter shape used to derive CCCM parameters according to an embodiment of the present specification.

[0041] FIG. 16 shows a type of conversion kernel that can be used for video coding according to one embodiment of the present specification.

[0042] FIG. 17 shows a conversion set table for LFNST and NSPT conversions according to one embodiment of the present specification.

[0043] Figure 18 shows an example of an ROI after LFNST transformation.

[0044] FIG. 19 illustrates a method for deriving multiple conversion sets and LFNST / NSPT sets according to one embodiment of the present specification.

[0045] FIG. 20 shows a mapping table according to one embodiment of the present specification.

[0046] FIG. 21 shows a conversion type set table according to one embodiment of the present specification.

[0047] FIG. 22 shows a conversion type combination table according to one embodiment of the present specification.

[0048] FIG. 23 shows a threshold value table for IDT conversion types according to one embodiment of the present specification.

[0049] FIG. 24 shows the block boundary and samples around the boundary in a deblocking filtering process according to one embodiment of the present specification.

[0050] FIG. 25 shows PDPC filtering according to an embodiment of the present invention.

[0051] FIG. 26 shows a gradient PDPC according to an embodiment of the present invention.

[0052] FIG. 27 shows a matrix-based intra-prediction method according to an embodiment of the present invention.

[0053] Figure 28 illustrates an offset difference sample-based CCCM method.

[0054] Figure 29 shows an arbitrary number of luminance samples around the luminance sample of the current block.

[0055] Figure 30 shows luminance samples before downsampling used to derive color difference samples in CCCM-ND mode.

[0056] FIGS. 31 and 32 show structural diagrams for the application of a cross-component residual model (CCRM) according to one embodiment of the present invention.

[0057] FIG. 33 shows the downsampling filter and the position of the sample applied to the luminance block according to one embodiment of the present invention.

[0058] FIG. 34 shows the coefficients of filtering applied to a prediction block according to one embodiment of the present invention and the filtering positions.

[0059] FIG. 35 illustrates a method for determining a CCP mode based on a template cost according to an embodiment of the present invention.

[0060] FIG. 36 shows peripheral blocks used to induce CCCM for a current color difference block according to one embodiment of the present invention.

[0061] FIG. 37 shows a reference area indicated by the block vector of the current luminance block according to one embodiment of the present invention.

[0062] FIG. 38 is a structural diagram showing a method for predicting the current block using a CCP merge mode according to one embodiment of the invention.

[0063] FIG. 39 shows the surrounding blocks of the current block according to one embodiment of the present invention.

[0064] FIG. 40 shows a co-located block and surrounding blocks according to one embodiment of the present invention.

[0065] FIG. 41 shows a CCLM illustrated in the form of a graph according to one embodiment of the present invention.

[0066] FIG. 42 shows a reference template position used to rearrange a CCP merge list according to one embodiment of the present invention.

[0067] FIGS. 43 and 44 show the generation of a color difference block by applying CCCM in an LMCS according to an embodiment of the present invention.

[0068] FIG. 45 shows color difference filtering using a CCP model according to an embodiment of the present invention.

[0069] FIG. 46 shows color difference filtering using a CCP model according to another embodiment of the present invention.

[0070] FIG. 47 shows a color difference compensation process using a CCP model according to an embodiment of the present invention.

[0071] FIG. 48 shows compensation between components using a CCP model according to an embodiment of the present invention.

[0072] FIG. 49 shows a video signal processing device according to an embodiment of the present invention mapping color difference components using a CCP method.

[0073] FIG. 50 shows the area of ​​a window used by a video signal processing device according to an embodiment of the present invention when deriving parameters of chroma mapping from luminance and color difference pictures when performing chroma mapping.

[0074] FIG. 51 shows a CCP model used in chroma mapping by a video signal processing device according to an embodiment of the present invention.

[0075] FIG. 52 shows a video signal processing device according to an embodiment of the present invention inducing samples outside the window area boundary in chroma mapping.

[0076] FIG. 53 shows a method in which a video signal processing device according to an embodiment of the present invention signals to the SPS whether chroma mapping is enabled.

[0077] FIG. 54 shows chroma mapping-related constraint flags in the general_constraint_info() syntax structure according to an embodiment of the present invention.

[0078] FIG. 55 shows whether the cross-mapping method according to an embodiment of the present invention is activated, signaled in the SPS.

[0079] The terms used in this specification have been selected to be as widely used as possible, taking into account their functions in the present invention; however, these may vary depending on the intent, convention, or emergence of new technologies of those skilled in the art. In addition, in certain cases, terms have been arbitrarily selected by the applicant, and in such cases, their meanings will be described in the relevant description of the invention. Therefore, it should be noted that the terms used in this specification should be interpreted based on their actual meanings and the overall content of this specification, rather than merely their names.

[0080] In this specification, 'A and / or B' may be interpreted as having the same meaning as 'comprising at least one of A or B'.

[0081] Some terms in this specification may be interpreted as follows. Depending on the case, "coding" may be interpreted as "encoding" or "decoding." In this specification, a device that performs encoding of a video signal to generate a video signal bitstream is referred to as an encoding device or an encoder, and a device that performs decoding of a video signal bitstream to restore a video signal is referred to as a decoding device or a decoder. Additionally, in this specification, "video signal processing device" is used as a conceptual term encompassing both encoders and decoders. "Information" is a term that includes values, parameters, coefficients, elements, etc., and since its meaning may be interpreted differently depending on the case, the present invention is not limited thereto. "Unit" is used to refer to a basic unit of image processing or a specific location of a picture, and refers to an image region that includes at least one of a luminance (luma) component and a chroma component. Additionally, 'block' refers to an image region containing specific components among luminance components and chrominance components (i.e., Cb and Cr). However, depending on the embodiment, terms such as 'unit', 'block', 'partition', 'signal', and 'region' may be used interchangeably. Furthermore, in this specification, 'current block' refers to a block scheduled for current encoding, and 'reference block' refers to a block that has already been encoded or decoded and is used as a reference in the current block. Additionally, in this specification, terms such as 'luma', 'luminance', and 'Y' may be used interchangeably. Furthermore, in this specification, terms such as 'chroma', 'chroma', 'chrominance', and 'Cb or Cr' may be used interchangeably, and since the chrominance component is divided into two types, Cb and Cr, each chrominance component may be used separately. Additionally, in this specification, 'unit' may be used as a concept that includes a coding unit, a prediction unit, and a transformation unit."Picture" refers to a field or a frame, and depending on the embodiment, these terms may be used interchangeably. Specifically, if the captured image is an interlaced image, a single frame is separated into an odd (or odd, top) field and an even (or even, bottom) field, and each field is configured as a single picture unit for encoding or decoding. If the captured image is a progressive image, a single frame is configured as a picture for encoding or decoding. Furthermore, in this specification, terms such as "error signal," "residual signal," "residual signal," "residual signal," and "difference signal" may be used interchangeably. Additionally, in this specification, terms such as "intra-prediction mode," "intra-prediction directional mode," "in-frame prediction mode," and "in-frame prediction directional mode" may be used interchangeably. Furthermore, in this specification, terms such as "motion" and "movement" may be used interchangeably. Additionally, in this specification, 'left', 'upper left', 'upper', 'upper right', 'right', 'lower right', 'lower side', and 'lower left' may be used interchangeably with 'left end', 'upper left end', 'top', 'upper right end', 'right end', 'lower right end', 'bottom', and 'lower left end'. Also, 'element' and 'member' may be used interchangeably. POC (Picture Order Count) represents the temporal position information of a picture (or frame), may be the playback order displayed on the screen, and each picture may have a unique POC. Additionally, in this specification, the size of a block may be the sum or product of the width and height of the block. Alternatively, the size of a block may represent the number of samples within the block. Bit depth may be a range of sample values ​​expressed in bit units. Specifically, if the bit depth is 8 bits, the range of sample values ​​can be from 0 to 255.Internal bit depth can represent the bit depth when the bit depth is expanded to effectively encode an image in a video signal processing device. A video signal processing device can expand the bit depth before encoding the image and reduce the bit depth to the bit depth of the original input image when outputting after decoding. For example, a video signal processing device expands an image with an 8-bit depth to an image with a 10-bit depth to perform encoding and decoding.

[0082] FIG. 1 is a schematic block diagram of a video signal encoding device (100) according to one embodiment of the present specification. Referring to FIG. 1, the encoding device (100) of the present invention includes a conversion unit (110), a quantization unit (115), an inverse quantization unit (120), an inverse conversion unit (125), a filtering unit (130), a prediction unit (150), and an entropy coding unit (160).

[0083] The conversion unit (110) converts the residual signal, which is the difference between the input video signal and the prediction signal generated by the prediction unit (150), to obtain a conversion coefficient value. For example, the Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), or Wavelet Transform may be used. The Discrete Cosine Transform and Discrete Sine Transform divide the input picture signal into blocks to perform the conversion. In the conversion, the coding efficiency may vary depending on the distribution and characteristics of the values ​​within the conversion area. The conversion kernel used for the conversion of the residual block may be a conversion kernel having separable characteristics of vertical conversion and horizontal conversion. In this case, the conversion of the residual block can be performed by separating it into vertical conversion and horizontal conversion. For example, the encoder can perform a vertical conversion by applying the conversion kernel in the vertical direction of the residual block. Additionally, the encoder may perform a horizontal transformation by applying a transformation kernel in the horizontal direction of the residual block. In this disclosure, the term "transformation kernel" may be used to refer to a set of parameters used for transforming a residual signal, such as a transformation matrix, a transformation array, a transformation function, or a transformation. For example, the transformation kernel may be any one of a plurality of available kernels. Furthermore, transformation kernels based on different transformation types may be used for the vertical transformation and the horizontal transformation, respectively.

[0084] Transformation coefficients are distributed such that higher coefficients are found towards the top-left corner of the block, while coefficients closer to '0' are found towards the bottom-right corner. As the current block size increases, there is a higher likelihood of '0' coefficients existing in the bottom-right region. To reduce the transformation complexity of large blocks, only an arbitrary top-left region can be retained, and the remaining regions can be reset to '0'.

[0085] Additionally, an error signal may exist only in some regions of a coding block. In this case, the conversion process may be performed only on some arbitrary regions. As an example of implementation, in a block of size 2Nx2N, an error signal may exist only in the first 2NxN block, and the conversion process may be performed only on the first 2NxN block, but the second 2NxN block may not be encoded or decoded without the conversion process being performed. Here, N can be any positive integer.

[0086] The encoder may perform an additional transformation before the transformation coefficients are quantized. The aforementioned transformation method is referred to as a primary transform, and the additional transformation may be referred to as a secondary transform. The secondary transform may be optional for each residual block. According to one embodiment, the encoder may improve coding efficiency by performing a secondary transform on regions where it is difficult to concentrate energy in the low-frequency region using only the primary transform. For example, a secondary transform may be additionally performed on blocks where residual values ​​appear significantly in directions other than the horizontal or vertical direction of the residual block. Unlike the primary transform, the secondary transform may not be performed separately into a vertical transform and a horizontal transform. Such a secondary transform may be referred to as a Low Frequency Non-Separable Transform (LFNST).

[0087] The quantization unit (115) quantizes the conversion coefficient value output from the conversion unit (110).

[0088] To increase coding efficiency, instead of coding the picture signal as is, a method is used to predict the picture using an already coded region through a prediction unit (150), and to obtain a restored picture by adding the residual value between the original picture and the predicted picture to the predicted picture. To prevent mismatches from occurring in the encoder and decoder, information that is also available in the decoder must be used when performing prediction in the encoder. To this end, the encoder performs a process of restoring the currently encoded block. The inverse quantization unit (120) inversely quantizes the transform coefficient value, and the inverse transform unit (125) restores the residual value using the inversely quantized transform coefficient value. Meanwhile, the filtering unit (130) performs filtering operations to improve the quality of the restored picture and enhance coding efficiency. For example, a deblocking filter, a Sample Adaptive Offset (SAO), and an adaptive loop filter may be included. The filtered picture is stored in a decoded picture buffer (DPB, 156) to be output or used as a reference picture.

[0089] A deblocking filter is a filter designed to remove distortion within blocks generated at the boundaries between blocks in a restored picture. The encoder can determine whether to apply a deblocking filter to a given boundary based on the distribution of pixels within a few columns or rows relative to an arbitrary edge within the block. When a deblocking filter is applied to a block, the encoder can apply a Long Filter, Strong Filter, or Weak Filter depending on the filtering intensity. Additionally, horizontal and vertical filtering can be processed in parallel. Sample Adaptive Offset (SAO) can be used to correct the offset from the original image on a pixel-by-pixel basis for residual blocks to which the deblocking filter has been applied. To correct the offset for a specific picture, the encoder can use a Band Offset method, which divides the pixels contained in the image into a certain number of regions, determines the region to be corrected, and applies the offset to that region. Alternatively, the encoder may use an Edge Offset method, which applies an offset by considering the edge information of each pixel. Cross-component SAO (CC-SAO) is a method for compensating samples. In CC-SAO, similar to conventional SAO, the video signal processor classifies the reconstructed samples into categories and derives an offset for each category to add to the reconstructed samples. While conventional SAO uses only the luminance and chrominance components, CC-SAO classifies categories using all three components: luminance and two chrominance components. An Adaptive Loop Filter (ALF) is a method that divides pixels included in an image into specific groups, determines a single filter to be applied to each group, and performs differential filtering for each group.Information regarding whether to apply ALF can be signaled at the coding unit level, and the shape and coefficients of the ALF filter to be applied may vary depending on each block. Additionally, an ALF filter of the same form (fixed form) may be applied regardless of the characteristics of the target block. Bilateral Filter (BF) filtering is a method that applies a sample-level offset derived from the difference in change between neighboring samples on a sample-by-sample basis.

[0090] The prediction unit (150) includes an intra prediction unit (152) and an inter prediction unit (154). The intra prediction unit (152) performs intra prediction within the current picture, and the inter prediction unit (154) performs inter prediction by predicting the current picture using a reference picture stored in the decoded picture buffer (156). The intra prediction unit (152) performs intra prediction from restored regions within the current picture and transmits intra encoding information to the entropy coding unit (160). The intra encoding information may include at least one of an intra prediction mode, an MPM (Most Probable Mode) flag, an MPM index, and information regarding a reference sample. The inter prediction unit (154) may again be configured to include a motion estimation unit (154a) and a motion compensation unit (154b). The motion estimation unit (154a) refers to a specific area of ​​the restored reference picture to find the part most similar to the current area and obtains a motion vector value, which is the distance between the areas. Motion information for the reference area obtained by the motion estimation unit (154a), such as reference direction indicator information (L0 prediction, L1 prediction, bidirectional prediction), reference picture index, motion vector information, etc., is transmitted to the entropy coding unit (160) so that it can be included in the bitstream. Using the motion information transmitted from the motion estimation unit (154a), the motion compensation unit (154b) performs inter-motion compensation to generate a prediction block for the current block. The inter-prediction unit (154) transmits inter-coding information containing motion information for the reference area to the entropy coding unit (160).

[0091] According to an additional embodiment, the prediction unit (150) may include an intra block copy (IBC) prediction unit (not shown). The IBC prediction unit performs IBC prediction from restored samples within the current picture and transmits IBC encoding information to the entropy coding unit (160). The IBC prediction unit obtains a block vector value indicating a reference area used for predicting the current area by referencing a specific area within the current picture. The IBC prediction unit may perform IBC prediction using the obtained block vector value. The IBC prediction unit transmits IBC encoding information to the entropy coding unit (160). The IBC encoding information may include at least one of size information of the reference area and block vector information (index information for predicting the block vector of the current block within the motion candidate list, block vector difference information).

[0092] When the above picture prediction is performed, the conversion unit (110) converts the residual value between the original picture and the predicted picture to obtain a conversion coefficient value. At this time, the conversion can be performed in units of specific blocks within the picture, and the size of the specific block can be varied within a preset range. The quantization unit (115) quantizes the conversion coefficient value generated by the conversion unit (110) and transmits the quantized conversion coefficient to the entropy coding unit (160).

[0093] The quantized transformation coefficients in the form of a two-dimensional array described above can be rearranged into a one-dimensional array for entropy coding. The method of scanning the quantized transformation coefficients can be determined by the size of the transformation block and the in-frame prediction mode, which scanning method will be used. As an example of implementation, diagonal, vertical, and horizontal scans may be applied. This scan information can be signaled in block units and can be derived according to pre-determined rules.

[0094] The entropy coding unit (160) generates a video signal bitstream by entropy coding information representing quantized conversion coefficients, intra-coding information, and inter-coding information. In the entropy coding unit (160), methods such as Variable Length Coding (VLC) and arithmetic coding may be used. Variable Length Coding (VLC) converts input symbols into a series of codewords, and the length of the codewords may be variable. For example, frequently occurring symbols are represented as short codewords, and infrequently occurring symbols are represented as long codewords. As a variable length coding method, Context-based Adaptive Variable Length Coding (CAVLC) may be used. Arithmetic coding converts a series of data symbols into a single prime number using the probability distribution of each data symbol, and arithmetic coding can obtain the optimal prime number bit required to represent each symbol. Context-based Adaptive Binary Arithmetic Coding (CABAC) can be used as arithmetic coding.

[0095] CABAC is a method of binary arithmetic encoding that utilizes multiple context models generated based on probabilities obtained through experiments. A context model can also be referred to as a context model. First, if a symbol is not in binary form, the encoder binarizes each symbol using tools such as exp-Golomb. Binarized 0s or 1s can be described as bins. The CABAC initialization process is divided into context initialization and arithmetic coding initialization. Context initialization is the process of initializing the occurrence probability of each symbol, which is determined by the symbol type, quantization parameters (QP), and slice type (whether it is I, P, or B). A context model possessing this initialization information can use probability-based values ​​obtained through experiments. The context model provides the occurrence probability of the LPS (Least Probable Symbol) or MPS (Most Probable Symbol) for the symbol currently being coded, as well as information (valMPS) regarding which bin value (0 or 1) corresponds to the MPS. One of several context models is selected through the context index (ctxIdx), and the context index can be derived from information about the block currently to be encoded or information about surrounding blocks. Initialization for binary arithmetic coding is performed based on the probability model selected from the context model. Binary arithmetic encoding proceeds by dividing into probability intervals based on the occurrence probabilities of 0 and 1, and then making the probability interval corresponding to the bin to be processed the entire probability interval for the next bin to be processed. Location information within the probability interval where the last bin was processed is output. However, since the probability interval cannot be divided indefinitely, if it shrinks to within a certain size, a renormalization process is performed to widen the probability interval and output the corresponding location information. Additionally, after each bin is processed, a probability update process may be performed to newly set the probability for the next bin to be processed based on the information of the processed bin.A video signal processing device can entropy-code by setting the probability of 0 or 1 occurring without a context model to '0.5', and this can be described as a bypass mode.

[0096] A bitstream may consist of one or more coded video sequences (CVS), and a single CVS may be encoded independently of the others. Each CVS may consist of one or more layers, and each layer may represent a specific image quality or resolution, or a general image, a depth information map, or a transparency map. Additionally, a coded layer video sequence (CLVS) may refer to a layer-wise CVS composed of consecutive (in decoding order) PUs within the same layer. For example, a CLVS for a specific image quality layer may exist, as well as a CLVS for a depth information map.

[0097] The bitstream generated above is encapsulated into Network Abstraction Layer (NAL) units as the basic unit. NAL units are classified into Video Coding Layer (VCL) NAL units containing video data and non-VCL NAL units containing parameter information for decoding video data, and various types of VCL or non-L NAL units exist. A NAL unit consists of NAL header information and Raw Byte Sequence Payload (RBSP), which is the data; the NAL header information includes summary information regarding the RBSP. The RBSP of a VCL NAL unit contains an integer number of encoded coding tree units. To decode the bitstream in a video decoder, the bitstream must first be separated into NAL units, and then each separated NAL unit must be decoded. Meanwhile, information required for decoding a video signal bitstream can be transmitted by including it in a Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptation Parameter Set (APS), etc. The RBSP of the VCL NAL unit may include an integer number of coding tree units. The VPS is a parameter set composed of common syntax by extracting duplicate parameters from the SPS parameter set signaled at each layer in a bitstream that supports image quality, resolution, frame rate scalability, or a bitstream that supports multiview.SPS is a parameter set that includes at least one of the following: a profile containing information on acceptable coding tools (or algorithms) and image formats; a level containing information on the decoder's processing capability regarding processable image resolution, frame rate, and acceptable memory size; a tier containing information on the maximum processable bit rate; and information on image resolution, bit depth, and whether a feature is enabled. PPS is a parameter set that includes at least one of the following: image resolution, tile partitioning information, information on whether weight prediction is enabled, quantization parameters, and filtering-related information. APS is a parameter set that includes one of the following information depending on the APS type: ALF filter coefficient information, LMCS-related parameters, and quantization scale parameters. APS is divided into a prefix APS that is signaled before the VCL NAL unit and a suffix APS that is signaled after the VCL NAL unit, and in the case of ALF APS, since it is efficient to apply the ALF filter coefficients derived from the previous picture to the next picture, it can be signaled as a suffix APS.

[0098] Meanwhile, the block diagram of FIG. 1 shows an encoding device (100) according to one embodiment of the present specification, and the separated blocks represent the elements of the encoding device (100) logically distinguished. Accordingly, the elements of the aforementioned encoding device (100) may be mounted as a single chip or as a plurality of chips depending on the design of the device. According to one embodiment, the operation of each element of the aforementioned encoding device (100) may be performed by a processor (not shown).

[0099] FIG. 2 is a schematic block diagram of a video signal decoding device (200) according to one embodiment of the present specification. Referring to FIG. 2, the decoding device (200) of the present invention includes an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (225), a filtering unit (230), and a prediction unit (250).

[0100] The entropy decoding unit (210) entropies decodes the video signal bitstream to extract transform coefficient information, intra-coding information, inter-coding information, etc. for each region. For example, the entropy decoding unit (210) can obtain a binary code for transform coefficient information of a specific region from the video signal bitstream. Additionally, the entropy decoding unit (210) inversely binarizes the binary code to obtain quantized transform coefficients. The inverse quantization unit (220) inversely quantizes the quantized transform coefficients, and the inverse transform unit (225) restores the residual value using the inversely quantized transform coefficients. The video signal processing device (200) restores the original pixel value by summing the residual value obtained from the inverse transform unit (225) with the predicted value obtained from the prediction unit (250).

[0101] Meanwhile, the filtering unit (230) performs filtering on the picture to improve image quality. This may include a deblocking filter to reduce block distortion and / or an adaptive loop filter to remove distortion from the entire picture. The filtered picture is stored in a decoded picture buffer (DPB, 256) to be output or used as a reference picture for the next picture.

[0102] The prediction unit (250) includes an intra prediction unit (252) and an inter prediction unit (254). The prediction unit (250) generates a prediction picture by utilizing the encoding type decoded through the aforementioned entropy decoding unit (210), the conversion coefficient for each region, and intra / inter encoding information. To restore the current block in which decoding is performed, the decoded regions of the current picture containing the current block or other pictures may be used. A picture (or tile / slice) that uses only the current picture for restoration, i.e., performs intra prediction or intra BC prediction, is called an intra picture or I picture (or tile / slice), and a picture (or tile / slice) capable of performing intra prediction, inter prediction, and intra BC prediction is called an inter picture (or tile / slice). A picture (or tile / slice) that uses at most one motion vector and reference picture index to predict sample values ​​of each block among inter-pictures (or tiles / slices) is called a predictive picture or P picture (or tile / slice), and a picture (or tile / slice) that uses at most two motion vectors and reference picture indices is called a bi-predictive picture or B picture (or tile / slice). In other words, a P picture (or tile / slice) uses at most one set of motion information to predict each block, and a B picture (or tile / slice) uses at most two sets of motion information to predict each block. Here, a set of motion information includes one or more motion vectors and one reference picture index.

[0103] The intra prediction unit (252) generates a prediction block using intra encoding information and restored samples within the current picture. Specifically, samples within the current block can be predicted using reference samples derived using the sample location within the current block and the directionality of the intra prediction mode. If the location of the reference sample is not an integer unit sample, the video signal processing device can predict the current block samples using interpolated reference samples through an interpolation method using adjacent reference samples. This may be referred to as linear-based intra prediction. As described above, the intra encoding information may include at least one of an intra prediction mode, a Most Probable Mode (MPM) flag, and an MPM index. The intra prediction unit (252) predicts the sample values ​​of the current block using restored samples located to the left and / or above the current block as reference samples. In the present disclosure, the restored samples, reference samples, and samples of the current block may represent pixels. Additionally, the sample values ​​may represent pixel values.

[0104] According to one embodiment, reference samples may be samples included in the surrounding blocks of the current block. For example, the reference samples may be samples adjacent to the left boundary of the current block and / or samples adjacent to the upper boundary. Additionally, the reference samples may be samples among the samples of the surrounding blocks of the current block located on a line within a preset distance from the left boundary of the current block and / or samples located on a line within a preset distance from the upper boundary of the current block. In this case, the surrounding blocks of the current block may include at least one of a left (L) block, an upper (A) block, a lower left (BL) block, an upper right (AR) block, or an upper left (AL) block adjacent to the current block. The surrounding blocks of the current block may be reference blocks for predicting the current block.

[0105] The inter-prediction unit (254) generates a prediction block using the reference picture and inter-coding information stored in the decoding picture buffer (256). The inter-coding information may include a set of motion information (reference picture index, motion vector information, etc.) of the current block for the reference block. Inter-prediction may include L0 prediction, L1 prediction, and bi-prediction. L0 prediction refers to a prediction using one reference picture included in the L0 picture list, and L1 prediction refers to a prediction using one reference picture included in the L1 picture list. For this, one set of motion information (e.g., motion vector and reference picture index) may be required. In the bi-prediction method, up to two reference regions may be used, and these two reference regions may exist in the same reference picture or in different pictures. That is, in the pair prediction method, up to two sets of motion information (e.g., motion vectors and reference picture indices) may be used, and the two motion vectors may correspond to the same reference picture index or different reference picture indices. In this case, the reference pictures are pictures located temporally before or after the current picture, and may be pictures that have already been restored and completed. According to one embodiment, the two reference regions used in the pair prediction method may be regions selected from the L0 picture list and the L1 picture list, respectively. Additionally, a prediction method that uses only reference pictures with a POC smaller than the current picture's POC or only reference pictures with a POC larger than the current picture's POC, based on the POC (picture order count) indicating the display order of the current picture, can be called unidirectional prediction.In addition, a prediction method that uses both a reference picture with a POC smaller than the current picture's POC and a reference picture with a POC larger than the current picture's POC, based on the POC (picture order count) representing the display order of the current picture, can be called bi-directional prediction. A prediction method that uses only one reference picture in unidirectional prediction can be called uni-prediction, and a prediction method that uses two reference pictures in unidirectional prediction can be called bi-prediction or paired prediction.

[0106] The inter prediction unit (254) can obtain a reference block of the current block using a motion vector and a reference picture index. The reference block exists within the reference picture corresponding to the reference picture index. Additionally, a sample value of the block specified by the motion vector or an interpolated value thereof can be used as a predictor of the current block. For motion prediction with sub-pel unit pixel accuracy, for example, an 8-tap interpolation filter may be used for the luminance signal and a 4-tap interpolation filter may be used for the chrominance signal. However, the interpolation filter for sub-pel unit motion prediction is not limited thereto. In this way, the inter prediction unit (254) performs motion compensation to predict the texture of the current unit from a previously restored picture. At this time, the inter prediction unit may use a set of motion information.

[0107] According to an additional embodiment, the prediction unit (250) may include an IBC prediction unit (not shown). The IBC prediction unit may restore a current region by referring to a specific region containing restored samples within the current picture. The IBC prediction unit may perform IBC prediction using IBC encoding information obtained from the entropy decoding unit (210). The IBC encoding information may include block vector information.

[0108] A restored video picture is generated by adding the predicted value output from the intra prediction unit (252) or the inter prediction unit (254) and the residual value output from the inverse transformation unit (225). That is, the video signal decoding device (200) restores the current block using the predicted block generated by the prediction unit (250) and the residual obtained from the inverse transformation unit (225).

[0109] Meanwhile, the block diagram of FIG. 2 shows a decoding device (200) according to one embodiment of the present specification, and the separated blocks represent the elements of the decoding device (200) logically distinguished. Accordingly, the elements of the aforementioned decoding device (200) may be mounted as a single chip or multiple chips depending on the design of the device. According to one embodiment, the operation of each element of the aforementioned decoding device (200) may be performed by a processor (not shown).

[0110] Meanwhile, the technology proposed in this specification is applicable to both the methods and devices of encoders and decoders, and the parts described as signaling and parsing may be described for convenience of explanation. Generally, signaling can be described as encoding each syntax from the perspective of an encoder, and parsing as interpreting each syntax from the perspective of a decoder. That is, each syntax can be signaled by being included in a bitstream from the encoder, and the decoder can parse the syntax and use it in the restoration process. In this case, a sequence of bits for each syntax arranged according to a defined hierarchical configuration can be referred to as a bitstream.

[0111] A single picture can be divided and encoded into sub-pictures, slices, tiles, etc. A sub-picture may contain one or more slices or tiles. When a single picture is divided and encoded into multiple slices or tiles, it can only be displayed on the screen after all slices or tiles within the picture have been decoded. Conversely, when a single picture is encoded into multiple sub-pictures, only any sub-picture may be decoded and displayed on the screen. A slice may contain multiple tiles or sub-pictures. Alternatively, a tile may contain multiple sub-pictures or slices. Since sub-pictures, slices, and tiles can be encoded or decoded independently of each other, they are effective for parallel processing and improving processing speed. However, there is a disadvantage in that the bit size increases because the encoded information of adjacent sub-pictures, slices, or tiles cannot be utilized. Sub-pictures, slices, and tiles can be divided and encoded into multiple Coding Tree Units (CTUs).

[0112] FIG. 3 illustrates an embodiment in which a Coding Tree Unit (CTU) is divided into Coding Units (CUs) within a picture. During the coding process of a video signal, the picture may be divided into a sequence of Coding Tree Units (CTUs). A Coding Tree Unit may consist of a Luminance Coding Tree Block (CTB), two Chroma Coding Tree Blocks, and its encoded syntax information. A single Coding Tree Unit may consist of a single Coding Unit, or a single Coding Tree Unit may be divided into multiple Coding Units. A single Coding Unit may consist of a Luminance Coding Block (CB), two Chroma Coding Blocks, and its encoded syntax information. A single Coding Block may be divided into multiple Sub-Coding Blocks. A single Coding Unit may consist of a Transform Unit (TU), or a single Coding Unit may be divided into multiple Transform Units. A single transform unit may consist of a luminance transform block (Transform Block, TB), two chrominance transform blocks, and their encoded syntax information. A coding tree unit may be divided into multiple coding units. A coding tree unit may not be divided and may become a leaf node. In this case, the coding tree unit itself may become a coding unit.

[0113] A coding unit refers to a basic unit for processing a picture during the video signal processing process described above, namely intra / inter prediction, transformation, quantization, and / or entropy coding. Within a single picture, the size and shape of the coding unit may not be constant. The coding unit may have a square or rectangular shape. A rectangular coding unit (or rectangular block) includes a vertical coding unit (or vertical block) and a horizontal coding unit (or horizontal block). In this specification, a vertical block is a block in which the height is greater than the width, and a horizontal block is a block in which the width is greater than the height. Additionally, in this specification, a non-square block may refer to a rectangular block, but the invention is not limited thereto.

[0114] Referring to FIG. 3, the coding tree unit is first divided into a Quad Tree (QT) structure. That is, in the Quad Tree structure, a single node with a size of 2NX2N can be divided into four nodes with a size of NXN. In this specification, a Quad Tree may also be referred to as a quaternary tree. Quad Tree division can be performed recursively, and not all nodes need to be divided to the same depth.

[0115] Meanwhile, the leaf nodes of the aforementioned quad tree can be further divided into a Multi-Type Tree (MTT) structure. According to an embodiment of the present invention, in a multi-type tree structure, a single node can be divided into a binary or ternary tree structure by horizontal or vertical division. That is, in a multi-type tree structure, there are four division structures: vertical binary division, horizontal binary division, vertical ternary division, and horizontal ternary division. According to an embodiment of the present invention, in each of the tree structures, the width and height of the node can both have powers of 2 values. For example, in a binary tree (BT) structure, a node of size 2NX2N can be divided into two NX2N nodes by vertical binary division and into two 2NXN nodes by horizontal binary division. In addition, in a ternary tree (TT) structure, a node of size 2NX2N can be divided into (N / 2)X2N, NX2N, and (N / 2)X2N nodes by vertical ternary partitioning, and into 2NX(N / 2), 2NXN, and 2NX(N / 2) nodes by horizontal ternary partitioning. This multi-type tree partitioning can be performed recursively.

[0116] Leaf nodes of a multi-type tree can be coding units. If a coding unit is not large compared to the maximum transformation length, the coding unit can be used as a unit of prediction and / or transformation without further splitting. In one embodiment, if the width or height of the current coding unit is greater than the maximum transformation length, the current coding unit can be split into multiple transformation units without explicit signaling regarding splitting. Meanwhile, in the aforementioned quad tree and multi-type tree, at least one of the following parameters may be predefined or transmitted via an RBSP of a higher-level set such as PPS, SPS, VPS, etc. 1) CTU size: Size of the quad tree root node, 2) MinQtSize: Allowed minimum QT leaf node size, 3) MaxBtSize: Allowed maximum BT root node size, 4) MaxTtSize: Allowed maximum TT root node size, 5) MaxMttDepth: Maximum allowed depth of MTT splitting from the QT leaf node, 6) MinBtSize: Allowed minimum BT leaf node size, 7) MinTtSize: Allowed minimum TT leaf node size.

[0117] FIG. 4 illustrates an embodiment of a method for signaling the splitting of a quad tree and a multi-type tree. Pre-configured flags may be used to signal the splitting of the aforementioned quad tree and multi-type tree. Referring to FIG. 4, at least one of the following may be used: a flag 'split_cu_flag' indicating whether a node is split, a flag 'split_qt_flag' indicating whether a quad tree node is split, a flag 'mtt_split_cu_vertical_flag' indicating the splitting direction of a multi-type tree node, or a flag 'mtt_split_cu_binary_flag' indicating the splitting shape of a multi-type tree node.

[0118] According to an embodiment of the present invention, a flag 'split_cu_flag' indicating whether the current node is split may be signaled first. If the value of 'split_cu_flag' is 0, it indicates that the current node is not split, and the current node becomes a coding unit. If the current node is a coding tree unit, the coding tree unit includes one unsplit coding unit. If the current node is a quad tree node 'QT node', the current node is a quad tree leaf node 'QT leaf node' and becomes a coding unit. If the current node is a multi-type tree node 'MTT node', the current node is a multi-type tree leaf node 'MTT leaf node' and becomes a coding unit.

[0119] When the value of 'split_cu_flag' is 1, the current node may be split into nodes of a quad tree or a multi-type tree according to the value of 'split_qt_flag'. A coding tree unit is the root node of a quad tree and may be split first into a quad tree structure. In a quad tree structure, 'split_qt_flag' is signaled for each node 'QT node'. When the value of 'split_qt_flag' is 1, the node is split into four square nodes, and when the value of 'split_qt_flag' is 0, the node becomes a quad tree leaf node 'QT leaf node' and is split into multi-type nodes. According to an embodiment of the present invention, quad tree splitting may be restricted depending on the type of the current node. Quad tree splitting may be allowed when the current node is a coding tree unit (root node of a quad tree) or a quad tree node, and quad tree splitting may not be allowed when the current node is a multi-type tree node. Each quad tree leaf node 'QT leaf node' can be further split into a multi-type tree structure. As described above, if 'split_qt_flag' is 0, the current node can be split into multi-type nodes. To indicate the split direction and split shape, 'mtt_split_cu_vertical_flag' and 'mtt_split_cu_binary_flag' can be signaled. If the value of 'mtt_split_cu_vertical_flag' is 1, a vertical split of node 'MTT node' is indicated, and if the value of 'mtt_split_cu_vertical_flag' is 0, a horizontal split of node 'MTT node' is indicated.Also, when the value of 'mtt_split_cu_binary_flag' is 1, the node 'MTT node' is split into 2 rectangular nodes, and when the value of 'mtt_split_cu_binary_flag' is 0, the node 'MTT node' is split into 3 rectangular nodes.

[0120] In a tree partitioning structure, luminance blocks and chrominance blocks can be partitioned in the same form. That is, a chrominance block can partition itself by referencing the partitioning form of the luminance block. If the current chrominance block is smaller than an arbitrarily defined size, the chrominance block may not be partitioned even if the luminance block has been partitioned.

[0121] The luminance block and the chrominance block can have the same tree partitioning structure, which can be referred to as a Single Tree. When the current block is encoded as a Single Tree, the partitioning structure, encoding mode information, and motion information of the luminance and chrominance blocks may be identical, while other information related to error signals may differ between the luminance and chrominance blocks. Additionally, the luminance and chrominance blocks can have different tree partitioning structures, which can be referred to as a Dual Tree. When the current block is encoded as a Dual Tree, at least one of the partitioning structure, encoding mode information, and motion information of the luminance and chrominance blocks may differ.

[0122] There may be a close correlation between the luminance block and the chrominance block corresponding to the luminance block. Therefore, in encoders and decoders, when the current block is encoded and decoded as a dual tree, the chrominance block can be encoded using the luminance block's partitioning information, encoding mode information, motion information, etc.

[0123] A node to be divided into the smallest unit can be processed as a single coding block. If the current block is a coding block, the coding block can be divided into multiple sub-blocks (sub-coding blocks), and the prediction information of each sub-block may be the same or different. As an example of implementation, if the coding unit is in intra mode, the intra prediction mode of each sub-block may be the same or different. Also, if the coding unit is in inter mode, the movement information of each sub-block may be the same or different. Additionally, each sub-block may be capable of encoding or decoding independently of one another. Each sub-block can be distinguished through a sub-block index (sbIdx). Furthermore, when the coding unit is divided into sub-blocks, it may be divided horizontally or vertically, or diagonally. The mode in which the current coding unit is divided into two or four sub-blocks in the horizontal or vertical direction in intra mode is called ISP (Intra Sub Partitions). The mode in which the current coding block is divided diagonally in inter mode is called GPM (Geometric partitioning mode). In GPM mode, the position and direction of the diagonal line are determined using a predetermined angle table, and the index information of the angle table is signaled.

[0124] Motion information may include one or more of reference direction indicator information, reference picture information, motion vector, motion resolution, affine model, CPMV (control point motion vector), block vector, block vector resolution, MHP information, LIC information, filtering information, BCW information, and RRIBC information.

[0125] Reference direction indicator information consists of L0 prediction, L1 prediction, and L0 and L1 predictions; L0 and L1 predictions are single predictions and unidirectional predictions, while L0 and L1 predictions are bidirectional predictions. Additionally, L0 and L1 predictions can be either unidirectional or bidirectional predictions. Here, L0 prediction is performed using reference pictures within the L0 reference picture list, and L1 prediction can be performed using reference pictures within the L1 reference picture list. In the L0 reference picture list, reference pictures with a POC smaller than the current picture's POC can be added based on the current picture's POC. Furthermore, the L0 reference picture list can be organized in order from reference pictures with a POC closer to the current picture's POC to reference pictures with a POC further away. In the L1 reference picture list, reference pictures with a POC larger than the current picture's POC can be added based on the current picture's POC. Additionally, the L1 reference picture list may be organized in order from reference pictures closer to the POC of the current picture to reference pictures further away from the POC. The L0 and L1 reference picture lists may vary for each slice, sub-picture, and picture. Furthermore, the L0 reference picture list may include reference pictures from the L1 reference picture list. And the L1 reference picture list may include reference pictures from the L0 reference picture list.

[0126] Reference picture information may vary per block and may be index information indicating which reference picture from the L0 reference picture list and / or L1 reference picture list the current block is predicted to use. Reference picture information may include one or more of L0 reference picture information and L1 reference picture information.

[0127] A motion vector is information indicating the block in the reference picture that best matches the current block; it is a value representing the distance from the top-left position of the current block within the picture to the reference block in horizontal and vertical coordinates.

[0128] Motion resolution represents the resolution of the motion vector, and motion resolution can be expressed in units of 4 pixels, 1 pixel (integer pixel), 1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel, 1 / 16 pixel, etc.

[0129] A block vector is information indicating the block that best matches the current block within the already restored area of ​​the current picture; it is a value representing the distance from the top-left position of the current block to the reference block in horizontal and vertical coordinates.

[0130] Block vector resolution can be expressed in units of 4 pixels, 1 pixel (integer pixel), 1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel, 1 / 16 pixel, etc.

[0131] MHP information may include whether additional motion information is applied and additional motion information.

[0132] LIC information may include whether LIC is applied to the current block.

[0133] Filtering information may include filtering type and coefficient information applied according to the movement resolution of the current block.

[0134] BCW information may include whether BCW is applied to the current block.

[0135] RRIBC information may include information on whether RRIBC is applied and the RRIBC type if the current block is encoded in IBC mode.

[0136] Picture prediction (motion compensation) for coding is performed on indivisible coding units (i.e., leaf nodes of coding tree units). The basic unit that performs this prediction is hereinafter referred to as a prediction unit or prediction block.

[0137] Hereinafter, the term "unit" as used in this specification may be used as a substitute for the prediction unit, which is the basic unit for performing predictions. However, the present invention is not limited thereto, and may be understood more broadly as a concept including the coding unit.

[0138] FIGS. 5 and 6 illustrate an intra prediction method according to an embodiment of the present invention in more detail. As described above, the intra prediction unit predicts sample values ​​of the current block using restored samples located to the left and / or above the current block as reference samples.

[0139] First, FIG. 5 illustrates an example of reference samples used for prediction of the current block in intra prediction mode. According to one example, the reference samples may be samples adjacent to the left boundary of the current block and / or samples adjacent to the upper boundary. As illustrated in FIG. 5, when the size of the current block is WXH and samples of a single reference line adjacent to the current block are used for intra prediction, the reference samples may be set using up to 2W+2H+1 surrounding samples located to the left and / or upper side of the current block.

[0140] Meanwhile, pixels of multiple reference lines may be used for intra prediction of the current block. A multiple reference line may consist of n lines located within a preset range from the current block. According to one embodiment, when pixels of multiple reference lines are used for intra prediction, separate index information indicating the lines to be set as reference pixels may be signaled, and this may be named a reference line index.

[0141] Additionally, if at least some of the samples to be used as reference samples have not yet been restored, the intra prediction unit may obtain reference samples by performing a reference sample padding process. Additionally, the intra prediction unit may perform a reference sample filtering process to reduce the error of the intra prediction. That is, filtered reference samples may be obtained by performing filtering on the surrounding samples and / or the reference samples obtained by the reference sample padding process. The intra prediction unit predicts the samples of the current block using the reference samples obtained in this manner. The intra prediction unit predicts the samples of the current block using unfiltered reference samples or filtered reference samples. In the present disclosure, surrounding samples may include samples on at least one reference line. For example, surrounding samples may include adjacent samples on a line adjacent to the boundary of the current block.

[0142] Next, FIG. 6 illustrates an example of prediction modes used for intra prediction. For intra prediction, intra prediction mode information indicating the direction of intra prediction may be signaled. The intra prediction mode information indicates one of a plurality of intra prediction modes constituting a set of intra prediction modes. If the current block is an intra prediction block, the decoder receives intra prediction mode information of the current block from the bitstream. The intra prediction unit of the decoder performs intra prediction for the current block based on the extracted intra prediction mode information.

[0143] According to an embodiment of the present invention, an intra prediction mode set may include all intra prediction modes used for intra prediction (e.g., a total of 67 intra prediction modes). More specifically, the intra prediction mode set may include a planar mode, a DC mode, and a plurality (e.g., 65) angle modes (i.e., direction modes). Each intra prediction mode may be indicated by a preset index (i.e., an intra prediction mode index). For example, as illustrated in FIG. 6, intra prediction mode index 0 indicates a planar mode, and intra prediction mode index 1 indicates a DC mode. Additionally, intra prediction mode indices 2 through 66 may each indicate different angle modes. Angle modes each indicate different angles within a preset angle range. For example, an angle mode may indicate an angle within an angle range between 45 degrees and -135 degrees clockwise (i.e., a first angle range). The angle mode may be defined based on the 12 o'clock direction. At this time, intra prediction mode index 2 indicates horizontal diagonal (HDIA) mode, intra prediction mode index 18 indicates horizontal (HOR) mode, intra prediction mode index 34 indicates diagonal (DIA) mode, intra prediction mode index 50 indicates vertical (VER) mode, and intra prediction mode index 66 indicates vertical diagonal (VDIA) mode.

[0144] Meanwhile, the pre-set angle ranges may be set differently depending on the shape of the current block. For example, if the current block is a rectangular block, a wide angle mode indicating an angle exceeding 45 degrees or less than -135 degrees clockwise may be additionally used. If the current block is a horizontal block, the angle mode may indicate an angle within an angle range between (45+offset1) degrees and (-135+offset1) degrees clockwise (i.e., a second angle range). In this case, angle modes 67 to 76 that fall outside the first angle range may be additionally used. Additionally, if the current block is a vertical block, the angle mode may indicate an angle within an angle range between (45-offset2) degrees and (-135-offset2) degrees clockwise (i.e., a third angle range). In this case, angle modes -10 to -1 that fall outside the first angle range may be additionally used. According to an embodiment of the present invention, the values ​​of offset1 and offset2 may be determined differently depending on the ratio between the width and height of the rectangular block. Also, offset1 and offset2 can be positive.

[0145] According to a further embodiment of the present invention, a plurality of angle modes constituting an intra-prediction mode set may include a basic angle mode and an extended angle mode. In this case, the extended angle mode may be determined based on the basic angle mode.

[0146] According to one embodiment, the basic angle mode is a mode corresponding to the angle used in the intra prediction of the existing HEVC (High Efficiency Video Coding) standard, and the extended angle mode may be a mode corresponding to the angle newly added in the intra prediction of the next-generation video codec standard. More specifically, the basic angle mode may be an angle mode corresponding to any one of the intra prediction modes {2, 4, 6, …, 66}, and the extended angle mode may be an angle mode corresponding to any one of the intra prediction modes {3, 5, 7, …, 65}. That is, the extended angle mode may be an angle mode between the basic angle modes within the first angle range. Accordingly, the angle indicated by the extended angle mode may be determined based on the angle indicated by the basic angle mode.

[0147] According to another embodiment, the basic angle mode is a mode corresponding to an angle within a preset first angle range, and the extended angle mode may be a wide angle mode that is outside the first angle range. That is, the basic angle mode is an angle mode corresponding to any one of the intra-prediction modes {2, 3, 4, …, 66}, and the extended angle mode may be an angle mode corresponding to any one of the intra-prediction modes {-14, -13, -12, …, -1} and {67, 68, …, 80}. The angle indicated by the extended angle mode may be determined as the opposite angle of the angle indicated by the corresponding basic angle mode. Accordingly, the angle indicated by the extended angle mode may be determined based on the angle indicated by the basic angle mode. Meanwhile, the number of extended angle modes is not limited thereto, and additional extended angles may be defined depending on the size and / or shape of the current block. Meanwhile, the total number of intra-prediction modes included in the set of intra-prediction modes may vary depending on the configuration of the aforementioned basic angle mode and extended angle mode.

[0148] In the above embodiment, the interval between the extended angle modes can be set based on the interval between the corresponding basic angle modes. For example, the interval between the extended angle modes {3, 5, 7, … , 65} can be determined based on the interval between the corresponding basic angle modes {2, 4, 6, … , 66}. Additionally, the interval between the extended angle modes {-14, -13, … , -1} is determined based on the interval between the corresponding opposite basic angle modes {53, 53, … , 66}, and the interval between the extended angle modes {67, 68, … , 80} can be determined based on the interval between the corresponding opposite basic angle modes {2, 3, 4, … , 15}. The angle interval between the extended angle modes can be set to be equal to the angle interval between the corresponding basic angle modes. Additionally, the number of extended angle modes in the intra-prediction mode set can be set to be less than or equal to the number of basic angle modes.

[0149] According to an embodiment of the present invention, an extended angle mode may be signaled based on a basic angle mode. For example, a wide angle mode (i.e., an extended angle mode) may replace at least one angle mode (i.e., a basic angle mode) within a first angle range. The replaced basic angle mode may be an angle mode corresponding to the opposite side of the wide angle mode. That is, the replaced basic angle mode is an angle mode corresponding to an angle in the opposite direction of the angle indicated by the wide angle mode, or an angle mode corresponding to an angle that differs from the angle in the opposite direction by a preset offset index. According to an embodiment of the present invention, the preset offset index is 1. An intra-prediction mode index corresponding to the replaced basic angle mode may be remapped to the wide angle mode to signal the corresponding wide angle mode. For example, wide angle modes {-14, -13, … , -1} may each be signaled by the intra-prediction mode index {52, 53, … , 66}, and wide angle modes {67, 68, … , 80} may be signaled by the intra-prediction mode index {2, 3, … Each can be signaled by { , 15}. In this way, by allowing the intra prediction mode index for the basic angle mode to signal the extended angle mode, the same set of intra prediction mode indexes can be used for signaling the intra prediction mode even if the configurations of the angle modes used for intra prediction in each block are different. Therefore, signaling overhead due to changes in the intra prediction mode configuration can be minimized.

[0150] Meanwhile, whether to use the extended angle mode may be determined based on at least one of the shape and size of the current block. According to one embodiment, if the size of the current block is larger than a preset size, the extended angle mode is used for intra prediction of the current block, and otherwise, only the basic angle mode is used for intra prediction of the current block. According to another embodiment, if the current block is a non-square block, the extended angle mode is used for intra prediction of the current block, and if the current block is a square block, only the basic angle mode is used for intra prediction of the current block.

[0151] The intra prediction unit determines the reference samples and / or interpolated reference samples to be used for intra prediction of the current block based on the intra prediction mode information of the current block. If the intra prediction mode index indicates a specific angle mode, the reference sample or interpolated reference sample corresponding to the specific angle from the current sample of the current block is used for the prediction of the current pixel. Therefore, different sets of reference samples and / or interpolated reference samples may be used for intra prediction depending on the intra prediction mode. Once the intra prediction of the current block is performed using the reference samples and intra prediction mode information, the decoder restores the sample values ​​of the current block by adding the residual signal of the current block obtained from the inverse transform unit to the intra prediction value of the current block.

[0152] The motion information used for inter-prediction may include reference direction indicator information (inter_pred_idc), reference picture indices (ref_idx_l0, ref_idx_l1), and motion vectors (mvL0, mvL1). Reference picture list utilization information (predFlagL0, predFlagL1) may be set according to the reference direction indicator information. As an example of implementation, in the case of unidirectional prediction using the L0 reference picture, predFlagL0 may be set to 1 and predFlagL1 to 0. In the case of unidirectional prediction using the L1 reference picture, predFlagL0 may be set to 0 and predFlagL1 to 1. In the case of bidirectional prediction using both the L0 and L1 reference pictures, predFlagL0 may be set to 1 and predFlagL1 to 1.

[0153] If the current block is a coding unit, the coding unit may be divided into multiple sub-blocks, and the prediction information of each sub-block may be the same or different. For example, if the coding unit is in intra mode, the intra prediction mode of each sub-block may be the same or different. Also, if the coding unit is in inter mode, the movement information of each sub-block may be the same or different. Additionally, each sub-block may be capable of encoding or decoding independently of each other. Each sub-block may be distinguished by a sub-block index (sbIdx).

[0154] The motion vector of the current block is highly likely to be similar to the motion vector of surrounding blocks. Therefore, the motion vectors of surrounding blocks can be used as motion vector predictors (mvp), and the motion vector of the current block can be derived using the motion vectors of surrounding blocks. Additionally, to improve the accuracy of the motion vector, the motion vector difference (mvd) between the optimal motion vector of the current block found from the original image by the encoder and the motion vector predictor can be signaled.

[0155] Motion vectors can have various resolutions, and the resolution of the motion vector can vary on a block-by-block basis. Motion vector resolution can be expressed in integer units, half-pixel units, quarter-pixel units, sixteenth-pixel units, or integer units of 4. Since images such as screen content are simple graphic forms like text, interpolation filters do not need to be applied; therefore, integer units and integer units of 4 can be selectively applied on a block-by-block basis. For blocks encoded in Affine mode, which can express rotation and scale, the shape changes significantly; therefore, integer units, quarter-pixel units, and sixteenth-pixel units can be selectively applied based on the block. Information regarding whether to selectively apply motion vector resolution on a block-by-block basis is signaled by amvr_flag. If applied, which motion vector resolution to apply to the current block is signaled by amvr_precision_idx.

[0156] For blocks where bidirectional prediction is applied, when applying the weighted average, the weights between the two prediction blocks can be applied equally or differently, and information about the weights is signaled through bcw_idx.

[0157] To improve the accuracy of motion prediction values, the Merge or AMVP (advanced motion vector prediction) methods can be selectively used on a block-by-block basis. The Merge method configures the motion information of the current block to be identical to the motion information of adjacent blocks; this method has the advantage of increasing the encoding efficiency of motion information by allowing motion information to propagate spatially without change within a homogeneous motion region. On the other hand, the AMVP method predicts motion information in the L0 and L1 prediction directions respectively to represent accurate motion information and signals the most optimal motion information. After deriving motion information for the current block through the AMVP or Merge method, the decoder uses the reference block located in the motion information derived from the reference picture as the prediction block for the current block.

[0158] In Merge or AMVP, the method for deriving motion information may involve constructing a motion candidate list using motion prediction values ​​derived from neighboring blocks of the current block, and then signaling index information for the optimal motion candidate. In the case of AMVP, since motion candidate lists are derived for L0 and L1 respectively, the optimal motion candidate indices (mvp_l0_flag, mvp_l1_flag) for L0 and L1 respectively are signaled. In the case of Merge, since a single motion candidate list is derived, a single merge index (merge_idx) is signaled. The motion candidate lists derived from a single coding unit can vary, and a motion candidate index or a merge index may be signaled for each motion candidate list. In this case, a mode in which there is no information regarding the remaining block in a block encoded in Merge mode can be referred to as MergeSkip mode.

[0159] Bidirectional motion information for the current block can be derived by combining AMVP and Merge modes. For example, motion information in the L0 direction can be derived using the AMVP method, while motion information in the L1 direction can be derived using the Merge method. Conversely, Merge can be applied to L0 and AMVP to L1. This encoding mode can be referred to as the AMVP-merge mode.

[0160] In this specification, the motion candidate and the motion information candidate may have the same meaning. Additionally, the motion candidate list and the motion information candidate list may have the same meaning.

[0161] SMVD (Symmetric MVD) is a method that reduces the amount of bits of motion information transmitted by making the Motion Vector Difference (MVD) values ​​in the L0 and L1 directions symmetrical in the case of bi-directional prediction. The MVD information in the L1 direction that is symmetrical to the L0 direction is not transmitted, and the reference picture information in the L0 and L1 directions is also not transmitted and can be derived during the decoding process.

[0162] OBMC (Overlapped Block Motion Compensation) is a method that generates prediction blocks for the current block using motion information from surrounding blocks when motion information between blocks differs, and then generates a final prediction block for the current block by weighted averaging the prediction blocks. This has the effect of reducing blocking phenomena occurring at block boundaries in motion-compensated images.

[0163] Generally, merge motion candidates have low motion accuracy. To improve the accuracy of these merge motion candidates, the MMVD (Merge mode with MVD) method can be used. The MMVD method corrects motion information using a single candidate selected from several motion difference value candidates. Information regarding the motion correction value obtained through the MMVD method (e.g., an index indicating a single candidate selected from the motion difference value candidates) can be included in the bitstream and transmitted to the decoder. Compared to the conventional method where motion information difference values ​​are included in the bitstream, the amount of bits can be saved by including information regarding the motion correction value in the bitstream.

[0164] MBVD (Merge Mode with Block Vector Differences) mode is a method for encoding difference values ​​for block vectors, similar to MMVD, which is a merge mode that encodes difference values ​​for motion vectors. The MBVD method determines a block vector by using a single candidate selected from several block vector difference value candidates. Information regarding the block vector difference value obtained through the MBVD method (e.g., an index indicating the single candidate selected from the block vector difference value candidates) can be included in the bitstream and transmitted to the decoder. Compared to the conventional method of including the block vector difference value in the bitstream, bit volume can be saved by including only a portion of the information regarding the block vector difference value in the bitstream.

[0165] Template Matching (TM) is a method that constructs a template using the surrounding pixels of the current block and finds the matching region with the highest similarity to the template to correct motion information. Template Matching is a method that performs motion prediction in the decoder without including motion information in the bitstream in order to reduce the size of the encoded bitstream. In this case, since the decoder does not have the original image, it can roughly derive motion information for the current block by using already restored surrounding blocks.

[0166] BM (Bilateral Matching) may be a method in which a video signal processing device corrects motion information based on the similarity between a reference block in a picture included in an L0 picture list derived based on motion information of the current block and a reference block in a picture included in an L1 picture list.

[0167] The Decoder-side Motion Vector Refinement (DMVR) method is a method that corrects motion information through the correlation of already restored reference images to find more accurate motion information. It uses the bidirectional motion information of the current block to select the point within the two reference pictures that best matches the reference blocks within the reference pictures within an arbitrary defined area of ​​the two reference pictures as the new bidirectional motion. When this DMVR is performed, the encoder can correct motion information by performing DMVR at the block level, and then divide the block into sub-blocks to perform DMVR at the sub-block level to correct the motion information of the sub-blocks. This can be referred to as Multi-pass DMVR (MP-DMVR).

[0168] The LIC (Local Illumination Compensation) method is a method for compensating for changes in luminance between blocks. It involves deriving a linear model using neighboring pixels adjacent to the current block, and then compensating for the luminance information of the current block through the linear model.

[0169] Existing video encoding methods perform motion compensation that considers only up, down, left, and right translations, so encoding efficiency is reduced when encoding videos that include movements such as enlargement, reduction, and rotation, which are commonly encountered in reality. To represent movements such as enlargement, reduction, and rotation, an Affine model-based motion prediction technique using a 4-parameter (rotation) or 6-parameter (enlargement, reduction, rotation) model can be applied.

[0170] Bi-Directional Optical Flow (BDOF) is used to correct a prediction block by estimating the amount of pixel change based on optical flow from a reference block of a block composed of bidirectional motion. The motion of the current block can be corrected using motion information derived from the BDOF of this VVC.

[0171] PROF (Prediction refinement with optical flow) is a technology designed to improve the accuracy of sub-block affine motion prediction to be comparable to that of pixel-level motion prediction. Similar to BDOF, PROF is a technique that obtains a final prediction signal by calculating pixel-level correction values ​​for sub-block affine motion-compensated pixel values ​​based on optical flow.

[0172] The CIIP (Combined Inter- / Intra-picture Prediction) method generates a final prediction block for the current block by taking a weighted average of the prediction blocks generated by the intra-picture prediction method and the prediction blocks generated by the inter-picture prediction method.

[0173] The Intra Block Copy (IBC) method locates the part most similar to the current block within an already restored area of ​​the current picture and uses that reference block as the prediction block for the current block. In this process, information related to the block vector—the distance between the current block and the reference block—can be included in the bitstream. The decoder can parse the information related to the block vector contained in the bitstream to calculate or set the block vector for the current block.

[0174] The BCW (Bi-prediction with CU-level Weights) method is a method that performs a weighted average of two motion-compensated prediction blocks by adaptively applying weights on a block-by-block basis, rather than generating a prediction block by averaging two motion-compensated prediction blocks from different reference pictures.

[0175] The Intra TMP (Template Matching Prediction) method is a method in which a video signal processing device constructs a reference template using pixel values ​​of neighboring blocks adjacent to the current block, finds the part most similar to the constructed reference template in an already restored area within the current picture, and then uses that reference block (the part found in the already restored area) as a prediction block for the current block.

[0176] The Multi-hypothesis Prediction (MHP) method performs weighted prediction through various prediction signals by transmitting additional motion information to unidirectional and bidirectional motion information during cross-frame prediction.

[0177] The Cross-component Linear Model (CCLM) is a method that predicts the chrominance signal by constructing a linear model utilizing the high correlation between a luminance signal and a chrominance signal located at the same position as that luminance signal. A template is constructed using a reconstructed block among the surrounding blocks adjacent to the current block, and parameters for the linear model are derived through this template. Next, the reconstructed current luminance block is selectively downsampled to fit the size of the chrominance block, depending on the image format. Finally, the chrominance block of the current block is predicted using the downsampled luminance block and the corresponding linear model. In this context, the method of using two or more linear models is called Multi-model Linear Mode (MMLM). Additionally, prediction methods that utilize the correlation between different signals, such as CCLM and MMLM, can be referred to as Cross-Component Prediction (CCP).

[0178] The Gradient Linear Model (GLM) is a method that predicts color difference signals by constructing a model that additionally incorporates the gradient between a luminance sample corresponding to a color difference sample and surrounding luminance samples adjacent to it, in addition to linear models such as CCLM.

[0179] Prediction methods that utilize the correlation between different signals, such as CCLM, MMLM, CCCM, and GLM, can be called Cross-Component Prediction (CCP). In other words, a method of predicting another signal (which can be a chrominance signal, such as Cb or Cr) from a single signal (e.g., a luminance signal) can be called Cross-Component Prediction (CCP).

[0180] The CCP merge method is a method that predicts the color difference block of the current block using CCP models (such as CCLM, MMLM, and CCCM) used in neighboring blocks.

[0181] The encoding modes (prediction modes) described in this specification may be described by omitting the term "mode" or by using the term "method" instead of "mode." For example, a CCLM mode may be described as a CCLM method or CCLM.

[0182]

[0183] Among intra-prediction coding techniques, the Matrix Intra Prediction (MIP) method is a matrix-based intra-prediction method that, unlike prediction methods that derive directionality from pixels of neighboring blocks adjacent to the current block, obtains a prediction signal by utilizing predefined matrix and offset values ​​for pixels to the left and top of neighboring blocks.

[0184] To derive the intra-mode derivation of the current block, an intra-mode derivation derived from the surrounding pixels of a template—which is an arbitrary region adjacent to the current block that has been restored—can be used for the restoration of the current block. First, the decoder generates a prediction template for the template using surrounding pixels (references) adjacent to the template, and the intra-mode derivation that generates the prediction template most similar to the already restored template can be used for the restoration of the current block. This method can be called TIMD (Template intra-mode derivation).

[0185] The TIMD merge mode is a mode used in the current block by inheriting the intra-prediction directional mode, information on whether to apply a weighted average, weight information for each intra-prediction mode, and transform type information from neighboring blocks encoded in TIMD or TIMD merge mode. The video signal processor can construct a TIMD merge list using the TIMD information used in neighboring blocks. The video signal processor can add a predetermined maximum number of TIMD information derived from neighboring blocks to the TIMD merge list in order of shortest distance. The predetermined maximum number may be a positive integer greater than or equal to 1. For example, the predetermined maximum number may be 5. The video signal processor may reorder the TIMD merge list based on the template cost and use the candidate with the minimum cost as the TIMD mode for the current block. If the current block is encoded in TIMD merge mode, the video signal processor may use an implicit DST7 transform for the residual signal if the size of the current block is between 4 and 16. Otherwise, the video signal processor may use the transform type derived from the neighboring blocks. The transformation type derived from neighboring blocks may be transformation type information within the TIMD candidate for the current block determined from the TIMD merge list. In TIMD mode, the template cost can be calculated via SATD. If the method of calculating the template cost changes, the derived intra-prediction mode may also change. TIMD SAD mode is similar to TIMD mode, but uses the MR-SAD method instead of SATD to calculate the template cost.

[0186] Generally, an encoder can determine a prediction mode for generating a prediction block and generate a bitstream containing information about the determined prediction mode. A decoder can parse the received bitstream to set an intra prediction mode. In this case, the bit amount of information regarding the prediction mode may be about 10% of the total bitstream size. To reduce the bit amount of information regarding the prediction mode, the encoder may not include information about the intra prediction mode in the bitstream. Accordingly, the decoder can derive (determine) an intra prediction mode for the restoration of the current block by utilizing the characteristics of surrounding blocks, and can restore the current block using the derived intra prediction mode. In this case, to derive the intra prediction mode, the decoder may use a method of inferring directional information by applying a Sobel filter in the horizontal and vertical directions to each surrounding pixel adjacent to the current block, and then mapping that directional information to the intra prediction mode. The method by which the decoder derives the intra prediction mode using surrounding blocks can be described as DIMD (Decoder-side intra mode derivation).

[0187] If the current block is a luminance block, the video signal processing device can derive an intra prediction mode through the DIMD and TIMD methods. If the current block is a chrominance block, since there is an already restored luminance block corresponding to the chrominance block, the video signal processing device can derive an intra prediction mode by applying the DIMD and TIMD methods using the restored luminance block, and then use the derived intra prediction mode as the intra prediction mode for the chrominance block. That is, the video signal processing device derives directional information when the current block is a luminance block, and may not derive directional information when the current block is a chrominance block, and may apply the directional information found in the luminance block to the chrominance block. This mode can be referred to as the 'DIMD Chroma' mode or the 'TIMD Chroma' mode.

[0188] Figure 7 is a diagram showing the locations of surrounding blocks used to construct a list of motion candidates in inter prediction.

[0189] The surrounding blocks may be blocks of spatial location or blocks of temporal location. A surrounding block spatially adjacent to the current block may be at least one of the Left (A1) block, Left Below (A0) block, Above (B1) block, Above Right (B0) block, or Above Left (B2) block. A surrounding block temporally adjacent to the current block may be a block containing the top-left pixel position of the bottom-right (BR) block of the current block in the corresponding collocated picture. If the surrounding block temporally adjacent to the current block is encoded in intra mode or if the surrounding block temporally adjacent to the current block exists in an unusable location, a block containing the horizontal and vertical center (Center, Ctr) pixel position of the current block in the collocated picture corresponding to the current block may be used as the temporal surrounding block. Motion candidate information derived from a corresponding picture can be referred to as TMVP (Temporal Motion Vector Predictor). Only one TMVP can be derived from a single block, or a single block can be divided into multiple sub-blocks, and a TMVP candidate can be derived for each sub-block. The method of deriving TMVP at the sub-block level can be referred to as sbTMVP (sub-block Temporal Motion Vector Predictor).

[0190] Whether the methods described herein may be applied may be determined based on at least one of the following information: slice type information (e.g., whether it is an I slice, P slice, or B slice), whether it is a tile, whether it is a sub-picture, the number of samples in the current block, the width of the current block, the height of the current block, the depth of the coding unit, whether the current block is a luminance block or a chrominance block, whether it is a reference frame or a non-reference frame, and the time hierarchy according to the reference order and hierarchy. The information used to determine whether the methods described herein may be applied may be information agreed upon in advance between the decoder and the encoder. Additionally, such information may be determined according to profiles and levels. Such information may be expressed as variable values, and the bitstream may contain information regarding variable values. That is, the decoder may determine whether the methods described above are applied by parsing information regarding variable values ​​included in the bitstream. The methods described herein may be performed in luminance blocks and chrominance blocks. The methods described herein may be performed only on the luminance block and not on the chrominance block. The methods described herein may be performed only on the chrominance block and not on the luminance block. Whether the methods described herein are performed may be determined based on information regarding whether the current picture is used as a reference picture or not. For example, whether the methods described above are applied may be determined based on the width or height of the coding unit. If the width or height is 32 or greater (e.g., 32, 64, 128, etc.), the methods described above may be applied. Additionally, if the width or height is less than 32 (e.g., 2, 4, 8, 16), the methods described above may be applied. Additionally, if the width or height is 4 or 8, the methods described above may be applied.

[0191] FIG. 8 illustrates a method for determining a reference pixel line based on a template according to one embodiment of the present specification.

[0192] The method for determining the optimal reference pixel line for the current block (for the restoration of the current block) based on the template below is described. Here, the reference pixel line may have the same meaning as the reference sample line.

[0193] Referring to FIG. 8, the video signal processing device can construct a reference template using reference pixel lines adjacent to the current block. The video signal processing device can generate prediction samples for the location of the reference template using reference pixel lines 1, 2, 3..., etc. The video signal processing device can calculate the cost between the generated prediction samples and the samples of the reference template. At this time, the cost can be calculated through methods such as SAD (Sum of Absolute Differences) or MRSAD (Mean-Removed SAD). The reference pixel corresponding to the minimum cost may be the optimal reference pixel. Additionally, the encoder can rearrange the calculated costs in ascending order, construct a list of reference pixel lines, and then generate and signal a bitstream containing information about the index of the optimal reference pixel line. The decoder can construct a list of reference pixel lines through the method described above, parse the index of the optimal reference pixel line included in the bitstream, and generate prediction samples using the reference pixel line indicated by the index. As described herein, a method for a video signal processing device to determine a reference pixel line based on a template may be described as a Template-based Multiple Reference Line (TMRL) method or a TMRL intra-prediction method. When the video signal processing device calculates the template cost, it may use any one of SAD, SATD, MR-SAD, or MR-SATD. SAD (Sum of Absolute Difference) may be the sum of the absolute values ​​of the differences between samples. Additionally, SATD (Sum of Absolute Transformed Difference) may be the sum of the absolute values ​​of the transformed differences between samples. Additionally, MR-SAD (Mean Removal SAD) is the value obtained by subtracting the mean from the sum of the absolute values ​​of the differences between samples.In addition, SATD (Mean Removal SATD) can be the sum of the absolute values ​​of the differences between samples minus the mean. In this case, various transformations can be used. For example, the Hadamard transformation can be used because it has low complexity.

[0194] FIG. 9 is a diagram showing block vectors related to an IBC encoding method according to one embodiment of the present specification.

[0195] The IBC encoding method (IBC mode) is a method that finds the part most similar to the current block (reference block) within an already restored area of ​​the current picture and uses the reference block as the prediction block for the current block. In this case, the encoder can generate a bitstream containing information related to the block vector, which is the distance between the current block and the reference block. The decoder can parse the information related to the block vector contained in the bitstream to calculate or set the block vector for the current block. The IBC encoding method can be applied to the chrominance block. In the chrominance block, instead of finding a new block vector, the block vector of the luminance block corresponding to the chrominance block can be used as the block vector for the chrominance block; this encoding method can be referred to as the DBV (Direct Block Vector) mode.

[0196] The RRIBC (Reconstruction-Reordered IBC) encoding mode can be used in IBC blocks (blocks to which the IBC encoding method is applied). RRIBC can consist of vertical flips and horizontal flips. In blocks to which RRIBC is applied, the reconstructed block is flipped according to the RRIBC type of the current block. The encoder may flip the original block to be encoded before finding the part most similar to the current block in the reference picture. That is, the most similar part in the reference picture is found using the flipped original block. Therefore, the prediction block uses the unflipped block, and the residual block may also be the unflipped block. The decoder may flip the final reconstructed block according to the RRIBC type of the current block.

[0197] Figure 10 shows a method for predicting the current block using RRIBC in the horizontal direction.

[0198] Figure 11 shows a method for predicting the current block using RRIBC in the vertical direction.

[0199] In FIGS. 10 and 11, (Xn, Yn) represents the center position of the surrounding blocks, and (Xc, Yc) represents the center position of the current block. BV n h , BV n v represent the horizontal block vector and vertical block vector of the surrounding block, respectively, and BV C h , BV C v and represent the horizontal block vector and vertical block vector of the current block, respectively. As shown in Fig. 10, when the RRIBC type of the current block is horizontal, BV C h is 2 * (Xn - Xc) + BV n hIt can be calculated as, and as shown in FIG. 53, if the RRIBC type of the current block is in the vertical direction, BV C v is 2 * (y n - y c ) + BV n v It can be calculated as. In this case, BV n h , BV n v Since it uses the restored block, BV n h , BV n v The sign of can be negative.

[0200] When the current block is encoded in RRIBC, the video signal processing device can determine the optimal block vector by constructing a block vector candidate list. In this case, the video signal processing device may construct the block vector candidate list according to the RRIBC type of the current block. For example, if the RRIBC type of the current block is horizontal, the video signal processing device may construct the block vector candidate list using only the surrounding blocks of the current block encoded in horizontal RRIBC. Additionally, the video signal processing device may construct the block vector candidate list regardless of the RRIBC type of the current block. For example, if the RRIBC type of the current block is horizontal, the block vector candidate list may be constructed using not only the surrounding blocks of the current block encoded in horizontal RRIBC, but also the surrounding blocks of the current block encoded in vertical RRIBC, / or blocks encoded in general motion, and / or blocks encoded in block vectors.

[0201] When a video signal processing device predicts the current block using general motion, the device may construct a motion candidate list based on whether the surrounding blocks of the current block are encoded in IBC mode or RRIBC mode. In this case, if the surrounding blocks of the current block are encoded in RRIBC mode, the device may construct the motion candidate list by additionally considering the RRIBC type. For example, when the video signal processing device constructs a motion candidate list for the current block, if the encoding mode of the surrounding blocks is RRIBC mode and the RRIBC type is vertical or horizontal, the block vector of the surrounding blocks may not be included in the motion candidate list. Alternatively, when the video signal processing device predicts the current block using general motion, the device may construct the motion candidate list regardless of the encoding mode of the surrounding blocks of the current block. That is, the video signal processing device may construct a motion candidate list regardless of whether the surrounding blocks of the current block are encoded in IBC mode or RRIBC mode. For example, when a video signal processing device constructs a motion candidate list for a current block, it may include the block vector of a surrounding block in the motion candidate list even if the encoding mode of the surrounding block of the current block is IBC mode and the RRIBC type is vertical or horizontal direction.

[0202] The RRIBC encoding method is effective for images with symmetrical characteristics. Symmetrical characteristics can refer to perfect horizontal (or vertical) symmetry, where the current block and the reference block are separated by equal distances along a single central axis. The vertical (or horizontal) direction (axis of symmetry) of the current block may lie on the same line as the vertical (or horizontal) direction (axis of symmetry) of the reference block. In this case, the block vector in the vertical (or horizontal) direction may be set to '0', and the block vector in the horizontal direction may be set to any negative value other than '0'. The current block and the reference block may be configured symmetrically, differing by different distances along a single central axis, and the current block and the reference block may lie on different vertical lines. In this case, the block vector in the vertical direction may be set to any negative value other than '0'. That is, a video signal processing device can encode or decode the current block using a symmetric block in which both the horizontal and vertical directions have block vector values ​​of any negative value other than '0'.

[0203] FIG. 12 shows a block vector of a block encoded in Intra TMP mode according to one embodiment of the present specification.

[0204] Referring to FIG. 12, the Intra TMP method (Intra TMP encoding mode) is a method in which a video signal processing device constructs a reference template using pixel values ​​of neighboring blocks adjacent to the current block, finds the part most similar to the reference template within an already restored area (reference block) in the current picture, and then uses the reference block (Ref. luma block in FIG. 12) as a prediction block for the current block. At this time, there may be one or more block vectors used to generate a prediction block for the current luminance block, and FIG. 12 shows the case where there are two block vectors. The video signal processing device can generate a prediction block for the luminance block by weighted averaging the reference blocks predicted from the two block vectors of the current luminance block. This can be referred to as a fusion mode. In the case of a chrominance block, the video signal processing device can derive a block vector from the luminance block corresponding to the chrominance block and then use the block vector to generate a chrominance prediction block. There may be one or more block vectors used to generate a prediction block for the current chrominance block, and FIG. 12 shows the case where there are two block vectors for the chrominance block. In this case, the block vector of the chrominance block may be the same as or different from the block vector derived from the luminance block. In Intra TMP mode, the video signal processing device derives filter coefficients using the correlation between samples around the reference block indicated by the block vector and samples around the current block, and can apply filtering to the reference block indicated by the block vector using the derived filter coefficients. The video signal processing device can use the filtered reference block as the prediction block for the current block. This can be referred to as Intra TMP filtering mode. In Intra TMP mode, the video signal processing device can apply compensation to the current prediction block through the block vector, similar to the LIC method using motion information.

[0205] FIG. 13 illustrates a case where the current block is divided by a GPM mode according to one embodiment of the present specification, and the divided area is encoded in an IBC mode.

[0206] The current block can be encoded or decoded in IBC-GPM mode. IBC-GPM mode may refer to a mode in which the current block is divided using GPM mode, and the divided regions are encoded in intra prediction mode or IBC prediction mode. Referring to Fig. 13, the current block can be divided into two regions based on the dotted line. At this time, one of the divided regions can be encoded in intra prediction mode, and the other can be encoded in IBC prediction mode. For example, among the divided regions, the left region can be encoded in intra prediction mode, and the right region can be encoded in IBC prediction mode. At this time, the region encoded in IBC prediction mode can be encoded in Intra TMP prediction mode, and this can be described as IntraTMP-GPM mode.

[0207] FIG. 14 illustrates a method in which a current block is encoded in IBC-CIIP mode according to one embodiment of the present specification.

[0208] The video signal processing device can predict the current block using a weighted average between the block predicted in intra mode and the block predicted in IBC mode, which can be referred to as the IBC-CIIP mode. In this case, the block predicted in Intra TMP mode may be used instead of the block predicted in IBC mode, which can be referred to as the IntraTMP-CIIP mode.

[0209] The video signal processing device can obtain a restored luminance block by adding the residual signal for the luminance prediction block of the current block and the luminance block, and can construct a CCP model using the correlation between the restored luminance block and the luminance prediction block. At this time, the CCP model can be one of CCLM, MMLM, GLM, CCCM, MM-CCCM, GL-CCCM, CCCM-ND, and CCCM-MDF. A first chrominance prediction block with the CCP model applied can be generated by applying the CCP model derived from the luminance block to the chrominance prediction block. A final chrominance block can be generated by adding the error signal for the first chrominance prediction block and the chrominance block.

[0210] When the current luminance block is encoded in Intra TMP or IBC mode, the video signal processing device can derive a CCP model through the correlation between the luminance block and the chrominance block of the reference block indicated by the block vector of the current luminance block. Using the derived CCP model and the currently restored luminance block, Cb and Cr prediction blocks, which are chrominance prediction blocks for the current block, can be generated. This is referred to as BVG CCP (block vector guided cross component prediction). The CCP model may include at least one of CCLM, MMLM, GLM, MM-GLM, CCCM, MM-CCCM, GL-CCCM, GL-MM-CCCM, CCCM-ND, MM-CCCM-ND, CCCM-MDF, MM-CCCM-MDF, and Chroma Fusion. When a CCP model is derived using the BV of the luminance block and the CCCM model among the CCP models is applied, it may be referred to as BVG CCCM. Currently, when a chrominance block derives a CCP model using the BV of the luminance block and a GL-CCCM model is applied among the CCP models, it can be referred to as BVG GL-CCCM. In addition, various CCP models can be applied to BVG CCP.

[0211] The types of CCP models may include at least one of CCLM, MMLM, GLM, MM-GLM, CCCM, MM-CCCM, GL-CCCM, MM-GL-CCCM, CCCM-ND, MM-CCCM-ND, CCCM-MDF, MM-CCCM-MDF, Chroma Fusion, BVG-CCCM, and CCP Merge. An MMLM may be composed of two CCLMs. If the types of CCP models between CCP models are different, they may be considered different CCP models. If the types of CCP models between CCP models are the same but the parameters between CCP models are different, they may be considered different CCP models. If the number of CCP models between CCP models is different, for example, if the first CCP model is composed of one CCLM and the second CCP model is composed of two CCLMs, they may be considered different CCP models. The types of CCP models may be referred to as CCP modes.

[0212] In chroma fusion (CF), the video signal processing unit can predict the chroma difference signal through a weighted average between the chroma difference block predicted via an intra-prediction mode and the current luminance block, without using a linear model (LM) mode. In this case, the weight parameters can be derived using the CCCM method.

[0213] In a gradient linear model (GLM), the video processing unit constructs the model by additionally reflecting the gradient between the luminance sample corresponding to the color difference sample in the linear model and the surrounding luminance samples adjacent to that luminance sample, and can predict the color difference signal through the model.

[0214] The CCP model used to predict color difference blocks constructs a list of CCP candidates using the CCP models used in the surrounding blocks of the current block, and can derive the optimal CCP candidate from this list. This can be referred to as CCP merge mode.

[0215] FIG. 15 shows an example of a reference region and filter shape used to derive CCCM parameters according to an embodiment of the present specification.

[0216] CCCM (Convolutional Cross-Component Intra Prediction Model) is a method for predicting chrominance signals using a non-linear model constructed by utilizing the high correlation between a luminance signal and a chrominance signal located at the same position as the luminance signal. Fig. 15(a) shows the positional relationship between reference samples (vertical diagonal lines, 1520) for applying CCCM to the current prediction block (diagonal diagonal lines, 1510) and side samples (horizontal diagonal lines) required when applying a cross-shaped filter. The current prediction block (MxN) can be composed of reference samples in the top 6 rows (2Mx6), reference samples in the left 6 rows (6x2N), and 6x6 reference samples in the top-left row, with a chroma to luminance sample count ratio of 1:1. When a video signal processing device applies a cross-shaped sample filter (Fig. 15(b)) for CCCM to the chroma sample prediction equation (Fig. 15(c)), cases occur where the reference sample area is exceeded. In this case, the additional samples required may be side samples. The chroma sample prediction relationship of Fig. 15(c) can be applied to each chroma component (i.e., Cb, Cr). The sample at position C (Center) in Fig. 15(b) may be a luminance sample corresponding to the Cb and Cr chroma samples, and N (North), E (East), S (South), and W (West) may be luminance samples adjacent to the luminance sample at position C. Side samples may be additionally required one sample at a time for regions other than the reference sample, depending on the position of the C sample. If the sample value at the position of the side sample is unavailable, the sample value at the unavailable position may be padded with the C sample value. The P value in Fig. 15(c) may be a nonlinear term. The P value can be calculated as P = ( C*C + midVal ) >> bitDepth, and in the case of 10-bit content, P = ( C*C + 512 ) >> 10.In FIG. 15(c), the value of B can be an integer offset value as a bias value. The value of B can be an intermediate value of bitDepth. For 10-bit content, the value of B can be 512.

[0217] MM-CCCM (multi-model CCCM) is a method that derives two CCCM parameters based on the average value of a reference area (or restored current luminance block).

[0218] GL-CCCM (Gradient and location based convolutional cross-component model) is an additional CCCM mode that utilizes gradient and location information. In the case of the existing CCCM mode, the video signal processing unit can derive a chrominance sample for the current block using the luminance sample at the location corresponding to the chrominance sample to be predicted, four samples surrounding that luminance sample, and coefficient information. In the case of the GL-CCCM mode, the video signal processing unit can derive a chrominance sample for the current block by reflecting the vertical and horizontal differences of the luminance sample at the location corresponding to the chrominance sample to be predicted and eight samples surrounding that luminance sample, and also using the location value of the current luminance sample and its coefficient information.

[0219] When CCCM mode is applied, the video signal processing device may apply downsampling filters to match the resolution difference between the luminance block and the chrominance block. This is to reduce the resolution of the luminance block to that of the chrominance block. The mode that applies downsampling filters can be described as CCCM-MDF (CCCM with multiple downsampling filters).

[0220] When an inter-coding mode is applied to the current block, the video signal processing device may derive linear and non-linear models between a luminance prediction block (Y') derived using motion information of the current block and a first chrominance prediction block (Cb', Cr') derived using motion information of the current block (Derive filter), generate a restored luminance block of the current block using the luminance residual block of the current block, and generate a second chrominance prediction block of the current block by applying the derived linear and non-linear models to the restored luminance block of the current block (Apply filter). Subsequently, the video signal processing device may generate the final chrominance block (Cb, Cr) of the current block by adding the chrominance residual block to the second chrominance prediction block of the current block predicted using the linear and non-linear models. This method may be described as a cross-component residual model (CCRM). CCRM, Inter CCCM, and Inter CCP may have the same meaning.

[0221] To improve coding efficiency, instead of coding the aforementioned residual signal as is, a method may be used in which the transformation coefficient values ​​obtained by transforming the residual signal are quantized, and the quantized transformation coefficients are coded. As described above, the transformation unit can obtain transformation coefficient values ​​by transforming the residual signal. In this case, the residual signal of a specific block may be distributed across the entire range of the current block. Accordingly, coding efficiency can be improved by concentrating energy in the low-frequency range through frequency domain transformation of the residual signal.

[0222] The encoder may acquire at least one residual block containing residual signals for the current block. The residual block may be either the current block or blocks partitioned from the current block. In this specification, the residual block may be described as a residual array or a residual matrix containing residual samples of the current block. Additionally, in this specification, the residual block may represent a block of the same size as the size of the conversion unit or the conversion block.

[0223] An encoder may transform a residual block using a transformation kernel. The transformation kernel used for transforming the residual block may be a transformation kernel having separable characteristics of vertical transformation and horizontal transformation. In this case, the transformation for the residual block may be performed by separating it into a vertical transformation and a horizontal transformation. For example, the encoder may perform a vertical transformation by applying a transformation kernel in the vertical direction of the residual block. Additionally, the encoder may perform a horizontal transformation by applying a transformation kernel in the horizontal direction of the residual block. In this specification, the term "transform kernel" may be used to refer to a set of parameters used for transforming a residual signal, such as a transformation matrix, a transformation array, a transformation function, or a transformation. According to one embodiment, the transformation kernel may be any one of a plurality of available kernels. Additionally, transformation kernels based on different transformation types may be used for each of the vertical transformation and the horizontal transformation. That is, before performing the first transformation, a transformation method for the vertical and horizontal directions can be derived using at least one of the current block's intra prediction mode, encoding mode, transformation method parsed from the bitstream, and size information of the current block. Additionally, to reduce computational complexity during the transformation process for large blocks, a process can be performed in which only the low-frequency region is retained and the high-frequency region is treated as '0'. This process is called high-frequency zeroing, and the transformation size during the actual first transformation can be set for this zeroing. In the high-frequency zeroing process, the low-frequency region can be set to an arbitrary fixed size; for example, the horizontal or vertical size can be a combination of 4, 8, 16, 32, etc.

[0224] The encoder can transmit the transformed block converted from the residual block to the quantization unit for quantization. At this time, the transformed block may include a plurality of transformation coefficients. Specifically, the transformed block may be composed of a plurality of transformation coefficients arranged in a two-dimensional array. The size of the transformed block may be the same as that of the residual block, either the current block or a block divided from the current block. The transformation coefficients transmitted to the quantization unit may be expressed as quantized values.

[0225] Additionally, the encoder may perform an additional transformation before the transformation coefficients are quantized. The aforementioned transformation method may be referred to as a primary transform, and the additional transformation may be referred to as a secondary transform. The secondary transform may be optional for each residual block. According to one embodiment, the encoder may improve coding efficiency by performing a secondary transform on regions where it is difficult to concentrate energy in the low-frequency region using only the primary transform. For example, a secondary transform may be added for blocks where residual values ​​appear significantly in directions other than the horizontal or vertical direction of the residual block. Residual values ​​of an intra-predicted block may have a higher probability of changing in directions other than the horizontal or vertical direction compared to residual values ​​of an inter-predicted block. Accordingly, the encoder may additionally perform a secondary transform on the residual signals of the intra-predicted block. Additionally, the encoder may omit the secondary transform for the residual signals of the inter-predicted block. High-frequency zeroing from the first conversion can also be performed during the second conversion process.

[0226] As another example, whether to perform a second transformation may be determined based on the size of the current block or the remaining block. Additionally, transformation kernels of different sizes may be used depending on the size of the current block or the remaining block. For example, an 8x8 second transformation may be applied to a block where the length of the shorter side (width or height) is greater than or equal to a first set length. Additionally, a 4x4 second transformation may be applied to a block where the length of the shorter side (width or height) is greater than or equal to a second set length and smaller than the first set length. In this case, the first set length may be a value greater than the second set length, but the present disclosure is not limited thereto. Furthermore, unlike the first transformation, the second transformation may not be performed by separating it into a vertical transformation and a horizontal transformation. Such a second transformation may be referred to as a Low Frequency Non-Separable Transform (LFNST).

[0227] Furthermore, in the case of video signals in specific regions, high-frequency band energy may not decrease even after frequency conversion due to abrupt changes in brightness. Consequently, compression performance through quantization may be degraded. Additionally, if conversion is performed on regions where residual values ​​are sparse, encoding and decoding times may increase unnecessarily. Accordingly, conversion for residual signals in specific regions may be omitted. Whether to perform conversion on residual signals in specific regions can be determined by syntax elements related to the conversion of those regions. For example, the syntax elements may include transform skip information. The transform skip information may be a transform skip flag. If the transform skip information for a residual block indicates a transform skip, the transformation for that residual block is not performed. In this case, the encoder can immediately quantize the residual signal for which the transformation was not performed.

[0228] The aforementioned transformation-related syntax elements may be information parsed from a video signal bitstream. A decoder can obtain the transformation-related syntax elements by entropy decoding the video signal bitstream. Additionally, an encoder can generate a video signal bitstream by entropy coding the transformation-related syntax elements.

[0229] The decoder can parse the transmitted bitstream to obtain encoding information necessary for decoding. At this time, information related to the conversion process includes index information for first and second conversion types and quantized conversion coefficients. The inverse conversion unit can obtain a residual signal by performing an inverse conversion on the inverse quantized conversion coefficients. First, the inverse conversion unit can detect whether an inverse conversion is performed for a specific region from the conversion-related syntax elements of that region. According to one embodiment, if the conversion-related syntax elements for a specific conversion block indicate a conversion skip, the conversion for that conversion block may be omitted. In this case, both the first inverse conversion and the second inverse conversion for the conversion block may be omitted. Additionally, the inverse quantized conversion coefficients may be used as a residual signal. For example, the decoder can use the inverse quantized conversion coefficients as a residual signal to restore the current block. Alternatively, a second inverse transform may be performed and a first inverse transform may be omitted, and the value of the second inverse transform may be used as a residual signal. The aforementioned first inverse transform represents the inverse transform for the first transform and may be referred to as the inverse first transform. The second inverse transform represents the inverse transform for the second transform and may be referred to as the inverse second transform or inverse LFNST. In the present invention, the first (inverse) transform may be referred to as the first (inverse) transform, and the second (inverse) transform may be referred to as the second (inverse) transform.

[0230] FIG. 16 shows a type of conversion kernel that can be used for video coding according to one embodiment of the present specification.

[0231] Figure 16 shows the formulas for the DCT-II, DCT-V (discrete cosine transform type-V), DCT-VIII (discrete cosine transform type-VIII), DST-I (discrete sine transform type-I), and DST-VII kernels applied to MTS. DCT and DST can be expressed as functions of cosine and sine, respectively. When the basis function of the transform kernel for the number of samples N is expressed as Ti(j), index i represents the index in the frequency domain, and index j represents the index within the basis function. That is, as i decreases, it represents a low-frequency basis function, and as i increases, it represents a high-frequency basis function. When the basis function Ti(j) is expressed as a two-dimensional matrix, it can represent the j-th element of the i-th row. Since all the transform kernels illustrated in Figure 16 have separable characteristics, transformations can be performed on the residual signal X in both the horizontal and vertical directions. In other words, if we denote the residual signal block as X and the transformation kernel matrix as T, the transformation for the residual signal X can be represented as TXT'. Here, T' represents the transpose of the transformation kernel matrix T. Since DCT and DST are decimal forms rather than integers, implementing them directly in hardware encoders and decoders poses a burden. Therefore, the decimal transformation kernel must be approximated into an integer form through scaling and rounding. The integer precision of the transformation kernel can be determined as 8-bit or 10-bit, but if precision is lowered, encoding efficiency may decrease. Although the orthonormal properties of DCT and DST may not be maintained due to the approximation, the resulting loss in encoding efficiency is not significant; therefore, approximating the transformation kernel into an integer form is advantageous in terms of hardware encoder and decoder implementation.An Identity Transform (IDTR) is a transformation in which the result is the original state itself; it is also called an identity transformation. Generally, an identity transformation constructs a transformation matrix by setting a '1' at positions where the row and column values ​​are identical. However, in this context, the identity transformation uses arbitrary fixed values ​​instead of '1' to equally increase or decrease the value of the input residual signal.

[0232] In the aforementioned first-order MTS transformation, the transformation is calculated by applying a transformation kernel to the vertical and horizontal directions of the error block, respectively, so it can be described as a separable transform method. On the other hand, in the aforementioned second-order LFNST transformation, the transformation is calculated by applying the transformation kernel only once, without applying a transformation kernel to the vertical and horizontal directions, so it can be described as a non-separable transform method. Furthermore, since the aforementioned second-order transformation is applied additionally to the first-order transformed transformation coefficients of the block to which the DCT-2 transformation has been applied, it can be described as a two-stage transformation technique. Although the aforementioned second-order transformation offers high encoding efficiency, it has the disadvantage of being complex because a total of three transformation kernels are applied. To reduce this complexity, the NSPT (Non-separable primary transform) method, which applies the transformation using only the second-order transformation method, can be applied. The NSPT transformation method is a non-separable transformation method, which is calculated by applying the transformation kernel only once, rather than applying transformation kernels to the vertical and horizontal directions of the error block separately. In a video signal processing device, the error block of the current block can be transformed or inversely transformed using one of the transformation methods among MTS, DCT2 + LFNST, and NSPT.

[0233] NSPT transformation may be a transformation method that replaces the existing DCT2 + LFNST transformation method. For a transformation block with a size equal to or smaller than 16x16, any one of the following kernels may be applied depending on the size of the transformation block: NSPT4x4 (16x16 kernel), NSPT4x8 (32x20 kernel), NSPT8x4 (32x20 kernel), NSPT8x8 (64x32 kernel), NSPT4x16 (64x24 kernel), NSPT16x4 (64x24 kernel), NSPT8x16 (128x40 kernel), NSPT16x8 (128x40 kernel), NSPT4x32 (128x20 kernel), NSPT32x4 (128x20 kernel), NSPT8x32 (256x24 kernel), NSPT32x8 (256x24 kernel). Similar to LFNST, NSPT consists of 35 sets of transformation kernels, each set consisting of 3 candidates. The encoder can derive a set of transformation kernels according to the intra-prediction mode, and then generate and signal a bitstream containing information on the index for the optimal candidate among the 3 candidates. The decoder can parse the information on the signaled index, then inversely transform the current transformation coefficients using the transformation kernel candidate indicated by the index information from the set of transformation kernels derived using the intra-prediction mode, and obtain a residual block. Zero-out may not be performed on a 4x4 block to which NSPT is applied. Additionally, the number of coefficients zeroed out may vary depending on the size of the NSPT kernel. For example, a 32x20 NSPT may be applied to a 4x8 block or an 8x4 block. Therefore, out of the 32 transformation coefficients, the remaining 12 coefficients may be zeroed out, excluding only 20 transformation coefficients.

[0234] FIG. 17 shows a conversion set table for LFNST and NSPT conversions according to one embodiment of the present specification.

[0235] There may be 35 types of transform sets used in LFNST and NSPT transforms, and they may vary depending on the intra-prediction mode (see FIG. 6). That is, the video signal processing device can derive the transform set index of the LFNST and NSPT transforms corresponding to the intra-prediction mode (see FIG. 6) by referring to the transform set table of FIG. 17. In addition, the LFNST and NSPT transform sets may vary depending on the intra-prediction mode, information on whether the current block is a luminance block or a chrominance block, the width and height of the current block, and whether the intra-prediction directional mode of the current block is an extended angle mode. For each transform set, there may be an arbitrary number of transform matrices. Here, the arbitrary number may be an integer greater than or equal to 1, or it may be 3. The encoder may signal by including index information for the optimal transform matrix among the multiple transform matrices within the transform set in the bitstream. The decoder may parse the index for the optimal transform matrix and then apply the inverse transform using the transform matrix corresponding to the index in the transform set.

[0236] The encoder and decoder can apply a total of three types of transformation kernels, LFNST4, LFNST8, and LFNST16, depending on the size of the transformation block. If the width and height of the current transformation block are greater than or equal to 16, the encoder and decoder can apply the LFNST16 transformation kernel. If the width and height of the current transformation block are greater than or equal to 8, the encoder and decoder can apply the LFNST8 transformation kernel. If the width and height of the current transformation block are less than 8, the encoder and decoder can apply the LFNST4 transformation kernel.

[0237] Figure 18 shows an example of an ROI after LFNST transformation.

[0238] The gray block area in FIG. 18, which is the low-frequency portion output after the encoder and decoder perform an LFNST transformation on the residual block, is referred to as the Region of Interest (ROI). Areas other than the pre-specified ROI can be zero-out processed by setting them to 0. In the case of LFNST16, only 96 samples are required. Therefore, as shown in FIG. 18 (a), the remaining white sub-blocks, excluding 6 NxN sub-blocks, can be zero-out processed. In the case of LFNST8, only 64 samples are required. Therefore, as shown in FIG. 18 (b), the remaining white sub-blocks, excluding 4 NxN sub-blocks, can be zero-out processed. In the case of LFNST4, there may be no zero-out region. Here, N is a positive integer and can be 4.

[0239] FIG. 19 illustrates a method for deriving multiple conversion sets and LFNST / NSPT sets according to one embodiment of the present specification.

[0240] Referring to FIG. 19(a), the encoder can select to apply one of MTS, DCT2 + LFNST, or NSPT to the residual block, and can transform the residual block and obtain transformation coefficients based on the selected transformation method. At this time, the encoder can signal by including information about which transformation method was applied in the bitstream. If MTS transformation is applied to the residual block, the transformation coefficients of the residual block can be obtained by applying MTS transformation, and LFNST and NSPT transformations may not be applied. If DCT2 + LFNST transformation is applied to the residual block, the encoder can obtain first-order transformation coefficients by applying DCT2 transformation to the residual block, and obtain second-order transformation coefficients by applying LFNST transformation to the first-order transformation coefficients. At this time, MTS and NSPT transformations may not be applied. When the NSPT transform is applied to the residual block, the video encoder can apply the NSPT transform to the residual block and output transform coefficients, and in this case, the MTS and DCT2 + LFNST transforms may not be applied to the residual block.

[0241] Referring to FIG. 19(b), the decoder parses information regarding which conversion method is applied from the bitstream and, based on the parsed information, determines whether to apply one of the inverse conversion methods among MTS, DCT2 + LFNST, and NSPT to the current conversion coefficient. Then, the decoder performs an inverse conversion on the conversion coefficient based on the determined conversion method and obtains a residual block. If the MTS conversion is applied to the current conversion coefficient, the decoder can perform an inverse MTS conversion on the conversion coefficient to obtain a residual block, and in this case, the LFNST and NSPT inverse conversions may not be applied. If the DCT2 + LFNST conversion is applied to the second conversion coefficient, the decoder can perform an inverse LFNST conversion on the second conversion coefficient to output a first conversion coefficient, and perform an inverse DCT2 conversion on the first conversion coefficient to obtain a residual block, and in this case, the MTS and NSPT conversions may not be applied. If the NSPT transform is applied to the current transform coefficients, the decoder can obtain the residual block by performing the inverse NSPT transform on the current transform coefficients, in which case the MTS and DCT2 + LFNST transforms may not be applied.

[0242] The video signal processing device can derive a transformation kernel for each of the MTS, LFNST, and NSPT transformations (or inverse transformations) using an intra prediction mode. Additionally, the video signal processing device can determine which transformation (or inverse transformation) among MTS, DCT2 + LFNST, and NSPT is applied. In this case, to determine the transformation, at least one of the following may be used: the width and height of the current block, whether the component of the current block is a luminance component or a chrominance component, whether the current block is a single tree or a dual tree, whether the current block is encoded in intra mode or inter mode, information on the encoding mode of the current block (e.g., IBC, Intra TMP, Merge, AMVP, GPM, SGPM, CCLM, CCCM), and information on the quantization parameters of the current block.

[0243] FIG. 20 shows a mapping table according to one embodiment of the present specification.

[0244] FIG. 21 shows a conversion type set table according to one embodiment of the present specification.

[0245] FIG. 22 shows a conversion type combination table according to one embodiment of the present specification.

[0246] FIG. 23 shows a threshold value table for IDT conversion types according to one embodiment of the present specification.

[0247] This describes how to select a set of multiple transforms available for the current block in a video signal processing device.

[0248] 1) First, the video signal processing device can derive the values ​​of nSzIdxW and nSzIdxH based on the size of the current block to map the width and height of the current block to a single variable. nSzIdxW may be the minimum value between 3 and the value obtained by calculating the logarithm of 2 of the width of the current block, discarding the decimal places, and then subtracting by 2. nSzIdxH may be the minimum value between 3 and the value obtained by calculating the logarithm of 2 of the height of the current block, discarding the decimal places, and then subtracting by 2.

[0249] 2) Next, the video signal processing device can derive the intra-directional mode (predMode) of the current block. In the case of TIMD mode, the intra-prediction mode expanded from the existing 67 to 131 can be used, reducing the precision to the existing 67 modes.

[0250] 3) Next, the video signal processing device can derive the values ​​ucMode, nMdIdx, and isTrTransposed.

[0251] A. If the current block is encoded in MIP mode, ucMode can be set to '0', nMdIdx to '35', and isTrTransposed to a value derived from MIP.

[0252] B. If the current block is not encoded in MIP mode, ucMode may be set to the intra-directional mode (predMode) of the current block. predMode may represent the index value of the intra-directional mode. predMode may be determined through the extended angle mode based on the aspect ratio of the current block. The video signal processing unit may clip predMode to a value between 2 and 66. If predMode is greater than the diagonal mode, angle mode 34, the isTrTransposed value may be set to 1, and if predMode is less than or equal to 34, the isTrTransposed value may be set to 0. If predMode is greater than 34, the video signal processing unit resets the value of predMode to the value obtained by subtracting predMode from the value obtained by adding 1 to 67 (the maximum value of the intra-directional mode index). For example, if predMode is 35, the video signal processing device can reset angle mode 35 to angle mode 33(67+1-35(predMode)). If predMode is 66, the video signal processing device can reset angle mode 66 to angle mode 2(67+1-66(predMode)). That is, by making it symmetrical with respect to angle mode 34, which is a diagonal mode, the size of the transformation mapping table in FIG. 20 is reduced by about half.

[0253] 4) The video signal processing device can derive the value of nSzIdx through the values ​​of nSzIdxW, nSzIdxH, and isTrTransposed. If the value of isTrTransposed is '1', the value of nSzIdx can be set by multiplying nSzIdxH by 4 and adding nSzIdxW. If the value of isTrTransposed is '0', the value of nSzIdx can be set by multiplying nSzIdxW by 4 and adding nSzIdxH.

[0254] 5) The video signal processing device can derive nTrSet, which is an index of a set of available transformation types according to a predefined table of FIG. 20, using nSzIdx, which is information on the size of the current block, and nMdIdx, which is information on the intra-directional mode of the current block. FIG. 20 defines an index of a set of transformation types according to the intra-directional mode (0 to 34 and MIP) of the current block and the size index (0 to 15) of the current block. Referring to FIG. 20, nTrSet can be 80, and if the size of the current block is 4x8 and the intra-directional mode of the current block is 13, nTrSet can be '7'.

[0255] 6) A video signal processing device can parse mts_idx included in the bitstream to derive a set of conversion types corresponding to nTrSet from the table of FIG. 21. The vertical and horizontal conversion types are set differently depending on whether predMode is greater than the diagonal mode, angle mode 34. Numbers 0 to 79 in the vertical column of FIG. 21 correspond to nTrSet, and numbers 0 to 3 in the horizontal column may correspond to mts_idx. Referring to FIG. 21, if nTrSet is 7 and the value of mts_idx is 3, then 22 may be selected from (2, 17, 18, 22). Then, DST1 and DCT5 corresponding to index 22 of the conversion type combination table of FIG. 22 are selected, and the vertical conversion type of the current block may be set to DST1 and the horizontal conversion type to DCT5. The horizontal columns 0 to 24 of FIG. 22 are indices selected through FIG. 21, and the vertical columns 0 to 1 may represent a vertical direction conversion type and a horizontal direction conversion type, respectively. If the in-screen predicted direction mode of the current block is greater than 34, which is a diagonal mode, the vertical and horizontal direction conversion types may be swapped.

[0256] If the mts_idx value is '3' and the width and height of the current block are both 16 or less, the vertical or horizontal direction conversion type can be reset to the IDT conversion type through the process described below.

[0257] If the absolute difference between the index of the in-frame predicted directional mode of the current block and the index of the horizontal directional mode, 18, is less than an arbitrary predetermined value, the vertical directional transformation type can be reset to the IDT transformation type. If the absolute difference between the index of the in-frame predicted directional mode of the current block and the index of the horizontal directional mode, 50, is less than an arbitrary predetermined value, the horizontal directional transformation type can be reset to the IDT transformation type. In this case, the arbitrary predetermined value is an integer and can be determined based on the width or height of the current block. For example, the arbitrary predetermined value can be determined through the table in FIG. 20. The table in FIG. 23 (a) shows a case where the threshold value is set differently whenever the width or height differs by 4, and the table in FIG. 23 (b) shows a case where the threshold value is set differently whenever the width or height differs by a factor of 2. When the size of the current block is 16x16, the video signal processing device may not reset the vertical directional transformation type to the IDT transformation type and may maintain the existing transformation type.

[0258] The encoder and decoder can reduce the amount of bits required to encode mts_idx by adaptively changing the number of transformation type sets per block. The encoder and decoder can determine the number of transformation type sets to be 1, 4, or 6 depending on the sum of the absolute values ​​of the transformation coefficients of the current transformation block. If the number of transformation type sets changes, the maximum number of bins for signaling mts_idx may change. If the sum of the absolute values ​​of the transformation coefficients is less than or equal to 6, the number of transformation type sets is 1. Therefore, the encoder may not signal mts_idx. In this case, the decoder may not parse mts_idx. If the sum of the absolute values ​​of the transformation coefficients is greater than 6 and less than or equal to 32, the number of transformation type sets may be 4. If the sum of the absolute values ​​of the transformation coefficients is greater than 32, the number of transformation kernel candidates may be 6. In this case, the encoder can signal mts_idx based on the maximum number of transformation type sets. Additionally, the decoder can parse mts_idx based on the maximum number of transformation type sets.

[0259] If the current block is predicted to be in Inter-Coding mode, the encoder and decoder may use one of four combinations of transform kernel sets, such as {(DST7, DST7), (DST7, DCT8), (DCT8, DST7), (DCT8, DCT8)}. The encoder may signal an index indicating which transform kernel set was used. The decoder may parse the index to determine the transform kernel set to use for the current transform block. If the set transform kernel set is (DST7, DCT8), the encoder and decoder transform or inversely transform the current block using DST7 in the horizontal direction and transform or inversely transform the current block using DCT8 in the vertical direction. For complexity optimization, the encoder and decoder may set the maximum CU size to which Inter MTS can be applied for images larger than 1920x1080 resolution to 32x32. The encoder and decoder set the maximum CU size to which Inter MTS can be applied for images with resolutions smaller than 1920x1080 to 16x16. Therefore, the encoder and decoder can apply DCT2 transformations in the horizontal and vertical directions without applying Inter MTS to transformation blocks larger than 16x16, such as transformation blocks of size 32x32. Additionally, the encoder and decoder can use separable KLT instead of DST7 and DCT8 for transformation and inverse transformation of transformation blocks that are equal to or smaller than 16x16.

[0260] FIG. 24 shows the block boundary and samples around the boundary in a deblocking filtering process according to one embodiment of the present specification.

[0261] Referring to Fig. 24 (a), the area indicated by the dotted line between block P and block Q may represent a block boundary. Block boundaries may exist at any given size, and block boundaries may exist at every size of 4.

[0262] Figure 24 (b) shows samples in which filtering is performed based on block boundaries. A video signal processing device can perform deblocking filtering on the currently restored block to mitigate blocking phenomena occurring at block boundaries. The deblocking filtering process may include determining the transform block boundary, determining the sub-block boundary, determining the length of the filter to be filtered, determining the filtering strength (bS), determining the filtering parameters, determining whether to perform filtering, and determining the type of filtering.

[0263] In independent scalar quantization, the restored coefficient t'k for an input coefficient tk depends only on the associated quantization index qk. That is, the quantization index for any restored coefficient has a different value from the quantization indices for other restored coefficients. Here, t'k may be a value containing the quantization error in tk, and may be different or the same depending on the quantization parameter. Here, t'k can be referred to as the restored transform coefficient or the inverse quantized transform coefficient, and the quantization index can be referred to as the quantized transform coefficient.

[0264] In Uniform Reconstruction Quantizers (URQ), the reconstructed coefficients are arranged at equal intervals. The distance between two adjacent reconstructed values ​​can be called the quantization step size. The reconstructed values ​​may include zero, and the entire set of available reconstructed values ​​can be uniquely defined according to the quantization step size. The quantization step size may vary depending on the quantization parameter.

[0265] In conventional methods, quantization reduces the set of acceptable reconstructed transformation coefficients, and the number of elements in this set can be finite. Consequently, there are limitations in minimizing the average error between the original image and the reconstructed image. Vector quantization can be used as a method to minimize this average error.

[0266] A simple form of vector quantization method used in video encoding is sign data hiding. This is a method in which the encoder does not encode the sign for a single non-zero coefficient, and the decoder determines the sign for that coefficient based on whether the sum of the absolute values ​​of all coefficients is even or odd. To achieve this, at least one coefficient in the encoder can be increased or decreased by '1', and at least one coefficient can be selected and adjusted to be optimal in terms of the cost of rate distortion. As an example of implementation, a coefficient having a value close to the boundary of the quantization interval can be selected.

[0267] Another vector quantization method is Trellis-Coded Quantization, which is utilized in video encoding as an optimal path search technique to obtain optimized quantization values ​​in dependent quantization. In block units, quantization candidates for all coefficients within a block are placed on a trellis graph, and the optimal trellis path between the optimized quantization candidates is searched while considering the cost of rate distortion. Specifically, dependent quantization applied in video encoding can be designed so that the set of acceptable restored transform coefficients depends on the value of the transform coefficient preceding the current transform coefficient in the restoration order. In this case, by selectively using multiple quantizers based on the transform coefficient, the average error between the original image and the restored image is minimized, thereby increasing encoding efficiency.

[0268]

[0269] FIG. 25 shows PDPC filtering according to an embodiment of the present invention.

[0270] Discontinuous changes in sample values ​​may occur at the boundary between an intra prediction block and an adjacent block. These discontinuities may occur at the top and left boundary portions of the prediction block. After the prediction block is generated, the video signal processing unit can perform filtering within the prediction block. Through this, the video signal processing unit can eliminate the discontinuities between the intra prediction block and the adjacent block. Specifically, the video signal processing unit can perform filtering by weighting the value predicted in intra mode with the value of one or more pre-restored surrounding samples. In this case, the surrounding samples may be selected from samples at pre-specified locations or based on the direction of the intra mode. This is referred to as PDPC (position-dependent intra prediction combination). In PDPC, the video signal processing unit can determine the filtering method according to the intra prediction directionality mode. Figure 25 (a) shows the PDPC filtering method when the intra prediction directionality mode is horizontal. In FIG. 25(a), when the encoder and decoder perform filtering on the (1, 0) sample within the prediction block, the video signal processing device may perform filtering on the (1, 0) sample by weighted averaging at least one of the restored surrounding samples R(-1, 0), R(-1, -1), and R(1, -1). At this time, the video signal processing device may apply the difference between the intra-reference samples to the (1, 0) sample within the prediction block. FIG. 25(b) shows a PDPC filtering method for cases where the intra-prediction directional mode is not a planar mode, DC mode, horizontal direction mode, and vertical direction mode. In FIG. 25(b), the encoder and decoder may perform filtering on the (1, 1) sample by weighted averaging at least one of the restored surrounding samples R(3, -1) and R(-1, 3) with the (1, 1) sample in the prediction block.Specifically, if the (1, 1) sample within the prediction block is predicted from the upper surrounding sample R(3, -1), the video signal processing device may apply compensation to the (1, 1) sample using the left surrounding sample R(-1, 3) indicated by the intra prediction directionality mode. Here, R(x, y) may be a restored sample adjacent to the current block. The video signal processing device may use the intra prediction mode to perform a prediction for the sample within the prediction block from the first surrounding sample corresponding to the sample within the prediction block in the first reference sample. The video signal processing device may perform PDPC using the sample within the prediction block and / or the second surrounding sample corresponding to the first reference sample in the second reference sample.

[0271] As with the embodiments described above, when PDPC is applied, filtering is performed on samples within the prediction block based on the difference between neighboring samples. Depending on the intra prediction mode, it may not be possible to derive a left neighboring sample corresponding to an upper neighboring sample. Depending on the intra prediction mode, it may not be possible to derive an upper neighboring sample corresponding to a left neighboring sample. The video signal processing device may not apply PDPC in an intra prediction where an intra prediction mode is applied in which the difference in change between neighboring samples from the upper neighboring sample and the left neighboring sample cannot be derived. The video signal processing device may determine whether to apply PDPC to the current block based on the width or height of the current block and at least one of the intra prediction mode. In this case, the upper neighboring sample and the left neighboring sample may include the upper-left neighboring sample. At least one of the first reference sample and the second reference sample may be a sample generated through filtering between neighboring samples. In this case, the filtering between neighboring samples may be one of a cubic filter, a Gaussian filter, a low-pass filter, a high-pass filter, and a smoothing filter. The smoothing filter may be a 1:2:1 filter.

[0272] FIG. 26 shows a gradient PDPC according to an embodiment of the present invention.

[0273] If the difference between surrounding samples cannot be derived according to the intra prediction mode, the video signal processing device may not use the surrounding samples corresponding to the sample within the prediction block, but may apply the difference between the surrounding samples corresponding to the intra prediction mode to the sample within the prediction block. In the embodiment of FIG. 46, the video signal processing device does not perform the PDPC process using the left reference sample corresponding to the sample (X) within the prediction block. The encoder and decoder may perform PDPC for the sample (X) within the prediction block by using the difference between the sample r(-1+d, -1) and the sample r(-1, y) among the reference samples that correspond to the current intra prediction mode. This is referred to as gradient PDPC. y may be the vertical coordinate of the sample within the prediction block, and d may be the decimal position of the sample pointed to by the intra prediction mode. Additionally, y may be changed to a horizontal coordinate according to the intra prediction mode. In this case, the surrounding samples may be the sample r(-1, -1+d) and the sample r(x, -1). The video signal processing device may perform PDPC by using an intra-prediction mode to select a first reference sample and a second reference sample that do not correspond to a sample within the prediction block, and applying the difference between the first reference sample and the second reference sample to the sample within the prediction block. At least one of the first reference sample and the second reference sample may be derived using one of the vertical and / or horizontal coordinates in the sample within the prediction block. Additionally, the video signal processing device may have at least one of the first reference sample and the second reference sample as a top-left sample. In this case, at least one of the first reference sample and the second reference sample may be a sample generated through filtering between neighboring samples. Additionally, the filtering between neighboring samples may be one of a cubic filter, a Gaussian filter, a low-frequency filter, a high-frequency filter, and a smoothing filter. The smoothing filter may be a 1:2:1 filter.

[0274] CIIP PDPC refers to a mode that applies PDPC to intra-prediction blocks in CIIP mode.

[0275] FIG. 27 shows a matrix-based intra-prediction method according to an embodiment of the present invention.

[0276] A method of performing intra prediction using a model (or matrix) that infers the relationship between the current block and a reference sample through a large amount of image data can be called a matrix-based position-dependent intra prediction (PDP). In the embodiment of FIG. 27, a video signal processing device can generate a prediction block P(x,y) for the current block through a calculation (Fig. 27 (b)) between a restored reference sample (r) r(k) adjacent to the current block (p) to be predicted and a predefined matrix F(x,y,k). The matrix may vary depending on the size of the current block and the intra prediction directionality mode. Additionally, the size of the reference sample may vary depending on the size of the current block. If the width and height of the current block are smaller than or equal to a predefined first size, T1 and T2 may be a predefined second size. If the width and height of the current block are larger than the predefined first size, T1 and T2 may be a predefined third size. In this case, the predefined first size may be 16. Additionally, a pre-specified second size may be 2. Additionally, a pre-specified third size may be 1. Also, W and H in FIG. 27 represent the width and height of the current block, respectively. Unlike MIP, when conditions are met for applying a PDP to the current block, the video signal processing device may generate a prediction block for the current block using matrix-based intra prediction via the PDP, rather than using intra prediction using an intra prediction directional mode. At this time, whether a PDP can be applied to the current block can be determined using at least one of the size of the current block, the intra prediction directional mode of the current block, whether the current block is a luminance block or a chrominance block, whether an inferred matrix exists, whether a reference sample exists, the encoding mode of the current block, and a reference sample line.If the encoding mode of the current block is one of DIMD, OBIC, CIIP, SGPM, TIMD, TMRL, Directional Planar, EIP, MIP, ISP, IBC, Intra TMP, BDPCM, Platte, Inter mode, LM, and CCCM, the video signal processing device may not be able to use a PDP-based method for the current block.

[0277] In CCCM, the autocorrelation matrix is ​​calculated using the reconstructed values ​​of luminance and chroma samples. Since these samples cover the full range (from 0 to 1023 for 10-bit content), the value of the autocorrelation matrix is ​​relatively large. This requires computational processes at high bit depths during model parameter calculation. Differentiating meanY and meanNonlinY values ​​can address this issue. However, the method of differentiating mean values ​​has the disadvantage of increasing implementation complexity because it requires additional pipeline steps to calculate the mean. To mitigate this drawback, we propose a method that reduces the magnitude of values ​​used for model generation and decreases the precision required for fixed-point operations by differentiating the luminance and chroma samples by arbitrary fixed values ​​for each model. Consequently, it becomes possible to use 16-bit decimal precision instead of the 22-bit precision of the original CCCM implementation.

[0278] Figure 28 illustrates an offset difference sample-based CCCM method.

[0279] Autocorrelation matrix

[0280] C' = C - offsetLuma

[0281] N' = N - offsetLuma

[0282] S' = S - offsetLuma

[0283] E' = E - offsetLuma

[0284] W' = W - offsetLuma

[0285] P' = nonLinear(C')

[0286] B = midValue = 1 << (bitDepth - 1)

[0287] The chroma value is predicted using the following mathematical formula 1, where offsetChroma uses offsetCr and offsetCb values ​​for the Cr and Cb components, respectively.

[0288] Mathematical formula 1

[0289] predChromaVal = c0C' + c1N' + c2S' + c3E' + c4W' + c5P' + c6B + offsetChroma

[0290] In this case, if the pixel value at position A in Fig. 28, which is used as the offsetLuma value, is '0', the difference process is omitted, so calculation may need to be performed again at the existing high bit depth. To solve this problem, various methods can be applied as follows.

[0291] For example, if the pixel value at position A in FIG. 28, which is used as the offsetLuma value, is smaller than an arbitrary first value or larger than an arbitrary second value, the encoder and decoder may use predefined default offset values ​​as offsetLuma, offsetCr, and offsetCb values. Here, the arbitrary first and second values ​​may be integers greater than or equal to 1 and may vary depending on the bit depth used to represent the current sample; if the bit depth is 10 bits, they may be values ​​such as 0, 128, 256, 512, 768, 1023, etc. Additionally, the predefined default offsetLuma, offsetCr, and offsetCb values ​​may be integers and may vary depending on the bit depth used to represent the current sample; if the bit depth is 10 bits, they may be values ​​such as -256, -128, 0, 128, 256, etc.

[0292] In another example of an embodiment, if the pixel value at position A in FIG. 28 is smaller than the defined arbitrary first value or larger than the defined arbitrary second value, the encoder and decoder can sequentially search for samples at positions B, C, D, E, F, G, a, b, c, d, e, f, and g in FIG. 28, check whether the sample at the corresponding position is within a valid range, and if valid, use the sample value at the corresponding position as the offsetLuma, offsetCr, and offsetCb values. Here, a predefined range can be used to set the valid range; for example, the encoder and decoder may determine that it is valid if it is greater than any integer, and the arbitrary integer may vary depending on the bit depth used to represent the current sample, and if the bit depth is 10 bits, it may be a value such as 128 or 256.

[0293] As another example of an embodiment, samples at positions A, B, C, D, E, F, G, a, b, c, d, e, f, g in FIG. 28 are searched in order, and after checking whether the sample at the corresponding position is within a valid range, the encoder can apply CCCM using all sample values ​​at the corresponding position if it is valid, and then signal by including information about the most optimal sample position in the bitstream. The decoder can parse the sample position information and then set offsetLuma, offsetCr, and offsetCb values ​​to be used for performing CCCM in the current block based on the corresponding sample.

[0294] FIG. 29 shows an arbitrary number of luminance samples surrounding the luminance sample of the current block. FIG. 29 (a) shows four surrounding luminance samples, and FIG. 29 (b) shows eight surrounding luminance samples.

[0295] In an exemplary embodiment, the encoder and decoder compare the luminance sample of the current block and an arbitrary number of luminance samples surrounding the luminance sample of the current block (Fig. 29 (b)) with the first threshold to calculate the number of samples belonging to the first linear model and the second linear model. If the number of samples in the first linear model is equal to or greater than the first arbitrary value, the color difference sample corresponding to the luminance sample of the current block can be predicted using only the first linear model. Otherwise, if the number of samples in the second linear model is equal to or greater than the first arbitrary value, the color difference sample corresponding to the luminance sample of the current block can be predicted using only the second linear model. Here, the first arbitrary value may be '6'. Otherwise, if the number of first linear models is greater than the number of second linear models, a first weight is set for the first linear models and a second weight is set for the second linear models; then, first color difference samples derived using the first linear models and second color difference samples derived using the second linear models are derived, and finally, a color difference sample corresponding to the luminance sample of the current block can be derived by weighted averaging using the first weight for the first color difference samples and the second weight for the second color difference samples. Otherwise (if the number of first linear models is equal to or less than the number of second linear models), a first weight is set for the second linear models and a second weight is set for the first linear models; then, first color difference samples derived using the first linear models and second color difference samples derived using the second linear models are derived, and finally, a color difference sample corresponding to the luminance sample of the current block can be derived by weighted averaging using the second weight for the first color difference samples and the first weight for the second color difference samples. Here, the first weight can be '13' and the second weight can be 3.

[0296] In another exemplary embodiment, the encoder and decoder may select and use one of the following methods: deriving a color difference sample using only a first linear model, deriving a color difference sample using only a second linear model, or deriving a color difference sample using a weighted average of both the first and second linear models, based on the difference between the luminance sample of the current block and the first threshold (average value of luminance samples of surrounding blocks) and the difference between the second threshold or the third threshold. If the difference between the luminance sample of the current block and the first threshold (average value of luminance samples of surrounding blocks) is greater than the second threshold, the luminance sample of the current block is compared with the first threshold. If the luminance sample of the current block is greater than the first threshold, the color difference sample corresponding to the luminance sample of the current block can be predicted using the second linear model. Otherwise, if the luminance sample of the current block is equal to or less than the first threshold, the color difference sample corresponding to the luminance sample of the current block can be predicted using the first linear model. Otherwise, if the difference between the luminance sample of the current block and the first threshold (average value of luminance samples of surrounding blocks) is equal to or smaller than the second threshold, and the difference between the luminance sample of the current block and the first threshold (average value of luminance samples of surrounding blocks) is smaller than the third threshold, then a third weight is set in the first linear model and a fourth weight is set in the second linear model; then, after deriving a first color difference sample derived using the first linear model and a second color difference sample derived using the second linear model, the first color difference sample is weighted averaged using the third weight and the second color difference sample is weighted averaged using the fourth weight to finally derive a color difference sample corresponding to the luminance sample of the current block. Here, the third weight can be '1' and the fourth weight can be '1'.Otherwise (if the difference between the luminance sample of the current block and the first threshold (average value of the luminance samples of surrounding blocks) is equal to or greater than the third threshold), the luminance sample of the current block is compared with the first threshold. If the luminance sample of the current block is less than or equal to the first threshold, a first weight is set for the first linear model and a second weight is set for the second linear model. Then, a first color difference sample derived using the first linear model and a second color difference sample derived using the second linear model are derived. Finally, a color difference sample corresponding to the luminance sample of the current block can be derived by applying the first weight to the first color difference sample and the second weight to the second color difference sample to perform a weighted average. Otherwise (if the luminance sample of the current block is greater than the first threshold), a first weight is set in the second linear model and a second weight is set in the first linear model, and then a first color difference sample derived using the first linear model and a second color difference sample derived using the second linear model are derived, and then the first color difference sample is weighted and averaged using the second weight and the second color difference sample is weighted and the first weight is applied to the first color difference sample to derive the final color difference sample corresponding to the luminance sample of the current block. Here, the first weight can be '13' and the second weight can be 3.

[0297] Alternatively, a color difference sample corresponding to the luminance sample of the current block can be derived by using at least one of a single linear model derived from the surrounding blocks of the current block and two linear models derived based on the average value of the surrounding blocks of the current block. That is, the encoder and decoder can predict a color difference sample corresponding to the luminance sample of the current block by comparing the luminance sample of the current block, an arbitrary number of luminance samples surrounding the luminance sample of the current block, and the luminance sample of the current block and an arbitrary number of luminance samples surrounding the luminance sample of the current block with the first threshold, using at least one of the number belonging to the first linear model and the second linear model, the difference between the luminance sample of the current block and the first threshold, the first linear model, the second linear model, the third linear model, the weight for the first linear model, the weight for the second linear model, and the weight for the third linear model. Here, the first linear model and the second linear model are linear models derived using the first threshold, and the third linear model is a linear model derived using all samples of the surrounding blocks.

[0298] In an exemplary embodiment, the encoder and decoder compare the luminance sample of the current block and an arbitrary number of luminance samples surrounding the luminance sample of the current block (Fig. 29 (b)) with the first threshold to calculate the number of samples belonging to the first linear model and the second linear model. If the number of samples in the first linear model is equal to or greater than the first arbitrary value, the color difference sample corresponding to the luminance sample of the current block can be predicted using only the first linear model. Otherwise, if the number of samples in the second linear model is equal to or greater than the first arbitrary value, the color difference sample corresponding to the luminance sample of the current block can be predicted using only the second linear model. Here, the first arbitrary value may be '6'. Otherwise, if the number of first linear models is greater than the number of second linear models, a first weight is set for the first linear models and a second weight is set for the third linear models; then, first color difference samples derived using the first linear models and second color difference samples derived using the third linear models are derived, and then, the first weight is applied to the first color difference samples and the second weight is applied to the second color difference samples to obtain a weighted average, thereby finally deriving a color difference sample corresponding to the luminance sample of the current block. Otherwise (if the number of first linear models is equal to or less than the number of second linear models), a first weight is set for the second linear models and a second weight is set for the third linear models; then, first color difference samples derived using the second linear models and second color difference samples derived using the third linear models are derived, and then, the first weight is applied to the first color difference samples and the second weight is applied to the second color difference samples to obtain a weighted average, thereby finally deriving a color difference sample corresponding to the luminance sample of the current block. Here, the first weight can be '13' and the second weight can be 3.

[0299] In another exemplary embodiment, the encoder and decoder may select and use one of the following methods: deriving a color difference sample using only a first linear model, deriving a color difference sample using only a second linear model, or deriving a color difference sample using a weighted average of at least one of the first linear model, the second linear model, and the third linear model, based on the difference between the luminance sample of the current block and the first threshold (average value of the luminance samples of surrounding blocks) and the difference between the second threshold or the third threshold. If the difference between the luminance sample of the current block and the first threshold (average value of the luminance samples of surrounding blocks) is greater than the second threshold, the luminance sample of the current block is compared with the first threshold. If the luminance sample of the current block is greater than the first threshold, the color difference sample corresponding to the luminance sample of the current block can be predicted using the second linear model. Otherwise, if the luminance sample of the current block is equal to or less than the first threshold, the color difference sample corresponding to the luminance sample of the current block can be predicted using the first linear model. Otherwise, if the difference between the luminance sample of the current block and the first threshold (average value of the luminance samples of surrounding blocks) is equal to or smaller than the second threshold, and the difference between the luminance sample of the current block and the first threshold (average value of the luminance samples of surrounding blocks) is smaller than the third threshold, then a third weight is set in the first linear model, a fourth weight is set in the second linear model, and a fifth weight is set in the third linear model. After deriving the first color difference sample derived using the first linear model, the second color difference sample derived using the second linear model, and the third color difference sample derived using the third linear model, the first color difference sample is weighted averaged using the third weight, the second color difference sample is weighted averaged using the fourth weight, and the third color difference sample is weighted averaged to finally derive the color difference sample corresponding to the luminance sample of the current block.Here, the third weight can be '1', the fourth weight can be '1', and the fifth weight can be '1'. Otherwise (if the difference between the luminance sample of the current block and the first threshold (average value of luminance samples of surrounding blocks) is equal to or greater than the third threshold), the luminance sample of the current block is compared with the first threshold. If the luminance sample of the current block is less than or equal to the first threshold, a first weight is set for the first linear model and a second weight is set for the third linear model. Then, a first color difference sample derived using the first linear model and a second color difference sample derived using the third linear model are derived. Finally, the first color difference sample is weighted and averaged using the first weight and the second weight, respectively, to derive the color difference sample corresponding to the luminance sample of the current block. Otherwise (if the luminance sample of the current block is greater than the first threshold), a first weight is set in the second linear model and a second weight is set in the third linear model, and then a first color difference sample derived using the second linear model and a second color difference sample derived using the third linear model are derived, and then the first weight is applied to the first color difference sample and the first weight is applied to the second color difference sample to derive a weighted average to finally derive a color difference sample corresponding to the luminance sample of the current block. Here, the first weight can be '13' and the second weight can be 3.

[0300] As shown in Equation 2, Chroma Fusion (CF) can be derived by calculating the weighted average between the chrominance block (pred'C) predicted by a general intra-prediction mode without using LM mode and the current luminance block (rec'L). In this case, the weight parameters (a0, a1, a2) can be derived using the CCCM method. The median (midValue) can be calculated as '1 << (bitDepth - 1)'. Additionally, whether to use one CCCM model or two CCCM models can be signaled and included in the bitstream. In the decoder, if the current chrominance block is encoded in CF mode, information regarding how many CCCM models are used can be parsed and used to predict the current chrominance block.

[0301] Mathematical formula 2

[0302]

[0303]

[0304] GL-CCCM (Gradient and location based convolutional cross-component model) is an additional CCCM mode that utilizes gradient and location information. The existing CCCM mode can derive a color difference sample for the current block using a luminance sample at a location corresponding to the color difference sample location to be predicted, four samples surrounding that luminance sample (Fig. 29 (a)), and coefficient information. At this time, as shown in Fig. 29 (b), the GL-CCCM mode can derive a color difference sample for the current block by reflecting the vertical and horizontal differences (Gx, Gy in Equation 3) for the luminance sample at a location corresponding to the color difference sample location to be predicted and eight samples surrounding that luminance sample, and also by using the location value of the current luminance sample (X, Y in Equation 3) and its coefficient information. At this time, the position value of the current luminance sample can be a value recalculated using the relative position of the current pixel when the top-left position of the reference template used to derive the CCCM model, the top-left position of the current block, and the top-left position of the current block are set to (0,0), an arbitrary offset value, and an arbitrary shift value. At this time, the position value of the current luminance sample can be calculated through (the relative position value of the current pixel when the top-left position of the current block is set to (0,0) + the difference value between the top-left position of the current block and the top-left position of the reference template used to derive the CCCM model + an arbitrary offset value) << an arbitrary shift value). At this time, the arbitrary offset value can be an integer and can be '8'. Also, the arbitrary shift value can be an integer and can be '3'. Additionally, Gx and Gy can be calculated through Equation 4 using the samples of FIG. 29 (b).

[0305] Mathematical formula 3

[0306] predChromaVal = C0 C + C1 Gy+ C2 Gx+C3 Y+ C4 X+c5P+ C6 B

[0307]

[0308] Mathematical formula 4

[0309] G y = (2N + NW + NE) - (2S + SW + SE)

[0310] G x =(2W + NW + SW) - (2E + NE + SE)

[0311]

[0312] Coefficient information (or weight parameter information) can be derived using an arbitrary region of already restored surrounding blocks around the current block. Additionally, the encoder can include information regarding whether GL-CCCM was used in the current block in the bitstream by signaling it via a flag. The decoder can parse the GL-CCCM information for the current block to determine whether to apply GL-CCCM to the current block to predict the chrominance block.

[0313] In GL-CCCM, the luminance samples used to calculate the vertical and horizontal differences of the luminance samples may include not only the area around the current luminance sample but also surrounding luminance samples located at an arbitrary position. In this case, the arbitrary position may be '1' in FIG. 19, and may have values ​​such as 2, 3, 4, etc.

[0314] The LM mode, which predicts the current chrominance block using the restored current luminance block, can apply a downsampling filter to the luminance block to make its resolution equal to that of the chrominance block due to the resolution difference between the luminance and chrominance blocks. Due to this downsampling, edge component information of the luminance block may be reduced. Therefore, the chrominance block can be predicted using samples of the luminance block before downsampling.

[0315]

[0316] Figure 30 shows luminance samples before downsampling used to derive color difference samples in CCCM-ND mode.

[0317] The CCCM-ND (CCCM using non-downsampled luma samples) mode is a method for predicting a chrominance block using samples of the luminance block before downsampling. FIG. 30 shows the chrominance sample location C to be predicted and the six luminance sample locations corresponding to that chrominance sample location, where the luminance samples are samples before downsampling. The chrominance samples are calculated using Equation 5 and can be calculated using a linear model using six luminance samples (samples L0, L1, L2, L3, L4, and L5 in FIG. 30) and a non-linear model using four luminance samples (samples L0, L3, L2, and L1 in FIG. 30). In Equation 5, offsetLuma and offsetchroma have the same meaning as the difference values ​​described in the specification. B can be calculated as '1 << (bitDepth - 1)'. The LDL decomposition method used in CCCM can be used to derive the coefficients (a0 ~ a10).

[0318]

[0319] Mathematical formula 5

[0320]

[0321]

[0322] The encoder can include information regarding whether CCCM-ND is applied to the current block in the bitstream by signaling it via a flag. The decoder can parse this information to set whether CCCM-ND is applied to the current block.

[0323] In the encoder, if CCCM-ND is applied to the current block, information regarding GL-CCCM may not be signaled. In the decoder, if CCCM-ND is applied to the current block, information related to GL-CCCM may not be parsed, and it may be configured not to apply GL-CCCM to the current block.

[0324] In GL-CCCM, downsampled luminance samples are used to derive the model, and downsampled luminance samples of the current block are used to predict color difference samples for the current block. In this case, as with CCCM-ND, GL-CCCM can also use luminance samples before downsampling to derive the model or predict color difference samples.

[0325] Mathematical formula 6

[0326]

[0327] The color difference sample is calculated using Equation 6, and a linear model using six luminance samples (L0, L1, L2, L3, L4, L5 luminance samples in FIG. 30) and a non-linear model using four luminance samples (L0, L3, L2, L1 luminance samples in FIG. 30) are used. The vertical and horizontal differences (Gx, Gy in Equation 6) between the luminance sample at the position corresponding to the color difference sample to be predicted and the samples surrounding that luminance sample (L1 to L8 luminance samples in FIG. 30) are reflected. Additionally, the color difference sample for the current block can be derived using the position value of the current luminance sample (X, Y in Equation 6) and its coefficient information. At this time, Gx and Gy can be calculated using Equation 7 with the samples in FIG. 30. In Equation 6, offsetLuma and offsetchroma have the same meaning as the difference values ​​described in the specification. B can be calculated as '1 << (bitDepth - 1)'. To derive the coefficients (a0 ~ a14), the LDL decomposition method used in CCCM can be used.

[0328] Mathematical formula 7

[0329] Gy = (2L6 + L7 + L8) - (2L3 + L4 + L5)

[0330] Gx = (2L1+ L7 + L4) - (2L2 + L8 + L5)

[0331]

[0332] FIGS. 31 and 32 show structural diagrams for the application of a cross-component residual model (CCRM) according to one embodiment of the present invention.

[0333] Referring to FIG. 31, when an inter-coding mode is applied to the current block, the video signal processing device can derive a CCP model (linear and non-linear model) between a luminance prediction block (Y') derived using motion information of the current block and a first chrominance prediction block (Cb', Cr') derived using motion information of the current block (Derive filter), and then generate a restored luminance block of the current block using the luminance error block of the current block. The video signal processing device can generate a second chrominance prediction block of the current block by applying the derived CCP model to the restored luminance block of the current block (Apply filter). The video signal processing device can finally generate (acquire) the chrominance block (Cb, Cr) of the current block by adding the chrominance error block to the second chrominance prediction block predicted using the CCP model.

[0334] Here, the CCP models include CCLM, MMLM, GLM, CCCM, MM-CCCM, GL-CCCM, CCCM-ND, and CCCM-MDF, where CCLM and MMLM are linear models, and GLM, CCCM, MM-CCCM, GL-CCCM, CCCM-ND, and CCCM-MDF may be non-linear models. That is, the CCP model may be one of CCLM, MMLM, GLM, CCCM, MM-CCCM, GL-CCCM, CCCM-ND, and CCCM-MDF. Additionally, referring to FIG. 31, when an inter-coding mode is applied to the current block, the video signal processing device can derive a CCP model between the luminance prediction block (Y') of the current block and the first chrominance prediction block (Cb', Cr') of the current block (Derive filter), and then generate a restored luminance block of the current block using the luminance error block of the current block. The video signal processing device can generate a second color difference prediction block of the current block by applying the derived CCP model to the restored luminance block of the current block. The video signal processing device can generate a third color difference prediction block through the weighted average between the second color difference prediction block and the first color difference prediction block of the current block predicted using the CCP model. At this time, the video signal processing device may apply a first weight to the first color difference prediction block and a second weight to the second color difference prediction block. The first weight and the second weight may be predefined integer values; for example, the first weight may be 3 and the second weight may be 13. The video signal processing device can finally generate the color difference block (Cb, Cr) of the current block by adding a color difference error block to the third color difference prediction block. The encoder may generate and signal a bitstream containing information about the first weight and the second weight (or information indicating the optimal weight value of the first weight and the second weight).When CCRM mode is applied to the current block, the decoder can parse information about the weights (or information indicating the optimal weight values ​​of the first and second weights) and use it to predict the color difference block of the current block.

[0335] Additionally, referring to FIG. 32, when the current block is in an inter-coding mode, the video signal processing device may derive one or more CCP models using motion information used to generate prediction blocks for the current block, and use them to adaptively / selectively predict the chrominance block of the current block. The encoder may acquire and signal a bitstream containing information about which model was used. The decoder may parse information about which model was used and use it to predict the chrominance block of the current block. When the current block is in an inter-coding mode and bidirectional motion prediction is applied to the current block, the video signal processing device may derive a CCP model using the predicted luminance block and chrominance block of the current block, which are weighted averaged through bidirectional motion. Alternatively, when the current block is in an inter-coding mode and bidirectional motion prediction is applied to the current block, the video signal processing device may derive a CCP model using the predicted luminance block and chrominance block of the current block, which are generated through L0 or L1 motion. When bidirectional motion prediction is applied to the current block, the encoder can signal motion information in the bitstream indicating which motion information, between L0 and L1, was used to derive the CCP model. The decoder can parse the motion information to determine L0 or L1, select one of the derived CCP models, and use it to generate the chrominance prediction block for the current block. Then, the decoder can add a chrominance error block to the generated chrominance prediction block to finally generate the chrominance block (Cb, Cr) for the current block. To reduce the amount of bits being signaled, the encoder may not signal information regarding which motion information, between L0 and L1, was used to derive the CCP model, but instead use the quantization parameter information used in the L0 and L1 reference blocks and derive the CCP model using a reference block composed of a higher quality.The decoder can derive a CCP model by using the quantization parameter information used in the L0 and L1 reference blocks and a reference block composed of a higher quality. Additionally, the video signal processing device can derive a CCP model by using the L0 and L1 motion information and a weighted averaged reference block (luminance reference block and chrominance reference block).

[0336] FIG. 33 shows the downsampling filter and the position of the sample applied to the luminance block according to one embodiment of the present invention.

[0337] In the CCCM method, which predicts chrominance blocks based on luminance blocks, the video signal processing unit may apply various downsampling filters to reduce the resolution of the luminance block to that of the chrominance block in order to match the resolution difference between the luminance block and the chrominance block. A CCCM mode that applies downsampling filters can be described as CCCM-MDF (CCCM with multiple downsampling filters).

[0338] FIG. 33(a) illustrates multiple downsampling filters that can be applied to a luminance block. Filtering using the H filter in FIG. 33(a) is a filtering method that considers the value of the current luminance sample location to be more important than surrounding samples. Filtering using filters G1, G2, and G3 in FIG. 33(a) may be a filtering method designed based on the degree of variation between adjacent samples. Filtering using the G1 filter may be filtering based on variation in the horizontal direction. Filtering using the G2 filter may be filtering based on variation in the vertical direction. Filtering using the G3 filter may be filtering based on variation in the diagonal direction. C in FIG. 33(b) is the location of the current color difference sample, and W, E, N, S, NW, NE, SW, and SE represent the locations of samples surrounding C.

[0339] Mathematical formula 8

[0340]

[0341] Mathematical formula 8 represents a calculation formula for deriving a color difference sample using at least one of the downsampled luminance samples obtained based on filtering using the H, G1, G2, and G3 filters of FIG. 33.

[0342] In Equation 8, P is a non-linear model and can be the square of the sample value at the corresponding position. In Equation 8, B can be half the maximum value of the current sample and can be the median value of the bit depth. For 10-bit content, the value of B can be 512.

[0343] Model 1 of Equation 8 may be an equation that derives color difference samples using downsampled luminance blocks obtained using four filters (using filters H, G1, G2, and G3 of FIG. 33). Model 2 of Equation 8 may be an equation that derives color difference samples using downsampled luminance blocks obtained using filters H and G1 of FIG. 33. Model 3 of Equation 16 may be an equation that derives color difference samples using downsampled luminance blocks obtained using filters H and G3 of FIG. 33.

[0344] The encoder can select the optimal model among the three color difference prediction blocks derived using Equation 8, and then signal by including information about which model is used when predicting the current color difference block in the bitstream. At this time, information about which model is used can be applied to the color difference signal Cb and the color difference signal Cr, respectively, on a block-by-block basis. The decoder can predict each color difference block using the model determined by parsing the information about which model is used for each color difference signal (Cb, Cr).

[0345] Alternatively, the encoder may select the optimal model among the three color difference prediction blocks derived through the above mathematical formula 8, and then integrally signal information regarding which model to use when predicting the current color difference blocks to the color difference signals (Cb, Cr), and include it in the bitstream in block units. The decoder may parse information regarding which model to use to generate the current color difference prediction block, and generate the color difference prediction block by applying it equally to the current Cb color difference block and the current Cr color difference block.

[0346] When two CCCM models are used in the current block, the video signal processing device can derive two models based on the average value of the downsampled luminance block generated through filtering using the H filter of FIG. 33 (a).

[0347] FIG. 34 shows the coefficients of filtering applied to a prediction block according to one embodiment of the present invention and the filtering positions.

[0348] The decoder can generate a filtered color difference prediction block by performing filtering using the color difference prediction block generated by applying the CCP model and the surrounding color difference samples of the restored current block. At this time, the video signal processing device can filter the current color difference prediction block using the filtering coefficients of FIG. 26(a).

[0349] In addition, to apply filtering to all sample locations of the current color difference prediction block, the video signal processing device may construct the samples surrounding the current color difference prediction block using restored samples adjacent to the current color difference block, as shown in Fig. 34 (b), and construct the unrestored parts by padding the boundary samples of the current color difference prediction block.

[0350] The encoder can acquire and signal a bitstream containing information regarding whether filtering is performed. The decoder can parse this information to determine whether to apply filtering to the current chrominance block. Filtering can be applied block by block; that is, filtering can be applied to the current luminance block, Cb chrominance block, and Cr chrominance block. In this case, the same filtering may be applied to the current Cb chrominance block and the Cr chrominance block.

[0351] If the current block is in GLM encoding mode or CCLM mode where offsets are added, filtering may not be applied to the current block. That is, if the current block is in GLM encoding mode or CCLM mode where offsets are added, the encoder may not include information in the bitstream regarding whether filtering is performed on the current chrominance block. If the current block is in GLM encoding mode or CCLM mode where offsets are added, the decoder may not parse information regarding whether filtering is performed on the current chrominance block and may determine that filtering is not applied to the current chrominance block.

[0352] FIG. 35 illustrates a method for determining a CCP mode based on a template cost according to an embodiment of the present invention.

[0353] The CCP coding efficiency may vary depending on the reference region used to derive the CCCM model. For example, the chrominance prediction blocks using CCCM models derived by using reference region A and reference region B of FIG. 35 (a), respectively, may differ from each other. The encoder can generate and signal a bitstream containing location information for the optimal reference region to derive the CCCM model used to predict the current chrominance block. The decoder can determine the reference region by parsing the location information and derive the CCCM model used to predict the current chrominance block using the determined reference region. At this time, the method for deriving the CCCM model can also be used to derive models in CCCM-LT, CCCM-L, and CCCM-T modes. CCCM-LT mode is a mode that derives the CCCM model using both the left reference region and the upper reference region of the current block, CCCM-L is a mode that derives the CCCM model using the left reference region of the current block, and CCCM-T is a mode that derives the CCCM model using the upper reference region of the current block.

[0354] To reduce the amount of bits being signaled, an optimal reference region can be derived based on a template. A video signal processing device can construct a reference template as shown in (c) of FIG. 35 using restored surrounding samples adjacent to the current chrominance block. The size of the reference template above the current block can be composed of the width and height M of the current block, and the size of the reference template to the left of the current block can be composed of the vertical length and width N of the current block. In this case, M and N can be integers greater than or equal to 1, and can be the same or different from each other; for example, they can be 1. In this case, the size of the reference template can be determined based on either the width and vertical dimensions of the current block or the number of samples in the current block. The video signal processing device can generate predicted samples for the reference template location using a CCCM model derived using various reference regions. After calculating the cost between the restored samples of the reference template and the predicted samples generated using the CCCM model derived from each reference region, the video signal processing device can determine the reference region with the minimum cost as the final reference region. Alternatively, the encoder may construct a list of reference regions using each reference region information, rearrange the list of reference regions based on each calculated cost, and then generate and signal a bitstream containing index information regarding which reference region to use. The decoder may parse the index information and then derive a CCCM model used to predict the current chrominance block using the reference region pointed to by the index information in the rearranged list of reference regions. The above template-based reference region determination method can be used to derive the optimal template form from three template forms (CCCM-LT, CCCM-L, and CCCM-T templates).

[0355] MM-CCCM (Multi-Model CCCM) is a method for predicting color difference blocks by deriving two CCCM models based on the average value of luminance samples calculated from a reference area. In this case, the average value may vary depending on which area it was calculated from, and the MM-CCCM model may vary depending on the average value. The encoder can calculate an average value to derive an optimal MM-CCCM model by using at least one of the reference area A, reference area B, the currently restored luminance block area, the reference area indicated by the block vector (BV) of the current block, and the reference area indicated by the motion vector (MV) of the current block in FIG. 35 (a). Additionally, the encoder can generate and signal a bitstream containing information about which area was used for the average value calculation (average value calculation area information). When the current color difference block is predicted using an MM-CCCM model, the decoder can parse the average value calculation area information and use it to derive the MM-CCCM used to predict the current color difference block based on the determined area.

[0356] To reduce the amount of bits being signaled, the region for calculating the optimal average value can be derived based on a template. A video signal processing device can construct a reference template as shown in FIG. 35 (c) using restored surrounding samples adjacent to the current chrominance block. The video signal processing device can obtain calculated average values ​​using reference region A or reference region B in FIG. 35 (a), the currently restored luminance block region in FIG. 35 (b), the reference region indicated by the BV of the current block, and the reference region indicated by the MV of the current block. The video signal processing device can generate prediction samples for the reference template location using an MM-CCCM model derived based on the obtained average values. After calculating the cost between the restored samples of the reference template and the prediction samples generated using each MM-CCCM model, the video signal processing device can determine the average value calculation region used to derive the MM-CCCM model having the minimum cost as the final average value calculation region. The encoder can construct a list of average calculation regions using region information used for average calculation, and then reorder the list of average calculation regions based on each calculated cost. The encoder can also generate and signal a bitstream containing indices for the used average calculation regions. After parsing the index information, the decoder can derive an MM-CCCM model used to predict the current color difference block using the average calculation regions pointed to by the index information in the reordered list of average calculation regions. This can also be applied to CCCM-LT, CCCM-L, and CCCM-T modes.

[0357] The video signal processing device can generate a final color difference prediction block through a weighted average between the block predicted using the intra-prediction directional mode and the block predicted using the CCP model, which can be referred to as the Chroma Fusion mode. In this case, the more effective model among CCLM or CCCM may be applied to the block predicted using the LM mode. The encoder can generate and signal a bitstream containing CCP model information used to predict the current color difference block. If the current color difference block is encoded in the Chroma Fusion mode, the decoder can generate a color difference prediction block based on the model determined by parsing the model information.

[0358] To reduce the amount of bits being signaled, an optimal CCP model can be derived based on a template. A video signal processing unit can construct a reference template as shown in Fig. 35(c) using reconstructed surrounding samples adjacent to the current chrominance block. The video signal processing unit can generate prediction samples for the reference template location using the CCP model. Then, the video signal processing unit can determine the CCP model with the minimum cost after calculating the cost between the reconstructed samples of the reference template and the prediction samples generated using the CCP model. Then, the video signal processing unit can predict the current chrominance block based on the determined CCP model. This can also be applied to CCCM-LT, CCCM-L, and CCCM-T modes.

[0359] The method of determining the CCP mode based on the template described in Fig. 35 can be called the Local-Boosting CCP (LB-CCP) model.

[0360] FIG. 36 shows peripheral blocks used to induce CCCM for a current color difference block according to one embodiment of the present invention.

[0361] FIG. 37 shows a reference area indicated by the block vector of the current luminance block according to one embodiment of the present invention.

[0362] The CCCM used to predict the current chrominance block can be derived from the surrounding blocks of the current block, which can be referred to as the CCP merge mode. FIG. 29 shows surrounding blocks A0, A1, B0, B1, B2 adjacent to the current block and surrounding blocks A0', A1', B0', B1', B2' not adjacent to the current block. The video signal processing device may use the CCCM model used in the surrounding blocks of FIG. 36 to predict the current chrominance block. Alternatively, if the current block is encoded in Intra TMP or IBC mode, the video signal processing device may generate a prediction block for the current chrominance block using the CCCM model used in the reference region indicated by the BV of the current luminance block and the currently restored luminance block, as shown in FIG. 37.

[0363] The types of CCP models may include CCLM, MMLM, GLM, MM-GLM, CCCM, MM-CCCM, GL-CCCM, MM-GL-CCCM, CCCM-ND, MM-CCCM-ND, CCCM-MDF, MM-CCCM-MDF, Chroma Fusion, BVG-CCCM, CCP Merge, etc. MMLM can be composed of two CCLMs. If the types of CCP models differ between CCP models, they can be considered different CCP models. If the types of CCP models are the same but the parameters differ between CCP models, they can be considered different CCP models. If the number of CCP models differs between CCP models (for example, if the first CCP model consists of one CCLM and the second CCP model consists of two CCLMs), they can be considered different CCP models.

[0364] FIG. 38 is a structural diagram showing a method for predicting the current block using a CCP merge mode according to one embodiment of the invention.

[0365] FIG. 39 shows the surrounding blocks of the current block according to one embodiment of the present invention.

[0366] FIG. 40 shows a co-located block and surrounding blocks according to one embodiment of the present invention.

[0367] First, the video signal processing device can configure a CCP merge list for the current block. At this time, the CCP merge list may include a maximum number of pre-set candidates. For example, the maximum number may be 12. Among surrounding blocks spatially adjacent to the current block (surrounding blocks A0, A1, B0, B1, B2 in FIG. 39), surrounding blocks that are not spatially adjacent and are far apart (surrounding blocks 1A0, 1A1, 1B0, 1B1, 1B2, …, NA0, NA1, NB0, NB1, NB2 in FIG. 39), and temporal surrounding blocks corresponding to the current block from other reference pictures that are not identical to the POC of the current picture, a CCP model used in at least one of these surrounding blocks may be included as a CCP model candidate in the CCP merge list. The location of a spatially non-adjacent and distant neighboring block can be determined based on at least one of the following: the width of the current block, the height of the current block, the ratio of the width to the height of the current block, the quantization parameters of the current block and the neighboring block, the encoding mode of the current block, information on whether the current block is encoded in Intra TMP or IBC mode, the block vector of the current block, and predefined locations. For example, the vertical location of neighboring blocks 1B0, 1B1, and 1B2 in FIG. 39 can be a location that is different from the top boundary of the current block by the height of the current block. The horizontal location of neighboring blocks 1A0, 1A1, and 1B2 in FIG. 39 can be a location that is different from the left boundary of the current block by the width of the current block. Neighboring block 1A1 can exist at a predefined location between 1A0 and 1B2, and neighboring block 1B1 can exist at a predefined location between 1B2 and 1B0. In addition, the vertical positions of the surrounding blocks NB0, NB1, and NB2 in FIG. 39 may be positions that are offset by a predefined distance from the top boundary of the current block.The horizontal positions of the surrounding blocks NA0, NA1, and NB2 in FIG. 39 may be positions differed by a predefined distance from the left boundary of the current block. The surrounding block NA1 may exist at a predefined position between NA0 and NB2, and the surrounding block NB1 may exist at a predefined position between NB2 and NB0. The predefined distance may be an integer greater than or equal to 1. Additionally, CCP models derived from surrounding blocks that are not spatially adjacent but are far apart may be added to the CCP merge list as CCP model candidates. In this case, the order of addition may be determined based on at least one of the quantization parameters of the current block and the distance between the current block and the surrounding blocks. For example, surrounding blocks farther from the current block may be added in order from closest surrounding blocks, or surrounding blocks closest to the current block may be added in order from farther surrounding blocks. The surrounding blocks may be considered reference blocks for deriving a CCP model for the current block.

[0368] If the CCP merge list is not filled, the video signal processing device may add CCP models used in temporal surrounding blocks (neighboring blocks corresponding to the current block from other reference pictures that are not identical to the POC of the current picture, and neighboring blocks of blocks co-located with the current block) to the CCP merge list as CCP model candidates. When deriving CCP model candidates for temporal surrounding blocks, the video signal processing device may derive CCP model candidates from various predefined locations based on the width and height of the current block. Additionally, the video signal processing device may derive CCP model candidates for temporal surrounding blocks as many times as a predefined number. The predefined number may be 4. For example, referring to FIG. 40, the video signal processing device may derive CCP model candidates from the BR0, Ctr0, BR1, and Ctr1 locations based on the width and height of the current block. If a predefined number of CCP model candidates has not been filled, the video signal processing unit may derive CCP model candidates from the BR0, Ctr0, BR1, and Ctr1 locations based on twice the width and height of the current block. This process may be repeated up to N times the width and height of the current block until a predefined number of CCP model candidates are filled. In this case, N may be 5.

[0369] The video signal processing device may use at least one of the encoding mode information, motion vector information, and block vector information of the temporal surrounding block to determine the availability of a CCP model candidate for the temporal surrounding block or to decide whether to add a CCP model candidate for the temporal surrounding block to the CCP merge list. For example, if the temporal surrounding block is of the RRIBC type, the video signal processing device may not use it as a CCP model candidate and may not add it to the CCP merge list.

[0370] If the CCP merge list is not filled, the video signal processing unit may add CCP models derived from a CCP model table stored in separate memory as CCP model candidates to the CCP merge list. In this case, the video signal processing unit may extract CCP model candidates in the order from the CCP model added first to the CCP model table to the CCP model added last and add them to the CCP merge list. Alternatively, the video signal processing unit may extract CCP model candidates in the order from the CCP model added last to the CCP model table to the CCP model added first and add them to the CCP merge list. CCP model candidates derived from a CCP model table stored in separate memory can be referred to as History-based CCP (HCCP) candidates. The CCP model table can be constructed from previously used CCP models. Additionally, the video signal processing device may determine whether to add a previously used CCP model to the CCP model table by using at least one of whether the block is encoded in dual tree mode, whether the block is encoded in ISP mode, whether the block is encoded in intra mode, whether the block is encoded in DIMD chroma, whether the block is encoded in DM mode, whether the CCP model already exists in the CCP model table, or whether the block is encoded in chroma fusion mode. The CCP model table may be reset when the encoding or decoding of each slice begins, or when the encoding or decoding of the first CTU in each CTU row begins, and if reset, the CCP model table may be emptied. A First In First Out (FIFO) method may be applied in which CCP models are removed in the order they were added to the CCP model table.When adding a new CCP model to the CCP model table, the video signal processing device may check whether the new CCP model exists in the CCP model table, and if the new CCP model does not exist in the CCP model table, the new CCP model may be added to the CCP model table. Additionally, when adding a new CCP model to the CCP model table, if the CCP model table is full, the first added CCP model may be deleted and the new CCP model may be added to the CCP model table.

[0371] If the CCP merge list is not filled, the video signal processing device may generate a combined CCP model candidate or a separated CCP model candidate using CCP models within the CCP merge list (e.g., CCLM, MMLM, GLM, CCCM, MM-CCCM, GL-CCCM, CCCM-ND, CCCM-MDF) and add them to the CCP merge list. In this case, the combined CCP model candidate may be a single CCP model candidate obtained by using two CCP model candidates within the CCP merge list. A combined CCP model based on CCLM may be CCLM + GLM, CCLM + CCCM, CCLM + GL-CCCM, CCLM + CCCM-ND, CCLM + CCCM-MDF, etc. Additionally, a combined CCP model based on GLM may be GLM + CCCM, GLM + GL-CCCM, GLM + CCCM-ND, GLM + CCCM-MDF, etc. Additionally, CCP models combined based on CCCM can be CCCM + GL-CCCM, CCCM + CCCM-ND, CCCM + CCCM-MDF, etc. Additionally, CCP models combined based on GL-CCCM can be GL-CCCM + CCCM-ND, GL-CCCM + CCCM-MDF, etc. Additionally, CCP models combined based on CCCM-ND can be CCCM-ND + CCCM-MDF, etc. For example, to predict the current chrominance block using CCCM-ND + CCCM-MDF candidates, the video signal processing device may divide the current luminance block into two regions based on a predefined threshold, predict current chrominance samples using the CCCM-ND mode in one region, and predict current chrominance samples using the CCCM-MDF mode in the other region, and then generate the final chrominance block. The predefined threshold may be the average value of surrounding luminance samples adjacent to the current luminance block or the average value of the current luminance block.A separated CCP model candidate may be two CCP model candidates separated from a single CCP model candidate within the CCP merge list. Alternatively, a separated CCP model candidate may be two newly created CCP model candidates based on each of the two separated CCP model candidates within the CCP merge list. For example, if a single CCP model candidate within the CCP merge list is a CCP model that combines two CCP models, such as MMLM or MM-CCCM, the combined CCP model may be separated into two CCLMs or two CCCMs, and a new CCP model candidate created based on the two separated CCP models may be a separated CCP model candidate and may be added to the CCP merge list.

[0372] If the CCP merge list is not filled, the video signal processing device checks for the presence of CCLM (or MMLM) within the CCP merge list and, depending on the presence of CCLM (or MMLM), can add a new CCP model candidate to the CCP merge list. For example, if CCLM (or MMLM) exists within the CCP merge list, the video signal processing device can regenerate a new CCLM (or MMLM) based on the CCLM (or MMLM) within the CCP merge list and add it to the CCP merge list. In this case, the video signal processing device may regenerate a new CCLM (or MMLM) using only a single CCLM (or MMLM) within the CCP merge list and add it to the CCP merge list. Alternatively, the video signal processing device may regenerate a new CCLM or a new MMLM using all CCLMs and MMLMs within the CCP merge list and add it to the CCP merge list. Specifically, the video signal processing device can add two CCP model candidates obtained by separating an MMLM in the CCP merge list into two CCLMs to the CCP merge list. Additionally, the video signal processing device can add two CCLMs in the CCP merge list to the CCP merge list by configuring them into a single MMLM. Alternatively, the video signal processing device can add one CCLM to the CCP merge list by configuring it through the weighted average of the parameter values ​​of two CCLMs in the CCP merge list. Or, the video signal processing device can add two new CCLMs obtained using each CCLM after separating an MMLM in the CCP merge list into two CCLMs as CCP model candidates.Additionally, the video signal processing device can acquire one MMLM through two new CCLMs obtained using each of the two CCLMs in the CCP merge list and add it to the CCP merge list. For example, if one CCP model in the CCP merge list is a CCP model that combines two CCP models, such as MMLM or MM-CCCM, the video signal processing device can separate the combined CCP model into two CCLMs or two CCCMs, and then create a new CCP model based on the separated two CCLMs or two CCCMs and add it to the CCP merge list.

[0373] A video signal processing device can generate a new CCP model based on CCP model candidates within a CCP merge list, wherein the CCP model to be generated for the new CCP model may be a specific CCP model. Therefore, the video signal processing device can determine whether a CCP model candidate within the CCP merge list is a specific CCP model. In this case, the specific CCP model may be CCLM, MMLM, etc.

[0374] FIG. 41 shows a CCLM illustrated in the form of a graph according to one embodiment of the present invention.

[0375] CCLM may be a CCP model that derives a color difference sample by multiplying a luminance sample by a slope value (a of (a) in Fig. 41) and then adding an offset value (b of (a) in Fig. 41).

[0376] The video signal processing device can obtain a new CCP model by adding a predefined value (u in (b) of Fig. 41) to the slope value of CCLM or MMLM in the CCP merge list.

[0377] The slope value (a') and offset value (b') for the new CCP model can be calculated using Equation 9.

[0378] Mathematical formula 9

[0379] a' = a + u

[0380] b' = b - u * y r

[0381]

[0382] In this specification, u in Equation 9 may be expressed as a predefined first value. Additionally, u may be an integer value, for example, -6, -5, …, 5, 6. y in Equation 9 r is an average value, which can be the average value of the current luminance block, the average value of the surrounding luminance blocks, or half the value of the bit depth used to represent the current luminance sample (e.g., 512 if the content is 10 bits).

[0383] The CCP model may include a Cb CCP model and a Cr CCP model. The video signal processing device may create a new Cb CCP model by adding a predefined first value to the slope value of the Cb CCP model. Additionally, the video signal processing device may create a new Cr CCP model by adding a predefined first value to the slope value of the Cr CCP model. Furthermore, since the new CCP model is identical to the existing CCP model when the u value is 0, the video signal processing device may exclude cases where the u value is 0 when creating a new CCP model based on CCLM or MMLM within the CCP merge list. That is, cases where the u value is 0 may not be added to the CCP merge list. Alternatively, the video signal processing device may not add the new CCP model to the CCP merge list if a' of the new CCP model is 0.

[0384] The video signal processing device may apply a new CCP model generated based on CCLM or MMLM within the CCP merge list to only one of the Cb or Cr chrominance components. For example, the new CCP model may include a Cb CCP model and a Cr CCP model. The video signal processing device may generate a new Cb CCP model by adding a predefined first value to the slope value of the Cb CCP model without modifying the Cr CCP model. Alternatively, the video signal processing device may generate a new Cr CCP model by adding a predefined second value to the slope value of the Cr CCP model without modifying the Cb CCP model. Here, the predefined first value and the second value may be the same or different from each other. For example, the first value and the second value may be -6, -5, ... 5, 6, respectively.

[0385] If no CCLM or MMLM exists in the CCP merge list, the video signal processing device may add a CCP model candidate using a predefined CCLM or MMLM to the CCP merge list. The slope value of the predefined CCLM can be an integer, for example, -6, -5, …, 5, 6. The offset value of the predefined CCLM can be half the value of the bit depth used to represent the current luminance sample (for example, 512 for 10-bit content).

[0386] When a new CCP model candidate is added to the CCP merge list, the video signal processing device may check whether the new CCP model candidate already exists in the CCP merge list. If the new CCP model candidate already exists in the CCP merge list, the new CCP model candidate may not be added to the CCP merge list. If the new CCP model candidate does not exist in the CCP merge list, the new CCP model candidate may be added to the CCP merge list. At this time, the video signal processing device may check whether the new CCP model candidate already exists in the CCP merge list by using at least one of the following: the type of the CCP model, the top-left position information of the block that derived the CCP model, the parameter information of the CCP model, the luminance offset value information of the CCP model, the luminance block average value information of the CCP model, information on whether the type of the new CCP model candidate already exists in the CCP merge list, and information on whether MMLM exists in the CCP merge list. The video signal processing device can check the identity of CCP models by comparing whether the parameters of a new CCP model candidate are identical to the parameters of a CCP model already existing in the CCP merge list. For example, if the parameters are identical, the video signal processing device can determine that the CCP models are identical. For instance, a method by which the video signal processing device checks whether a new CCP model candidate already exists in the CCP merge list may involve checking the identity of the CCP model candidate by comparing all parameters of the new CCP model with all parameters of all CCP candidates within the CCP merge list.Alternatively, a method for a video signal processing device to check whether a new CCP model candidate already exists in a CCP merge list may check the identity of the CCP model candidate by comparing a predefined number of parameters among all parameters of the new CCP model with a predefined number of parameters among all parameters of all CCP candidates in the CCP merge list. In this case, the predefined number may be an integer greater than or equal to 1. Alternatively, a method for a video signal processing device to check whether a new CCP model candidate already exists in a CCP merge list may check the identity of the CCP model candidate by comparing the first three parameters among all parameters of the new CCP model with the first three parameters among all parameters of all CCP candidates in the CCP merge list. Alternatively, a method for a video signal processing device to check whether a new CCP model candidate already exists in a CCP merge list may check the identity of the CCP model candidate by determining whether the top-left position information of the block that derived the new CCP model is the same or similar to the top-left position information of the block that derived each CCP candidate in the CCP merge list. The video signal processing device may determine that they are similar if the absolute value of the difference between the top-left position information of the block that derived the new CCP model and the top-left position information of the block that derived each CCP candidate in the CCP merge list is smaller than a predefined value. The predefined value may be an integer greater than or equal to 1.

[0387] Among the CCP model candidates, the top-left position information of the block that derived the HCCP (History-based CCP) and the predefined CCP model may not exist. That is, when a new CCP model candidate composed of the HCCP and the predefined CCP model checks for identity with the CCP model candidate in the CCP merge list, the top-left position information of the block cannot be used. Therefore, the video signal processing device may not perform a check on whether the top-left position information of the block is the same for the HCCP and the predefined CCP model candidates. Specifically, when a prediction block of a chrominance block is generated using the HCCP and the predefined CCP model, the video signal processing device may store the top-left position information of the block as (0, 0) or (-1, -1) when storing the CCP model for the current chrominance block. When checking whether a new CCP model candidate exists in the CCP merge list, the video signal processing device may not perform the check process using the top-left position information if the top-left candidate position value of the block for the new CCP model candidate is (0, 0) or (-1, -1).

[0388] Alternatively, when generating a prediction block using an HCCP and a predefined CCP model in the encoder and decoder, when saving the CCP model for the current chrominance block, the top-left position information of the current block can be saved in the top-left position information of the block.

[0389] When a video signal processing device adds a new CCP model to the CCP merge list, it may add only one type of CCP model to the CCP merge list for each type of CCP model. That is, each type of CCP model in the CCP merge list may not be identical, except for HCCP models or predefined CCP models.

[0390] FIG. 42 shows a reference template position used to rearrange a CCP merge list according to one embodiment of the present invention.

[0391] The video signal processing device can configure a CCP merge list and rearrange the CCP merge list based on the template cost. Whether the rearrangement process is performed can be determined based on the width and / or height of the current block.

[0392] Additionally, the template for rearranging the CCP merge list can be determined based on the width and / or height of the current block. For example, if the width of the current block is greater than the height, the top template adjacent to the top of the current block can be used for rearranging the CCP merge list. If the width and height of the current block are equal, the templates of FIG. 42 (the top template adjacent to the top of the current block and the left template adjacent to the left of the current block) can be used for rearranging the CCP merge list.

[0393] The video signal processing device can construct a reference template using restored surrounding samples adjacent to the current block as shown in FIG. 42. The video signal processing device can generate color difference prediction samples by predicting color difference samples of the reference template using luminance samples of each CCP model in the CCP merge list and the reference template. The video signal processing device can generate a cost by calculating the SAD (or MR-SAD) between the color difference prediction samples and the color difference samples of the reference template. The video signal processing device can calculate the cost for each CCP model and rearrange the CCP merge list in ascending order of cost.

[0394] The encoder can select the optimal CCP model for the current chrominance block based on the reordered CCP merge list and generate and signal a bitstream containing index information indicating the optimal CCP model. The decoder can parse the index information and generate a prediction block for the current chrominance block using the CCP model indicated by the index information within the reordered CCP merge list.

[0395] When the optimal CCP model selected from the reordered CCP merge list is used to predict the current chrominance block, the video signal processing unit may derive a compensation value for some CCP models and then apply the compensation value to the chrominance block predicted using the CCP models. Some CCP models may be any one of CCLM, MMLM, GLM, CCCM, MM-CCCM, GL-CCCM, CCCM-ND, or CCCM-MDF. For example, if the optimal CCP model selected from the reordered CCP merge list is CCCM, the video signal processing unit may construct a reference template using restored surrounding samples adjacent to the current block. The video signal processing unit may apply CCCM to the luminance samples of the reference template to predict the chrominance samples of the reference template and generate chrominance prediction samples. The video signal processing unit may calculate the average value between the chrominance prediction samples and the chrominance samples of the reference template and use the average value as a compensation value. The compensation value may be at least one of the compensation value for the Cb chrominance block and the compensation value for the Cr chrominance block. Additionally, if the optimal CCP model is a two-model such as MMLM or MM-CCCM, two compensation values ​​for the Cb chrominance block and two compensation values ​​for the Cr chrominance block may be used as compensation values. If the optimal CCP model is CCCM, the video signal processing device may apply CCCM to the currently restored luminance block to predict the current chrominance block, and then add the derived compensation values ​​to the chrominance prediction block to generate the final chrominance prediction block. The video signal processing device may determine whether a compensation value is derived by using at least one of the width and height of the current block, the encoding mode of the current luminance block, and the type of CCP model. For example, if the width and height of the current block are smaller than a predefined number, a compensation value may not be derived.Or, if the width and height of the current block are greater than a predefined number, the reward value may not be derived. The predefined number may be an integer greater than or equal to 1, or the product of the width and height of the current block. For example, the predefined number may be 16.

[0396] The CCP model for the current block used to predict the chrominance block in the current block can be stored for encoding and decoding the next block, and can be stored in a CCP model table to be used as an HCCP candidate. If the video signal processing device has derived a CCP model for predicting the chrominance block in the current block through a CCP merge method, it can store the CCP model used to predict the current chrominance block. The video signal processing device can use the stored CCP model information to construct the CCP merge list for the next block.

[0397] When a CCP merge list for a current chrominance block is configured, the video signal processing device may determine whether to add a predefined CCP model to the CCP merge list by using at least one of the encoding mode information of the current luminance block, motion vector information of the current luminance block, block vector information of the current luminance block, transform coefficient information of the current luminance block, and quantization parameter information of the current luminance block. For example, when the encoding mode of the current luminance block is GPM mode, the video signal processing device may not add CCP candidates that use only one CCP model (e.g., CCLM, CCCM, etc.) to the CCP merge list when configuring the CCP merge list, but may add CCP candidates that use two CCP models (e.g., MMLM, MM-CCCM, etc.) to the CCP merge list.

[0398] When a video signal processing unit stores CCP model parameters for the current block in memory, it may store them with a reduced scale of the values ​​to decrease memory usage. When using the stored CCP model parameters, the video signal processing unit may increase the scale of the values ​​and then use them to predict the chrominance block. The descalement and increase processes can be handled by quantization and inverse quantization methods. Alternatively, the descalement and increase processes can be handled through shift operations using predefined integer values, where the predefined integers can be integers greater than or equal to 1. For example, if the predefined integer value is 4, the descalement operation can be 'P >> 4' and the scale increase operation can be 'P << 4'. P can be any one of the CCP model parameters. The shift operation A >> B means dividing A by 2 B times, and "A < <B"는 A를 B번 2로 곱하는 것을 의미한다.

[0399] The CCP merge list can be configured integrally for two color difference signals. That is, any one CCP model candidate within the CCP merge list may include a CCP model for the Cb color difference block and a CCP model for the Cr color difference block. The encoder can generate and signal a bitstream containing information on whether to apply the CCP merge method to the current color difference block and index information indicating the optimal CCP model for the current color difference block among the CCP models within the CCP merge list. The decoder can parse the information on whether to apply the CCP merge method to the current color difference block, and if the CCP method is applied to the current color difference block, it can parse the index information and generate a Cb prediction block and a Cr prediction block for the current color difference block using the CCP model indicated by the index information within the CCP merge list.

[0400] Additionally, the CCP merge list can be applied independently to each chrominance signal. That is, the CCP merge list for the Cb chrominance block (Cb CCP merge list) and the CCP merge list for the Cr chrominance block (Cr CCP merge list) can be generated independently, and the Cb CCP merge list and the Cr CCP merge list can be different from each other. The encoder can construct a merge list for each chrominance signal and generate and signal a bitstream containing information on whether the CCP merge method is currently applied to the Cb chrominance block, index information indicating the optimal CCP model index for the Cb chrominance block, information on whether the CCP merge method is currently applied to the Cr chrominance block, and index information indicating the optimal CCP model index for the Cr chrominance block. The decoder parses information regarding whether the CCP merge method is currently applied to the Cb color difference block; if the CCP merge method is currently applied to the Cb color difference block, it can parse index information representing the optimal CCP model index for the Cb color difference block, and generate a Cb color difference prediction block based on the CCP model represented by the index information. Additionally, the decoder parses information regarding whether the CCP merge method is applied to the Cr color difference block; if the CCP merge method is currently applied to the Cr color difference block, it can parse index information representing the optimal CCP model index for the Cr color difference block, and generate a Cr color difference prediction block based on the CCP model represented by the index information.

[0401] The information for each CCP model in the above CCP merge list may include at least one of the following: the type of the CCP model, parameter coefficients for the CCP model, coordinate information of the block that derived the CCP model, average value, luminance offset value, encoding mode information, and quantization parameter.

[0402] A video signal processing device can generate a CCP merge list based on CCP model information used in the surrounding blocks of the current block. This is because the CCP models between the surrounding blocks and the current block have a correlation. If the correlation between the current block and the surrounding blocks is low, the encoding efficiency of the method of deriving the CCP model through CCP merge may be low. To prevent this, information regarding which CCP model was used can be signaled on a block-by-block basis. However, if there are too many types of selectable CCP models, the amount of bits that need to be signaled may increase excessively. Therefore, the encoder can construct an explicit CCP list. In this case, the encoder can include the index of the CCP list corresponding to the CCP model used in the current block in the bitstream.

[0403] Since the CCP model within the CCP merge list is the same as the CCP model used in neighboring blocks, it may not be suitable for predicting the current chrominance block. Therefore, the video signal processing unit can predict the current chrominance block by re-deriving CCP model parameters through neighboring blocks adjacent to the current chrominance block. Specifically, the video signal processing unit can derive new parameters for the CCP model using neighboring blocks (or samples) adjacent to the current block, utilizing only the type of the CCP model from the optimal CCP model derived from the CCP merge list (or an explicit CCP list). The video signal processing unit can predict the current chrominance block using the parameters for the newly derived CCP model. The encoder may include CCP model selection information in the bitstream that indicates whether to use the newly derived CCP model or the CCP model derived from the CCP merge list. The decoder can predict the current chrominance block using the CCP model indicated by the CCP model selection information.

[0404] The video signal processing device may rearrange the CCP merge list using at least one of the CCP model information of a neighboring block adjacent to the current block, the encoding mode of the restored luminance block, the encoding mode of the restored luminance block, or the width and height of the current block. In a specific embodiment, if the type of the CCP model of a neighboring block adjacent to the current block is CCCM, the video signal processing device may change the order of the CCP merge list so that a CCP model of type CCCM is placed at the beginning of the list. In another specific embodiment, if the encoding mode of the restored current luminance block is a divided encoding mode such as GPM or SGPM, the video signal processing device may change the order of the CCP merge list so that a CCP model of type GLM or GL-CCCM is placed at the beginning of the list. In another specific embodiment, if the encoding mode of the restored current luminance block is a split encoding mode, such as GPM or SGPM, the video signal processing device may change the order of the CCP merge list so that a CCP model using two CCP models is placed at the beginning of the list. A CCP model using two CCP models may include at least one of MMLM, MM-CCCM, MM-GLM, MM-GL-CCCM, or MM-CCCM-ND.

[0405] Alternatively, the video signal processing device may change the order so that the type of CCP model with generally superior encoding efficiency among the types (or kinds) of CCP models is placed at the top of the list.

[0406] Since there are various types of CCP models, the amount of bits required to signal information regarding the CCP model used for the current chrominance block may be large. To reduce this amount of bits, the video signal processing device may construct an explicit CCP list and then select a CCP model for the current chrominance block from within the explicit CCP list. Specifically, after constructing the explicit CCP list, the encoder may generate and signal a bitstream containing index information for the optimal CCP model. The decoder may parse the index information and generate a prediction block for the current chrominance block using the CCP model indicated by the index information within the explicit CCP list. The video signal processing device may rearrange the explicit CCP merge list using at least one of the CCP model information of neighboring blocks adjacent to the current block, the encoding mode of the restored luminance block, the encoding mode of the restored luminance block, and the width and height of the current block. For example, if the type (or kind) of the CCP model of a neighboring block adjacent to the current block is CCCM, the video signal processing device may change the order of the CCCM type CCP model in the explicit CCP merge list to place it at the beginning of the list. Alternatively, if the encoding mode of the restored current luminance block is a partitioned encoding mode such as GPM or SGPM, the video signal processing device may change the order of the GLM or GL-CCCM type CCP models in the explicit CCP merge list to place them at the beginning of the list.

[0407] The explicit CCP list may include CCLM, MMLM, GLM, MM-GLM, CCCM, MM-CCCM, GL-CCCM, GL-MM-CCCM, CCCM-ND, MM-CCCM-ND, CCCM-MDF, MM-CCCM-MDF, Chroma Fusion, LB-CCP, BVG-CCCM, and CCP Merging.

[0408] In the encoder and decoder, the optimal CCP model for the current color difference block can be determined through the above method by using a pre-specified explicit CCP list.

[0409] Specifically, an explicit CCP list may be configured in the order of CCLM, MMLM, GLM, MM-GLM, CCCM, MM-CCCM, GL-CCCM, GL-MM-CCCM, CCCM-ND, MM-CCCM-ND, CCCM-MDF, MM-CCCM-MDF, Chroma Fusion, LB-CCP, BVG-CCCM, MM-BVG-CCCM, and CCP Merge. Alternatively, an explicit CCP list may be configured in the order of CCP Merge, CCLM, MMLM, GLM, MM-GLM, CCCM, MM-CCCM, GL-CCCM, GL-MM-CCCM, CCCM-ND, MM-CCCM-ND, CCCM-MDF, MM-CCCM-MDF, Chroma Fusion, LB-CCP, BVG-CCCM, and MM-BVG-CCCM. An explicit CCP list can be constructed using some or all of the aforementioned explicit CCP models, and the order of each model can be configured in various forms.

[0410] The characteristics of the image may differ from block to block, and the degree of change between samples may be gradual, complex patterns may exist within the block, or sharp edge components may exist between samples. Therefore, the video signal processing device may select a different pre-specified explicit CCP list depending on the characteristics of the block. Specifically, if sharp edge components exist between samples within the block, the video signal processing device may use an explicit CCP list composed of CCP models having gradient characteristics. An explicit CCP list composed of CCP models having gradient characteristics may consist of GLM, MM-GLM, GL-CCCM, GL-MM-CCCM, CCCM-MDF, and MM-CCCM-MDF. The characteristics of the block may be the CCP model information of surrounding blocks, the encoding mode of the reconstructed surrounding luminance block, the encoding mode of the reconstructed current luminance block, and the width and height of the current block. The explicit CCP list may be reordered based on the template cost. The encoder can generate and signal a bitstream containing type information of the explicit CCP list regarding which characteristic explicit CCP list has been applied to the current block. Additionally, the encoder can generate and signal a bitstream containing index information for the optimal CCP model within the explicit CCP list. After parsing the type information of the explicit CCP list, the decoder can determine the explicit CCP list used to predict the current chrominance block from among multiple explicit CCP lists based on the type information. Furthermore, the decoder can determine the optimal CCP model for predicting the current chrominance block within the determined explicit CCP list by parsing the index information for the optimal CCP model.

[0411] A video signal processing device can derive type information of an explicit CCP list by utilizing at least one of the following: CCP model information of a neighboring block adjacent to the current block, the encoding mode of the restored luminance block, the encoding mode of the neighboring block, the width and height of the current block, the quantization parameters of the current block and the neighboring block, the encoding mode of the current block or the neighboring block, and information regarding whether the current block was encoded in an intra-TMP or IBC mode. In this case, the encoder may use the derived type information without signaling the type information of the explicit CCP list regarding which characteristics of the explicit CCP list were applied to the current block. Furthermore, the encoder may generate and signal a bitstream containing index information for the optimal CCP model within the explicit CCP list configured according to the derived type information. The decoder can parse the index information for the optimal CCP model, construct an explicit CCP list using the derived type information, and then determine the optimal CCP model for predicting the current chrominance block in the explicit CCP list using the index information.

[0412] When configuring an explicit CCP list, the video signal processing device may select a different pre-specified explicit CCP list depending on whether one CCP model or two or more CCP models is used. Specifically, when one CCP model is used, the video signal processing device may use an explicit CCP list composed of one CCP model. An explicit CCP list composed of one CCP model may consist of CCLM, GLM, CCCM, GL-CCCM, CCCM-ND, CCCM-MDF, BVG-CCCM, and CCP Merge. An explicit CCP list composed of two or more CCP models may consist of MMLM, MM-GLM, MM-CCCM, GL-MM-CCCM, MM-CCCM-ND, MM-CCCM-MDF, Chroma Fusion, LB-CCP, MM-BVG-CCCM, and CCP Merge. The configured explicit CCP list may be reordered based on template costs. The encoder can generate and signal a bitstream containing type information of the explicit CCP list regarding which explicit CCP list has been applied to the current block. Additionally, the encoder can generate and signal a bitstream containing index information for the optimal CCP model within the explicit CCP list. After parsing the type information of the explicit CCP list, the decoder can determine the explicit CCP list used to predict the current chrominance block from among multiple explicit CCP lists based on the type information. Furthermore, the decoder can parse the index information for the optimal CCP model to determine the optimal CCP model for predicting the current chrominance block from the determined explicit CCP list.

[0413] When constructing an explicit CCP list, the video signal processing device may construct an explicit combined CCP list such that two or more CCP models become a single candidate. The explicit combined CCP list may include a candidate consisting of only one CCP model. Specifically, the explicit combined CCP list may be composed of {(CCLM, CCCM), (GLM, CCCM), (GL-CCCM, MM-CCCM-ND), (CCCM-ND, MM-CCCM-MDF), (BVG-CCCM, 0), (CCP merge, 0). The constructed explicit combined CCP list may be reordered based on the template cost. The encoder may generate and signal a bitstream containing index information for the optimal CCP model candidate in the explicit combined CCP list. The decoder may parse the index information for the optimal CCP model candidate to determine the optimal CCP model candidate for predicting the current chrominance block in the explicit combined CCP list. If the optimal CCP model candidate is (CCCM-ND, MM-CCCM-MDF), the video signal processing device can generate a final chrominance block by weighted averaging the chrominance block predicted using CCCM-ND and the chrominance block predicted using MM-CCCM-MDF. If the optimal CCP model candidate is (BVG-CCCM, 0), the video signal processing device can generate a chrominance prediction block using only BVG-CCCM.

[0414] A video signal processing device can determine whether to predict a chrominance block by constructing an explicit CCP list using at least one of the following: CCP model information of a neighboring block adjacent to the current block, the encoding mode of the restored luminance block, the encoding mode of the neighboring block, the width and height of the current block, the quantization parameters of the current block and the neighboring block, the encoding mode of the current block or the neighboring block, and information on whether the current block was encoded in an intra-TMP or IBC mode.

[0415] The types of acceptable CCP models may vary depending on the width and height of the current block. For example, if the width and height of the current block are smaller than a predefined number, predefined CCP models among CCLM, MMLM, GLM, CCCM, MM-CCCM, GL-CCCM, CCCM-ND, CCCM-MDF, etc., may not be allowed. The predefined number may be an integer greater than or equal to 1, or the product of the width and height. For example, the predefined number may be 8, 16, or 32. For example, if the product of the width and height of the current block is 16 or less, GL-CCCM and CCCM-MDF, which require many samples from surrounding blocks when predicting the current color difference block, may not be allowed. If the product of the width and height of the current block is 16 or less, the encoder may not include information related to GL-CCCM and CCCM-MDF in the bitstream and may not signal them. If the product of the width and height of the current block is 16 or less, the decoder may not parse information related to GL-CCCM and CCCM-MDF, and may be configured so that GL-CCCM and CCCM-MDF are not applied.

[0416] When the encoding mode of the restored current luminance block is a divided encoding mode such as GPM or SGPM, the video signal processing device may configure only CCP models of the MMLM, MM-GLM, MM-CCCM, GL-MM-CCCM, MM-CCCM-ND, MM-CCCM-MDF, and MM Chroma Fusion types using two CCP models to be used for predicting the current chroma difference block. The encoder may generate and signal a bitstream containing information on the number of CCP models (whether there is one or two), information on the template form of the CCP models (L, T, LT), and information on the type of the CCP models (CCLM, GLM, CCCM, GL-CCCM, CCCM-ND, CCCM-MDF, Chroma Fusion). The decoder may first parse information on the number of CCP models and information on the template form of the CCP models, and then parse information on the type of the CCP models. Alternatively, the decoder may first parse the type information of the CP model, and then parse the information regarding the number of CCP models and the information regarding the template form of the CCP model. If the encoding mode of the recovered current luminance block is a split encoding mode such as GPM or SGPM, the encoder may disable one CCP model and may not include information related to one CCP model in the bitstream and may not signal it. Specifically, if the encoding mode of the recovered current luminance block is a split encoding mode such as GPM or SGPM, the encoder may not include information regarding the number of CCP models in the bitstream and may include only information regarding the template form of the CCP model and the type of the CCP model in the bitstream and signal it.When the encoding mode of the recovered current luminance block is a divided encoding mode such as GPM or SGPM, the decoder can set the number of CCP models to 2 without parsing information about the number of CCP models, and can determine CCP model information for the current chrominance block by parsing information about the template form of the CCP models and information about the type of the CCP models.

[0417]

[0418] LMCS (Luma Mapping with Chroma Scaling) is a preprocessing step in which a video signal processing unit dynamically changes the representation range of an input video signal to improve encoding performance and subjective image quality. In the LMCS operation, the video signal processing unit dynamically changes the representation range of pixel values ​​and can perform luminance component mapping and chrominance component scaling. Luminance component mapping refers to the redistribution of luminance samples to increase encoding efficiency, while chrominance component scaling refers to compensating for the difference between the luminance and chrominance components that has changed due to redistribution. Through this, the video signal processing unit can improve encoding efficiency. Specifically, when the encoder encodes the current video, the encoder can convert the input video into a dynamic range through forward mapping and perform inverse mapping to convert the restored video back to its original representation range. The encoder can perform forward mapping by dividing the existing dynamic region into 16 identical segments and redistributing the codewords of the input video using a linear model for each segment. The encoder can perform reverse mapping, which performs inverse mapping from a mapped dynamic region to an existing dynamic region. The encoder can generate a bitstream containing parameters related to forward mapping and reverse mapping.

[0419] Mathematical Equation 10 is an equation that calculates a linear model of forward mapping and maps the codeword of the input luminance component to the codeword of the transmitted dynamic area.

[0420] Mathematical formula 10

[0421] FwdMap(Y pred ) = b1 + ((b2 - b1) / (a2 - a1)) * (Y pred -1)

[0422] Y in mathematical formula 10 pred represents the codeword for the luminance component of the input image, and a1 and a2 are Y pred These are the first and last codewords of a specific dynamic range segment containing, and b1 and b2 are the first and last codewords of the corresponding mapped dynamic range segment. Forward mapping and reverse mapping can be calculated directly using Equation 10 or a Look-Up Table (LUT) method can be used.

[0423] The encoder can convert the input original image into a reconstructed dynamic region image by forward mapping. At this time, the encoder can perform encoding in the reconstructed dynamic region. For convenience of explanation, the original dynamic region is referred to as the first domain, and the reconstructed dynamic region as the second domain. After encoding of the image is completed, the decoder must store an image identical to the restored image as a reference image. The encoder can output the restored image through the same decoding process as the decoder and store it as a reference image. At this time, the video signal processing device can convert the dynamic region image reconstructed in the second domain into the original dynamic region image of the first domain by reverse mapping before it is used as a reference image. When the video signal processing device performs intra prediction for the current block, the device may not perform a forward mapping process on the reconstructed reference samples surrounding the current block. This is because the reconstructed reference samples surrounding the current block are samples in the second domain. When the video signal processing device performs inter prediction for the current block, the video signal processing device forward maps a reference sample of the reference image to convert it into a sample in the second domain and can use the converted sample as a reference sample. This is because the reference sample of the reference image is a sample in the first domain.

[0424] As previously explained, the video signal processing device can construct the dynamic range of the chrominance component using chrominance component scaling. Specifically, the video signal processing device can correct the chrominance component based on the interrelationship between the luminance component and the chrominance component. The video signal processing device can derive scaling parameters used for chrominance component scaling based on the average value of the luminance samples. The video signal processing device can obtain the average value of the luminance samples in the first domain by inverse mapping the average value of the luminance samples in the second domain, and derive the scaling parameters according to a table defining the relationship between the luminance component and the scaling parameters. The video signal processing device can generate a scaled chrominance residual signal by multiplying the chrominance residual signal by the scaling parameters. At this time, the video signal processing device can restore the chrominance signal by adding the chrominance prediction signal and the scaled chrominance residual signal.

[0425] When a video signal processing unit restores a chrominance signal in LMCS, the video signal processing unit applies the same scaling parameter to all chrominance samples within a block. Therefore, although the average image quality of the chrominance component may be improved, the difference in variation between chrominance samples may not be taken into account. To compensate for this drawback, the video signal processing unit can perform mapping or inverse mapping of the chrominance signal by applying CCCM instead of scaling the chrominance component in LMCS.

[0426] FIGS. 43 and 44 show the generation of a color difference block by applying CCCM in an LMCS according to an embodiment of the present invention.

[0427] As previously explained, the video signal processing device can convert the representation range of the original image in the first domain to the second domain. At this time, the video signal processing device can convert the luminance component of the first domain, OrgY in FIG. 43, into the luminance component of the second domain, Y in FIG. 43, by applying forward mapping of LMCS. Additionally, the video signal processing device can obtain filter coefficients, Dervie filter coefficients in FIG. 43, by utilizing the correlation between the luminance component and the chrominance component of the first domain. The video signal processing device can obtain the chrominance components of the second domain, Cb and Cr in FIG. 43, by applying a filter using the luminance component of the second domain, the chrominance component of the first domain, and the filter coefficients in FIG. 43. At this time, the luminance component of the second domain may be one that has been converted through the LMCS forward mapping described above.

[0428] The video signal processing device can derive filter coefficients using the model used to derive the previously described CCP model. In this case, the model used to derive the CCP model may be one of CCLM, MMLM, GLM, MM-GLM, CCCM, MM-CCCM, GL-CCCM, GL-MM-CCCM, CCCM-ND, MM-CCCM-ND, CCCM-MDF, MM-CCCM-MDF, and BVG-CCCM, or CCP merge. Additionally, the video signal processing device may use the CCP model used to derive the filter coefficients as the CCP model used when applying filtering.

[0429] When encoding, the image (Y, Cb, Cr) of the second domain may be degraded due to quantization, and the restored image (Y', Cb', Cr') may differ from the image (Y, Cb, Cr) before degradation. When decoding, the restored image (Y', Cb', Cr') is an image of the second domain, so it can be converted into an image of the first domain. At this time, the decoder can generate the luminance component of the first domain, RecY of FIG. 44, by applying LMCS inverse mapping to the luminance component of the second domain, Y' of FIG. 44.

[0430] The video signal processing device can obtain filter coefficients, Derive filter coefficients of FIG. 44, by utilizing the correlation between the luminance component and the chrominance component of the second domain. The video signal processing device can generate chrominance components of the first domain, RecCb and RecCr of FIG. 44, by applying a filter using the luminance component of the first domain, the chrominance component of the second domain, and the filter coefficients of FIG. 44. The video signal processing device can obtain the luminance component of the first domain using LMCS inverse mapping.

[0431] A video signal processing device can derive filter coefficients using the CCP model used during encoding. In a specific embodiment, the CCP model used during encoding may include one of CCLM, MMLM, GLM, MM-GLM, CCCM, MM-CCCM, GL-CCCM, GL-MM-CCCM, CCCM-ND, MM-CCCM-ND, CCCM-MDF, MM-CCCM-MDF, BVG-CCCM, or CCP merge. Additionally, the CCP model used when applying filtering may be the same CCP model used when deriving filter coefficients.

[0432] The video signal processing device may apply the previously described embodiment in a sample unit or a block unit. When the video signal processing device applies the previously described embodiment in a block unit, the luminance component and the chrominance component may each be a luminance block and a chrominance block, respectively. Additionally, the video signal processing device may determine whether to apply the transformation of the chrominance component according to the previously described embodiment in at least one of a block unit, a slice unit, a tile unit, a picture unit, or an image unit. The encoder may include information in the bitstream indicating whether to apply the transformation of the chrominance component according to the previously described embodiment according to the application unit. The decoder may parse the said information and determine whether to apply the chrominance component transformation using the CCP model in a block unit, a slice unit, a tile unit, a picture unit, or an image unit.

[0433] The CCP model is derived by utilizing the correlation between the luminance component and the chrominance component. Therefore, a video signal processing device can use the CCP model to predict the difference in the chrominance component changed by signal processing based on the difference in the luminance component changed by the signal processing. In this case, the signal processing may include filtering. Additionally, the CCP model may include CCRM.

[0434] A video signal processing device can derive a relationship between a luminance component and a chrominance component in this manner, and obtain a changed chrominance component by applying the derived relationship to a luminance component changed by signal processing. The changed chrominance component may be a chrominance component obtained from a changed luminance component using the derived relationship. At this time, the video signal processing device can reconstruct the final block using the luminance component and the changed chrominance component changed by signal processing. Through this, the video signal processing device can obtain a changed chrominance component without signal processing on the chrominance component, thereby reducing the complexity of the video signal processing device. The relationship between the luminance component and the chrominance component can be obtained using a CCP model. In this specification, a filter using a CCP model or a CCP model-based filter represents a filter that takes a filtered luminance component as input to a CCP model having coefficients obtained using the relationship between the luminance component before filtering and the chrominance component before filtering. Specifically, a filter using a CCP model or a CCP model-based filter may be a filter that obtains the value of a filtered chrominance component by inputting the filtered luminance component into a CCP model having coefficients obtained using the relationship between the luminance component and the chrominance component before filtering. In another specific embodiment, a filter using a CCP model or a CCP model-based filter may be a filter that performs a weighted average of the value obtained by inputting the filtered luminance component into the CCP model and the chrominance component before filtering.

[0435] Furthermore, even if the relationship between the luminance component and the chrominance component is described as being obtained based on the CCP model in the embodiments described below, a relationship between the luminance component and the chrominance component obtained by a different method may be used. Additionally, in the embodiments described below, the component may include a sample or a block.

[0436] FIG. 45 shows color difference filtering using a CCP model according to an embodiment of the present invention.

[0437] A video signal processing device can derive filter coefficients of a CCP model using luminance components and chrominance components before filtering. The video signal processing device can generate a filtered luminance component, FilteredY of FIG. 45, by applying filtering to the luminance component, Y of FIG. 45. The video signal processing device can generate filtered chrominance components, FilteredCb and FilteredCr of FIG. 45, by applying a filter having the derived filter coefficients to the filtered luminance component. Filtering may include at least one of deblocking filtering, SAO (sample adaptive offset) filtering, CC-SAO (cross component-sample adaptive offset) filtering, ALF (adaptive loop filter) filtering, CC-ALF (cross component-adaptive loop filter) filtering, BF (bilateral filter) filtering, or film grain filtering. BF filtering is a filtering method that applies a sample-unit offset derived based on the difference in change between surrounding samples on a sample-unit basis. CC-SAO is a method for compensating samples and, similar to SAO, is a filtering method that classifies reconstructed samples into categories, derives an offset for each category, and adds it to the reconstructed samples. While SAO uses only the luminance and chrominance components, CC-SAO classifies categories using a total of three components, including luminance and two chrominances. CC-ALF is a filtering method that modifies reconstructed chrominance samples using values ​​derived by applying a linear filter to the reconstructed luminance sample values. In a specific embodiment, the video signal processing device can derive filter coefficients of the CCP model using the luminance and chrominance components before deblocking filtering. At this time, the video signal processing device can generate a deblocked luminance component by applying deblocking filtering to the luminance component.A video signal processing device can generate a filtered chrominance component by applying a filter utilizing a CCP model with derived filter coefficients to a deblocked chrominance component. The filtered chrominance component may be a chrominance component modified using the CCP model, rather than the deblocked chrominance component. The video signal processing device can output a final filtered chrominance component by blending the chrominance component before deblocking filtering with the filtered chrominance component. The video signal processing device can use the output final filtered chrominance component as the deblocked chrominance component. Film grain filtering is a filtering method that overlays noise onto a picture to provide a weathered-looking image. A video signal processing device can derive the filter coefficients of a CCP model using the luminance and chrominance components before film grain filtering. A video signal processing device can generate a film grain-filtered luminance component by applying film grain filtering to the luminance component. A video signal processing device can generate a film grain-filtered chrominance component by applying a filter utilizing a CCP model with derived filter coefficients to the film grain-filtered luminance component. The film grain-filtered chrominance component may be a chrominance component modified using a CCP model, rather than the film grain-filtered chrominance component itself. The video signal processing device may output a final modified chrominance component by blending the chrominance component before film grain filtering with the modified chrominance component. In this case, the video signal processing device may use the final modified chrominance component as the film grain-filtered chrominance component.

[0438] FIG. 46 shows color difference filtering using a CCP model according to another embodiment of the present invention.

[0439] As previously explained, filtering may include at least one of deblocking, SAO, CC-SAO, ALF, or CC-ALF and BF, and is applied to the luminance component and the chrominance component, respectively. Therefore, filtering may require high complexity. The chrominance component has milder characteristics than the luminance component. Accordingly, a video signal processing device may apply multiple filtering methods to the luminance component and not perform filtering on the chrominance component, or apply CCP model-based filtering to the chrominance component. For example, a video signal processing device may apply three filtering methods (A, B, C) to the luminance component and apply CCP model-based filtering to the chrominance component instead of the three filtering methods. This allows for a reduction in the complexity of filtering for the chrominance component.

[0440] Additionally, the video signal processing device may apply multiple filters to the luminance component, perform some of the filters on the chrominance component, and apply CCP model-based filtering to the chrominance component instead of the remaining filters. For example, in the embodiment of FIG. 46, the video signal processing device applies three filters (A, B, C) to the luminance sample. At this time, the video signal processing device may perform one filter (A) on the chrominance component and apply CCP model-based filtering instead of the remaining two filters (B, C). In these embodiments, the three filters may include a deblocking filter, SAO, CC-SAO, ALF filter, or CC-ALF. The one filter (A) applied to the chrominance component may be a deblocking filter. The video signal processing device may apply a deblocking filter, SAO, and ALF to the luminance component, apply only the deblocking filter to the chrominance component, and then apply CCP model-based filtering instead of applying SAO and ALF to the chrominance component.

[0441] The video signal processing device can apply the prediction of chrominance components using a CCP model to inter-prediction, intra-prediction, reference picture resampling (RPR), upsampling, or downsampling, in addition to the in-loop filtering of the previously described embodiments. In this case, RPR involves encoding and decoding by adjusting the spatial resolution of the picture. In RPR, when the video signal processing device encodes a picture at a small spatial resolution, it may downsample the restored picture at a large resolution and use it as a reference picture. Additionally, the video signal processing device may upsample the restored picture at a small resolution and use it as a reference picture to encode the picture restored at a large spatial resolution.

[0442] In a specific embodiment, the video signal processing device may apply the difference of the luminance sample modified through PDPC by performing CCP model-based filtering on the chrominance sample during intra prediction. The video signal processing device may derive the filter coefficients of the CCP model through the relationship between the luminance sample and the chrominance sample before the application of PDPC. The video signal processing device may apply CCP model-based filtering to the chrominance sample. Specifically, the video signal processing device may generate a filtered chrominance sample by applying the luminance sample to which PDPC has been applied to a filter utilizing the derived CCP model. When the video signal processing device performs inter prediction, the video signal processing device may apply PROF (prediction refinement with optical flow), which calculates a sample-unit correction value based on spatial sample changes, and BDOF (Bi-directional optical flow), which corrects the current sample by estimating the amount of pixel change in a reference block, to the luminance sample to improve the accuracy of sample-unit motion prediction. The video signal processing device may not apply such sample correction to the chrominance sample. Therefore, when the video signal processing device performs inter-prediction, it can obtain filter coefficients of the CCP model through the relationship between the luminance sample and the chrominance sample before the application of PROF and / or BDOF. The video signal processing device can generate filtered chrominance samples by applying a filter utilizing the obtained CCP model to the luminance samples to which the PROF and / or BDOF scheme has been applied. Through this, the video signal processing device can correct the chrominance samples, similar to the luminance samples corrected by PROF and / or BDOF. Additionally, the video signal processing device may apply upsampling and / or downsampling to the chrominance samples through the upsampling and / or downsampling CCP model.As an example of implementation, filter coefficients of a CCP model can be derived through the relationship between a luminance sample and a chrominance sample before upsampling and / or downsampling, and then filtering using the CCP model can be applied to the chrominance sample through the upsampling and / or downsampling luminance sample and the filter coefficients of the CCP model to generate an upsampled and / or downsampled chrominance sample. A video signal processing device can use such upsampling and / or downsampling in RPR. When the video signal processing device predicts the current block, the video signal processing device can obtain filter coefficients of a CCP model through the relationship between a luminance sample and a chrominance sample before the first method is applied. The video signal processing device can generate a filtered chrominance sample by applying a filter using the CCP model to the luminance sample to which the first method is applied. The first method may include at least one of LIC applied to a luminance sample in Intra TMP mode, a Wiener filter applied to a luminance sample in Intra TMP mode, LIC applied to a luminance sample in IBC mode, blending applied to the boundary between two regions in SGPM or GPM mode, OBMC using a block vector of a surrounding block in Intra TMP or IBC mode, LIC applied to a sample predicted through motion information in Inter-prediction mode, OBMC using motion information of a surrounding block in Inter-prediction mode, or BCW applied to a sample predicted through motion information in Inter-prediction mode.

[0443] FIG. 47 shows a color difference compensation process using a CCP model according to an embodiment of the present invention.

[0444] There may be a high correlation between color difference components. A video signal processing device can predict color difference samples using the correlation between color difference components. Specifically, the video signal processing device can derive filter coefficients of a CCP model using the correlation between the Cb color difference prediction sample, the Cb and Cr color difference sample of FIG. 47, and Cr of FIG. 47. At this time, any one of the previously described CCP models may be used as the CCP model. The video signal processing device can generate a restored Cb color difference sample, Cb' of FIG. 47, by adding a Cb color difference residual sample to the Cb color difference sample. A corrected Cr color difference sample can be generated by applying filtering using the CCP model through the restored Cb color difference sample and the derived CCP model. At this time, the video signal processing device can use the corrected Cr color difference sample as the restored Cr color difference sample, Cr' of FIG. 47. In another specific embodiment, a restored Cr color difference sample, Cr' of FIG. 47, can be generated by adding a Cr color difference residual sample to the corrected Cr color difference sample. Although the embodiments described above were explained as being processed on a sample basis, the operation of the embodiments described above can be applied on a block basis.

[0445] FIG. 48 shows compensation between components using a CCP model according to an embodiment of the present invention.

[0446] The video signal processing device can convert the Y, Cb, and Cr components into R, G, and B domains and perform encoding and decoding for each of the R, G, and B components. Additionally, the video signal processing device can convert each of the R, B, and G components back into Y, Cb, and Cr. At this time, when the video signal processing device performs encoding and decoding in the R, G, and B domains, it can generate predicted samples using a CCP model obtained by utilizing the correlation between the R, G, and B components. Specifically, the video signal processing device can first derive first filter coefficients for the CCP model using the correlation between the R sample and the G sample. Additionally, the video signal processing device can generate a restored R sample by adding the R residual sample to the R sample. The video signal processing device can generate a corrected G sample by applying a filter using the CCP model that uses the derived first filter coefficients to the restored R sample. The video signal processing device can use the corrected G sample as the restored G sample. In another specific embodiment, the video signal processing device can generate a restored G sample by adding the G residual sample to the corrected G sample. The video signal processing device can obtain second filter coefficients for the CCP model using the correlation between the R sample and the B sample. Additionally, the video signal processing device can obtain third filter coefficients for the CCP model using the correlation between the G sample and the B sample. The video signal processing device can generate a corrected first B sample by applying a filter utilizing the CCP model using the second filter coefficients derived from the reconstructed R sample. Additionally, the video signal processing device can generate a corrected second B sample by applying a filter utilizing the CCP model using the third filter coefficients derived from the reconstructed G sample. The video signal processing device may use either the corrected first B sample or the corrected second B sample as the reconstructed B sample.In another specific embodiment, the video signal processing device may generate a restored B sample by weighted averaging at least two of the B sample before correction, the first corrected B sample, and the second corrected B sample. The restored B sample may be generated by adding a B residual sample to either the first corrected B sample or the second corrected B sample. In another specific embodiment, the video signal processing device may generate a restored B sample by adding a B residual sample to the third B sample generated by the weighted average above. Although the embodiments described above were explained as being processed on a sample-by-sample basis, the operation of the embodiments described above may be applied on a block-by-block basis. Furthermore, the processing order of the R, G, and B components may be applied in various ways.

[0447] FIG. 49 shows a video signal processing device according to an embodiment of the present invention mapping color difference components using a CCP method.

[0448] FIG. 50 shows the area of ​​a window used by a video signal processing device according to an embodiment of the present invention when deriving chroma mapping parameters from a luminance and color difference picture when performing chroma mapping.

[0449] FIG. 51 shows a CCP model used in chroma mapping by a video signal processing device according to an embodiment of the present invention.

[0450] FIG. 52 shows a video signal processing device according to an embodiment of the present invention inducing samples outside the window area boundary in chroma mapping.

[0451] A video signal processing device can determine the variables used for chroma mapping based on signaling information. In this case, the signaling information is referred to as chroma mapping information. Chroma mapping information may indicate at least one of whether downsampling is applied to the luminance component in chroma mapping, the shape of the chroma mapping filter, whether blending is applied, the blending type, or the size of the chroma mapping window. The chroma mapping window may indicate the area of ​​the component used when deriving or applying chroma mapping parameters. The shape of the chroma mapping filter indicates the shape of the chroma filter and, for example, may indicate the type of the CCP model. Whether blending is applied indicates whether to blend the mapped color difference sample with the color difference sample before mapping after chroma mapping. Additionally, the blending type may indicate how to set the blending weights between the color difference sample before mapping and the mapped color difference sample.

[0452] The video signal processing device can generate mapped luminance samples by applying LMCS-based luminance sample mapping to samples (Y, Cb, Cr) prior to mapping. The video signal processing device can generate downsampled luminance samples by downsampling the luminance samples prior to mapping and the mapped luminance samples, respectively. The video signal processing device can determine whether to downsample the luminance component based on chroma mapping information. The video signal processing device can set at least one of the type of the chroma mapping filter, whether to apply blending, the blending type, or the window size of the chroma mapping using the chroma mapping information.

[0453] The video signal processing device may set the window size for deriving chroma mapping parameters to be smaller than and / or equal to the maximum size of the CTU. In another specific embodiment, the window size for chroma mapping may be larger than the maximum size of the CTU. The video signal processing device may divide the CTU area into chroma mapping window areas. In this case, the video signal processing device may derive chroma mapping parameters or apply chroma mapping to each chroma mapping window area. The video signal processing device may obtain a correlation between a downsampled luminance sample and an input chroma difference sample by applying a CCP to each set chroma mapping window. Based on the obtained correlation, the video signal processing device may derive chroma mapping parameters for the Cb chroma difference sample and chroma mapping parameters for the Cr chroma difference sample.

[0454] To derive chroma mapping parameters, the video signal processing device may use a chroma mapping filter such as the embodiment of FIG. 51. It may be a filter that uses luminance samples L0, L1, L2, and L3 of FIG. 51 and a chrominance sample C corresponding to the L0 luminance sample. The video signal processing device may use one of the luminance sample locations of L0, L1, L2, and L3 as the location of the chrominance sample C.

[0455] The video signal processing device may not access samples outside the right boundary and samples outside the bottom boundary of the window area in order to process each window area of ​​the chroma mapping in parallel. As shown in FIG. 52, the video signal processing device may pad the values ​​at sample locations outside the window area with adjacent sample values ​​inside the window area boundary. The process of generating padded samples can be applied to luminance samples before mapping and luminance samples after mapping. When deriving chroma mapping parameters, the video signal processing device may use padded sample values ​​using luminance samples before mapping. When applying chroma mapping, the video signal processing device may use padded sample values ​​using luminance samples after mapping. The video signal processing device may apply chroma mapping using mapped luminance samples, chroma mapping parameters for derived Cb chroma difference samples, and chroma mapping parameters for derived Cr chroma difference samples. The video signal processing device may generate mapped Cb chroma difference samples and mapped Cr chroma difference samples. As previously described, the video signal processing device may decide whether to apply blending based on the chroma mapping information. When blending is applied, the video signal processing device may obtain a blending type from chroma mapping information and then blend the color difference sample before mapping with the mapped color difference sample to generate a final mapped color difference sample. At this time, the video signal processing device may perform blending by multiplying the color difference sample before mapping by J, multiplying the mapped color difference sample by K, and then right-shifting by L. J, K, and L may be integers greater than or equal to 1. Specifically, (J, K, L) may be (2, 6, 3), (3, 5, 3), (4, 4, 3), (5, 3, 3), or (6, 2, 3).

[0456] The video signal processing device may apply clipping to the final mapped chroma component. The range of the minimum and maximum values ​​of the clipping may be the range of the values ​​of the chroma component, e.g., the sample values. In another specific embodiment, the minimum and maximum values ​​of the clipping may be set as the minimum value obtained by subtracting a preset value from the values ​​of the chroma samples before chroma mapping, and as the maximum value obtained by adding a preset value to the values ​​of the chroma samples before chroma mapping. The preset value may be an integer greater than or equal to 1. For example, the preset value may be 10 or 20. The video signal processing device may perform chroma mapping before filtering. In another specific embodiment, the video signal processing device may perform chroma mapping after filtering. The filtering may be in-loop filtering. Specifically, the filtering may include at least one of deblocking filtering, CC-SAO, BF, or CC-ALF.

[0457] When the range of luminance samples and chroma difference samples used to derive chroma mapping parameters or apply chroma mapping is 10 bits deep, the luminance samples and chroma difference samples can have values ​​from 0 to 1023. If the number of chroma mapping filter types and chroma mapping parameter coefficients is large, the number of chroma mapping filter types and chroma mapping parameter coefficients may exceed the range that the video signal processing device can allow during calculation. The video signal processing device may subtract a pre-specified offset from the luminance samples and chroma difference samples. The video signal processing device may use the subtracted luminance samples and subtracted chroma difference samples when deriving chroma mapping parameters or applying chroma mapping. The pre-specified offset value may be one of the median of N bit depth, the average value of the luminance samples in the window area, and the average value of the chroma difference samples in the window area. When N is 10, the median may be 512.

[0458] The encoder may include chroma mapping information in the bitstream that indicates at least one of whether downsampling is applied, the filter type of the chroma mapping, whether blending is applied, the blending type, or the window size of the chroma mapping. The decoder may parse the chroma mapping information and set at least one of whether downsampling is applied, the filter type of the chroma mapping, whether blending is applied, the blending type, or the window size. The video signal processing device may derive at least one of whether downsampling is applied, the filter type of the chroma mapping, whether blending is applied, the blending type, and the window size from information on surrounding blocks. In this case, the information on surrounding blocks may be information derived using one of the following: information on surrounding blocks spatially adjacent to the current block, information on surrounding blocks that are not adjacent to the current block but are separated, information on temporal surrounding blocks corresponding to the current block in a reference picture, or information within a history-based table. The video signal processing device may derive chroma ...

Claims

1. In a video signal decoding device, Includes a processor, The above processor Generate a restoration block for the current block of the current picture included in the above video signal, and A relationship equation representing the correlation between the luminance component and the color difference component of the above restoration block is obtained, and Signal processing is performed on the above luminance component, and A color difference component is obtained by performing signal processing using the luminance component to which the above signal processing was performed and the above relationship equation, and Generating a final restoration block for the current block based on the luminance component for which the above signal processing has been performed and the chrominance component for which the above signal processing has been performed. Video signal decoding device.

2. In Paragraph 1, The above processor When obtaining a relationship equation representing the correlation between the luminance component and the color difference component of the above-mentioned restoration block, using a downsampled luminance component Video signal decoding device.

3. In Paragraph 2, The above processor Information is obtained indicating whether to use the downsampled luminance component from the video signal, and When obtaining a relationship equation representing the correlation between the luminance component and the color difference component of the above-mentioned restoration block, determining whether to use the downsampled luminance component according to information indicating whether to use the downsampled luminance component. Video signal decoding device.

4. In Paragraph 1, The above processor Information regarding the format of the relationship equation is obtained from the above video signal, and Obtaining a relationship based on information regarding the format of the above relationship Video signal decoding device.

5. In Paragraph 1, The above processor Information regarding the blending of the color difference component before the signal processing and the color difference component after the signal processing is performed is obtained from the above video signal, and Blending the color difference sample before the signal processing and the color difference sample after the signal processing according to information regarding blending Video signal decoding device.

6. In Paragraph 5, The information regarding the above blending is including information regarding whether the above blending is applied Video signal decoding device.

7. In Paragraph 5, The information regarding the above blending is including information regarding the type of the above blending Video signal decoding device.

8. In Paragraph 1, The above processor Information regarding a window that specifies the range of components used to obtain the relationship from the video signal or the range of components to which the relationship is applied, and Based on information regarding the above window, the above relationship formula is obtained, or the range of color difference components to which the above relationship formula is applied is determined based on information regarding the above window. Video signal decoding device.

9. In Paragraph 1, The color difference component on which the above signal processing has been performed is clipped to obtain the clipped color difference component, and Generating a final restored block for the current block using the luminance component on which the above signal processing was performed and the clipped chrominance component. Video signal processing device.

10. In Paragraph 9, The range of the clipping above is determined according to the range of color difference sample values ​​before the signal processing is performed. Video signal processing device.

11. In Paragraph 1, The above signal processing includes an in-loop filter. Video signal processing device.

12. In a video signal encoding device, Includes a processor, The above processor Generate a restoration block for the current block of the current picture included in the above video signal, and A relationship equation representing the correlation between the luminance component and the color difference component of the above restoration block is obtained, and Signal processing is performed on the above luminance component, and A color difference component is obtained by performing signal processing using the luminance component to which the above signal processing was performed and the above relationship equation, and Generating a final restoration block for the current block based on the luminance component for which the above signal processing has been performed and the chrominance component for which the above signal processing has been performed. Video signal encoding device.

13. In Paragraph 12, The above processor When obtaining a relationship equation representing the correlation between the luminance component and the color difference component of the above-mentioned restoration block, using a downsampled luminance component Video signal encoding device.

14. In Paragraph 13, The above processor Including information indicating whether to use the downsampled luminance component in a bit stream containing the video signal Video signal encoding device.

15. In Paragraph 12, The above processor Including information regarding the format of the above relationship in a bit stream containing the video signal Video signal encoding device.

16. In Paragraph 12, The above processor Including information regarding the blending of the color difference component before signal processing and the color difference component after signal processing from the video signal in a bit stream containing the video signal. Video signal encoding device.

17. In Paragraph 16, The information regarding the above blending is including information regarding whether the above blending is applied Video signal encoding device.

18. In Paragraph 16, The information regarding the above blending is including information regarding the type of the above blending Video signal encoding device.

19. In the method of operation of a video signal decoding device, A step of generating a restoration block for the current block of the current picture included in the above video signal; A step of obtaining a relationship equation representing the correlation between the luminance component and the color difference component of the above-mentioned restoration block; A step of performing signal processing on the above luminance component; A step of obtaining a color difference component for which signal processing has been performed using the luminance component for which signal processing has been performed and the relationship equation; and A step comprising generating a final restored block for the current block based on the luminance component on which the signal processing was performed and the chrominance component on which the signal processing was performed. Method of operation.

20. In a bitstream containing a video signal contained in a storage medium, The method for generating a bit stream A step of generating a restoration block for the current block of the current picture included in the above video signal; A step of obtaining a relationship equation representing the correlation between the luminance component and the color difference component of the above-mentioned restoration block; A step of performing signal processing on the above luminance component; A step of obtaining a color difference component for which signal processing has been performed using the luminance component for which signal processing has been performed and the relationship equation; and A step comprising generating a final restored block for the current block based on the luminance component on which the signal processing was performed and the chrominance component on which the signal processing was performed. Bitstream.