Method for processing video signal and device therefor

The video signal processing method generates mixed vectors for improved prediction accuracy and efficiency by using block and motion information, weighted averaging, and filtering, addressing inefficiencies in existing video compression methods.

WO2026023972A1PCT designated stage Publication Date: 2026-01-29WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/010290
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-14
Filing Date
2025-07-14
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing video signal processing methods are inefficient in terms of coding efficiency, particularly in handling spatial, temporal, and probabilistic correlations in video compression.

Method used

A video signal processing method and device that generates a mixed vector using block and motion information for a current block, performs weighted averaging of prediction blocks, and constructs a mixed vector list to improve prediction accuracy, while avoiding certain prediction modes like intrafusion and PDPC, and includes filtering to enhance quality.

Benefits of technology

Enhances coding efficiency by improving prediction accuracy and quality of video signals through advanced vector-based prediction techniques and filtering, leading to more effective video compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025010290_29012026_PF_FP_ABST
    Figure KR2025010290_29012026_PF_FP_ABST
Patent Text Reader

Abstract

A video signal decoding device is disclosed. The video signal decoding device includes a processor. The processor generates a mixed vector on the basis of a block vector for a current block and motion information for the current block, and generates a prediction block for the current block by using the mixed vector.
Need to check novelty before this filing date? Find Prior Art

Description

Video signal processing method and device therefor

[0001] The present invention relates to a method and device for processing a video signal, and more particularly, to a method and device for processing a video signal for encoding or decoding a video signal.

[0002] Compression coding refers to a series of signal processing techniques for transmitting digitized information over communication lines or storing it in a format suitable for storage media. Objects subject to compression coding include audio, video, and text, and the technology that performs compression coding on video in particular is called video compression. Compression coding of video signals is achieved by removing redundant information by considering spatial, temporal, and probabilistic correlations. However, with the recent advancements in various media and data transmission media, there is a growing demand for more efficient video signal processing methods and devices.

[0003] The purpose of this specification is to provide a video signal processing method and a device therefor to improve the coding efficiency of a video signal.

[0004] The present specification provides a video signal processing method and a device therefor. A video signal decoding device according to an embodiment of the present invention includes a processor. The processor generates a mixed vector based on a block vector for a current block and motion information for the current block, and generates a predicted block for the current block using the mixed vector.

[0005] The above mixed vector may include a block vector indicated by block vector information for the current block and a motion vector indicated by motion information for the current block. The processor may obtain a first prediction block based on the block vector, obtain a second prediction block based on the motion vector, and then perform a weighted average of the first prediction block and the second prediction block to generate a prediction block for the current block.

[0006] The processor can construct a mixed vector list using block vector information derived from a first surrounding block of the current block and motion information derived from a second surrounding block of the current block, and obtain a block vector for the current block and motion information for the current block from the mixed vector list.

[0007] The above processor can derive a block vector from a reference block of a reference picture indicated by motion information of a surrounding block of the current block.

[0008] The processor can determine whether to use information of a block surrounding the current block as a candidate of the mixed vector candidate list according to an encoding mode of the block surrounding the current block.

[0009] A video signal decoding device according to an embodiment of the present invention includes a processor.

[0010] The processor generates a first prediction block using first motion information of a current block, generates a second prediction block using neighboring blocks of the current block and an intra prediction mode derived from the neighboring blocks, and generates a final prediction block by weighting the first prediction block and the second prediction block.

[0011] The processor may generate the second prediction block using a decoder side intra mode derivation (DIMD) mode derived using restored samples of the surrounding blocks.

[0012] The above processor may not use intrafusion when generating the second prediction block.

[0013] The above processor may not use PDPC (position dependent intra prediction combination) when generating the second prediction block.

[0014] The above processor can perform filtering on the final prediction block.

[0015] A video signal encoding device according to an embodiment of the present invention includes a processor. The processor generates a mixture vector based on a block vector for a current block and motion information for the current block, and generates a prediction block for the current block using the mixture vector.

[0016] The above mixed vector may include a block vector indicated by block vector information for the current block and a motion vector indicated by motion information for the current block. The processor may obtain a first prediction block based on the block vector, obtain a second prediction block based on the motion vector, and then perform a weighted average of the first prediction block and the second prediction block to generate a prediction block for the current block.

[0017] The processor can construct a mixed vector list using block vector information derived from a first surrounding block of the current block and motion information derived from a second surrounding block of the current block, and obtain a block vector for the current block and motion information for the current block from the mixed vector list.

[0018] The above processor can derive a block vector from a reference block of a reference picture indicated by motion information of a surrounding block of the current block.

[0019] The processor can determine whether to use information of a block surrounding the current block as a candidate of the mixed vector candidate list according to an encoding mode of the block surrounding the current block.

[0020] A video signal encoding device according to an embodiment of the present invention includes a processor, wherein the processor can generate a first prediction block using first motion information of a current block, generate a second prediction block using a neighboring block of the current block and an intra prediction mode derived from the neighboring block, and generate a final prediction block by weighting the first prediction block and the second prediction block.

[0021] The processor can generate the second prediction block using a decoder side intra mode derivation (DIMD) mode derived using restored samples of the surrounding blocks.

[0022] The above processor may not use intrafusion when generating the second prediction block.

[0023] The above processor may not use PDPC (position dependent intra prediction combination) when generating the second prediction block.

[0024] The above processor can perform filtering on the final prediction block.

[0025] An operating method of a video signal decoding device according to an embodiment of the present invention includes the steps of generating a mixture vector based on a block vector for a current block and motion information for the current block; and the step of generating a prediction block for the current block using the mixture vector.

[0026] An operating method of a video signal decoding device according to an embodiment of the present invention comprises the steps of: generating a first prediction block using first motion information of a current block; generating a second prediction block using a neighboring block of the current block and an intra prediction mode derived from the neighboring block; and generating a final prediction block by weighting the first prediction block and the second prediction block.

[0027] According to an embodiment of the present invention, a bitstream is disclosed, which is included in a storage medium and includes a video signal. A method for generating a bitstream includes the steps of generating a mixture vector based on a block vector for a current block and motion information for the current block; and the step of generating a prediction block for the current block using the mixture vector.

[0028] According to an embodiment of the present invention, a bitstream is disclosed, which is included in a storage medium and includes a video signal. A method of generating a bitstream includes the steps of: generating a first prediction block using first motion information of a current block; generating a second prediction block using neighboring blocks of the current block and intra-prediction modes derived from the neighboring blocks; and generating a final prediction block by weighting the first prediction block and the second prediction block.

[0029] This specification provides a method for efficiently processing a video signal.

[0030] The effects that can be obtained from this specification are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by those skilled in the art to which the present invention pertains from the description below.

[0031] FIG. 1 is a schematic block diagram of a video signal encoding device according to one embodiment of the present specification.

[0032] FIG. 2 is a schematic block diagram of a video signal decoding device according to one embodiment of the present specification.

[0033] Figure 3 illustrates an embodiment in which a coding tree unit within a picture is divided into coding units.

[0034] Figure 4 illustrates one embodiment of a method for signaling the splitting of a quad tree and a multi-type tree.

[0035] Figures 5 and 6 illustrate the intra prediction method according to an embodiment of the present invention in more detail.

[0036] Figure 7 is a diagram showing the locations of surrounding blocks used to construct a motion candidate list in inter prediction.

[0037] FIG. 8 illustrates a method for determining a reference pixel line based on a template according to one embodiment of the present specification.

[0038] FIG. 9 is a diagram illustrating a block vector related to an IBC encoding method according to one embodiment of the present specification.

[0039] Figure 10 illustrates a method for predicting the current block using RRIBC in the horizontal direction.

[0040] Figure 11 illustrates a method for predicting the current block using RRIBC in the vertical direction.

[0041] FIG. 12 illustrates a block vector of a block encoded in Intra TMP mode according to one embodiment of the present specification.

[0042] FIG. 13 illustrates a case where a current block is divided by GPM mode according to one embodiment of the present specification and the divided area is encoded by IBC mode.

[0043] FIG. 14 illustrates how a current block is encoded in IBC-CIIP mode according to one embodiment of the present specification.

[0044] FIG. 15 illustrates an example of a reference region and filter shape used to derive CCCM parameters according to one embodiment of the present specification.

[0045] FIG. 16 illustrates types of transform kernels that can be used in video coding according to one embodiment of the present specification.

[0046] FIG. 17 illustrates a transformation set table for LFNST and NSPT transformations according to one embodiment of the present specification.

[0047] Figure 18 shows an example of ROI after LFNST transformation.

[0048] FIG. 19 illustrates a method for deriving a multi-transform set and a LFNST / NSPT set according to one embodiment of the present specification.

[0049] Figure 20 illustrates a mapping table according to one embodiment of the present specification.

[0050] Figure 21 illustrates a conversion type set table according to one embodiment of the present specification.

[0051] Figure 22 illustrates a conversion type combination table according to one embodiment of the present specification.

[0052] FIG. 23 illustrates a threshold value table for an IDT conversion type according to one embodiment of the present specification.

[0053] FIG. 24 illustrates block boundaries and samples around the boundaries in a deblocking filtering process according to one embodiment of the present specification.

[0054] FIG. 25 is a diagram illustrating a process of generating a prediction block using DIMD (Decoder side intra mode derivation) according to one embodiment of the present invention.

[0055] FIG. 26 is a diagram showing the locations of surrounding pixels used to derive directional information according to one embodiment of the present invention.

[0056] FIG. 27 is a diagram illustrating a method for mapping directional modes according to one embodiment of the present invention.

[0057] FIG. 28 is a diagram showing a histogram for deriving an intra prediction directional mode according to one embodiment of the present invention.

[0058] FIG. 29 is a diagram illustrating a method for generating a prediction sample using intra prediction directional mode information and weights according to one embodiment of the present invention.

[0059] FIG. 30 and FIG. 31 are diagrams showing templates used to derive an intra prediction mode of a current block according to one embodiment of the present invention.

[0060] FIG. 32 is a diagram illustrating a method for generating prediction samples (pixels) based on a plurality of reference pixel lines according to one embodiment of the present invention.

[0061] FIG. 33 illustrates a method for predicting a sample using a plurality of reference pixel lines according to one embodiment of the present invention.

[0062] FIG. 34 is a structural diagram illustrating a method for determining an optimal reference pixel line using a plurality of reference pixel lines based on a template according to one embodiment of the present invention.

[0063] FIG. 35 illustrates a method for generating prediction samples using a planar mode according to one embodiment of the present invention.

[0064] Figure 36 illustrates an intra prediction mode based on an extrapolation filter according to one embodiment of the present invention.

[0065] Figure 37 illustrates a method for selecting an optimal intra prediction mode combination according to one embodiment of the present invention.

[0066] Figures 38 and 39 illustrate neighboring blocks of a current block according to one embodiment of the present invention.

[0067] Figure 40 illustrates a method of using a candidate list based on a prediction mode according to one embodiment of the present invention.

[0068] Figure 41 shows reference sample filtering according to an embodiment of the present invention.

[0069] Figure 42 shows PDPC filtering according to an embodiment of the present invention.

[0070] Figure 43 shows a gradient PDPC according to an embodiment of the present invention.

[0071] FIG. 44 is a diagram illustrating a process of performing OBMC according to one embodiment of the present specification.

[0072] FIG. 45 is a diagram illustrating a method for performing OMBC of CU units according to one embodiment of the present specification.

[0073] FIG. 46 is a diagram illustrating a method for performing OBMC in sub-block units according to one embodiment of the present invention.

[0074] FIG. 47 is a diagram showing a table in which weights are defined according to one embodiment of the present specification.

[0075] FIG. 48 is a diagram illustrating a method for configuring a template for performing OBMC according to one embodiment of the present specification.

[0076] FIG. 49 is a diagram illustrating a method for generating a prediction block according to each OBMC mode according to one embodiment of the present specification.

[0077] FIG. 50 shows a method for a video signal processing device according to an embodiment of the present invention to perform OBMC Intra.

[0078] FIG. 51 shows a method for determining the strength of deblocking filtering for a block to which OBMC is applied by a video signal processing device according to an embodiment of the present invention.

[0079] Figure 52 shows a matrix-based intra prediction method according to an embodiment of the present invention.

[0080] FIGS. 53 and 54 illustrate a method of compensating motion information using DMVR according to one embodiment of the present disclosure.

[0081] Figure 55 illustrates a process of performing multiple DMVR according to one embodiment of the present specification.

[0082] FIG. 56 illustrates a search method for obtaining a cost value related to corrected motion information of a coding block according to one embodiment of the present specification.

[0083] Figure 57 illustrates a 3x3 Square search method according to one embodiment of the present specification.

[0084] Figure 58 illustrates a search area divided into zones for an integer unit global search process according to one embodiment of the present specification.

[0085] Figures 59 and 60 illustrate a global search process in integer units according to one embodiment of the present specification.

[0086] FIG. 61 and FIG. 62 illustrate a method for performing correction of motion information based on BDOF according to one embodiment of the present specification.

[0087] FIG. 63 and FIG. 64 illustrate a method for generating a prediction block for a current block based on BDOF according to one embodiment of the present specification.

[0088] FIG. 65 illustrates a chrominance block and a luminance block corresponding to the chrominance block according to one embodiment of the present specification.

[0089] Figure 66 shows a location for deriving a peripheral block according to one embodiment of the present specification.

[0090] Figures 67 to 70 illustrate affine motion prediction according to one embodiment of the present specification.

[0091] Figures 71 and 72 illustrate modes of affine motion prediction according to one embodiment of the present specification.

[0092] Figures 73 and 74 illustrate the derivation of an affine motion predictor according to one embodiment of the present disclosure.

[0093] FIG. 75 illustrates a method for deriving TMVP motion candidates according to one embodiment of the present specification.

[0094] FIG. 76 illustrates a method for deriving a TMVP based on a motion shift candidate list according to one embodiment of the present specification.

[0095] FIG. 77 illustrates a method for reordering a motion candidate list using template cost according to one embodiment of the present specification.

[0096] FIG. 78 illustrates a method for deriving a TMVP candidate based on a motion shift candidate list according to one embodiment of the present specification.

[0097] FIGS. 79 and 80 illustrate a method for deriving affine candidates from a history parameter-based affine model according to one embodiment of the present disclosure.

[0098] FIG. 81 illustrates the locations of non-adjacent surrounding blocks according to one embodiment of the present specification.

[0099] Figure 82 illustrates a method for deriving automatic relocation motion candidates according to one embodiment of the present specification.

[0100] Figure 83 shows a reference position for deriving an automatic relocation movement candidate according to one embodiment of the present specification.

[0101] FIG. 84 shows a video signal processing device according to an embodiment of the present invention deriving automatic rearrangement motion information from temporal surrounding blocks.

[0102] Figure 85 shows a video signal processing device according to an embodiment of the present invention deriving a block vector.

[0103] FIG. 86 shows a video signal processing device according to an embodiment of the present invention predicting a current block using a mixed vector.

[0104] FIG. 87 illustrates a method for signaling whether automatic repositioning motion vectors are activated according to one embodiment of the present specification.

[0105] Figure 88 illustrates a GCI syntax structure according to one embodiment of the present specification.

[0106] The terms used in this specification have been selected from widely used and current terms, taking into account the functions of the present invention. However, these terms may vary depending on the intentions of those skilled in the art, customs, or the emergence of new technologies. Furthermore, in certain cases, the applicant may arbitrarily select terms, in which case their meanings will be described in the description of the relevant invention. Therefore, it should be noted that the terms used in this specification should be interpreted based on their substantive meaning and the overall content of this specification, rather than simply their names.

[0107] In this specification, 'A and / or B' may be interpreted to mean 'comprising at least one of A or B'.

[0108] In this specification, some terms may be interpreted as follows. Coding may be interpreted as encoding or decoding, depending on the case. In this specification, a device that encodes a video signal to generate a video signal bitstream is referred to as an encoding device or encoder, and a device that decodes a video signal bitstream to restore a video signal is referred to as a decoding device or decoder. In addition, in this specification, a video signal processing device is used as a term that includes both an encoder and a decoder. Information is a term that includes values, parameters, coefficients, elements, etc., and since the meaning may be interpreted differently depending on the case, the present invention is not limited thereto. 'Unit' is used to mean a basic unit of image processing or a specific location of a picture, and refers to an image area that includes at least one of a luminance component and a chroma component. In addition, 'block' refers to an image area including specific components among luminance components and chrominance components (i.e., Cb and Cr). However, depending on the embodiment, terms such as 'unit', 'block', 'partition', 'signal', and 'region' may be used interchangeably. In addition, in this specification, 'current block' means a block that is currently scheduled to be encoded, and 'reference block' means a block that has already been encoded or decoded and is used as a reference in the current block. In addition, in this specification, terms such as 'luma', 'luminance', and 'Y' may be used interchangeably. In addition, in this specification, terms such as 'chroma', 'chroma', 'color difference', and 'Cb or Cr' may be used interchangeably, and since chrominance components are divided into two, Cb and Cr, each chrominance component may be used separately. In addition, in this specification, a unit may be used as a concept including all of a coding unit, a prediction unit, and a transformation unit.A picture refers to a field or a frame, and depending on the embodiment, the terms may be used interchangeably. Specifically, if the captured image is an interlace image, one frame is divided into an odd (or odd, top) field and an even (or even, bottom) field, and each field is configured as one picture unit and can be encoded or decoded. If the captured image is a progressive image, one frame is configured as a picture and can be encoded or decoded. In addition, in this specification, the terms 'error signal', 'residual signal', 'residual signal', 'residual signal', and 'differential signal' may be used interchangeably. In addition, in this specification, the terms 'intra prediction mode', 'intra prediction directional mode', 'intra-screen prediction mode', and 'intra-screen prediction directional mode' may be used interchangeably. In addition, in this specification, the terms 'motion', 'movement', and the like may be used interchangeably. In addition, in this specification, 'left', 'upper left', 'upper left', 'upper right', 'right', 'lower right', 'lower left', and 'lower left' can be used interchangeably with 'left', 'upper left', 'top', 'upper right', 'right', 'lower right', 'bottom', and 'lower left'. In addition, element and member can be used interchangeably with each other. POC (Picture Order Count) represents temporal position information of a picture (or frame), can be the playback order displayed on the screen, and can have a unique POC for each picture. In addition, in this specification, the size of a block can be the sum or product of the horizontal length and the vertical length of the block. Alternatively, the size of a block can represent the number of samples in the block. Bit depth can be an expression of the range of sample values ​​in bit units. Specifically, if the bit depth is 8 bits, the range of sample values ​​can be from 0 to 255.The internal bit depth can represent the bit depth when the bit depth is expanded to effectively encode an image in a video signal processing device. The video signal processing device can expand the bit depth before encoding the image, and reduce the bit depth to the bit depth of the input original image when outputting after decoding. For example, the video signal processing device performs encoding and decoding by expanding an image with an 8-bit depth to an image with a 10-bit depth.

[0109] FIG. 1 is a schematic block diagram of a video signal encoding device (100) according to one embodiment of the present specification. Referring to FIG. 1, the encoding device (100) of the present invention includes a transform unit (110), a quantization unit (115), an inverse quantization unit (120), an inverse transform unit (125), a filtering unit (130), a prediction unit (150), and an entropy coding unit (160).

[0110] The transform unit (110) obtains a transform coefficient value by transforming the residual signal, which is the difference between the input video signal and the prediction signal generated by the prediction unit (150). For example, a discrete cosine transform (DCT), a discrete sine transform (DST), or a wavelet transform may be used. The discrete cosine transform and the discrete sine transform divide the input picture signal into blocks and perform the transformation. The coding efficiency may vary depending on the distribution and characteristics of the values ​​within the transform domain during the transformation. The transform kernel used for the transformation on the residual block may be a transform kernel having separable characteristics of vertical transformation and horizontal transformation. In this case, the transformation on the residual block may be performed separately as vertical transformation and horizontal transformation. For example, the encoder may perform vertical transformation by applying the transform kernel in the vertical direction of the residual block. Additionally, the encoder can perform horizontal transformation by applying a transformation kernel in the horizontal direction of the residual block. In the present disclosure, the transformation kernel may be used as a term referring to a set of parameters used for transformation of the residual signal, such as a transformation matrix, a transformation array, a transformation function, or a transformation. For example, the transformation kernel may be any one of a plurality of available kernels. Additionally, transformation kernels based on different transformation types may be used for each of the vertical transformation and the horizontal transformation.

[0111] The transformation coefficients are distributed in a way that increases toward the upper left corner of the block, and decreases toward 0 toward the lower right corner. As the current block size increases, there is a high probability of a significant number of 0 coefficients in the lower right corner. To reduce the transformation complexity of large blocks, the remaining regions can be reset to 0, leaving only the upper left corner as an arbitrary region.

[0112] Additionally, error signals may exist only in some regions of a coding block. In this case, the conversion process may be performed only on some arbitrary regions. For example, in a block of size 2Nx2N, an error signal may exist only in the first 2NxN block, and the conversion process may be performed only on the first 2NxN block, but the conversion process may not be performed on the second 2NxN block and may not be encoded or decoded. Here, N can be any positive integer.

[0113] The encoder may perform an additional transform before the transform coefficients are quantized. The aforementioned transform method may be referred to as a primary transform, and the additional transform may be referred to as a secondary transform. The secondary transform may be optional for each residual block. In one embodiment, the encoder may improve coding efficiency by performing the secondary transform on areas where it is difficult to concentrate energy in the low-frequency region using only the primary transform. For example, the secondary transform may be additionally performed on blocks where residual values ​​appear large in directions other than the horizontal or vertical direction of the residual block. Unlike the primary transform, the secondary transform may not be performed separately into a vertical transform and a horizontal transform. Such a secondary transform may be referred to as a Low Frequency Non-Separable Transform (LFNST).

[0114] The quantization unit (115) quantizes the transformation coefficient value output from the transformation unit (110).

[0115] In order to increase coding efficiency, rather than coding the picture signal as it is, a method is used to predict a picture using an already coded area through a prediction unit (150), and to obtain a restored picture by adding the residual value between the original picture and the predicted picture to the predicted picture. In order to prevent mismatches from occurring in the encoder and decoder, when the encoder performs prediction, information that is also available to the decoder must be used. To this end, the encoder performs a process of restoring the encoded current block. The inverse quantization unit (120) inversely quantizes the transform coefficient value, and the inverse transform unit (125) restores the residual value using the inverse quantized transform coefficient value. Meanwhile, the filtering unit (130) performs a filtering operation to improve the quality of the restored picture and enhance the coding efficiency. For example, a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter may be included. The filtered picture is stored in the Decoded Picture Buffer (DPB, 156) to be output or used as a reference picture.

[0116] A deblocking filter is a filter that removes distortion within blocks generated at the boundaries between blocks in a reconstructed picture. The encoder can determine whether to apply a deblocking filter to a certain boundary based on the distribution of pixels in several columns or rows based on an arbitrary boundary (edge) within the block. When applying a deblocking filter to a block, the encoder can apply a long filter, a strong filter, or a weak filter depending on the deblocking filtering strength. Additionally, horizontal and vertical filtering can be processed in parallel. Sample adaptive offset (SAO) can be used to correct the offset from the original image on a pixel-by-pixel basis for the residual block to which the deblocking filter has been applied. To correct the offset for a specific picture, the encoder can divide the pixels contained in the image into a certain number of regions, determine the regions to perform offset correction, and apply the offset to those regions (Band Offset). Alternatively, the encoder can use a method (Edge Offset) that applies an offset by considering the edge information of each pixel. CC-SAO (cross component SAO) is a method for compensating samples. In CC-SAO, the video signal processing device can classify the restored samples into categories similar to the existing SAO, derive an offset for each category, and add it to the restored samples. In the existing SAO, the video signal processing device uses only each luminance and chrominance component, but in CC-SAO, the video signal processing device classifies the category using all three components: luminance and two chrominance components. The adaptive loop filter (ALF) is a method that divides the pixels included in the image into a predetermined group, determines one filter to be applied to the group, and performs filtering differentially for each group.Information regarding whether to apply ALF can be signaled on a coding unit basis, and the shape and filter coefficients of the ALF filter to be applied can vary depending on each block. Furthermore, an ALF filter of the same shape (fixed shape) can be applied regardless of the characteristics of the target block. Bilateral Filter (BF) filtering is a method of applying a sample-by-sample offset derived from the difference in variation between neighboring samples.

[0117] The prediction unit (150) includes an intra prediction unit (152) and an inter prediction unit (154). The intra prediction unit (152) performs intra prediction within a current picture, and the inter prediction unit (154) performs inter prediction to predict the current picture using a reference picture stored in a decoded picture buffer (156). The intra prediction unit (152) performs intra prediction from reconstructed regions within the current picture and transfers intra encoding information to the entropy coding unit (160). The intra encoding information may include at least one of an intra prediction mode, an MPM (Most Probable Mode) flag, an MPM index, and information about a reference sample. The inter prediction unit (154) may be configured to include a motion estimation unit (154a) and a motion compensation unit (154b). The motion estimation unit (154a) refers to a specific area of ​​the restored reference picture to find the part most similar to the current area and obtains a motion vector value, which is the distance between the areas. The motion information (reference direction indication information (L0 prediction, L1 prediction, bidirectional prediction), reference picture index, motion vector information, etc.) for the reference area obtained by the motion estimation unit (154a) is transferred to the entropy coding unit (160) so that it can be included in the bitstream. Using the motion information transferred from the motion estimation unit (154a), the motion compensation unit (154b) performs inter-motion compensation to generate a prediction block for the current block. The inter-prediction unit (154) transfers inter-encoding information including motion information for the reference area to the entropy coding unit (160).

[0118] According to an additional embodiment, the prediction unit (150) may include an intra block copy (IBC) prediction unit (not shown). The IBC prediction unit performs IBC prediction on reconstructed samples in the current picture and transfers IBC encoding information to the entropy coding unit (160). The IBC prediction unit obtains a block vector value indicating a reference region used for prediction of the current region by referring to a specific region in the current picture. The IBC prediction unit may perform IBC prediction using the obtained block vector value. The IBC prediction unit transfers the IBC encoding information to the entropy coding unit (160). The IBC encoding information may include at least one of size information of the reference region, block vector information (index information for block vector prediction of the current block within a motion candidate list, and block vector difference information).

[0119] When the picture prediction as above is performed, the transformation unit (110) obtains a transformation coefficient value by transforming the residual value between the original picture and the predicted picture. At this time, the transformation can be performed in units of specific blocks within the picture, and the size of the specific block can be varied within a preset range. The quantization unit (115) quantizes the transformation coefficient value generated by the transformation unit (110) and transfers the quantized transformation coefficient to the entropy coding unit (160).

[0120] The quantized transform coefficients in the form of a two-dimensional array can be rearranged into a one-dimensional array for entropy coding. The method of scanning the quantized transform coefficients can be determined by which scanning method is used depending on the size of the transform block and the prediction mode within the screen. For example, diagonal, vertical, and horizontal scanning can be applied. This scanning information can be signaled on a block-by-block basis and can be derived according to predetermined rules.

[0121] The entropy coding unit (160) entropy-codes information representing quantized transform coefficients, intra-coding information, inter-coding information, etc. to generate a video signal bitstream. The entropy coding unit (160) may use a variable length coding (VLC) method and an arithmetic coding method. The variable length coding (VLC) method converts input symbols into continuous codewords, and the length of the codewords may be variable. For example, frequently occurring symbols are expressed as short codewords, and infrequently occurring symbols are expressed as long codewords. A context-based adaptive variable length coding (CAVLC) method may be used as a variable length coding method. Arithmetic coding converts continuous data symbols into a single prime number by using the probability distribution of each data symbol, and arithmetic coding can obtain the optimal prime number bits required to express each symbol. Context-based Adaptive Binary Arithmetic Coding (CABAC) can be used as an arithmetic coding.

[0122] CABAC is a binary arithmetic coding method that uses multiple context models generated based on experimentally obtained probabilities. The context models can also be referred to as context models. First, if the symbols are not in binary form, the encoder binarizes each symbol using exp-Golomb, etc. The binarized 0 or 1 can be described as a bin. The CABAC initialization process is divided into context initialization and arithmetic coding initialization. Context initialization is the process of initializing the occurrence probability of each symbol, and is determined by the symbol type, quantization parameter (QP), and slice type (I, P, B). A context model with this initialization information can use probability-based values ​​obtained through experiments. The context model provides the occurrence probability of the Least Probable Symbol (LPS) or Most Probable Symbol (MPS) for the symbol to be currently encoded, as well as information (valMPS) on which bin value corresponds to the MPS between 0 and 1. One of several context models is selected through the context index (ctxIdx), and the context index can be derived from information about the current block to be encoded or information about the surrounding blocks. Initialization for binary arithmetic coding is performed based on the probability model selected from the context model. Binary arithmetic coding is performed by dividing the data into probability intervals based on the occurrence probabilities of 0 and 1, and then encoding is performed through a process in which the probability interval corresponding to the bin to be processed becomes the entire probability interval for the bin to be processed next. The location information within the probability interval in which the last bin has been processed is output. However, since the probability interval cannot be divided infinitely, if it is reduced to a certain size, a renormalization process is performed to expand the probability interval and output the corresponding location information. In addition, after each bin is processed, a probability update process can be performed in which the probability for the next bin to be processed is newly set based on the information of the processed bin.

[0123] A bitstream may consist of one or more coded video sequences (CVSs), and a CVS may be encoded independently of other CVSs. Each CVS may consist of one or more layers, and each layer may represent a specific quality level, a specific resolution, or a general image, a depth map, or a transparency map. Furthermore, a CLVS may mean a layer-wise CVS composed of consecutive (in decoding order) PUs within the same layer. For example, there may be a CLVS for a specific quality layer, and there may be a CLVS for a depth map.

[0124] The above generated bitstream is encapsulated into NAL (Network Abstraction Layer) units as basic units. NAL units are divided into VCL (Video Coding Layer) NAL units containing video data and non-VCL NAL units containing parameter information for decoding video data, and there are various types of VCL or non-VCL NAL units. A NAL unit consists of NAL header information and RBSP (Raw Byte Sequence Payload) data, and the NAL header information includes summary information about the RBSP. The RBSP of a VCL NAL unit includes an integer number of encoded coding tree units. In order to decode a bitstream in a video decoder, the bitstream must first be divided into NAL unit units, and then each divided NAL unit must be decoded. Meanwhile, information required for decoding a video signal bitstream can be transmitted as included in a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), an Adaptation Parameter Set (APS), etc. An RBSP of a VCL NAL unit can include an integer number of coding tree units. A VPS is a parameter set configured with a common syntax by extracting duplicate parameters from an SPS parameter set signaled for each layer in a bitstream that supports image quality, resolution, and frame rate scalability or a bitstream that supports multi-view.SPS is a parameter set that includes at least one of the following: Profile, which contains information about acceptable coding tools (or algorithms) and video formats; Level, which contains information about the decoder's processing capabilities, such as the resolution and frame rate of processable video and the allowable memory size; Tier, which contains information about the maximum bit rate that can be processed; and information about the resolution, bit depth, and whether or not a function can be enabled. PPS is a parameter set that includes at least one of the following: resolution of the video, tile division information, whether or not to enable weight prediction, quantization parameters, and filtering-related information. APS is a parameter set that includes one of the following: ALF filter coefficient information, LMCS-related parameters, and quantization scale parameters, depending on the APS type. APS is divided into prefix APS, which is signaled before the VCL NAL unit, and suffix APS, which is signaled after the VCL NAL unit. In the case of ALF APS, it is efficient to apply the ALF filter coefficients derived from the previous picture to the next picture, so it can be signaled as suffix APS.

[0125] Meanwhile, the block diagram of FIG. 1 illustrates an encoding device (100) according to one embodiment of the present specification, and the blocks shown separately logically distinguish elements of the encoding device (100). Accordingly, the elements of the encoding device (100) described above may be mounted as one chip or as multiple chips depending on the design of the device. According to one embodiment, the operations of each element of the encoding device (100) described above may be performed by a processor (not shown).

[0126] Fig. 2 is a schematic block diagram of a video signal decoding device (200) according to one embodiment of the present specification. Referring to Fig. 2, the decoding device (200) of the present invention includes an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (225), a filtering unit (230), and a prediction unit (250).

[0127] The entropy decoding unit (210) entropy decodes the video signal bitstream to extract transform coefficient information, intra-coding information, inter-coding information, etc. for each region. For example, the entropy decoding unit (210) can obtain a binarization code for transform coefficient information of a specific region from the video signal bitstream. In addition, the entropy decoding unit (210) inversely binarizes the binarization code to obtain a quantized transform coefficient. The inverse quantization unit (220) inversely quantizes the quantized transform coefficient, and the inverse transform unit (225) restores the residual value using the inverse quantized transform coefficient. The video signal processing device (200) restores the original pixel value by adding the residual value obtained by the inverse transform unit (225) and the prediction value obtained by the prediction unit (250).

[0128] Meanwhile, the filtering unit (230) performs filtering on the picture to improve the image quality. This may include a deblocking filter to reduce block distortion and / or an adaptive loop filter to remove distortion from the entire picture. The filtered picture is output or stored in the decoded picture buffer (DPB, 256) to be used as a reference picture for the next picture.

[0129] The prediction unit (250) includes an intra prediction unit (252) and an inter prediction unit (254). The prediction unit (250) generates a prediction picture by utilizing the encoding type decoded through the entropy decoding unit (210) described above, the transform coefficients for each region, intra / inter encoding information, etc. In order to restore the current block on which decoding is performed, the decoded region of the current picture or other pictures including the current block may be used. A picture (or tile / slice) that uses only the current picture for restoration, i.e., performs intra prediction or intra BC prediction, is called an intra picture or I picture (or tile / slice), and a picture (or tile / slice) that can perform all of intra prediction, inter prediction, and intra BC prediction is called an inter picture (or tile / slice). A picture (or tile / slice) that uses at most one motion vector and reference picture index to predict sample values ​​of each block among inter-pictures (or tiles / slices) is called a predictive picture or P-picture (or tile / slice), and a picture (or tile / slice) that uses at most two motion vectors and reference picture indices is called a bi-predictive picture or B-picture (or tile / slice). In other words, a P-picture (or tile / slice) uses at most one motion information set to predict each block, and a B-picture (or tile / slice) uses at most two motion information sets to predict each block. Here, a motion information set includes one or more motion vectors and one reference picture index.

[0130] The intra prediction unit (252) generates a prediction block using intra encoding information and reconstructed samples within the current picture. Specifically, a sample within the current block can be predicted using a reference sample derived using the sample position within the current block and the directionality of the intra prediction mode. If the position of the reference sample is not an integer unit sample, the video signal processing device can predict the current block sample using an interpolated reference sample through an interpolation method using adjacent reference samples. This may be referred to as linear-based intra prediction. As described above, the intra encoding information may include at least one of an intra prediction mode, a Most Probable Mode (MPM) flag, and an MPM index. The intra prediction unit (252) predicts sample values ​​of the current block using reconstructed samples located on the left and / or above the current block as reference samples. In the present disclosure, the reconstructed samples, the reference samples, and the samples of the current block may represent pixels. Additionally, the sample values ​​may represent pixel values.

[0131] According to one embodiment, the reference samples may be samples included in a neighboring block of the current block. For example, the reference samples may be samples adjacent to the left boundary of the current block and / or samples adjacent to the upper boundary. In addition, the reference samples may be samples located on a line within a preset distance from the left boundary of the current block and / or samples located on a line within a preset distance from the upper boundary of the current block among the samples of the neighboring blocks of the current block. In this case, the neighboring blocks of the current block may include at least one of a left (L) block, an upper (A) block, a lower left (BL) block, an above right (AR) block, or an above left (AL) block adjacent to the current block. The neighboring blocks of the current block may be reference blocks for predicting the current block.

[0132] The inter prediction unit (254) generates a prediction block using the reference picture and inter encoding information stored in the decoded picture buffer (256). The inter encoding information may include a set of motion information (reference picture index, motion vector information, etc.) of the current block with respect to the reference block. Inter prediction may include L0 prediction, L1 prediction, and bi-prediction. L0 prediction is prediction using one reference picture included in the L0 picture list, and L1 prediction means prediction using one reference picture included in the L1 picture list. For this, one set of motion information (e.g., motion vector and reference picture index) may be required. In the bi-prediction method, up to two reference areas can be used, and these two reference areas may exist in the same reference picture or may exist in different pictures, respectively. That is, in the bi-prediction method, up to two sets of motion information (e.g., motion vectors and reference picture indices) can be used, and the two motion vectors may correspond to the same reference picture index or may correspond to different reference picture indices. At this time, the reference pictures are pictures that are located temporally before or after the current picture, and may be completed pictures that have already been restored. According to one embodiment, the two reference areas used in the bi-prediction method may be areas selected from each of the L0 picture list and the L1 picture list. In addition, a prediction method that uses only reference pictures having a POC smaller than the POC of the current picture or uses only reference pictures having a POC larger than the POC of the current picture based on the POC (picture order count) indicating the display order of the current picture can be called uni-directional prediction.Also, a prediction method that uses both a reference picture with a POC (picture order count) smaller than that of the current picture and a reference picture with a POC larger than that of the current picture based on the POC indicating the display order of the current picture can be called bi-directional prediction. A prediction method that uses only one reference picture in uni-directional prediction can be called uni-prediction, and a prediction method that uses two reference pictures in uni-directional prediction can be called bi-prediction or bi-prediction.

[0133] The inter prediction unit (254) can obtain a reference block of the current block using a motion vector and a reference picture index. The reference block exists in a reference picture corresponding to the reference picture index. In addition, a sample value of a block specified by the motion vector or an interpolated value thereof can be used as a predictor of the current block. For motion prediction with sub-pel unit pixel accuracy, for example, an 8-tap interpolation filter can be used for a luminance signal and a 4-tap interpolation filter can be used for a chrominance signal. However, the interpolation filter for sub-pel unit motion prediction is not limited thereto. In this way, the inter prediction unit (254) performs motion compensation to predict the texture of the current unit from a previously restored picture. At this time, the inter prediction unit can use a motion information set.

[0134] According to an additional embodiment, the prediction unit (250) may include an IBC prediction unit (not shown). The IBC prediction unit may reconstruct the current region by referring to a specific region including reconstructed samples within the current picture. The IBC prediction unit may perform IBC prediction using IBC encoding information obtained from the entropy decoding unit (210). The IBC encoding information may include block vector information.

[0135] A restored video picture is generated by adding the predicted value output from the intra prediction unit (252) or inter prediction unit (254) and the residual value output from the inverse transformation unit (225). That is, the video signal decoding device (200) restores the current block using the predicted block generated from the prediction unit (250) and the residual obtained from the inverse transformation unit (225).

[0136] Meanwhile, the block diagram of FIG. 2 illustrates a decoding device (200) according to one embodiment of the present specification, and the blocks shown separately illustrate logically distinguishing elements of the decoding device (200). Accordingly, the elements of the aforementioned decoding device (200) may be mounted as one chip or as multiple chips depending on the design of the device. According to one embodiment, the operations of each element of the aforementioned decoding device (200) may be performed by a processor (not shown).

[0137] Meanwhile, the technology proposed in this specification is applicable to both the methods and devices of the encoder and decoder, and the parts described as signaling and parsing may be described for convenience of explanation. In general, signaling can be described as encoding each syntax from the encoder's perspective, and parsing can be described as interpreting each syntax from the decoder's perspective. That is, each syntax can be included in the bitstream from the encoder and signaled, and the decoder can parse the syntax and use it in the restoration process. At this time, the sequence of bits for each syntax listed in the prescribed hierarchical structure can be referred to as a bitstream.

[0138] A picture can be encoded by dividing it into sub-pictures, slices, tiles, etc. A sub-picture can include one or more slices or tiles. When a picture is encoded by dividing it into multiple slices or tiles, all slices or tiles within the picture must be decoded before it can be displayed on the screen. On the other hand, when a picture is encoded into multiple sub-pictures, only any sub-picture can be decoded and displayed on the screen. A slice can include multiple tiles or sub-pictures, or a tile can include multiple sub-pictures or slices. Sub-pictures, slices, and tiles can be encoded or decoded independently, which is effective for parallel processing and improving processing speed. However, there is a disadvantage in that the amount of bits increases because the encoded information of adjacent sub-pictures, slices, and tiles cannot be used. Sub-pictures, slices, and tiles can be encoded by dividing them into multiple coding tree units (CTUs).

[0139] FIG. 3 illustrates an embodiment in which a Coding Tree Unit (CTU) within a picture is divided into Coding Units (CUs). In the process of coding a video signal, a picture may be divided into a sequence of Coding Tree Units (CTUs). A Coding Tree Unit may be composed of a luminance (luma) Coding Tree Block (CTB), two chroma (chroma) Coding Tree Blocks, and their encoded syntax information. One Coding Tree Unit may be composed of one Coding Unit, or one Coding Tree Unit may be split into multiple Coding Units. One Coding Unit may be composed of a luminance Coding Block (CB), two chroma Coding Blocks, and their encoded syntax information. One Coding Block may be split into multiple Sub-Coding Blocks. One Coding Unit may be composed of one Transform Unit (TU), or one Coding Unit may be split into multiple Transform Units. A single transform unit may consist of a luminance transform block (TB), two chrominance transform blocks, and their encoded syntax information. A coding tree unit may be divided into multiple coding units. A coding tree unit may not be divided and may also be a leaf node. In this case, the coding tree unit itself may be a coding unit.

[0140] A coding unit refers to a basic unit for processing a picture in the video signal processing described above, i.e., intra / inter prediction, transformation, quantization, and / or entropy coding. The size and shape of a coding unit within a picture may not be constant. A coding unit may have a square or rectangular shape. A rectangular coding unit (or rectangular block) includes a vertical coding unit (or vertical block) and a horizontal coding unit (or horizontal block). In the present specification, a vertical block is a block whose height is greater than its width, and a horizontal block is a block whose width is greater than its height. In addition, a non-square block in the present specification may refer to a rectangular block, but the present invention is not limited thereto.

[0141] Referring to Fig. 3, the coding tree unit is first partitioned into a Quad Tree (QT) structure. That is, in the Quad Tree structure, one node with a size of 2NX2N can be partitioned into four nodes with a size of NXN. In this specification, the Quad Tree may also be referred to as a quaternary tree. The Quad Tree partitioning can be performed recursively, and not all nodes need to be partitioned to the same depth.

[0142] Meanwhile, the leaf node of the aforementioned quad tree can be further split into a multi-type tree (MTT) structure. According to an embodiment of the present invention, in the multi-type tree structure, one node can be split into a binary or ternary tree structure of horizontal or vertical split. That is, the multi-type tree structure has four split structures: vertical binary split, horizontal binary split, vertical ternary split, and horizontal ternary split. According to an embodiment of the present invention, in each of the above tree structures, the width and height of the node can both have a power of 2 value. For example, in the binary tree (BT) structure, a node of size 2NX2N can be split into two NX2N nodes by vertical binary split, and can be split into two 2NXN nodes by horizontal binary split. Also, in a Ternary Tree (TT) structure, a node of size 2NX2N can be split into nodes of size (N / 2)X2N, NX2N, and (N / 2)X2N by vertical ternary splitting, and into nodes of size 2NX(N / 2), 2NXN, and 2NX(N / 2) by horizontal ternary splitting. This multi-type tree splitting can be performed recursively.

[0143] A leaf node of a multi-type tree can be a coding unit. If the coding unit is not larger than the maximum transformation length, the coding unit can be used as a unit of prediction and / or transformation without further splitting. In one embodiment, if the width or height of the current coding unit is larger than the maximum transformation length, the current coding unit can be split into multiple transformation units without explicit signaling regarding the splitting. Meanwhile, in the quad tree and multi-type tree described above, at least one of the following parameters can be predefined or transmitted through an RBSP of a higher-level set, such as a PPS, an SPS, or a VPS. 1) CTU size: The size of the root node of the quad tree, 2) MinQtSize: The minimum allowed QT leaf node size, 3) MaxBtSize: The maximum allowed BT root node size, 4) MaxTT size (MaxTtSize): The maximum allowed TT root node size, 5) MaxMttDepth: The maximum allowed depth of an MTT split from a leaf node of a QT, 6) MinBT size (MinBtSize): The minimum allowed BT leaf node size, 7) MinTT size (MinTtSize): The minimum allowed TT leaf node size.

[0144] Fig. 4 illustrates one embodiment of a method for signaling splitting of a quad tree and a multi-type tree. Pre-configured flags may be used to signal splitting of the quad tree and the multi-type tree described above. Referring to Fig. 4, at least one of a flag 'split_cu_flag' indicating whether a node is split, a flag 'split_qt_flag' indicating whether a quad tree node is split, a flag 'mtt_split_cu_vertical_flag' indicating a splitting direction of a multi-type tree node, or a flag 'mtt_split_cu_binary_flag' indicating a splitting shape of a multi-type tree node may be used.

[0145] According to an embodiment of the present invention, a flag 'split_cu_flag' indicating whether a current node is split may be signaled first. If the value of 'split_cu_flag' is 0, it indicates that the current node is not split, and the current node becomes a coding unit. If the current node is a coding tree unit, the coding tree unit includes one coding unit that is not split. If the current node is a quad tree node 'QT node', the current node is a leaf node 'QT leaf node' of the quad tree and becomes a coding unit. If the current node is a multi-type tree node 'MTT node', the current node is a leaf node 'MTT leaf node' of the multi-type tree and becomes a coding unit.

[0146] When the value of 'split_cu_flag' is 1, the current node can be split into nodes of a quad tree or a multi-type tree depending on the value of 'split_qt_flag'. The coding tree unit is the root node of the quad tree and can be first split into a quad tree structure. In the quad tree structure, 'split_qt_flag' is signaled for each node 'QT node'. When the value of 'split_qt_flag' is 1, the node is split into four square nodes, and when the value of 'split_qt_flag' is 0, the node becomes a leaf node 'QT leaf node' of the quad tree, and the node is split into multi-type nodes. According to an embodiment of the present invention, quad tree splitting can be limited depending on the type of the current node. Quad tree splitting may be allowed if the current node is a coding tree unit (root node of a quad tree) or a quad tree node, and quad tree splitting may not be allowed if the current node is a multi-type tree node. Each quad tree leaf node 'QT leaf node' may be further split into a multi-type tree structure. As described above, if 'split_qt_flag' is 0, the current node may be split into multi-type nodes. To indicate the splitting direction and splitting shape, 'mtt_split_cu_vertical_flag' and 'mtt_split_cu_binary_flag' may be signaled. If the value of 'mtt_split_cu_vertical_flag' is 1, a vertical split of the node 'MTT node' is indicated, and if the value of 'mtt_split_cu_vertical_flag' is 0, a horizontal split of the node 'MTT node' is indicated.Additionally, if the value of 'mtt_split_cu_binary_flag' is 1, the node 'MTT node' is split into two rectangular nodes, and if the value of 'mtt_split_cu_binary_flag' is 0, the node 'MTT node' is split into three rectangular nodes.

[0147] The tree partitioning structure allows luminance blocks and chrominance blocks to be partitioned in the same manner. That is, the chrominance blocks can be partitioned by referring to the partitioning form of the luminance blocks. If the current chrominance block is smaller than a predetermined size, the chrominance block may not be partitioned even if the luminance block is partitioned.

[0148] The luminance block and the chrominance block may have the same tree partitioning structure, which may be referred to as a single tree. If the current block is encoded with a single tree, the partitioning structure, encoding mode information, motion information, etc. of the luminance block and the chrominance block may be the same, and information related to other error signals may be different between the luminance block and the chrominance block. In addition, the luminance block and the chrominance block may have different tree partitioning structures, which may be referred to as a dual tree. If the current block is encoded with a dual tree, the partitioning structure, encoding mode information, motion information, etc. of the luminance block and the chrominance block may be different in at least one or more.

[0149] There may be a close correlation between a luminance block and its corresponding chrominance block. Therefore, when the current block is encoded and decoded using a dual tree, the encoder and decoder can use the segmentation information, encoding mode information, and motion information of the luminance block to encode the chrominance block.

[0150] A node to be divided into the smallest unit can be processed as a single coding block. If the current block is a coding block, the coding block can be divided into multiple sub-blocks (sub-coding blocks), and the prediction information of each sub-block can be the same or different. For example, if the coding unit is an intra mode, the intra prediction modes of each sub-block can be the same or different. Furthermore, if the coding unit is an inter mode, the motion information of each sub-block can be the same or different. Furthermore, each sub-block can be encoded or decoded independently. Each sub-block can be distinguished by a sub-block index (sbIdx). Furthermore, when a coding unit is divided into sub-blocks, it can be divided horizontally, vertically, or diagonally. In intra mode, the mode that divides the current coding unit into two or four sub-blocks horizontally or vertically is called ISP (Intra Sub-Partitions). In inter mode, the mode that divides the current coding block diagonally is called GPM (Geometric partitioning mode). In GPM mode, the position and direction of the diagonal line are derived using a predefined angle table, and the index information of the angle table is signaled.

[0151] The motion information may include one or more of reference direction indication information, reference picture information, motion vector, motion resolution, affine model, CPMV (control point motion vector), block vector, block vector resolution, MHP information, LIC information, filtering information, BCW information, and RRIBC information.

[0152] The reference direction indication information is composed of L0 prediction, L1 prediction, L0 and L1 prediction, and L0 prediction and L1 prediction are uni-prediction and unidirectional prediction, and L0 and L1 prediction are bi-prediction. And L0 and L1 prediction can be uni-prediction or bi-directional prediction. Here, L0 prediction is predicted using reference pictures in the L0 reference picture list, and L1 prediction can be predicted using reference pictures in the L1 reference picture list. In the L0 reference picture list, a reference picture with a POC smaller than the POC of the current picture can be added to the reference picture list based on the POC of the current picture. In addition, the L0 reference picture list can be composed in order from a reference picture with a POC closer to the POC of the current picture to a reference picture with a POC farther from the POC. In the L1 reference picture list, a reference picture with a POC larger than the POC of the current picture can be added to the reference picture list based on the POC of the current picture. Additionally, the L1 reference picture list can be organized in the order of reference pictures with a POC closer to the POC of the current picture to reference pictures with a POC farther away. The L0 and L1 reference picture lists can vary for each slice, subpicture, and picture. Additionally, the L0 reference picture list can include reference pictures of the L1 reference picture list. And the L1 reference picture list can include reference pictures of the L0 reference picture list.

[0153] Reference picture information may vary for each block and may be index information indicating which reference picture in the L0 reference picture list and / or the L1 reference picture list is used to predict the current block. The reference picture information may include one or more of L0 reference picture information and L1 reference picture information.

[0154] A motion vector is information that indicates the block that best matches the current block in a reference picture. It is a value that represents the distance to the reference block in horizontal and vertical coordinates based on the upper left position of the current block in the picture.

[0155] Motion resolution represents the resolution of the motion vector, and motion resolution can be expressed in units of 4 pixels, 1 pixel (integer pixel), 1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel, 1 / 16 pixel, etc.

[0156] A block vector is information indicating the block that best matches the current block in the already restored area within the current picture. It is a value that represents the distance to the reference block in horizontal and vertical coordinates based on the upper left position of the current block within the picture.

[0157] Block vector resolution can be expressed in units of 4 pixels, 1 pixel (integer pixel), 1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel, 1 / 16 pixel, etc.

[0158] MHP information may include whether additional motion information is applied and additional motion information.

[0159] LIC information may include whether LIC applies to the current block.

[0160] Filtering information may include filtering type and coefficient information applied according to the motion resolution of the current block.

[0161] BCW information may include whether BCW is applied to the current block.

[0162] RRIBC information may include information about whether RRIBC is applied and the RRIBC type if the current block is encoded in IBC mode.

[0163] Picture prediction (motion compensation) for coding is performed on coding units that are no longer divisible (i.e., leaf nodes of a coding tree unit). The basic unit performing this prediction is referred to below as a prediction unit or prediction block.

[0164] Hereinafter, the term "unit" used in this specification may be used as a replacement for the prediction unit, which is the basic unit for performing prediction. However, the present invention is not limited thereto, and can be understood more broadly as a concept that includes the coding unit.

[0165] Figures 5 and 6 illustrate an intra prediction method according to an embodiment of the present invention in more detail. As described above, the intra prediction unit predicts sample values ​​of the current block using restored samples located to the left and / or above the current block as reference samples.

[0166] First, FIG. 5 illustrates an embodiment of reference samples used for prediction of a current block in an intra prediction mode. According to an embodiment, the reference samples may be samples adjacent to a left boundary of the current block and / or samples adjacent to an upper boundary. As illustrated in FIG. 5, when the size of the current block is WXH and samples of a single reference line adjacent to the current block are used for intra prediction, reference samples may be set using at most 2W+2H+1 surrounding samples located on the left and / or upper sides of the current block.

[0167] Meanwhile, pixels of multiple reference lines may be used for intra prediction of the current block. The multiple reference lines may be composed of n lines located within a preset range from the current block. According to one embodiment, when pixels of multiple reference lines are used for intra prediction, separate index information indicating the lines to be set as reference pixels may be signaled, and this may be referred to as a reference line index.

[0168] In addition, if at least some of the samples to be used as reference samples have not yet been restored, the intra prediction unit may perform a reference sample padding process to obtain a reference sample. In addition, the intra prediction unit may perform a reference sample filtering process to reduce an error in intra prediction. That is, filtering may be performed on the surrounding samples and / or the reference samples obtained by the reference sample padding process to obtain filtered reference samples. The intra prediction unit predicts samples of the current block using the reference samples obtained in this manner. The intra prediction unit predicts samples of the current block using unfiltered reference samples or filtered reference samples. In the present disclosure, the surrounding samples may include samples on at least one reference line. For example, the surrounding samples may include adjacent samples on a line adjacent to a boundary of the current block.

[0169] Next, FIG. 6 illustrates an embodiment of prediction modes used for intra prediction. For intra prediction, intra prediction mode information indicating an intra prediction direction may be signaled. The intra prediction mode information indicates any one of a plurality of intra prediction modes constituting an intra prediction mode set. If the current block is an intra prediction block, the decoder receives intra prediction mode information of the current block from the bitstream. The intra prediction unit of the decoder performs intra prediction on the current block based on the extracted intra prediction mode information.

[0170] According to an embodiment of the present invention, an intra prediction mode set may include all intra prediction modes used for intra prediction (e.g., a total of 67 intra prediction modes). More specifically, the intra prediction mode set may include a planar mode, a DC mode, and a plurality of (e.g., 65) angular modes (i.e., directional modes). Each intra prediction mode may be indicated by a preset index (i.e., an intra prediction mode index). For example, as illustrated in FIG. 6, an intra prediction mode index 0 indicates a planar mode, and an intra prediction mode index 1 indicates a DC mode. In addition, intra prediction mode indices 2 to 66 may each indicate different angular modes. The angular modes each indicate different angles within a preset angular range. For example, an angular mode may indicate an angle within an angular range from 45 degrees to -135 degrees clockwise (i.e., a first angular range). The angular mode may be defined based on the 12 o'clock direction. At this time, intra prediction mode index 2 indicates horizontal diagonal (HDIA) mode, intra prediction mode index 18 indicates horizontal (HOR) mode, intra prediction mode index 34 indicates diagonal (DIA) mode, intra prediction mode index 50 indicates vertical (VER) mode, and intra prediction mode index 66 indicates vertical diagonal (VDIA) mode.

[0171] Meanwhile, the preset angle range may be set differently depending on the shape of the current block. For example, if the current block is a rectangular block, a wide-angle mode indicating an angle exceeding 45 degrees or less than -135 degrees in a clockwise direction may be additionally used. If the current block is a horizontal block, the angle mode may indicate an angle within an angle range (i.e., a second angle range) between (45+offset1) degrees and (-135+offset1) degrees in a clockwise direction. At this time, angle modes 67 to 76 that are outside the first angle range may be additionally used. In addition, if the current block is a vertical block, the angle mode may indicate an angle within an angle range (i.e., a third angle range) between (45-offset2) degrees and (-135-offset2) degrees in a clockwise direction. At this time, angle modes -10 to -1 that are outside the first angle range may be additionally used. According to an embodiment of the present invention, the values ​​of offset1 and offset2 may be determined differently depending on the ratio between the width and height of the rectangular block. Also, offset1 and offset2 can be positive.

[0172] According to an additional embodiment of the present invention, the plurality of angular modes constituting the intra prediction mode set may include a basic angular mode and an extended angular mode. In this case, the extended angular mode may be determined based on the basic angular mode.

[0173] According to one embodiment, the basic angle mode may be a mode corresponding to an angle used in intra prediction of the existing HEVC (High Efficiency Video Coding) standard, and the extended angle mode may be a mode corresponding to an angle newly added in intra prediction of the next-generation video codec standard. More specifically, the basic angle mode may be an angle mode corresponding to any one of the intra prediction modes {2, 4, 6, … , 66}, and the extended angle mode may be an angle mode corresponding to any one of the intra prediction modes {3, 5, 7, … , 65}. That is, the extended angle mode may be an angle mode between the basic angle modes within the first angle range. Therefore, the angle indicated by the extended angle mode may be determined based on the angle indicated by the basic angle mode.

[0174] According to another embodiment, the basic angle mode may be a mode corresponding to an angle within a preset first angle range, and the extended angle mode may be a wide-angle mode outside the first angle range. That is, the basic angle mode may be an angle mode corresponding to any one of the intra prediction modes {2, 3, 4, … , 66}, and the extended angle mode may be an angle mode corresponding to any one of the intra prediction modes {-14, –13, –12, … , –1} and {67, 68, … , 80}. The angle indicated by the extended angle mode may be determined as an opposite angle to the angle indicated by the corresponding basic angle mode. Accordingly, the angle indicated by the extended angle mode may be determined based on the angle indicated by the basic angle mode. Meanwhile, the number of extended angle modes is not limited thereto, and additional extended angles may be defined according to the size and / or shape of the current block. Meanwhile, the total number of intra prediction modes included in the intra prediction mode set may vary depending on the configuration of the basic angle mode and the extended angle mode described above.

[0175] In the above embodiment, the spacing between the extended angular modes may be set based on the spacing between the corresponding basic angular modes. For example, the spacing between the extended angular modes {3, 5, 7, … , 65} may be determined based on the spacing between the corresponding basic angular modes {2, 4, 6, … , 66}. In addition, the spacing between the extended angular modes {-14, -13, … , -1} may be determined based on the spacing between the corresponding opposite basic angular modes {53, 53, … , 66}, and the spacing between the extended angular modes {67, 68, … , 80} may be determined based on the spacing between the corresponding opposite basic angular modes {2, 3, 4, … , 15}. The angular spacing between the extended angular modes may be set to be equal to the angular spacing between the corresponding basic angular modes. In addition, the number of the extended angular modes in the intra prediction mode set may be set to be less than or equal to the number of the basic angular modes.

[0176] According to an embodiment of the present invention, an extended angular mode may be signaled based on a base angular mode. For example, a wide-angle mode (i.e., an extended angular mode) may replace at least one angular mode (i.e., a base angular mode) within a first angular range. The replaced base angular mode may be an angular mode corresponding to an opposite side of the wide-angle mode. That is, the replaced base angular mode is an angular mode corresponding to an angle in an opposite direction to an angle indicated by the wide-angle mode or an angle that is different from the angle in the opposite direction by a preset offset index. According to an embodiment of the present invention, the preset offset index is 1. An intra prediction mode index corresponding to the replaced base angular mode may be remapped to the wide-angle mode to signal the corresponding wide-angle mode. For example, the wide-angle modes {-14, -13, … , -1} may be signaled by intra prediction mode indices {52, 53, … , 66}, and the wide-angle modes {67, 68, … , 69} may be signaled by intra prediction mode indices {52, 53, … , 66}, respectively. , 80} can be signaled by intra prediction mode indices {2, 3, … , 15}, respectively. By having the intra prediction mode index for the basic angular mode signal the extended angular mode in this way, even if the configurations of the angular modes used for intra prediction of each block are different, the same set of intra prediction mode indices can be used to signal the intra prediction mode. Therefore, the signaling overhead due to changes in the intra prediction mode configuration can be minimized.

[0177] Meanwhile, whether to use the extended angle mode may be determined based on at least one of the shape and size of the current block. In one embodiment, if the size of the current block is larger than a preset size, the extended angle mode may be used for intra prediction of the current block, and otherwise, only the default angle mode may be used for intra prediction of the current block. In another embodiment, if the current block is a non-square block, the extended angle mode may be used for intra prediction of the current block, and if the current block is a square block, only the default angle mode may be used for intra prediction of the current block.

[0178] The intra prediction unit determines reference samples and / or interpolated reference samples to be used for intra prediction of the current block based on intra prediction mode information of the current block. If the intra prediction mode index indicates a specific angular mode, the reference sample or interpolated reference sample corresponding to the specific angle from the current sample of the current block is used for prediction of the current pixel. Therefore, different sets of reference samples and / or interpolated reference samples can be used for intra prediction depending on the intra prediction mode. After intra prediction of the current block is performed using the reference samples and intra prediction mode information, the decoder restores the sample values ​​of the current block by adding the residual signal of the current block obtained from the inverse transform unit to the intra prediction value of the current block.

[0179] Motion information used for inter prediction may include reference direction indication information (inter_pred_idc), reference picture indexes (ref_idx_l0, ref_idx_l1), and motion vectors (mvL0, mvL1). Reference picture list utilization information (predFlagL0, predFlagL1) may be set according to the reference direction indication information. As an example, in case of unidirectional prediction using an L0 reference picture, predFlagL0=1, predFlagL1=0 may be set. In case of unidirectional prediction using an L1 reference picture, predFlagL0=0, predFlagL1=1 may be set. In case of bidirectional prediction using both L0 and L1 reference pictures, predFlagL0=1, predFlagL1=1 may be set.

[0180] If the current block is a coding unit, the coding unit can be divided into multiple sub-blocks, and the prediction information of each sub-block can be the same or different. For example, if the coding unit is an intra mode, the intra prediction modes of each sub-block can be the same or different. In addition, if the coding unit is an inter mode, the motion information of each sub-block can be the same or different. In addition, each sub-block can be encoded or decoded independently. Each sub-block can be distinguished through a sub-block index (sbIdx).

[0181] The motion vector of the current block is likely to be similar to that of the surrounding blocks. Therefore, the motion vectors of the surrounding blocks can be used as motion vector predictors (mvp), and the motion vector of the current block can be derived using the motion vectors of the surrounding blocks. Furthermore, to improve the accuracy of the motion vector, the difference in motion vectors (mvd) between the optimal motion vector of the current block, found from the original image by the encoder, and the motion predictor can be signaled.

[0182] Motion vectors can have various resolutions, and the resolution of motion vectors can vary on a block-by-block basis. Motion vector resolution can be expressed in integer units, half-pixel units, quarter-pixel units, sixteenth-pixel units, and integer-4 pixel units. Since images such as screen content are in simple graphical forms such as characters, interpolation filters do not need to be applied, integer units and integer-4 pixel units can be selectively applied on a block-by-block basis. Blocks encoded in affine mode, which can express rotation and scale, have significant shape changes, so integer units, quarter-pixel units, and sixteenth-pixel units can be selectively applied on a block-by-block basis. Information on whether to selectively apply motion vector resolution on a block-by-block basis is signaled with amvr_flag. If applicable, which motion vector resolution to apply to the current block is signaled with amvr_precision_idx.

[0183] For blocks to which bidirectional prediction is applied, the weights between the two prediction blocks can be applied equally or differently when applying weighted average, and information about the weights is signaled through bcw_idx.

[0184] To improve the accuracy of the motion prediction value, Merge or advanced motion vector prediction (AMVP) methods can be selectively used on a block-by-block basis. The Merge method configures the motion information of the current block to be identical to the motion information of the adjacent blocks to the current block, and has the advantage of increasing the encoding efficiency of the motion information by spatially propagating the motion information without change in the homogeneous motion region. On the other hand, the AMVP method predicts motion information in the L0 and L1 prediction directions respectively to express accurate motion information and signals the most optimal motion information. After the decoder derives the motion information for the current block through the AMVP or Merge method, it uses the reference block located in the motion information derived from the reference picture as the prediction block for the current block.

[0185] A method for deriving motion information in Merge or AMVP may be to construct a motion candidate list using motion prediction values ​​derived from neighboring blocks of the current block, and then signal the index information for the optimal motion candidate. In the case of AMVP, since motion candidate lists are derived for each of L0 and L1, the optimal motion candidate indices (mvp_l0_flag, mvp_l1_flag) for each of L0 and L1 are signaled. In the case of Merge, since one motion candidate list is derived, one merge index (merge_idx) is signaled. The motion candidate lists derived from one coding unit may vary, and a motion candidate index or merge index may be signaled for each motion candidate list. In this case, a mode in which there is no information about the remaining blocks in a block encoded in Merge mode can be referred to as MergeSkip mode.

[0186] Bidirectional motion information for the current block can be derived using a combination of AMVP and Merge modes. For example, motion information in the L0 direction can be derived using AMVP, while motion information in the L1 direction can be derived using Merge. Conversely, Merge can be applied to L0, while AMVP can be applied to L1. This encoding mode can be referred to as AMVP-merge mode.

[0187] The motion candidates and motion information candidates in this specification may have the same meaning. Furthermore, the motion candidate list and motion information candidate list in this specification may have the same meaning.

[0188] SMVD (Symmetric MVD) is a method for reducing the bit volume of transmitted motion information by ensuring that the MVD (Motion Vector Difference) values ​​in the L0 and L1 directions are symmetrical in the case of bidirectional prediction. The MVD information in the L1 direction, which is symmetrical to the L0 direction, is not transmitted, and reference picture information in the L0 and L1 directions is also not transmitted and can be derived during the decoding process.

[0189] OBMC (Overlapped Block Motion Compensation) is a method that generates prediction blocks for the current block using motion information from surrounding blocks when the motion information between blocks differs. It then weights and averages these prediction blocks to generate a final prediction block for the current block. This method effectively reduces blocking artifacts that occur at block boundaries in motion-compensated images.

[0190] Typically, merge motion candidates have low motion accuracy. To improve the accuracy of these merge motion candidates, the MMVD (Merge mode with MVD) method can be used. The MMVD method is a method of compensating motion information using one candidate selected from several motion difference value candidates. Information about the compensation value of the motion information obtained through the MMVD method (e.g., an index indicating one candidate selected from among the motion difference value candidates) can be included in the bitstream and transmitted to the decoder. Compared to the conventional method of including the motion information difference value in the bitstream, the inclusion of information about the compensation value of the motion information in the bitstream can save bits.

[0191] MBVD (merge mode with block vector differences) mode is a mode that encodes the differential value for a block vector, similar to MMVD, which is a mode that encodes the differential value for a motion vector in merge mode. The MBVD method is a method that determines a block vector by using one candidate selected from several block vector difference value candidates. Information about the differential value of the block vector obtained through the MBVD method (e.g., an index indicating one candidate selected from among the block vector difference value candidates) can be included in the bitstream and transmitted to the decoder. Compared to including the existing block vector difference value in the bitstream, the amount of bits can be saved by including only a part of the information about the differential value of the block vector in the bitstream.

[0192] The TM (Template Matching) method constructs a template using the surrounding pixels of the current block, finds the matching area with the highest similarity to the template, and then compensates for motion information. TM (Template Matching) is a method that performs motion prediction in the decoder without including motion information in the bitstream to reduce the size of the encoded bitstream. Since the decoder does not have the original image, it can roughly derive motion information for the current block using already reconstructed surrounding blocks.

[0193] BM (Bilateral Matching) may be a method in which a video signal processing device corrects motion information based on the similarity between a reference block in a picture included in an L0 picture list derived based on motion information of a current block and a reference block in a picture included in an L1 picture list.

[0194] The DMVR (Decoder-side Motion Vector Refinement) method is a method of compensating motion information through the correlation of already restored reference images to find slightly more accurate motion information. It is a method of using the bidirectional motion information of the current block to use the best matching point between reference blocks within an arbitrary set area of ​​two reference pictures as a new bidirectional motion. When this DMVR is performed, the encoder performs DMVR on a block basis to compensate for the motion information, and then divides the block into sub-blocks and performs DMVR on each sub-block basis to compensate for the motion information of the sub-block again. This can be called MP-DMVR (Multi-pass DMVR).

[0195] The LIC (Local Illumination Compensation) method is a method of compensating for luminance changes between blocks. It is a method of deriving a linear model using neighboring pixels adjacent to the current block, and then compensating for the luminance information of the current block through the linear model.

[0196] Existing video encoding methods perform motion compensation that only considers parallel translation in all directions, which reduces encoding efficiency when encoding videos containing motions commonly encountered in the real world, such as zooming in, zooming out, and rotation. To express such motions, an affine model-based motion prediction technique utilizing a four-parameter (rotation) or six-parameter (zoom in, zoom out, rotation) model can be applied.

[0197] Bi-Directional Optical Flow (BDOF) is used to compensate for predicted blocks by estimating pixel changes based on optical flow from a reference block of a block composed of bidirectional motion. The motion information derived from BDOF in VVC can be used to compensate for the motion of the current block.

[0198] PROF (Prediction Refinement with Optical Flow) is a technique for improving the accuracy of sub-block-level affine motion prediction to be similar to that of pixel-level motion prediction. Similar to BDOF, PROF calculates a correction value on a pixel-by-pixel basis for affine motion-compensated pixel values ​​on a sub-block-by-subblock basis based on optical flow to obtain a final prediction signal.

[0199] The CIIP (Combined Inter- / Intra-picture Prediction) method is a method of generating a final prediction block by weighting and averaging prediction blocks generated by the intra-picture prediction method and prediction blocks generated by the inter-picture prediction method when generating a prediction block for the current block.

[0200] The Intra Block Copy (IBC) method finds the most similar part to the current block in a previously restored area within the current picture and uses that reference block as a prediction block for the current block. At this time, information related to the block vector, which represents the distance between the current block and the reference block, can be included in the bitstream. The decoder can parse the block vector information contained in the BeastStream to calculate or set the block vector for the current block.

[0201] The BCW (Bi-prediction with CU-level Weights) method is a method that performs a weighted average on two motion-compensated prediction blocks by adaptively applying weights on a block-by-block basis, rather than generating a prediction block by averaging two motion-compensated prediction blocks from different reference pictures.

[0202] The Intra TMP (Template Matching Prediction) method is a method in which a video signal processing device constructs a reference template using pixel values ​​of neighboring blocks adjacent to the current block, finds the part most similar to the constructed reference template in an already restored area within the current picture, and then uses the reference block (the part found in the already restored area) as a prediction block for the current block.

[0203] The MHP (Multi-hypothesis prediction) method is a method of performing weight prediction using various prediction signals by transmitting additional motion information to unidirectional and bidirectional motion information during inter-screen prediction.

[0204] CCLM (Cross-component linear model) is a method of predicting chrominance signals by constructing a linear model using the high correlation between a luminance signal and the chrominance signals located at the same location as the luminance signal. A template is constructed using blocks adjacent to the current block that have been reconstructed, and the parameters for the linear model are derived from the template. Next, the current luminance block, which has been reconstructed to fit the size of the chrominance block, is selectively downsampled depending on the image format. Finally, the chrominance block of the current block is predicted using the downsampled luminance block and the linear model. In this case, a method using two or more linear models is called MMLM (Multi-model Linear mode). In addition, a prediction method that utilizes the correlation between different signals, such as CCLM and MMLM, can be referred to as CCP (Cross-Component Prediction).

[0205] GLM (Gradient Linear Model) is a method of predicting a color difference signal through a linear model such as CCLM, which constructs a model by additionally reflecting the gradient between the luminance sample corresponding to the color difference sample and the surrounding luminance samples adjacent to that luminance sample.

[0206] Prediction methods that utilize the correlation between different signals, such as CCLM, MMLM, CCCM, and GLM, can be called Cross-Component Prediction (CCP). In other words, a method of predicting another signal (a chrominance signal, such as Cb or Cr) from one signal (e.g., a luminance signal) can be called Cross-Component Prediction (CCP).

[0207] The CCP merge method is a method of predicting the chrominance block of the current block using the CCP model (models such as CCLM, MMLM, and CCCM) used in the surrounding blocks.

[0208] The encoding modes (prediction modes) described herein may be described without the term "mode" or with the term "method" instead of "mode." For example, the CCLM mode may be described as the CCLM method or CCLM.

[0209]

[0210] Among intra prediction encoding techniques, the MIP (Matrix Intra Prediction) method is a matrix-based intra prediction method. Unlike the prediction method that has directionality from the pixels of the surrounding blocks adjacent to the current block, it is a method of obtaining a prediction signal using a predefined matrix and offset values ​​for the pixels on the left and top of the surrounding blocks.

[0211] To derive the intra-prediction mode of the current block, a template, which is a random region adjacent to the current block and reconstructed, is used. The intra-prediction mode derived from the template's surrounding pixels can be used to reconstruct the current block. First, the decoder generates a prediction template for the template using the surrounding pixels (reference) adjacent to the template. The intra-prediction mode that generates the prediction template most similar to the already reconstructed template can then be used to reconstruct the current block. This method is called TIMD (Template intra-mode derivation).

[0212] TIMD merge mode is a mode that inherits the intra prediction directionality mode, information on whether to apply a weighted average, weighting information for each intra prediction mode, and transformation type information from neighboring blocks encoded with TIMD or TIMD merge mode, and uses them in the current block. The video signal processing device can construct a TIMD merge list using the TIMD information used in the neighboring blocks. The video signal processing device can add a predetermined maximum number of TIMD information derived from neighboring blocks to the TIMD merge list in order of distance. The predetermined maximum number can be a positive integer greater than or equal to 1. For example, the predetermined maximum number can be 5. The video signal processing device can rearrange the TIMD merge list based on the template cost and use the candidate with the minimum cost as the TIMD mode of the current block. If the current block is encoded with TIMD merge mode, the video signal processing device can use an implicit DST7 transform for the residual signal if the size of the current block is greater than or equal to 4 and less than or equal to 16. In other cases, the video signal processing device can use a transformation type derived from a neighboring block. The transformation type derived from a neighboring block may be transformation type information within a TIMD candidate for the current block determined from the TIMD merge list. In TIMD mode, the template cost can be calculated through SATD. If the calculation method of the template cost changes, the derived intra prediction mode may also change. The TIMD SAD mode is similar to the TIMD mode, but the method of calculating the template cost uses the MR-SAD method instead of SATD.

[0213] Typically, an encoder can determine a prediction mode for generating a prediction block and generate a bitstream containing information about the determined prediction mode. A decoder can parse the received bitstream to set an intra-prediction mode. At this time, the bit amount of information about the prediction mode may be about 10% of the total bitstream size. To reduce the bit amount of information about the prediction mode, the encoder may not include information about the intra-prediction mode in the bitstream. Accordingly, the decoder can derive (determine) an intra-prediction mode for restoring the current block using the characteristics of the surrounding blocks, and can restore the current block using the derived intra-prediction mode. At this time, the decoder can use a method of applying a Sobel filter in the horizontal and vertical directions to each neighboring pixel (pixel) of the current block to infer directional information, and then mapping the directional information to the intra-prediction mode to derive the intra-prediction mode. The method by which the decoder derives the intra prediction mode using surrounding blocks can be described as Decoder side intra mode derivation (DIMD).

[0214] If the current block is a luminance block, the video signal processing device can derive an intra prediction mode through the DIMD and TIMD methods. If the current block is a chroma block, since there is an already reconstructed luminance block corresponding to the chroma block, the video signal processing device can derive an intra prediction mode by applying the DIMD and TIMD methods using the reconstructed luminance block, and then use the derived intra prediction mode as the intra prediction mode of the chroma block. That is, the video signal processing device can derive directional information if the current block is a luminance block, and can not derive directional information if the current block is a chroma block, and can apply the directional information found in the luminance block to the chroma block. This mode can be called a 'DIMD Chroma' mode or a 'TIMD Chroma' mode.

[0215] Figure 7 is a diagram showing the locations of surrounding blocks used to construct a motion candidate list in inter prediction.

[0216] The neighboring blocks can be blocks of spatial position or blocks of temporal position. The neighboring block spatially adjacent to the current block can be at least one of a left (A1) block, a left below (A0) block, an above (B1) block, an above right (B0) block, and an above left (B2) block. The neighboring block temporally adjacent to the current block can be a block that includes an upper left pixel position of a bottom right (BR) block of the current block in a corresponding picture (Collocated picture). If the neighboring block temporally adjacent to the current block is encoded in intra mode or the neighboring block temporally adjacent to the current block exists in an unusable position, a block that includes a horizontal and vertical center (Center, Ctr) pixel position of the current block in the picture corresponding to the current picture (Collocated picture) can be used as the temporal neighboring block. Motion candidate information derived from corresponding pictures can be referred to as a Temporal Motion Vector Predictor (TMVP). Only one TMVP can be derived from a single block, or a block can be divided into multiple sub-blocks, and then a separate TMVP candidate can be derived for each sub-block. The method for deriving TMVPs at the sub-block level can be referred to as a sub-block Temporal Motion Vector Predictor (sbTMVP).

[0217] Whether the methods described in this specification are applied may be determined based on at least one of information about the slice type (e.g., whether it is an I slice, a P slice, or a B slice), whether it is a tile, whether it is a subpicture, the number of samples of the current block, the horizontal size of the current block, the vertical size of the current block, the depth of the coding unit, whether the current block is a luminance block or a chrominance block, whether it is a reference frame or a non-reference frame, the temporal layer according to the reference order and layer, etc. The information used to determine whether the methods described in this specification are applied may be information agreed upon in advance between the decoder and the encoder. In addition, such information may be determined according to a profile and a level. Such information may be expressed as a variable value, and the bitstream may include information about the variable value. That is, the decoder may parse the information about the variable value included in the bitstream to determine whether the above-described methods are applied. The methods described in this specification may be performed on a luminance block and a chrominance block. The methods described in this specification may be performed only on luminance blocks and not on chrominance blocks. The methods described in this specification may not be performed on luminance blocks and may only be performed on chrominance blocks. Whether the methods described in this specification are performed may be determined based on information about whether the current picture is used as a reference picture or not. For example, whether the above-described methods are applied may be determined based on the horizontal length or the vertical length of the coding unit. If the horizontal length or the vertical length is 32 or more (e.g., 32, 64, 128, etc.), the above-described methods may be applied. In addition, when the horizontal length or the vertical length is less than 32 (e.g., 2, 4, 8, 16), the above-described methods may be applied. In addition, when the horizontal length or the vertical length is 4 or 8, the above-described methods may be applied.

[0218] FIG. 8 illustrates a method for determining a reference pixel line based on a template according to one embodiment of the present specification.

[0219] Below, we describe a method for determining the optimal reference pixel line for the current block (for restoring the current block) based on a template. Here, the reference pixel line may have the same meaning as the reference sample line.

[0220] Referring to FIG. 8, a video signal processing device can construct a reference template using reference pixel lines adjacent to a current block. The video signal processing device can generate prediction samples for the positions of the reference template using reference pixel lines 1, 2, 3, etc. The video signal processing device can calculate a cost between the generated prediction samples and samples of the reference template. At this time, the cost can be calculated using a method such as SAD (Sum of Absolute Differences) or MRSAD (Mean-Removed SAD). The reference pixel corresponding to the minimum cost may be the optimal reference pixel. In addition, the encoder can rearrange the calculated costs in ascending order, construct a list for reference pixel lines, and then generate and signal a bitstream including information on an index for the optimal reference pixel line. The decoder can construct a list for reference pixel lines using the above-described method, parse the index for the optimal reference pixel line included in the bitstream, and generate a prediction sample using the reference pixel line indicated by the index. As described herein, a method for a video signal processing device to determine a reference pixel line based on a template can be described as a Template-based Multiple Reference Line (TMRL) method or a TMRL intra prediction method.

[0221] FIG. 9 is a diagram illustrating a block vector related to an IBC encoding method according to one embodiment of the present specification.

[0222] The IBC coding method (IBC mode) is a method of finding the part (reference block) that is most similar to the current block in an already reconstructed area within the current picture and using the reference block as a prediction block for the current block. At this time, the encoder can generate a bitstream that includes information related to a block vector, which is the distance between the current block and the reference block. The decoder can parse information related to the block vector included in the BeastStream to calculate or set the block vector for the current block. The IBC coding method can be applied to chrominance blocks. In the chrominance block, the block vector of the luminance block corresponding to the chrominance block can be used as the block vector of the chrominance block without finding a new block vector, and this coding method can be called DBV (Direct Block Vector) mode.

[0223] The RRIBC (Reconstruction-Reordered IBC) encoding mode can be used in IBC blocks (blocks to which the IBC encoding method is applied). RRIBC can be composed of vertical flips and horizontal flips. The reconstructed block of a block to which RRIBC is applied is flipped according to the RRIBC type of the current block. The encoder can flip the original block to be encoded before finding the most similar part to the current block from the reference picture. That is, the most similar part from the reference picture is found using the flipped original block. Therefore, the prediction block uses a block that has not been flipped, and the residual block can also be a block that has not been flipped. The decoder can flip the final reconstructed block according to the RRIBC type of the current block.

[0224] Figure 10 illustrates a method for predicting the current block using RRIBC in the horizontal direction.

[0225] Figure 11 illustrates a method for predicting the current block using RRIBC in the vertical direction.

[0226] In Figures 10 and 11, (Xn, Yn) represents the central position of the surrounding blocks, and (Xc, Yc) represents the central position of the current block. BV n h , BV n v represents the horizontal block vector and vertical block vector of the surrounding blocks, respectively, and BV C h , BV C v represent the horizontal block vector and the vertical block vector of the current block, respectively. As shown in Fig. 10, if the RRIBC type of the current block is horizontal, BV C h is 2 * (Xn - Xc) + BV n h It can be calculated as, and as in Fig. 53, if the RRIBC type of the current block is in the vertical direction, BV C v is 2 * (y n - y c ) + BV n v can be calculated as . At this time, BV n h , BV n v Since it uses restored blocks, BV n h , BV n v The sign of can be negative.

[0227] When the current block is encoded with RRIBC, the video signal processing device can determine an optimal block vector by constructing a block vector candidate list. At this time, the video signal processing device can construct the block vector candidate list according to the RRIBC type of the current block. For example, when the RRIBC type of the current block is horizontal, the video signal processing device can construct the block vector candidate list using only the neighboring blocks of the current block encoded with RRIBC in the horizontal direction. In addition, the video signal processing device can construct the block vector candidate list regardless of the RRIBC type of the current block. For example, when the RRIBC type of the current block is horizontal, the block vector candidate list can be constructed using not only the neighboring blocks of the current block encoded with RRIBC in the horizontal direction, but also the neighboring blocks of the current block encoded with RRIBC in the vertical direction, and / or blocks encoded with general motion, and / or blocks encoded with block vectors.

[0228] When a video signal processing device predicts a current block using general motion, the video signal processing device can construct a motion candidate list depending on whether the neighboring blocks of the current block are encoded in the IBC mode or the RRIBC mode. At this time, if the neighboring blocks of the current block are encoded in the RRIBC mode, the video signal processing device can construct the motion candidate list by additionally considering the RRIBC type. For example, when the video signal processing device constructs a motion candidate list for the current block, if the encoding mode of the neighboring blocks of the current block is the RRIBC mode and the RRIBC type is the vertical direction or the horizontal direction, the block vector of the neighboring blocks may not be included in the motion candidate list. Alternatively, when the video signal processing device predicts the current block using general motion, the video signal processing device can construct a motion candidate list regardless of the encoding mode of the neighboring blocks of the current block. In other words, the video signal processing device can construct a motion candidate list regardless of whether the neighboring blocks of the current block are encoded in the IBC mode or the RRIBC mode. For example, when a video signal processing device constructs a motion candidate list for a current block, the video signal processing device may include block vectors of the surrounding blocks in the motion candidate list even if the encoding mode of the surrounding blocks of the current block is the IBC mode and the RRIBC type is the vertical direction or the horizontal direction.

[0229] The RRIBC encoding method is effective for images with symmetrical characteristics. Symmetry can mean perfect horizontal (or vertical) symmetry with equal distances between the current block and the reference block around a central axis. The vertical (or horizontal) direction (symmetry axis) of the current block can be colinear with the vertical (or horizontal) direction (symmetry axis) of the reference block. In this case, the block vector in the vertical (or horizontal) direction can be set to '0', and the block vector in the horizontal direction can be set to any negative value other than '0'. The current block and the reference block can be configured to be symmetric with different distances apart from one central axis between the current block and the reference block, and the current block and the reference block can be located on different vertical lines. In this case, the block vector in the vertical direction can be set to any negative value other than '0'. That is, the video signal processing device can encode or decode the current block using a symmetric block having a block vector value of an arbitrary negative value other than '0' in both the horizontal and vertical directions.

[0230] FIG. 12 illustrates a block vector of a block encoded in Intra TMP mode according to one embodiment of the present specification.

[0231] Referring to FIG. 12, the Intra TMP method (Intra TMP encoding mode) is a method in which a video signal processing device constructs a reference template using pixel values ​​of neighboring blocks adjacent to a current block, then finds the most similar part to the reference template in an already reconstructed area (reference block) within the current picture, and then uses the reference block (Ref. luma block of FIG. 12) as a prediction block for the current block. At this time, there may be more than one block vector used to generate a prediction block for the current luminance block, and FIG. 12 shows a case in which there are two block vectors. The video signal processing device can generate a prediction block for the luminance block by weighting and averaging the reference blocks predicted from the two block vectors of the current luminance block. This can be referred to as a fusion mode. In the case of a chrominance block, the video signal processing device can derive a block vector from a luminance block corresponding to the chrominance block, and then generate a chrominance prediction block using the block vector. There may be more than one block vector used to generate a prediction block for the current chrominance block, and FIG. 12 shows a case where there are two block vectors for the chrominance block. In this case, the block vector of the chrominance block may be the same as or different from the block vector derived from the luminance block. In the Intra TMP mode, the video signal processing device can derive filter coefficients using the correlation between samples around the reference block indicated by the block vector and samples around the current block, and apply filtering to the reference block indicated by the block vector using the derived filter coefficients. The video signal processing device can use the filtered reference block as a prediction block of the current block. This can be referred to as the Intra TMP filtering mode. In the Intra TMP mode, the video signal processing device can apply compensation to the current prediction block through the block vector similar to the LIC method using motion information.

[0232] FIG. 13 illustrates a case where a current block is divided by GPM mode according to one embodiment of the present specification and the divided area is encoded by IBC mode.

[0233] The current block can be encoded or decoded in IBC-GPM mode. IBC-GPM mode may mean a mode in which the current block is divided using GPM mode, and the divided region is encoded in intra prediction mode or IBC prediction mode. Referring to Fig. 13, the current block can be divided into two regions based on a dotted line. At this time, one of the divided regions can be encoded in intra prediction mode, and the other can be encoded in IBC prediction mode. For example, the left region among the divided regions can be encoded in intra prediction mode, and the right region can be encoded in IBC prediction mode. At this time, the region encoded in IBC prediction mode can be encoded in Intra TMP prediction mode, and this can be described as IntraTMP-GPM mode.

[0234] FIG. 14 illustrates how a current block is encoded in IBC-CIIP mode according to one embodiment of the present specification.

[0235] A video signal processing device can predict the current block using a weighted average between a block predicted in Intra mode and a block predicted in IBC mode, which can be referred to as IBC-CIIP mode. In this case, a block predicted in Intra TMP mode can be used instead of a block predicted in IBC mode, which can be referred to as Intra TMP-CIIP mode.

[0236] A video signal processing device can obtain a reconstructed luminance block by adding a residual signal for a luminance prediction block of a current block and a luminance block, and can configure a CCP model by using a correlation between the reconstructed luminance block and the luminance prediction block. At this time, the CCP model can be one of CCLM, MMLM, GLM, CCCM, MM-CCCM, GL-CCCM, CCCM-ND, and CCCM-MDF. The CCP model derived from the luminance block can be applied to a chrominance prediction block to generate a first chrominance prediction block to which the CCP model is applied. A final chrominance block can be generated by adding an error signal for the first chrominance prediction block and the chrominance block.

[0237] If the current luminance block is encoded in Intra TMP or IBC mode, the video signal processing device can derive a CCP model through the correlation between the luminance block and the chrominance block of the reference block indicated by the block vector of the current luminance block. The derived CCP model and the currently reconstructed luminance block can be used to generate Cb and Cr prediction blocks, which are chrominance prediction blocks for the current block. This is referred to as BVG CCP (block vector guided cross component prediction). The CCP model can include at least one of CCLM, MMLM, GLM, MM-GLM, CCCM, MM-CCCM, GL-CCCM, GL-MM-CCCM, CCCM-ND, MM-CCCM-ND, CCCM-MDF, MM-CCCM-MDF, and Chroma Fusion. When the CCP model is derived using the BV of the luminance block, and the CCCM model is applied among the CCP models, it can be referred to as BVG CCCM. If the current chrominance block derives a CCP model using the BV of the luminance block and the GL-CCCM model is applied among the CCP models, it can be referred to as BVG GL-CCCM. In addition, various CCP models can be applied to the BVG CCP.

[0238] The type of a CCP model may include at least one of CCLM, MMLM, GLM, MM-GLM, CCCM, MM-CCCM, GL-CCCM, MM-GL-CCCM, CCCM-ND, MM-CCCM-ND, CCCM-MDF, MM-CCCM-MDF, chroma fusion, BVG-CCCM, and CCP merge. An MMLM may be composed of two CCLMs. If the types of CCP models between CCP models are different, they may be considered different CCP models. If the types of CCP models between CCP models are the same but the parameters between the CCP models are different, they may be considered different CCP models. If the number of CCP models between CCP models is different, for example, if a first CCP model is composed of one CCLM and a second CCP model is composed of two CCLMs, they may be considered different CCP models. The type of a CCP model may be referred to as a CCP mode.

[0239] In chroma fusion (CF), a video signal processing device can predict a chrominance signal through a weighted average between a chrominance block predicted through intra prediction mode and the current luminance block, without using the linear model (LM) mode. At this time, the weight parameters can be derived using the CCCM method.

[0240] In a GLM (gradient linear model), a video processing device constructs a model by additionally reflecting the gradient between a luminance sample corresponding to a chrominance sample in a linear model and surrounding luminance samples adjacent to the luminance sample, and can predict a chrominance signal through the model.

[0241] The CCP model used to predict a chroma block can construct a CCP candidate list using the CCP models used in the surrounding blocks of the current block, and derive the optimal CCP candidate from the CCP candidate list. This can be called CCP merge mode.

[0242] FIG. 15 illustrates an example of a reference region and filter shape used to derive CCCM parameters according to one embodiment of the present specification.

[0243] CCCM (Convolutional cross-component intra prediction model) is a method of predicting a chrominance signal using a nonlinear model that utilizes the high correlation between a luminance signal and a chrominance signal located at the same location as the luminance signal. Fig. 15(a) shows the positional relationship between reference samples (vertical hatching, 1520) for applying CCCM to the current prediction block (diagonal hatching, 1510) and side samples (horizontal hatching) required when applying a cross-shaped filter. The current prediction block (MxN) can be composed of reference samples in the upper 6 rows (2Mx6), reference samples in the left 6 rows (6x2N), and 6x6 reference samples in the upper left, where the ratio of the number of chroma to luma samples is 1:1. When a video signal processing device applies a cross-shaped sample (Fig. 15(b)) filter to a chroma sample prediction relationship (Fig. 15(c)) for CCCM, cases where the reference sample area is exceeded may occur. At this time, additionally required samples may be side samples. The chroma sample prediction relationship of Fig. 15(c) may be applied to each chroma component (i.e., Cb, Cr). The sample at the C (Center) position in Fig. 15(b) may be a luma sample corresponding to the Cb and Cr chroma samples, and N (North), E (East), S (South), and W (West) may be luma samples adjacent to the luma sample at the C position. Side samples may additionally require one sample for an area other than the reference sample depending on the position of the C sample. If the sample value existing at the position of the side sample is not available, the sample value at the unavailable position may be padded with the C sample value. The P value of Fig. 15(c) may be a nonlinear term. The P value may be calculated as P = ( C*C + midVal ) >> bitDepth , and in the case of 10-bit content, P = ( C*C + 512 ) >> 10 .In Fig. 15(c), the B value can be an integer offset value as a bias value. The B value can be an intermediate value of bitDepth (bit depth). In the case of 10-bit content, the B value can be 512.

[0244] MM-CCCM (multi-model CCCM) is a method that derives two CCCM parameters based on the average value of the reference area (or the restored current luminance block).

[0245] GL-CCCM (Gradient and location based convolutional cross-component model) is an additional CCCM mode that uses gradient and location information. In the case of the existing CCCM mode, the video signal processing device can derive the chrominance sample for the current block using the luminance sample at the corresponding position from the predicted chrominance sample position, four samples around the luminance sample, and coefficient information. At this time, in the case of the GL-CCCM mode, the video signal processing device can derive the chrominance sample for the current block by reflecting the vertical and horizontal differences for the luminance sample at the corresponding position from the predicted chrominance sample position and eight samples around the luminance sample, and also using the position value of the current luminance sample and its coefficient information.

[0246] When CCCM mode is applied, the video signal processing device can apply a downsampling filter to match the resolution difference between the luminance block and the chrominance block. This is to reduce the resolution of the luminance block to that of the chrominance block. The mode that applies the downsampling filter can be described as CCCM-MDF (CCCM with multiple downsampling filters).

[0247] When the inter encoding mode is applied to the current block, the video signal processing device can derive linear and nonlinear models between the luminance prediction block (Y') derived using the motion information of the current block and the first chrominance prediction block (Cb', Cr') derived using the motion information of the current block (Derive filter), and then generate a reconstructed luminance block of the current block using the luminance residual block of the current block, and apply the derived linear and nonlinear models to the reconstructed luminance block of the current block (Apply filter) to generate a second chrominance prediction block of the current block. Thereafter, the video signal processing device can add the chrominance residual block to the second chrominance prediction block of the current block predicted using the linear and nonlinear models to finally generate the chrominance block (Cb, Cr) of the current block. This method can be described as a cross-component residual model (CCRM). CCRM, Inter CCCM, and Inter CCP may have the same meaning.

[0248] To improve coding efficiency, rather than coding the aforementioned residual signal as is, a method may be used in which the transform coefficient values ​​obtained by transforming the residual signal are quantized and the quantized transform coefficients are coded. As described above, the transform unit may obtain the transform coefficient values ​​by transforming the residual signal. At this time, the residual signal of a specific block may be distributed across the entire region of the current block. Accordingly, by performing a frequency-domain transformation on the residual signal, energy can be concentrated in the low-frequency region, thereby improving coding efficiency.

[0249] The encoder can obtain at least one residual block containing a residual signal for the current block. The residual block may be the current block or one of the blocks split from the current block. In this specification, the residual block may be described as a residual array or a residual matrix containing residual samples of the current block. Furthermore, in this specification, the residual block may represent a block having the same size as the size of a transform unit or a transform block.

[0250] An encoder can transform a residual block using a transform kernel. The transform kernel used for transforming the residual block may be a transform kernel having separable vertical and horizontal transform properties. In this case, the transform for the residual block may be performed separately as a vertical transform and a horizontal transform. For example, the encoder may perform a vertical transform by applying a transform kernel in the vertical direction of the residual block. Additionally, the encoder may perform a horizontal transform by applying a transform kernel in the horizontal direction of the residual block. In this specification, a transform kernel may be used as a term referring to a set of parameters used for transforming a residual signal, such as a transform matrix, a transform array, a transform function, or a transform. In one embodiment, the transform kernel may be any one of a plurality of available kernels. Additionally, transform kernels based on different transform types may be used for each of the vertical and horizontal transforms. That is, before performing the first transformation, the transformation method for the vertical and horizontal directions can be derived using at least one or more of the intra prediction mode of the current block, the encoding mode, the transformation method parsed from the bitstream, and the size information of the current block. In addition, in order to reduce the computational complexity in the transformation process for large blocks, a process of processing the high-frequency region as '0' while leaving only the low-frequency region can be performed. This process is called high-frequency zeroing, and the transformation size at the time of the actual first transformation can be set for this zeroing. In the high-frequency zeroing process, the low-frequency region can be set to an arbitrary fixed size, and for example, the horizontal or vertical size can be a combination of 4, 8, 16, 32, etc.

[0251] The encoder can quantize the transform block transformed from the residual block by passing it to the quantization unit. At this time, the transform block can include multiple transform coefficients. Specifically, the transform block can be composed of multiple transform coefficients arranged in a two-dimensional array. The size of the transform block, like the residual block, can be the same as either the current block or a block split from the current block. The transform coefficients passed to the quantization unit can be expressed as quantized values.

[0252] Additionally, the encoder may perform an additional transform before the transform coefficients are quantized. The aforementioned transform method may be referred to as a primary transform, and the additional transform may be referred to as a secondary transform. The secondary transform may be optional for each residual block. In one embodiment, the encoder may perform the secondary transform for areas where it is difficult to focus energy in the low-frequency region using only the primary transform, thereby improving coding efficiency. For example, the secondary transform may be added for blocks in which the residual values ​​appear significantly in directions other than the horizontal or vertical directions of the residual block. The residual values ​​of an intra-predicted block may be more likely to change in directions other than the horizontal or vertical directions compared to the residual values ​​of an inter-predicted block. Accordingly, the encoder may additionally perform the secondary transform on the residual signal of the intra-predicted block. Additionally, the encoder may omit the secondary transform on the residual signal of an inter-predicted block. High-frequency zeroing in the first conversion can also be performed in the second conversion process.

[0253] As another example, whether to perform a secondary transform may be determined based on the size of the current block or the remaining block. Furthermore, transform kernels having different sizes may be used based on the size of the current block or the remaining block. For example, an 8X8 secondary transform may be applied to a block in which the length of a shorter side among the width or the height is greater than or equal to a first preset length. Furthermore, a 4X4 secondary transform may be applied to a block in which the length of a shorter side among the width or the height is greater than or equal to a second preset length and less than the first preset length. In this case, the first preset length may be a value greater than the second preset length, but the present disclosure is not limited thereto. Furthermore, unlike the first transform, the secondary transform may not be performed separately into a vertical transform and a horizontal transform. Such a secondary transform may be referred to as a low frequency non-separable transform (LFNST).

[0254] In addition, for video signals in a specific region, high-frequency band energy may not be reduced even if frequency transform is performed due to rapid brightness changes. Accordingly, compression performance due to quantization may deteriorate. In addition, when transform is performed on a region where residual values ​​rarely exist, encoding time and decoding time may unnecessarily increase. Accordingly, transform for the residual signal in a specific region may be omitted. Whether or not to perform transform for the residual signal in a specific region may be determined by a syntax element related to the transform of the specific region. For example, the syntax element may include transform skip information. The transform skip information may be a transform skip flag. If the transform skip information for a residual block indicates transform skip, transform for the corresponding residual block is not performed. In this case, the encoder can immediately quantize the residual signal in the region where the transform is not performed.

[0255] The aforementioned transformation-related syntax elements may be information parsed from a video signal bitstream. A decoder may entropy decode the video signal bitstream to obtain the transformation-related syntax elements. Additionally, an encoder may entropy code the transformation-related syntax elements to generate a video signal bitstream.

[0256] The decoder can obtain encoding information necessary for decoding by parsing the transmitted bitstream. At this time, information related to the transformation process includes index information for the first and second transformation types and quantized transformation coefficients. The inverse transformation unit can obtain a residual signal by inversely transforming the inverse quantized transformation coefficients. First, the inverse transformation unit can detect whether an inverse transformation is performed for a specific region from a transformation-related syntax element of the specific region. According to one embodiment, if a transformation-related syntax element for a specific transformation block indicates a transformation skip, the transformation for the corresponding transformation block may be skipped. In this case, both the first inverse transformation and the second inverse transformation for the transformation block may be skipped. In addition, the inverse quantized transformation coefficients can be used as a residual signal. For example, the decoder can use the inverse quantized transformation coefficients as a residual signal to reconstruct the current block. Alternatively, the second inverse transform may be performed and the first inverse transform may be omitted, and the second inverse transformed value may be used as the residual signal. The first inverse transform described above represents the inverse transform for the first transform and may be referred to as the inverse primary transform. The second inverse transform represents the inverse transform for the second transform and may be referred to as the inverse secondary transform or the inverse LFNST. In the present invention, the first (inverse) transform may be referred to as the first (inverse) transform, and the second (inverse) transform may be referred to as the second (inverse) transform.

[0257] FIG. 16 illustrates types of transform kernels that can be used in video coding according to one embodiment of the present specification.

[0258] Fig. 16 shows the formulas of DCT-II, DCT-V (discrete cosine transform type-V), DCT-VIII (discrete cosine transform type-VIII), DST-I (discrete sine transform type-I), and DST-VII kernels applied to MTS. DCT and DST can be expressed as functions of cosine and sine, respectively, and when the basis function of the transform kernel for the number of samples N is expressed as Ti(j), the index i represents the index in the frequency domain, and the index j represents the index within the basis function. That is, as i becomes smaller, it represents a low-frequency basis function, and as i becomes larger, it represents a high-frequency basis function. When the basis function Ti(j) is expressed as a two-dimensional matrix, it can represent the j-th element of the i-th row, and since the transform kernels illustrated in Fig. 16 all have a separable characteristic, they can perform transformations in the horizontal and vertical directions respectively for the residual signal X. That is, when the residual signal block is X and the transform kernel matrix is ​​T, the transform for the residual signal X can be expressed as TXT'. Here, T' means the transpose of the transform kernel matrix T. Since DCT and DST are in decimal form, not integer form, it is burdensome to implement them as they are in hardware encoders and decoders. Therefore, the decimal form transform kernel must be approximated to an integer form transform kernel through scaling and rounding. The integer precision of the transform kernel can be determined as 8-bit or 10-bit, but if the precision is low, the encoding efficiency may decrease. Depending on the approximation, the orthonormal property of DCT and DST may not be maintained, but the encoding efficiency loss due to this is not large, so approximating the transform kernel to an integer form is advantageous in terms of implementing a hardware encoder and decoder.IDTR (Identity Transform) is a transformation whose result is the same as the original transformation. This is called an identity transformation. Typically, an identity transformation constructs a transformation matrix by setting "1" in positions where rows and columns have the same value. However, here, the identity transformation uses an arbitrary fixed value, not "1," to uniformly increase or decrease the value of the input residual signal.

[0259] In the above-mentioned first-order transform, the MTS transform, the transform is calculated by applying the transform kernel to the vertical and horizontal directions of the error block respectively, so it can be said to be a separable transform method. On the other hand, in the above-mentioned second-order transform, the LFNST transform, the transform is calculated by applying the transform kernel only once without applying the transform kernel to the vertical and horizontal directions respectively, so it can be said to be a non-separable transform method. In addition, the above-mentioned second-order transform is additionally applied to the first-order transformed transform coefficients of the block to which the DCT-2 transform is applied, so it can be said to be a two-step transform technique. The above-mentioned second-order transform has high encoding efficiency, but it has the disadvantage of being complex because the transform kernel is applied a total of three times. To reduce this complexity, the NSPT (Non-separable primary transform) method, which is a method of applying the transform using only the second-order transform method, can be applied. The NSPT transform method is a non-separable transform method, and is calculated by applying the transform kernel only once to the vertical and horizontal directions of the error block, without applying the transform kernel separately. In a video signal processing device, the error block of the current block can be transformed or inversely transformed using one of the following transform methods: MTS, DCT2 + LFNST, or NSPT transform.

[0260] NSPT transform can be a transform method that replaces the existing DCT2 + LFNST transform method. For blocks whose size is equal to or smaller than 16x16, one of the following kernels can be applied depending on the size of the transform block: NSPT4x4 (16x16 kernel), NSPT4x8 (32x20 kernel), NSPT8x4 (32x20 kernel), NSPT8x8 (64x32 kernel), NSPT4x16 (64x24 kernel), NSPT16x4 (64x24 kernel), NSPT8x16 (128x40 kernel), NSPT16x8 (128x40 kernel), NSPT4x32 (128x20 kernel), NSPT32x4 (128x20 kernel), NSPT8x32 (256x24 kernel), NSPT32x8 (256x24 kernel). NSPT, similar to LFNST, consists of 35 sets of transform kernels, each of which can have three candidates. The encoder can derive a set of transform kernels according to the intra prediction mode, and then generate and signal a bitstream containing information about the index of the optimal candidate among the three candidates. The decoder can parse the information about the signaled index, and then use the transform kernel candidate indicated by the index information among the set of transform kernels derived using the intra prediction mode to inversely transform the current transform coefficients and obtain a residual block. Zero-out may not be performed on a 4x4 block to which NSPT is applied. In addition, the number of coefficients to be zeroed out may vary depending on the size of the NSPT kernel. For example, NSPT with a size of 32x20 can be applied to a 4x8 block or an 8x4 block. Therefore, out of the 32 transform coefficients, only 20 transform coefficients can be zeroed out, while the remaining 12 transform coefficients can be zeroed out.

[0261] FIG. 17 illustrates a transformation set table for LFNST and NSPT transformations according to one embodiment of the present specification.

[0262] There can be 35 transform sets used in LFNST and NSPT transforms, and they can vary depending on the intra prediction mode (see FIG. 6). That is, the video signal processing device can derive the transform set index of the LFNST and NSPT transforms corresponding to the intra prediction mode (see FIG. 6) by referring to the transform set table of FIG. 17. In addition, the LFNST and NSPT transform sets can vary depending on the intra prediction mode, information on whether the current block is a luminance block or a chrominance block, the horizontal and vertical sizes of the current block, and whether the intra prediction directional mode of the current block is an extended angle mode. There can be an arbitrary number of transform matrices for each transform set. Here, the arbitrary number can be an integer greater than or equal to 1, and can be 3. The encoder can signal the index information for the optimal transform matrix among multiple transform matrices in the transform set by including it in the bitstream. After the decoder parses the index for the optimal transformation matrix, it can apply the inverse transformation using the transformation matrix corresponding to the index in the transformation set.

[0263] The encoder and decoder can apply three types of transformation kernels, LFNST4, LFNST8, and LFNST16, depending on the size of the transformation block. If the width and height of the current transformation block are greater than or equal to 16, the encoder and decoder can apply the LFNST16 transformation kernel. If the width and height of the current transformation block are greater than or equal to 8, the encoder and decoder can apply the LFNST8 transformation kernel. If the width and height of the current transformation block are less than 8, the encoder and decoder can apply the LFNST4 transformation kernel.

[0264] Figure 18 shows an example of ROI after LFNST transformation.

[0265] The gray block area in Fig. 18, which is the low-frequency part output after the encoder and decoder LFNST-convert the residual block, is referred to as the ROI (Region of Interest). Areas other than the pre-designated ROI can be zero-out processed, which is set to 0. In the case of LFNST16, only 96 samples are required. Therefore, as shown in (a) of Fig. 18, the remaining white sub-blocks except for 6 NxN sub-blocks can be zero-out processed. In the case of LFNST8, only 64 samples are required. Therefore, as shown in (b) of Fig. 18, the remaining white sub-blocks except for 4 NxN sub-blocks can be zero-out processed. In the case of LFNST4, there may not be a zero-out area. In this case, N is a positive integer and may be 4.

[0266] FIG. 19 illustrates a method for deriving a multi-transform set and a LFNST / NSPT set according to one embodiment of the present specification.

[0267] Referring to Fig. 19(a), the encoder can select which one of MTS, DCT2 + LFNST, and NSPT to apply to the residual block, and can transform the residual block and obtain transform coefficients based on the selected transform method. At this time, the encoder can signal information on which transform method was applied by including it in the bitstream. If the MTS transform is applied to the residual block, the transform coefficients of the residual block can be obtained by applying the MTS transform, and LFNST and NSPT transforms may not be applied. If the DCT2 + LFNST transform is applied to the residual block, the encoder can apply the DCT2 transform to the residual block to obtain the first transform coefficients, and apply the LFNST transform to the first transform coefficients to obtain the second transform coefficients. At this time, the MTS and NSPT transforms may not be applied. When the NSPT transform is applied to the residual block, the video encoder can output transform coefficients by applying the NSPT transform to the residual block, and at this time, the MTS and DCT2 + LFNST transforms may not be applied to the residual block.

[0268] Referring to Fig. 19(b), the decoder can parse information about which transform method was applied from the bitstream, and determine whether to apply one of the inverse transform methods among MTS, DCT2 + LFNST, and NSPT to the current transform coefficient based on the parsed information. Then, the decoder can perform an inverse transform on the transform coefficient based on the determined transform method and obtain a residual block. If the MTS transform is applied to the current transform coefficient, the decoder can perform an inverse MTS transform on the transform coefficient to obtain a residual block, and at this time, the LFNST and NSPT inverse transforms may not be applied. If the DCT2 + LFNST transform is applied to the secondary transform coefficient, the decoder can perform an inverse LFNST transform on the secondary transform coefficient to output a primary transform coefficient, and perform an inverse DCT2 transform on the primary transform coefficient to obtain a residual block, and at this time, the MTS and NSPT transforms may not be applied. If the NSPT transform is applied to the current transform coefficients, the decoder can obtain a residual block by performing the NSPT inverse transform on the current transform coefficients, and at this time, the MTS and DCT2 + LFNST transforms may not be applied.

[0269] A video signal processing device can derive a transform kernel for each of MTS, LFNST, and NSPT transforms (or inverse transforms) using an intra prediction mode. In addition, the video signal processing device can determine which transform (or inverse transform) is applied among MTS, DCT2 + LFNST, and NSPT. At this time, in order to determine the transform, at least one or more of the horizontal and vertical sizes of the current block, whether the components of the current block are luminance components or chrominance components, whether the current block is a single tree or a dual tree, information on whether the current block is encoded in intra mode or inter mode, information on the encoding mode of the current block (e.g., IBC, Intra TMP, Merge, AMVP, GPM, SGPM, CCLM, CCCM), and information on the quantization parameters of the current block may be used.

[0270] Figure 20 illustrates a mapping table according to one embodiment of the present specification.

[0271] Figure 21 illustrates a conversion type set table according to one embodiment of the present specification.

[0272] Figure 22 illustrates a conversion type combination table according to one embodiment of the present specification.

[0273] FIG. 23 illustrates a threshold value table for an IDT conversion type according to one embodiment of the present specification.

[0274] A method for selecting a set of multiple transforms available to the current block in a video signal processing device is described.

[0275] 1) First, the video signal processing device can derive the nSzIdxW and nSzIdxH values ​​based on the size of the current block to map the horizontal and vertical sizes of the current block into a single variable. nSzIdxW can be the minimum value among the logarithm of 2 calculated for the width of the current block, the difference of 2 with the decimal places discarded, and the three. nSzIdxH can be the minimum value among the logarithm of 2 calculated for the height of the current block, the difference of 2 with the decimal places discarded, and the three.

[0276] 2) Next, the video signal processing device can derive the intra-directional mode (predMode) of the current block. In the case of TIMD mode, the number of intra-prediction modes expanded from the existing 67 to 131 can be used, reducing the precision of the existing 67 modes.

[0277] 3) Next, the video signal processing device can derive the ucMode, nMdIdx, and isTrTransposed values.

[0278] A. If the current block is encoded in MIP mode, ucMode can be set to '0', nMdIdx can be set to '35', and isTrTransposed can be set to a value derived from MIP.

[0279] B. If the current block is not encoded in MIP mode, ucMode can be set to the intra directional mode (predMode) of the current block. predMode can mean an index value of the intra directional mode. predMode can be determined through an extended angle mode according to the aspect ratio of the current block. The video signal processing device can clip predMode to a value between 2 and 66. If predMode is greater than the 34th angle mode, which is a diagonal mode, the isTrTransposed value can be set to 1, and if predMode is less than or equal to 34, the isTrTransposed value can be set to 0. If predMode is greater than 34, the video signal processing device resets the value of predMode to a value that is a difference between 67 (the maximum value of the intra directional mode index) plus 1. For example, if predMode is 35, the video signal processing device can reset the 35th angle mode to the 33rd angle mode (67+1-35(predMode)). If predMode is 66, the video signal processing device can reset the 66th angle mode to the 2nd angle mode (67+1-66(predMode)). That is, by making it symmetrical with respect to the 34th angle mode, which is a diagonal mode, there is an effect of reducing the size of the transformation mapping table of Fig. 35 by about half.

[0280] 4) The video signal processing device can derive the nSzIdx value through nSzIdxW, nSzIdxH, and isTrTransposed values. If the isTrTransposed value is '1', the value obtained by multiplying nSzIdxH by 4 and then adding nSzIdxW can be set to nSzIdx. If the isTrTransposed value is '0', the value obtained by multiplying nSzIdxW by 4 and then adding nSzIdxH can be set to nSzIdx.

[0281] 5) The video signal processing device can derive nTrSet, which is an index of an available transformation type set, according to the predefined table of FIG. 20 using nSzIdx, which is the size information of the current block, and nMdIdx, which is the intra-directional mode information of the current block. FIG. 20 defines an index of a transformation type set according to the intra-screen directional mode (0 to 34 and MIP) of the current block and the size index (0 to 15) of the current block. Referring to FIG. 20, nTrSet can be 80, and if the size of the current block is 4x8 and the intra-screen directional mode of the current block is 13, nTrSet can be '7'.

[0282] 6) The video signal processing device can derive a set of transformation types corresponding to nTrSet from the table of FIG. 21 by parsing mts_idx included in the bitstream. The transformation types in the vertical and horizontal directions are set differently depending on whether predMode is a value greater than the 34th angle mode, which is a diagonal mode. 0 to 79 in the vertical column of FIG. 21 may correspond to nTrSet, and 0 to 3 in the horizontal column may correspond to mts_idx. Referring to FIG. 21, if nTrSet is 7 and the value of mts_idx is 3, 22 may be selected from (2, 17, 18, 22). In addition, DST1 and DCT5 corresponding to index 22 of the transformation type combination table of FIG. 22 are selected, and the vertical transformation type of the current block may be set to DST1 and the horizontal transformation type may be set to DCT5. 0 to 24 in the horizontal column of FIG. 22 are indices selected through FIG. 21, and 0 to 1 in the vertical column may represent vertical and horizontal transformation types, respectively. If the predicted directionality mode within the screen of the current block is greater than 34, which is a diagonal mode, the vertical and horizontal transformation types may be interchanged.

[0283] If the mts_idx value is '3' and the width and height of the current block are both 16 or less, the vertical or horizontal transformation type can be reset to the IDT transformation type through the process described below.

[0284] If the absolute difference between the index of the on-screen prediction directional mode of the current block and the index of the horizontal mode, 18, is less than any predetermined value, the vertical transformation type may be reset to the IDT transformation type. If the absolute difference between the index of the on-screen prediction directional mode of the current block and the index of the horizontal mode, 50, is less than any predetermined value, the horizontal transformation type may be reset to the IDT transformation type. At this time, the predetermined value is an integer and may be determined based on the horizontal or vertical size of the current block. For example, the predetermined value may be determined through the table of Fig. 20. The table of Fig. 23 (a) shows a case where the threshold value is set differently whenever the horizontal or vertical size differs by 4, and the table of Fig. 23 (b) shows a case where the threshold value is set differently whenever the horizontal or vertical size differs by double. If the size of the current block is 16x16, the video signal processing device may not reset the vertical transformation type to the IDT transformation type, but may maintain the existing transformation type as is.

[0285] The encoder and decoder can reduce the amount of bits for encoding mts_idx by adaptively changing the number of transform type sets for each block. The encoder and decoder can determine the number of transform type sets to be one, four, or six, depending on the sum of the absolute values ​​of the transform coefficients of the current transform block. If the number of transform type sets changes, the maximum number of bins for signaling mts_idx may change. If the sum of the absolute values ​​of the transform coefficients is less than or equal to 6, the number of transform type sets is one. Therefore, the encoder may not signal mts_idx. In this case, the decoder may not parse mts_idx. If the sum of the absolute values ​​of the transform coefficients is greater than 6 and less than or equal to 32, the number of transform type sets may be four. If the sum of the absolute values ​​of the transform coefficients is greater than 32, the number of transform kernel candidates may be six. At this time, the encoder can signal mts_idx based on the maximum number of transformation type sets. Additionally, the decoder can parse mts_idx based on the maximum number of transformation type sets.

[0286] When the current block is predicted in inter coding mode, the encoder and decoder can use one of four transform kernel sets: {(DST7, DST7), (DST7, DCT8), (DCT8, DST7), (DCT8, DCT8)}. The encoder can signal an index of which transform kernel set it used. The decoder can parse the index to determine which transform kernel set to use for the current transform block. When the configured transform kernel set is (DST7, DCT8), the encoder and decoder transform or inversely transform the current block using DST7 in the horizontal direction, and transform or inversely transform the current block using DCT8 in the vertical direction. To optimize complexity, the encoder and decoder can set the maximum CU size to which Inter MTS can be applied to 32x32 for images larger than 1920x1080 resolution. The encoder and decoder set the maximum CU size to which Inter MTS can be applied to images with a resolution other than 1920x1080 to 16x16. Therefore, the encoder and decoder can apply DCT2 transform in the horizontal and vertical directions without applying Inter MTS to transform blocks larger than 16x16, for example, transform blocks of size 32x32. In addition, the encoder and decoder can use separable KLT instead of DST7 and DCT8 in transform and inverse transform of transform blocks of size equal to or smaller than 16x16.

[0287] FIG. 24 illustrates block boundaries and samples around the boundaries in a deblocking filtering process according to one embodiment of the present specification.

[0288] Referring to (a) of Fig. 24, the part indicated by the dotted line between the P block and the Q block may denote a block boundary. The block boundary may exist at any fixed size, and a block boundary may exist at every size of 4.

[0289] Figure 24 (b) shows samples for which filtering is performed based on block boundaries. The video signal processing device can perform a process of alleviating blocking artifacts occurring at block boundaries by performing deblocking filtering on the currently restored block. The deblocking filtering process can include a process of determining a transform block boundary, determining a sub-block boundary, determining a length of a filter to perform filtering, determining a filtering strength (bS), determining a filtering parameter, determining whether to perform filtering, and determining a filtering type.

[0290] In independent scalar quantization, the reconstructed coefficient t'k for an input coefficient tk depends only on the associated quantization index qk. That is, the quantization index for any reconstructed coefficient has a different value from the quantization indices for other reconstructed coefficients. Here, t'k may be a value of tk that includes quantization error, and may be different or the same depending on the quantization parameter. Here, t'k can be named a reconstructed transform coefficient or an inverse quantized transform coefficient, and the quantization index can also be named a quantized transform coefficient.

[0291] In Uniform Reconstruction Quantizers (URQ), the reconstructed coefficients are spaced at equal intervals. The distance between two adjacent reconstructed values ​​is referred to as the quantization step size. Reconstructed values ​​can include zeros, and the entire set of available reconstructed values ​​can be uniquely defined by the quantization step size. The quantization step size can vary depending on the quantization parameter.

[0292] In conventional methods, quantization reduces the set of acceptable reconstructed transform coefficients, and the number of elements in this set can be finite. This limits the ability to minimize the average error between the original and reconstructed images. Vector quantization can be used as a method to minimize this average error.

[0293] A simple form of vector quantization used in video encoding is sign data hiding. This is a method in which the encoder does not encode the sign of a non-zero coefficient, and the decoder determines the sign of the coefficient based on whether the sum of the absolute values ​​of all coefficients is even or odd. To achieve this, the encoder may increase or decrease the value of at least one coefficient, and at least one coefficient may be selected and adjusted to be optimal in terms of rate-distortion cost. In one implementation, a coefficient having a value close to the boundary of the quantization interval may be selected.

[0294] Another vector quantization method is Trellis-Coded Quantization, which is used in video coding as an optimal path search technique to obtain optimized quantization values ​​in dependent quantization. Block-wise, quantization candidates for all coefficients in the block are arranged in a trellis graph, and the optimal trellis path between the optimized quantization candidates is searched for by considering the cost for rate-distortion. Specifically, dependent quantization applied to video coding can be designed so that the set of allowable reconstructed transform coefficients for a transform coefficient depends on the value of the transform coefficient preceding the current transform coefficient in the restoration order. At this time, by selectively using multiple quantizers according to the transform coefficients, the average error between the original image and the reconstructed image can be minimized, thereby increasing the encoding efficiency.

[0295]

[0296] FIG. 25 is a diagram illustrating a process of generating a prediction block using DIMD according to one embodiment of the present invention.

[0297] Referring to Figure 25, the decoder can derive a predicted block using surrounding samples (blocks, pixels). At this time, the surrounding samples may be blocks (pixels) surrounding the current block. Specifically, the decoder can determine intra-prediction modes and weight information for restoring the current block through a histogram of directional information (angle information) using the surrounding samples as input.

[0298] FIG. 26 is a diagram showing the locations of surrounding pixels used to derive directional information according to one embodiment of the present invention.

[0299] Figure 26 (a) shows when all neighboring blocks of the current block are available for deriving directional information, Figure 26 (b) shows when the upper boundary of the current block is a sub-picture, slice, tile, or CTU boundary, and Figure 26 (c) shows when the left boundary of the current block is a sub-picture, slice, tile, or CTU boundary. On the other hand, if the neighboring blocks and the current block do not belong to the same sub-picture, slice, tile, or CTU, the neighboring blocks may not be used for deriving directional information. The gray dots in Figure 26 indicate the locations of pixels actually used for deriving directional information, and the dotted lines indicate sub-picture, slice, tile, and CTU boundaries. Additionally, referring to (d) to (f) of FIG. 26, pixels located at the boundary may be padded by one pixel outside the boundary to derive directional information. Such padding may enable more accurate directional information to be derived.

[0300] In order to derive direction information for a pixel at a specific location, a 3x3 Sobel filter of Equation 1 can be applied in the horizontal and vertical directions, respectively. A in Equation 1 can denote pixel information (values) of the reconstructed neighboring blocks of the current block of 3x3 size. And the direction information (θ) can be determined using Equation 2. In order to reduce the computational complexity for deriving direction information, the decoder can derive the direction information (θ) only by calculating Gy / Gx of Equation 1 without calculating the atan function of Equation 2.

[0301]

[0302]

[0303] Referring to Fig. 26, directional information can be calculated for each gray dot shown in Fig. 26, and the directional information can be mapped to an angle of an intra prediction mode. The intra prediction mode set can include a planar mode, a DC mode, and a plurality of (e.g., 65) angular modes (i.e., directional modes). The intra prediction modes can be 67 modes, and the directional information (angle, θ) calculated through mathematical expression 2 can be a value in real units. Therefore, a process of mapping directional information to a specific intra prediction directional mode is required. The directional mode described herein can be the same as the angular mode.

[0304] FIG. 27 is a diagram illustrating a method for mapping directional modes according to one embodiment of the present invention.

[0305] Referring to FIG. 27, the intra prediction directional mode can be divided into four sections based on 0 degrees (index 18), 45 degrees (index 34), 90 degrees (index 50), and 135 degrees (index 66) (see FIG. 6). Referring to FIG. 10, the sections for determining the intra prediction directional mode can be divided into four sections from section 0 to section 3. Section 0 can be from -45 degrees to 0 degrees, section 1 can be from 0 degrees to 45 degrees, section 2 can be from 45 degrees to 90 degrees, and section 3 can be from 90 degrees to 135 degrees. In this case, each section can include 16 intra prediction directional modes. The directional mode can be determined as one of the four sections by comparing the signs and magnitudes of Gx and Gy calculated through mathematical expression 1. For example, if Gx and Gy are positive and the absolute value of Gx is greater than the absolute value of Gy, interval 1 can be selected. The intra prediction directional mode mapped to each interval can be determined through the directional information (θ) calculated from mathematical expression 2. Specifically, the decoder extends the value by multiplying the directional information (θ) by 2^16. Then, the decoder can compare the extended value with the numbers in the predefined table to find the value closest to the extended value and determine the intra prediction directional mode based on the closest value. At this time, there can be 17 values ​​in the predefined table. Specifically, the values ​​of the predefined table can be {0, 2048, 4096, 6144, 8192, 12288, 16384, 20480, 24576, 28672, 32768, 36864, 40960, 47104, 53248, 59392, 65536}. At this time, the difference between the predefined table values ​​can be set differently depending on the difference between the angles of the intra prediction directional mode.

[0306] Meanwhile, if the directional angle is obtained using only Gy / Gx without performing atan calculation to reduce the computational complexity, the difference between the predefined table values ​​may not match the distance between the angles of the intra prediction directional mode. atan has a characteristic that the slope gradually decreases as the input value increases. Therefore, the above-defined table should also be set in numerical value by considering not only the difference between the angles of the intra prediction directional mode but also the nonlinear characteristic of atan. For example, the difference between the above-defined table values ​​may be set to gradually decrease. Conversely, the difference between the above-defined table values ​​may be set to gradually increase.

[0307] If the width and height of the current block are different, the available intra prediction directional modes may vary. That is, if the width and height of the current block are different, the interval for deriving the intra prediction directional mode may vary. In other words, the interval for deriving the intra prediction directional mode may change based on the width and height of the current block (e.g., the ratio of the width and height). For example, if the width of the current block is longer than the height, the intra prediction mode may be remapped to 67 to 80, and the intra prediction mode in the opposite direction may be excluded to 2 to 15. For example, if the width of the current block is n (an integer) times longer than the height (e.g., 2), the intra prediction modes {3, 4, 5, 6, 7, 8} may be remapped (mapped) to {67, 68, 69, 70, 71, 72}, respectively. Additionally, if the horizontal length of the current block is longer than the vertical length, the intra prediction mode may be reset to a value obtained by adding '65' to the intra prediction mode. On the other hand, if the horizontal length of the current block is shorter than the vertical length, the intra prediction mode may be reset to a value obtained by subtracting '67' from the intra prediction mode.

[0308] A histogram can be used to derive an intra-prediction directional mode for restoring the current block. If, after obtaining directional information for surrounding blocks, there are more non-directional blocks than directional blocks, the prediction mode for the non-directional block may have the highest cumulative value in the histogram. However, since a directional mode must be derived for restoring the current block, the prediction mode for the non-directional block may be excluded even if it has the highest cumulative value in the histogram. That is, a smooth region with no gradient or directionality between surrounding pixels may not be used to derive an intra-prediction directional mode. For example, the prediction mode for a block without directional information may be planar or DC. If the left neighboring block is planar or DC, the left neighboring block may not be used to derive directional information, and directional information may be derived using only the upper neighboring block. When the surrounding blocks of the current block contain a mixture of smooth regions and directional regions, the decoder can generate a histogram using the G value calculated as in Equation 3 to emphasize the directionality. In this case, the histogram may be a cumulative value in which the calculated G value is added to each generated intra prediction directional mode, rather than a frequency-based one in which '1' is added to each generated intra prediction directional mode.

[0309]

[0310] FIG. 28 is a diagram showing a histogram for deriving an intra prediction directional mode according to one embodiment of the present invention.

[0311] The X-axis of Fig. 28 represents an intra prediction directional mode, and the Y-axis represents the accumulated value of G values. The decoder can select an intra prediction directional mode having the largest accumulated value of G values ​​among the intra prediction directional modes. In other words, the decoder can select an intra prediction directional mode for the current block based on the accumulated value. Referring to Fig. 35, modeA having the largest accumulated value and modeB having the second largest accumulated value can be selected as the intra prediction directional modes. The decoder can generate a final prediction block by weighting the prediction block generated with modeA, the prediction block generated with modeB, and finally the prediction block generated with the planar mode to generate a prediction block for the current block. At this time, the weight of each prediction block can be determined using the accumulated value of modeA and modeB. For example, the weight for the prediction block generated with the planar mode can be set to 1 / 3 of the total weight. The weight for the prediction block generated by modeA can be set to a weight corresponding to the value obtained by dividing the modeA accumulated value by the sum of the accumulated values ​​of modeA and modeB. The weight for the prediction block generated by modeB can be determined as a difference value between the modeA weight and 1 / 3 of the total weight. To make the calculation of the weight more accurate, the decoder can expand the range of the weight by multiplying the weight for the prediction block generated by modeA by an arbitrary value. The weight for the prediction block generated by modeB and the weight for the prediction block generated in the planar mode can also be expanded in the same way.

[0312] FIG. 29 is a diagram illustrating a method for generating a prediction sample using intra prediction directional mode information and weights according to one embodiment of the present invention.

[0313] Specifically, Fig. 29 illustrates the 'Intra prediction' and 'weighted prediction' processes described above. Referring to Fig. 29, when there are multiple intra prediction directional modes induced by the decoder, the decoder can obtain a prediction sample by performing weighted prediction using the weight information of each of the multiple intra prediction directional modes. The weight information can be reset based on at least one of the horizontal length of the current block, the vertical length, the quantization parameter information, and information on whether the current block is luminance or chrominance (Additional information).

[0314] When DIMD is applied to the current block, the video signal processing device can perform the following process.

[0315] 1) The video signal processing device can generate left, top, and top-left histograms by deriving directionality from the restored samples at the left, top, and top-left positions adjacent to the current block, respectively. The video signal processing device can generate an entire histogram by deriving directionality from all of the restored samples at the left, top, and top-left positions. At this time, the video signal processing device can select the top five directionality with high frequency values ​​using the entire histogram.

[0316] 2) The video signal processing device can classify each of the top five directional modes into vertical, horizontal, and diagonal characteristics, depending on whether the directional mode is frequently induced from the left sample or from the upper sample. The video signal processing device can classify each directional mode into vertical, horizontal, and diagonal characteristics by comparing the frequency values ​​in the entire histogram of each directional mode with the frequency values ​​in the left, upper, and upper-left histograms of each directional mode. This can be referred to as a position-specific feature weight.

[0317] 3) The video signal processing device may determine whether to apply a weighted average to the current block by using at least one of the directional modes, the frequency value for the directional modes, and whether the directional modes have a characteristic (vertical, horizontal, or diagonal). If the frequency value of the upper second directional mode is greater than 0, and the first and second upper directional modes are angular modes (not planar mode or DC mode), the video signal processing device may apply the weighted average to the current block. Otherwise, the video signal processing device may not apply the weighted average to the current block. If the upper first mode is not a diagonal characteristic, the frequency value of the upper second directional mode is '0', and the upper first directional mode is angular mode, the video signal processing device may set the upper second directional mode to the planar mode and apply the weighted average by using a weight of 3:1 between blocks predicted by the first directional mode and the second planar mode. The video signal processing device can set the weight for each directional mode based on the frequency value for the directional mode and the sum of the frequency values.

[0318] 4) After the video signal processing device generates a prediction block using each directional mode, the video signal processing device can perform a first block-by-block weighted average between the prediction blocks of the directional modes classified with the same characteristic using the previously calculated weights. Through this, the video signal processing device can generate a vertical characteristic prediction block, a horizontal characteristic prediction block, and a diagonal characteristic prediction block.

[0319] 5) The video signal processing device can calculate a vertical feature weight, a horizontal feature weight, and a diagonal feature weight by calculating a weighted sum between the weights for the directional modes classified by the same feature. The video signal processing device can perform a second sample-by-sample weighted average between the first weighted averaged blocks. At this time, the video signal processing device can determine the mode of the second sample-by-sample weighted average by using at least one of a comparison between the horizontal and vertical size difference of the current block, the size of the current block, the vertical feature weight, the horizontal feature weight, and the diagonal feature weight. If both the vertical feature weight and the horizontal feature weight are greater than '0', the mode of the second sample-by-sample weighted average can be determined as the diagonal mode. If the vertical feature weight is greater than '0', the video signal processing device can determine the mode of the second sample-by-sample weighted average as the vertical mode. If the vertical feature weight is not greater than '0', the video signal processing device can determine the mode of the second sample-by-sample weighted average as the horizontal mode. When the mode of the second sample unit weighted average is a vertical mode, the video signal processing device can apply a sample unit weighted average using a vertical feature prediction block, a diagonal feature prediction block, and the vertical feature weights and the diagonal feature weights. At this time, the vertical feature weights may decrease and the diagonal feature weights may increase as the distance from the upper boundary of the current block increases. Alternatively, when the mode of the second sample unit weighted average is a horizontal mode, the video signal processing device can apply a sample unit weighted average using a horizontal feature prediction block, a diagonal feature prediction block, and the horizontal feature weights and the diagonal feature weights. At this time, the horizontal feature weights may decrease and the diagonal feature weights may increase as the distance from the left boundary of the current block increases.Alternatively, if the mode of the second sample-by-sample weighted average is a diagonal mode, the video signal processing device may apply the sample-by-sample weighted average using a horizontal feature prediction block, a vertical feature prediction block, a diagonal feature prediction block, and horizontal feature weights, vertical feature weights, and diagonal feature weights. At this time, the horizontal feature weights and vertical feature weights may decrease and the diagonal feature weights may increase as the distance from the left boundary and the upper boundary of the current block increases.

[0320] FIG. 30 and FIG. 31 are diagrams showing templates used to derive an intra prediction mode of a current block according to one embodiment of the present invention.

[0321] Referring to FIG. 30, the decoder can use a template, which is a reconstructed arbitrary region (pixel(s)) adjacent to the current block, to derive an intra-prediction mode of the current block. First, the decoder can generate a prediction template for the template using neighboring pixels (reference) adjacent to the template. Then, the decoder can use the intra-prediction mode for the prediction template that is most similar to the already reconstructed template to reconstruct the current block. The method of deriving the intra-prediction mode of the current block using the above-described template can be described as TIMD (Template intra-mode derivation). At this time, the intra-prediction mode can be a mode with an index of 0 to 67, and may only be an intra-prediction mode within the MPM list derived from the neighboring blocks of the current block. At this time, the intra-prediction mode can be an intra-prediction mode within the MPM list derived from the neighboring blocks of the current block and modes that differ from the intra-prediction mode by an arbitrary number. The arbitrary number can be 1, 2, 3, ... Alternatively, the intra prediction mode for the template may only correspond to a directional mode, and may not correspond to a non-directional mode (planar mode, DC mode).

[0322] Below, a method for deriving an intra prediction directional mode using the TIMD mode is described.

[0323] i) The decoder can set the size of the template. The width or height of the template can be 4, and if the width or height of the current block is 8 or less, the width or height of the template can be set to 2. ii) The decoder can set the type of the template. The type of the template can be classified into a type that uses only left samples, a type that uses only upper samples, and a type that uses all of the left, upper, and upper-left samples. The decoder can determine the type of the template based on whether the neighboring blocks are valid or whether the neighboring blocks can be used to derive an intra-prediction directional mode. Meanwhile, if the neighboring blocks cannot be used to derive an intra-prediction directional mode, the TIMD mode can be set to a planar mode, and weighted averaging may not be performed. iii) The decoder can configure a template for the current block. iv) The decoder can derive intra prediction directional modes for neighboring blocks located to the left, above, upper-left, upper-right, and lower-left of the current block to determine whether the current block has directionality. v) If none of the neighboring blocks of the current block have directionality (e.g., non-directional mode (DC mode, planar mode, MIP mode, etc.)), the decoder can select one intra prediction directional mode with the minimum cost and not perform the TIMD mode. In this case, weighted averaging using multiple prediction blocks may not be performed. vi) If there is one or more blocks among the neighboring blocks of the current block that have directionality, the process described below can be performed. The process described below can be performed based on the intra prediction directional modes existing in the MPM list. This is because complexity may increase if all 67 intra prediction directional modes are checked. a. The decoder can construct an MPM list. b.Next, the decoder can modify the MPM list by adding DC mode, horizontal mode, and vertical mode to the MPM list if they do not exist in the MPM list. c. The decoder can compare costs by evaluating all intra prediction directional modes in the modified list. The decoder can select a first mode with a lowest cost and a second mode with a second lowest cost. d. The decoder can additionally perform an evaluation for an intra prediction directional mode corresponding to an index that is one less than or greater than the intra prediction directional mode index of the first mode and the intra prediction directional mode index of the second mode to improve accuracy. The decoder can perform an additional evaluation and again select a third mode with a lowest cost and a fourth mode with a second lowest cost. Meanwhile, the first mode and the third mode may be the same, and the second mode and the fourth mode may be the same. e. The decoder can determine whether to perform weighted averaging based on the costs of the third mode and the fourth mode. If the difference between the cost of the third mode and the cost of the fourth mode is less than a specific value, the decoder can perform weighted averaging, and the weights of the third and fourth modes can be determined based on the costs of the third and fourth modes. If the difference between the cost of the third mode and the cost of the fourth mode is greater than a specific value, the decoder can generate a prediction block using only the third mode without performing weighted averaging. In this case, the specific value may be a predetermined value.

[0324] The size of the template may vary depending on the horizontal or vertical length of the current block. For example, as shown in (a) of Fig. 30, an above template may be configured that is longer than the horizontal length of the current block. At this time, the vertical length of the above template may be a predetermined length. Similarly, a left template may be configured that is longer than the vertical length of the current block. At this time, the horizontal length of the left template may be a predetermined length. The predetermined length may be 1, 2, 3, ...

[0325] If the current block is located at the CTU boundary (if any of the upper, lower, left, or right boundaries of the current block is included in the boundary of the CTU), the reference pixels for deriving / predicting the template used for the TIMD mode can be changed. Referring to Fig. 31, if the upper boundary of the current block is included in the boundary of the CTU, there can be only one reference line located at the upper side of the current block used for template construction. This is to minimize line buffer memory usage. Therefore, the decoder can perform the TIMD mode by constructing only the left template of the current block, without constructing the upper template of the current block. At this time, the reference pixels for predicting the left template can be the above reference pixel and the left reference pixel of the current block. At this time, the height of the left template can be the same as the height of the current block, as shown in Fig. 31 (a). In addition, as shown in (b) of Fig. 31, the decoder can check whether the block adjacent to the left of the current block is a block that has already been restored, and if it is a block that has been restored, the height of the left template can be configured to be greater than the height of the current block.

[0326] In general, the accuracy of prediction samples for the current block can be increased as the decoder references more adjacent pixels of the current block. However, referencing more adjacent pixels increases the required memory. Furthermore, if there are blocks adjacent to the current block that have not yet been restored, those areas cannot be used as templates. To effectively handle this increased memory and the unrestored areas, as shown in Figure 30 (b), the length of the upper template can be set to be equal to the horizontal length of the current block, and the length of the left template can be set to be equal to the vertical length of the current block.

[0327] The decoder can use an intra prediction mode derived from a template to obtain prediction samples for the current block. The decoder can generate prediction samples using neighboring pixels adjacent to the current block and adaptively select which neighboring pixels to use for prediction sample generation. Furthermore, the decoder can use multiple reference lines to generate prediction samples, and the index information for these multiple reference lines can be included in the bitstream.

[0328] For entropy coding, a new context for the indices of multiple reference lines for TIMD mode can be defined. The increased number of contexts can be associated with increased memory and context switching complexity. Therefore, the context used to encode and decode the indices of multiple reference lines in TIMD mode can be reused from the existing context for the indices of multiple reference lines.

[0329] The transformation of the residual signal of the current block can be performed in two steps. The first transform can be a transform such as DCT-II, DST-VII, DCT-VIII, DCT5, DST4, DST1, identity transformation (IDT), etc., which are adaptively applied horizontally and vertically, respectively. A second transform can be additionally applied to the transform coefficients for which the first transform has been completed, and the second transform can be calculated as a matrix multiplication between the first-transformed transform coefficients and a predefined matrix. The second transform can be described as a low frequency non-separable transform (LFNST). The matrix transform set for the second transform can vary depending on the intra prediction mode of the current block. Coefficient information of the transform matrix used for the second transform can be included in the bitstream.

[0330] When a secondary transform is applied to a current block to which the DIMD mode or the TIMD mode is applied, a transform set for the secondary transform can be determined based on an intra prediction mode derived from the DIMD mode or the TIMD mode. Coefficient information of a transform matrix used for the secondary transform can be included in a bitstream. A decoder can parse the coefficient information included in the bitstream to set matrix coefficient information of the secondary transform for the DIMD mode or the TIMD mode. At this time, one of the two intra prediction modes derived from the TIMD mode can be used to select the primary transform or the secondary transform set. The costs of each of the two intra prediction directional modes can be compared, and the intra prediction directional mode with the smallest cost can be used to select the primary transform or the secondary transform set. In addition, one of the two intra prediction directional modes derived from DIMD can be used to select the primary transform or the secondary transform set. The weights of each of the two intra prediction modes can be compared, and the intra prediction directional mode with the highest weight can be used to select the first or second transformation set.

[0331] TIMD mode is highly complex because it predicts the template of the current block and uses the intra prediction mode derived from the template to generate the prediction block of the current block. Therefore, when the decoder generates the prediction template for the template region, the existing reference sample filtering process may not be performed. In addition, TIMD mode may not be applied if the ISP mode is applied to the current block or if the CIIP mode is applied to the current block. The ISP mode or CIIP mode may not be applied to the current block to which TIMD mode is applied, or the syntax related to the ISP or CIIP may not be parsed. In this case, the value of the syntax related to the ISP or CIIP that is not parsed can be inferred as a predetermined value.

[0332] Template prediction can be performed by dividing the current block into a left template region and an upper template region, and an intra prediction mode can be derived for each template. In addition, two or more intra prediction modes can be derived for each template, and there can be four or more intra prediction modes for the current block. When there are two or more intra prediction modes, a prediction sample for the current block can be generated using all of the derived intra prediction modes, and the decoder can generate a final prediction block for the current block by performing a weighted average of the generated prediction samples. At this time, at least three or more of the two or more intra prediction modes derived from template prediction, a planar mode, a DC mode, and a MIP mode can be used to generate the prediction sample. For example, when the decoder generates (obtains) a prediction sample for the current block, the decoder can generate a final prediction sample by performing a weighted average of the prediction samples generated using the intra prediction modes derived from template prediction and the planar mode.

[0333] Even when CIIP mode is applied, prediction samples can be generated using the methods described above. CIIP mode utilizes both intra-prediction and inter-prediction when generating prediction samples (blocks) for the current block. The prediction sample for the current block can be generated as a weighted average of the intra-prediction samples and inter-prediction samples.

[0334] When generating intra prediction samples by applying the CIIP mode, the DIMD mode or the TIMD mode can be used. In this case, when the DIMD mode is used, the intra prediction sample can be generated based on the DIMD combination information. For example, the decoder can generate the first prediction sample using the intra prediction mode with the highest weight and the second prediction sample using the intra prediction mode with the second highest weight. Then, the decoder can generate the final intra prediction block by performing a weighted average of the first prediction sample and the second prediction sample. In this case, the decoder can generate the final intra prediction block by performing a weighted average of three prediction samples, including the planar mode predicted sample, the first prediction sample, and the second prediction sample among the neighboring blocks of the current block. In this case, the intra prediction sample can be generated based on the TIMD combination information. For example, the decoder can generate two prediction samples using each of the two intra prediction modes. Then, the decoder can generate a final intra prediction sample by weighting the two prediction samples. At this time, the decoder can generate a final intra prediction sample by weighting the two prediction samples and the sample predicted in planar mode.

[0335] Intra prediction samples may have varying accuracy depending on their location. That is, pixels located farther away from the surrounding pixels used for prediction within a prediction sample may contain more residual signals than pixels located closer to the surrounding pixels. Therefore, the decoder can divide the prediction sample into vertical, horizontal, and diagonal directions depending on the direction of the intra prediction mode, and set different weights depending on the distance from the surrounding pixels used for prediction. This can be applied to intra prediction blocks generated using the CIIP mode or intra prediction blocks generated using two or more intra prediction modes, and different weights can be set for each pixel within the prediction block depending on the distance between the location of the reference pixel and the pixel location within the prediction block. For example, if the intra prediction mode of the current block is a mode having a vertical or vertical-like direction, a higher weight can be set for each pixel location within the prediction block as the pixel location is closer to the top pixel, and a lower weight can be set for each pixel location as the pixel location is farther away from the top pixel.

[0336] When the current block is encoded in CIIP mode, the decoder can generate a final prediction block by weighting the intra-prediction samples and inter-prediction samples. The pixel-wise weights of the inter-prediction samples can be set by considering the pixel-wise weights of the intra-prediction samples. For example, the pixel-wise weights of the inter-prediction samples can be a value obtained by subtracting the pixel-wise weights of the intra-prediction samples from the sum of the total weights. In this case, the sum of the total weights can be a value obtained by adding the weights of the intra-prediction samples and the weights of the inter-prediction samples at the pixel level.

[0337] When two or more intra prediction modes are used to generate prediction samples, the decoder can generate prediction samples based on each intra prediction mode, and then weight and average the generated prediction samples to generate a final prediction sample. When generating prediction samples for each intra prediction mode, pixel-wise weights according to the intra prediction mode can be applied.

[0338] The pixel-level weights may be set based on at least one of the following: an intra prediction mode, the horizontal length and vertical length of the current block, a quantization parameter, information about whether the current block is luminance or chrominance, whether the surrounding block is intra coded, and information about the presence or absence of residual transform coefficients of the surrounding block.

[0339] FIG. 32 is a diagram illustrating a method for generating prediction samples (pixels) based on a plurality of reference pixel lines according to one embodiment of the present invention.

[0340] Referring to (a) of FIG. 32, the video signal processing device can generate a prediction sample (4102) within the current block based on a first reference pixel line (reference line 1) adjacent to the current block (4101) and a second reference pixel line (reference line 2) adjacent above the first reference pixel line. The prediction sample (4102) of FIG. 32 (a) is only a sample corresponding to a position according to an embodiment of the present invention, and the position of the pixel is not limited thereto. In the present specification, the meaning of “generate” by the video signal processing device may be the same as the meaning of “acquire” by the video signal processing device. (b) of FIG. 32 is a drawing showing (a) of FIG. 32 in more detail. For example, the video signal processing device can generate a first prediction pixel (4103) using a smoothing filter, a cubic filter, or a Gaussian filter according to an intra prediction mode through six reference pixels of the first reference pixel line. And, the video signal processing device can generate a second prediction pixel (4104) using a smoothing filter, a cubic filter, or a Gaussian filter according to the intra prediction mode through the six reference pixels of the second reference pixel line. The video signal processing device can generate a third prediction pixel (4105) by performing a weighted average through an arbitrary set of weights on the generated first prediction pixel (4103) and second prediction pixel (4104). At this time, the six reference pixels of the second reference pixel line may be reference pixels at a position shifted by one pixel to the right of each pixel of the first reference pixel line in consideration of the intra prediction mode of the current block, the position of the pixel to be generated, the position of the reference pixel line, etc. The video signal processing device can generate a prediction sample (4102) within the current block based on the third prediction pixel (4105).Alternatively, the video signal processing device may generate a prediction sample (4102) within the current block based on the third prediction pixel (4105) and the distance between the third prediction pixel (4105) and the prediction sample (4102) within the current block. The weight used by the video signal processing device to generate the third prediction pixel (4105) may be an integer greater than or equal to 0. For example, the weight of the first prediction pixel (4103) may be 3, and the weight of the second prediction pixel (4104) may be 1. At this time, the positions of the reference pixels used to generate the prediction sample (the positions of the six reference pixels of the first reference pixel line and the six reference pixels of the second reference pixel line in FIG. 41) may vary based on at least one of the following: the intra prediction mode of the current block, the positions of the pixels to be generated (e.g., the positions of the first prediction pixel (4103), the second prediction pixel (4104), the third prediction pixel (4105) in FIG. 41), the positions of the reference pixel lines (e.g., the positions of the first reference pixel line and the second reference pixel line in FIG. 41).

[0341] The video signal processing device can obtain a prediction sample within a current block according to the directionality of the intra prediction mode by using a sample of a pre-specified or signaled reference sample line. Here, the pre-specified reference sample line may be a reference sample adjacent to the current block. In addition, the signaled reference sample line may be a sample at a position -X samples from the boundary of the current block, and X may be a signaled value. X may be 1, 2, 3, etc. The video signal processing device can generate a prediction block by using a plurality of reference sample lines. In the present specification, the pre-specified reference sample line or the signaled reference sample line may be referred to as a main reference sample line. The video signal processing device can set a sub-reference sample line based on the main reference sample line. In the present specification, the sub-reference sample line may be a reference sample line at a position Y samples apart from the main reference sample line. In this case, Y may be an integer, and may be -3, -2, -1, 1, 2, 3, etc. An encoder can signal information about a sub-reference sample line and include it in the bitstream, and a decoder can parse the information to set the sub-reference sample line. In a video signal processing device, a prediction block can be generated from each of the main reference pixel line and the sub-reference sample line using one intra prediction mode, and then a final prediction block can be generated by weighting and averaging each prediction block. This method can be called intra fusion.

[0342] FIG. 33 illustrates a method for predicting a sample using a plurality of reference pixel lines according to one embodiment of the present invention.

[0343] Referring to FIG. 33, a video signal processing device may receive a plurality of reference pixel lines as input and perform intra prediction to generate a prediction block within a current block. Depending on which reference pixel line is used, a different prediction block may be generated, and the video signal processing device may generate a final prediction block by performing a weighted average according to the weights input for each prediction block. At this time, the weights may be preset values. For example, the weight of a sample predicted by a main reference pixel line may be 3, and the weight of a sample predicted by a sub-reference pixel line may be 1. At this time, the weights may be determined based on at least one or more of the size of the current block, the horizontal or vertical size of the current block, the intra-prediction mode of the current block, quantization parameter information, and the distance (or difference) between the main reference pixel line and the sub-reference pixel line. In addition, the reference pixel line may be determined based on at least one or more of the size of the current block, the horizontal or vertical size of the current block, the intra-prediction mode of the current block, quantization parameter information, and MRL information. For example, the main reference pixel line may be a reference pixel line adjacent to the current block, and the sub-reference pixel line may be a reference pixel line indicated by the MRL. As another example, the main reference pixel line may be a reference pixel line indicated by the MRL, and the sub-reference pixel line may be a reference pixel line that is located at an arbitrary fixed position relative to the reference pixel line indicated by the MRL, and the arbitrary fixed position may be an integer from -N to +N, where N may be an integer greater than 0.

[0344] The intra prediction mode used in the method for generating prediction samples within the current block described above may be the same for each reference pixel line. Or, conversely, the intra prediction mode used in the method for generating prediction samples within the current block described above may be different for each reference pixel line. That is, a signaled intra prediction mode may be used in the main reference pixel line, and a prediction mode (corresponding to the index) that is added or subtracted by an arbitrary value from (the index of) the intra prediction mode used in the main reference pixel line may be used in the sub-reference pixel line. At this time, the arbitrary value may be an integer greater than or equal to 1. In addition, the video signal processing device may determine whether to increase or decrease the arbitrary value depending on the value of the intra prediction mode used in the main reference pixel line. For example, the video signal processing device may increase the intra prediction mode by an arbitrary value when the angle is negative, and may decrease the intra prediction mode by an arbitrary value when the angle is positive.

[0345] FIG. 34 is a structural diagram illustrating a method for determining an optimal reference pixel line using a plurality of reference pixel lines based on a template according to one embodiment of the present invention.

[0346] Referring to FIG. 34, an encoder can receive multiple reference pixel lines as input and perform intra prediction to generate prediction blocks for a template. Depending on which reference pixel line is used, different prediction blocks can be generated. The encoder can perform weighted averaging based on various weight information input to each prediction block to ultimately generate a prediction block for the template. Depending on which reference pixel lines are used and which weights are used, multiple prediction blocks can be generated. After calculating the cost between each prediction block and the reference template, the encoder can rearrange the prediction blocks in ascending order based on the cost corresponding to each prediction block and construct a separate list using only a predetermined number of top candidates. In this case, the predetermined number can be an integer greater than or equal to 2, and can be 10. The encoder can generate a prediction block for the current block using the combination information used to generate the prediction blocks in the separate list. The encoder can then select the optimal candidate from the list in terms of image quality and bit rate, and then generate and signal a bitstream containing information about the index of the optimal candidate. The decoder can then construct a separate list identical to the above-described method, parse the information about the optimal candidate index contained in the bitstream, and use the optimal combination information indicated by the index of the optimal candidate to generate prediction samples.

[0347] When the current block is encoded using an intra prediction mode, the encoded intra prediction mode can be any one of angular mode, planar mode, DC mode, and MIP mode. Prediction according to the angular mode can be performed according to 65 different angles, and prediction according to the MIP mode can be performed based on a predefined matrix. The angular mode can be effective in blocks that have characteristics such as edges within the current block. However, if the current block has smooth characteristics, blocks predicted using the angular mode can generate discontinuous edges at the boundary between blocks or visible outlines within the block. This can be a factor that reduces encoding efficiency. In addition, the DC mode can have the disadvantage of generating visible edges at the boundary between blocks at low bit rates. The planar mode can improve the edge problems that occur in the angular and DC modes, and can generate predicted blocks without discontinuities.

[0348] FIG. 35 illustrates a method for generating prediction samples using a planar mode according to one embodiment of the present invention.

[0349] Referring to FIG. 35, according to the planar mode, the video signal processing device can generate a linearly predicted value in the vertical direction and a linearly predicted value in the horizontal direction to generate a prediction sample within the current block. The video signal processing device can generate a prediction sample (value) within the current block by weighting the linearly predicted value in the vertical direction and the linearly predicted value in the horizontal direction.

[0350] A linearly predicted value in the vertical direction (predV (x, y)) can be generated based on Equation 4, and a linearly predicted value in the horizontal direction (predH (x, y)) can be generated based on Equation 5. And, a new predicted value (pred (x, y)) can be generated based on Equation 6. In Equations 4 to 6, W may be the horizontal size (width) of the current block, and H may be the vertical size (height) of the current block. rec(x, y) may mean a pixel value at the (x, y) coordinate. Predicted values ​​(predV(x, y), predH(x, y), pred(x, y)) may mean a predicted pixel value at the (x, y) coordinate.

[0351]

[0352]

[0353]

[0354] A video signal processing device may use only vertical linear prediction when performing prediction related to a current block according to a planar mode. Alternatively, the video signal processing device may use only horizontal linear prediction when performing prediction related to a current block according to a planar mode. Therefore, the planar mode can be divided into three modes. That is, in addition to the method of weighting and averaging prediction blocks generated using conventional vertical and horizontal linear prediction, it can be divided into a vertical planar mode that uses only vertical linear prediction, and a horizontal planar mode that uses only horizontal linear prediction. An encoder can generate and signal a bitstream including information on which of the three planar modes a current block used. A decoder can generate a prediction block for the current block based on a planar mode determined by parsing information on which prediction mode was used included in the bitstream.

[0355]

[0356] Figure 36 illustrates an intra prediction mode based on an extrapolation filter according to one embodiment of the present invention.

[0357] Specifically, FIG. 36 illustrates an EIP (Extrapolation filter-based Intra Prediction) mode, which is an intra prediction mode based on an extrapolation filter according to one embodiment of the present invention.

[0358] A method for performing EIP mode by a video signal processing device according to an embodiment of the present invention is described through FIG. 36.

[0359] Referring to (a) of Fig. 36, a video signal processing device can derive non-linear model parameters using a template configured using neighboring samples adjacent to a current block. At this time, there may be three types of templates, and the video signal processing device can derive non-linear model parameters using one of the three types. The template may be configured (L-shape) with left neighboring samples and upper neighboring samples of the current block, the template may be configured (left-only) with left samples of the current block, and the template may be configured (above-only) with upper samples of the current block. The non-linear model may be CCCM. Referring to (b) of Fig. 36, a filter shape used to derive non-linear model parameters may be one of a square shape, a horizontally elongated (longer than wide) rectangular shape, and a vertically elongated (longer than wide) rectangular shape. An encoder can obtain a bitstream including first information indicating a template type and a filter shape, and a decoder can determine the template type and the filter shape through the first information of the bitstream. Referring to (c) of FIG. 36, a video signal processing device can derive non-linear model parameters based on the template type and the filter shape, and can predict a current block sample by sample using the derived non-linear model parameters. When predicting the current block sample by sample, the video signal processing device can predict in a diagonal direction. When predicting the current sample, a sample predicted before the current sample can be used. The video signal processing device can perform filtering on a predicted block using an EIP mode. The encoder can generate and signal a bitstream including information indicating whether to apply filtering to a block predicted in the EIP mode.In the decoder, if the current block is in EIP mode, information indicating whether filtering is applied to the predicted block can be parsed to set whether filtering is applied to the predicted block. At this time, the filtering may be as shown in Table 1 below. Table 1 shows filtering related to a 3x3 kernel and a 5x5 kernel according to one embodiment of the present invention.

[0360]

[0361] A video signal processing device can derive an intra prediction mode using a block predicted in the EIP mode for a residual signal of a block encoded in the EIP mode, and then derive a transform kernel for MTS, NSPT, LFNST, etc. using the derived intra prediction mode. At this time, in order to derive the intra prediction mode, the video signal processing device can use the DIMD method or a method similar to DIMD.

[0362] The EIP merge mode may be a mode in which a video signal processing device constructs an EIP merge list using filter shapes and filter coefficients of neighboring blocks previously encoded in the EIP mode, and then determines an EIP merge candidate having an optimal filter shape and filter coefficients for the current block. At this time, the EIP merge candidate (optimal filter shape and filter coefficient) may be determined through an index included in the bitstream. At this time, the EIP merge list may be constructed using adjacent spatial candidates, non-adjacent spatial candidates, temporal candidates, shifted temporal candidates, and history-based candidates of the current block. In addition, the EIP merge list may be reordered based on a template cost.

[0363] When the current block is encoded in any one of the intra TMP, MIP, and EIP modes, the video signal processing device may store any one of the planar mode, the DC mode, the derived DIMD mode, and the intra prediction mode at the position indicated by the block vector (BV) for the intra prediction mode of the current block. The derived DIMD mode may be an intra prediction mode derived by the DIMD method using pre-reconstructed samples adjacent to the current block. Alternatively, the derived DIMD mode may be an intra prediction mode derived by the DIMD method using the current prediction block. Alternatively, the derived DIMD mode may be an intra prediction mode derived by the DIMD method using samples of the current block that have been reconstructed after the reconstruction of the current block is completed. The intra prediction mode indicated by the BV may be an intra prediction mode of a reference block indicated by the BV used for intra TMP prediction. The stored intra prediction mode may be used to derive an MPM list for intra prediction of the next block. Additionally, the video signal processing device can use the intra prediction mode of the reference block indicated by the BV to derive an intra prediction mode to be stored for a block encoded in the IBC or intra TMP mode.

[0364] When a video signal processing device derives an MPM list for a current block, if the encoding mode of a block adjacent to the current block is any one of the intra TMP, MIP, and EIP modes, the intra prediction mode of the block adjacent to the current block can be derived as any one of the planar mode, the DC mode, the derived DIMD mode, and the intra prediction mode at the position indicated by the BV.

[0365] The video signal processing device may set a value obtained by converting an angle of a diagonal line dividing the current block into an intra prediction mode as the intra prediction mode of the current block when the current block is encoded in either the Spatial Geometry Partitioning Mode (SGPM) or the Geometry Partitioning Mode (GPM). When the current block is encoded in either the SGPM or the GPM, and one of the two regions of the current block is encoded in the intra prediction mode, the intra prediction mode representing the current block may be the intra prediction mode applied to the region encoded in the intra prediction mode among the two regions. At this time, when the intra prediction mode is a directional mode (any one of modes 2 to 66), the intra prediction mode representing the current block may be an intra prediction mode corresponding to the directional mode (any one of modes 2 to 66). If the intra prediction mode is not a directional mode (one of modes 2 to 66), the intra prediction mode representing the current block can be a value obtained by converting the angle of the diagonal line used in the segmentation process of the current block into the intra prediction mode.

[0366] The encoding mode of a neighboring block of a current block is one of SGPM and GPM, and one of two divided regions of the current block can be encoded with an intra prediction mode. At this time, when the video signal processing device derives an MPM list for the current block, the intra prediction mode of a region encoded with an intra prediction mode among two divided regions of the neighboring block of the current block can be set to the intra prediction mode of the neighboring block. At this time, if the intra prediction mode is a directional mode (any one of modes 2 to 66), the intra prediction mode of the neighboring block can be an intra prediction mode of the directional mode (any one of modes 2 to 66). If the intra prediction mode is not a directional mode (any one of modes 2 to 66), the value can be obtained by converting the angle of the diagonal line used in the segmentation process of the current block into the intra prediction mode. The derived intra prediction mode of the neighboring block can be added to the MPM list.

[0367] If the current block is encoded in the IBC mode, the intra prediction mode of the current block can be one of the derived DIMD mode and the intra prediction mode at the position indicated by the BV. The derived DIMD mode can be an intra prediction mode derived by the DIMD method using already reconstructed samples adjacent to the current block. Alternatively, the derived DIMD mode can be an intra prediction mode derived by the DIMD method using the current prediction block. The derived DIMD mode can be an intra prediction mode derived by the DIMD method using samples of the current block that have been reconstructed after the current block has been reconstructed. The intra prediction mode indicated by the BV can be an intra prediction mode of a reference block indicated by the BV that is used for IBC prediction. If the coding mode of the current block is the RRIBC coding mode, the intra prediction mode of the reference block indicated by the BV can be converted according to the flip direction of the current block, and the intra prediction mode of the current block can be the converted intra prediction mode.

[0368]

[0369] Figure 37 illustrates a method for selecting an optimal intra prediction mode combination according to one embodiment of the present invention.

[0370] Figures 38 and 39 illustrate neighboring blocks of a current block according to one embodiment of the present invention.

[0371] The optimal combination of intra prediction modes can be determined based on the occurrence frequency of the intra prediction modes. OBIC (Occurrence-based intra coding), an intra coding mode based on the occurrence frequency of the intra prediction modes, can be performed in the following order.

[0372] First, the video signal processing device can determine whether the OBIC mode is applicable to the current block. Whether the OBIC mode is applicable can be determined using at least one of the position of the current block, whether the current block is a luminance block or a chrominance block, whether the current block is encoded in an intra prediction mode, whether DIMD is applicable, the horizontal and vertical sizes of the current block, and the number of blocks used to derive the OBIC mode among the neighboring blocks of the current block. If the number of blocks used to derive the OBIC mode among the neighboring blocks of the current block is less than a predetermined number, the OBIC mode may not be applied to the current block, and syntax elements related to OBIC may not be signaled or parsed.

[0373] When the OBIC mode is applicable to the current block, the video signal processing device can derive the intra prediction mode from the neighboring blocks of the current block. At this time, the neighboring blocks may be blocks adjacent to the current block. For example, the blocks at positions A0, A1, B0, B1, and B2 in FIG. 38 may be neighboring blocks. Alternatively, the neighboring blocks may be blocks that are not adjacent to the current block. For example, the blocks at positions Ax, Bx, Cx, ax, bx, and cx in FIG. 39 may be neighboring blocks. When the neighboring blocks of the current block are encoded in the IBC or Intra TMP mode, the reference block indicated by the BV of the current block can be used as the neighboring block for the OBIC mode. In addition, when the current picture is not an I picture but a P picture or a B picture, a temporal neighboring block corresponding to the current block from another reference picture that is not identical to the POC (picture order count) of the current picture can also be used as the neighboring block for the OBIC mode. A temporal neighboring block may be a block whose position has been moved using motion information from blocks adjacent to the current block.

[0374] Whether neighboring blocks can be used as neighboring blocks for deriving an intra prediction mode for the OBIC mode can be determined based on one or more of the encoding mode of the neighboring block and the distance between the neighboring block and the current block. If the neighboring block is not encoded in the intra mode, the neighboring block cannot be used as neighboring blocks for the OBIC mode. Even if the neighboring block is encoded in the intra mode, if the encoding mode is one of the EIP, Intra TMP, and MIP modes, the neighboring block cannot be used as neighboring blocks for the OBIC mode. Alternatively, if the neighboring block is encoded in the intra mode and the encoding mode is one of the EIP, Intra TMP, and MIP modes, the neighboring block can be used as neighboring blocks for the OBIC mode. Alternatively, if the neighboring block is encoded in the IBC or Inter encoding mode, the neighboring block can be used as neighboring blocks for the OBIC mode.

[0375] The neighboring blocks available for OBIC mode can be sorted based on their distance from the current block. Only a pre-specified number of neighboring blocks in the sorted order can be used as neighboring blocks for OBIC mode. The pre-specified number is an integer greater than or equal to 1, and can be 20.

[0376] Next, the video signal processing device can derive an intra prediction mode from the selected surrounding block. Then, the video signal processing device can generate a histogram for the intra prediction mode based on the value of the derived intra prediction mode and the product of the horizontal length and the vertical length of the surrounding block (or the sum of the horizontal and vertical lengths of the surrounding block). At this time, the value of the derived intra prediction mode can be an X-axis component (which can be described as a class in this specification) value of the histogram, and the product of the horizontal length and the vertical length of the surrounding block (or the sum of the horizontal and vertical lengths of the surrounding block) can be a Y-axis component (which can be described as a degree in this specification) value of the histogram.

[0377] If the intra prediction mode of the surrounding block is DIMD or TIMD mode, the intra prediction mode used for combination (blending) can also be added to the histogram. Alternatively, if the intra prediction mode of the surrounding block is OBIC mode, a predetermined number of intra prediction modes can be added to the histogram, and the added intra prediction mode can be an intra prediction mode other than the first intra prediction mode (the mode with the highest frequency in the histogram). In this case, when the intra prediction mode is added to the histogram, the frequency value can be the product of the horizontal length and the vertical length of the surrounding block (or the sum of the horizontal and vertical lengths of the surrounding block). Alternatively, the frequency can be a value smaller (or larger) than the product of the horizontal length and the vertical length of the surrounding block (or the sum of the horizontal and vertical lengths of the surrounding block). For example, a small value (or large value) can be half (or twice) the product of the horizontal length and the vertical length of the surrounding block (or the sum of the horizontal and vertical lengths of the surrounding block).

[0378] If the intra prediction mode of the surrounding blocks is the SGPM mode, the intra prediction mode of each region can be added to the histogram, and additionally, a value obtained by converting the segmentation angle into the intra prediction mode can be added to the histogram.

[0379] If the encoding mode of the surrounding blocks is inter-prediction mode or IBC mode, the intra-prediction mode stored in the surrounding blocks can be added to the histogram. If the intra-prediction mode of the surrounding blocks is GPM mode, the intra-prediction mode of each region can be added to the histogram, and additionally, the value of the segmentation angle converted to the intra-prediction mode can be added to the histogram.

[0380] When a video signal processing device generates a histogram, the histogram frequency value for the intra prediction mode derived from the surrounding blocks may vary depending on the encoding mode of the surrounding blocks. The encoding mode and the frequency value may be agreed upon in advance. Here, the pre-agreed encoding mode may be one of the intra prediction modes based on DIMD, TIMD, TMRL, SGPM, and MPM. In addition, the pre-agreed frequency value may be the product of the horizontal length and the vertical length of the surrounding blocks (or the sum of the horizontal and vertical lengths of the surrounding blocks). Alternatively, the frequency may be a value smaller (or larger) than the product of the horizontal length and the vertical length of the surrounding blocks (or the sum of the horizontal and vertical lengths of the surrounding blocks). For example, the small value (or large value) may be half (or twice) the product of the horizontal length and the vertical length of the surrounding blocks (or the sum of the horizontal and vertical lengths of the surrounding blocks). When the encoding mode of the current block is EIP, Intra TMP, or MIP, the frequency value can be the product of the width and height of the surrounding blocks (or the sum of the width and height of the surrounding blocks). Alternatively, the frequency can be a value smaller (or larger) than the product of the width and height of the surrounding blocks (or the sum of the width and height of the surrounding blocks). For example, a small value (or large value) can be half (or twice) the product of the width and height of the surrounding blocks (or the sum of the width and height of the surrounding blocks). In addition to the intra prediction modes derived from the surrounding blocks, intra prediction modes derived based on history can also be added to the histogram.

[0381] When a video signal processing device constructs a histogram, the histogram frequency value for the intra prediction mode derived from the surrounding blocks may vary depending on the distance between the surrounding blocks and the current block. The closer the distance between the surrounding blocks and the current block, the higher the histogram frequency value. Furthermore, the farther the distance between the surrounding blocks and the current block, the lower the histogram frequency value. For example, a value obtained by subtracting the distance between the surrounding blocks and the current block from the histogram frequency value for the intra prediction mode derived from the surrounding blocks (the product of the horizontal and vertical lengths of the surrounding blocks (or the sum of the horizontal and vertical lengths of the surrounding blocks)) may be used as the frequency value.

[0382] A method based on the occurrence frequency of an intra prediction mode can be used by a video signal processing device to construct a motion candidate list, a block vector candidate list, or a CCP candidate list. During the process of constructing a candidate list (motion candidate list, block vector candidate list, or CCP candidate list), the video signal processing device can construct the candidate list based on the occurrence frequency of surrounding blocks or reorder the candidates in order of occurrence frequency.

[0383] Additionally, a method based on occurrence frequency can be used by a video signal processing device to derive a BV. For example, a video signal processing device can construct a BV histogram based on the occurrence frequency of BVs used in surrounding blocks, and then use the top N most frequently occurring BVs based on the BV histogram as the BV for the current block. This can be called OBVC (Occurrence-based BV coding) mode. Here, N can be an integer greater than or equal to 1, such as 1, 2, or 3.

[0384] Additionally, a method based on occurrence frequency can be used by a video signal processing device to derive LIC information. For example, the video signal processing device can construct an LIC histogram based on the occurrence frequency of LIC information used in surrounding blocks, and then use the most frequently occurring LIC based on the LIC histogram as the LIC information for the current block. The above-mentioned occurrence frequency-based method can be used to determine one of the CCP model, MV, EIP model, and conversion method (MTS, LFNST, NSPT).

[0385] The video signal processing device can determine whether to apply the prediction method using multiple reference pixel lines described based on FIG. 32 based on whether the AIDR mode is applied to the current block. For example, if the AIDR mode is applied to the current block and the AIDR mode is the third range mode, the prediction method using multiple reference pixel lines may not be applied.

[0386] To encode the intra prediction mode of the current block, an MPM list may be used. Candidates in the MPM list may be reordered based on a template cost. When reordering the candidates in the MPM list, the video signal processing device may perform scaling and inverse scaling processes for the intra prediction mode of each candidate in the MPM list depending on whether the AIDR mode of the current block is applied. For example, if the AIDR mode of the current block is the first range mode, the intra prediction mode of each candidate in the MPM list may also be the first range mode. When reordering the MPM list, the video signal processing device may scale the intra prediction mode of each candidate in the MPM list to the third range mode, and then generate a prediction template to calculate a template cost for the prediction template. If the AIDR mode of the current block is the third range mode, the intra prediction mode of each candidate in the MPM list may also be the third range mode. When rearranging the MPM list, the video signal processing device can generate a prediction template by using the intra prediction mode for each candidate in the MPM list as is without a scaling process, and can calculate a template cost for the prediction template.

[0387] The video signal processing device can determine whether the AIDR mode is allowed by using at least one or more of the encoding mode of the current block (DIMD, TIMD, Intra TMP, MIP, EIP, SGPM, TMRL, BDPCM, ISP), reference sample line information, horizontal and vertical size information of the current block, and information on whether the current block is a luminance component block. For example, if the encoding mode of the current block is not DIMD, TIMD, Intra TMP, MIP, EIP, SGPM, TMRL, BDPCM, or ISP mode, the reference sample line is 0, and the current block is a luminance component block, the AIDR mode may be allowed for the current block. If the AIDR mode is allowed, the encoder can generate and signal a bitstream including information related to AIDR, and the decoder can parse the information related to AIDR to determine whether the AIDR mode is applied to the current block.

[0388] Figure 40 illustrates a method of using a candidate list based on a prediction mode according to one embodiment of the present invention.

[0389] Referring to FIG. 40, a video signal processing device can generate a prediction block of a current block using a candidate list based on a prediction mode. The video signal processing device can configure a prediction mode candidate list for the current block. At this time, the video signal processing device can configure the prediction mode candidate list using at least one of neighboring block information, DIMD information of the current block, TIMD information of the current block, MPM list of the current block, and Non-MPM list of the current block. The neighboring block information may be a block adjacent to the current block, a block that is not adjacent to the current block and is separated by a predetermined distance, a neighboring block stored in a separate memory, etc. The neighboring block information may be intra-prediction directional mode information of the neighboring block and a weight value for each mode. The predetermined distance may be determined based on at least one of the upper-left sample position of the current block, the horizontal and vertical sizes of the current block, and the predetermined horizontal and vertical sizes. For example, a video signal processing device can derive a neighboring block based on a position shifted horizontally to the left by the width of the current block based on the upper left sample position of the current block and a position shifted vertically upward based on the upper left sample position of the current block.

[0390] A method for a video signal processing device to construct a candidate list based on a prediction mode is described. The video signal processing device can generate a prediction mode candidate using at least one of neighboring block information, DIMD information of a current block, TIMD information of the current block, an MPM list of the current block, a Non-MPM list of the current block, DIMD information derived from a cumulative histogram of multiple histograms used to derive DIMD from neighboring blocks, and a top N (predetermined number) intra prediction modes with a high occurrence frequency among intra prediction modes used in neighboring blocks. The predetermined number may be an integer of 5. For example, a prediction mode candidate based on occurrence frequency may include five intra prediction modes. In this case, the prediction mode may be composed of two or more intra prediction directional modes, and the intra prediction directional modes of one prediction mode candidate may be different from each other. In addition, the prediction mode may only include intra prediction directional modes that have directionality. In addition, the prediction mode may not include a planar mode, which is a non-directional mode. If the intra prediction directional mode derived from the neighboring block is a planar mode, the intra prediction directional mode derived from the neighboring block may be excluded from the prediction mode. In addition, the prediction mode may not include the DC mode, which is a non-directional mode. If the intra prediction directional mode derived from the neighboring block is a DC mode, the intra prediction directional mode derived from the neighboring block may be excluded from the prediction mode. The prediction mode candidate generated by the above-described method may be added to a prediction mode candidate list. The video signal processing device may determine whether the same prediction mode candidate exists in the prediction mode candidate list, and only if it does not exist, may add the generated prediction mode candidate to the prediction mode candidate list. The maximum size of the prediction mode candidate list may be a pre-specified number, and the pre-specified number may be an integer, such as 25.If the maximum size of the prediction mode candidate list is a pre-specified number, the video signal processing device may not add any more prediction mode candidates.

[0391] A video signal processing device can generate a reordered prediction mode candidate list by reordering a prediction mode candidate list based on template costs. Alternatively, the reordering process may not be performed. Reordering based on template costs may be performed in the above-described manner using the templates of FIG. 30 or FIG. 31. The video signal processing device can configure a reference template using reconstructed neighboring samples adjacent to a current block. The video signal processing device can generate a prediction template for the reference template using prediction mode candidates in the prediction mode candidate list and pre-designated reference sample lines (e.g., Reference lines 1, 2, 3, … of FIG. 8). The pre-designated reference sample line may be Reference line 1 (e.g., a line adjacent to the current block template). The video signal processing device can calculate a cost between the reference template and the prediction template. Here, the cost may be the sum of the absolute values ​​of the differences (SAD) between samples. Additionally, the cost may be the sum of the absolute values ​​of the differences (MR-SAD) between the mean difference sample values ​​obtained by subtracting the mean of each template from each sample value. The intra-prediction directional mode of the prediction mode candidate used to generate the prediction template may be converted to an intra-prediction mode range that is extended beyond the range of the existing intra-prediction mode, and then used to generate the prediction template. For example, if the intra-prediction directional mode of the prediction mode candidate is N, the extended intra-prediction directional mode may be M = (N * 2) - 2. In addition, since the prediction mode candidate is composed of two or more intra-prediction directional modes, the prediction templates generated using each intra-prediction directional mode may be weighted averaged by applying a predetermined weight, and the video signal processing device may generate a weighted average prediction template. At this time, the weight may be a predetermined value, and the weights for each intra-prediction directional mode may be the same or different from each other.The list of prediction mode candidates can be reordered based on the above calculated cost, and can be reordered in ascending order.

[0392] A video signal processing device can reconstruct a prediction mode candidate list using an MPM list and a rearranged prediction mode candidate list. At this time, the video signal processing device can construct a prediction mode candidate list using a portion of the MPM list and a portion of the rearranged prediction mode candidate list. Candidates starting from the first candidate of the MPM list up to a predetermined number can be added to the prediction mode candidate list, and the predetermined number can be 5. Candidates from among the prediction mode candidates of the rearranged prediction mode candidate list can be added to the prediction mode candidate list in a predetermined number in order from lowest cost to highest cost, and the predetermined number can be an integer and can be 16.

[0393] The encoder can determine an optimal prediction mode candidate for the current block from the derived prediction mode candidate list (or the prediction mode candidate list if the prediction mode candidate list reconstruction process is not performed), and then generate and signal a bitstream including index information for the optimal prediction mode candidate. The decoder can parse the index information for the optimal prediction mode candidate from the bitstream and obtain a prediction block for the current block using the determined optimal prediction mode candidate. The video signal processing device can obtain a final prediction block of the current block through a weighted average between a prediction block obtained based on a planar mode (or a DC mode) and a prediction block obtained using the optimal prediction mode candidate.

[0394] Whether a method for generating a prediction block using a candidate list based on a prediction mode is activated may be determined based on at least one of whether the current block is a luminance block or a chrominance block, whether an intra prediction mode is applied to the current block, whether one of DIMD, TIMD, SGPM, ISP, MIP, Intra TMP, and TMRL modes is applied to the current block, reference pixel line information of the current block, horizontal and vertical sizes of the current block, and whether the current block is located at the upper boundary of a CTU including the current block. For example, if one of DIMD, TIMD, SGPM, ISP, MIP, Intra TMP, and TMRL modes is applied to the current block, the method for generating a prediction block using a candidate list based on a prediction mode may not be activated, and the video signal processing device may not signal or parse syntax related to the method for generating a prediction block using a candidate list based on a prediction mode.

[0395] The candidate list based on the derived prediction mode can be used when the coding mode of the current block is one of DIMD, TIMD, SGPM, MIP, Intra TMP, and TMRL.

[0396] When DIMD is applied to the current block, the video signal processing device can add the prediction mode derived using DIMD to a candidate list based on the derived prediction mode. At this time, the candidate list based on the derived prediction mode can be reordered based on the template cost. The encoder can generate and signal a bitstream including index information for the optimal prediction mode within the candidate list based on the derived prediction mode, and the decoder can parse the index information to determine the prediction mode for the current block within the candidate list based on the derived prediction mode.

[0397] When TIMD is applied to the current block, the video signal processing device can add the prediction mode derived using TIMD to a candidate list based on the derived prediction mode, and the candidate list based on the derived prediction mode can be reordered based on the template cost. The encoder can generate and signal a bitstream including index information for an optimal prediction mode in the candidate list based on the derived prediction mode, and the decoder can parse the index information for the optimal prediction mode to determine the optimal prediction mode in the candidate list based on the derived prediction mode.

[0398] When SGPM is applied to the current block, the video signal processing device can use a candidate list based on the derived prediction mode to derive an intra prediction mode of the two divided regions, the encoder can generate and signal a bitstream including index information on an optimal prediction mode for one or more regions among the two divided regions, and the decoder can parse the index information to determine a prediction mode for one or more regions among the two divided regions from the list.

[0399] When DIMD is applied to the current block, the video signal processing device can add block vector candidates derived using intra-TMP to a candidate list based on the derived prediction mode, and the candidate list based on the derived prediction mode can be reordered based on the template cost. The encoder can generate and signal a bitstream including index information for an optimal prediction mode within the candidate list based on the derived prediction mode, and the decoder can parse the index information to determine the optimal prediction mode within the candidate list based on the derived prediction mode.

[0400] When TMRL is applied to the current block, the video signal processing device can construct a unified candidate list using the candidate list based on the induced prediction mode and the reference sample line. The unified list can be reordered based on the template cost. The encoder can generate and signal a bitstream including index information for the optimal prediction mode and reference sample line in the candidate list based on the induced prediction mode, and the decoder can parse the information for the index to determine the optimal prediction mode reference sample line in the candidate list based on the induced prediction mode from the unified list.

[0401] The Non-MPM list can be reordered based on template cost using templates, and the Non-MPM list can be used to signal or parse the optimal intra prediction mode.

[0402] The MPM list (e.g., Primary MPM list, Secondary MPM list) or Non-MPM list for the current block can be organized or reordered according to the occurrence frequency of intra prediction modes used in the neighboring blocks. The intra prediction modes used in the neighboring blocks can be at most J intra prediction modes if the neighboring blocks are DIMD mode, at most K intra prediction modes if the neighboring blocks are TIMD mode, at most L intra prediction modes if the neighboring blocks are SGPM mode, and if the neighboring blocks are intra TMP or IBC mode, they can be intra prediction modes stored in the block to which the intra TMP or IBC mode is applied (or intra prediction modes of the block indicated by the block vector of the block to which the intra TMP or IBC mode is applied). Information about the neighboring blocks can include blocks adjacent to the current block, blocks that are not adjacent to the current block but are separated by a predetermined distance, neighboring blocks stored in a separate memory, etc. In addition, information about the surrounding blocks may include intra prediction directional mode information of the surrounding blocks and weight values ​​for each mode. The video signal processing device may generate a histogram for the intra prediction mode using the intra prediction mode derived from the surrounding blocks. At this time, J, K, and L may be integers, and may be 5, 2, and 2, respectively. The video signal processing device may construct an MPM list for the current block using the histogram for the intra prediction mode. MPMs may be added to the primary MPM list in order of highest occurrence frequency to lowest occurrence frequency, and when the primary MPM list is filled to a predetermined maximum number, they may be added to the secondary MPM list in order of the next occurrence frequency. When the secondary MPM list is filled to a predetermined maximum number, they may be added to the non-MPM list in order of the next occurrence frequency.The histogram for the intra prediction mode is updated for each block and can be used to construct the MPM list or the Non-MPM list for the block.

[0403] In addition, the video signal processing device can first configure the MPM list of the current block, and then reorder the MPM list or the Non-MPM list using the histogram for the intra prediction mode. Specifically, the video signal processing device can sequentially set the reference intra prediction mode in the order of high occurrence frequency to low occurrence frequency in the histogram for the intra prediction mode, and then reorder the MPM list or the Non-MPM list based on the reference intra prediction mode. In addition, the video signal processing device can change the order of the intra prediction mode that is the same as the reference intra prediction mode or has a difference of a predetermined value in the MPM list or the Non-MPM list to a higher order in the MPM list or the Non-MPM list. The predetermined value can be an integer of 3.

[0404] The order in which the video signal processing device adds intra prediction modes to the Non-MPM list may vary based on the horizontal and vertical sizes of the current block. For example, if the horizontal size of the current block is greater than the vertical size, intra prediction modes (e.g., the prediction modes of FIG. 6) may be added in the order from a pre-specified first intra prediction mode to a pre-specified second intra prediction mode. The first intra prediction mode may be the prediction mode of index 66, and the second intra may be the prediction mode of index 2.

[0405] In addition, if the horizontal size of the current block is greater than the vertical size, the intra prediction modes may be added in the order of a pre-specified first intra prediction mode to a pre-specified second intra prediction mode, and then in the order of a pre-specified third intra prediction mode to a pre-specified fourth intra prediction mode. The first intra prediction mode may be a prediction mode of index 34, the second intra prediction mode may be a prediction mode of index 66, the third intra prediction mode may be a prediction mode of index 33, and the fourth intra prediction mode may be a prediction mode of index 2.

[0406] A video signal processing device can use one intra prediction mode as a reference intra prediction mode and additionally derive one or more intra prediction modes through a predefined rule. The video signal processing device can generate one or more prediction blocks using the reference intra prediction mode and the derived intra prediction mode. In addition, when the video signal processing device generates a plurality of prediction blocks, the video signal processing device can generate a final prediction block by weighted averaging (fusion) the plurality of prediction blocks. The predefined rule may configure an additional intra prediction mode by adding or subtracting a predefined value based on the reference intra prediction mode. The predefined value may be an integer greater than or equal to 1. For example, when the reference intra prediction mode is mode 18, the additionally derived intra prediction modes may be mode 17 (18-1) or mode 19 (18+1). The video signal processing device can generate a plurality of prediction blocks using one or more reference pixel lines and then weighted averaging (fusion) the prediction blocks to generate a final prediction block. The method of generating a final prediction block by weighting and averaging multiple prediction blocks can be described as a fusion method based on a reference intra prediction mode.

[0407] A fusion method based on a reference intra prediction mode can be applied in various ways. For example, it can be applied based on a candidate list for the reference intra prediction mode. A video signal processing device can determine whether a fusion method based on a reference intra prediction mode is applied to the current block using one or more of the position of the current block, the horizontal and vertical sizes of the current block, whether the current block is a luminance block or a chrominance block, whether the coding mode of the current block is an intra mode, and whether the coding mode of the current block is one of BDPCM, DIMD, TIMD, Intra TMP, MIP, SGPM, TMRL, EIP, and IBC modes. When a fusion method based on a reference intra prediction mode is applied to the current block, the video signal processing device can configure a candidate list for the reference intra prediction mode. The candidate list can include all intra prediction modes. Alternatively, the candidate list can include only pre-specified intra prediction modes. For example, the candidate list can include only directional intra prediction modes excluding the planar mode and the DC mode. Alternatively, the candidate list may include only directional intra prediction modes excluding the planar mode, the DC mode, the 2nd mode (see FIG. 6), and the 66th mode (see FIG. 6). The video signal processing device may calculate a cost based on the template cost for each intra prediction mode in the candidate list, and then reorder the candidate list based on the cost. The video signal processing device may construct a final candidate list using only a pre-specified number of candidates with low template costs in the candidate list. The pre-specified number may be an integer greater than or equal to 2. The encoder may generate and signal a bitstream including information indicating whether a fusion mode based on a reference intra prediction mode is applied to a current block.In addition, the encoder can generate and signal a bitstream including an index indicating an optimal candidate within a candidate list when a fusion mode based on a reference intra prediction mode is applied. The decoder can parse information indicating whether a fusion mode based on a reference intra prediction mode is applied to a current block, and determine whether a fusion mode based on a reference intra prediction mode is applied to the current block based on the parsing result. In addition, when a fusion mode based on a reference intra prediction mode is applied, the decoder can parse an index indicating an optimal candidate, and determine an optimal candidate within the candidate list based on the parsing result. The video signal processing device can additionally derive one or more intra prediction modes through a predefined rule using the optimal reference intra prediction mode for the current block. The video signal processing device can then generate a plurality of prediction blocks using the optimal reference intra prediction mode and the derived intra prediction modes, and then perform a weighted average (fusion) of the plurality of prediction blocks to generate a final prediction block. Alternatively, the video signal processing device can determine a candidate having a minimum template cost among the candidate list as the optimal reference intra prediction mode. At this time, the encoder may not include an index indicating the optimal candidate in the bitstream, and the decoder may not parse the index indicating the optimal candidate. When the video signal processing device performs weighted averaging (fusion) of multiple prediction blocks, the weights used may be predefined. The weight for the reference intra prediction mode may be greater than the weight for the derived intra prediction mode. For example, the weight for the reference intra prediction mode may be 32, and the weight for the derived intra prediction mode may be 16. Since the range of samples is expanded as the weights are multiplied to each prediction block, the video signal processing device may perform a process of adjusting it back to the range of the original samples.To this end, the video signal processing device can apply (X >> 6) to each sample to fit the range. X may represent the value of the sample. Alternatively, the video signal processing device may perform weighted averaging by applying a predefined weight derivation method when performing weighted averaging (fusion) on a plurality of prediction blocks. The predefined weight derivation method may be a method of determining a weight for each prediction block based on a template cost. In order to set a higher weight as the template cost is smaller, the video signal processing device may calculate an integer weight for each intra prediction mode by multiplying the total weight value by the value obtained by subtracting each template cost from the sum of all template costs for each intra prediction mode (A) and dividing the value by A. The total weight may be 64.

[0408] A video signal processing device may use a predefined probability model when encoding and / or decoding information indicating whether a fusion mode based on a reference intra prediction mode is applied to a current block. The predefined probability models may vary, and the video signal processing device may select one of the various probability models based on information of neighboring blocks adjacent to the current block and encoding information of the current block. For example, if both a left neighboring block and an upper neighboring block adjacent to the current block have a fusion mode based on a reference intra prediction mode applied, the video signal processing device may use a predefined first probability model. If only one of the left neighboring block and the upper neighboring block adjacent to the current block has a fusion mode based on a reference intra prediction mode applied, the video signal processing device may use a predefined second probability model. If neither a left neighboring block nor an upper neighboring block adjacent to the current block has a fusion mode based on a reference intra prediction mode applied, the video signal processing device may use a predefined third probability model. At this time, the first probability model, the second probability model, and the third probability model may be different from each other. At this time, the initial values ​​of the first probability model, the second probability model, and the third probability model may be the same.

[0409] Information indicating whether a fusion mode based on a reference intra prediction mode is applied to a current block can be signaled as a sub-mode of a pre-specified encoding mode. The pre-specified encoding mode can be one of DIMD, TIMD, Intra TMP, MIP, EIP, TMRL, SGPM, and ISP. For example, if the pre-specified encoding mode is a DIMD mode, the video signal processing device can signal or parse information indicating whether a fusion mode based on a reference intra prediction mode is applied to a current block only when the DIMD mode is applied to the current block. If the DIMD mode is applied to the current block and the fusion mode based on the reference intra prediction mode is applied, the video signal processing device can apply the fusion mode based on the reference intra prediction mode to the intra prediction mode derived by the DIMD method to generate a prediction block.

[0410] Alternatively, the video signal processing device may apply the fusion method based on the reference intra prediction mode to all encoding modes in which a prediction block is generated based on the intra prediction mode, without signaling or parsing information indicating whether the fusion method based on the reference intra prediction mode is applied to the current block. For example, the fusion method based on the reference intra prediction mode may be applied to all modes in which a prediction block is generated based on the intra prediction mode among intra prediction modes based on DIMD, TIMD, Intra TMP, MIP, EIP, TMRL, SGPM, ISP, and MPM.

[0411] When a fusion mode based on a reference intra prediction mode is applied to a current block, a video signal processing device can convert the intra prediction mode into a wider range of intra prediction modes to generate a prediction block. The conversion to a wider range can be calculated as "(MODE<<1) - 2", and the planar mode and DC mode may not be converted. The conversion to the original range can be calculated as "(MODE>>1) + 1". Here, MODE can be an index of the intra prediction mode of FIG. 6. The wider range of intra prediction modes can also be used when calculating template costs. The fusion method based on the reference intra prediction mode can be applied to each intra prediction mode when the intra prediction mode is used in DIMD, TIMD, intra TMP, MIP, EIP, TMRL, SGPM, and ISP. The fusion method based on the reference intra prediction mode can be applied only to a pre-specified reference sample line. For example, when a fusion method based on a reference intra prediction mode is applied, the reference sample line may be a reference sample adjacent to the current block. PDPC based on the intra prediction mode may be applied to each block predicted using the fusion method based on the reference intra prediction mode.

[0412] When a fusion method based on a reference intra prediction mode is applied to a current block, a prediction block generated using the reference intra prediction mode and the derived intra prediction mode can be obtained (generated) using one or more reference sample lines. In other words, a video signal processing device can obtain multiple prediction blocks using one or more reference sample lines and a reference intra prediction mode, and the video signal processing device can obtain a fused first prediction block by performing a weighted average on each of the obtained prediction blocks. In addition, the video signal processing device can generate a plurality of prediction blocks using one or more reference sample lines and the derived intra prediction mode A, and then perform a weighted average on each of the prediction blocks to generate a fused second prediction block. Alternatively, the video signal processing device can generate a plurality of prediction blocks using one or more reference sample lines and the derived intra prediction mode B, and then perform a weighted average on each of the prediction blocks to generate a fused third prediction block. Weights used when generating the first, second, and third prediction blocks can be defined in advance. The weight for a block predicted using the main reference sample line may be greater than (or equal to or less than) the weight for a block predicted using the sub-reference sample line. For example, the weight for a block predicted using the main reference sample line may be 3, and the weight for a block predicted using the sub-reference sample line may be 1.

[0413] A prediction block for a current block acquired by a video signal processing device may be a prediction block fused using one or more intra prediction modes within an MPM list. At this time, the video signal processing device may acquire the fused prediction block using one or more intra prediction modes from among a Primary MPM list, a Secondary MPM list, and a Non-MPM list. Alternatively, the video signal processing device may acquire the fused prediction block using only all candidates within the Primary MPM list. Alternatively, the video signal processing device may configure an intra prediction mode combination candidate using one or more intra prediction modes from among the Primary MPM list, the Secondary MPM list, and the Non-MPM list, and then configure a combination candidate list using the combination candidate. The encoder may acquire and signal a bitstream including information indicating whether a fusion-based intra prediction mode is applied and an index indicating an optimal combination candidate within the combination candidate list. The decoder can parse information indicating whether a fusion-based intra prediction mode is applied and an index indicating an optimal combination candidate in a combination candidate list, configure a combination candidate based on the parsing result, and then obtain a prediction block for the current block based on the combination candidate. A weight for an intra prediction mode or an intra prediction mode in a combination candidate can be a preset weight or a weight derived using the weight derivation method described above. The number of intra prediction modes in a combination candidate can be two or more. For example, a combination candidate composed of two intra prediction modes can exist, or a combination candidate composed of 3, 4, 5, … N intra prediction modes can exist.

[0414] The range of intra prediction modes can be directional modes from modes 2 to 66, excluding the planar mode (mode 0 in Fig. 6) and the DC mode (mode 1 in Fig. 6). In the case of modes that derive intra prediction modes and use them for intra prediction, such as TIMD and TMRL, the range of intra prediction modes can be expanded by two times from modes 2 to 130 by narrowing the angle between each mode and used for prediction. The DIMD mode is a method of calculating directional modes from modes 2 to 66 from reconstructed surrounding samples and using them as intra prediction modes. In the DIMD mode, the range of directional modes calculated from reconstructed surrounding samples can be calculated as directional modes from modes 2 to 130 and used as intra prediction modes for the current block. The video signal processing apparatus can apply a method of expanding the range of directional modes in the DIMD mode to all blocks. The method of expanding the range of directional modes in the DIMD mode can be selectively applied to each block. An encoder can include and signal information indicating a directional range in a bitstream when a current block is encoded in DIMD mode. A decoder can parse the information indicating the directional range when encoded in DIMD mode and use it to calculate (derive) a DIMD intra prediction mode for the current block. The information indicating the directional range can indicate whether the range of directional modes from 2 to 66 is used or the range of directional modes from 2 to 130 is used. If the coding mode of the current block is DIMD and the directional mode from 2 to 130 is used, the video signal processing device can reduce the range of directional modes of the current block to the range from 2 to 66 for the next block to be encoded and store the intra prediction mode.Alternatively, the video signal processing device may store the intra prediction mode according to the range of the extended intra prediction mode if the encoding mode of the current block is DIMD and the directional mode is in the range of 2 to 130. For calculating the range of the extended intra prediction mode in the DIMD mode, the range of the predefined table may be extended from 17 to 33. The values ​​in the predefined table are {0, 2048, 4096, 6144, 8192, 12288, 16384, 20480, 24576, 28672, 32768, 36864, 40960, 47104, 53248, 59392, 65536} to {0, 1024, 2048, 3072, 4096, 5120, 6144, 7168, 8192, 10240, 12288, 14336, 16384, 18432, 20480, 22528, 24576, 26624, 28672, {30720, 32768, 34816, 36864, 38912, 40960, 44032, 47104, 50176, 53248, 56320, 59392, 62464, 65536} can be used to calculate the directional mode.

[0415] The video signal processing device can set the range of the intra prediction mode differently based on at least one of the size of the current block, the difference (or ratio) between the horizontal size and the vertical size of the current block, and the intra prediction directional mode. For example, if the horizontal size of the current block is greater than the vertical size and the intra prediction directional mode is greater than mode 34 (see FIG. 6), the video signal processing device can scale the range of the current intra prediction mode so that the range of the intra prediction mode of the current block is used in the range from 2 to 130. At this time, the scaling can be performed using "(MODE<<1) - 2". Alternatively, the video signal processing device can set the type of the angle of the intra prediction mode and the number of intra prediction modes differently based on at least one of the size of the current block, the difference (or ratio) between the horizontal size and the vertical size of the current block. For example, if the horizontal size of the current block is greater than the vertical size, the video signal processing device can reduce the number of modes from 32 to 16 by deleting odd modes and leaving only even modes among modes 2 to 34 of FIG. 6. In addition, if the horizontal size of the current block is greater than the vertical size, the video signal processing device can expand the number of modes from 31 to 47 by adding 16 new angle modes between each of modes 34 to 66 of FIG. 6. At this time, the new angle modes can be pre-specified angle modes, and can be modes such as angles between modes 34 and 35, or angles between modes 35 and 36.

[0416] When the TIMD (or TMRL) mode is applied to the current block, the range of the extended intra prediction mode can be used. Therefore, the video signal processing device can scale the range of the intra prediction mode to the range of the extended intra prediction mode if the intra prediction mode derived from the surrounding block is a mode that does not use the range of the extended intra prediction mode or is a mode stored in the range of the basic intra prediction mode (range from 2 to 66). At this time, the scaling method can be calculated as "(MODE<<1) - 2", and the planar mode and DC mode may not be converted. The inverse scaling to the original range can be calculated as "(MODE>>1) + 1", and the planar mode and DC mode may not be converted. The mode that does not use the range of the extended intra prediction mode may be an encoding mode other than DIMD, TIMD, and TMRL, and may be a mode based on SGPM, ISP, MIP, Intra TMP, IBC, EIP, or MPM. Therefore, the encoding mode of the surrounding blocks can be scaled to the range of extended intra prediction modes in the case of SGPM, ISP, MIP, Intra TMP, IBC, EIP, and MPM-based modes, and can be used to derive the TIMD intra prediction mode of the current block.

[0417] When the SGPM (or GPM or MPM-based mode) mode is applied to the current block, the range of the basic intra prediction mode (the range from 2 to 66) can be used. Accordingly, the video signal processing device can perform inverse scaling of the range of the extended intra prediction mode to the range of the basic intra prediction mode if the intra prediction mode derived from the surrounding block is a mode that uses the range of the extended intra prediction mode. When the encoding mode of the surrounding block is a DIMD, TIMD, or TMRL-based mode, the video signal processing device can inversely scale the range of the intra prediction mode to the range of the basic intra prediction mode, which can be used to derive the intra prediction mode of the current block.

[0418] When the Derived Mode (DM) mode is applied to the current chrominance block, the range of the basic intra prediction mode (the range from 2 to 66) can be used. Accordingly, the video signal processing device can, if the intra prediction mode derived from the luminance block corresponding to the current chrominance block is a mode that uses the range of the extended intra prediction mode, reverse-scale the range of the extended intra prediction mode to the range of the basic intra prediction mode. The video signal processing device can, if the encoding mode of the luminance block corresponding to the current chrominance block is a DIMD, TIMD, or TMRL-based mode, reverse-scale the range of the extended intra prediction mode to the range of the basic intra prediction mode to derive the intra prediction mode of the current chrominance block.

[0419] When the DIMD mode is applied to a current block, the video signal processing device may use a range of an extended intra prediction mode for the current block. Accordingly, when the DIMD mode is applied to the current block, the video signal processing device may descale the range of the extended intra prediction mode to a range of a basic intra prediction mode when deriving a transform kernel. The video signal processing device may derive one of a set of transform kernels, such as MTS, LFNST, and NSPT, a transform kernel, and a transform type based on the intra prediction mode descaled to the range of the basic intra prediction mode.

[0420] A video signal processing device can determine a final intra prediction mode by compensating an intra prediction mode for a current block. At this time, the video signal processing device can perform compensation for the intra prediction mode by deriving an additional intra prediction mode using the intra prediction mode of the current block and then determining an intra prediction mode with a lower cost based on a template cost. The additional intra prediction mode can be obtained by adding or subtracting a predetermined compensation value from the intra prediction mode, and the predetermined compensation value can be an integer greater than or equal to 1. For example, if the intra prediction mode for the current block is 24, the video signal processing device can derive additional intra prediction modes 23 (24-1) and 25 (24+1). In addition, the video signal processing device can calculate a template cost using intra prediction modes 23, 24, and 25, and apply the intra prediction mode with the smallest template cost as the intra prediction mode of the current block. At this time, if the range of the intra prediction mode of the current block is from 2 to 66, the video signal processing device can generate intra prediction modes from 2 to 130 and set them as additional intra prediction modes. In addition, the video signal processing device can set the additional intra prediction mode as the final intra prediction mode. The video signal processing device can obtain the prediction block of the current block based on the final intra prediction mode. The method for correcting the intra prediction mode for the current block can be applied to modes based on DIMD, TIMD, TMRL, SGPM, ISP, MIP, Intra TMP, IBC, EIP, and MPM.

[0421] The range of the intra prediction mode can be selectively applied to each block, and the method of selectively applying the intra prediction mode to each block can be called Adaptive Intra Directional mode Resolution (AIDR). At this time, the range of the intra prediction mode can be configured in various ways. For example, the first range mode can be the range from 2 to 34, the second range mode can be the range from 2 to 66, and the third range mode can be the range from 2 to 130. If the coding mode of the current block is not a mode that implicitly induces an intra prediction mode (e.g., DIMD, TIMD, TMRL, SGPM, ISP, MIP, Intra TMP, IBC, EIP) (i.e., an coding mode based on MPM that explicitly signals an intra prediction mode), the encoder can signal information indicating the range mode by including it in the bitstream. The information indicating the range mode can be signaled before the Primary MPM index is signaled. Alternatively, the information indicating the range mode may be signaled before the Secondary MPM index is signaled. Alternatively, the information indicating the range mode may be signaled before the Non-MPM index is signaled. Alternatively, the information indicating the range mode may be signaled after any one of the Primary MPM index, the Secondary MPM index, and the Non-MPM index is signaled. The decoder can parse the information indicating the range mode to set the range of the intra prediction mode for the current block.

[0422] If the encoding mode of the current block is a mode that implicitly induces an intra prediction mode (e.g., DIMD, TIMD, TMRL, SGPM, ISP, MIP, Intra TMP, IBC, EIP), the range of the extended intra prediction mode, which is the third range mode, may be implicitly used. If the encoding mode of the current block is a mode that implicitly induces an intra prediction mode, the intra prediction mode of a neighboring block may be used. If the intra prediction mode derived from the neighboring block is a mode that does not use the range of the extended intra prediction mode, a mode stored in the range of the basic intra prediction mode (ranges from 2 to 66), the first range mode, or the second range mode, the video signal processing device may scale the range of the intra prediction mode to the range of the extended intra prediction mode.

[0423] When the SGPM (or GPM or MPM-based mode) mode is applied to the current block, the range of the basic intra prediction mode (the range from 2 to 66) can be used. Therefore, when the intra prediction mode derived from the surrounding block is the first range mode or the third range mode, the video signal processing device can scale or inversely scale the derived intra prediction mode to the range of the basic intra prediction mode. That is, when a mode that is encoded in the range of the basic intra prediction mode is applied to the current block, and the encoding mode of the surrounding block is the first range mode, the video signal processing device can scale the encoding mode of the surrounding block to the range of the basic intra prediction mode, and the scaled intra prediction mode range can be used to derive the intra prediction mode of the current block. Alternatively, if a mode that is encoded in the range of the basic intra prediction mode is applied to the current block, and the encoding mode of the surrounding block is a third range mode, the video signal processing device can descale the encoding mode of the surrounding block to the range of the basic intra prediction mode, and the descaled intra prediction mode range can be used to derive the intra prediction mode of the current block.

[0424] When the Derived Mode (DM) mode is applied to the current chrominance block, the range of the basic intra prediction mode (the range from 2 to 66) can be used. Accordingly, when the intra prediction mode derived from the luminance block corresponding to the current chrominance block is the first range mode or the third range mode, the video signal processing device can scale or inversely scale the derived intra prediction mode to the range of the basic intra prediction mode. When the coding mode of the luminance block corresponding to the current chrominance block is the first range mode, the first range mode can be scaled to the range of the basic intra prediction mode, and the scaled prediction mode range can be used to derive the intra prediction mode of the current chrominance block. Alternatively, when the coding mode of the luminance block corresponding to the current chrominance block is the third range mode, the third range mode can be inversely scaled to the range of the basic intra prediction mode, and the inversely scaled prediction mode range can be used to derive the intra prediction mode of the current chrominance block.

[0425] When deriving a transformation kernel for a current block to which a first range mode is applied, the video signal processing device may scale the intra prediction mode to the range of the basic intra prediction mode and then derive the transformation kernel based on the scaled intra prediction mode. When deriving a transformation kernel for a current block to which a third range mode is applied, the video signal processing device may descale the intra prediction mode to the range of the basic intra prediction mode and then derive the transformation kernel based on the descaled intra prediction mode. The scaling and descale processes for the intra prediction mode may be applied to MTS, LFNST, NSPT, etc.

[0426] The extended intra prediction mode range can be applied to all blocks. In addition, the video signal processing device can use one or more of the information about the encoding mode of the current block or the range mode of the current block to perform scaling and descaling on the intra prediction mode derived from the surrounding blocks, thereby deriving the intra prediction mode for the current block.

[0427] If the encoding mode of the current block is a mode that explicitly sets the intra prediction mode (e.g., MPM-based intra prediction mode), the method of configuring the Primary MPM, Secondary MPM, and Non-MPM lists may vary depending on the information about the range mode. Alternatively, the number of Primary MPM, Secondary MPM, and Non-MPM lists may vary depending on the information about the range mode. In the case of the first range mode (range from 2 to 34), the number of Primary MPM lists may be A excluding the planar mode, the Secondary MPM is not used, and the number of Non-MPM lists may be B. In this case, A and B may be integers greater than or equal to 1, and A may be 5 and B may be 29. Alternatively, A may be 3 and B may be 31. For the second range mode (range from 2 to 66), the number of Primary MPM lists can be C excluding the planar mode, the number of Secondary MPM lists can be D, and the number of Non-MPM lists can be E. In this case, C can be 5, D can be 16, and E can be 45. For the third range mode (range from 2 to 130), the number of Primary MPM lists can be F excluding the planar mode, the number of Secondary MPM lists can be G, and the number of Non-MPM lists can be H. In this case, F can be 8, G can be 32, and H can be 91. Alternatively, for the third range mode, the number of Primary MPM lists can be F excluding the planar mode, the number of Secondary MPM lists can be G, the number of Third MPM lists can be H, and the number of Non-MPM lists can be I. At this time, F can be 8, G can be 27, H can be 32, and I can be 64.Alternatively, in the case of the third range mode, the method for constructing the Primary MPM, Secondary MPM, and Non-MPM lists may be the same as the method for constructing the second range mode. The encoder may generate and signal a bitstream including information indicating an additional correction value, and the decoder may parse the information indicating the additional correction value to determine the correction value. Therefore, in the case of the third range mode, the video signal processing device may determine the final intra prediction mode by adding an additional correction value to the intra prediction mode obtained from the MPM list. The range of the additional correction value may be one of -1, 0, and +1. Alternatively, in the case of the third range mode, the method for constructing the Primary MPM, Secondary MPM, and Non-MPM lists may be the same as the method for constructing the second range mode. The video signal processing device may derive an additional intra prediction mode to become the mode of the third range using the determined intra prediction mode, and then determine the final intra prediction mode of the third range based on the template cost. At this time, the intra prediction mode with the lowest template cost may become the final intra prediction mode. Even at this time, additional correction values ​​may be applied. Furthermore, the video signal processing device may generate a prediction block using multiple intra prediction modes configured based on the additional correction values, and then perform a weighted average on the multiple prediction blocks to generate a final prediction block.

[0428]

[0429] Figure 41 shows reference sample filtering according to an embodiment of the present invention.

[0430] A video signal processing device uses the restored surrounding samples in intra prediction. At this time, the video signal processing device can perform filtering on the reference samples. Through this, the video signal processing device can prevent the subjective image quality reduction of the prediction block caused by noise in the surrounding samples. Fig. 41 (a) shows reference sample filtering using a {1, 2, 1} filter. When filtering is performed using surrounding samples A and B for the reference sample position (X), the video signal processing device can calculate the filtered sample (X') as '(A + 2 * X + B) / 4'. The video signal processing device can perform filtering on all reference samples adjacent to the block to be currently predicted. Fig. 41 (b) shows reference sample filtering for the side array. To facilitate memory access to the reference samples, the video signal processing device can copy the reference samples of the filtered side array to the main array and use them for intra prediction. At this time, the video signal processing device can determine the main array and the side array according to the intra prediction directional mode of the current block. When the intra prediction directionality mode is 34 or more, the video signal processing device may determine the upper sample as the main array and the left sample as the side array. When the intra prediction directionality mode is less than 34, the video signal processing device may determine the upper sample as the side array and the left sample as the main array. Depending on the size of the current block and the intra prediction directionality mode, the video signal processing device may copy some of the reference samples located in the side array to the main array without modification. Alternatively, depending on the size of the current block and the intra prediction directionality mode, the video signal processing device may filter the reference samples located in the side array using adjacent reference samples, and then copy the filtered reference samples of the side array to the main array.At this time, the filtering may be any one of a cubic filter, a Gaussian filter, a low-pass filter, a high-pass filter, and a smoothing filter. The smoothing filter may be a 1:2:1 filter. Specifically, the video signal processing device may derive a reference sample position in units of decimal samples by using at least one of the horizontal and vertical sizes of the current block, the intra prediction directional mode, the position of the reference sample line, and the position of the main array to which the reference sample derived from the side array is to be copied. The video signal processing device may perform filtering by using a reference sample adjacent to the reference sample position in units of decimal samples. In addition, the video signal processing device may copy the filtered reference sample to the main array to use it for intra prediction. The unit of decimal samples may be one of 1 / 2 pixel, 1 / 4 pixel, 1 / 16 pixel, and 1 / 32 pixel. The video signal processing device may determine whether to perform filtering on a reference sample by using at least one of the horizontal and vertical sizes of the current block, the intra prediction directional mode, the position of the reference sample line, and the position of the main array to which the reference sample derived from the side array is to be copied. Specifically, if the reference sample position in the derived decimal sample unit is an integer sample position, the video signal processing device may not perform filtering. In this case, the video signal processing device may copy the integer-unit reference sample to the main array and use it for intra prediction.

[0431]

[0432] Figure 42 shows PDPC filtering according to an embodiment of the present invention.

[0433] Discontinuous sample value changes may occur at the boundary between an intra-prediction block and an adjacent block. The discontinuity may occur at the top and left borders of the prediction block. After the prediction block is generated, the video signal processing device may perform filtering within the prediction block. Through this, the video signal processing device may remove the discontinuity between the intra-prediction block and the adjacent block. Specifically, the video signal processing device may perform filtering that weights the predicted value in the intra mode and the values ​​of one or more pre-reconstructed surrounding samples. At this time, the surrounding samples may be selected from samples at a predetermined position or based on the direction of the intra mode. This is referred to as position-dependent intra-prediction combination (PDPC). In PDPC, the video signal processing device may determine a filtering method according to the intra-prediction directional mode. Figure 42 (a) shows a PDPC filtering method when the intra-prediction directional mode is horizontal. In (a) of Fig. 42, when the encoder and decoder perform filtering on the (1, 0) sample within the prediction block, the video signal processing device can perform filtering on the (1, 0) sample by weighting and averaging at least one of the reconstructed surrounding samples R(-1, 0), R(-1, -1), and R(1, -1) with the (1, 0) sample. At this time, the video signal processing device can apply the difference between the intra reference samples to the (1, 0) sample within the prediction block. (b) of Fig. 42 shows a PDPC filtering method for a case where the intra prediction directional mode is not a planar mode, a DC mode, a horizontal mode, and a vertical mode. In (b) of Fig. 42, the encoder and decoder can perform filtering on the (1, 1) sample by weighting and averaging at least one of the reconstructed surrounding samples R(3, -1), and R(-1, 3) with the (1, 1) sample of the prediction block.Specifically, if the (1, 1) sample within the prediction block is predicted from the upper peripheral sample R(3, -1), the video signal processing device can apply compensation for the (1, 1) sample using the left peripheral sample R(-1, 3) indicated by the intra prediction directional mode. Here, R(x, y) may be a reconstructed sample adjacent to the current block. The video signal processing device can perform prediction for a sample within the prediction block from a first peripheral sample corresponding to the sample within the prediction block using the intra prediction mode, at a first reference sample. The video signal processing device can perform PDPC using a second reference sample, at a sample within the prediction block and / or a second peripheral sample corresponding to the first reference sample.

[0434] As in the embodiments described above, when PDPC is applied, filtering is performed on samples within a prediction block based on differences between neighboring samples. Depending on the intra prediction mode, a left neighboring sample corresponding to an upper neighboring sample may not be derived. Depending on the intra prediction mode, an upper neighboring sample corresponding to a left neighboring sample may not be derived. The video signal processing device may not apply PDPC in intra prediction where an intra prediction mode that cannot derive a difference in change between neighboring samples from the upper neighboring sample and the left neighboring sample is applied. The video signal processing device may determine whether to apply PDPC to the current block based on at least one of the horizontal or vertical size of the current block and the intra prediction mode. At this time, the upper neighboring sample and the left neighboring sample may include an upper-left neighboring sample. At least one of the first reference sample and the second reference sample may be a sample generated through filtering between neighboring samples. At this time, the filtering between neighboring samples may be one of a cubic filter, a Gaussian filter, a low-pass filter, a high-pass filter, and a smoothing filter. The smoothing filter can be a 1:2:1 filter.

[0435] Figure 43 shows a gradient PDPC according to an embodiment of the present invention.

[0436] If the difference between surrounding samples cannot be derived according to the intra prediction mode, the video signal processing device may apply the difference between the samples corresponding to the intra prediction mode among the surrounding samples to the sample within the prediction block, without using the surrounding samples corresponding to the samples within the prediction block. In the embodiment of FIG. 43, the video signal processing device does not perform the PDPC process using the left reference sample corresponding to the sample (X) within the prediction block. The encoder and decoder may perform PDPC on the sample (X) within the prediction block using the difference between the r(-1+d, -1) sample and the r(-1, y) sample corresponding to the current intra prediction mode among the reference samples. This is referred to as gradient PDPC. y may be the vertical coordinate of the sample within the prediction block, and d may be the decimal point position of the sample indicated by the intra prediction mode. In addition, y may be changed to a horizontal coordinate according to the intra prediction mode. At this time, the surrounding samples may be the r(-1, -1+d) sample and the r(x, -1) sample. A video signal processing device may select a first reference sample and a second reference sample that do not correspond to a sample within a prediction block using an intra prediction mode, and perform PDPC applying the difference between the first reference sample and the second reference sample to the sample within the prediction block. At least one of the first reference sample and the second reference sample may be derived from a sample within the prediction block using one of vertical and / or horizontal coordinates. In addition, the video signal processing device may determine that at least one of the first reference sample and the second reference sample may be an upper-left sample. In this case, at least one of the first reference sample and the second reference sample may be a sample generated through filtering between neighboring samples. In addition, the filtering between neighboring samples may be one of a cubic filter, a Gaussian filter, a low-pass filter, a high-pass filter, and a smoothing filter. The smoothing filter may be a 1:2:1 filter.

[0437] CIIP PDPC refers to a mode that applies PDPC to intra prediction blocks in CIIP mode.

[0438] Screen content images generated by computer graphics or captured from screen images tend to have sharp edge characteristics, unlike images captured by a camera. In particular, edge components may appear in letters or shapes. When reference sample filtering or PDPC is applied to edge component images, the edge components may change smoothly, and the characteristics of the original image may disappear. A video signal processing device may not apply smoothing filtering applied in intra prediction to screen content images. In this case, smoothing filtering may include at least one of reference sample filtering, PDPC, and intra fusion. In this case, PDPC may include at least one of gradient PDPC and CIIP PDPC. In addition, a plurality of flags indicating whether to disable each of gradient PDPC and CIIP PDPC may be included in the bitstream. In the embodiments described below, the flag indicating whether to disable PDPC may be a plurality of flags indicating whether to disable each of gradient PDPC and CIIP PDPC.

[0439] The encoder can include smoothing filter information in the bitstream, indicating whether smoothing filtering is disabled. The decoder can obtain the smoothing filter information from the bitstream and determine whether the smoothing filter is disabled. At this time, the smoothing filter information can be applied at a pre-specified level. At this time, the pre-specified level can be any one of the video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), picture, slice tile, and block.

[0440] In this embodiment, if reference sample filtering is not performed, the video signal processing device may not perform the filtering applied to the reference sample of the side array described through (b) of FIG. 41. At this time, the video signal processing device may use the reference sample of the side array to be used in the main array as a reference sample in integer sample units. Due to the angle of the intra prediction directional mode, it may be difficult for the encoder and decoder to obtain a reference sample that matches the integer sample unit. In order to derive the reference sample in integer sample units, the video signal processing device may derive the reference sample position in decimal sample units by using at least one of the horizontal and vertical sizes of the current block, the intra prediction directional mode, the position of the reference sample line, and the position in the main array to which the reference sample derived from the side array is to be copied. At this time, the video signal processing device may copy the reference sample of the side array in integer sample units that is closest to the derived decimal sample unit reference sample position to the main array and use it for intra prediction. The reference sample position in decimal sample units can be one of 1 / 2 pixel, 1 / 4 pixel, 1 / 16 pixel, and 1 / 32 pixel.

[0441] The smoothing filter information described above can be set as one or more flags, and one or more flags can be set for each smoothing filter.

[0442]

[0443] FIG. 44 is a diagram illustrating a process of performing OBMC according to one embodiment of the present specification.

[0444] Referring to FIG. 44, the decoder can obtain motion information (e.g., motion vector) of the current block (e.g., current coding unit) and motion information of surrounding blocks of the current block (S810). The decoder can determine whether OBMC is applied to the sub-blocks of the current block (S820). At this time, the decoder can determine whether OMBC is applied to each sub-block. If an affine mode (e.g., merge-based affine mode, AMVP-based affine mode), sbTMVP (subblock-based temporal motion vector predictors) mode, or MP-DMVR (multi-pass decoder-side motion vector refinement) mode is applied to the current block, the decoder can determine that OBMC is applied to the sub-blocks of the current block. This is because motion information between sub-blocks may be different when the affine mode, sbTMVP mode, or MP-DMVR mode is applied to the current block. The decoder can perform CU-based OBMC (S830). CU-based OBMC can be performed independently for each sub-block including the upper boundary and the left boundary of the current block. Specifically, the decoder can divide the current block into sub-blocks. At this time, the size of each sub-block can be any size and have a positive value. For example, the arbitrary size can be 4. The decoder can perform OBMC on a sub-block basis based on whether OBMC is applied to the sub-block of the current block (S840). That is, if the decoder determines that OBMC is applied to the sub-block, it can perform OBMC on a sub-block basis. OBMC in step S840 can be performed on sub-blocks of the current block unit that do not include the upper / left boundary. That is, OBMC can be performed on the remaining sub-blocks except for the sub-blocks on which OBMC is performed in step S830.The decoder can perform OBMC on the sub-blocks of the current block and obtain a prediction block of the current block (S850).

[0445] Hereinafter, step S830 will be described in more detail. The decoder can determine whether the motion information of the first sub-bloc...

Claims

1. In a video signal decoding device, Includes a processor, The above processor Generate a mixed vector based on the block vector for the current block and the motion information for the current block, Generate a prediction block for the current block using the above mixed vector. Video signal decoding device.

2. In paragraph 1, The above mixed vector includes a block vector indicated by block vector information for the current block and a motion vector indicated by motion information for the current block, The above processor After obtaining a first prediction block based on the block vector and obtaining a second prediction block based on the motion vector, a prediction block for the current block is generated by weighting the first prediction block and the second prediction block. Video signal decoding device.

3. In paragraph 1, The above processor A mixed vector list is formed using block vector information derived from the first surrounding block of the current block and motion information derived from the second surrounding block of the current block, Obtaining a block vector for the current block and motion information for the current block from the above mixed vector list. Video signal decoding device.

4. In paragraph 2, The above processor Deriving a block vector from a reference block of a reference picture indicated by motion information of a surrounding block of the current block Video signal decoding device.

5. In paragraph 2, The above processor Decide whether to use the information of the surrounding blocks of the current block as a candidate of the mixed vector candidate list according to the encoding mode of the surrounding blocks of the current block. Video signal decoding device.

6. In the video signal decoding device, Includes a processor, The above processor Generating a first prediction block using the first motion information of the current block, generating a second prediction block using the surrounding blocks of the current block and the intra prediction mode derived from the surrounding blocks, and generating a final prediction block by weighting the first prediction block and the second prediction block. Video signal decoding device.

7. In paragraph 6, The above processor Generating the second prediction block using the DIMD (decoder side intra mode derivation) mode derived using the restored samples of the surrounding blocks. Video signal decoding device.

8. In paragraph 6, The above processor When generating the second prediction block, intra fusion is not used. Video signal decoding device.

9. In paragraph 6, The above processor When generating the second prediction block, PDPC (position dependent intra prediction combination) is not used. Video signal processing device.

10. In paragraph 6, The above processor Perform filtering on the final prediction block above. Video signal processing device.

11. In a video signal encoding device, Includes a processor, The above processor Generate a mixed vector based on the block vector for the current block and the motion information for the current block, Generate a prediction block for the current block using the above mixed vector. Video signal encoding device.

12. In paragraph 11, The above mixed vector includes a block vector indicated by block vector information for the current block and a motion vector indicated by motion information for the current block, The above processor After obtaining a first prediction block based on the block vector and obtaining a second prediction block based on the motion vector, a prediction block for the current block is generated by weighting the first prediction block and the second prediction block. Video signal encoding device.

13. In paragraph 11, The above processor A mixed vector list is formed using block vector information derived from the first surrounding block of the current block and motion information derived from the second surrounding block of the current block, Obtaining a block vector for the current block and motion information for the current block from the above mixed vector list. Video signal encoding device.

14. In paragraph 12, The above processor Deriving a block vector from a reference block of a reference picture indicated by motion information of a surrounding block of the current block Video signal encoding device.

15. In paragraph 12, The above processor Decide whether to use the information of the surrounding blocks of the current block as a candidate of the mixed vector candidate list according to the encoding mode of the surrounding blocks of the current block. Video signal encoding device.

16. In a video signal encoding device, Includes a processor, The above processor Generating a first prediction block using the first motion information of the current block, generating a second prediction block using the surrounding blocks of the current block and the intra prediction mode derived from the surrounding blocks, and generating a final prediction block by weighting the first prediction block and the second prediction block. Video signal encoding device.

17. In paragraph 16, The above processor Generating the second prediction block using the DIMD (decoder side intra mode derivation) mode derived using the restored samples of the surrounding blocks. Video signal encoding device.

18. In paragraph 16, The above processor When generating the second prediction block, intra fusion is not used. Video signal encoding device.

19. In paragraph 16, The above processor When generating the second prediction block, PDPC (position dependent intra prediction combination) is not used. Video signal encoding device.

20. In paragraph 16, The above processor Perform filtering on the final prediction block above. Video signal encoding device.

21. In the operating method of the video signal decoding device, A step of generating a mixed vector based on a block vector for the current block and motion information for the current block; and A step of generating a prediction block for the current block using the above mixture vector is included. How it works.

22. How the video signal decoding device operates A step of generating a first prediction block using first motion information of the current block; A step of generating a second prediction block using a neighboring block of the current block and an intra prediction mode derived from the neighboring block; and a step of generating a final prediction block by weighting the first prediction block and the second prediction block. How it works.

23. In a bit stream contained in a storage medium and containing a video signal, How to generate a bit stream A step of generating a mixed vector based on a block vector for the current block and motion information for the current block; and A step of generating a prediction block for the current block using the above mixture vector. Bit stream.

24. In a bit stream contained in a storage medium and containing a video signal, How to generate a bit stream A step of generating a first prediction block using first motion information of the current block; A step of generating a second prediction block using a neighboring block of the current block and an intra prediction mode derived from the neighboring block; and a step of generating a final prediction block by weighting the first prediction block and the second prediction block. Bit stream.

Citation Information

Patent Citations

  • Methods and systems for intra block copy coding with block vector derivation

    KR1020170023086A

  • Nonaqueous electrolyte secondary battery

    KR1020250010553A

  • Handcart with variable loading and transport structure

    KR102897197B1

  • CIIP-based prediction method and device

    WO2023055172A1

  • KR20190107581A