Video signal processing method and apparatus

By parsing MMVD information from a bitstream to generate a modified motion vector and applying it to video blocks, the method addresses inefficiencies in existing video signal processing, enhancing coding efficiency and performance.

JP2025111809AActive Publication Date: 2025-07-30WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025078081
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-09
Filing Date
2025-05-08
Publication Date
2025-07-30
Estimated Expiration
2040-07-24

AI Technical Summary

Technical Problem

Existing video signal processing methods lack efficiency in encoding and decoding processes, particularly in handling motion vector differences, leading to suboptimal coding performance.

Method used

The method involves parsing MMVD activation information, merge information, distance, and direction from a bitstream to generate a modified motion vector, and applying it to enhance coding efficiency by generating a merge candidate list and restoring blocks based on this vector, within specified POC differences and reference picture types.

Benefits of technology

This approach improves coding efficiency by optimizing motion vector handling, resulting in enhanced video signal processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111809000001_ABST
    Figure 2025111809000001_ABST
Patent Text Reader

Abstract

To provide a method for increasing the coding efficiency of a video signal.SOLUTION: A video signal processing method includes steps of: parsing, from a bit stream, upper level merge with motion vector difference (MMVD) activity information (sps_mmvd_enabled_flag) indicating whether an upper level MMVD including a current block is used; parsing, from the bit stream, MMVD merge information indicating whether to use the MMVD in the current block when the upper level MMVD activity information indicates the activity of the MMVD; parsing information related to the distance of the MMVD and information related to the direction of the MMVD when the MMVD merge information indicates using the MMVD in the current block; and acquiring information related to the MMVD (mMvdLX) on the basis of the pieces of information. -2^17≤mMvdLX≤2^17-1.SELECTED DRAWING: Figure 38
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method and apparatus for processing video signals, and more particularly, to a method and apparatus for encoding or decoding video signals.

Background Art

[0002] Compression encoding refers to a series of signal processing techniques for transmitting digitized information via a communication line or storing it in a form suitable for a storage medium. The targets of compression encoding include audio, video, characters, etc. In particular, the technique of performing compression encoding on video is called video compression. Compression encoding of video signals is performed by removing redundant information in consideration of spatial correlation, temporal correlation, probabilistic correlation, etc. However, with the recent development of various media and data transmission media, more efficient video signal processing methods and apparatuses are desired.

Summary of the Invention

Problems to be Solved by the Invention

[0003] An object of the present disclosure is to improve the coding efficiency of video signals.

Means for Solving the Problems

[0004] A method for decoding a video signal according to an embodiment of the present disclosure includes parsing, from a bitstream, upper-level MMVD activation information (sps_mmvd_enabled_flag) indicating whether upper-level MMVD (Merge with MVD) including a current block is used; parsing, from the bitstream, MMVD merge information (mmvd_merge_flag) indicating whether to use MMVD for the current block when the upper-level MMVD activation information indicates activation of MMVD; parsing information related to the MMVD distance (mmvd_distance_idx) and information related to the MMVD direction (mmvd_direction_idx) when the MMVD merge information indicates using MMVD for the current block; and obtaining information about MMVD (mMvdLX) based on the information related to the MMVD distance and the information related to the MMVD direction, wherein the information about MMVD is greater than or equal to -2^17 and less than or equal to 2^17-1.

[0005] In a method for decoding a video signal according to an embodiment of the present disclosure, the upper level is any one of a Coding Tree Unit, a slice, a tile, a tile group, a picture, or a sequence unit.

[0006] A method for decoding a video signal according to an embodiment of the present disclosure further includes generating a merge candidate list for a current block; selecting a motion vector from the merge candidate list based on a merge index parsed from a bitstream; adding information about MMVD to the motion vector to obtain a modified motion vector when the MMVD merge information indicates using MMVD for the current block; and restoring the current block based on the modified motion vector, wherein the modified motion vector is greater than or equal to -2^17 and less than or equal to 2^17-1.

[0007] In a method for decoding a video signal according to an embodiment of the present disclosure, the step of obtaining information (mMvdLX) related to MMVD includes obtaining an MMVD offset based on information related to the distance of MMVD and information related to the direction of MMVD, obtaining a difference in POC (Picture Order Count) between a current picture including a current block and a first reference picture based on the first reference list as a first POC difference when the first reference list and the second reference list are used, and obtaining a difference in POC (Picture Order Count) between the current picture and a second reference picture based on the second reference list as a second POC difference, and obtaining information related to a first MMVD related to the first reference list and information related to a second MMVD related to the second reference list based on at least one of the MMVD offset, the first POC difference, and the second POC difference, wherein the information related to MVD includes the information related to the first MMVD and the information related to the second MMVD.

[0008] A method for decoding a video signal according to an embodiment of the present disclosure includes obtaining the MMVD offset as information related to the first MMVD and obtaining the MMVD offset as information related to the second MMVD when the first POC difference and the second POC difference are the same.

[0009] A method for decoding a video signal according to an embodiment of the present disclosure includes obtaining the MMVD offset as information related to the first MMVD when the absolute value of the first POC difference is greater than or equal to the absolute value of the second POC difference, obtaining information related to the second MMVD by scaling the information related to the first MMVD when the first reference picture is not a long-term reference picture and the second reference picture is not a long-term reference picture, and obtaining information related to the second MMVD without scaling the absolute value of the information related to the first MMVD when the first reference picture is a long-term reference picture or the second reference picture is a long-term reference picture.

[0010] A method for decoding a video signal according to an embodiment of the present disclosure includes obtaining an MMVD offset as information related to a second MMVD when an absolute value of a first POC difference is smaller than an absolute value of a second POC difference; obtaining information related to a first MMVD by scaling information related to the second MMVD when a first reference picture is not a long-term reference picture and a second reference picture is not a long-term reference picture; and obtaining information related to the first MMVD without scaling an absolute value of the information related to the second MMVD when the first reference picture is a long-term reference picture or the second reference picture is a long-term reference picture.

[0011] In a method for decoding a video signal according to an embodiment of the present disclosure, obtaining information (mMvdLX) related to MMVD includes obtaining an MMVD offset based on information related to a distance of MMVD and information related to a direction of MMVD; obtaining, without scaling the MMVD offset, as information related to a first MMVD related to a first reference list when only the first reference list is used; and obtaining, without scaling the MMVD offset, as information related to a second MMVD related to a second reference list when only the second reference list is used.

[0012] A method for decoding a video signal according to an embodiment of the present disclosure includes obtaining chroma component format information from a higher-level bitstream, obtaining information regarding width (SubWidthC) and information regarding height (SubHeightC) based on the chroma component format information, obtaining x-axis scale information based on the information regarding width or information regarding the color component of the current block, obtaining y-axis scale information based on the information regarding height or information regarding the color component of the current block, determining the position of the left block based on the y-axis scale information, determining the position of the upper block based on the x-axis scale information, determining a weighting value based on the left block and the upper block, obtaining a first sample by predicting the current block in merge mode, obtaining a second sample by predicting the current block in intra mode, and obtaining a combined prediction sample for the current block based on the weighting value, the first sample, and the second sample.

[0013] In a method for decoding a video signal according to an embodiment of the present disclosure, the step of determining a weighting value includes setting the code information (isIntraCodedNeighbourA) regarding the left block to TRUE when the left block is available and the prediction mode of the left block is intra prediction, setting the code information regarding the left block to FALSE when the left block is not available or the prediction mode of the left block is not intra prediction, setting the code information (isIntraCodedNeighbourB) regarding the upper block to TRUE when the upper block is available and the prediction mode of the upper block is intra prediction, and setting the code information regarding the upper block to FALSE when the upper block is not available or the prediction mode of the upper block is not intra prediction.

[0014] In a method for decoding a video signal according to an embodiment of the present disclosure, the step of determining a weighting value includes determining the weighting value as 3 when both the code information regarding the left block and the code information regarding the upper block are TRUE; determining the weighting value as 1 when both the code information regarding the left block and the code information regarding the upper block are FALSE; and determining the weighting value as 2 when only one of the code information regarding the left block and the code information regarding the upper block is TRUE.

[0015] In a method for decoding a video signal according to an embodiment of the present disclosure, the step of obtaining a combined prediction sample includes predicting a current block based on predSamplesComb[x][y]=(w*predSamplesIntra[x][y]+(4 - w)*predSamplesInter[x][y]+2)>>2, where predSamplesComb means a combined prediction sample, w means a weighting value, predSamplesIntra means a second sample, predSamplesInter means a first sample, [x] means the x-axis coordinate of a sample included in the current block, and [y] means the y-axis coordinate of a sample included in the current block.

[0016] In a method for decoding a video signal according to an embodiment of the present disclosure, the step of obtaining x-axis scale information includes determining the x-axis scale information as 0 when the color component of the current block is 0 or the information regarding the width is 1, and determining the x-axis scale information as 1 when the color component of the current block is not 0 and the information regarding the width is not 1. The step of obtaining y-axis scale information includes determining the y-axis scale information as 0 when the color component of the current block is 0 or the information regarding the height is 1, and determining the y-axis scale information as 1 when the color component of the current block is not 0 and the information regarding the height is not 1.

[0017] In a method for decoding a video signal according to an embodiment of the present disclosure, the position of the left block is (xCb - 1, yCb - 1+(cbHeight << scallFactHeight)), where xCb is the x-axis coordinate of the upper-left sample of the current luma block, yCb is the y-axis coordinate of the upper-left sample of the current luma block, cbHeight is the size of the height of the current block, scallFactHeight is the scale information of the y-axis, the position of the upper block is (xCb - 1+(cbWidth << scallFactWidth), yCb - 1), where xCb is the x-axis coordinate of the upper-left sample of the current luma block, yCb is the y-axis coordinate of the upper-left sample of the current luma block, cbWidth is the size of the width of the current block, and scallFactWidth is the scale information of the x-axis.

[0018] An apparatus for decoding a video signal according to an embodiment of the present disclosure includes a processor and a memory. The processor parses from a bitstream upper-level MMVD (Merge with MVD) active information (sps_mmvd_enabled_flag) indicating whether upper-level MMVD used for the current block is used based on instruction words stored in the memory. When the upper-level MMVD active information indicates the activation of MMVD, the processor parses from the bitstream MMVD merge information (mmvd_merge_flag) indicating whether to use MMVD for the current block. When the MMVD merge information indicates using MMVD for the current block, the processor parses information related to the distance of MMVD (mmvd_distance_idx) and information related to the direction of MMVD (mmvd_direction_idx), and obtains information about MMVD (mMvdLX) based on the information related to the distance of MMVD and the information related to the direction of MMVD. The information about MMVD is greater than or equal to -2^17 and less than or equal to 2^17 - 1.

[0019] In an apparatus for decoding a video signal according to an embodiment of the present disclosure, the upper level is any one of a Coding Tree Unit, a slice, a tile, a tile group, a picture, or a sequence unit.

[0020] In an apparatus for decoding a video signal according to an embodiment of the present disclosure, the processor generates a merge candidate list for a current block based on instruction words stored in a memory, selects a motion vector from the merge candidate list based on a merge index parsed from a bitstream, adds information regarding MMVD to the motion vector when MMVD merge information indicates that MMVD is to be used for the current block, obtains a modified motion vector, restores the current block based on the modified motion vector, and the modified motion vector is greater than or equal to -2^17 and less than or equal to 2^17 - 1.

[0021] In an apparatus for decoding a video signal according to an embodiment of the present disclosure, the processor obtains an MMVD offset based on information related to the distance of MMVD and information related to the direction of MMVD based on instruction words stored in a memory, obtains a difference in POC (Picture Order Count) between a current picture including the current block and a first reference picture based on the first reference list as a first POC difference when the first reference list and the second reference list are used, obtains a difference in POC (Picture Order Count) between the current picture and a second reference picture based on the second reference list as a second POC difference, obtains information regarding a first MMVD related to the first reference list and information regarding a second MMVD related to the second reference list based on at least one of the MMVD offset, the first POC difference, and the second POC difference, and the information regarding MMVD includes the information regarding the first MMVD and the information regarding the second MMVD.

[0022] In an apparatus for decoding a video signal according to an embodiment of the present disclosure, the processor obtains the MMVD offset as information regarding the first MMVD when the first POC difference and the second POC difference are the same, based on instruction words stored in a memory, and obtains the MMVD offset as information regarding the second MMVD.

[0023] In an apparatus for decoding a video signal according to an embodiment of the present disclosure, the processor obtains the MMVD offset as information regarding the first MMVD when the absolute value of the first POC difference is greater than or equal to the absolute value of the second POC difference, based on instruction words stored in a memory. When the first reference picture is not a long-term reference picture and the second reference picture is not a long-term reference picture, the information regarding the first MMVD is scaled to obtain the information regarding the second MMVD. When the first reference picture is a long-term reference picture or the second reference picture is a long-term reference picture, the information regarding the second MMVD is obtained without scaling the absolute value of the information regarding the first MMVD.

[0024] In an apparatus for decoding a video signal according to an embodiment of the present disclosure, the processor obtains the MMVD offset as information regarding the second MMVD when the absolute value of the first POC difference is less than the absolute value of the second POC difference, based on instruction words stored in a memory. When the first reference picture is not a long-term reference picture and the second reference picture is not a long-term reference picture, the information regarding the second MMVD is scaled to obtain the information regarding the first MMVD. When the first reference picture is a long-term reference picture or the second reference picture is a long-term reference picture, the information regarding the first MMVD is obtained without scaling the absolute value of the information regarding the second MMVD.

[0025] In an apparatus for decoding a video signal according to an embodiment of the present disclosure, a processor obtains an MMVD offset based on information related to the distance of the MMVD and information related to the direction of the MMVD based on instruction words stored in a memory, and when only the first reference list is used, the MMVD offset is not scaled and is obtained as information related to a first MMVD related to the first reference list, and when only the second reference list is used, the MMVD offset is not scaled and is obtained as information related to a second MMVD related to the second reference list.

[0026] An apparatus for decoding a video signal according to an embodiment of the present disclosure includes a processor and a memory. The processor obtains chroma component format information from a higher-level bitstream based on instruction words stored in the memory, obtains information related to width (SubWidthC) and information related to height (SubHeightC) based on the chroma component format information, obtains x-axis scale information based on the information related to width or information related to the color component of the current block, obtains y-axis scale information based on the information related to height or information related to the color component of the current block, determines the position of the left block based on the y-axis scale information, determines the position of the upper block based on the x-axis scale information, determines a weighting value based on the left block and the upper block, obtains a first sample obtained by predicting the current block in the merge mode, obtains a second sample obtained by predicting the current block in the intra mode, and obtains a combined prediction sample for the current block based on the weighting value, the first sample, and the second sample.

[0027] In an apparatus for decoding a video signal according to an embodiment of the present disclosure, based on instruction words stored in a memory, when a left block is available and the prediction mode of the left block is intra prediction, the processor sets code information (isIntraCodedNeighbourA) regarding the left block to TRUE, and when the left block is not available or the prediction mode of the left block is not intra prediction, sets the code information regarding the left block to FALSE. When an upper block is available and the prediction mode of the upper block is intra prediction, the processor sets code information (isIntraCodedNeighbourB) regarding the upper block to TRUE, and when the upper block is not available or the prediction mode of the upper block is not intra prediction, sets the code information regarding the upper block to FALSE.

[0028] In an apparatus for decoding a video signal according to an embodiment of the present disclosure, based on instruction words stored in a memory, when both the code information regarding the left block and the code information regarding the upper block are TRUE, the processor determines a weighting value to be 3, when both the code information regarding the left block and the code information regarding the upper block are FALSE, determines the weighting value to be 1, and when only one of the code information regarding the left block and the code information regarding the upper block is TRUE, determines the weighting value to be 2.

[0029] In an apparatus for decoding a video signal according to an embodiment of the present disclosure, a processor predicts a current block based on predSamplesComb[x][y]=(w*predSamplesIntra[x][y]+(4 - w)*predSamplesInter[x][y]+2)>>2 based on instruction words stored in a memory, where predSamplesComb means combined prediction samples, w means a weighting value, predSamplesIntra means second samples, predSamplesInter means first samples, [x] means the x-axis coordinate of a sample included in the current block, and [y] means the y-axis coordinate of a sample included in the current block.

[0030] In an apparatus for decoding a video signal according to an embodiment of the present disclosure, a processor determines that the scale information of the x-axis is 0 when the color component of the current block is 0 or the information regarding the width is 1, determines that the scale information of the x-axis is 1 when the color component of the current block is not 0 and the information regarding the width is not 1, determines that the scale information of the y-axis is 0 when the color component of the current block is 0 or the information regarding the height is 1, and determines that the scale information of the y-axis is 1 when the color component of the current block is not 0 and the information regarding the height is not 1.

[0031] The position of the left block of the apparatus for decoding a video signal according to an embodiment of the present disclosure is (xCb - 1, yCb - 1+(cbHeight << scallFactHeight)), where xCb is the x - axis coordinate of the upper - left sample of the current luma block, yCb is the y - axis coordinate of the upper - left sample of the current luma block, cbHeight is the size of the height of the current block, scallFactHeight is the scale information of the y - axis. The position of the upper block is (xCb - 1+(cbWidth << scallFactWidth), yCb - 1), where xCb is the x - axis coordinate of the upper - left sample of the current luma block, yCb is the y - axis coordinate of the upper - left sample of the current luma block, cbWidth is the size of the width of the current block, scallFactWidth is the scale information of the x - axis, which is characterized by the above.

[0032] A method for encoding a video signal according to an embodiment of the present disclosure includes generating information (mMvdLX) regarding the MMVD of the current block, generating information related to the distance of the MMVD and information related to the direction of the MMVD based on the information (mMvdLX) regarding the MMVD, generating MMVD merge information (mmvd_merge_flag) indicating whether to use the MMVD for the current block, generating upper - level MMVD activation information (sps_mmvd_enabled_flag) indicating whether the upper - level MMVD (Merge with MVD) including the current block is used, and generating a bitstream based on the information related to the distance of the MMVD, the information related to the direction of the MMVD, the MMVD merge information (mmvd_merge_flag), and the upper - level MMVD activation information (sps_mmvd_enabled_flag). The information regarding the MMVD is greater than or equal to - 2^17 and less than or equal to 2^17 - 1, which is characterized by the above.

[0033] An apparatus for encoding a video signal according to an embodiment of the present disclosure includes a processor and a memory. The processor generates information (mMvdLX) regarding the MMVD of the current block based on instruction words stored in the memory, generates information related to the distance of the MMVD and information related to the direction of the MMVD based on the information (mMvdLX) regarding the MMVD, generates MMVD merge information (mmvd_merge_flag) indicating whether to use the MMVD for the current block, generates upper-level MMVD activation information (sps_mmvd_enabled_flag) indicating whether to use the upper-level MMVD (Merge with MVD) including the current block, generates a bitstream based on the information related to the distance of the MMVD, the information related to the direction of the MMVD, the MMVD merge information (mmvd_merge_flag), and the upper-level MMVD activation information (sps_mmvd_enabled_flag), and the information regarding the MMVD is greater than or equal to -2^17 and less than or equal to 2^17-1.

[0034] A method for encoding a video signal according to an embodiment of the present disclosure includes generating upper-level chroma component format information, obtaining information (SubWidthC) regarding the width and information (SubHeightC) regarding the height based on the chroma component format information, obtaining x-axis scale information based on the information regarding the width or information regarding the color component of the current block, obtaining y-axis scale information based on the information regarding the height or information regarding the color component of the current block, determining the position of the left block based on the y-axis scale information, determining the position of the upper block based on the x-axis scale information, determining a weighting value based on the left block and the upper block, obtaining a first sample predicting the current block in the merge mode, obtaining a second sample predicting the current block in the intra mode, and obtaining a combined prediction sample for the current block based on the weighting value, the first sample, and the second sample.

[0035] An apparatus for encoding a video signal according to an embodiment of the present disclosure includes a processor and a memory. The processor generates upper-level chroma component format information based on instruction words stored in the memory, obtains information related to width (SubWidthC) and information related to height (SubHeightC) based on the chroma component format information, obtains x-axis scale information based on the information related to width or information related to the color component of the current block, obtains y-axis scale information based on the information related to height or information related to the color component of the current block, determines the position of the left block based on the y-axis scale information, determines the position of the upper block based on the x-axis scale information, determines a weighting value based on the left block and the upper block, obtains a first sample obtained by predicting the current block in merge mode, obtains a second sample obtained by predicting the current block in intra mode, and obtains a combined prediction sample for the current block based on the weighting value, the first sample, and the second sample.

Advantages of the Invention

[0036] According to an embodiment of the present invention, the coding efficiency of a video signal can be improved.

Brief Description of the Drawings

[0037]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Mode for Carrying Out the Invention

[0038] In consideration of the functions in the present disclosure, the terms used in this specification are, as much as possible, general terms that are currently widely used. However, this may change depending on the intentions of those skilled in the art, conventions, or the emergence of new technologies. Also, in certain cases, there are terms arbitrarily selected by the applicant, and in this case, their meanings will be described in the description part of the corresponding invention. Therefore, it is clear that the terms used in this specification should be interpreted based on the substantial meanings of the terms rather than simply their names and the content throughout this specification.

[0039] In this specification, some terms may be interpreted as follows. Coding may, in some cases, be interpreted as encoding or decoding. In this specification, a device that encodes (codes) a video signal to generate a video signal bitstream is referred to as an encoding device or an encoder, and a device that decodes (decodes) a video signal bitstream to restore the video signal is referred to as a decoding device or a decoder. Also, in this specification, a video signal processing device is used as a term for a concept that includes both an encoder and a decoder. Information is a term that includes any of values, parameters, coefficients, elements, etc., and may be interpreted with different meanings depending on the case, so the present disclosure is not limited thereto. "Unit" is used to mean a basic unit of video processing or a specific position in a picture, and refers to an image area that includes both a luma component and a chroma component. Also, "block" refers to an image area that includes a specific component among the luma component and the chroma components (i.e., Cb and Cr). However, depending on the embodiment, terms such as "unit", "block", "partition", and "area" may be used with the same meaning. Also, in this specification, a unit may be used as a concept that includes any of a coding unit, a prediction unit, and a transform unit. A picture refers to a field or a frame, and those terms may be used with the same meaning depending on the embodiment.

[0040] Figure 1 is a schematic block diagram of a video signal encoding device according to an embodiment of the present disclosure. Referring to Figure 1, the encoding device 100 of the present disclosure includes a transform unit 110, a quantization unit 115, an inverse quantization unit 120, an inverse transform unit 125, a filtering unit 130, a prediction unit 150, and an entropy coding unit 160.

[0041] The conversion unit 110 converts the residual signal, which is the difference between the received video signal and the prediction signal generated by the prediction unit 150, to obtain conversion coefficient values. For example, a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or a Wavelet Transform may be used. The Discrete Cosine Transform and the Discrete Sine Transform divide the input picture signal into block forms for conversion. In the conversion, the coding efficiency may vary depending on the distribution and characteristics of the values within the conversion region. The quantization unit 115 quantizes the conversion coefficient values output from the conversion unit 110.

[0042] To improve the coding efficiency, instead of directly coding the picture signal, the prediction unit 150 uses the already coded region to predict the picture, and adds the residual value between the original picture and the predicted picture to the predicted picture to obtain the restored picture. To prevent a mismatch from occurring between the encoder and the decoder, when making a prediction in the encoder, information that can also be used by the decoder must be used. For this purpose, the encoder performs a process of restoring the encoded current block again. The inverse quantization unit 120 inverse quantizes the conversion coefficient values, and the inverse conversion unit 125 restores the residual values using the inverse quantized conversion coefficient values. On the other hand, the filtering unit 130 performs filtering operations for improving the quality of the restored picture and enhancing the coding efficiency. For example, it may include a deblocking filter, a Sample Adaptive Offset (SAO), and an adaptive loop filter. The picture that has undergone filtering is either output or stored in the Decoded Picture Buffer (DPB) 156 for use as a reference picture.

[0043] The prediction unit 150 includes an intra prediction unit 152 and an inter prediction unit 154. The intra prediction unit 152 performs intra prediction within the current picture, and the inter prediction unit 154 performs inter prediction for predicting the current picture using the reference pictures stored in the decoded picture buffer 156. The intra prediction unit 152 performs intra prediction from the restored samples within the current picture and transmits the intra coding information to the entropy coding unit 160. The intra coding information can include at least one of an intra prediction mode, an MPM (Most Probable Mode) flag, and an MPM index. The intra coding information can include information regarding reference samples. The inter prediction unit 154 may be configured to include a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains a motion vector value of the current region by referring to a specific region of the restored reference picture. The motion estimation unit 154a transmits a set of motion information regarding the reference region (such as a reference picture index, motion vector information, etc.) to the entropy coding unit 160. The motion compensation unit 154b performs motion compensation using the motion vector value transmitted from the motion estimation unit 154a. The inter prediction unit 154 transmits inter coding information including motion information regarding the reference region to the entropy coding unit 160.

[0044] According to a further embodiment, the prediction unit 150 can include an intra block copy (BC) prediction unit (not shown). The intra BC prediction unit performs intra BC prediction from the restored samples in the current picture and transmits the intra BC coding information to the entropy coding unit 160. The intra BC prediction unit refers to a specific region in the current picture and obtains a block vector value indicating a reference region used for predicting the current region. The intra BC prediction unit can perform intra BC prediction using the obtained block vector value. The intra BC prediction unit transmits the intra BC coding information to the entropy coding unit 160. The intra BC coding information can include block vector information.

[0045] When the picture prediction as described above is performed, the conversion unit 110 converts the residual value between the original picture and the predicted picture to obtain a conversion coefficient value. At this time, the conversion may be performed in units of specific blocks within the picture, and the size of the specific block may vary within a preset range. The quantization unit 115 quantizes the conversion coefficient value generated by the conversion unit 110 and transmits it to the entropy coding unit 160.

[0046] The entropy coding unit 160 performs entropy coding on information indicating quantized transform coefficients, intra-coding information, inter-coding information, etc. to generate a video signal bitstream. In the entropy coding unit 160, a variable length coding (VLC) method, an arithmetic coding method, etc. may be used. The variable length coding (VLC) method converts an input symbol into a continuous codeword, and the length of the codeword may be variable. For example, a frequently occurring symbol is represented by a short codeword, and a symbol that does not occur frequently is represented by a long codeword. As the variable length coding method, a context-based adaptive variable length coding (CAVLC) method may be used. Arithmetic coding converts continuous data symbols into a single prime number, and arithmetic coding can obtain the optimal prime number bits required to represent each symbol. As arithmetic coding, context-based adaptive binary arithmetic code (CABAC) may be used. For example, the entropy coding unit 160 can binaryize the information indicating the quantized transform coefficients. Further, the entropy coding unit 160 can perform arithmetic coding on the binaryized information to generate a bitstream.

[0047] The generated bitstream is encapsulated in units of NAL (Network Abstraction Layer) units. An NAL unit contains an integer number of encoded coding tree units. In order to decode the bitstream with a video decoder, first, the bitstream must be separated into NAL unit units, and then each separated NAL unit must be decoded. On the other hand, the information necessary for decoding the video signal bitstream may be transmitted in the RBSP (Raw Byte Sequence Payload) of higher-level sets such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS).

[0048] On the other hand, the block diagram of FIG. 1 shows an encoding device 100 according to an embodiment of the present disclosure. However, the separately shown blocks logically distinguish the elements of the encoding device 100. Therefore, the elements of the encoding device 100 described above may be mounted as one chip or a plurality of chips according to the design of the device. According to one embodiment, the operations of each element of the encoding device 100 described above may be performed by a processor (not shown).

[0049] FIG. 2 is a schematic block diagram of a video signal decoding device 200 according to an embodiment of the present disclosure. Referring to FIG. 2, the decoding device 200 of the present disclosure includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 225, a filtering unit 230, and a prediction unit 250.

[0050] The entropy decoding unit 210 entropy-decodes the video signal bitstream and extracts transform coefficient information, intra-coding information, inter-coding information, etc. for each region. For example, the entropy decoding unit 210 can obtain the binary code for the transform coefficient information of a specific region from the video signal bitstream. Also, the entropy decoding unit 210 inverse-binarizes the binary code to obtain the quantized transform coefficients. The inverse quantization unit 220 inverse-quantizes the quantized transform coefficients, and the inverse transform unit 225 restores the residual value using the inverse-quantized transform coefficients. The video signal processing apparatus 200 adds the residual value obtained from the inverse transform unit 225 to the predicted value obtained from the prediction unit 250 to restore the original pixel value.

[0051] On the other hand, the filtering unit 230 performs filtering on the picture to improve the picture quality. This may include a deblocking filter for reducing the block distortion phenomenon and / or an adaptive loop filter for removing distortion of the entire picture. The picture that has undergone filtering is either output or stored in the decoded picture buffer (DPB) 256 for use as a reference picture for the next picture.

[0052] The prediction unit 250 includes an intra prediction unit 252 and an inter prediction unit 254. The prediction unit 250 generates a predicted picture using the encoding type decoded by the entropy decoding unit 210 described above, the transform coefficients for each region, the intra / inter encoding information, and the like. To restore the current block for which decoding is performed, the decoded regions of the current picture or other pictures including the current block may be used. A picture (or tile / slice) that uses only the current picture for restoration, that is, performs only intra prediction (or intra prediction and intra BC prediction), is called an intra picture or I picture (or tile / slice), and a picture (or tile / slice) that can perform both intra prediction and inter prediction is called an inter picture (or tile / slice). Among the inter pictures (or tiles / slices), a picture (or tile / slice) that uses at most one motion vector and a reference picture index to predict the sample values of each block is called a predictive picture or P picture (or tile / slice), and a picture (or tile / slice) that uses at most two motion vectors and a reference picture index is called a Bi-predictive picture or B picture (or tile / slice). In other words, a P picture (or tile / slice) uses at most one set of motion information to predict each block, and a B picture (or tile / slice) uses at most two sets of motion information to predict each block. Here, a set of motion information includes one or more motion vectors and one reference picture index.

[0053] The intra prediction unit 252 generates a prediction block using the intra-coded information and the restored samples in the current picture. As described above, the intra-coded information can include at least one of an intra prediction mode, an MPM (Most Probable Mode) flag, and an MPM index. The intra prediction unit 252 predicts the sample values of the current block using the restored samples located on the left side and / or the upper side of the current block as reference samples. In the present disclosure, the restored samples, the reference samples, and the samples of the current block can represent pixels. Also, the sample value can represent a pixel value.

[0054] According to an embodiment, the reference samples may be samples included in the peripheral blocks of the current block. For example, the reference samples may be samples adjacent to the left boundary of the current block and / or samples adjacent to the upper boundary of the current block. Also, the reference samples may be samples located on a line within a preset distance from the left boundary of the current block and / or samples located on a line within a preset distance from the upper boundary of the current block among the samples of the peripheral blocks of the current block. At this time, the peripheral blocks of the current block can include at least one of a left (L) block adjacent to the current block, an upper (A) block, a below left (BL) block, an above right (AR) block, or an above left (AL) block.

[0055] The inter prediction unit 254 generates a prediction block using the reference pictures stored in the decoded picture buffer 256 and the inter-coded information. The inter-coded information can include a set of motion information (such as a reference picture index, motion vector information, etc.) for the current block with respect to the reference block. Inter prediction can include L0 prediction, L1 prediction, and bi-prediction. L0 prediction is a prediction using one reference picture included in the L0 picture list, and L1 prediction means a prediction using one reference picture included in the L1 picture list. For this purpose, a set of motion information (e.g., a motion vector and a reference picture index) may be required. In the bi-prediction method, up to two reference regions can be utilized, and these two reference regions may exist in the same reference picture or in different pictures respectively. That is, in the bi-prediction method, up to two sets of motion information (e.g., a motion vector and a reference picture index) may be used, but the two motion vectors may correspond to the same reference picture index or different reference picture indexes. At this time, the reference picture may be displayed (or output) either before or after the current picture in terms of time. Also, in the bi-prediction method, the two reference regions may exist in two reference picture lists.

[0056] The inter prediction unit 254 can obtain a reference block of the current block using the motion vector and the reference picture index. The reference block exists in a reference picture corresponding to the reference picture index. Also, the sample value of the block specified by the motion vector or its interpolated value may be used as a predictor of the current block. For motion prediction with pixel accuracy at the sub-pel unit, for example, an 8-tap interpolation filter may be used for the luma signal, and a 4-tap interpolation filter may be used for the chroma signal. However, the interpolation filter for motion prediction at the sub-pel unit is not limited to this. In this way, the inter prediction unit 254 performs motion compensation to predict the texture of the current unit from a previously reconstructed picture. At this time, the inter prediction unit can use the motion information set.

[0057] According to a further embodiment, the prediction unit 250 can include an intra BC prediction unit (not shown). The intra BC prediction unit performs intra BC prediction from the reconstructed samples in the current picture and transmits the intra BC coding information to the entropy coding unit 160. The intra BC prediction unit obtains a block vector value of the current region that indicates a specific region in the current picture. The intra BC prediction unit can perform intra BC prediction using the obtained block vector value. The intra BC prediction unit transmits the intra BC coding information to the entropy coding unit 160. The intra BC coding information can include block vector information.

[0058] The predicted value output from the intra prediction unit 252 or the inter prediction unit 254 and the residual value output from the inverse transform unit 225 are added to generate a reconstructed video picture. That is, the video signal decoding apparatus 200 restores the current block using the predicted block generated by the prediction unit 250 and the residual obtained from the inverse transform unit 225.

[0059] On the one hand, the block diagram of FIG. 2 shows a decoding device 200 according to an embodiment of the present disclosure. The separately shown blocks logically distinguish the elements of the decoding device 200. Therefore, the elements of the decoding device 200 described above may be mounted as one chip or multiple chips depending on the device design. According to one embodiment, the operation of each element of the decoding device 200 described above may be performed by a processor (not shown).

[0060] FIG. 3 shows an embodiment in which a Coding Tree Unit (CTU) is divided into Coding Units (CUs) within a picture. In the coding process of a video signal, a picture may be divided into a sequence of Coding Tree Units (CTUs). A Coding Tree Unit is composed of an NXN block of luma samples and two blocks of corresponding chroma samples. A Coding Tree Unit may be divided into multiple Coding Units. A Coding Tree Unit may also become a leaf node without being divided. In this case, the Coding Tree Unit itself can become a Coding Unit. A Coding Unit refers to a basic unit for processing a picture in the above-described video signal processing process, that is, processes such as intra / inter prediction, conversion, quantization, and / or entropy coding. The size and shape of Coding Units within one picture do not have to be constant. A Coding Unit may have a square or rectangular shape. A rectangular Coding Unit (or rectangular block) includes a vertical Coding Unit (or vertical block) and a horizontal Coding Unit (or horizontal block). In this specification, a vertical block is a block whose height is greater than its width, and a horizontal block is a block whose width is greater than its height. Also, in this specification, a non-square block can represent a rectangular block, but the present disclosure is not limited thereto.

[0061] Referring to FIG. 3, the coding tree unit is first divided into a Quad Tree (QT) structure. That is, in the quad tree structure, one node having a size of 2N×2N may be divided into four nodes having a size of N×N. In this specification, the quad tree can also be called a quaternary tree. The quad tree division may be performed recursively, and it is not necessary for all nodes to be divided to the same depth.

[0062] On the other hand, the leaf node of the aforementioned quad tree may be further divided into a Multi-Type Tree (MTT) structure. According to an embodiment of the present disclosure, in the multi-type tree structure, one node may be divided into a binary or ternary tree structure of horizontal or vertical division. That is, there are four division structures in the multi-type tree structure: vertical binary division, horizontal binary division, vertical ternary division, and horizontal ternary division. According to an embodiment of the present disclosure, in each of the above tree structures, both the width and height of the node may have a value that is a power of 2. For example, in a Binary Tree (BT) structure, a node with a size of 2N×2N may be divided into two nodes of N×2N by vertical binary division and into two nodes of 2N×N by horizontal binary division. Also, in a Ternary Tree (TT) structure, a node with a size of 2N×2N may be divided into nodes of (N / 2)×2N, N×2N, and (N / 2)×2N by vertical ternary division and into nodes of 2N×(N / 2), 2N×N, and 2N×(N / 2) by horizontal binary division. Such multi-type tree division may be performed recursively.

[0063] The leaf nodes of a multi-type tree can be coding units. When no splitting is indicated for a coding unit or the coding unit is not larger than the maximum transform length, the corresponding coding unit is used as a prediction and transformation unit without further splitting. On the other hand, in the aforementioned quad-tree and multi-type tree, at least one of the following parameters may be predefined or transmitted in the RBSP of a higher-level set such as PPS, SPS, VPS, etc. 1) CTU size: the size of the root node of the quad-tree, 2) Minimum QT size (MinQtSize): the minimum allowable QT leaf node size, 3) Maximum BT size (MaxBtSize): the maximum allowable BT root node size, 4) Maximum TT size (MaxTtSize): the maximum allowable TT root node size, 5) Maximum MTT depth (MaxMttDepth): the maximum allowable depth of MTT splitting from the leaf node of the QT, 6) Minimum BT size (MinBtSize): the minimum allowable BT leaf node size, 7) Minimum TT size (MinTtSize): the minimum allowable TT leaf node size.

[0064] Figure 4 shows an example of a method for signaling the splitting of a quad-tree and a multi-type tree. To signal the splitting of the aforementioned quad-tree and multi-type tree, already set flags may be used. Referring to Figure 4, at least one of the flag 'qt_split_flag' indicating whether a quad-tree node is split, the flag'mtt_split_flag' indicating whether a multi-type tree node is split, the flag'mtt_split_vertical_flag' indicating the splitting direction of a multi-type tree node, or the flag'mtt_split_binary_flag' indicating the splitting shape of a multi-type tree node may be used.

[0065] According to an embodiment of the present disclosure, the coding tree unit is the root node of a quad tree and may first be divided into a quad tree structure. In the quad tree structure, a 'qt_split_flag' is signaled for each node 'QT_node'. When the value of 'qt_split_flag' is 1, the corresponding node is divided into four square nodes. When the value of 'qt_split_flag' is 0, the corresponding node becomes a leaf node 'QT_leaf_node' of the quad tree.

[0066] Each quad tree leaf node 'QT_leaf_node' may be further divided into a multi-type tree structure. In the multi-type tree structure, an'mtt_split_flag' is signaled for each node 'MTT_node'. When the value of'mtt_split_flag' is 1, the corresponding node is divided into a plurality of rectangular nodes. When the value of'mtt_split_flag' is 0, the corresponding node becomes a leaf node 'MTT_leaf_node' of the multi-type tree. When the multi-type tree node 'MTT_node' is divided into a plurality of rectangular nodes (i.e., when the value of'mtt_split_flag' is 1), an'mtt_split_vertical_flag' and an'mtt_split_binary_flag' for the node 'MTT_node' may be further signaled. When the value of'mtt_split_vertical_flag' is 1, a vertical split of the node 'MTT_node' is indicated. When the value of'mtt_split_vertical_flag' is 0, a horizontal split of the node 'MTT_node' is indicated. Also, when the value of'mtt_split_binary_flag' is 1, the node 'MTT_node' is divided into two rectangular nodes. When the value of'mtt_split_binary_flag' is 0, the node 'MTT_node' is divided into three rectangular nodes.

[0067] Picture prediction (motion compensation) for coding is performed on coding units that are not further divided (i.e., leaf nodes of the coding unit tree). The basic unit for performing such prediction is hereinafter referred to as a prediction unit or a prediction block.

[0068] Hereinafter, the term "unit" used in this specification may be used as a term substituting for the prediction unit which is the basic unit for performing prediction. However, the present disclosure is not limited thereto, and more broadly, it may be understood in a concept including the coding unit.

[0069] FIGS. 5 and 6 show more specifically the intra prediction method according to an embodiment of the present disclosure. As described above, the intra prediction unit predicts the sample value of the current block using the restored samples located on the left side and / or the upper side of the current block as reference samples.

[0070] First, FIG. 5 shows an example of reference samples used for predicting the current block in the intra prediction mode. According to an example, the reference samples may be samples adjacent to the left boundary of the current block and / or samples adjacent to the upper boundary of the current block. As shown in FIG. 5, when the size of the current block is W×H and samples of a single reference line adjacent to the current block are used for intra prediction, the reference samples may be set using up to 2W+2H+1 peripheral samples located on the left side and / or the upper side of the current block.

[0071] According to a further embodiment of the present disclosure, samples on a plurality of reference lines may be used for intra prediction of the current block. The plurality of reference lines may be composed of n lines located within a preset distance from the boundary of the current block. In this case, another reference line information indicating at least one reference line used for intra prediction of the current block may be signaled. Specifically, the reference line information may include an index indicating any one of the plurality of reference lines.

[0072] Also, when at least some of the samples used as reference samples have not yet been restored, the intra prediction unit can perform a reference sample padding process to obtain reference samples. Further, the intra prediction unit can perform a reference sample filtering process to reduce the error of intra prediction. That is, filtering can be performed on the surrounding samples and / or the reference samples obtained by the reference sample padding process to obtain filtered reference samples. The intra prediction unit predicts the samples of the current block using the reference samples thus obtained. The intra prediction unit predicts the samples of the current block using the unfiltered reference samples or the filtered reference samples. In the present disclosure, the surrounding samples can include samples on at least one reference line. For example, the surrounding samples can include adjacent samples on a line adjacent to the boundary of the current block.

[0073] Next, FIG. 6 shows an embodiment of a prediction mode used for intra prediction. For intra prediction, intra prediction mode information indicating an intra prediction direction may be signaled. The intra prediction mode information indicates any one of a plurality of intra prediction modes constituting an intra prediction mode set. When the current block is an intra prediction block, the decoder receives the intra prediction mode information of the current block from the bitstream. The intra prediction unit of the decoder performs intra prediction on the current block based on the extracted intra prediction mode information.

[0074] According to an embodiment of the present disclosure, the intra prediction mode set may include all intra prediction modes used for intra prediction (for example, a total of 67 intra prediction modes). More specifically, the intra prediction mode set may include a planar mode, a DC mode, and a plurality of (for example, 65) angular modes (i.e., direction modes). Each intra prediction mode may be indicated by a pre-set index (i.e., an intra prediction mode index). For example, as shown in FIG. 6, the intra prediction mode index 0 indicates the planar mode, and the intra prediction mode index 1 indicates the DC mode. Also, the intra prediction mode indices 2 to 66 may each indicate different angular modes. The angular modes each indicate different angles within a pre-set angular range. For example, the angular modes may indicate angles within an angular range of 45° to -135° (i.e., a first angular range) in the clockwise direction. The angular mode may be defined based on the 12 o'clock direction. At this time, the intra prediction mode index 2 indicates the Horizontal Diagonal (HDIA) mode, the intra prediction mode index 18 indicates the Horizontal (HOR) mode, the intra prediction mode index 34 indicates the Diagonal (DIA) mode, the intra prediction mode index 50 indicates the Vertical (VER) mode, and the intra prediction mode index 66 indicates the Vertical Diagonal (VDIA) mode.

[0075] On the one hand, the already set angle range may be set to be different according to the shape of the current block. For example, when the current block is a rectangular block, a wide-angle mode indicating an angle exceeding 45° or less than -135° in the clockwise direction may be further used. When the current block is a horizontal block, the angle mode can indicate an angle within the angle range of (45 + offset1)° to (-135 + offset1)° (i.e., the second angle range) in the clockwise direction. At this time, angle modes 67 to 80 that deviate from the first angle range may be further used. Also, when the current block is a vertical block, the angle mode can indicate an angle within the angle range of (45 - offset2)° to (-135 - offset2)° (i.e., the third angle range) in the clockwise direction. At this time, angle modes -10 to -1 that deviate from the first angle range may be further used. According to an embodiment of the present disclosure, the values of offset1 and offset2 may be individually determined according to the ratio of the width and height of the rectangular block. Also, offset1 and offset2 may be positive numbers.

[0076] According to a further embodiment of the present disclosure, the plurality of angle modes constituting the intra prediction mode set may include a basic angle mode and an extended angle mode. At this time, the extended angle mode may be determined based on the basic angle mode.

[0077] According to one embodiment, the basic angle mode is a mode corresponding to the angle used in the intra prediction of the existing HEVC (High Efficiency Video Coding) standard, and the extended angle mode may be a mode corresponding to the angle newly added in the intra prediction of the next-generation video codec standard. More specifically, the basic angle mode is an angle mode corresponding to any one of the intra prediction modes {2, 4, 6, …, 66}, and the extended angle mode may be an angle mode corresponding to any one of the intra prediction modes {3, 5, 7, …, 65}. That is, the extended angle mode may be an angle mode between the basic angle modes within the first angle range. Therefore, the angle indicated by the extended angle mode may be determined based on the angle indicated by the basic angle mode.

[0078] According to another embodiment, the basic angle mode may be a mode corresponding to an angle within a preset first angle range, and the extended angle mode may be a wide-angle mode outside the first angle range. That is, the basic angle mode may be an angle mode corresponding to any one of the intra prediction modes {2, 3, 4, …, 66}, and the extended angle mode may be an angle mode corresponding to any one of the intra prediction modes {-10, -9, …, -1} and {67, 68, …, 76}. The angle indicated by the extended angle mode may be determined as the opposite angle of the angle indicated by the corresponding basic angle mode. Therefore, the angle indicated by the extended angle mode may be determined based on the angle indicated by the basic angle mode. On the other hand, the number of extended angle modes is not limited thereto, and further extended angles may be defined according to the size and / or shape of the current block. For example, the extended angle mode may be defined as an angle mode corresponding to any one of the intra prediction modes {-14, -13, …, -1} and {67, 68, …, 80}. On the other hand, the total number of intra prediction modes included in the intra prediction mode set may vary according to the configurations of the aforementioned basic angle mode and extended angle mode.

[0079] In the above embodiment, the interval between the extended angle modes may be set based on the interval between the corresponding basic angle modes. For example, the interval between the extended angle modes {3, 5, 7, …, 65} may be determined based on the interval between the corresponding basic angle modes {2, 4, 6, …, 66}. Also, the interval between the extended angle modes {-10, -9, …, -1} may be determined based on the interval between the corresponding opposite basic angle modes {56, 57, …, 65}, and the interval between the extended angle modes {67, 68, …, 76} may be determined based on the interval between the corresponding opposite basic angle modes {3, 4, …, 12}. The angular interval between the extended angle modes may be set to be the same as the angular interval between the corresponding basic angle modes. Also, the number of extended angle modes in the intra prediction mode set may be set to be less than or equal to the number of basic angle modes.

[0080] According to an embodiment of the present disclosure, the extended angle mode may be signaled based on the basic angle mode. For example, the wide-angle mode (i.e., the extended angle mode) can replace at least one angle mode (i.e., the basic angle mode) within the first angle range. The basic angle mode to be replaced may be an angle mode corresponding to the opposite side of the wide-angle mode. That is, the basic angle mode to be replaced is an angle mode corresponding to an angle in the opposite direction of the angle indicated by the wide-angle mode or an angle offset from the angle in the opposite direction by a preset offset index. According to an embodiment of the present disclosure, the preset offset index is 1. The intra prediction mode index corresponding to the basic angle mode to be replaced is remapped to the wide-angle mode, and the corresponding wide-angle mode can be signaled. For example, the wide-angle modes {-10, -9, …, -1} may be signaled by the intra prediction mode indices {57, 58, …, 66} respectively, and the wide-angle modes {67, 68, …, 76} may be signaled by the intra prediction mode indices {2, 3, …, 11} respectively. In this way, by enabling the intra prediction mode index for the basic angle mode to signal the extended angle mode, even if the configurations of the angle modes used for intra prediction of each block are different from each other, the same set of intra prediction mode indices can be used for signaling the intra prediction mode. Therefore, the signaling overhead due to changes in the intra prediction mode configuration can be minimized.

[0081] On the one hand, whether the extended angle mode is used or not may be determined based on at least one of the shape and size of the current block. According to one embodiment, when the size of the current block is larger than a preset size, the extended angle mode is used for intra prediction of the current block; otherwise, only the basic angle mode may be used for intra prediction of the current block. According to other embodiments, when the current block is a non-square block, the extended angle mode is used for intra prediction of the current block; when the current block is a square block, only the basic angle mode may be used for intra prediction of the current block.

[0082] The intra prediction unit determines the reference samples and / or interpolated reference samples used for intra prediction of the current block based on the intra prediction mode information of the current block. When the intra prediction mode index indicates a specific angle mode, the reference sample or interpolated reference sample corresponding to the specific angle from the current sample of the current block is used for prediction of the current sample. Therefore, different sets of reference samples and / or interpolated reference samples may be used for intra prediction depending on the intra prediction mode. When intra prediction of the current block is performed using the reference samples and intra prediction mode information, the decoder adds the residual signal of the current block obtained from the inverse conversion unit to the intra prediction value of the current block to restore the sample value of the current block.

[0083] Hereinafter, with reference to FIG. 7, an inter prediction method according to an embodiment of the present disclosure will be described. In the present disclosure, the inter prediction method may include a general inter prediction method optimized for translational motion and an affine model-based inter prediction method. Also, the motion vector may include at least one of a general motion vector for motion compensation by the general inter prediction method and a control point motion vector for affine motion compensation.

[0084] FIG. 7 shows an inter-prediction method according to an embodiment of the present disclosure. As described above, the decoder can predict the current block by referring to the restored samples of other decoded pictures. Referring to FIG. 7, the decoder obtains a reference block 702 in a reference picture 720 based on the motion information set of the current block 701. At this time, the motion information set can include a reference picture index and a motion vector 703. The reference picture index indicates the reference picture 720 in which the reference block for the inter-prediction of the current block is included in the reference picture list. According to an embodiment, the reference picture list can include at least one of the L0 picture list or the L1 picture list described above. The motion vector 703 indicates an offset between the coordinate value of the current block 701 in the current picture 710 and the coordinate value of the reference block 702 in the reference picture 720. The decoder obtains a predictor of the current block 701 based on the sample value of the reference block 702 and restores the current block 701 using the predictor.

[0085] Specifically, the encoder can search for a block similar to the current block from the pictures with an earlier restoration order and obtain the reference block described above. For example, the encoder can search for a reference block with the minimum sum of the differences in sample values between the current block within a preset search area. At this time, at least one of SAD (Sum Of Absolute Difference) or SATD (Sum of Hadamard Transformed Difference) may be used to measure the similarity between the samples of the current block and the reference block. Here, SAD may be a value obtained by summing up all the absolute values of the differences in sample values included in the two blocks. Also, SATD may be a value obtained by summing up all the absolute values of the Hadamard transform coefficients obtained by performing a Hadamard transform on the differences in sample values included in the two blocks.

[0086] On the one hand, the current block may be predicted using one or more reference regions. As described above, the current block may be inter-predicted by a dual prediction method that utilizes two or more reference regions. According to one embodiment, the decoder may obtain two reference blocks based on two motion information sets of the current block. Further, the decoder may obtain a first predictor and a second predictor of the current block based on the sample values of each of the two obtained reference blocks. Further, the decoder may restore the current block using the first predictor and the second predictor. For example, the decoder may restore the current block based on the per-sample average of the first predictor and the second predictor.

[0087] As described above, one or more motion information sets may be signaled for motion compensation of the current block. At this time, the similarity between the motion information sets for motion compensation of each of the plurality of blocks may be used. For example, the motion information set used for prediction of the current block may be derived from the motion information set used for prediction of any one of the other samples that have already been restored. Thereby, the encoder and the decoder can reduce the signaling overhead.

[0088] For example, there may be a plurality of candidate blocks predicted based on a motion information set identical or similar to the motion information set of the current block. The decoder can generate a merge candidate list. The decoder can generate a merge candidate list based on the corresponding plurality of candidate blocks. Here, the merge candidate list can include candidates corresponding to samples predicted based on a motion information set related to the motion information set of the current block among the samples restored prior to the current block. The merge candidate list can include spatial candidates or temporal candidates. The merge candidate list may be generated based on the positions of the samples restored prior to the current block. The samples restored prior to the current block may be neighboring blocks of the current block. The neighboring blocks of the current block can mean the blocks adjacent to the current block. The encoder and the decoder can configure the merge candidate list of the current block according to predefined rules. At this time, the merge candidate lists configured by the encoder and the decoder respectively may be identical to each other. For example, the encoder and the decoder can configure the merge candidate list of the current block based on the position of the current block within the current picture. In the present disclosure, the position of a specific block represents the relative position of the top-left sample of the specific block within the picture including the specific block.

[0089] FIG. 8 is a diagram showing a method by which a motion vector of a current block is signaled according to an embodiment of the present disclosure. According to an embodiment of the present disclosure, the motion vector of the current block may be derived from a motion vector predictor (MVP) of the current block. According to an embodiment, the motion vector predictor referred to for deriving the motion vector of the current block may be obtained using a motion vector predictor (MVP) candidate list. The MVP candidate list may include a preset number of MVP candidates (Candidate1, Candidate2,..., Candidate N).

[0090] According to an embodiment, the MVP candidate list may include at least one of a spatial candidate or a temporal candidate. The spatial candidate may be a set of motion information used for prediction of neighboring blocks within a certain range from the current block in the current picture. The spatial candidate may be configured based on available neighboring blocks among the neighboring blocks of the current block. Also, the temporal candidate may be a set of motion information used for prediction of blocks in the current picture and other pictures. For example, the temporal candidate may be configured based on a specific block corresponding to the position of the current block in a specific reference picture. At this time, the position of the specific block represents the position of the top-left sample of the specific block in the reference picture. According to a further embodiment, the MVP candidate list may include a zero motion vector. According to a further embodiment, a rounding process may be performed on the MVP candidates included in the MVP candidate list of the current block. At this time, the resolution of the motion vector difference value of the current block, which will be described later, may be used. For example, each of the MVP candidates of the current block may be rounded based on the resolution of the motion vector difference value of the current block.

[0091] In the present disclosure, the MVP candidate list can include an advanced temporal motion vector prediction (ATMVP) list, a subblock-based temporal motion vector prediction (SbTMVP) list, a merge candidate list for merge inter prediction, a control point motion vector candidate list for affine motion compensation, a subblock-based temporal motion vector prediction (STMVP) list for subblock-based motion compensation, and combinations thereof.

[0092] According to one embodiment, the encoder 810 and the decoder 820 can construct an MVP candidate list for motion compensation of a current block. For example, among the samples restored prior to the current block, there may be candidates corresponding to samples that are predicted based on a motion information set that is the same as or similar to the motion information set of the current block. The encoder 810 and the decoder 820 can construct the MVP candidate list of the current block based on the corresponding plurality of candidate blocks. At this time, the encoder 810 and the decoder 820 can construct the MVP candidate list according to a rule defined in advance between the encoder 810 and the decoder 820. That is, the MVP candidate lists constructed in the encoder 810 and the decoder 820 respectively may be the same as each other.

[0093] Also, the predefined rules may vary depending on the prediction mode of the current block. For example, when the prediction mode of the current block is the affine prediction mode based on the affine model, the encoder and the decoder can construct the MVP candidate list of the current block using the first method based on the affine model. The first method may be a method of obtaining the control point motion vector candidate list. On the other hand, when the prediction mode of the current block is the general inter prediction mode not based on the affine model, the encoder and the decoder can construct the MVP candidate list of the current block using the second method not based on the affine model. At this time, the first method and the second method may be different from each other.

[0094] The decoder 820 can derive the motion vector of the current block based on any one of at least one MVP candidate included in the MVP candidate list of the current block. For example, the encoder 810 can signal an MVP index indicating a motion vector predictor referred to for deriving the motion vector of the current block. Signaling can mean that the encoder generates the signal as a bitstream and the decoder 820 parses it from the bitstream. The MVP index can include a merge index for the merge candidate list. The decoder 820 can obtain the motion vector predictor of the current block based on the signaled MVP index. The decoder 820 can derive the motion vector of the current block using the motion vector predictor. According to one embodiment, the decoder 820 can use the motion vector predictor obtained from the MVP candidate list as the motion vector of the current block without another motion vector difference value. For example, the decoder 820 can select a motion vector from the merge list based on the merge index.

[0095] The decoder 820 can restore the current block based on the motion vector of the current block. An inter prediction mode in which the motion vector predictor obtained from the MVP candidate list is used as the motion vector of the current block without another motion vector difference value can be called the merge mode.

[0096] According to other embodiments, the decoder 820 can obtain another motion vector difference value for the motion vector of the current block. The decoder 820 can obtain the motion vector of the current block by adding the motion vector predictor obtained from the MVP candidate list and the motion vector difference value of the current block. In this case, the encoder 810 can signal a motion vector difference value (MV difference) representing the difference between the motion vector of the current block and the motion vector predictor. The method by which the motion vector difference value is signaled will be specifically described with reference to FIG. 9. The decoder 820 can obtain the motion vector of the current block based on the motion vector difference value (MV difference). The decoder 820 can restore the current block based on the motion vector of the current block.

[0097] Furthermore, a reference picture index for motion compensation of the current block may be signaled. The encoder 810 can signal a reference picture index indicating a reference picture including the reference block. The decoder 820 can obtain the POC of the reference picture referred to for restoring the current block based on the signaled reference picture index. At this time, the POC of the reference picture may be different from the POC of the reference picture corresponding to the MVP referred to for deriving the motion vector of the current block. In this case, the decoder 820 can perform motion vector scaling. That is, the decoder 820 can scale the MVP to obtain an MVP'. At this time, the motion vector scaling may be performed based on the POC of the current picture, the POC of the signaled reference picture of the current block, and the POC of the reference picture corresponding to the MVP. Also, the decoder 820 can use the MVP' as the motion vector predictor of the current block.

[0098] As described above, the motion vector of the current block can be obtained by adding the motion vector predictor of the current block and the motion vector difference value. At this time, the motion vector difference value may be signaled from the encoder. The encoder can encode the motion vector difference value to generate and signal information indicating the motion vector difference value. Hereinafter, a method of signaling the motion vector difference value according to an embodiment of the present disclosure will be described.

[0099] FIG. 9 is a diagram showing a method of signaling the motion vector difference value of the current block according to an embodiment of the present disclosure. According to an embodiment, the information indicating the motion vector difference value can include at least one of the absolute value information of the motion vector difference value or the sign information of the motion vector difference value. The absolute value and sign of the motion vector difference value may be encoded separately.

[0100] According to an embodiment, the absolute value of the motion vector difference value may not be signaled by the value itself. The encoder can reduce the size of the signaled value by using at least one flag indicating the characteristic of the absolute value of the motion vector difference value. The decoder can derive the absolute value of the motion vector difference value from the signaled value by using at least one flag.

[0101] For example, at least one flag can include a first flag indicating whether the absolute value of the motion vector difference value is greater than N. At this time, N may be an integer. When the size of the absolute value of the motion vector difference value is greater than N, the value (absolute value of the motion vector difference value - N) may be signaled together with the activated first flag. At this time, the activated flag can indicate that the size of the absolute value of the motion vector difference value is greater than N. The decoder can obtain the absolute value of the motion vector difference value based on the activated first flag and the signaled value.

[0102] Referring to FIG. 9, a second flag (abs_mvd_greater0_flag) indicating whether the absolute value of the motion vector difference value is greater than '0' may be signaled. When the second flag (abs_mvd_greater0_flag[]) indicates that the absolute value of the motion vector difference value is not greater than '0', the absolute value of the motion vector difference value may be '0'. Also, when the second flag (abs_mvd_greater0_flag) indicates that the absolute value of the motion vector difference value is greater than '0', the decoder can obtain the absolute value of the motion vector difference value by using other information about the motion vector difference value.

[0103] According to one embodiment, a third flag (abs_mvd_greater1_flag) indicating whether the absolute value of the motion vector difference value is greater than '1' may be signaled. When the third flag (abs_mvd_greater1_flag) indicates that the absolute value of the motion vector difference value is not greater than '1', the decoder can determine that the absolute value of the motion vector difference value is '1'.

[0104] Conversely, when the third flag (abs_mvd_greater1_flag) indicates that the absolute value of the motion vector difference value is greater than '1', the decoder can obtain the absolute value of the motion vector difference value using further other information for the motion vector difference value. For example, the (absolute value of the motion vector difference value - 2) value (abs_mvd_minus2) may be signaled. This is because when the absolute value of the motion vector difference value is greater than '1', the absolute value of the motion vector difference value can be a value of 2 or more.

[0105] As described above, the absolute value of the motion vector difference value of the current block may be transformed into at least one flag. For example, the transformed absolute value of the motion vector difference value can indicate (absolute value of the motion vector difference value - N) according to the size of the motion vector difference value. According to one embodiment, the transformed absolute value of the motion vector difference value may be signaled with at least one bit. At this time, the number of bits signaled to indicate the transformed absolute value of the motion vector difference value may be variable. The encoder can encode the transformed absolute value of the motion vector difference value using a variable-length binary method. For example, as the variable-length binary method, the encoder can use at least one of truncated unary binary, unary binary, truncated rice, or exp-Golomb binary.

[0106] Also, the sign of the motion vector difference value may be signaled by a sign flag (mvd_sign_flag). On the other hand, the sign of the motion vector difference value may be implicitly signaled by sign-bit-hiding.

[0107] Also, in FIG. 9, [0] and [1] can indicate component indexes. For example, they can indicate the x-component and y-component.

[0108] FIG. 10 is a diagram showing adaptive motion vector resolution signaling according to an embodiment of the present disclosure.

[0109] According to an embodiment of the present disclosure, the resolution indicating the motion vector or the motion vector difference value can be various. In other words, the resolution at which the motion vector or the motion vector difference value is encoded can be various. For example, the resolution can be indicated based on pixels (pixel, pel). For example, the motion vector or the motion vector difference value can be signaled in units such as 1 / 4 (quarter), 1 / 2 (half), 1 (integer), 2, and 4 pixels. For example, when wanting to indicate 16, if in 1 / 4 unit, it is encoded as 64 (1 / 4 * 64 = 16), if in 1 unit, it is encoded as 16 (1 * 16 = 16), and if in 4 units, it can be encoded as 4 (4 * 4 = 16). That is, the value can be determined as follows.

[0110] valueDetermined = resolution * valuePerResolution

[0111] Here, valueDetermined may be a value to be transmitted, which may be a motion vector or a motion vector difference value in this embodiment. Also, valuePerResolution may be a value indicated in units of [ / resolution].

[0112] At this time, when the value signaled by the motion vector or the motion vector difference value is not divisible by the resolution, an inaccurate value may be sent instead of the motion vector or the motion vector difference value with the best prediction performance by means of rounding or the like. Using a high resolution can reduce inaccuracy, but more bits may be used because the value to be coded is large. Using a low resolution may increase inaccuracy, but fewer bits can be used because the value to be coded is small.

[0113] Also, it is possible to individually set the resolution in units such as blocks, CUs, slices, etc. Therefore, the resolution can be adaptively applied according to the unit.

[0114] The resolution may be signaled from the encoder to the decoder. At this time, the signaling for the resolution may be the binarized signaling in the aforementioned variable length. In such a case, signaling overhead can be reduced by signaling with the index corresponding to the smallest value (the value at the front).

[0115] As an example, it can be matched to the signaling index in the order from high resolution (fine-grained signaling) to low resolution.

[0116] Figure 10 shows the signaling for three resolutions. In such a case, the three signalings can be 0, 10, 11, and each of the three signalings can correspond to Resolution 1, Resolution 2, and Resolution 3. One bit is required to signal Resolution 1, and two bits are required to signal the remaining resolutions. Therefore, there is less signaling overhead when signaling Resolution 1. In the example of Figure 10, Resolution 1, Resolution 2, and Resolution 3 are 1 / 4, 1, and 4 pels respectively.

[0117] In the following invention, the motion vector resolution can mean the resolution of the motion vector difference value.

[0118] Figure 11 is a diagram showing the inter prediction related syntax according to an embodiment of the present disclosure.

[0119] According to an embodiment of the present disclosure, the inter prediction method can include a skip mode, a merge mode, an inter mode, etc. According to an embodiment, in the skip mode, a residual signal may not be transmitted. Also, an MV determination method such as the merge mode can be used in the skip mode. Whether the skip mode is used or not may be determined by a skip flag. Referring to Figure 11, whether the skip mode is used or not may be determined by the cu_skip_flag value.

[0120] According to one embodiment, in the merge mode, it may not be necessary to use the motion vector difference value. The motion vector can be determined based on the motion vector index. Whether the merge mode is used or not may be determined by a merge flag. Referring to FIG. 11, whether the merge mode is used or not may be determined by the merge_flag value. Also, it is possible to use the merge mode when the skip mode is not used.

[0121] It is possible to selectively use one or more types of candidate lists in the skip mode or the merge mode. For example, it is possible to use a merge candidate or a sub-block merge candidate. Also, the merge candidate can include a spatial neighboring candidate, a temporal candidate, etc. Also, the merge candidate can include a candidate that uses a motion vector for the entire current block (CU). That is, the motion vectors of each sub-block belonging to the current block can include the same candidate. Also, the sub-block merge candidate can include a sub-block based temporal MV, an affine merge candidate, etc. Also, the sub-block merge candidate can include a candidate that can use different motion vectors for each sub-block of the current block (CU). The affine merge candidate may be a method created by a method of determining the control point motion vector of affine motion prediction without using the motion vector difference value. Also, the sub-block merge candidate can include a method of determining a motion vector in units of sub-blocks in the current block. For example, the sub-block merge candidate can include a planar MV, a regression based MV, STMVP, etc. in addition to the aforementioned sub-block based temporal MV and affine merge candidate.

[0122] According to one embodiment, in the inter mode, motion vector difference values can be used. A motion vector predictor can be determined based on a motion vector index, and a motion vector can be determined based on the motion vector predictor and the motion vector difference value. Whether the inter mode is used or not may be determined by whether other modes are used or not. As yet another embodiment, whether the inter mode is used or not may be determined by a flag. FIG. 11 shows an example in which the inter mode is used when skip mode and merge mode, which are other modes, are not used.

[0123] The inter mode can include an AMVP mode, an affine inter mode, etc. The inter mode may be a mode that determines a motion vector based on a motion vector predictor and a motion vector difference value. The affine inter mode may be a method that uses a motion vector difference value when determining a control point motion vector for affine motion prediction.

[0124] Referring to FIG. 11, after it is determined to be in skip mode or merge mode, it is possible to determine whether to use sub-block merge candidates or merge candidates. For example, a merge_subblock_flag indicating whether to use sub-block merge candidates when specific conditions are met can be parsed. Further, the specific conditions may be conditions related to the block size. For example, they may be conditions regarding width, height, area, etc., and these can be used in combination. Referring to FIG. 11, for example, it may be a condition when the width and height of the current block (CU) are greater than or equal to specific values. When parsing the merge_subblock_flag, its value can be inferred as 0. It may be that sub-block merge candidates are used if merge_subblock_flag is 1, and merge candidates are used if it is 0. When using sub-block merge candidates, a merge_subblock_idx, which is a candidate index, can be parsed, and when using merge candidates, a merge_idx, which is a candidate index, can be parsed. At this time, if the maximum number of candidate lists is 1, it may not be necessary to parse. If merge_subblock_idx or merge_idx is not parsed, it can be inferred as 0.

[0125] FIG. 11 shows the coding_unit function, but the content related to intra prediction may be omitted, and FIG. 11 may show the case when inter prediction is determined.

[0126] FIG. 12 is a diagram showing multi-hypothesis prediction according to an embodiment of the present disclosure.

[0127] As described above, encoding and decoding can be performed based on a prediction block. According to an embodiment of the present disclosure, a prediction block can be generated based on a number of predictions. This can be called multi-hypothesis (MH) prediction. The prediction can mean a block generated by a certain prediction method. Also, the prediction methods in the number of predictions can include methods such as intra prediction and inter prediction. Or, the prediction methods in the number of predictions may further be subdivided to mean a merge mode, an AMVP mode, a specific mode of intra prediction, etc.

[0128] Also, as a method of generating a prediction block based on a number of predictions, it is possible to perform a weighted sum of the number of predictions.

[0129] Or, the maximum number of the number of predictions may already be set. For example, the maximum number of the number of predictions may be 2. Therefore, in the case of uni prediction, from 2 predictions, and in the case of bi prediction, 2 (when using the number of predictions only for predictions from one reference list) or 4 (when using the number of predictions for predictions from two reference lists) predictions can be used to generate a prediction block.

[0130] Alternatively, prediction modes that can be used in multiple predictions may already be set. Alternatively, combinations of prediction modes that can be used in multiple predictions may already be set. For example, it is possible to use predictions generated by inter prediction and intra prediction. At this time, it is possible to use only some modes of inter prediction or intra prediction for multi-hypothesis prediction. For example, it is possible to use only the merge mode in inter prediction for multi-hypothesis prediction. Alternatively, it is possible to use only merge modes other than sub-block merge in inter prediction for multi-hypothesis prediction. Alternatively, it is possible to use only specific intra modes in intra prediction for multi-hypothesis prediction. For example, it is possible to use only modes including planar, DC, vertical, and horizontal modes in intra prediction for multi-hypothesis prediction.

[0131] Therefore, for example, it is possible to generate a prediction block based on prediction to the merge mode and intra prediction. At this time, for intra prediction, it is possible to allow only planar, DC, vertical, and horizontal modes.

[0132] Referring to FIG. 12, a prediction block is generated based on prediction 1 and prediction 2. At this time, the prediction block is generated by the weighted sum of prediction 1 and prediction 2, and at this time, the weighting values of prediction 1 and prediction 2 are w1 and w2, respectively.

[0133] Also, according to an embodiment of the present disclosure, when generating a prediction block based on multiple predictions, the weighting values of the multiple predictions may be based on the position within the block. Alternatively, the weighting values of the multiple predictions may be based on the mode for generating the prediction.

[0134] For example, when any one of the modes for generating a prediction is an intra prediction, a weighting value can be determined based on the prediction mode. For example, when any one of the modes for generating a prediction is an intra prediction and is a directional mode, the weighting value of a position far from the reference sample can be increased. More specifically, when any one of the modes for generating a prediction is an intra prediction and is a directional mode, and the other modes for generating predictions are inter predictions, the weighting value of a prediction generated based on the intra prediction on the far side of the reference sample can be increased. This is because in the case of inter prediction, motion compensation can be performed based on spatial neighboring candidates, and in that case, there is a probability that the motion of the spatial neighboring blocks referred to for the current block and MC is the same or similar, and there is a probability that the prediction of the region including the object with motion and the prediction near the spatial neighboring blocks are more accurate than other parts. Then, there may be more residual signals remaining near the opposite side of the spatial neighboring blocks compared to other parts, but this can be offset by using intra prediction in multi-hypothesis prediction. Since the reference sample position of intra prediction can be near the spatial neighboring candidates of inter prediction, the weighting value of the far side can be increased.

[0135] As yet another example, when any one of the modes for generating a prediction is an intra prediction and is a directional mode, the weighting value of a position close to the reference sample can be increased. More specifically, when any one of the modes for generating a prediction is an intra prediction and is a directional mode, and the other modes for generating predictions are inter predictions, the weighting value of a prediction generated based on the intra prediction on the near side of the reference sample can be increased. This is because in intra prediction, high prediction accuracy can be obtained near the reference sample.

[0136] As yet another example, when any one of the modes for generating a prediction is an intra prediction and is not a directional mode (for example, in the case of planar or DC mode), the weighting value can be constant regardless of the position in the block.

[0137] Also, in multiple hypothesis prediction, the weight value for prediction 2 can be determined based on the weight value for prediction 1.

[0138] Next, an equation showing an example of determining prediction samples pbSamples based on a number of predictions is presented.

[0139] pbSamples[ x ][ y ] = Clip3( 0, ( 1 << bitDepth ) - 1, ( w * predSamples [ x ][ y ] + (8- w) * predSamplesIntra [ x ][ y ] ) >> 3 )

[0140] In the above equation, x and y can indicate the coordinates of samples within the block and may be in the following ranges: x = 0..nCbW-1 and y = 0..nCbH-1. Also, nCbW and nCbH may be the width and height of the current block, respectively. Also, predSamples may be a block / sample generated by inter prediction, and predSamplesIntra may be a block / sample generated by intra prediction. Also, the weight value w may be determined as follows.

[0141] - If predModeIntra is INTRA_PLANAR or INTRA_DC or nCbW < 4 or nCbH < 4 or cIdx > 0, w is set equal to 4.

[0142] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and y < (nCbH / 4), w is set equal to 6.

[0143] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and (nCbH / 4) <= y < (nCbH / 2), w is set equal to 5.

[0144] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and (nCbH / 2) <= y < (3 * nCbH / 4), w is set equal to 4.

[0145] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and (3 * nCbH / 4) <= y < nCbH, w is set equal to 3.

[0146] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and x < (nCbW / 4), w is set equal to 6.

[0147] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and (nCbW / 4) <= x < (nCbW / 2), w is set equal to 5.

[0148] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and (nCbW / 2) <= x < (3 * nCbW / 4), w is set equal to 4.

[0149] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and (3 * nCbW / 4) <= x < nCbW, w is set equal to 3.

[0150] FIG. 13 is a diagram showing multi-hypothesis prediction related syntax according to an embodiment of the present disclosure.

[0151] Referring to FIG. 13, mh_intra_flag may be a flag indicating whether multiple hypothesis prediction is used. According to one embodiment, multiple hypothesis prediction can be used only when mh_conditions are satisfied, and when mh_conditions are not satisfied, it is possible to infer 0 without parsing mh_intra_flag. For example, mh_conditions can include a condition regarding the block size. Also, mh_conditions can include a condition regarding whether to use a specific mode. For example, when merge_flag is 1 and sub-block_merge_flag is 0, it is possible to parse mh_intra_flag.

[0152] Also, according to one embodiment of the present disclosure, in order to determine the mode in multiple hypothesis prediction, candidate modes can be divided into multiple lists and it is possible to signal which list to use. Referring to FIG. 13, mh_intra_luma_mpm_flag may be a flag indicating which list among multiple lists to use. If mh_intra_luma_mpm_flag does not exist, it is possible to infer 1. Also, as one embodiment of the present disclosure, the multiple lists may be an MPM list and a non-MPM list.

[0153] Also, according to one embodiment of the present disclosure, an index can be signaled to indicate which index candidate in which list among the multiple lists is used. Referring to FIG. 13, mh_intra_luma_mpm_idx may be that index. Also, as one embodiment, the index can be signaled only when a specific list is selected. Referring to FIG. 13, mh_intra_luma_mpm_idx can be parsed only when any list is determined by mh_intra_luma_mpm_flag.

[0154] For example, as in the embodiment described with reference to FIG. 12, multiple hypothesis prediction can be performed based on the prediction generated by inter prediction and the prediction generated by intra prediction. Also, multiple hypothesis prediction can be performed only when it is signaled to use inter prediction. Or, multiple hypothesis prediction can be performed only when it is signaled to use a specific mode of inter prediction such as the merge mode. In that case, signaling for inter prediction does not need to be performed separately. Also, as an example, when generating a prediction by intra prediction, there may be a total of four candidate modes. And the total of four candidate modes can be divided into three in List1 and one in List2. At this time, when List2 is selected, it is not necessary to signal an index. Also, when List1 is selected, an index can be signaled. However, since there are three candidates included in List1, it is signaled by variable length coding, and one or two bits may be required for index signaling.

[0155] FIG. 14 is a diagram showing multi-hypothesis prediction related syntax according to an embodiment of the present disclosure.

[0156] As described with reference to FIG. 13, there may be signaling indicating which list to use, and in FIGS. 13 to 14, mh_intra_luma_mpm_flag may correspond to it.

[0157] According to an embodiment of the present disclosure, signaling indicating which list to use may be explicitly signaled only in a specific case. Also, when not explicitly signaled, the value of the signaling can be inferred according to the method already set. Referring to FIG. 14, when the condition of mh_mpm_infer_condition is satisfied, there is no explicit signaling, and when the condition is not satisfied, there may be explicit signaling. Also, when the mh_mpm_infer_condition condition is satisfied and mh_intra_luma_mpm_flag does not exist, in that case, it can be inferred as 1. That is, it can be inferred that the MPM list is used.

[0158] FIG. 15 is a diagram showing multi - hypothesis prediction - related syntax according to an embodiment of the present disclosure.

[0159] As described with reference to FIGS. 13 to 14, there may be signaling indicating which list to use, and the value can be inferred when a certain condition is met.

[0160] According to an embodiment of the present disclosure, the condition for inferring the signaling value indicating which list to use may be based on the current block size. For example, it may be based on the width and height of the current block. More specifically, when the larger of the width and height of the current block is greater than n times the smaller one, it is possible to infer the signaling value. For example, n may be 2.

[0161] Referring to FIG. 15, the condition for inferring the signaling value indicating which list to use may be that the larger of the width and height of the current block is greater than twice the smaller one. If the width and height of the current block are cbWidth and cbHeight respectively, the Abs(Log2(cbWidth / cbHeight)) value is 0 when cbWidth and cbHeight are the same, and 1 when the difference is twice. Therefore, when the difference between cbWidth and cbHeight is greater than twice, the Abs(Log2(cbWidth / cbHeight)) value is greater than 1 (2 or more).

[0162] FIG. 16 is a diagram showing a multi - hypothesis prediction mode determination method according to an embodiment of the present disclosure.

[0163] As described with reference to FIGS. 13 to 15, mode determination may be made based on a number of lists. The mode may mean an intra mode that generates a prediction based on intra prediction. Also, the number of lists may include List1 and List2. Referring to FIG. 16, it is possible to determine whether to use List1 with List1_flag. Also, there may be a number of candidates that can belong to List1, and there may be one candidate that can belong to List2. Also, the number of lists may be two lists.

[0164] If List1_flag is inferred, its value can be inferred as using List1. In that case, it is possible to parse List1_index, which is an index indicating which candidate to use in List1. Also, if List1_flag is not inferred, it is possible to parse List1_flag. If List1_flag is 1, it is possible to parse List1_index, and if List1_flag is not 1, it is not necessary to parse the index. Also, when List1_flag is 1, it is possible to determine the mode actually used from among the candidate modes of List1 based on the index. Also, when List1_flag is not 1, it is possible to determine the mode actually used as the candidate mode of List2 without an index. That is, the mode may be determined based on the flag and index in List1, and the mode may be determined based on the flag in List2.

[0165] FIG. 17 is a diagram showing a candidate list generation method according to an embodiment of the present disclosure.

[0166] According to an embodiment of the present disclosure, when variable-length coding an index for determining candidate modes in a list, there may be a method for determining the mode order included in the candidate list to increase coding efficiency. For example, there may be a method for determining the mode order included in List1. At this time, in order to determine the mode order, the modes around the current block can be referred to. Also, List2 can be determined without referring to the modes around the current block. For example, List1 can be generated by referring to the modes around the current block, and modes not included in List1 can be included in List2.

[0167] Also, List1 may be the MPM mode and List2 may be the non-MPM mode. Also, the total number of candidate modes may be four, with three modes included in List1 and one mode included in List2.

[0168] Referring to FIG. 17, there may be a List1_flag which is a signaling indicating whether List1 is used. If List1 is used, List1 can be generated and a mode in List1 can be selected. At this time, the generation of List1 and the confirmation of whether List1 is used can be in any order. However, in the situation where List1 is used, List1 can be generated before or after confirming whether List1 is used. Also, when List1 is used, the process of generating List2 does not need to be performed. If List1 is not used, List2 can be generated and a mode in List2 can be selected. At this time, it is possible to generate List1 to generate List2. And among the candidate modes, candidates not included in List1 can be included in List2.

[0169] Also, according to an embodiment of the present disclosure, the method for generating List1 may be the same regardless of whether List1 is used (List1_flag value), whether to infer whether List1 is used, etc.

[0170] At this time, the List signaling and the mode signaling can follow the embodiments described above with reference to FIG. 16 and the like.

[0171] Next, a number of list generation methods described with reference to FIGS. 16 to 17 will be further described. The number of lists can include List1 and List2. Further, the number of lists may be a list used particularly in the multi-hypothesis prediction process.

[0172] According to an embodiment of the present disclosure, a number of lists can be generated by referring to the mode around the current block. Further, intra prediction can be performed using the mode selected from the number of lists, and the intra prediction can be combined with inter prediction (multi-hypothesis prediction) and used as a prediction block. As an example, the modes (candidate modes) that may be included in the number of lists may be a planar mode, a DC mode, a vertical mode, and a horizontal mode of the intra prediction method. Further, the vertical mode may be the mode of index 50 in FIG. 6, and the horizontal mode may be the mode of index 18 in FIG. 6. Further, the planar mode and the DC mode may be indices 0 and 1, respectively.

[0173] According to an embodiment of the present disclosure, candModeList can be generated by referring to the mode around the current block. Further, candModeList may be List1 described in the previous embodiment. As an example, there may be candIntraPredModeX which is a mode around the current block or a mode based on the mode around the current block. Here, X may be a character for indicating correspondence to a specific position around the current block, such as A or B.

[0174] As an example, a candModeList can be generated based on whether a number of candIntraPredModeX values match. For example, candIntraPredModeX may exist for two positions, which can be represented as candIntraPredModeA and candIntraPredModeB. If candIntraPredModeA and candIntraPredModeB are the same, the candModeList can include a planar mode and a DC mode.

[0175] If candIntraPredModeA and candIntraPredModeB are the same and their value indicates the planar mode or the DC mode, the mode indicated by candIntraPredModeA and candIntraPredModeB can be added to candModeList. Also, in this case, among the planar mode and the DC mode, the mode not indicated by candIntraPredModeA and candIntraPredModeB can be added to candModeList. Also, in this case, a mode that has already been set and is not the planar mode or the DC mode can be added to candModeList. As an example, in this case, the order of the planar mode, the DC mode, and the already set mode within candModeList may already be set. For example, it may be in the order of planar, DC, and the already set mode. That is, candModeList[0] = planar mode, candModeList[1] = DC mode, and candModeList[2] = the already set mode may be possible. Also, the already set mode may be the vertical mode. As yet another example, in this case, among the planar mode, the DC mode, and the already set mode, the mode indicated by candIntraPredModeA and candIntraPredModeB is located at the head of candModeList, among the planar mode and the DC mode, the mode not indicated by candIntraPredModeA and candIntraPredModeB is located next in candModeList, and the already set mode can be located after that.

[0176] Also, if candIntraPredModeA and candIntraPredModeB are the same and their values do not indicate the planar mode or the DC mode, the modes indicated by candIntraPredModeA and candIntraPredModeB can be added to candModeList. Also, the planar mode and the DC mode may be added to candModeList. Also, in this case, the order of the modes indicated by candIntraPredModeA and candIntraPredModeB and the planar mode and the DC mode within candModeList may already be set. Also, the already set order may be in the order of the modes indicated by candIntraPredModeA and candIntraPredModeB, the planar mode, and the DC mode. That is, candModeList[0] = candIntraPredModeA, candModeList[1] = planar mode, and candModeList[2] = DC mode may be possible.

[0177] Also, if candIntraPredModeA and candIntraPredModeB are different, both candIntraPredModeA and candIntraPredModeB can be added to candModeList. Also, candIntraPredModeA and candIntraPredModeB may be included in candModeList according to a specific order. For example, they may be included in candModeList in the order of candIntraPredModeA, candIntraPredModeB. Also, there may already be a set order among the candidate modes, and among the modes following the already set order, modes other than candIntraPredModeA and candIntraPredModeB can be added to candModeList. Also, the modes other than candIntraPredModeA and candIntraPredModeB may be positioned after candIntraPredModeA and candIntraPredModeB in candModeList. Also, the already set order may be the planar mode, the DC mode, the vertical mode. Or, the already set order may be the planar mode, the DC mode, the vertical mode, the horizontal mode. That is, candModeList[0]=candIntraPredModeA, candModeList[1]=candIntraPredModeB may be the case, and candModeList[2] may be the mode that is not candIntraPredModeA and not candIntraPredModeB and is the first among the planar mode, the DC mode, the vertical mode.

[0178] Also, among the candidate modes, a mode not included in candModeList may become candIntraPredModeC. Also, candIntraPredModeC may be included in List2. Also, this makes it possible to determine candIntraPredModeC when the signaling indicating whether or not the aforementioned List1 is used indicates not to use it.

[0179] Also, when using List1, the mode can be determined by the index in candModeList, and when not using List1, the mode of List2 can be used.

[0180] Also, as described, after generating candModeList, a process of modifying candModeList may be added. For example, depending on the current block size condition, the process of modifying may or may not be further performed. For example, the current block size condition may be based on the width and height of the current block. For example, when the larger one of the width and height of the current block is greater than n times the smaller one, the process of modifying can be further performed. n may be 2.

[0181] Also, the process of modifying may be a process of substituting one mode in candModeList with another mode when any mode is included in candModeList. For example, when the vertical mode is included in candModeList, the horizontal mode can be put into candModeList instead of the vertical mode. Or, when the vertical mode is included in candModeList, candIntraPredModeC can be put into candModeList instead of the vertical mode. However, when generating candModeList as described above, the planar mode and the DC mode may always be included in candModeList. In this case, candIntraPredModeC may be the horizontal mode. Also, using such a process of modifying may be when the height of the current block is greater than n times the width. For example, n may be 2. This is because when the height is greater than the width, the lower part of the block is far from the reference sample for intra prediction, and the accuracy of the vertical mode may be low. Or, using such a process of modifying may be when it is inferred that List1 is used.

[0182] As yet another example of the above-described process of correction, when the horizontal mode is included in the candModeList, the vertical mode can be put into the candModeList instead of the horizontal mode. Or, when the horizontal mode is included in the candModeList, the candIntraPredModeC can be put into the candModeList instead of the horizontal mode. However, when generating the above-described candModeList, the planar mode and the DC mode may always be included in the candModeList. In this case, the candIntraPredModeC may be the vertical mode. Also, using such a process of correction may be when the width of the current block is larger than n times the height. For example, n may be 2. This is because when the width is larger than the height, the right side of the block may be far from the reference samples for intra prediction, and thus the accuracy of the horizontal mode may be low. Or, using such a process of correction may be when it is inferred that List1 is used.

[0183] Examples of the above-described list setting method are described again below. In the following, IntraPredModeY may be a mode used for intra prediction in multi-hypothesis prediction. Also, this may be a mode of the luma component. As an example, in multi-hypothesis prediction, the intra prediction mode of the chroma component may follow the mode of the luma component. Also, mh_intra_luma_mpm_flag may be signaling indicating which list to use. That is, for example, it may be the mh_intra_luma_mpm_flag in FIGS. 13 to 15, or the List1_flag in FIGS. 16 to 17. Also, mh_intra_luma_mpm_idx may be an index indicating which candidate in the list to use. That is, for example, it may be the mh_intra_luma_mpm_idx in FIGS. 13 to 15, or the List1_index in FIG. 16. Also, xCb and yCb may be the x and y coordinates of the top-left of the current block. Also, cbWidth and cbHeight may be the width and height of the current block.

[0184] The candModeList[ x ] with x = 0..2 is derived as follows:

[0185] A. If candIntraPredModeB is equal to candIntraPredModeA, the following applies:

[0186] a. If candIntraPredModeA is less than 2 (i.e., equal to INTRA_PLANAR or INTRA_DC), candModeList[ x ] with x = 0..2 is derived as follows:

[0187] candModeList

[0000] = INTRA_PLANAR

[0188] candModeList

[0001] = INTRA_DC

[0189] candModeList

[0002] = INTRA_ANGULAR50

[0190] b. Otherwise, candModeList[ x ] with x = 0..2 is derived as follows:

[0191] candModeList

[0000] = candIntraPredModeA

[0192] candModeList

[0001] = INTRA_PLANAR

[0193] candModeList

[0002] = INTRA_DC

[0194] B. Otherwise (candIntraPredModeB is not equal to candIntraPredModeA), the following applies:

[0195] a. candModeList

[0000] and candModeList

[0001] are derived as follows:

[0196] candModeList

[0000] = candIntraPredModeA

[0197] candModeList

[0001] = candIntraPredModeB

[0198] b. If neither of candModeList

[0000] and candModeList

[0001] is equal to INTRA_PLANAR, candModeList

[0002] is set equal to INTRA_PLANAR,

[0199] c. Otherwise, if neither of candModeList

[0000] and candModeList

[0001] is equal to INTRA_DC, candModeList

[0002] is set equal to INTRA_DC,

[0200] d. Otherwise, candModeList

[0002] is set equal to INTRA_ANGULAR50.

[0201] IntraPredModeY[ xCb ][ yCb ] is derived by applying the following procedure:

[0202] A. If mh_intra_luma_mpm_flag[ xCb ][ yCb ] is equal to 1, the IntraPredModeY[ xCb ][ yCb ] is set equal to candModeList[ intra_luma_mpm_idx[ xCb ][ yCb ] ].

[0203] B. Otherwise, IntraPredModeY[ xCb ][ yCb ] is set to equal to candIntraPredModeC, derived by applying the following steps:

[0204] a. If neither of candModeList[ x ], x = 0..2 is equal to INTRA_PLANAR, candIntraPredModeC is set equal to INTRA_PLANAR,

[0205] b. Otherwise, if neither of candModeList[ x ], x = 0..2 is equal to INTRA_DC, candIntraPredModeC is set equal to INTRA_DC,

[0206] c. Otherwise, if neither of candModeList[ x ], x = 0..2 is equal to INTRA_ANGULAR50, candIntraPredModeC is set equal to INTRA_ANGULAR50,

[0207] d. Otherwise, if neither of candModeList[ x ], x = 0..2 is equal to INTRA_ANGULAR18. candIntraPredModeC is set equal to INTRA_ANGULAR18,

[0208] The variable IntraPredModeY[ x ][ y ] with x = xCb..xCb + cbWidth - 1 and y = yCb..yCb + cbHeight - 1 is set to be equal to IntraPredModeY[ xCb ][ yCb ]. One additional setting is when cbHeight is larger than double of cbWidth, mh_intra_luma_mpm_flag[ xCb ][ yCb ] is inferred to be 1 and if candModeList[ x ], x = 0..2 is equal to INTRA_ANGULAR50, candModeList[ x ] is replaced with candIntraPredModeC. Another additional setting is when cbWidth is larger than double of cbHeight, mh_intra_luma_mpm_flag[ xCb ][ yCb ] is inferred to be 1 and if candModeList[ x ], x = 0..2 is equal to INTRA_ANGULAR18, candModeList[ x ] is replaced with candIntraPredModeC.

[0209] FIG. 18 is a diagram showing peripheral positions referred to in multi-hypothesis prediction according to an embodiment of the present disclosure.

[0210] As described above, peripheral positions can be referred to in the process of creating a candidate list for multi-hypothesis prediction. For example, the aforementioned candIntraPredModeX may be required. At this time, the positions of A and B in the periphery of the current block to be referred to may be NbA and NbB shown in FIG. 18. That is, they may be immediately to the left and immediately above the upper left end of the current block. If the position of the upper left end of the current block is Cb as shown in FIG. 18 and its coordinates are (xCb, yCb), NbA may be (xNbA, yNbA) = (xCb - 1, yCb), and NbB may be (xNbB, yNbB) = (xCb, yCb - 1).

[0211] FIG. 19 is a diagram showing a method of referring to peripheral modes according to an embodiment of the present disclosure.

[0212] As described above, peripheral positions can be referred to in the process of creating a candidate list for multi-hypothesis prediction. Also, it is possible to use the peripheral modes as they are or to generate a list using modes based on the peripheral modes. The mode obtained by referring to the peripheral positions may be candIntraPredModeX.

[0213] As an example, when the peripheral positions are not available, candIntraPredModeX may be the already set mode. Cases where the peripheral positions are not available may include when the inter prediction is used, or when the mode determination has not been made in the defined decoding and encoding order.

[0214] Or, when the peripheral positions do not use multi-hypothesis prediction, candIntraPredModeX may be the already set mode.

[0215] Or, when the peripheral positions are above the CTU to which the current block belongs, candIntraPredModeX may be the already set mode. As another example, when the peripheral positions are outside the CTU to which the current block belongs, candIntraPredModeX may be the already set mode.

[0216] Also, as an example, the already set mode may be a DC mode. As yet another example, the already set mode may be a planar mode.

[0217] Also, candIntraPredModeX can be set according to whether the mode at the peripheral position exceeds a threshold angle or whether the index of the mode at the peripheral position exceeds a threshold. For example, when the index of the mode at the peripheral position is larger than the diagonal mode index, candIntraPredModeX can be set to the vertical mode index. Also, when the index of the mode at the peripheral position is less than or equal to the diagonal mode index and it is a direction mode, candIntraPredModeX can be set to the horizontal mode index. The diagonal mode index may be mode 34 in FIG. 6.

[0218] Also, if the mode at the peripheral position is a planar mode or a DC mode, candIntraPredModeX may be set to the planar mode or the DC mode as it is.

[0219] Referring to FIG. 19, mh_intra_flag may be a signaling indicating whether multiple hypothesis prediction is used (has been used). Also, the intra prediction mode used in the neighboring block may be X. Also, the current block can use multiple hypothesis prediction and can generate a candidate list using candIntraPredMode based on the mode of the peripheral blocks. However, since the periphery does not use multiple hypothesis prediction, regardless of the intra prediction mode of the peripheral blocks and regardless of whether the peripheral blocks use intra prediction or not, candIntraPredMode can be set to the DC mode which is the already set mode.

[0220] An example of the above-described peripheral mode reference method is described again below.

[0221] For X being replaced by either A or B, the variables candIntraPredModeX are derived as follows:

[0222] 1. The availability derivation process for a block as specified in Neighbouring blocks availability checking process is invoked with the location ( xCurr, yCurr ) set equal to ( xCb, yCb ) and the neighbouring location ( xNbY, yNbY ) set equal to ( xNbX, yNbX ) as inputs, and the output is assigned to availableX.

[0223] 2. The candidate intra prediction mode candIntraPredModeX is derived as follows:

[0224] A. If one or more of the following conditions are true, candIntraPredModeX is set equal to INTRA_DC.

[0225] a. The variable availableX is equal to FALSE.

[0226] b. mh_intra_flag[ xNbX ][ yNbX ] is not equal to 1.

[0227] c. X is equal to B and yCb - 1 is less than ( ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY ).

[0228] B. Otherwise, if IntraPredModeY[ xNbX ][ yNbX ] > INTRA_ANGULAR34, candIntraPredModeX is set equal to INTRA_ANGULAR50.

[0229] C. Otherwise, if IntraPredModeY[ xNbX ][ yNbX ] <= INTRA_ANGULAR34 and IntraPredModeY[ xNbX ][ yNbX ] > INTRA_DC, candIntraPredModeX is set equal to INTRA_ANGULAR18.

[0230] D. Otherwise, candIntraPredModeX is set equal to IntraPredModeY[ xNbX ][ yNbX ].

[0231] In the above-mentioned list setting method, candIntraPredModeX may be determined by the above-mentioned peripheral mode reference method.

[0232] FIG. 20 is a diagram showing a candidate list generation method according to an embodiment of the present disclosure.

[0233] According to the List1 and List2 generation methods described with reference to FIGS. 13 to 17, List1 can be generated by referring to the modes around the current block, and the modes not included in List1 among the candidate modes can be put into List2. Since there is spatial similarity in the picture, those referring to the surrounding modes can be those with higher priorities. That is, List1 may have a higher priority than List2. However, according to the List1 and List2 signaling methods described with reference to FIGS. 13 to 17, since the mode of List1 is used when the signaling for determining the list is not inferred, signals can be sent using flags and indexes, and since the mode of List2 is used, only flags can be used. That is, relatively fewer bits can be used for the signaling of List2. However, it may not be good in terms of coding efficiency that relatively more bits are used for signaling the modes of the list with higher priorities. Therefore, there may be a method of using signaling with relatively fewer bits for the list and modes with higher priorities as in the present disclosure.

[0234] According to an embodiment of the present disclosure, the candidate list generation method may vary depending on whether only List1 can be used. Whether only List1 can be used can indicate whether signaling indicating the list to be used is inferred. For example, when there is List3 that has a candidate mode and is generated by a preset method, List3 can be divided into List1 and List2 and inserted. For example, List3 and its generation method generated by a preset method may be the aforementioned candModeList and its generation method. If signaling indicating the list to be used is inferred, only List1 can be used, and in that case, it can be filled into List1 from the head of List3. Also, if signaling indicating the list to be used is not inferred, List1 or List2 may be used, and in that case, it can be filled into List2 from the head of List3, and the rest can be filled into List1. Also, in this case, when filling List1, it is also possible to fill in the order of List3. That is, candIntraPredModeX can be put into the candidate list by referring to the mode around the current block, but candIntraPredModeX can be put into List1 when signaling indicating the list is inferred, and into List2 when not inferred. The size of List2 may be 1, and in that case, candIntraPredModeA can be put into List1 when signaling indicating the list is inferred, and into List2 when not inferred. candIntraPredModeA may be List3[0] which is the first mode of the aforementioned List3. Therefore, in the present disclosure, in some cases, candIntraPredModeA which is a mode based on the surrounding mode may be directly included in both List1 and List2. On the other hand, in the method described with reference to FIGS. 13 to 17, candIntraPredModeA could only be included in List1. Also, in the present disclosure, the List1 generation method varies depending on whether signaling indicating the list to be used is inferred.

[0235] Referring to FIG. 20, the candidate mode may be a candidate that can be used to generate an intra prediction of multiple hypothesis prediction. The candidate list generation method may vary depending on whether a signaling List1_flag indicating the list to be used is inferred. If it is inferred, it may be inferred that List1 is used, and since only List1 can be used, it can be inserted into List1 from the top of List3. When generating List3, a mode based on the surrounding modes can be placed at the top. Also, if it is not inferred, since both List1 and List2 are used, it can be inserted into List2 with less signaling from the top of List3. And when List1 is necessary, for example, when it is signaled that List1 is used, those other than those included in List2 in List3 can be inserted into List1.

[0236] In the present disclosure, there is a part described using List3, but this may be a conceptual explanation, and it is possible to generate List1 and List2 without actually storing List3.

[0237] According to the following embodiments, the method of creating a candidate list and the method of creating a candidate list described in FIGS. 16 to 17 can be used depending on the case. For example, depending on whether signaling indicating which list to use is inferred, either one of the two methods of creating a candidate list can be selected. Also, this may be the case of multiple hypothesis prediction. Also, the following List1 may contain three modes, and List2 may contain one mode. Also, as described in FIGS. 13 to 16, the mode signaling method can signal the mode of List1 with a flag and an index, and signal the mode of List2 with a flag.

[0238] As an example, when candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is the planar mode or the DC mode, List2[0] = planar mode, List1[0] = DC mode, List1[1] = vertical mode, and List1[2] = horizontal mode may be used.

[0239] As yet another example, when candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is the planar mode or the DC mode, List2[0] = candIntraPredModeA, List1[0] =!candIntraPredModeA, List1[1] = vertical mode, and List1[2] = horizontal mode may be used.

[0240] As an example, when candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is the directional mode, List2[0] = candIntraPredModeA, List1[0] = planar mode, List1[1] = DC mode, and List1[2] = vertical mode may be used.

[0241] As an example, when candIntraPredModeA and candIntraPredModeB are different, List2[0] = candIntraPredModeA and List1[0] = candIntraPredModeB may be used. List1[1] and List1[2] can be filled with modes other than candIntraPredModeA and candIntraPredModeB among the planar mode, DC mode, vertical mode, and horizontal mode.

[0242] FIG. 21 is a diagram showing a candidate list generation method according to an embodiment of the present disclosure.

[0243] In the previous embodiment, a method for determining a mode based on a number of lists was described. In the invention of FIG. 21, the mode can be determined based on one list instead of a number of lists.

[0244] Referring to FIG. 21, as shown in FIG. 21(a), one candidate list including all candidate modes of multi-hypothesis prediction can be generated. Since there is only one candidate list, referring to FIG. 21(b), there is no signaling for selecting a list, and there may be index signaling indicating which mode among the modes of the candidate list is to be used. Therefore, when the mh_intra_flag indicating whether multi-hypothesis prediction is used is 1, the mh_intra_luma_idx which is a candidate index can be parsed.

[0245] According to one embodiment, the method for generating a candidate list for multi-hypothesis prediction may be based on the MPM list generation method in existing intra prediction.

[0246] According to one embodiment, the method for generating a candidate list for multi-hypothesis prediction may be in a form where List1 and List2 are concatenated in the order of List1 and List2 in the List1 and List2 generation methods described in FIG. 17 above.

[0247] That is, if the candidate list for multi-hypothesis prediction is candModeList, the size of candModeList may be 4 in this embodiment. If candIntraPredModeA and candIntraPredModeB are the same and are the planar mode or the DC mode, candModeList may be determined according to the already set order. For example, candModeList[0] = planar mode, candModeList[1] = DC mode, candModeList[2] = vertical mode, candModeList[3] = horizontal mode may be possible.

[0248] As yet another example, if candIntraPredModeA and candIntraPredModeB are the same and are in the planar mode or the DC mode, then candModeList[0]=candIntraPredModeA, candModeList[1]=!candIntraPredModeA, candModeList[2]=vertical mode, and candModeList[3]=horizontal mode may be used.

[0249] If candIntraPredModeA and candIntraPredModeB are the same and are in a directional mode, then candModeList[0]=candIntraPredModeA, candModeList[1]=planar mode, candModeList[2]=DC mode, and candModeList[3] may be a mode other than candIntraPredModeA, the planar mode, and the DC mode.

[0250] If candIntraPredModeA and candIntraPredModeB are different, then candModeList[0]=candIntraPredModeA and candModeList[1]=candIntraPredModeB may be used. Also, candModeList[2] and candModeList[3] may be sequentially filled with modes other than candIntraPredModeA and candIntraPredModeB according to the already set order of candidate modes. The already set order may be the planar mode, the DC mode, the vertical mode, and the horizontal mode.

[0251] According to an embodiment of the present disclosure, candidate lists may vary according to block size conditions. For example, if the larger of the block width and height is greater than n times the remaining part, the candidate list may be shorter. For example, when the width is greater than n times the height, in the candidate list described in FIG. 21, the horizontal mode can be removed from the candidate list, and the next mode can be moved up and filled. Also, when the height is greater than n times the width, in the candidate list described in FIG. 21, the vertical mode can be removed from the candidate list, and the next mode can be moved up and filled. Therefore, when the width is greater than n times the height, the candidate list size may be 3. Also, when the width is greater than n times the height, the size of the candidate list may be smaller or the same as in the case where it is not.

[0252] According to an embodiment of the present disclosure, the candidate index of the embodiment of FIG. 21 may be variable length coded. This may be for increasing the signaling efficiency by filling the modes with relatively high probabilities of use at the head of the list.

[0253] According to still another embodiment, the candidate index of the embodiment of FIG. 21 may be fixed length coded. The number of modes used in multi-hypothesis prediction may be a power of 2. For example, as described above, it can be used from among 4 intra prediction modes. In this case, since no unassignable values occur even with fixed length coding, no unnecessary parts occur in signaling. Also, when fixed length coding is used, the number of cases of list construction may be 1. This is because the number of bits is the same regardless of which index is signaled.

[0254] According to one embodiment, the candidate index may be variable-length coded or fixed-length coded depending on the case. For example, as in the previous embodiment, the candidate list size may vary depending on the case. As one embodiment, the candidate index may be variable-length coded or fixed-length coded depending on the candidate list size. For example, it may be fixed-length coded when the candidate list size is a power of 2, and variable-length coded when it is not a power of 2. That is, according to the previous embodiment, the coding method may vary depending on the block size condition.

[0255] According to an embodiment of the present disclosure, when using the DC mode when using multiple hypothesis prediction, since the weighting values among a large number of predictions are the same for the entire block, it may be the same as adjusting the weighting value of the prediction block. Therefore, the DC mode can be excluded in multiple hypothesis prediction.

[0256] As an example, in multi-hypothesis prediction, it is possible to use only one of the planar mode, vertical mode, and horizontal mode. In this case, as shown in FIG. 21, it is possible to signal multi-hypothesis prediction using one list. Also, variable-length coding can be used for index signaling. As an example, the list can be generated in a fixed order. For example, it may be in the order of planar mode, vertical mode, and horizontal mode. As yet another example, the list can be generated by referring to the mode around the current block. For example, if candIntraPredModeA and candIntraPredModeB are the same, candModeList[0] = candIntraPredModeA may be used. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is the planar mode, candModeList[1] and candModeList[2] can be set according to the already set order. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is not the planar mode, candModeList[1] = planar mode, and candModeList[2] may be a mode other than the planar mode and not candIntraPredModeA. If candIntraPredModeA and candIntraPredModeB are different, candModeList[0] = candIntraPredModeA, candModeList[1] = candIntraPredModeB, and candModeList[2] may be a mode other than candIntraPredModeB and not candIntraPredModeA.

[0257] As yet another example, in multi-hypothesis prediction, it is possible to use only any one of three modes. The three modes may include a planar mode and a DC mode. Further, the three modes may conditionally include either a vertical mode or a horizontal mode. The condition may be a condition related to a block size. For example, depending on whether the width or the height of the block is larger, it is possible to determine whether to include the horizontal mode or the vertical mode. For example, when the width of the block is larger than the height, the vertical mode can be included. When the height of the block is larger than the width, the horizontal mode can be included. When the height and the width of the block are the same, either the vertical mode or the horizontal mode, whichever is the agreed mode, can be included.

[0258] As an example, a list can be generated in a fixed order. For example, it may be in the order of planar mode, DC mode, vertical or horizontal mode. As yet another example, a list can be generated by referring to the mode around the current block. For example, if candIntraPredModeA and candIntraPredModeB are the same, candModeList[0] may be candIntraPredModeA. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is not a direction mode, candModeList[1] and candModeList[2] can be set according to the already set order. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is a direction mode, candModeList[1] = planar mode, candModeList[2] = DC mode may be possible. If candIntraPredModeA and candIntraPredModeB are different, candModeList[0] = candIntraPredModeA, candModeList[1] = candIntraPredModeB, candModeList[2] may be a mode other than candIntraPredModeA and other than candIntraPredModeB.

[0259] According to still other embodiments of the present disclosure, it is possible to use only one of two modes in multi-hypothesis prediction. The two modes may include a planar mode. Also, the two modes may include either a vertical mode or a horizontal mode depending on a condition. The condition may be a condition related to the block size. For example, it is possible to determine whether to include the horizontal mode or the vertical mode depending on which of the width and height of the block is larger. For example, when the width of the block is larger than the height, the vertical mode can be included. When the height of the block is larger than the width, the horizontal mode can be included. When the height and width of the block are the same, the predefined mode of either the vertical mode or the horizontal mode can be included. In this case, a flag indicating which mode to use in multi-hypothesis prediction may be signaled. According to one embodiment, a specific mode can be excluded depending on the block size. For example, when the block size is small, a specific mode can be excluded. For example, when the block size is small, it is possible to use only the planar mode in multi-hypothesis prediction. If a specific mode is excluded, it is possible to omit or reduce the mode signaling and signal it.

[0260] According to still other embodiments of the present disclosure, it is possible to use only one mode in multi-hypothesis prediction. The one mode may be a planar mode. As still other embodiments, the one mode can be determined based on a block size from among a vertical mode and a horizontal mode. For example, it may be determined from among the vertical mode and the horizontal mode depending on which of the width and height of the block is larger. For example, when the width of the block is larger than the height, it may be determined as the vertical mode, and when the height of the block is larger than the width, it may be determined as the horizontal mode. If the width and height of the block are the same, it is possible to determine as the already set mode. If the width and height of the block are the same, it is possible to determine as the already set mode among the horizontal mode or the vertical mode. If the width and height of the block are the same, it is possible to determine as the already set mode among the planar mode or the DC mode.

[0261] Also, according to an embodiment of the present disclosure, there may be a flipping signaling that flips the prediction generated by multi-hypothesis prediction. Thereby, even if one mode is selected in multi-hypothesis prediction, it can have the effect of eliminating the opposite residual by flipping. Also, thereby, it can have the effect of reducing the candidate modes available in multi-hypothesis prediction. More specifically, for example, in the case of using only one mode among the above embodiments, flipping can be used. Thereby, the prediction performance can be improved. The flipping can mean flipping with respect to the x-axis or flipping with respect to the y-axis or flipping with respect to both the x and y axes. As an example, based on the mode selected in multi-hypothesis prediction, the flipping direction can be determined. For example, when the mode selected in multi-hypothesis prediction is a planar mode, it can be determined that it is a flip with respect to both the x and y axes. Also, flipping with respect to both the x and y axes may be due to the block shape. For example, when the block is not square, it can be determined not to flip with respect to both the x and y axes. For example, when the mode selected in multi-hypothesis prediction is a horizontal mode, it can be determined that it is a flip with respect to the x-axis. For example, when the mode selected in multi-hypothesis prediction is a vertical mode, it can be determined that it is a flip with respect to the y-axis. Also, when the mode selected in multi-hypothesis prediction is a DC mode, it can be determined that there is no flipping and no explicit signaling is required.

[0262] Also, in multi-hypothesis prediction, the DC mode may have an effect similar to illumination compensation. Therefore, according to an embodiment of the present disclosure, when using either the DC mode or the illumination compensation method in multi-hypothesis prediction, it is not necessary to use the other.

[0263] Also, multi-hypothesis prediction may have an effect similar to that of generalized bi-prediction (GBi). For example, in multi-hypothesis prediction, the DC mode may have an effect similar to that of GBi. GBi may be a method of adjusting the weighting value between two reference blocks of bi-prediction in block units or CU units. Therefore, according to an embodiment of the present disclosure, either multi-hypothesis prediction (or the DC mode in multi-hypothesis prediction) or the GBi method may be used without using the other. Also, this may be the case when including those that are bi-prediction among the predictions of multi-hypothesis prediction. For example, when the selected merge candidate of multi-hypothesis prediction is bi-prediction, the GBi may not be used. In these embodiments, the relationship between multi-hypothesis prediction and GBi may be limited to when using a specific mode of multi-hypothesis prediction, for example, the DC mode. Or, when GBi-related signaling exists before multi-hypothesis prediction-related signaling, using GBi may not require using multi-hypothesis prediction or a specific mode of multi-hypothesis prediction.

[0264] Not using a certain method may mean not signaling for the certain method and not parsing the related syntax.

[0265] FIG. 22 is a diagram showing peripheral positions referred to in multi-hypothesis prediction according to an embodiment of the present disclosure.

[0266] As described above, the peripheral positions can be referred to in the process of creating the candidate list for multi-hypothesis prediction. For example, the aforementioned candIntraPredModeX may be required. At this time, the positions of A and B around the current block to be referred to may be NbA and NbB shown in FIG. 22. If the upper left end position of the current block is Cb as shown in FIG. 18 and its coordinates are (xCb, yCb), then NbA may be (xNbA, yNbA) = (xCb - 1, yCb + cbHeight - 1), and NbB may be (xNbB, yNbB) = (xCb + cbWidth - 1, yCb - 1). Here, cbWidth and cbHeight may be the width and height of the current block, respectively. Also, in the process of creating the candidate list for multi-hypothesis prediction, the peripheral positions may be the same as the peripheral positions referred to in the generation of the MPM list for intra prediction.

[0267] As yet another example, the peripheral positions referred to in the process of creating the candidate list for multi-hypothesis prediction may be near the center of the left side and the center of the upper side of the current block. For example, NbA and NbB may be (xCb - 1, yCb + cbHeight / 2 - 1), (xCb + cbWidth / 2 - 1, yCb - 1). Or, NbA and NbB may be (xCb - 1, yCb + cbHeight / 2), (xCb + cbWidth / 2, yCb - 1).

[0268] FIG. 23 is a diagram showing a method of referring to peripheral modes according to an embodiment of the present disclosure.

[0269] As described above, the peripheral positions can be referred to in the process of creating the candidate list for multi-hypothesis prediction. However, in the embodiment of FIG. 19, when the peripheral positions do not use multi-hypothesis prediction, candIntraPredModeX is set to a pre-set mode. This is because the mode of the peripheral positions may not directly become candIntraPredModeX when setting candIntraPredModeX.

[0270] Therefore, according to an embodiment of the present disclosure, even if the surrounding position does not use multiple hypothesis prediction, when the mode used by the surrounding position is a mode used for multiple hypothesis prediction, candIntraPredModeX can be set to the mode used by the surrounding position. The mode used for multiple hypothesis prediction may be a planar mode, a DC mode, a vertical mode, or a horizontal mode.

[0271] Alternatively, even if the surrounding position does not use multiple hypothesis prediction, when the mode used by the surrounding position is a specific mode, candIntraPredModeX can be set to the mode used by the surrounding position.

[0272] Alternatively, even if the surrounding position does not use multiple hypothesis prediction, when the mode used by the surrounding position is a vertical mode or a height mode, candIntraPredModeX can be set to the mode used by the surrounding position. Alternatively, even if the surrounding position does not use multiple hypothesis prediction when the surrounding position is above the current block, when the mode used by the surrounding position is a vertical mode, candIntraPredModeX can be set to the mode used by the surrounding position. Also, when the surrounding position is to the left of the current block, even if the surrounding position does not use multiple hypothesis prediction, when the mode used by the surrounding position is a horizontal mode, candIntraPredModeX can be set to the mode used by the surrounding position.

[0273] Referring to FIG. 23, mh_intra_flag may be a signaling indicating whether (or not) to use multiple hypothesis prediction. Also, the intra prediction mode used in the adjacent block may be a horizontal mode. Also, the current block can use multiple hypothesis prediction and a candidate list can be generated using candIntraPredMode based on the mode of the surrounding blocks. However, even if the surrounding does not use multiple hypothesis prediction, since the intra prediction mode of the surrounding blocks is a specific mode, for example, a horizontal mode, candIntraPredMode can be set to the horizontal mode.

[0274] The above-described example of the peripheral mode referencing method is restated below in conjunction with another embodiment of FIG.

[0275] For X being replaced by either A or B, the variables candIntraPredModeX are derived as follows:

[0276] 1. The availability derivation process for a block as specified in Neighboring blocks availability checking process is invoked with the location ( xCurr, yCurr ) set equal to ( xCb, yCb ) and the neighboring location ( xNbY, yNbY ) set equal to ( xNbX, yNbX ) as inputs, and the output is assigned to availableX.

[0277] 2. The candidate intra prediction mode candIntraPredModeX is derived as follows:

[0278] A. If one or more of the following conditions are true, candIntraPredModeX is set equal to INTRA_DC.

[0279] a. The variable availableX is equal to FALSE.

[0280] b. mh_intra_flag[ xNbX ][ yNbX ] is not equal to 1, and IntraPredModeY[ xNbX ][ yNbX ] is neither INTRA_ANGULAR50 nor INTRA_ANGULAR18.

[0281] c. X is equal to B and yCb - 1 is less than ( ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY ).

[0282] B. Otherwise, if IntraPredModeY[ xNbX ][ yNbX ] > INTRA_ANGULAR34, candIntraPredModeX is set equal to INTRA_ANGULAR50.

[0283] C. Otherwise, if IntraPredModeY[ xNbX ][ yNbX ] <= INTRA_ANGULAR34 and IntraPredModeY[ xNbX ][ yNbX ] > INTRA_DC, candIntraPredModeX is set equal to INTRA_ANGULAR18.

[0284] D. Otherwise, candIntraPredModeX is set equal to IntraPredModeY[ xNbX ][ yNbX ].

[0285] Next, it is another example.

[0286] For X being replaced by either A or B, the variables candIntraPredModeX are derived as follows:

[0287] 1. The availability derivation process for a block as specified in Neighbouring blocks availability checking process is invoked with the location ( xCurr, yCurr ) set equal to ( xCb, yCb ) and the neighbouring location ( xNbY, yNbY ) set equal to ( xNbX, yNbX ) as inputs, and the output is assigned to availableX.

[0288] 2. The candidate intra prediction mode candIntraPredModeX is derived as follows:

[0289] A. If one or more of the following conditions are true, candIntraPredModeX is set equal to INTRA_DC.

[0290] a. The variable availableX is equal to FALSE.

[0291] b. mh_intra_flag[ xNbX ][ yNbX ] is not equal to 1, and IntraPredModeY[ xNbX ][ yNbX ] is neither INTRA_PLANAR, INTRA_DC, INTRA_ANGULAR50 nor INTRA_ANGULAR18.

[0292] c. X is equal to B and yCb - 1 is less than ( ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY ).

[0293] B. Otherwise, if IntraPredModeY[ xNbX ][ yNbX ] > INTRA_ANGULAR34, candIntraPredModeX is set equal to INTRA_ANGULAR50.

[0294] C. Otherwise, if IntraPredModeY[ xNbX ][ yNbX ] <= INTRA_ANGULAR34 and IntraPredModeY[ xNbX ][ yNbX ] > INTRA_DC, candIntraPredModeX is set equal to INTRA_ANGULAR18.

[0295] D. Otherwise, candIntraPredModeX is set equal to IntraPredModeY[ xNbX ][ yNbX ].

[0296] In the above-described list setting method, candIntraPredModeX may be determined by the peripheral mode reference method.

[0297] FIG. 24 is a diagram showing the use of peripheral samples according to an embodiment of the present disclosure.

[0298] As described above, when using multi-hypothesis prediction, it is possible to combine intra prediction with other predictions. Therefore, when using multi-hypothesis prediction, samples around the current block can be used as reference samples to generate intra prediction.

[0299] According to an embodiment of the present disclosure, when using multi-hypothesis prediction, a mode using reconstructed samples can be used. Also, when not using multi-hypothesis prediction, the mode using reconstructed samples does not have to be used. The reconstructed samples may be reconstructed samples around the current block.

[0300] An example of a mode using the reconstructed sample can be template matching. A reconstructed sample at a position already set based on a certain block can be defined as a template. Template matching can be an operation of comparing the cost of the template of the block to be compared with the template of the current block and searching for the block with a smaller cost. At this time, the cost can be defined as the sum of the absolute difference values of the templates, the sum of the squares of the difference values, etc. For example, it is possible to search for a block expected to be similar to the current block using template matching between the current block and the block of the reference picture, and based on this, set a motion vector or refine the motion vector.

[0301] Examples of modes using the reconstructed sample can include motion compensation using the reconstructed sample, motion vector refinement, etc.

[0302] In order to use the reconstructed samples around the current block, when decoding the current block, it is necessary to wait for the decoding of the surrounding blocks to be completed. In this case, it may be difficult to perform parallel processing of the current block and the surrounding blocks. Therefore, when not using multi-hypothesis prediction, in order to enable parallel processing, it may not be necessary to use the mode of using the reconstructed samples around the current block. Also, when using multi-hypothesis prediction, since intra prediction can be generated using the reconstructed samples around the current block, other modes of using the reconstructed samples around the current block can also be used.

[0303] Also, according to an embodiment of the present disclosure, even when using multi-hypothesis prediction, it may or may not be necessary to use the reconstructed samples around the current block by the candidate index. As an example, when the candidate index is smaller than the threshold, the reconstructed samples around the current block may be used. If the candidate index is small, the number of candidate index signaling bits may be small and the candidate accuracy may be high, but it is possible to further increase the accuracy by using the reconstructed samples with high coding efficiency. As another example, when the candidate index is larger than the threshold, the reconstructed samples around the current block may be used. If the candidate index is large, the number of candidate index signaling bits may be large and the candidate accuracy may be low, but the accuracy can be complemented by using the reconstructed samples around the current block for candidates with low accuracy.

[0304] According to an embodiment of the present disclosure, when using multi-hypothesis prediction, it is possible to generate an inter prediction by using the reconstructed samples around the current block and combine the inter prediction with the intra prediction of the multi-hypothesis prediction to generate a predicted block.

[0305] Referring to FIG. 24, the mh_intra_flag value, which is a signaling indicating whether the current block uses multi-hypothesis prediction, is 1. Since the current block uses multi-hypothesis prediction, a mode of using the reconstructed samples around the current block can be used.

[0306] FIG. 25 is a diagram showing a transform mode according to an embodiment of the present disclosure.

[0307] According to an embodiment of the present disclosure, there may be a conversion mode that transforms only a sub-part of a block. Such a conversion mode can be referred to as, for example, a sub-block transform (SBT) or a spatially varying transform (SVT). For example, a CU or a PU can be divided into a number of TUs, and only some of the number of TUs can be transformed. For example, only one of the number of TUs can be transformed. Among the number of TUs, the TUs that are not transformed can be assumed to have a residual of 0.

[0308] Referring to FIG. 25, as types of dividing one CU or PU into a number of TUs, there can be two types: SVT-V and SVT-H. SVT-V may be a type in which the height of a number of TUs is the same as the height of the CU or PU, and the width of the number of TUs is different from the width of the CU or PU. SVT-H may be a type in which the height of a number of TUs is different from the height of the CU or PU, and the width of the number of TUs is the same as the width of the CU or PU. As an embodiment, the width and position of the TU to be transformed by SVT-V may be signaled. Also, the height and position of the TU to be transformed by SVT-H may be signaled.

[0309] According to an embodiment, a conversion kernel may already be set according to the SVT type and position, width, or height.

[0310] In this way, there is a mode of transforming only a part of the CU or PU because, after prediction, the residual may mainly exist in a part of the CU or PU.

[0311] That is, SBT has the same concept as the skip mode for a TU. The existing skip mode may be the skip mode for a CU.

[0312] Referring to FIG. 25, it is defined that there are two positions to be transformed, each of which is denoted as A, for each type of SVT-V and SVT-H, and the width or height is defined as 1 / 2 or 1 / 4 of the width or height of the CU. Also, the portions other than the regions denoted as A can have their residues set to 0 values.

[0313] Also, conditions under which SBT can be used may exist. For example, the conditions for SBT can include conditions related to the block size, signaling values indicating whether it is used in a high-level (e.g., sequence, slice, tile, etc.) syntax, and the like.

[0314] FIG. 26 is a diagram showing the relationship between multi-hypothesis prediction and the transformation mode according to an embodiment of the present disclosure.

[0315] According to an embodiment of the present disclosure, there may be a correlation between multi-hypothesis prediction and the transformation mode. For example, depending on whether either one is used or not, it may be determined whether the other is used. Or, depending on one of the modes, it may be determined whether the other mode is used. Or, depending on whether one of the modes is used or not, it may be determined whether the other mode is used.

[0316] As an example, the transformation mode may be SVT described with reference to FIG. 25. That is, depending on whether multi-hypothesis prediction is used or not, it may be determined whether SVT is used. Or, depending on whether SVT is used or not, it may be determined whether multi-hypothesis prediction is used. This is because multi-hypothesis prediction can improve the prediction performance for the entire block and reduce the occurrence of the phenomenon where residues gather only in a part of the block.

[0317] According to an embodiment of the present disclosure, depending on whether multi-hypothesis prediction is used or the mode of multi-hypothesis prediction, the position of the TU to be converted for SVT may be restricted. Alternatively, depending on whether multi-hypothesis prediction is used or the mode of multi-hypothesis prediction, the width (SVT-V) or height (SVT-H) of the TU to be converted for SVT may be restricted. Therefore, signaling regarding the position, width, or height can be reduced. For example, the position of the TU to be converted for SVT may not be on the side where the weight value of intra prediction may be large in multi-hypothesis prediction. This is because the residual on the side with a large weight value can be reduced by multi-hypothesis prediction. Therefore, when using multi-hypothesis prediction, there may be no mode of converting the side with a large weight value in SVT. For example, when using the horizontal mode or the vertical mode in multi-hypothesis prediction, it is possible to omit position 1 in FIG. 25. As yet another example, when using the planar mode in multi-hypothesis prediction, the position of the TU to be converted for SVT may be restricted. For example, when using the planar mode in multi-hypothesis prediction, it is possible to omit position 0 in FIG. 25. This is because when using the planar mode in multi-hypothesis prediction, the portion near the reference sample of intra prediction may produce a value similar to the reference sample value, and thus the residual near the reference sample may be small.

[0318] As yet another example, when using multi-hypothesis prediction, the possible values of the width or height of the TU to be converted for SVT may change. Alternatively, when using a specific mode in multi-hypothesis prediction, the possible values of the width or height of the TU to be converted for SVT may change. For example, when using multi-hypothesis prediction, since less residual remains in the wider portion of the block, larger values of the width or height of the TU to be converted for SVT may be excluded. Alternatively, when using multi-hypothesis prediction, the values of the width or height of the TU to be converted for SVT that are the same as the unit in which the weight value changes in multi-hypothesis prediction may be excluded.

[0319] Referring to FIG. 26, there may be a cu_sbt_flag indicating whether SBT is used and an mh_intra_flag indicating whether multi-hypothesis prediction is used. Referring to the drawings, when mh_intra_flag is 0, cu_sbt_flag can be parsed. Also, when cu_sbt_flag does not exist, it can be inferred as 0.

[0320] Both combining intra prediction in multi-hypothesis prediction and SBT are for solving the problem that a large amount of residue may remain only in part of the CU or PU when the corresponding technology is not used. Therefore, since there may be a correlation between the two technologies, it is possible to determine whether one technology is used or whether a specific mode of one technology is used, etc., based on that of the other technology.

[0321] Also, in FIG. 26, sbtblockConditions can indicate the conditions under which SBT is possible. The conditions under which SBT is possible can include conditions related to the block size, signaling values indicating whether it is used in a high-level (e.g., sequence, slice, tile, etc.) syntax, and the like.

[0322] FIG. 27 is a diagram showing the relationship of color components according to an embodiment of the present disclosure.

[0323] Referring to FIG. 27, the color format may be indicated by chroma_format_idc, Chroma format, separate_colour_plane_flag, etc.

[0324] If it is monochrome, only one sample array may exist. Also, both SubWidthC and SubHeightC may be 1.

[0325] In the case of 4:2:0 sampling, two chroma arrays may exist. Also, the chroma array may have a half width and a half height of the luma array. Both the information regarding the width (SubWidthC) and the information regarding the height (SubHeightC) may be 2.

[0326] The information regarding the width (SubWidthC) and the information regarding the height (SubHeightC) can indicate what size the chroma array has compared to the luma array. When the width or height of the chroma array is half the size of the luma array, the information regarding the width (SubWidthC) or the information regarding the height (SubHeightC) is 2, and when the width or height of the chroma array is the same size as the luma array, the information regarding the width (SubWidthC) or the information regarding the height (SubHeightC) may be 1.

[0327] In the case of 4:2:2 sampling, two chroma arrays may exist. Also, the chroma array may have a half width and the same height as the luma array. SubWidthC and SubHeightC may be 2 and 1 respectively.

[0328] In the case of 4:4:4 sampling, the chroma array may have the same width and the same height as the luma array. Both SubWidthC and SubHeightC may be 1. At this time, the processing may differ based on the separate_colour_plane_flag. If the separate_colour_plane_flag is 0, the chroma array may have the same width and height as the luma array. If the separate_colour_plane_flag is 1, the three color planes (luma, Cb, Cr) can be processed respectively. Regardless of the separate_colour_plane_flag, when it is 4:4:4, both SubWidthC and SubHeightC may be 1.

[0329] If the separate_colour_plane_flag is 1, it is possible that only one thing corresponding to one colour component exists in one slice. If the separate_colour_plane_flag is 0, it is possible that there are things corresponding to multiple colour components in one slice.

[0330] Referring to Figure 27, SubWidthC and SubHeightC can be different only when it is 4:2:2. Therefore, when it is 4:2:2, the relationship between the luma reference width and height and the relationship between the chroma reference width and height may be different.

[0331] For example, when the luma sample reference width is width L and the chroma sample reference width is width C, if width L and width C are corresponding, the relationship between the two may be as follows.

[0332] widthC = widthL / SubWidthC

[0333] That is, widthL = widthC * SubWidthC

[0334] Similarly, when the luma sample reference height is heightL and the chroma sample reference height is heightC, if heightL and heightC are corresponding, the relationship between the two may be as follows.

[0335] heightC = heightL / SubHeightC

[0336] That is, heightL = heightC * SubHeightC

[0337] Also, there may be values indicating color components. For example, cIdx can indicate a color component. For example, cIdx may be a color component index. If cIdx is 0, it can indicate a luma component. Also, if cIdx is not 0, it can indicate a chroma component. Also, if cIdx is 1, it can indicate a chroma Cb component. Also, if cIdx is 2, it can indicate a chroma Cr component.

[0338] FIG. 28 is a diagram showing the relationship of color components according to an embodiment of the present disclosure.

[0339] FIGS. 28(a), (b), and (c) show the cases of 4:2:0, 4:2:2, and 4:4:4, respectively.

[0340] Referring to FIG. 28(a), one chroma sample (one Cb and one Cr) may be positioned per two luma samples in the horizontal direction. Also, one chroma sample (one Cb and one Cr) may be positioned per two luma samples in the vertical direction.

[0341] Referring to FIG. 28(b), one chroma sample (one Cb and one Cr) may be positioned per two luma samples in the horizontal direction. Also, one chroma sample (one Cb and one Cr) may be positioned per one luma sample in the vertical direction.

[0342] Referring to FIG. 28(c), one chroma sample (one Cb and one Cr) may be positioned per one luma sample in the horizontal direction. Also, one chroma sample (one Cb and one Cr) may be positioned per one luma sample in the vertical direction.

[0343] As described above, such a relationship may determine SubWidthC and SubHeightC described in FIG. 27, and based on SubWidthC and SubHeightC, conversion between the luma sample standard and the chroma sample standard can be performed.

[0344] FIG. 29 is a diagram showing a peripheral reference position according to an embodiment of the present disclosure.

[0345] According to an embodiment of the present disclosure, a peripheral position can be referred to when making a prediction. For example, as described above, a peripheral position can be referred to when performing CIIP. CIIP may be the multi-hypothesis prediction described above. CIIP may be combined inter-picture merge and intra-picture prediction. That is, CIIP may be a prediction method that combines inter prediction (e.g., merge mode inter prediction) and intra prediction.

[0346] According to an embodiment of the present disclosure, it is possible to combine inter prediction and intra prediction by referring to a peripheral position. For example, it is possible to determine the ratio of inter prediction to intra prediction by referring to a peripheral position. Or, it is possible to determine a weighting value when combining inter prediction and intra prediction by referring to a peripheral position. Or, it is possible to determine a weighting value when weighted-summing (weighted-averaging) inter prediction and intra prediction by referring to a peripheral position.

[0347] According to an embodiment of the present disclosure, the peripheral positions to be referred to can include NbA and NbB. The coordinates of NbA and NbB may be (xNbA, yNbA) and (xNbB, yNbB), respectively.

[0348] Also, NbA may be at the left position of the current block. More specifically, when the upper left coordinates of the current block are (xCb, yCb), and the width and height of the current block are cbWidth and cbHeight respectively, NbA may be (xCb - 1, yCb + cbHeight - 1). The upper left coordinates (xCb, yCb) of the current block may be values based on luma samples. Or, the upper left coordinates (xCb, yCb) of the current block may be the upper left sample luma position of the current luma coding block relative to the upper left luma sample of the current picture. Also, the cbWidth and cbHeight may be values based on the corresponding color component. The coordinates described above may be for the luma component (luma block). For example, cbWidth and cbHeight may be values based on the luma component.

[0349] Also, NbB may be at the above position of the current block. More specifically, when the upper left coordinates of the current block are (xCb, yCb), and the width and height of the current block are cbWidth and cbHeight respectively, NbB may be (xCb + cbWidth - 1, yCb - 1). The upper left coordinates (xCb, yCb) of the current block may be values based on luma samples. Or, the upper left coordinates (xCb, yCb) of the current block may be the upper left sample luma position of the current luma coding block relative to the upper left luma sample of the current picture. Also, the cbWidth and cbHeight may be values based on the corresponding color component. The coordinates described above may be for the luma component (luma block). For example, cbWidth and cbHeight may be values based on the luma component.

[0350] Referring to FIG. 29, the upper left end, the coordinates of NbA, the coordinates of NbB, etc. are indicated for the block displayed with the luma block at the upper end.

[0351] Also, NbA may be at the left position of the current block. More specifically, when the upper left coordinates of the current block are (xCb, yCb), and the width and height of the current block are cbWidth and cbHeight respectively, NbA may be (xCb - 1, yCb + 2 * cbHeight - 1). The upper left coordinates (xCb, yCb) of the current block may be values based on luma samples. Or, the upper left coordinates (xCb, yCb) of the current block may be the upper left sample luma position of the current luma coding block relative to the upper left luma sample of the current picture. Also, the cbWidth and cbHeight may be values based on the corresponding color component. The coordinates described above may be for the chroma component (chroma block). For example, cbWidth and cbHeight may be values based on the chroma component. Also, this coordinate may be applicable when in the 4:2:0 format.

[0352] Also, NbB may be at the above position of the current block. More specifically, when the upper left coordinates of the current block are (xCb, yCb), and the width and height of the current block are cbWidth and cbHeight respectively, NbB may be (xCb + 2 * cbWidth - 1, yCb - 1). The upper left coordinates (xCb, yCb) of the current block may be values based on luma samples. Or, the upper left coordinates (xCb, yCb) of the current block may be the upper left sample luma position of the current luma coding block relative to the upper left luma sample of the current picture. Also, the cbWidth and cbHeight may be values based on the corresponding color component. The coordinates described above may be for the chroma component (chroma block). For example, cbWidth and cbHeight may be values based on the chroma component. Also, this coordinate may be applicable when in the 4:2:0 format or the 4:2:2 format.

[0353] Referring to FIG. 29, for the block displayed with a chroma block at the lower end, the upper left end, the coordinates of NbA, the coordinates of NbB, etc. are indicated.

[0354] FIG. 30 is a diagram showing a weighted sample prediction process according to an embodiment of the present disclosure.

[0355] The embodiment of FIG. 30 may relate to a method of combining two or more prediction signals. Also, the embodiment of FIG. 30 is applicable when using CIIP. Also, the embodiment of FIG. 30 may include the peripheral position reference method described in FIG. 29.

[0356] Referring to equation (8-838) in FIG. 30, the scale information (scallFact) can be explained as follows.

[0357] scallFact = (cIdx == 0)? 0 : 1

[0358] That is, when the information (cIdx) regarding the color component of the current block is 0, the scale information (scallFact) can be set to 0, and when the information (cIdx) regarding the color component of the current block is not 0, the scale information (scallFact) can be set to 1. In the embodiment of the present disclosure, x?y:z can indicate the y value when x is true (or x is not 0), and can indicate the z value otherwise (when x is false (or x is 0)).

[0359] Also, the coordinates (xNbA, yNbA) and (xNbB, yNbB) of the surrounding positions NbA and NbB to be referred to can be set. According to the embodiment described in FIG. 29, for the luma component, (xNbA, yNbA) and (xNbB, yNbB) are (xCb - 1, yCb + cbHeight - 1) and (xCb + cbWidth - 1, yCb - 1) respectively, and for the chroma component, (xNbA, yNbA) and (xNbB, yNbB) may be (xCb - 1, yCb + 2 * cbHeight - 1) and (xCb + 2 * cbWidth - 1, yCb - 1) respectively. Also, the operation of multiplying by 2^n may be the same as left-shifting n bits. For example, the operation of multiplying by 2 may be the same as left-shifting 1 bit. Also, left-shifting x by n bits can be expressed as "x << n". Also, the operation of dividing by 2^n may be the same as right-shifting n bits. Also, the operation of dividing by 2^n and discarding the decimal part may be the same as right-shifting n bits. For example, the operation of dividing by 2 may be the same as right-shifting 1 bit. Also, right-shifting x by n bits can be expressed as "x >> n". Therefore, (xCb - 1, yCb + 2 * cbHeight - 1) and (xCb + 2 * cbWidth - 1, yCb - 1) can be expressed as (xCb - 1, yCb + (cbHeight << 1) - 1) and (xCb + (cbWidth << 1) - 1, yCb - 1). Therefore, when representing both the coordinates for the above-described luma component and the coordinates for the chroma component, it may be as follows.

[0360] (xNbA, yNbA) = (xCb - 1, yCb + (cbHeight << scallFact) - 1)

[0361] (xNbB, yNbB) = (xCb + (cbWidth<<scallFact) - 1, yCb - 1)

[0362] Here, the scale information (scallFact) may be (cIdx == 0)? 0 : 1 as described above. At this time, cbWidth and cbHeight may be indicated based on each color component. For example, when the width and height based on the luma component are cbWidthL and cbHeightL respectively, cbWidth and cbHeight may be cbWidthL and cbHeightL respectively when performing the weighted sample prediction process for the luma component. Also, when the width and height based on the luma component are cbWidthL and cbHeightL respectively, cbWidth and cbHeight may be cbWidthL / SubWidthC and cbHeightL / SubHeightC respectively when performing the weighted sample prediction process for the chroma component.

[0363] Also, according to one embodiment, it is possible to determine the prediction mode of a corresponding position with reference to the peripheral position. For example, it is possible to determine whether the prediction mode is intra prediction. Also, the prediction mode may be indicated by CuPredMode. When CuPredMode is MODE_INTRA, intra prediction may be used. Also, the CuPredMode value may be MODE_INTRA, MODE_INTER, MODE_IBC, MODE_PLT. When CuPredMode is MODE_INTER, inter prediction may be used. Also, when CuPredMode is MODE_IBC, intra block copy (IBC) may be used. Also, when CuPredMode is MODE_PLT, palette mode may be used. Also, the CuPredMode may be indicated by the channel type (chType) and the position. For example, it may be indicated as CuPredMode[chType][x][y], and this value may be the CuPredMode value for the channel type chType at the (x, y) position. Also, the chType may be based on the tree type. For example, the tree type may be set to values such as SINGLE_TREE, DUAL_TREE_LUMA, DUAL_TREE_CHROMA. In the case of SINGLE_TREE, there may be a part where the block partitioning of the luma component and the chroma component is shared. For example, in the case of SINGLE_TREE, the block partitioning of the luma component and the chroma component may be the same. Or, in the case of SINGLE_TREE, the block partitioning of the luma component and the chroma component may be the same or partially the same.Alternatively, when it is SINGLE_TREE, the block partitioning of the luma component and the chroma component may be performed with the same syntax element value. Also, when it is DUAL TREE, the block partitioning of the luma component and the chroma component may be independent. Or, when it is DUAL TREE, the block partitioning of the luma component and the chroma component may be performed with individual syntax element values. Also, when it is DUAL TREE, the treeType value may be DUAL_TREE_LUMA or DUAL_TREE_CHROMA. If the treeType is DUAL_TREE_LUMA, it can be indicated that DUAL TREE is used and it is a process for the luma component. If the treeType is DUAL_TREE_CHROMA, it can be indicated that DUAL TREE is used and it is a process for the chroma component. Also, chType may be determined based on whether the tree type is DUAL_TREE_CHROMA. For example, chType may be set to 1 when the treeType is DUAL_TREE_CHROMA and set to 0 when the treeType is not DUAL_TREE_CHROMA. Therefore, referring to FIG. 30, the CuPredMode[0][xNbX][yNbY] value can be determined. X may be replaced with A and B. That is, the CuPredMode values for the NbA and NbB positions can be determined.

[0364] Also, the isIntraCodedNeighbourX value can be set based on the determination of the prediction mode for the peripheral position. For example, the isIntraCodedNeighbourX value can be set according to whether the CuPredMode for the peripheral position is MODE_INTRA or not. If the CuPredMode for the peripheral position is MODE_INTRA, the isIntraCodedNeighbourX value can be set to TRUE, and if the CuPredMode for the peripheral position is not MODE_INTRA, the isIntraCodedNeighbourX value can be set to FALSE. As described above and in the present disclosure to be described later, X may be replaced with A or B, etc. Also, what is written as X can indicate corresponding to the X position.

[0365] Also, according to one embodiment, it is possible to determine whether the corresponding position is available by referring to the peripheral position. It is possible to set availableX according to whether the corresponding position is available. Also, isIntraCodedNeighbourX can be set based on availableX. For example, when availableX is TRUE, isIntraCodedNeighbourX can be set to TRUE, and for example, when availableX is FALSE, isIntraCodedNeighbourX can be set to FALSE. Referring to FIG. 30, whether the corresponding position is available may be determined by "The derivation process for neighbouring block availability". Also, whether the corresponding position is available can be determined based on whether the corresponding position is inside the current picture. When the corresponding position is (xNbY, yNbY), if xNbY or yNbY is less than 0, it is outside the current picture and availableX may be set to FALSE. Also, when xNbY is greater than or equal to the picture width, it is outside the current picture and availableX may be set to FALSE. The picture width may be indicated by pic_width_in_luma_samples. Also, when yNbY is greater than or equal to the picture height, it is outside the current picture and availableX may be set to FALSE. The picture height may be indicated by pic_height_in_luma_samples. Also, when the corresponding position is in a different block (brick) or another slice from the current block, availableX may be set to FALSE. Also, when the reconstruction of the corresponding position is not completed, availableX may be set to FALSE. Whether the final configuration is completed may be indicated by IsAvailable[cIdx][xNbY][yNbY].Therefore, in short, when any one of the following conditions is satisfied, availableX can be set to FALSE, and otherwise (when all of the following conditions are not satisfied), availableX can be set to TRUE.

[0366] Condition 1: xNbY < 0

[0367] Condition 2: yNbY < 0

[0368] Condition 3: xNbY >= pic_width_in_luma_samples

[0369] Condition 4: yNbY >= pic_height_in_luma_samples

[0370] Condition 5: IsAvailable[cIdx][xNbY][yNbY] == FALSE

[0371] Condition 6: When the corresponding position (peripheral position (xNbY, yNbY) position) belongs to a block (or a different slice) different from the current block

[0372] Also, optionally, it can be determined whether the current position and the corresponding position are in the same CuPredMode, and availableX can be set.

[0373] The two described conditions can be combined to set isIntraCodedNeighbourX. For example, isIntraCodedNeighbourX can be set to TRUE when all of the following conditions are satisfied, and otherwise (when any one of the following conditions is not satisfied), isIntraCodedNeighbourX can be set to FALSE.

[0374] Condition 1: availableX == TRUE

[0375] Condition 2: CuPredMode[0][xNbX][yNbX] == MODE_INTRA

[0376] Also, according to an embodiment of the present disclosure, the weighting of CIIP can be determined based on a number of code information (isIntraCodedNeighbourX). For example, the weighting can be determined when combining inter prediction and intra prediction based on a number of code information (isIntraCodedNeighbourX). For example, it can be determined based on the code information (isIntraCodedNeighbourA) regarding the left block and the code information (isIntraCodedNeighbourB) regarding the upper block. According to one embodiment, when both the code information (isIntraCodedNeighbourA) regarding the left block and the code information (isIntraCodedNeighbourB) regarding the upper block are TRUE, the weighting value (w) can be set to 3. For example, the weighting value (w) may be the weighting of CIIP or a value for determining the weighting. Also, when both the code information (isIntraCodedNeighbourA) regarding the left block and the code information (isIntraCodedNeighbourB) regarding the upper block are FALSE, w can be set to 1. Also, when either one of the code information (isIntraCodedNeighbourA) regarding the left block and the code information (isIntraCodedNeighbourB) regarding the upper block is FALSE (the same as when one of the two is TRUE), the weighting value (w) can be set to 2. That is, the weighting value (w) may be set based on whether the surrounding position is predicted by intra prediction or based on how much the surrounding position is predicted by intra prediction.

[0377] Also, the weighting value (w) may be the weighting corresponding to intra prediction. Also, the weighting corresponding to inter prediction may be determined based on the weighting value (w). For example, the weighting corresponding to inter prediction may be (4 - w). Referring to equation (8 - 840) in FIG. 30, when combining two or more prediction signals, it can be done as follows.

[0378] predSampleComb[x][y] = (w*predSamplesIntra[x][y] + (4 - w)*predSamplesInter[x][y] + 2 ) >> 2

[0379] Here, the second sample (predSamplesIntra) and the first sample (predSamplesInter) may be prediction signals. For example, the second sample (predSamplesIntra) and the first sample (predSamplesInter) may be prediction signals predicted by intra prediction and prediction signals predicted by inter prediction (e.g., merge mode, more specifically, regular merge mode), respectively. Also, the combined prediction sample (predSampleComb) may be a prediction signal used in CIIP.

[0380] Also, a process of updating the prediction signal before performing the equation (8 - 840) in FIG. 30 may be included. For example, it can be updated in a process such as the equation (8 - 839) in FIG. 30. For example, the process of updating the prediction signal may be a process of updating the inter prediction signal of CIIP.

[0381] FIG. 31 is a diagram showing peripheral reference positions according to an embodiment of the present disclosure.

[0382] Although the peripheral reference positions have been described with reference to FIGS. 29 to 30, there may be problems when using the described positions in all cases (e.g., all chroma blocks), and this problem will be described with reference to FIG. 31.

[0383] The embodiment of FIG. 31 shows a chroma block. In FIGS. 29 to 30, the NbA and NbB coordinates based on the luma sample for the chroma block were (xCb - 1, yCb + 2 * cbHeight - 1) and (xCb + 2 * cbWidth - 1, yCb - 1), respectively. However, when SubWidthC or SubHeightC is 1, it may indicate a position in FIG. 31 different from the position shown in FIG. 29. Multiplying cbWidth and cbHeight by 2 in the above coordinates means that cbWidth and cbHeight are shown based on each color component (chroma component in this embodiment), and since the coordinates are shown based on the luma standard, it may be for compensating the number of luma samples to chroma samples in the case of 4:2:0. That is, it may be for showing the coordinates in the case where there is 1 chroma sample corresponding to 2 luma samples based on the x-axis of the luma sample and 1 chroma sample corresponding to 2 luma samples based on the y-axis of the luma sample. Therefore, when SubWidthC or SubHeightC is 1, it may indicate a different position. Therefore, if the positions of (xCb - 1, yCb + 2 * cbHeight - 1) and (xCb + 2 * cbWidth - 1, yCb - 1) are always used for the chroma block in this way, it may refer to a position far from the current chroma block. Also, in this case, the relative position used in the luma block of the current block may not match the relative position used in the chroma block. Further, since different positions are referred to for the chroma block, weights may be set by referring to positions with little relevance to the current block, or decoding / reconstruction may not be performed in the block decoding order.

[0384] Referring to FIG. 31, in the case of 4:4:4, that is, when both SubWidthC and SubHeightC are 1, it shows the positions of the above-described luma-based coordinates. NbA and NbB may exist at positions away from the chroma block shown by the solid line.

[0385] FIG. 32 is a diagram showing a weighted sample prediction process according to an embodiment of the present disclosure.

[0386] The embodiment of FIG. 32 may be an embodiment for solving the problems described in FIGS. 29 to 31. Also, the above-described content may be omitted.

[0387] In FIG. 30, the peripheral position is set based on the scale information (scallFact). As described in FIG. 31, the scale information (scallFact) was a value for converting the position when SubWidthC and SubHeightC were 2.

[0388] However, as described above, problems may occur depending on the color format, and the ratio of chroma samples to luma samples may be different horizontally and vertically. According to the embodiments of the present disclosure, the scale information (scallFact) can be separated into horizontal (width) and vertical (height).

[0389] According to the embodiments of the present disclosure, there may be scale information for the x-axis (scallFactWidth) and scale information for the y-axis (scallFactHeight), and the peripheral position can be set based on the scale information for the x-axis (scallFactWidth) and the scale information for the y-axis (scallFactHeight). Also, the peripheral position may be set based on luma samples (luma blocks).

[0390] The video signal processing apparatus may perform a step of acquiring information regarding the width (SubWidthC) and information regarding the height (SubHeightC) based on the chroma format information (chroma_format_idc). Here, the chroma format information (chroma_format_idc) may be signaled in any one of a coding tree unit, a slice, a tile, a tile group, a picture, or a sequence unit.

[0391] The video signal processing apparatus can obtain information (SubWidthC) regarding width and information (SubHeightC) regarding height based on chroma format information (chroma_format_idc) according to the table shown in FIG. 27.

[0392] The video signal processing apparatus can perform a step (8-838) of obtaining x-axis scale information (scallFactWidth) based on information (SubWidthC) regarding width or information (cIdx) regarding the color component of the current block. More specifically, the x-axis scale information (scallFactWidth) may be set based on information (cIdx) regarding the color component of the current block and information (SubWidthC) regarding width. For example, when the information (cIdx) regarding the color component of the current block is 0 or the information (SubWidthC) regarding width is 1, the x-axis scale information (scallFactWidth) can be set to 0, and otherwise (when the information (cIdx) regarding the color component of the current block is not 0 and the information (SubWidthC) regarding width is not 1 (when SubWidthC is 2)), the x-axis scale information (scallFactWidth) can be set to 1. Referring to the equation (8-838) in FIG. 32, it can be shown as follows.

[0393] scallFactWidth = ( cIdx == 0 || SubWidthC == 1)? 0 : 1

[0394] Also, the video signal processing apparatus can perform a step (8-839) of obtaining y-axis scale information (scallFactHeight ) based on information (SubHeightC) regarding height or information (cIdx) regarding the color component of the current block.

[0395] Also, the scale information of the y-axis (scallFactHeight) can be set based on the information about the color component of the current block (cIdx) and the information about the height (SubHeightC). For example, when the information about the color component of the current block (cIdx) is 0 or the information about the height (SubHeightC) is 1, the scale information of the y-axis (scallFactHeight) can be set to 0. Otherwise (when the information about the color component of the current block (cIdx) is not 0 and the information about the height (SubHeightC) is not 1 (when SubHeightC is 2)), the scale information of the y-axis (scallFactHeight) can be set to 1. Referring to Equation (8-839) in Figure 32, it can be shown as follows.

[0396] scallFactHeight = ( cIdx == 0 || SubHeightC == 1)? 0 : 1

[0397] Also, the x - coordinate of the peripheral position can be indicated based on the scale information (scallFactWidth) of the x - axis, and the y - coordinate of the peripheral position can be indicated based on the scale information (scallFactHeight) of the y - axis. The video signal processing apparatus can perform a step of determining the position of the left - hand block (NbA) based on the scale information (scallFactHeight) of the y - axis. Also, the video signal processing apparatus can perform a step of determining the position of the upper - hand block (NbB) based on the scale information (scallFactWidth) of the x - axis. For example, the coordinates of the upper - hand block (NbB) can be set based on the scale information (scallFactWidth) of the x - axis. For example, the coordinates of the left - hand block (NbA) can be set based on the scale information (scallFactHeight) of the y - axis. Also, as described above, here, being based on the scale information (scallFactWidth) of the x - axis may mean being based on the information regarding the width (SubWidthC), and being based on the scale information (scallFactHeight) of the y - axis may mean being based on the information regarding the height (SubHeightC). For example, the peripheral position coordinates may be as follows.

[0398] (xNbA, yNbA) = (xCb - 1, yCb + (cbHeight << scallFactHeight) - 1)

[0399] (xNbB, yNbB) = (xCb + (cbWidth<<scallFactWidth) - 1, yCb - 1)

[0400] (xCb,yCb) may be the upper - left sample luma position of the current luma block (luma coding block) with respect to the upper - left luma sample of the current picture. cbWidth and cbHeight may be the width (width) and height (height) of the current block, respectively. At this time, xCb and yCb may be the coordinates shown based on the luma sample reference as described above. Also, cbWidth and cbHeight may be shown based on each color component.

[0401] Therefore, in the case of a chroma block where SubWidthC is 1, (xNbB, yNbB) may be (xCb + cbWidth - 1, yCb - 1). That is, in this case, the NbB coordinates for the luma block and the NbB coordinates for the chroma block may be the same. Also, in the case of a chroma block where SubHeightC is 1, (xNbA, yNbA) may be (xCb - 1, yCb + cbHeight - 1). That is, in this case, the NbA coordinates for the luma block and the NbA coordinates for the chroma block may be the same.

[0402] Therefore, the embodiment of FIG. 32 can set the same peripheral coordinates as the embodiments of FIGS. 29 to 30 in the case of the 4:2:0 format, and can set different peripheral coordinates from the embodiments of FIGS. 29 to 30 in the case of the 4:2:2 format or the 4:4:4 format.

[0403] The video signal processing apparatus can perform a step of determining a weighting value (w) based on the left block (NbA) and the upper block (NbB). FIG. 32 may be similar to that described in FIG. 30. That is, based on the peripheral position coordinates described in FIG. 32, a prediction mode or availability can be determined, and the weighting of CIIP can be determined. Among the descriptions regarding FIG. 32, some descriptions overlapping with FIG. 30 may be omitted.

[0404] Referring to FIG. 32, the video signal processing apparatus can combine two conditions to set the code information (isIntraCodedNeighbourX). For example, when all of the following conditions are satisfied, the code information (isIntraCodedNeighbourX) can be set to TRUE, and otherwise (when at least one of the following conditions is not satisfied), the code information (isIntraCodedNeighbourX) can be set to FALSE.

[0405] Condition 1: availableX == TRUE

[0406] Condition 2: CuPredMode[0][xNbX][yNbX] == MODE_INTRA

[0407] More specifically, referring to line 3210, the video signal processing apparatus can perform a step of setting the code information (isIntraCodedNeighbourA) regarding the left block to TRUE when the left block is available (availableA == TRUE) and the prediction mode of the left block is intra prediction (CuPredMode[0][xNbA][yNbA] is equal to MODE_INTRA).

[0408] Also, the video signal processing apparatus can perform a step of setting the code information regarding the left block to FALSE when the left block is not available or the prediction mode of the left block is not intra prediction.

[0409] Also, the video signal processing apparatus can perform a step of setting the code information (isIntraCodedNeighbourB) regarding the upper block to TRUE when the upper block is available (availableB == TRUE) and the prediction mode of the upper block is intra prediction (CuPredMode[0][xNbB][yNbB] is equal to MODE_INTRA).

[0410] Also, the video signal processing apparatus can perform a step of setting the code information regarding the upper block to FALSE when the upper block is not available or the prediction mode of the upper block is not intra prediction.

[0411] Referring to line 3220, when both the code information (isIntraCodedNeighbourA) regarding the left block and the code information (isIntraCodedNeighbourB) regarding the upper block are TRUE, the video signal processing apparatus can perform the step of determining the weighting value (w) to be 3. Also, when both the code information (isIntraCodedNeighbourA) regarding the left block and the code information (isIntraCodedNeighbourB) regarding the upper block are FALSE, the video signal processing apparatus can perform the step of determining the weighting value (w) to be 1. Further, when only one of the code information (isIntraCodedNeighbourA) regarding the left block and the code information (isIntraCodedNeighbourB) regarding the upper block is TRUE, the video signal processing apparatus can perform the step of determining the weighting value (w) to be 2.

[0412] The video signal processing apparatus can perform the step of obtaining a first sample (predSamplesInter) that predicts the current block in merge mode. Also, the video signal processing apparatus can perform the step of obtaining a second sample (predSamplesIntra) that predicts the current block in intra mode.

[0413] The video signal processing apparatus can perform the step (8-841) of obtaining a combined prediction sample (predSampleComb) for the current block based on the weighting value (w), the first sample (predSamplesInter), and the second sample (predSamplesIntra). For example, the video signal processing apparatus can obtain the combined prediction sample (predSampleComb) based on the following formula.

[0414] predSampleComb[x][y] = (w*predSamplesIntra[x][y] + (4-w)*predSamplesInter[x][y] + 2 ) >> 2

[0415] Here, predSamplesComb means combined prediction samples, w means weight values, predSamplesIntra means the second samples, predSamplesInter means the first samples, [x] can mean the x-axis coordinate of the samples included in the current block, and [y] can mean the y-axis coordinate of the samples included in the current block.

[0416] In the present disclosure, the peripheral position and the peripheral position coordinates may be used with the same meaning.

[0417] FIG. 33 is a diagram showing a weighted sample prediction process according to an embodiment of the present disclosure.

[0418] The embodiment of FIG. 33 represents the peripheral position coordinates described in FIG. 32 in another way. Therefore, the content overlapping with the foregoing may be omitted.

[0419] As described above, bit shift can be expressed by multiplication. FIG. 32 may be shown using bit shift, and FIG. 33 may be shown using multiplication.

[0420] According to an embodiment, the scale information (scallFactWidth) of the x-axis can be set based on the information (cIdx) regarding the color component of the current block and the information (SubWidthC) regarding the width. For example, when the information (cIdx) regarding the color component of the current block is 0 or the information (SubWidthC) regarding the width is 1, the scale information (scallFactWidth) of the x-axis can be set to 1, and otherwise (when cIdx is not 0 and the information (SubWidthC) regarding the width is not 1 (when SubWidthC is 2)), the scale information (scallFactWidth) of the x-axis can be set to 2. Referring to equation (8-838) in FIG. 33, it can be shown as follows.

[0421] scallFactWidth = (cIdx == 0 || SubWidthC == 1)? 1 : 2

[0422] Also, the scale information of the y-axis (scallFactHeight) can be set based on the information about the color component of the current block (cIdx) and the information about the height (SubHeightC). For example, when the information about the color component of the current block (cIdx) is 0 or the information about the height (SubHeightC) is 1, the scale information of the y-axis (scallFactHeight) can be set to 1, and when it is not the case (when cIdx is not 0 and the information about the height (SubHeightC) is not 1 (when SubHeightC is 2)), the scale information of the y-axis (scallFactHeight) can be set to 2. Referring to Equation (8-839) in FIG. 33, it can be shown as follows.

[0423] scallFactHeight = (cIdx == 0 || SubHeightC == 1)? 1 : 2

[0424] Also, the x-coordinate of the peripheral position can be indicated based on the scale information of the x-axis (scallFactWidth), and the y-coordinate of the peripheral position can be indicated based on the scale information of the y-axis (scallFactHeight). For example, the coordinates of NbB can be set based on the scale information of the x-axis (scallFactWidth). For example, the coordinates of NbA can be set based on the scale information of the y-axis (scallFactHeight). Also, as described above, the basis of the scale information of the x-axis (scallFactWidth) here may be based on SubWidthC, and the basis of the scale information of the y-axis (scallFactHeight) may be based on SubHeightC. For example, the peripheral position coordinates may be as follows.

[0425] (xNbA, yNbA) = (xCb - 1, yCb + (cbHeight * scallFactHeight) - 1)

[0426] (xNbB, yNbB) = (xCb + (cbWidth * scallFactWidth) - 1, yCb - 1)

[0427] At this time, xCb and yCb may be the coordinates shown in the luma sample reference as described above. Also, cbWidth and cbHeight may be those shown based on each color component.

[0428] FIG. 34 is a diagram showing a weighted sample prediction process according to an embodiment of the present disclosure.

[0429] In embodiments such as FIGS. 30, 32, and 33, it was determined whether the corresponding position was available by referring to the peripheral position. At this time, cIdx, which is an index indicating a color component, was set to 0 (luma component). Also, when determining whether the corresponding position is available, information (cIdx) regarding the color component of the current block can be used to determine whether the reconstruction of the corresponding position cIdx is complete. That is, when determining whether the corresponding position is available, cIdx can be used to determine the IsAvailable[cIdx][xNbY][yNbY] value. However, when performing the weighted sample prediction process for a chroma block, referring to the IsAvailable value corresponding to cIdx0 may result in an incorrect determination. For example, when the reconstruction of the luma component of the block including the peripheral position is not complete but the reconstruction of the chroma component is complete, when cIdx is 0, IsAvailable[0][xNbY][yNbY] may be FALSE and IsAvailable[cIdx][xNbY][yNbY] may be TRUE. Therefore, there may be a case where it is determined that the periphery is not available even though it is actually available. In the embodiment of FIG. 34, in order to solve this problem, when determining whether the corresponding position is available by referring to the peripheral position, the cIdx of the current coding block can be used as an input. That is, when invoking "the derivation process for neighbouring block availability", the cIdx of the current coding block can be put as the input cIdx.

[0430] Also, the content described in FIGS. 32 to 33 may be omitted.

[0431] Also, when determining the prediction mode of the peripheral position described above, CuPredMode[0][xNbX][yNbY], which is the CuPredMode corresponding to chType0, is referred to. However, if the chType for the current block does not match, incorrect parameters may be referred to. Therefore, according to the embodiments of the present disclosure, when determining the prediction mode of the peripheral position, CuPredMode[chType][xNbX][yNbY] corresponding to the chType value corresponding to the current block can be referred to.

[0432] FIG. 35 is a diagram showing CIIP weight derivation according to an embodiment of the present disclosure.

[0433] In the embodiments described with reference to FIGS. 29 to 34, the content related to CIIP weight derivation is described, and the above-described content may be omitted in this embodiment.

[0434] Also, in the embodiments described with reference to FIGS. 29 to 34, it was possible to determine the weighting used in CIIP based on relatively the same positions for a number of color components. Therefore, this can mean that the weighting used in CIIP can be determined based on the neighboring locations described for the luma component with respect to the chroma component. Also, therefore, this can mean that the weighting used in CIIP can be determined based on the neighboring locations described for the luma component with respect to a number of color components. Also, therefore, this can mean that the weighting used in CIIP can be determined based on the same neighboring locations for a number of color components. Also, this can mean that the weighting used in CIIP for a number of color components is the same. This is because, as described above, it can refer to the prediction mode of the neighboring locations. Or because the prediction modes of relatively the same positions with respect to the color component field can be the same. More specifically, when the prediction modes corresponding to the already set positions for each color component are the same with respect to the color component field, the weighting used in CIIP for a number of color components may be the same. Also, when inter prediction can be used, the prediction modes corresponding to the already set positions for each color component may be the same with respect to the color component field. Or, when it is a P slice or a B slice, the prediction modes corresponding to the already set positions for each color component may be the same with respect to the color component field. This is because it can be SINGLE_TREE when it is a P slice or a B slice.

[0435] Also, as described above, the neighboring locations described for the luma component may be as follows.

[0436] (xNbA, yNbA) = (xCb - 1, yCb + cbHeight - 1)

[0437] (xNbB, yNbB) = (xCb + cbWidth - 1, yCb - 1)

[0438] That is, according to the embodiments of the present disclosure, the weightings used for CIIP for a number of color components may be the same. More specifically, it is possible to use the weighting derived for the luma component for the chroma component in CIIP. This may be for preventing the process of deriving weightings for a number of color components from being performed multiple times. Also, for the chroma component, the weighting value based on (xNbA, yNbA), (xNbB, yNbB) can be used. At this time, cbHeight and cbWidth in (xNbA, yNbA), (xNbB, yNbB) may be values based on luma samples. Also, the xCb, yCb values in (xNbA, yNbA), (xNbB, yNbB) may be values based on luma samples. That is, also for the chroma component, it is possible to use the weighting value determined based on adjacent positions based on the width, height, and coordinates based on luma samples for CIIP.

[0439] Referring to FIG. 35, the derived weighting value is used for all color components, that is, Y, Cb, and Cr. That is, the same weighting value is used for CIIP for all color components.

[0440] FIG. 36 is a diagram showing a CIIP process according to an embodiment of the present disclosure.

[0441] The embodiment of FIG. 36 may be related to the embodiment described in FIG. 35. Also, the embodiment of FIG. 36 can show the structure of the embodiment described in FIG. 35. Also, in the embodiment of FIG. 36, the content described in FIGS. 29 to 35 may be omitted.

[0442] As described in FIG. 35, when using CIIP for each color component, it is possible to use the same weighting value.

[0443] Referring to FIG. 36, ciip_flag may be a signaling indicating whether CIIP is used. If CIIP is used, the process mentioned in FIG. 36 can be performed. If CIIP is used, it is possible to perform a weighting value derivation process. At this time, the weighting value derivation process can be performed regardless of the color component. The weighting value derivation process can include the process described in FIG. 30. At this time, the adjacent position can be defined regardless of the color component. For example, the following adjacent positions can be used.

[0444] (xNbA, yNbA) = (xCb - 1, yCb + cbHeight - 1)

[0445] (xNbB, yNbB) = (xCb + cbWidth - 1, yCb - 1)

[0446] Also, referring to FIG. 36, xCb, yCb, cbWidth, cbHeight which are the input of the weighting value derivation process and the xCb, yCb, cbWidth, cbHeight of the adjacent position may be component criteria that have already been set regardless of the color component currently being processed. For example, they may be values based on the luma sample criteria. That is, when performing the weighting value derivation process for the chroma component, it can be based on the xCb, yCb, cbWidth, cbHeight of the luma sample criteria. More specifically, when performing the weighting value derivation process for the chroma component, the adjacent position can be set based on the xCb, yCb, cbWidth, cbHeight of the luma sample criteria, and based on this, the weighting value can be determined.

[0447] The weighting value w can be set based on the weighting value derivation process. The weighting value derivation process in FIG. 36 may be the process of setting the weighting value w described in FIG. 30. At this time, as described, adjacent positions that have nothing to do with the color component can be used.

[0448] Referring to FIG. 36, when using CIIP, a general intra sample prediction process and a weighted sample prediction process can be performed. Also, this can be done for each color component field. The weighted sample prediction process can include a process of combining inter prediction and intra prediction using the weighting described in FIG. 30. For example, the weighted sample prediction process can include a process of combining using equation (8-840) of FIG. 30. Also, according to an embodiment of the present disclosure, the weighted sample prediction process can be performed based on a weighting value, and the same weighting value can be used for a number of color component fields. Also, the weighting value used in weighted sample prediction may be the weighting value w determined based on the weighting value derivation process described above.

[0449] Referring to FIG. 36, the width and height of the coding block, which is the input of the weighted sample prediction process, may be values indicated based on each color component. Referring to FIG. 36, in the case of the luma component (when cIdx is 0), the width and height are cbWidth and cbHeight, respectively, and in the case of the chroma component (when cIdx is not 0; when cIdx is 1 or 2), the width and height may be cbWidth / SubWidthC and cbHeight / SubHeightC, respectively.

[0450] That is, for a certain color component, a weighting value determined based on other color components can be used, and a weighted sample prediction process can be performed using the width and height based on the certain color component. Or, for a certain color component, a weighting value determined based on the width and height based on other color components can be used, and a weighted sample prediction process can be performed using the width and height based on the certain color component. At this time, the certain color component may be a chroma component, and the other color components may be luma components. As yet another example, the certain color component may be a luma component, and the other color components may be chroma components.

[0451] This may be for simplifying implementation. For example, the process shown in FIG. 36 may be performed for each color component, but by making the weighting value derivation process independent of the color components, the implementation of the weighting value derivation process can be simplified. Or, this may be for preventing a process from being repeated.

[0452] FIG. 37 is a diagram showing the ranges of MV and MVD according to an embodiment of the present disclosure.

[0453] According to an embodiment of the present disclosure, there may be a number of MVD generation methods or MVD determination methods. The MVD may be the motion vector difference value described above. Also, there may be a number of motion vector (MV) generation methods or MV determination methods.

[0454] For example, the MVD determination method may include a method of determining from the value of a syntax element. For example, the MVD can be determined based on the syntax element as described in FIG. 9. This will be further described with reference to FIG. 40.

[0455] In addition, the MVD determination method can include a method for determination when using the MMVD (merge with MVD) mode. For example, there may be an MVD used when using the merge mode. This will be further described with reference to FIGS. 38 to 39.

[0456] In an embodiment of the present disclosure, the MV or MVD can include those for a control point motion vector (CPMV) for performing affine motion compensation. That is, the MV and MVD can include the CPMV and CPMVD, respectively.

[0457] According to an embodiment of the present disclosure, the range that the MV or MVD can indicate may be limited. Thereby, the MV or MVD can be expressed and stored using limited resources and a limited number of bits, and operations using the MV or MVD can be performed. In an embodiment of the present disclosure, the range that the MV or MVD can indicate can be referred to as an MV range and an MVD range, respectively.

[0458] According to an embodiment, the MV range or the MVD range may be from -2^N to (2^N - 1). At this time, it may be a range including -2^N and (2^N - 1). According to still another embodiment, the MV range or the MVD range may be from (-2^N + 1) to 2^N. At this time, it may be a range including (-2^N + 1) and 2^N. In this embodiment, N is an integer, for example, a positive integer. More specifically, N may be 15 or 17. Also, at this time, the MV range or the MVD range can be indicated using N + 1 bits.

[0459] Also, in order to implement the described restricted MV range or restricted MVD range, it is possible to perform clipping and modulus operations. For example, the MV or MVD at a certain stage (e.g., before the final stage) can be clipped or modulus-operated so that a value within the MV range or MVD range is derived. Alternatively, in order to implement the restricted MV range or restricted MVD range, the binary representation range of the syntax element indicating the MV or MVD can be restricted.

[0460] According to an embodiment of the present disclosure, the MV range and the MVD range may be different. Also, there are a number of MVD determination methods, and MVD1 and MVD2 may be MVDs determined by different methods. According to still other embodiments of the present disclosure, the MVD1 range and the MVD2 range may be different. Using different ranges from each other may be for using different resources or different numbers of bits from each other, whereby it is not necessary to show an overly wide range for a certain element.

[0461] According to one embodiment, MVD1 may be the MVD described in FIG. 40 or FIG. 9. Alternatively, MVD1 may be the MVD of AMVP, inter mode, or affine inter mode. Alternatively, MVD1 may be the MVD when the merge mode is not used. Whether the merge mode is used or not may be indicated by merge_flag or general_merge_flag. For example, when merge_flag or general_merge_flag is 1, the merge mode is used, and when it is 0, the merge mode is not used.

[0462] According to one embodiment, MVD2 may be the MVD described in FIGS. 38 to 39. Alternatively, MVD2 may be the MVD in the MMVD mode. Alternatively, MVD2 may be the MVD when the merge mode is used.

[0463] According to one embodiment, the MV may be an MV used for final motion compensation or prediction. Or, the MV may be an MV that enters a candidate list. Or, the MV may be a collocated motion vector (temporal motion vector). Or, the MV may be a value obtained by adding the MVD to the MVP. Or, the MV may be a CPMV. Or, the MV may be a value obtained by adding the CPMVD to the CPMVP. Or, the MV may be an MV for each sub-block in affine MC. The MV for the sub-block may be an MV derived from the CPMV. Or, the MV may be the MV of the MMVD.

[0464] According to an embodiment of the present disclosure, the MVD2 range may be narrower than the MVD1 range or the MV range. Also, the MVD1 range and the MV range may be the same. According to one embodiment, the MVD1 range and the MV range may be from -2^17 to (2^17 - 1) (inclusive). Also, the MVD2 range may be from -2^15 to (2^15 - 1) (inclusive).

[0465] Referring to FIG. 37, the quadrilateral represented by the dotted line indicates the MV range or the MVD range. Also, the inner quadrilateral may indicate the MVD2 range. The outer quadrilateral may indicate the MVD1 range or the MV range. That is, the MVD2 range may be different from the MVD1 range or the MV range. The range in the drawing may indicate the vector range that can be indicated from the points in the drawing. For example, as described above, the MVD2 range may be from -2^15 to (2^15 - 1) (inclusive). Also, the MVD1 range or the MV range may be from -2^17 to (2^17 - 1) (inclusive). If MV or MVD is a value in x-pel units, the value indicated by MV or MVD may actually be (MV * x) or (MVD * x) pixels. For example, if it is a value in 1 / 16-pel units, it may indicate (MV / 16) or (MVD / 16) pixels. In this embodiment, in the case of the MVD2 range, the maximum absolute value is 32768. If this is a value in 1 / 16-pel units, the value of (MVD2 value) / 16 is at most 2048 pixels. Therefore, it cannot completely cover an 8K picture. In the case of the MVD1 range or the MV range, the maximum absolute value is 131072, and when using 1 / 16-pel units, it can indicate at most 8192 pixels. Therefore, it can completely cover an 8K picture. The 8K resolution can indicate a resolution where the horizontal length (or the length of the longer one of the horizontal or vertical) is 7680 pixels or 8192 pixels. For example, a picture such as 7680×4320 may be an 8K resolution.

[0466] FIG. 38 is a diagram showing MMVD according to an embodiment of the present disclosure.

[0467] According to an embodiment of the present disclosure, MMVD may be merge mode with MVD, merge with MVD, which is a method of using MVD in the merge mode. For example, an MV can be generated based on a merge candidate and MVD. Also, the MVD of MMVD may have a limited range that can be shown compared to the MVD of FIG. 9 or FIG. 40, or the aforementioned MVD1. For example, the MVD of MMVD may have only one of a horizontal component and a vertical component. Also, the absolute values of the values that the MVD of MMVD can show may not be equally spaced from each other. FIG. 38(a) shows the points that the MVD of MMVD can show from the central dotted-line-shaped point.

[0468] FIG. 38(b) shows MMVD-related syntax. According to one embodiment, there may be high-level signaling indicating whether MMVD is used. Here, signaling can mean parsing from a bitstream. The high level may be a unit including the current block and the current coding block, and may be, for example, a slice, a sequence, a tile, a tile group, a CTU, etc. The high-level signaling indicating whether MMVD is used may be high-level MMVD activation information (sps_mmvd_enabled_flag). When the high-level MMVD activation information (sps_mmvd_enabled_flag) is 1, it can indicate that MMVD is activated, and when the high-level MMVD activation information (sps_mmvd_enabled_flag) is 0, it can indicate that MMVD is not activated. However, it is not limited thereto. When the high-level MMVD activation information (sps_mmvd_enabled_flag) is 0, it can indicate that MMVD is activated, and when the high-level MMVD activation information (sps_mmvd_enabled_flag) is 1, it can indicate that MMVD is not activated.

[0469] When the high-level MMVD activity information (sps_mmvd_enabled_flag), which is high-level signaling indicating whether MMVD is used or not, is 1, it is possible to indicate whether MMVD is used for the current block by additional signaling. That is, when the high-level MMVD activity information indicates the activation of MMVD, the video signal processing device can perform a step of parsing MMVD merge information (mmvd_merge_flag) indicating whether to use MMVD for the current block from the bitstream. The signaling indicating whether MMVD is used or not may be the MMVD merge information (mmvd_merge_flag). When the MMVD merge information (mmvd_merge_flag) is 1, it can mean that MMVD is used for the current block. When the MMVD merge information (mmvd_merge_flag) is 0, it can mean that MMVD is not used for the current block. However, it is not limited to this. When the MMVD merge information (mmvd_merge_flag) is 0, it can mean that MMVD is used for the current block, and when the MMVD merge information (mmvd_merge_flag) is 1, it can mean that MMVD is not used for the current block.

[0470] If the MMVD merge information (mmvd_merge_flag) is 1 (a value indicating use), it is possible to parse MMVD-related syntax elements. The MMVD-related syntax elements can include at least one of mmvd_cand_flag, information related to the distance of MMVD (mmvd_distance_idx), and information related to the direction of MMVD (mmvd_direction_idx).

[0471] As an example, mmvd_cand_flag can indicate the MVP used in the MMVD mode. Or, mmvd_cand_flag can indicate the merge candidate used in the MMVD mode. Also, the MVD of MMVD may be determined based on the information related to the distance of MMVD (mmvd_distance_idx) and the information related to the direction of MMVD (mmvd_direction_idx). Here, the MVD of MMVD can indicate the information about MVD (mMvdLX). For example, the information related to the distance of MMVD (mmvd_distance_idx) can indicate a value related to the absolute value of the MVD of MMVD, and the information related to the direction of MMVD (mmvd_direction_idx) can indicate a value related to the direction of the MVD of MMVD.

[0472] FIG. 39 is a diagram showing the MVD derivation of MMVD according to an embodiment of the present disclosure.

[0473] Referring to the left side of FIG. 39, the video signal processing apparatus can determine whether MMVD is used for the current block based on the MMVD merge information (mmvd_merge_flag). Also, when MMVD is used for the current block, the MVD derivation process can be performed. The MVD derivation process may be 8.5.2.7 shown on the right side of FIG. 39. Also, the MVD that is the output of 8.5.2.7 may be the MVD of MMVD or the information about MMVD (mMvdLX). As already described, the information about MMVD (mMvdLX) can be obtained based on the information related to the distance of MMVD (mmvd_distance_idx) and the information related to the direction of MMVD (mmvd_direction_idx). Here, X may be replaced by 0 or 1, corresponding to reference list L0 and reference list L1 respectively.

[0474] Also, as shown in FIG. 39, information (mMvdLX) regarding the MVD of MMVD or MMVD can be added to mvLX, which is the motion vector (MV) derived in the previous process ((8-281), (8-282)). Here, X may be replaced by 0, 1, etc., and may correspond to the first reference list (reference list L0) and the second reference list (reference list L1), respectively. Also, here, the MVD of MMVD can indicate information (mMvdLX) regarding MMVD. Also, the video signal processing apparatus can perform a process of restricting the range of the modified motion vector (mvLX) to which the MVD is added. For example, clipping can be performed ((8-283), (8-284)). For example, it can be restricted to the above-described MV range. For example, it can be restricted from -2^17 to (2^17-1) (inclusive).

[0475] In the present disclosure, Clip3(x, y, z) may indicate clipping. For example, the result of Clip3(x, y, z) can be 1) x when z < x, 2) y when z > y, and 3) z in other cases. Therefore, the range of the result of Clip3(x, y, z), the result, may be x <= result <= y.

[0476] In the present embodiment, in mvLX[0][0][comp] and mMvdLX[comp], comp can indicate the x-component or the y-component. For example, they may indicate the horizontal component and the vertical component, respectively.

[0477] Also, the video signal processing apparatus can predict the current block based on the modified motion vector (mvLX). Also, the video signal processing apparatus can restore the current block based on the modified motion vector (mvLX). As already described, the modified motion vector (mvLX) may be a value obtained by adding information (mMvdLX) regarding MMVD to mvLX, which is the motion vector (MV).

[0478] Also, the MVD derivation process of MMVD may be the same as that on the right side of FIG. 39. The MMVD offset (MmvdOffset) in the drawing may be a value based on the above-described MMVD-related syntax elements. For example, the MMVD offset (MmvdOffset) may be a value based on information related to the distance of MMVD (mmvd_distance_idx) and information related to the direction of MMVD (mmvd_direction_idx). In short, the MMVD offset (MmvdOffset) may be obtained based on at least one of the information related to the distance (mmvd_distance_idx) and the information related to the direction of MMVD (mmvd_direction_idx). Also, information about MMVD (mMvdLX) may be derived based on the MMVD offset (MmvdOffset).

[0479] According to one embodiment, in the case of dual prediction (predFlagLX can indicate which reference list to use, X may be replaced by 0 or 1, and L0 and L1 may correspond to the first reference list (reference list L0) and the second reference list (reference list L1) respectively.), it can be determined whether to directly use the MMVD offset (MmvdOffset) as the MVD of MMVD based on the POC (picture order count), whether to use a value calculated based on the MMVD offset (MmvdOffset), whether to directly use the MMVD offset (MmvdOffset) for the value for a certain reference list, whether to use a value calculated based on the MMVD offset (MmvdOffset) for the value for a certain reference list, etc.

[0480] According to one embodiment, an MMVD offset (MmvdOffset) may be used to obtain the MVD for both the first reference list (reference list L0) and the second reference list (reference list L1) ((8 - 350) to (8 - 353)). For example, it may be the case where the first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1) are the same. The POC difference (currPocDiffLX) may be the difference between the POC of the current picture and the POC of the reference picture in the reference list LX, and X may be replaced by 0 and 1. The first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1) may be determined by (8 - 348) and (8 - 349) of FIG. 39, respectively. DiffPicOrderCnt can be calculated as follows.

[0481] DiffPicOrderCnt( picA, picB ) = PicOrderCnt( picA ) - PicOrderCnt( picB )

[0482] PicOrderCnt(picX) may indicate the POC value of the picture picX.

[0483] Also, in FIG. 39, currPic may indicate the current picture. Also, RefPicList[X][refIdxLX] may indicate the reference picture when using refIdxLX in the reference list LX.

[0484] According to an embodiment, for the MVD with respect to the second reference list (reference list L1), a value calculated based on the MmvdOffset may be used ((8-354) to (8-363)). At this time, the MMVD offset (MmvdOffset) value can be used as it is as the MVD with respect to the first reference list (reference list L0). In this case, Abs(currPocDiffL0) may be the same as or greater than Abs(currPocDiffL1). Or, Abs(currPocDiffL0) may be greater than Abs(currPocDiffL1).

[0485] According to an embodiment, for the MVD with respect to the first reference list (reference list L0), a value calculated based on the MMVD offset (MmvdOffset) may be used ((8-364) to (8-373)). At this time, the MMVD offset (MmvdOffset) value can be used as it is as the MVD with respect to the second reference list (reference list L1). In this case, Abs(currPocDiffL0) may be smaller than Abs(currPocDiffL1). Or, Abs(currPocDiffL0) may be smaller than or the same as Abs(currPocDiffL1).

[0486] According to one embodiment, performing an operation based on the MMVD offset (MmvdOffset) may indicate MV scaling. The MV scaling may be (8-356) to (3-361) or (8-366) to (3-371) in FIG. 39. The former is a scaling based on information (mMvdL0) regarding the first MMVD or the MMVD offset (MmvdOffset), and may be a process of creating information (mMvdL1) regarding the second MMVD which is the scaled MV. The latter is a scaling based on information (mMvdL1) regarding the second MMVD or the MMVD offset (MmvdOffset), and may be a process of creating information (mMvdL0) regarding the first MMVD which is the scaled motion vector (MV). The MV scaling may be an operation based on a scale factor (distScaleFactor ((8-359), (8-369))) which is a value based on the first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1). Also, the MV scaling may be an operation based on a value obtained by multiplying the scale factor (distScaleFactor) by the MV to be scaled ((8-360), (8-361), (8-370), (8-371)). Also, the MV scaling may include a process of restricting the MV range. For example, the range restriction process may be included in (8-360), (8-361), (8-370), (8-371) of FIG. 39. For example, the MV scaling may include a clipping process. According to an embodiment of the present disclosure, at this time, the above-described MVD2 range can be used. For example, the range can be restricted to from -2^15 to (2^15-1) (inclusive). That is, the MVD range of the MMVD can be restricted to from -2^15 to (2^15-1) (inclusive). For example, that is, the MVD derivation process of the MMVD can include a process of restricting a value based on a value obtained by multiplying the distScaleFactor by the MV to be scaled to from -2^15 to (2^15-1) (inclusive). For example, the Clip3(-2^15, 2^15-1, x) process may be included in the MVD derivation process of the MMVD.More specifically, mMvdLX may be determined as follows.

[0487] mMvdLX = Clip3(-2^15, 2^15 - 1, (distScaleFactor*mMvdLY + 128 - (distScaleFactor*mMvdLY >=0 )) >> 8 )

[0488] At this time, Y may be 0 or 1 and may be!X. Also, mMvdLY may be MmvdOffset. distScaleFactor may be a value shown in (8-359) or (8-369) and may be a value based on currPocDiffL0 and currPocDiffL1. However, the MVD of MMVD, the information regarding MMVD (mMvdLX), or the range of the scaled MV is not limited to -2^15 to (2^15-1). The MVD of MMVD, the information regarding MMVD (mMvdLX), or the range of the scaled MV may be -2^17 to (2^17-1). This will be described with reference to FIG. 43.

[0489] This process can be performed when neither the reference pictures for L0 and L1 are long-term reference pictures.

[0490] According to one embodiment, the operation based on the MMVD offset (MmvdOffset) may indicate the MMVD offset (MmvdOffset) or the negative MMVD offset (-MmvdOffset) based on the first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1) ((8-362), (8-363), (8-372), (8-373)). This process can be performed when at least one of the reference pictures for L0 and L1 is a long-term reference picture.

[0491] Also, in the case of uni-prediction, it is possible to directly use MmvdOffset with mMvdLX ((8-374), (8-375)).

[0492] FIG. 40 is a diagram showing MVD and MV derivation according to an embodiment of the present disclosure.

[0493] The embodiment of FIG. 40 may be an example using the MV range or MVD1 range of FIG. 37. The embodiment of FIG. 40 may be an example using a range from -2^17 to (2^17 - 1) (inclusive).

[0494] Referring to FIG. 40, lMvd can be derived based on the syntax elements described in FIG. 9 etc. (7-161). Also, lMvd may be in the range from -2^17 to (2^17 - 1) (inclusive). Also, values such as MvdLX and MvdCpLX may be set to lMvd. Also, lMvd, MvdLX, MvdCpLX, etc. may be the aforementioned MVD1.

[0495] Also, when using AMVP, inter mode, etc., the MV range can be restricted to -2^17 to (2^17 - 1) (inclusive). In this case, it may be a case where the merge mode is not used. Referring to FIG. 40, (8-264) to (8-267) may include the MV range restriction process. Referring to the drawing, uLX may be a value based on the sum of MVP (mvpLX) and MVD (mvdLX). Also, based on ULX, the final MV, mvLX, can be calculated. FIG. 40 includes a modulus operation, but the value indicated by this operation may be restricted so as to be expressed in a certain number of bits or less. For example, it can be made possible to be expressed in 18 bits by calculating (%2^18). Therefore, the MV range can be restricted to -2^17 to (2^17 - 1) (inclusive). (8-265), (8-267) can play a role in resolving the overflow that may occur when adding MVP and MVD.

[0496] Figure 41 is a diagram showing MV, CPMV derivation according to an embodiment of the present disclosure.

[0497] The embodiment of Figure 41 may be an example using the MV range or MVD1 range of Figure 37. The embodiment of Figure 40 may be an example using the range from -2^17 to (2^17 - 1) (inclusive).

[0498] Figure 41(a) may show a method of deriving a collated motion vector (temporal motion vector). This process may include MV scaling. At this time, it may include an operation of clipping a value based on distScaleFactor and the MV (mvCol) to be scaled ((8 - 398)). At this time, clipping can be performed so that the result is in the range of -131072 to 131071 (inclusive). -131072 is the same value as -2^17. 131071 is the same value as 2^17 - 1.

[0499] Figure 41(b) is a diagram showing a part of the process of deriving CPMV (cpMvLX). Similarly, a clipping process for restricting the range may be included at this time. At this time, the range may be from -2^17 to (2^17 - 1) (inclusive).

[0500] Figure 41(c) is a diagram showing a part of the process of deriving a sub-block-based MV. The sub-block-based MV may be an MV when using affine MC, sub-block-based temporal MV prediction, etc. Similarly, a clipping process for restricting the range may be included at this time. At this time, the range may be from -2^17 to (2^17 - 1) (inclusive). The xSbIdx and ySbIdx in the drawing can respectively indicate the sub-block index for the x-axis and the sub-block index for the y-axis.

[0501] FIG. 42 is a diagram showing the ranges of MV and MVD according to an embodiment of the present disclosure.

[0502] The above-described part may be omitted in this embodiment.

[0503] As described above, there may be a number of MVD generation methods or MVD determination methods. The MVD may be the above-described motion vector difference value. Also, there may be a number of MV generation methods or MV determination methods.

[0504] And at this time, according to an embodiment of the present disclosure, the range that the MV or MVD can indicate may be restricted.

[0505] According to an embodiment of the present disclosure, the MV range and the MVD range may be the same. Also, there is an MVD generation method, the MVD generated by the MVD generation method 1 is MVD1, and the MVD generated by the MVD generation method 2 may be MVD2. At this time, the MVD1 range and the MVD2 range may be the same. Also, the MV range, the MVD range, the MVD1 range, and the MVD2 range may be from -2^N to (2^N - 1). At this time, it may be a range including -2^N and (2^N - 1). Further, according to another embodiment, the MV range or the MVD range may be from (-2^N + 1) to 2^N. At this time, it may be a range including (-2^N + 1) and 2^N. In these embodiments, N is an integer, for example, a positive integer. More specifically, N may be 17. Also, at this time, the MV range or the MVD range can be indicated using N + 1 bits. Therefore, the MV range, the MVD range, the MVD1 range, and the MVD2 range may all be from -2^17 to (2^17 - 1).

[0506] Also, as described in the previous drawings, MV, MVD1, MVD2, etc. may be defined. That is, MV can mean an MV used for final motion compensation or prediction, etc., and reference can be made to the description in FIG. 37 for this. Also, MVD1 may be the MVD described in FIG. 40 or FIG. 9. Or, MVD1 may be the MVD of AMVP, inter mode, affine inter mode, or the MVD when the merge mode is not used, and reference can be made to the description in FIG. 37 for this. Also, MVD2 may be the MVD described in FIGS. 38 to 39, or the MVD of the MMVD mode, and reference can be made to the description in FIG. 37 for this. Here, the MVD of the MMVD mode may be information (mMvdLX) regarding MMVD.

[0507] Therefore, according to one embodiment, the MVDs of AMVP, inter mode, affine inter mode, and MMVD may all have the same representable range. Furthermore, the final MV may also have the same range. More specifically, the MVDs of AMVP, inter mode, affine inter mode, and MMVD may all have a range from -2^17 to (2^17 - 1) (inclusive). Furthermore, the final MV may also have a range from -2^17 to (2^17 - 1) (inclusive).

[0508] Thereby, a picture of a certain size can be completely covered by MV and MVD. Also, when there is a range that a certain MV or MVD1 can represent, MVD2 can also represent the same range. Therefore, for example, the method of using MVD2 does not have to be restricted compared to other methods depending on the representation range. Also, when MV, MVD1, etc. can represent any range, making MVD2 represent the same range does not require additional resources in hardware or software. If the MVD of MVD2 or MMVD can represent from -2^17 to (2^17 - 1) (inclusive), the maximum absolute value is 131072, and when this is a value in 1 / 16 - pel units, it can show up to 8192 pixels at most. Therefore, an 8K picture can be completely covered.

[0509] Referring to FIG. 42, the outer dotted line indicates the range that MV, MVD1, and MVD2 can indicate from the central point. And the MV range, MVD1 range, and MVD2 range may all be the same. For example, the MV range, MVD1 range, and MVD2 range may be in the range from -2^17 to (2^17 - 1) (inclusive). Also, this range can include all 8K pictures. That is, in the worst case, even when showing from end to end of the picture, the MV, MVD1, and MVD2 ranges can indicate it. Therefore, this can show a better motion vector, and it can be said that motion compensation is improved, the residual is reduced, and the coding efficiency can be increased.

[0510] FIG. 43 is a diagram showing the MVD derivation of MMVD according to an embodiment of the present disclosure.

[0511] FIG. 43 may be a modification of a part of FIG. 39. In the embodiment of FIG. 43, the content described in FIG. 39 may be omitted. Or, the above-described content may be omitted.

[0512] When using MMVD, the MVD derivation process can be performed. The MVD derivation process may be 8.5.2.7 shown in FIG. 43. Also, the MVD that is the output of 8.5.2.7 may be the MVD of MMVD and may be the information (mMvdLX) regarding MMVD. As already described, the information (mMvdLX) regarding MMVD may be obtained based on the information (mmvd_distance_idx) related to the distance of MMVD and the information (mmvd_direction_idx) related to the direction of MMVD.

[0513] More specifically, it will be described together with FIG. 43. Referring to FIG. 43, the video signal processing apparatus can obtain an MMVD offset (MmvdOffset). The video signal processing apparatus can obtain an MMVD offset (MmvdOffset) to obtain information (mMvdLX) regarding MMVD. The MMVD offset (MmvdOffset) may be a value based on the MMVD-related syntax elements described above. For example, the MMVD offset (MmvdOffset) may be a value based on information (mmvd_distance_idx) related to the distance of MMVD and information (mmvd_direction_idx) related to the direction of MMVD.

[0514] According to an embodiment, the video signal processing apparatus can determine whether the current block is predicted by dual prediction based on predFlagLX. predFlagLX can indicate which reference list is used. Here, X may be replaced with 0 or 1, and L0 and L1 may correspond to reference lists L0 and L1, respectively. When predFlagL0 is 1, it can be indicated that the first reference list is used. Also, when predFlagL0 is 0, it can be indicated that the first reference list is not used. When predFlagL1 is 1, it can be indicated that the second reference list is used. Also, when predFlagL1 is 0, it can be indicated that the second reference list is not used.

[0515] In FIG. 43, when both predFlagL0 and predFlagL1 are 1, dual prediction can be indicated. Dual prediction can indicate that both the first reference list and the second reference list are used. When the first reference list and the second reference list are used, the video signal processing apparatus can perform a step (8-348) of obtaining the difference in POC (Picture Order Count) between the current picture (currPic) including the current block and the first reference picture (RefPicList[0][refIdxL0]) based on the first reference list as the first POC difference (currPocDiffL0). Also, when the first reference list and the second reference list are used, the video signal processing apparatus can perform a step (8-349) of obtaining the difference in POC (Picture Order Count) between the current picture (currPic) and the second reference picture (RefPicList[1][refIdxL1]) based on the first list as the second POC difference (currPocDiffL1).

[0516] The POC difference (currPocDiffLX) may be the difference between the POC of the current picture and the POC of the reference picture in the reference list (reference list LX), and X may be replaced by 0 and 1. The first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1) may be determined by (8-348) and (8-349) in FIG. 43, respectively. DiffPicOrderCnt can be calculated as follows.

[0517] DiffPicOrderCnt( picA, picB ) = PicOrderCnt( picA ) - PicOrderCnt( picB )

[0518] PicOrderCnt(picX) may indicate the POC value of the picture picX.

[0519] The video signal processing apparatus can perform a step of obtaining information (mMvdL0) regarding a first MMVD related to a first reference list and information (mMvdL1) regarding a second MMVD related to a second reference list based on at least one of an MMVD offset (MmvdOffset), a first POC difference (currPocDiffL0), and a second POC difference (currPocDiffL1). Here, the information (mMvdLX) regarding the MMVD can include the information (mMvdL0) regarding the first MMVD and the information (mMvdL1) regarding the second MMVD. The process by which the video signal processing apparatus obtains the information (mMvdL0) regarding the first MMVD and the information (mMvdL1) regarding the second MMVD will be described in detail below.

[0520] The video signal processing apparatus can determine whether to directly use the MMVD offset (MmvdOffset) as the MVD of the MMVD based on the POC (picture order count), use a value calculated based on the MMVD offset (MmvdOffset), directly use the MMVD offset (MmvdOffset) for a value for a certain reference list, or use a value calculated based on the MMVD offset (MmvdOffset) for a value for a certain reference list.

[0521] When the first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1) are the same, the video signal processing apparatus can perform a step (8-350, 8-351) of obtaining the MMVD offset (MmvdOffset) as the information (mMvdL0) regarding the first MMVD. Also, when the first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1) are the same, the video signal processing apparatus can perform a step (8-352, 8-353) of obtaining the MMVD offset (MmvdOffset) as the information (mMvdL1) regarding the second MMVD.

[0522] The video signal processing apparatus can determine whether the absolute value of the first POC difference (Abs(currPocDiffL0)) is greater than the absolute value of the second POC difference (Abs(currPocDiffL1)). Also, when the absolute value of the first POC difference (Abs(currPocDiffL0)) is greater than or equal to the absolute value of the second POC difference (Abs(currPocDiffL1)), the video signal processing apparatus can perform steps (8-354, 8-355) of obtaining the MMVD offset (MmvdOffset) as information (mMvdL0) regarding the first MMVD.

[0523] When the first reference picture (RefPicList[0][refIdxL0]) is not a long-term reference picture and the second reference picture (RefPicList[1][refIdxL1]) is not a long-term reference picture, the video signal processing apparatus can perform steps (8-356 to 8-361) of scaling the information (mMvdL0) regarding the first MMVD to obtain the information (mMvdL1) regarding the second MMVD. A scale factor (distScaleFactor) may be used for scaling. The scale factor (distScaleFactor) may be obtained based on at least one of the first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1).

[0524] When the first reference picture is a long-term reference picture or the second reference picture is a long-term reference picture, the video signal processing apparatus can perform a step (8-362, 8-363) of obtaining information (mMvdL1) regarding the second MMVD without scaling the absolute value of the information (mMvdL0) regarding the first MMVD. Here, not scaling may mean not changing the absolute value of the information (mMvdL0) regarding the first MMVD. That is, the video signal processing apparatus can obtain the information (mMvdL1) regarding the second MMVD with or without changing the sign of the information (mMvdL0) regarding the first MMVD. The video signal processing apparatus can set the information (mMvdL1) regarding the second MMVD to the information (mMvdL0) regarding the first MMVD when the signs of the first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1) are the same. Also, the video signal processing apparatus can set the information (mMvdL1) regarding the second MMVD by changing the sign of the information (mMvdL0) regarding the first MMVD when the signs of the first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1) are different.

[0525] The video signal processing apparatus can determine whether the absolute value of the first POC difference (Abs(currPocDiffL0)) is smaller than the absolute value of the second POC difference (Abs(currPocDiffL1)). When the absolute value of the first POC difference (Abs(currPocDiffL0)) is smaller than the absolute value of the second POC difference (Abs(currPocDiffL1)), the video signal processing apparatus can perform a step (8-364, 8-365) of obtaining the MMVD offset (MmvdOffset) as the information (mMvdL1) regarding the second MMVD.

[0526] When the first reference picture (RefPicList[0][refIdxL0]) is not a long-term reference picture and the second reference picture (RefPicList[1][refIdxL1]) is not a long-term reference picture, the video signal processing apparatus can perform steps (8-366 to 8-371) of scaling the information (mMvdL1) related to the second MMVD and obtaining the information (mMvdL0) related to the first MMVD. A scale factor (distScaleFactor) may be used for scaling. The scale factor (distScaleFactor) may be obtained based on at least one of the first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1).

[0527] When the first reference picture (RefPicList[0][refIdxL0]) is a long-term reference picture or the second reference picture (RefPicList[1][refIdxL1]) is a long-term reference picture, the video signal processing apparatus can perform steps (8-372, 8-373) of obtaining the information (mMvdL0) related to the first MMVD without scaling the absolute value of the information (mMvdL1) related to the second MMVD. Here, not scaling may mean not changing the absolute value of the information (mMvdL1) related to the second MMVD. That is, the video signal processing apparatus can obtain the information (mMvdL0) related to the first MMVD with or without changing the sign of the information (mMvdL1) related to the second MMVD. The video signal processing apparatus can set the information (mMvdL0) related to the first MMVD to the information (mMvdL1) related to the second MMVD when the signs of the first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1) are the same. Also, the video signal processing apparatus can set the information (mMvdL0) related to the first MMVD by changing the sign of the information (mMvdL1) related to the second MMVD when the signs of the first POC difference (currPocDiffL0) and the second POC difference (currPocDiffL1) are different.

[0528] In FIG. 43, when at least one of predFlagL0 and predFlagL1 is not 1, uni-prediction can be indicated. When it is uni-prediction, the MMVD offset (MmvdOffset) can be directly used as the information regarding MMVD (mMvdLX) ((8-374), (8-375)).

[0529] As already described, the video signal processing apparatus can perform a step of obtaining the MMVD offset based on the information related to the distance of MMVD and the information related to the direction of MMVD in order to obtain the information regarding MMVD (mMvdLX).

[0530] When only the first reference list is used (predFlagL0 == 1), the video signal processing apparatus can perform a step (8-374, 8-375) of obtaining the MMVD offset (MmvdOffset) without scaling it and using it as the information regarding the first MMVD (mMvdL0) related to the first reference list.

[0531] When only the second reference list is used (predFlagL1 == 1), the video signal processing apparatus can perform a step (8-374, 8-375) of obtaining the MMVD offset (MmvdOffset) without scaling it and using it as the information regarding the second MMVD (mMvdL1) related to the second reference list.

[0532] As described above, the range of the MVD of MMVD may be restricted. Here, the MVD of MMVD may be the information regarding MMVD (mMvdLX). According to an embodiment of the present disclosure, the range of the MVD of MMVD may be the same as the range of the MVD determined by other methods (for example, the MVD based on the syntax element in FIG. 9). Also, the range of the MVD of MMVD may be the same as the range of the final MV. According to an embodiment of the present disclosure, the range of the MVD of MMVD or the range of the information regarding MMVD (mMvdLX) may be from -2^17 to 2^17 - 1, and at this time, it may be a range including -2^17 and 2^17 - 1.

[0533] Also, as described above, the MV scaling process can be included in the MVD derivation process of MMVD. Also, a clipping operation for restricting a range may be included in the MVD derivation process of MMVD.

[0534] Referring to FIG. 43, the information (mMvdLX) regarding MMVD may be based on the clipping operation. At this time, X may be replaced with 0 or 1, which may correspond to the first reference list (reference list L0) and the second reference list (reference list L1), respectively. For example, the information (mMvdLX) regarding MMVD may be based on the Clip3 operation. For example, the information (mMvdLX) regarding MMVD may be based on Clip3(-2^17, 2^17-1, x). At this time, x may be a value based on the first POC difference (currPocDiffL0), the second POC difference (currPocDiffL1), or the MMVD offset (MmvdOffset). For example, x may be a value based on the scale factor (distScaleFactor) and the MMVD offset (MmvdOffset). (8-360), (8-361), (8-370), and (8-371) in FIG. 43 include a clipping operation for restricting the MVD range.

[0535] As a further example, in FIGS. 42 to 43, the unified range among MV, MVD, MVD1, and MVD2, the MVD by a number of MVD generation methods, and the MVD range of MMVD may be from -2^N + 1 to 2^N, and at this time, it may be a range including -2^N + 1 and 2^N. More specifically, N may be 17. That is, the unified range among MV, MVD, MVD1, and MVD2, the MVD by a number of MVD generation methods, and the MVD range of MMVD may be from -2^17 + 1 to 2^17, and at this time, it may be a range including -2^17 + 1 and 2^17. In this case, the MVD can be represented by 18 bits.

[0536] According to an embodiment of the present disclosure, the range of the chroma MV may be different from the range of the luma MV. For example, the chroma MV may have a higher resolution than the luma MV. That is, one unit of the chroma MV can represent a smaller pixel than one unit of the luma MV. For example, the luma MV may be in 1 / 16-pel units. Also, the chroma MV may be in 1 / 32-pel units.

[0537] Also, the chroma MV may be based on the luma MV. For example, the chroma MV may be determined by multiplying any value to the luma MV. The any value may be 2 / SubWidthC for the horizontal component and 2 / SubHeightC for the vertical component. The 2 included in the any value may be a value included because the resolution of the chroma MV is twice as high as the resolution of the luma MV. Also, SubWidthC and SubHeightC may be values determined by a color format. Also, SubWidthC and SubHeightC may be values for chroma sample sampling. For example, SubWidthC and SubHeightC may be values related to how many chroma samples exist for a luma sample. Also, SubWidthC and SubHeightC may be 1 or 2.

[0538] According to an embodiment of the present disclosure, when the resolution of the chroma MV is higher than the resolution of the luma MV, the range of the chroma MV may be wider than the range of the luma MV. This may be for covering the area covered by the luma MV with the chroma MV. For example, the range of the chroma MV may be twice the range of the luma MV. For example, the chroma MV range may be from -2^18 to 2^18 - 1, and at this time, it may be a range including -2^18 and 2^18 - 1. Also, the luma MV range may be from -2^17 to 2^17 - 1, and at this time, it may be a range including -2^17 and 2^17 - 1.

[0539] Also, the range of the chroma MV may vary depending on the color format. For example, when SubWidthC or SubHeightC is 1, the range of the chroma MV may be different from the range of the luma MV. At this time, the range of the chroma MV may be twice the range of the luma MV. For example, the chroma MV range may be from -2^18 to 2^18 - 1, and at this time, it may be a range including -2^18 and 2^18 - 1, and the luma MV range may be from -2^17 to 2^17 - 1, and at this time, it may be a range including -2^17 and 2^17 - 1. Also, when SubWidthC and SubHeightC are 2, the range of the chroma MV may be the same as the range of the luma MV. At this time, the range may be from -2^17 to 2^17 - 1, and at this time, it may be a range including -2^17 and 2^17 - 1. This may be for unifying the ranges that the chroma MV and the luma MV can represent.

[0540] In the embodiments of the present disclosure, in the case of the 4:2:0 format, SubWidthC and SubHeightC may be 2 and 2 respectively. In the case of the 4:2:2 format, SubWidthC and SubHeightC may be 2 and 1 respectively. In the case of the 4:4:4 format, SubWidthC and SubHeightC may be 1 and 1 respectively.

[0541] The embodiments of the present disclosure described above may be implemented by various means. For example, the embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof.

[0542] In the implementation by hardware, the method according to the embodiments of the present disclosure may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, etc.

[0543] In the implementation by firmware or software, the method according to the embodiments of the present disclosure may be implemented in the form of modules, procedures, functions, etc. that execute the functions or operations described above. The software code may be stored in a memory and driven by a processor. The memory may be arranged inside or outside the processor and can exchange data with the processor by various means already known.

[0544] The above description of the present disclosure is for illustrative purposes, and those of ordinary skill in the art to which the present disclosure pertains will understand that it can be easily deformed into another specific form without changing the technical idea and essential features of the present disclosure. Therefore, the embodiments described above should be construed as illustrative in all respects and not restrictive. For example, each component described as a single type may be implemented distributively, and similarly, components described as being distributed may also be implemented in a combined form.

[0545] The scope of the present disclosure is indicated by the claims described below rather than the above detailed description, and any changes or modifications derived from the meaning and scope of the claims and the equivalent concept thereof should be construed as being included within the scope of the present disclosure.

Description of Reference Numerals

[0546] 110 Conversion Unit 115 Quantization Unit 120 Inverse Quantization Unit 125 Inverse Transformation Unit 130 Filtering Unit 150 Prediction Unit 152 Intra Prediction Unit 154 Inter Prediction Unit 154a Motion Estimation Unit 154b Motion Compensation Unit 160 Entropy Coding Unit 210 Entropy Decoding Unit 220 Inverse Quantization Unit 225 Inverse Transformation Unit 230 Filtering Unit 250 Prediction Unit 252 Intra Prediction Unit 254 Inter Prediction Unit 810 Encoder 820 Decoder

Claims

1. Obtaining chroma component format information from a higher-level bitstream; Obtaining information regarding width (SubWidthC) and information regarding height (SubHeightC) based on the chroma component format information; Obtaining x-axis scale information based on the information regarding width or information regarding the color component of the current block; Obtaining y-axis scale information based on the information regarding height or information regarding the color component of the current block; Determining the position of the left block based on the y-axis scale information; Determining the position of the upper block based on the x-axis scale information; Determining a weighting value based on the left block and the upper block; Obtaining a first sample by predicting the current block in merge mode; Obtaining a second sample by predicting the current block in intra mode; and Obtaining a combined prediction sample for the current block based on the weighting value, the first sample, and the second sample, the method for decoding a video signal being characterized by including these steps.

2. The step of determining the weighting value includes setting code information (isIntraCodedNeighbourA) regarding the left block to TRUE when the left block is available and the prediction mode of the left block is intra prediction; setting the code information regarding the left block to FALSE when the left block is not available or the prediction mode of the left block is not intra prediction; setting code information (isIntraCodedNeighbourB) regarding the upper block to TRUE when the upper block is available and the prediction mode of the upper block is intra prediction; and setting the code information regarding the upper block to FALSE when the upper block is not available or the prediction mode of the upper block is not intra prediction, the method for decoding a video signal according to Claim 1 being characterized by including these steps.

3. The step of determining the weighting value includes determining the weighting value to be 3 when both the code information regarding the left block and the code information regarding the upper block are TRUE; When both the code information regarding the left block and the code information regarding the upper block are FALSE, determining the weighted value as 1; and When only one of the code information regarding the left block and the code information regarding the upper block is TRUE, including the step of determining the weighted value as 2. A method for decoding a video signal according to claim 2, characterized in that it comprises the above steps.

4. The step of obtaining the combined prediction sample is predSamplesComb[x][y] = (w * predSamplesIntra[x][y] + (4 - w) * predSamplesInter[x][y] + 2) >> 2 including the step of predicting the current block based on this, where predSamplesComb means the combined prediction sample, w means the weighted value, predSamplesIntra means the second sample, predSamplesInter means the first sample, [x] means the x-axis coordinate of the sample included in the current block, and [y] means the y-axis coordinate of the sample included in the current block. A method for decoding a video signal according to claim 1, characterized in that it has the above features.

5. The step of obtaining the scale information of the x-axis is When the information regarding the color component of the current block is 0 or the information regarding the width is 1, determining the scale information of the x-axis as 0; and When the information regarding the color component of the current block is not 0 and the information regarding the width is not 1, including the step of determining the scale information of the x-axis as 1, The step of obtaining the scale information of the y-axis is When the information regarding the color component of the current block is 0 or the information regarding the height is 1, determining the scale information of the y-axis as 0; and When the information regarding the color component of the current block is not 0 and the information regarding the height is not 1, including the step of determining the scale information of the y-axis as 1. A method for decoding a video signal according to claim 1, characterized in that it has the above features.

6. The position of the left block is (xCb - 1, yCb - 1 + (cbHeight << scallFactorHeight)) The xCb is the x-axis coordinate of the left-upper sample of the current luma block, the yCb is the y-axis coordinate of the left-upper sample of the current luma block, the cbHeight is the size of the height of the current block, and the scaleFactHeight is the scale information of the y-axis. The position of the upper block is (xCb - 1 + (cbWidth << scaleFactWidth), yCb - 1). The method for decoding a video signal according to claim 1, wherein the xCb is the x-axis coordinate of the left-upper sample of the current luma block, the yCb is the y-axis coordinate of the left-upper sample of the current luma block, the cbWidth is the size of the width of the current block, and the scaleFactWidth is the scale information of the x-axis.

7. The apparatus for decoding a video signal includes a processor and a memory. Based on the instruction words stored in the memory, the processor obtains chroma component format information from a higher-level bitstream, obtains information related to width (SubWidthC) and information related to height (SubHeightC) based on the chroma component format information, obtains scale information of the x-axis based on the information related to width or the information related to the color component of the current block, obtains scale information of the y-axis based on the information related to height or the information related to the color component of the current block, determines the position of the left block based on the scale information of the y-axis, determines the position of the upper block based on the scale information of the x-axis, determines a weighting value based on the left block and the upper block, obtains a first sample by predicting the current block in merge mode, obtains a second sample by predicting the current block in intra mode, The apparatus for decoding a video signal, characterized in that a combined prediction sample for the current block is obtained based on the weighting value, the first sample, and the second sample.

8. Based on the instruction words stored in the memory, the processor sets the code information (isIntraCodedNeighbourA) related to the left block to TRUE when the left block is available and the prediction mode of the left block is intra prediction. When the left block is not available or the prediction mode of the left block is not intra prediction, set the code information regarding the left block to FALSE. When the upper block is available and the prediction mode of the upper block is intra prediction, set the code information (isIntraCodedNeighbourB) regarding the upper block to TRUE. The apparatus for decoding a video signal according to claim 7, characterized in that when the upper block is not available or the prediction mode of the upper block is not intra prediction, the code information regarding the upper block is set to FALSE.

9. Based on the instruction words stored in the memory, the processor When both the code information regarding the left block and the code information regarding the upper block are TRUE, determine the weighting value to be 3. When both the code information regarding the left block and the code information regarding the upper block are FALSE, determine the weighting value to be 1. The apparatus for decoding a video signal according to claim 8, characterized in that when only one of the code information regarding the left block and the code information regarding the upper block is TRUE, determine the weighting value to be 2.

10. Based on the instruction words stored in the memory, the processor Predict the current block based on predSamplesComb[x][y] = (w * predSamplesIntra[x][y] + (4 - w) * predSamplesInter[x][y] + 2) >> 2 where predSamplesComb means the combined prediction samples, w means the weighting value, predSamplesIntra means the second samples, predSamplesInter means the first samples, [x] means the x-axis coordinate of the samples included in the current block, and [y] means the y-axis coordinate of the samples included in the current block. The apparatus for decoding a video signal according to claim 7, characterized in that

11. Based on the instruction words stored in the memory, the processor When the information regarding the color component of the current block is 0 or the information regarding the width is 1, determine the scale information of the x-axis to be 0. When the information regarding the color component of the current block is not 0 and the information regarding the width is not 1, determine the scale information of the x-axis as 1. When the information regarding the color component of the current block is 0 or the information regarding the height is 1, determine the scale information of the y-axis as 0. The apparatus for decoding a video signal according to claim 7, characterized in that when the information regarding the color component of the current block is not 0 and the information regarding the height is not 1, determine the scale information of the y-axis as 1.

12. The position of the left block is (xCb - 1, yCb - 1 + (cbHeight << scallFactHeight)). The xCb is the x-axis coordinate of the left-upper sample of the current luma block, the yCb is the y-axis coordinate of the left-upper sample of the current luma block, the cbHeight is the size of the height of the current block, and the scallFactHeight is the scale information of the y-axis. The position of the upper block is (xCb - 1 + (cbWidth << scallFactWidth), yCb - 1). The apparatus for decoding a video signal according to claim 7, characterized in that the xCb is the x-axis coordinate of the left-upper sample of the current luma block, the yCb is the y-axis coordinate of the left-upper sample of the current luma block, the cbWidth is the size of the width of the current block, and the scallFactWidth is the scale information of the x-axis.

13. Generating upper-level chroma component format information; Obtaining information regarding the width (SubWidthC) and information regarding the height (SubHeightC) based on the chroma component format information; Obtaining scale information of the x-axis based on the information regarding the width or the information regarding the color component of the current block; Obtaining scale information of the y-axis based on the information regarding the height or the information regarding the color component of the current block; Determining the position of the left block based on the scale information of the y-axis; Determining the position of the upper block based on the scale information of the x-axis; Determining a weighting value based on the left block and the upper block; Obtaining a first sample by predicting the current block in merge mode; Obtaining a second sample by predicting the current block in the intra mode; A method for encoding a video signal, comprising: obtaining a combined prediction sample for the current block based on the weighting value, the first sample, and the second sample.