Video signal processing method and apparatus using multi-assumption prediction

By constructing a merge candidate list with HMVP and applying combined inter and intra prediction, the method addresses inefficiencies in video signal processing, improving coding efficiency and compression.

JP2025106503AActive Publication Date: 2025-07-15WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025065101
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-09-06
Filing Date
2025-04-10
Publication Date
2025-07-15
Estimated Expiration
2039-10-14

AI Technical Summary

Technical Problem

Existing video signal processing methods lack efficiency in coding, particularly in handling spatial and temporal correlations, necessitating improved techniques for enhanced compression.

Method used

The method involves constructing a merge candidate list using spatial candidates and incorporating history-based motion vector predictors (HMVP) to predict current blocks, with specific updates to the HMVP table based on the decoding order and block location, and employing combined inter and intra prediction modes.

Benefits of technology

This approach increases coding efficiency by utilizing a conversion kernel suitable for current blocks, enhancing compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106503000001_ABST
    Figure 2025106503000001_ABST
Patent Text Reader

Abstract

To provide a video signal processing method and apparatus for encoding or decoding a video signal.SOLUTION: A video signal processing method comprises the steps of: receiving information for prediction of a current block; determining whether a merge mode is applied to the current block on the basis of the information for prediction; when a merge mode is applied to the current block, obtaining a first syntax element indicating whether a combined prediction is applied to the current block; generating an inter-prediction block and an intra-prediction block of the current block when the first syntax element indicates that the combined prediction is applied to the current block; and generating a combined prediction block of the current block by weighted-summing the inter-prediction block and the intra-prediction block.SELECTED DRAWING: Figure 45
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video signal processing method and apparatus, and more particularly, to a video signal processing method and apparatus for encoding or decoding a video signal.

Background Art

[0002] Compression encoding refers to a series of signal processing techniques for transmitting digitized information via a communication line or storing it in a form suitable for a storage medium. The targets of compression encoding include audio, video, characters, etc. In particular, the technique of performing compression encoding on video is called video compression. Compression encoding of a video signal is performed by removing redundant information in consideration of spatial correlation, temporal correlation, probabilistic correlation, etc. However, due to the recent development of various media and data transmission media, a more efficient video signal processing method and apparatus are required.

Summary of the Invention

Problems to be Solved by the Invention

[0003] An object of the present invention is to increase the coding efficiency of a video signal. Specifically, the present invention has an object to increase the coding efficiency by using a conversion kernel suitable for a conversion block.

Means for Solving the Problems

[0004] In order to solve the above problems, the present invention provides a video signal processing apparatus and a video signal processing method as follows.

[0005] According to an embodiment of the present invention, there is provided a video signal processing method, comprising: constructing a merge candidate list using spatial candidates; adding a specific history-based motion vector predictor (HMVP) in an HMVP table including at least one history-based motion vector predictor (HMVP) to the merge candidate list, wherein the HMVP indicates motion information of a block encoded before the plurality of coding blocks; obtaining index information indicating a merge candidate used for prediction of a current block in the merge candidate list; and generating a predicted block of the current block based on motion information of the merge candidate determined based on the index information. When the current block is located in a merge sharing node including a plurality of coding blocks, the merge candidate list is constructed using spatial candidates adjacent to the merge sharing node, and motion information of at least one coding block among the plurality of coding blocks included in the merge sharing node is not updated in the HMVP table.

[0006] Also, according to an embodiment of the present invention, there is provided a video signal processing apparatus including a processor, wherein the processor configures a merge candidate list using spatial candidates, adds a specific HMVP in an HMVP table including at least one history-based motion vector predictor (HMVP) to the merge candidate list, wherein the HMVP indicates motion information of a block encoded before the plurality of coding blocks, obtains index information indicating a merge candidate used for prediction of a current block in the merge candidate list, generates a predicted block of the current block based on the motion information of the merge candidate determined based on the index information, and when the current block is located in a merge sharing node including a plurality of coding blocks, the merge candidate list is configured using spatial candidates adjacent to the merge sharing node, and motion information of at least one coding block among the plurality of coding blocks included in the merge sharing node is not updated in the HMVP table.

[0007] As an example, the method may further include updating the HMVP table using motion information of a predefined number of coding blocks having a relatively late decoding order among the plurality of coding blocks included in the merge sharing node.

[0008] As an example, the method may further include updating the HMVP table using motion information of a coding block having the relatively latest decoding order among the plurality of coding blocks included in the merge sharing node.

[0009] As an example, when the current block is not located within the merge shared node, the method may further include updating the HMVP table using the motion information of the merge candidate.

[0010] As an example, the step of adding the HMVP to the merge candidate list may include: checking whether the HMVP having a specific index defined in advance in the HMVP table has motion information overlapping with candidates in the merge candidate list; and adding the HMVP having the specific index to the merge candidate list when the HMVP having the specific index does not have motion information overlapping with candidates in the merge candidate list.

[0011] Also, according to an embodiment of the present invention, there is provided a video signal processing method, including: receiving information for prediction of a current block; determining whether a merge mode is applicable to the current block based on the information for prediction; when the merge mode is applicable to the current block, obtaining a first syntax element indicating whether combined prediction is applicable to the current block, where the combined prediction indicates a prediction mode combining inter prediction and intra prediction; when the first syntax element indicates that the combined prediction is applicable to the current block, generating an inter prediction block and an intra prediction block of the current block; and generating a combined prediction block of the current block by weighted-summing the inter prediction block and the intra prediction block.

[0012] Also, according to an embodiment of the present invention, there is provided a video signal processing apparatus including a processor. When a merge mode is applied to a current block, the processor acquires a first syntax element indicating whether combined prediction is applied to the current block. Here, the combined prediction indicates a prediction mode combining inter prediction and intra prediction. When the first syntax element indicates that the combined prediction is applied to the current block, the processor generates an inter prediction block and an intra prediction block of the current block, and generates a combined prediction block of the current block by weighted-summing the inter prediction block and the intra prediction block.

[0013] As an example, it may further include decoding a residual block of the current block; and restoring the current block using the combined prediction block and the residual block.

[0014] As an example, when the first syntax element indicates that the combined prediction is not applied to the current block, the step of decoding the residual block may further include acquiring a second syntax element indicating whether sub-block transform is applied to the current block. The sub-block transform may indicate a transform mode in which transform is applied to only one of the sub-blocks of the current block divided in the horizontal or vertical direction.

[0015] As an example, when the second syntax element does not exist, the value of the second syntax element may be inferred as 0.

[0016] As an example, when the first syntax element indicates that the combined prediction is applied to the current block, the intra prediction mode for intra prediction of the current block may be set to a planar mode.

[0017] As an example, it may further include a step of setting positions of a left-side peripheral block and an upper-side peripheral block referred to for the combination prediction, and the positions of the left-side peripheral block and the upper-side peripheral block may be the same as the positions referred to by the intra prediction.

[0018] As an example, the positions of the left-side peripheral block and the upper-side peripheral block may be determined using a scaling factor variable determined by a color component index value of the current block.

Advantages of the Invention

[0019] According to an embodiment of the present invention, the coding efficiency of a video signal can be increased. Also, according to an embodiment of the present invention, a conversion kernel suitable for a current conversion block can be selected.

Brief Description of the Drawings

[0020]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Best Mode for Carrying Out the Invention

[0021] The terms used in this specification are selected as generally widely used terms as much as possible while considering the functions in the present invention, but these may vary depending on the intentions, conventions of those skilled in the art, or the emergence of new technologies. Also, in certain cases, there are terms arbitrarily selected by the applicant, and in this case, the meaning is described in the part of the embodiment for carrying out the corresponding invention. Therefore, it is clarified that the terms used in this specification should be interpreted based not only on the name of the terms but also on the substantial meaning of the terms and the content throughout this specification.

[0022] In this specification, some terms are interpreted as follows. Coding may be interpreted as encoding or decoding in some cases. In this specification, a device that encodes (codes) a video signal to generate a bitstream of the video signal is referred to as an encoding device or an encoder, and a device that decodes (decodes) the video signal bitstream to restore the video signal is referred to as a decoding device or a decoder. Also, in this specification, a video signal processing device is used as a term for a concept that includes both an encoder and a decoder. Information is a term that includes values, parameters, coefficients, elements, etc., and may be interpreted differently in some cases, so the present invention is not limited thereto. 'Unit' is used to represent a basic unit of video processing or a specific position in a picture, and refers to an image region including at least one of a luma component and a chroma component. Also, 'block' refers to an image region including a specific component among a luminance component and color difference components (that is, Cb and Cr). However, depending on the embodiment, terms such as 'unit', 'block', 'partition', and'region' may be used interchangeably. Also, in this specification, a unit is used as a concept that includes both a coding unit, a prediction unit, and a transformation unit. A picture refers to a field or a frame, and depending on the embodiment, the above terms may be used interchangeably.

[0023] FIG. 1 is a schematic block diagram of a video signal encoding device 100 according to an embodiment of the present invention. Referring to FIG. 1, the encoding device 100 of this specification includes a transformation unit 110, a quantization unit 115, an inverse quantization unit 120, an inverse transformation unit 125, a filtering unit 130, a prediction unit 150, and an entropy coding unit 160.

[0024] The conversion unit 110 converts the residual signal, which is the difference between the input video signal and the prediction signal generated by the prediction unit 150, to obtain conversion coefficient values. For example, a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or a Wavelet Transform may be used. The Discrete Cosine Transform and the Discrete Sine Transform perform the conversion by dividing the input picture signal into blocks. In the conversion, the coding efficiency may vary depending on the distribution and characteristics of the values within the conversion region. The quantization unit 115 quantizes the values of the conversion coefficients output within the conversion unit 110.

[0025] To improve the coding efficiency, instead of coding the picture signal as it is, a method is used in which the picture is predicted using the region that has been pre-coded through the prediction unit 150, and the residual value between the original picture and the predicted picture is added to the predicted picture to obtain the restored picture. To prevent a mismatch from occurring between the encoder and the decoder, information that can also be used by the decoder should be used when making a prediction in the encoder. For this purpose, the encoder performs a process of further restoring the encoded current block. The inverse quantization unit 120 inverse-quantizes the conversion coefficient values, and the inverse conversion unit 125 restores the residual values using the inverse-quantized conversion coefficient values. On the other hand, the filtering unit 130 performs filtering operations for improving the quality of the restored picture and enhancing the coding efficiency. For example, it may include a deblocking filter, a Sample Adaptive Offset (SAO), and an adaptive loop filter. The picture that has undergone filtering is stored in the Decoded Picture Buffer (DPB) 156 either for output or for use as a reference picture.

[0026] In order to improve coding efficiency, instead of directly coding the picture signal, a method is used in which the picture is predicted using the area that has already been coded by the prediction unit 150, and a residual value between the original picture and the predicted picture is added to the predicted picture to obtain a restored picture. In the intra prediction unit 152, in-picture prediction is performed within the current picture, and in the inter prediction unit 154, the current picture is predicted using the reference picture stored in the decoded picture buffer 156. The intra prediction unit 152 performs in-picture prediction from the restored area within the current picture and transmits the in-picture coding information to the entropy coding unit 160. The inter prediction unit 154 may further include a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains the motion vector value of the current area by referring to the restored specific area. The motion estimation unit 154a transmits the position information (such as the reference frame, motion vector, etc.) of the reference area to the entropy coding unit 160 so that it can be included in the bitstream. Using the motion vector value transmitted from the motion estimation unit 154a, the motion compensation unit 154b performs inter-picture motion compensation.

[0027] The prediction unit 150 includes an intra prediction unit 152 and an inter prediction unit 154. The intra prediction unit 152 performs intra prediction within the current picture, and the inter prediction unit 154 performs inter prediction that predicts the current picture using the reference buffer stored in the decoded picture buffer 156. The intra prediction unit 152 performs intra prediction from the restored samples within the current picture and transmits the intra-coded information to the entropy coding unit 160. The intra-coded information includes at least one of an intra prediction mode, an MPM (Most Probable Mode) flag, and an MPM index. The intra-coded information can include information regarding reference samples. The intra-coded information includes information regarding reference samples. The inter prediction unit 154 includes a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains a motion vector value of the current region by referring to a specific region of the restored reference signal picture. The motion estimation unit 154a transmits a set of motion information (reference picture index, motion vector information) for the reference region to the entropy coding unit 160. The motion compensation unit 154b performs motion compensation using the motion vector value transmitted from the motion compensation unit 154a. The inter prediction unit 154 transmits inter-coded information including motion information for the reference region to the entropy coding unit 160.

[0028] According to a further embodiment, the prediction unit 150 includes an intra block copy (BC) prediction unit (not shown). The intra BC prediction unit performs intra BC prediction from the restored samples within the current picture and transmits the intra BC coded information to the entropy coding unit 160. The intra BC prediction unit obtains a block vector value indicating a reference region used for predicting the current region by referring to a specific region within the current picture. The intra BC prediction unit performs intra BC prediction using the obtained block vector value. The intra BC prediction unit transmits the intra BC coded information to the entropy coding unit 160. The intra BC prediction unit includes block vector information.

[0029] When the above-described picture prediction is performed, the conversion unit 110 converts the residual value between the original picture and the predicted picture to obtain a conversion coefficient value. At this time, the conversion is performed in units of specific blocks within the picture, and the size of the specific block varies within a preset range. The quantization unit 115 quantizes the value of the conversion coefficient generated by the conversion unit 110 and transmits it to the entropy coding unit 160.

[0030] The entropy coding unit 160 performs entropy coding on information indicating the quantized conversion coefficient, intra-coding information, inter-coding information, etc. to generate a video signal bitstream. In the entropy coding unit 160, a variable length coding (VLC) method, an arithmetic coding method, etc. are used. The variable length coding (VLC) method maps the input symbols to consecutive codewords, but the length of the codewords is variable. For example, frequently occurring symbols are represented by short codewords, and infrequently occurring symbols are represented by long codewords. As the variable length coding method, a context-based adaptive variable length coding (CAVLC) method is used. Arithmetic coding converts consecutive data symbols into a single prime number, and arithmetic coding obtains the optimal prime number bits required to represent each symbol. As arithmetic coding, a context-based adaptive binary arithmetic coding (CABAC) method is used. For example, the entropy coding unit 160 can binarize the information indicating the quantized conversion coefficient. Also, the entropy coding unit 160 can perform arithmetic coding on the binarized information to generate a bitstream.

[0031] The generated bitstream is encapsulated in units of NAL (Network Abstraction Layer) units. An NAL unit contains an integer number of coded coding tree units that have been encoded. In order to decode the bitstream with a video decoder, first the bitstream should be separated into NAL units, and then each of the separated NAL units should be decoded. On the other hand, information necessary for decoding the video signal bitstream is transmitted via upper-level sets of RBSP (Raw Byte Sequence Payload) such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS).

[0032] On the other hand, the block diagram of FIG. 1 shows an encoding apparatus 100 according to an embodiment of the present invention, and the separately shown blocks logically distinguish the elements of the encoding apparatus 100. Therefore, the elements of the encoding apparatus 100 described above are attached to one chip or a plurality of chips according to the design of the device. According to one embodiment, the operations of each of the elements of the encoding apparatus 100 described above are performed by a processor (not shown).

[0033] FIG. 2 is a schematic block diagram of a video signal decoding apparatus 200 according to an embodiment of the present invention. Referring to FIG. 2, the decoding apparatus 200 described in this specification includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 225, a filtering unit 230, and a prediction unit 250.

[0034] The entropy decoding unit 210 entropy decodes the video signal bitstream to extract transform coefficient information, intra-coding information, inter-coding information, etc. for each region. For example, the entropy decoding unit 210 may obtain a binary code for transform coefficient information of a specific region from the video signal bitstream. The entropy decoding unit 210 also de-binarizes the binary code to obtain quantized transform coefficients. The inverse quantization unit 220 inverse quantizes the quantized transform coefficients, and the inverse transform unit 225 restores residual values using the inverse quantized transform coefficients. The video signal processing device 200 combines the residual values obtained from the inverse transform unit 225 with the predicted values obtained from the prediction unit 250 to restore original pixel values.

[0035] Meanwhile, the filtering unit 230 performs filtering on the picture to improve image quality, including a deblocking filter for reducing block distortion and / or an adaptive loop filter for removing distortion of the entire picture, etc. The filtered picture is output or stored in a decoded picture buffer (DPB) 256 to be used as a reference picture for the next picture.

[0036] The prediction unit 250 includes an intra prediction unit 252 and an inter prediction unit 254. The prediction unit 250 generates a predicted picture by utilizing the encoded type decoded through the entropy decoding unit 210 described above, the conversion coefficients for each region, the intra / inter coding information, and the like. To restore the current block for which decoding is performed, the decoded regions of the current picture or other pictures containing the current block can be used. A picture (or tile / slice) that uses only the current picture for restoration, that is, performs intra prediction or intra BC prediction, is called an intra picture or I picture (or tile / slice). A picture (or tile / slice) that can perform all of intra prediction, inter prediction, and intra BC prediction is called an inter picture (or tile / slice). Among the inter pictures (or tiles / slices), a picture (or tile / slice) that uses at most one motion vector and a reference picture index to predict the sample values of each block is called a predictive picture or P picture (or tile / slice), and a picture (or tile / slice) that uses at most two motion vectors and a reference picture index is called a Bi-predictive picture or B picture (or tile / slice). In other words, the P picture (or tile / slice) uses at most one set of motion information to predict each block, and the B picture (or tile / slice) uses at most two sets of motion information to predict each block. Here, a set of motion information includes one or more motion vectors and one reference picture index.

[0037] The intra prediction unit 252 generates a prediction block by using the intra-coded information and the restored samples within the current picture. As described above, the intra-coded information includes at least one of an intra prediction mode, an MPM (MOST Probable Mode) flag, and an MPM index. The intra prediction unit 252 predicts the sample values of the current block by using the restored samples located on the left side and / or the upper side of the current block as reference samples. In the present disclosure, the restored samples, the reference samples, and the samples of the current block indicate pixels. Also, the sample value indicates a pixel value.

[0038] In one embodiment, the reference samples are the samples included in the peripheral blocks of the current block. For example, the reference samples are the samples adjacent to the left boundary of the current block and / or the samples adjacent to the upper boundary of the current block. Also, the reference samples are the samples located on the lines within a preset distance from the left boundary of the current block and / or the samples located on the lines within a preset distance from the upper boundary of the current block among the samples of the peripheral blocks of the current block. At this time, the peripheral blocks of the current block include at least one of the left (L) block adjacent to the current block, the upper (A) block, the below left (BL) block, the above right (AR) block, or the above left (AL) block.

[0039] The inter prediction unit 254 generates a prediction block by using the reference pictures and inter-coded information stored in the decoded picture buffer 256. The inter-coded information includes a set of motion information (such as a reference picture index, a motion vector, etc.) for the current block with respect to the reference block. There are L0 prediction, L1 prediction, and bi-prediction for inter prediction. The L0 prediction is a prediction using one reference picture included in the L0 picture list, and the L1 prediction means a prediction using one reference picture included in the L1 picture list. For this purpose, a set of motion information (for example, a motion vector and a reference picture index) is required. In the bi-prediction method, a maximum of two reference areas are used. These two reference areas may exist in the same reference picture or may exist in different pictures respectively. That is, in the bi-prediction method, a maximum of two sets of motion information (for example, a motion vector and a reference picture index) are used, but the two motion vectors may correspond to the same reference picture index or may correspond to different reference picture indexes. At this time, the reference picture is displayed (or output) either before or after the current picture in terms of time. According to one embodiment, in the bi-prediction method, the two reference areas used may be areas selected from the L0 picture list and the L1 picture list respectively.

[0040] The inter prediction unit 254 acquires a current reference block by using a motion vector and a reference picture index. The reference block exists in a reference picture corresponding to the reference picture index. Also, the sample value of the block specified by the motion vector or its interpolated value is used as a predictor of the current block. For motion prediction with pixel accuracy in sub-pel units, for example, an 8-tap interpolation filter is used for the luminance signal, and a 4-tap interpolation filter is used for the chrominance signal. However, the interpolation filter for motion prediction in sub-pel units is not limited to this. In this way, the inter prediction unit 254 performs motion compensation for predicting the texture of the current unit from a previously restored picture. At this time, the inter prediction unit uses a motion information set.

[0041] According to a further embodiment, the prediction unit 250 can include an intra BC prediction unit (not shown). The intra BC prediction unit can restore the current region by referring to a specific region including the restored samples in the current picture. The intra BC prediction unit acquires intra BC coding information for the current region from the entropy decoding unit 210. The intra BC prediction unit acquires a block vector value of the current region that indicates a specific region in the current picture. The intra BC prediction unit can perform intra BC prediction by using the acquired block vector value. The intra BC coding information can include block vector information.

[0042] The predicted value output from the intra prediction unit 252 or the inter prediction unit 254 and the residual value output from the inverse conversion unit 225 are added together to generate a restored video picture. That is, the video signal decoding device 200 restores the current block by using the predicted block generated by the prediction unit 250 and the residual acquired from the inverse conversion unit 225.

[0043] On the one hand, the block diagram of FIG. 2 shows a decoding apparatus 200 according to an embodiment of the present invention, and the blocks shown separately logically distinguish the elements of the decoding apparatus 200. Therefore, the elements of the decoding apparatus 200 described above are attached to one chip or a plurality of chips according to the design of the device. According to one embodiment, the operations of each element of the decoding apparatus 200 described above are performed by a processor (not shown).

[0044] FIG. 3 shows an example in which a Coding Tree Unit (CTU) is divided into Coding Units (CUs) within a picture. In the coding process of a video signal, a picture is divided into a sequence of Coding Tree Units (CTUs). A coding tree unit consists of an N×N block of luminance samples and two blocks of corresponding chrominance samples. A coding tree unit is divided into a plurality of coding units. The coding tree unit may become a leaf node without being divided. In this case, the coding tree unit itself can become a coding unit. A coding unit refers to a basic unit for processing a picture in the above-described video signal processing process, that is, processes such as intra / inter prediction, conversion, quantization, and / or entropy coding. Within one picture, the size and pattern of the coding units are not constant. The coding unit has a square or rectangular pattern. The rectangular coding unit (or rectangular block) includes a vertical coding unit (or vertical block) and a horizontal coding unit (or horizontal block). In this specification, a vertical block is a block whose height is greater than its width, and a horizontal block is a block whose width is greater than its height. Also, in this specification, a non-square block refers to a rectangular block, but the present invention is not limited thereto.

[0045] Referring to FIG. 3, the coding tree unit is first divided into a Quad Tree (QT) structure. That is, in the quad tree structure, one node having a size of 2N×2N is divided into four nodes having a size of N×N. In this specification, the quad tree is also referred to as a quaternary tree. The quad tree division is performed recursively, and it is not necessary for all nodes to be divided to the same depth.

[0046] On the other hand, the leaf nodes of the above-mentioned quad tree are further divided into a Multi-Type Tree (MTT) structure. According to an embodiment of the present invention, in the multi-type tree structure, one node is divided into a binary (binary) or ternary (ternary) tree structure of horizontal or vertical division. That is, in the multi-type tree structure, there are four division structures: vertical binary division, horizontal binary division, vertical ternary division, and horizontal ternary division. According to an embodiment of the present invention, in each of the above tree structures, both the width and height of the node have a value that is a power of 2. For example, in a binary Tree (BT) structure, a node with a size of 2N×2N is divided into two nodes of N×2N by vertical binary division and into two nodes of 2N×N by horizontal binary division. Also, in a Ternary Tree (TT) structure, a node with a size of 2N×2N is divided into nodes of (N / 2)×2N, N×2N, and (N / 2)×2N by vertical ternary division and into nodes of 2N×(N / 2), 2N×N, and 2N×(N / 2) by horizontal ternary division. Such multi-type tree division is performed recursively.

[0047] The leaf nodes of a multi-type tree can be coding units. If no splitting is indicated for a coding unit or the coding unit is not larger than the maximum transform length, the corresponding coding unit is used as a prediction and transformation unit without further splitting. On the other hand, in the above-mentioned quad tree and multi-type tree, at least one of the following parameters is predefined or transmitted via a higher-level set of RBSPs such as PPS, SPS, VPS, etc. 1) CTU size: the size of the root node of the quad tree; 2) Minimum QT size (MinQtSize): the size of the smallest allowable QT leaf node; 3) Maximum BT size (MaxBtSize): the size of the largest allowable BT root node; 4) Maximum TT size (MaxTtSize): the size of the largest allowable TT root node; 5) Maximum MTT depth (MaxMttDepth): the maximum allowable depth of MTT splitting from the leaf node of the QT; 6) Minimum BT size (MinBtSize): the size of the smallest allowable BT leaf node; 7) Minimum TT size: the size of the smallest allowable TT leaf node.

[0048] Figure 4 illustrates an embodiment of a method for signaling the splitting of a quad tree and a multi-type tree. Flags that have been set can be used to signal the splitting of the above-mentioned quad tree and multi-type tree. Referring to Figure 4, at least one of the flag 'qt_split_flag' for indicating whether to split a quad tree node, the flag'mtt_split_flag' for indicating whether to split a multi-type tree node, the flag'mtt_split_vertical_flag' for indicating the splitting direction of a multi-type tree node, or the flag'mtt_split_binary_flag' for indicating the splitting form of a multi-type tree node can be used.

[0049] According to an embodiment of the present invention, the coding tree unit is the root node of a quad tree and can be split into a quad tree structure first. In the quad tree structure, a 'qt_split_flag' is signaled for each node 'QT_node'. When the value of 'qt_split_flag' is 1, the corresponding node is split into four square nodes, and when the value of 'qt_split_flag' is 0, the corresponding node becomes a leaf node 'QT_leaf_node' of the quad tree.

[0050] Each quad tree leaf node 'QT_leaf_node' can be further split into a multi-type tree structure. In the multi-type tree structure, an'mtt_split_flag' is signaled for each node 'MTT_node'. When the value of'mtt_split_flag' is 1, the corresponding node is split into a plurality of rectangular nodes, and when the value of'mtt_split_flag' is 0, the corresponding node becomes a leaf node 'MTT_leaf_node' of the multi-type tree. When the multi-type tree node 'MTT_node' is split into a plurality of rectangular nodes (i.e., when the value of'mtt_split_flag' is 1), an'mtt_split_vertical_flag' and an'mtt_split_binary_flag' for the node 'MTT_node' can be additionally signaled. When the value of'mtt_split_vertical_flag' is 1, a vertical split of the node 'MTT_node' is indicated, and when the value of'mtt_split_vertical_flag' is 0, a horizontal split of the node 'MTT_node' is indicated. Also, when the value of'mtt_split_binary_flag' is 1, the node 'MTT_node' is split into two rectangular nodes, and when the value of'mtt_split_binary_flag' is 0, the node 'MTT_node' is split into three rectangular nodes.

[0051] Picture prediction (motion compensation) for coating is performed on a coding unit that cannot be further divided (i.e., a leaf node of a coding unit tree). The basic unit for performing such prediction is hereinafter referred to as a prediction unit or a prediction block.

[0052] Hereinafter, the term "unit" used in this specification is used as a term that substitutes for the prediction unit which is the basic unit for performing prediction. However, the present invention is not limited thereto, and in a broader sense, it is understood as a concept including the coding unit.

[0053] FIGS. 5 and 6 are diagrams showing in more detail an intra prediction method according to an embodiment of the present invention. As described above, the intra prediction unit uses the restored samples located on the left side and / or the upper side of the current block as reference samples to predict the sample values of the current block.

[0054] First, FIG. 5 shows an example of reference samples used for predicting a current block in the intra prediction mode. According to an example, the reference samples are samples adjacent to the left boundary of the current block and / or samples adjacent to the upper boundary of the current block. As shown in FIG. 5, if the size of the current block is W×H and the samples of a single reference line adjacent to the current block are used for intra prediction, the reference samples are set using up to 2W+2H+1 peripheral samples located on the left side and / or the upper side of the current block.

[0055] Also, if at least some of the samples used as reference samples have not yet been restored, the intra prediction unit performs a reference sample padding process to obtain reference samples. Further, the intra prediction unit performs a reference sample filtering process to reduce the error of intra prediction. That is, filtering is performed on the reference samples obtained by the peripheral samples and / or the reference sample padding process to obtain filtered reference samples. The intra prediction unit predicts the samples of the current block using the reference samples thus obtained. The intra prediction unit predicts the samples of the current block using either unfiltered reference samples or filtered reference samples. In the present disclosure, the peripheral samples can include samples on at least one reference line. For example, the peripheral samples can include adjacent samples on the line adjacent to the boundary of the current block.

[0056] Next, FIG. 6 illustrates an embodiment of a prediction mode used for intra prediction. For intra prediction, intra prediction mode information indicating an intra prediction direction can be signaled. The intra prediction mode information indicates any one of a plurality of intra prediction modes that make up the intra prediction mode set. When the current block is an intra-predicted block, the decoder receives the intra prediction mode information of the current block from the bitstream. The intra prediction unit of the decoder performs intra prediction on the current block based on the extracted intra prediction mode information.

[0057] According to embodiments of the present invention, the intra prediction mode set includes all intra prediction modes used for intra prediction (e.g., a total of 67 intra prediction modes). More specifically, the intra prediction mode set includes a planar mode, a DC mode, and a plurality of (e.g., 65) angular modes (i.e., direction modes). Each intra prediction mode is indicated via a preset index (i.e., an intra prediction mode index). For example, as shown in FIG. 6, the intra prediction mode index 0 indicates the planar mode, and the intra prediction mode index 1 indicates the DC mode. Also, the intra prediction mode indices 2 to 66 each indicate different angular modes. The angular modes each indicate different angles within a preset angle range. For example, the angular mode can indicate an angle within an angle range of 45° to -135° in the clockwise direction (i.e., a first angle range). The angular mode may be defined based on the 12 o'clock direction. At this time, the intra prediction mode index 2 indicates the Horizontal Diagonal (HDIA) mode, the intra prediction mode index 18 indicates the Horizontal (HOR) mode, the intra prediction mode index 34 indicates the Diagonal (DIA) mode, the intra prediction mode index 50 indicates the Vertical (VER) mode, and the intra prediction mode index 66 indicates the Vertical Diagonal (VDIA) mode.

[0058] Hereinafter, with reference to FIG. 7, an inter prediction method according to an embodiment of the present invention will be described. In the present invention, the inter prediction method can include a general inter prediction method optimized for translational motion and an affine model-based inter prediction method. Also, the motion vector can usually include at least one of a general motion vector for motion compensation based on the inter prediction method and a control point motion vector for affine motion compensation.

[0059] FIG. 7 illustrates an inter prediction method according to an embodiment of the present invention. As described above, the decoder can predict the current block by referring to the restored samples of other decoded pictures. Referring to FIG. 7, the decoder obtains a reference block 702 in a reference picture 720 based on the motion information set of the current block 701. At this time, the motion information set can include a reference picture index and a motion vector 703. The reference picture index indicates the reference picture 720 in which the reference block for inter prediction of the current block is included in the reference picture list. According to one embodiment, the reference picture list can include at least one of the L0 picture list or the L1 picture list described above. The motion vector 703 indicates an offset between the coordinate value of the current block 701 in the current picture 710 and the coordinate value of the reference block 702 in the reference picture 720. The decoder obtains a predictor of the current block 701 based on the sample value of the reference block 702, and restores the current block 701 using the predictor.

[0060] Specifically, the encoder can search for a block similar to the current block in a picture with an earlier restoration order to obtain the reference block described above. For example, the encoder can search for a reference block in a preset search area where the sum of the differences between the current block and the sample values is minimized. At this time, at least one of SAD (Sum Of Absolute Difference) or SATD (Sum of Hadamard Transformed Difference) can be used to measure the similarity between the samples of the current block and the reference block. Here, SAD can be a value obtained by adding up all the absolute values of the differences between the sample values included in the two blocks. Also, SATD can be a value obtained by adding up all the absolute values of the Hadamard transform coefficients obtained by performing a Hadamard transform on the differences between the sample values included in the two blocks.

[0061] On the other hand, the current block can also be predicted using one or more reference regions. As described above, the current block can be inter-predicted by a dual prediction method using two or more reference regions. According to one embodiment, the decoder can obtain two reference blocks based on two motion information sets of the current block. Further, the decoder can obtain a first predictor and a second predictor of the current block based on the sample values of each of the two obtained reference blocks. Further, the decoder can restore the current block using the first predictor and the second predictor. For example, the decoder can restore the current block based on the sample-by-sample average of the first predictor and the second predictor.

[0062] As described above, for motion compensation of the current block, one or more motion information sets can be signaled. At this time, the similarity between the motion information sets for motion compensation of each of the plurality of blocks can be utilized. For example, the motion information set used for prediction of the current block can be derived from the motion information set used for prediction of any one of the already restored other samples. Through this, the encoder and the decoder can reduce the signaling overhead. Hereinafter, various embodiments in which the motion information set of the current block is signaled will be described.

[0063] For example, there may be a plurality of candidate blocks that may be predicted based on a motion information set identical or similar to the motion information set of the current block. The decoder can generate a merge candidate list based on the plurality of candidate blocks. Here, the merge candidate list can include candidates corresponding to samples that may be predicted based on a motion information set related to the motion information set of the current block among the samples restored prior to the current block. The encoder and the decoder can configure the merge candidate list of the current block based on a predefined rule. At this time, the merge candidate lists configured by the encoder and the decoder may be the same as each other. For example, the encoder and the decoder can configure the merge candidate list of the current block based on the position of the current block within the current picture. A method for the encoder and the decoder to configure the merge candidate list of the current block will be described later with reference to FIG. 9. In the present disclosure, the position of a specific block represents the relative position of the top-left sample of the specific block within the picture including the specific block.

[0064] On the other hand, in order to improve coding efficiency, instead of directly coding the residual signal described above, a method can be used in which the transform coefficient values obtained by transforming the residual signal are quantized and the quantized transform coefficients are coded. As described above, the transform unit can obtain transform coefficient values by transforming the residual signal. At this time, the residual signal of a specific block may be distributed over the entire area of the current block. Thereby, the energy can be concentrated in the low-frequency region by using frequency-domain transformation for the residual signal, and the coding efficiency can be improved. Hereinafter, a method in which the residual signal is transformed or inverse-transformed will be specifically described.

[0065] FIG. 8 is a diagram specifically showing a method by which an encoder converts a residual signal. As described above, the residual signal in the spatial domain may be converted into the frequency domain. The encoder can convert the acquired residual signal to obtain conversion coefficients. First, the encoder can obtain at least one residual block including the residual signal for the current block. The residual block may be either the current block or one of the blocks divided from the current block. In the present disclosure, the residual block may be referred to as a residual array or a residual matrix including the residual samples of the current block. Also, in the present disclosure, the residual block represents a block having the same size as the size of the conversion unit or the conversion block.

[0066] Next, the encoder can convert the residual block using a conversion kernel. The conversion kernel used for the conversion of the residual block may be a conversion kernel having separable characteristics of vertical conversion and horizontal conversion. In this case, the conversion of the residual block may be performed separately for vertical conversion and horizontal conversion. For example, the encoder can perform vertical conversion by applying a conversion kernel in the vertical direction of the residual block. Also, the encoder can perform horizontal conversion by applying a conversion kernel in the horizontal direction of the residual block. In the present disclosure, the conversion kernel may be used as a term representing a parameter set used for the conversion of the residual signal such as a conversion matrix, a conversion array, a conversion function, and a conversion. According to one embodiment, the conversion kernel may be any one of a plurality of available kernels. Also, conversion kernels based on different conversion types may be used for each of the vertical conversion and the horizontal conversion. A method of selecting any one of the plurality of available conversion kernels will be described later with reference to FIGS. 12 to 26.

[0067] The encoder can transmit the transformed block converted from the residual block to the quantization unit for quantization. At this time, the transformed block can include a plurality of transform coefficients. Specifically, the transformed block may be composed of a plurality of transform coefficients arranged in a two-dimensional array. The size of the transformed block may be the same as any one of the current block or a block divided from the current block, similar to the residual block. The transform coefficients transmitted to the quantization unit may be represented by quantized values.

[0068] In addition, the encoder can perform additional transformation before the transform coefficients are quantized. As shown in FIG. 8, the aforementioned transformation method can be called a primary transform, and the additional transformation can be called a secondary transform. The secondary transform can be selective for each residual block. According to one embodiment, the encoder can perform a secondary transform on a region where it is difficult to concentrate energy in the low-frequency region only by the primary transform, thereby improving the coding efficiency. For example, a secondary transform may be added to a block in which the residual value appears largely in a direction other than the horizontal or vertical direction of the residual block. The probability that the residual value of an intra-predicted block changes in a direction other than the horizontal or vertical direction may be higher than that of an inter-predicted block. Accordingly, the encoder can further perform a secondary transform on the residual signal of the intra-predicted block. Also, the encoder can omit the secondary transform on the residual signal of the inter-predicted block.

[0069] As another example, whether to perform the secondary conversion may be determined according to the size of the current block or the residual block. Also, conversion kernels of different sizes may be used according to the size of the current block or the residual block. For example, an 8×8 secondary conversion may be applied to a block whose length of the shorter side among the width or the height is shorter than a first preset length. Also, a 4×4 secondary conversion may be applied to a block whose length of the shorter side among the width or the height is longer than a second preset length. At this time, the first preset length may be a value larger than the second preset length, but the present disclosure is not limited thereto. Also, unlike the primary conversion, the secondary conversion does not have to be performed separately for the vertical conversion and the horizontal conversion. Such a secondary conversion can be called a Low Frequency Non-Separable Transform (LFNST).

[0070] Also, in the case of a video signal in a specific region, due to a sudden brightness change, the high-frequency band energy may not decrease even when frequency conversion is performed. As a result, the compression performance due to quantization may decrease. Also, when conversion is performed on a region where residual values rarely exist, the encoding time and the decoding time may increase unnecessarily. Therefore, the conversion for the residual signal in the specific region may be omitted. Whether to perform the conversion for the residual signal in the specific region may be determined by a syntax element related to the conversion in the specific region. For example, the syntax element may include transform skip information. The transform skip information may be a transform skip flag. When the transform skip information for the residual block indicates transform skip, the conversion for the residual block is not performed. In this case, the encoder can immediately quantize the residual signal for which the conversion in the region is not performed. The operation of the encoder described with reference to FIG. 8 can be performed in the conversion unit of FIG. 1.

[0071] The above-described conversion-related syntax elements may be information parsed from a video signal bitstream. The decoder can entropy-decode the video signal bitstream to obtain the conversion-related syntax elements. Also, the encoder can entropy-encode the conversion-related syntax elements to generate a video signal bitstream.

[0072] FIG. 9 is a diagram specifically showing a method in which an encoder and a decoder inverse-transform conversion coefficients to obtain a residual signal. Hereinafter, for convenience of explanation, it will be described assuming that the inverse-transform operation is performed in each inverse-transform unit of the encoder and the decoder. The inverse-transform unit can inverse-transform the inverse-quantized conversion coefficients to obtain a residual signal. First, the inverse-transform unit can detect whether inverse transformation for a specific region is to be performed from the conversion-related syntax elements of the specific region. According to one embodiment, when the conversion-related syntax elements for a specific transform block indicate transform skip, the transform for the transform block may be omitted. In this case, all of the above-described first inverse transformation and second inverse transformation may be omitted for the transform block. Also, the inverse-quantized conversion coefficients may be used as the residual signal. For example, the decoder can restore the current block using the inverse-quantized conversion coefficients as the residual signal.

[0073] According to another embodiment, the conversion-related syntax elements for a specific transform block may not be able to represent transform skip. In this case, the inverse-transform unit can determine whether to perform a second inverse transform for the second transform. For example, when the transform block is a transform block of an intra-predicted block, a second inverse transform for the transform block may be performed. Also, based on the intra-prediction mode corresponding to the transform block, the second transform kernel used for the transform block may be determined. As another example, it may be determined whether to perform a second inverse transform based on the size of the transform block. The second inverse transform may be performed after the inverse-quantization process and before the first inverse transform is performed.

[0074] The inverse transformation unit can perform an inverse quantization of the transformed coefficients or a first inverse transformation on the secondarily inverse-transformed coefficients. In the case of the first inverse transformation, similar to the first transformation, it may be performed separately as a vertical transformation and a horizontal transformation. For example, the inverse transformation unit can perform a vertical inverse transformation and a horizontal inverse transformation on a transformation block to obtain a residual block. The inverse transformation unit can inverse-transform the transformation block based on the transformation kernel used for the transformation of the transformation block. For example, the encoder can explicitly or implicitly signal information indicating the transformation kernel currently applied to the transformation block among a plurality of available transformation kernels. The decoder can select, from among the plurality of available transformation kernels, the transformation kernel used for the inverse transformation of the transformation block using the information indicating the signaled transformation kernel. The inverse transformation unit can restore the current block using the residual signal obtained by the inverse transformation on the transformed coefficients.

[0075] FIG. 10 is a diagram illustrating a motion vector signaling method according to an embodiment of the present invention. According to an embodiment of the present invention, a motion vector (MV) may be generated based on a motion vector prediction (or predictor) (MVP). As an example, as shown in the following mathematical formula 1, the MV may be determined as the MVP. In other words, the MV may be determined (or set, derived) to have the same value as the MVP.

[0076]

Equation

[0077] As another example, as shown in the following mathematical formula 2, the MV may be determined based on the MVP and a motion vector difference (MVD). The encoder can signal the MVD information to the decoder in order to indicate a more accurate MV, and the decoder can derive the MV by adding the obtained MVD to the MVP.

[0078]

Number

[0079] According to an embodiment of the present invention, the encoder transmits the determined motion information to the decoder, and the decoder can generate (or derive) an MV from the received motion information and generate a prediction block based thereon. For example, the motion information may include MVP information and MVD information. At this time, the components of the motion information may change depending on the inter prediction mode. As an example, in the merge mode, the motion information may include MVP information and may not include MVD information. As another example, in the AMVP (advanced motion vector prediction) mode, the motion information may include MVP information and MVD information.

[0080] To determine, transmit, and receive information regarding the MVP, the encoder and the decoder can generate an MVP candidate (or an MVP candidate list) in the same way. For example, the encoder and the decoder can generate the same MVP candidate in the same order. Then, the encoder transmits an index indicating (or designating) the determined (or selected) MVP from among the generated MVP candidates to the decoder, and the decoder can derive the determined MVP and / or MV based on the received index.

[0081] According to an embodiment of the present invention, the MVP candidate may include a spatial candidate, a temporal candidate, etc. When the merge mode is applied, the MVP candidate can be called a merge candidate, and when the AMVP mode is applied, it can be called an AMVP candidate. The spatial candidate may be an MV (or motion information) for a block at a specific position based on the current block. For example, the spatial candidate may be an MV of a block adjacent to or not adjacent to the current block. The temporal candidate may be an MV corresponding to a block in the current picture and other pictures. Also, for example, the MVP candidate can include an affine MV, an ATMVP, an STMVP, a combination of the aforementioned MVs (or candidates), an average MV of the aforementioned MVs (or candidates), a zero MV, etc.

[0082] In one embodiment, the encoder can signal information indicating a reference picture to the decoder. As an example, when the reference picture of the MVP candidate is different from the reference picture of the current block (or the current processing block), the encoder / decoder can scale the MV of the MVP candidate. At this time, the MV scaling may be performed based on the picture order count (POC) of the current picture, the POC of the reference picture of the current block, and the POC of the reference picture of the MVP candidate.

[0083] Hereinafter, specific embodiments related to the MVD signaling method will be described. Table 1 below illustrates a syntax structure for MVD signaling.

[0084] [Table 1]

[0085] Referring to Table 1, according to one embodiment of the present invention, the MVD may be coded separately for the sign and the absolute value of the MVD. That is, the sign and the absolute value of the MVD may be different syntaxes (or syntax elements). Also, the absolute value of the MVD may be directly coded for its value, or may be coded stepwise based on a flag indicating whether the absolute value is greater than N as in Table 1. If the absolute value is greater than N (absolute value - N), both values may be signaled. Specifically, an abs_mvd_greater0_flag indicating whether the absolute value is greater than 0 may be transmitted in the example of Table 1. If the abs_mvd_greater0_flag indicates (or indicates) that the absolute value is not greater than 0, the absolute value of the MVD may be determined to be 0. Also, if the abs_mvd_greater0_flag indicates that the absolute value is greater than 0, additional syntax (or syntax elements) may be present.

[0086] For example, an abs_mvd_greater1_flag indicating whether the absolute value is greater than 1 may be transmitted. If the abs_mvd_greater1_flag indicates (or indicates) that the absolute value is not greater than 1, the absolute value of the MVD may be determined to be 1. If the abs_mvd_greater1_flag indicates that the absolute value is greater than 1, additional syntax may be present. For example, abs_mvd_minus2 may be present. abs_mvd_minus2 may be the value of (absolute value - 2). Since the absolute value is determined to be greater than 1 (i.e., 2 or more) by the abs_mvd_greater0_flag and abs_mvd_greater1_flag values, the (absolute value - 2) value may be signaled. In this way, by hierarchically syntax-signaling the information regarding the absolute value, fewer bits are used compared to the case of directly binarizing and signaling the absolute value.

[0087] In one embodiment, the syntax related to the absolute value described above may be coded by applying a variable length binary method such as Exponential-Golomb, truncated unary, truncated Rice, etc. Also, the flag indicating the sign of the MVD may be signaled by mvd_sign_flag.

[0088] In the above-described embodiment, the coding method for MVD has been described. However, information other than MVD can also be signaled separately for the sign and the absolute value. And the absolute value may be coded into a flag indicating whether the absolute value is greater than a predefined specific value and a value obtained by subtracting the specific value from the absolute value. In Table 1 above, [0] and [1] can represent a component index. For example, it can represent an x-component (i.e., a horizontal component) and a y-component (i.e., a vertical component).

[0089] FIG. 11 is a diagram illustrating a signaling method for adaptive motion vector resolution information according to an embodiment of the present invention. According to an embodiment of the present invention, the resolution for indicating MV or MVD can be various. For example, the resolution may be expressed based on a pixel (or, pel). For example, MV or MVD may be signaled in units such as 1 / 4 (quarter), 1 / 2 (half), 1 (integer), 2, 4 pixels, etc. And the encoder can signal the resolution information of MV or MVD to the decoder. Also, for example, 16 may be coded as 64 when in 1 / 4 unit (1 / 4 * 64 = 16), 16 when in 1 unit (1 * 16 = 16), and 4 when in 4 units (4 * 4 = 16). That is, the MV or MVD value may be determined by the following mathematical formula 3.

[0090]

Equation

[0091] In Mathematical formula 3, valueDetermined represents the MV or MVD value. Also, valuePerResolution represents the value signaled based on the determined resolution. At this time, when the value signaled by the MV or MVD is not divisible by the determined resolution, a rounding process or the like may be applied. Using a high resolution can improve accuracy, but since the coded value is large, many bits may be used. Using a low resolution may result in low accuracy, but since the coded value is small, fewer bits may be used. As an example, the resolution described above may be individually set in units such as a sequence, a picture, a slice, a coding tree unit (CTU), and a coding unit (CU). That is, the encoder / decoder can adaptively determine / apply the resolution according to a predefined unit among the above-mentioned units.

[0092] According to an embodiment of this specification, the above-described resolution information may be signaled from the encoder to the decoder. At this time, the information regarding the resolution may be binary-coded and signaled based on the variable length described above. In such a case, when signaling is performed based on the index corresponding to the minimum value (i.e., the value at the front), the signaling overhead can be reduced. As an example, it may be mapped to the signaling index in descending order from high resolution to low resolution.

[0093] According to an embodiment of the present specification, FIG. 11 shows a signaling method assuming a case where three resolutions are used among a plurality of different resolutions. In this case, the three signaling bits may be 0, 10, and 11, and the three signaling indexes may respectively represent the first resolution, the second resolution, and the third resolution. Since 1 bit is required to signal the first resolution and 2 bits are required to signal the remaining resolutions, the signaling overhead can be relatively reduced when signaling the first resolution. In the illustration of FIG. 11, the first resolution, the second resolution, and the third resolution may be defined as 1 / 4 and 1,4 pixel resolutions, respectively. In the following embodiments, the MV resolution may mean the resolution of the MVD.

[0094] FIG. 12 is a diagram illustrating a history-based motion vector prediction (HMVP) method according to an embodiment of the present invention. As described above, the encoder / decoder can use spatial candidates, temporal candidates, etc. as motion vector candidates, and in an embodiment of the present invention, a history-based motion vector, that is, an HMVP can be further used as a motion vector candidate.

[0095] According to an embodiment of the present specification, the encoder / decoder can store the motion information of previously encoded blocks in a table. In the present specification, the HMVP represents the motion information of previously encoded blocks. That is, the encoder / decoder can store the HMVP in the table. In the present specification, the table for storing the HMVP is referred to as a table or an HMVP table, but the present invention is not limited to such a name. As an example, the table (or HMVP table) can be referred to as a buffer, an HMVP buffer, an HMVP candidate buffer, an HMVP list, an HMVP candidate list, etc.

[0096] The motion information stored in the HMVP table can include at least one of MV, reference list, reference index, or utilization flag. For example, the motion information can include at least one of the MV of reference list L0, the MV of L1, the L0 reference index, the L1 reference index, the L0 prediction list utilization flag, or the L1 prediction list utilization flag. At this time, the prediction list utilization flag can indicate whether the information is available for the list, whether it is significant information, and so on.

[0097] Also, according to an embodiment of the present specification, the motion information stored in the HMVP table may be generated / stored in a history-based manner. The history-based motion information represents the motion information of blocks coded before the current block in the coding order. For example, the encoder / decoder can store the motion information of blocks coded before the current block in the HMVP table. At this time, the block may be a coding unit (CU) or a prediction unit (PU), etc. The motion information of the block can mean the motion information used for motion compensation of the block or the candidate motion information used for motion compensation. Then, the motion information (i.e., HMVP) stored in the HMVP table may be used for motion compensation of blocks to be encoded / decoded later. For example, the motion information stored in the HMVP table may be used for motion candidate list construction.

[0098] Referring to FIG. 12, the encoder / decoder can fetch one or more HMVP candidates from the HMVP table (S1201). Then, the encoder / decoder can add the HMVP candidate to the motion candidate list. For example, the motion candidate list may be a merge candidate list (or, merge list), an AMVP candidate list (or, AMVP list). And the encoder / decoder can perform motion compensation and encoding / decoding for the current block based on the motion candidate list (S1202). The encoder / decoder can update the HMVP table using the information used for motion compensation or decoding of the current block (S1203).

[0099] According to an embodiment of the present invention, the HMVP candidate may be used in the merge candidate list construction process. A plurality of the latest HMVP candidates in the HMVP table may be sequentially confirmed and inserted (or, added) into the merge candidate list in the order after the temporal motion vector prediction (or, predictor) (TMVP) candidate. Also, when adding an HMVP candidate, a pruning process (or, pruning check) may be performed on the spatial or temporal merge candidates included in the merge list except for the sub-block motion candidates (i.e., ATMVP). In one embodiment, the following embodiment may be applied to reduce the number of pruning operations.

[0100] 1) For example, the number of HMPV candidates may be set as in the following Mathematical Formula 4.

[0101]

Equation

[0102] In Mathematical Expression 4, L represents the number of HMPV candidates. And N represents the number of available non-sub block merge candidates, and M represents the number of HMVP candidates available in the table. For example, N can indicate the number of non-sub block merge candidates included in the merge candidate list. When N is less than or equal to 4, the number of HMVP candidates may be determined as M, and otherwise, the number of HMVP candidates may be determined as (8 - N).

[0103] 2) Also, for example, if the total number of available merge candidates reaches a value obtained by subtracting 1 from the number of signaling-maximum allowable merge candidates, the merge candidate list construction process from the HMVP list may end.

[0104] 3) Also, for example, the number of candidate pairs for the derivation of combined bi-predictive merge candidates may decrease from 12 to 6.

[0105] Also, according to an embodiment of the present invention, HMVP candidates may also be used in the AMVP candidate list construction process. In the table, the motion vectors of HMVP candidates having the last K indices may be inserted following the TMVP candidates. In one embodiment, only HMVP candidates having the same reference picture as the AMVP target reference picture can be used to construct the AMVP candidate list. Also in this case, the pruning process described above may be applied to the HMVP candidates. For example, K described above may be set to 4, and the size (or length) of the AMVP list may be set to 2.

[0106] FIG. 13 is a diagram illustrating a method for updating an HMVP table according to an embodiment of the present invention. According to an embodiment of the present invention, the HMVP table may be maintained / managed in a FIFO (first-in, first-out) manner. That is, when there is a new input, the oldest element (or candidate) may be output first. For example, when adding the motion information used in the current block to the HMVP table, if the number of elements equal to the maximum number of elements in the HMVP table is filled, the encoder / decoder may output the motion information added most recently from the HMVP table and add the motion information used in the current block to the HMVP table. At this time, when the motion information existing in the existing HMVP table is output, the encoder / decoder may move the motion information at the next position (or index) while filling the output position. For example, when the HMVP table index of the output motion information is m and the oldest element is located at the HMVP table index 0, the motion information corresponding to the index n greater than m may be moved to the HMVP table index (n - 1) position respectively. Thereby, the position with a higher index in the HMVP table becomes empty, and the encoder / decoder can assign the highest index in the HMVP table to the motion information used in the current block. That is, the motion information of the current block may be inserted at the position of index (M + 1) when the maximum index including valid elements is M.

[0107] Also, according to an embodiment of the present invention, when updating the HMVP table based on specific motion information, the encoder / decoder can apply a pruning process. That is, the encoder / decoder can check whether the specific motion information or information corresponding to the specific motion information is included in the HMVP table. Then, the HMVP table update methods can be individually defined for the case where duplicate motion information is included and the case where it is not. Thereby, it is possible to prevent the HMVP table from including duplicate motion information, and various candidates can be considered for motion compensation.

[0108] In one embodiment, when updating the HMVP table based on the motion information used for the current block, if the motion information is already included in the HMVP table, the encoder / decoder can delete the duplicate motion information already included in the HMVP table and newly add the motion information to the HMVP table. At this time, prior to the method of deleting the existing candidate and newly adding it, the method described using FIFO can be used. That is, the index of the motion information to be currently added in the HMVP table is determined to be m, and the oldest motion information may be output. If the motion information is not already included in the HMVP table, the encoder / decoder can delete the most recently added motion information and add the motion information to the HMVP table.

[0109] In another embodiment, when updating the HMVP table based on the motion information used for the current block, if the motion information is already included in the HMVP table, the encoder / decoder does not change the HMVP table, and if it is not already included, the HMVP table can be updated in a FIFO manner.

[0110] Also, according to an embodiment of the present invention, the encoder / decoder can initialize (or reset) the HMVP table at a predefined time point or a predefined position. Since the same motion candidate list must be used by the encoder and the decoder, the encoder and the decoder need to use the same HMVP table. At this time, if the HMVP table is continuously used without initialization, there is a problem that dependencies between coding blocks occur. Therefore, it is necessary to reduce the dependencies between blocks to support parallel processing, and an operation of initializing the HMVP table may be preset according to a unit that supports parallel processing. For example, the encoder / decoder can be set to initialize the HMVP table at the slice level, CTU row level, CTU level, etc. For example, when the initialization for the HMVP table is defined at the CTU row level, the encoder / decoder can perform encoding / decoding in a state where the HMVP table is empty when starting coding for each CTU row.

[0111] FIG. 14 is a diagram for explaining a method of updating an HMVP table according to an embodiment of the present invention. Referring to FIG. 14, in an embodiment of the present invention, the encoder / decoder can update the HMVPCandList based on the mvCand. In this specification, the HMVPCandList represents the HMVP table, and the mvCand represents the motion information of the current block. The process shown in FIG. 14 can take as input a motion candidate mvCand having two motion vectors mvL0 and mvL1, two reference indices refIdxL0 and refIdxL1, and two prediction list utilization flags predFlagL0 and predFlagL1. Then, the process shown in FIG. 14 can output a modified HMVPCandList array.

[0112] In the first step (Step1), the encoder / decoder can check whether mvCand is the same as HMVPCandList[HMVPIdx] while changing the HMVPIdx value, which is a variable indicating the index of the HMVP in the HMVP table, from 0 to (HMVPCandNum - 1). Here, HMVPCandNum indicates the number of HMVP included in the HMVP table. And HMVPCandList[HMVPIdx] indicates the candidate in the HMVP table having the HMVPIdx value. When mvCand is the same as HMVPCandList[HMVPIdx], the encoder / decoder can set the variable sameCand, which indicates whether the candidates are the same, to true. In one embodiment, when checking whether mvCand is the same as HMVPCandList[HMVPIdx], the encoder / decoder can perform a comparison on the motion information for which the prediction list utilization flag is 1 among the MV and reference index related to L0, or the MV and reference index related to L1.

[0113] Then, in the second step (Step2), the encoder / decoder can set the variable tempIdx, which indicates the temporary index, to HMVPCandNum. In the third step (Step3), when sameCand is true or HMVPCandNum is the same as the size (or length) of the maximum HMVP table, the encoder / decoder can copy HMVPCandList[tempIdx] to HMVPCandList[tempIdx - 1] while changing tempIdx from (sameCand? HMVPIdx:1) to (HMVPCandNum - 1). That is, when sameCand is true, tempIdx starts from HMVPIdx, and the encoder / decoder can set HMVPIdx to the value obtained by adding 1 to the element index that is the same as mvCand in the HMVP table. Also, when sameCand is false, tempIdx starts from 1, and HMVPCandList[0] may be written with the content of HMVPCandList[1].

[0114] In the fourth stage (Step4), the encoder / decoder can copy mvCand, which is the motion information to be updated, to HMVPCandList[tempIdx]. In the fifth stage (Step5), if HMVPCandNum is smaller than the maximum HMVP table size, the encoder / decoder can increment HMVPCandNum by 1.

[0115] FIG. 15 is a diagram for explaining a method for updating an HMVP table according to an embodiment of the present invention. Referring to FIG. 15, in an embodiment of the present invention, the encoder / decoder can start comparing HMVPIdx from a value greater than 0 in the process of checking whether the motion information of the current block is included in the HMVP table. For example, when checking whether the motion information of the current block is included in the HMVP table, the encoder / decoder can compare excluding those corresponding to HMVPIdx0. In other words, the encoder / decoder can compare with mvCand from the candidates corresponding to HMVPIdx1. The process shown in FIG. 15 can take as input a motion candidate mvCand having two motion vectors mvL0 and mvL1, two reference indexes refIdxL0 and refIdxL1, and two prediction list utilization flags predFlagL0 and predFlagL1. And the process shown in FIG. 15 can output a modified HMVPCandList array.

[0116] To further explain the HMVP table update method described in FIG. 14 above, at the first stage, it is possible to check whether there is the same motion information as mvCand starting from HMVPIdx0. However, in the above-described update method, when mvCand exists in HMVPCandList[0] or does not exist in HMVPCandList[HMVPIdx] corresponding to HMVPCandNum from HMVPIdx0, the motion information of HMVPCandList[0] is output. When mvCand exists in HMVPCandList[HMVPIdx] among HMVPIdx values other than 0, the content of HMVPCandList[0] is not updated, so there is no need to compare mvCand with HMVPCandList[0]. Also, since the motion information of the current block is similar to the motion information of spatially adjacent blocks, it is highly likely to be similar to the motion information most recently added in the HMVP table. According to such an assumption, mvCand may have motion information similar to candidates with an HMVPIdx greater than 0 rather than the candidate motion information corresponding to HMVPIdx0. When similar (or identical) motion information is searched, the pruning process can be terminated. Therefore, by performing the first stage of FIG. 14 above with a value of HMVPIdx greater than 0, the number of comparisons can be reduced compared to the embodiment of FIG. 14 described above.

[0117] FIG. 16 is a diagram for explaining a method of updating an HMVP table according to an embodiment of the present invention. Referring to FIG. 16, in an embodiment of the present invention, an encoder / decoder can compare the HMVPIdx from a value greater than 0 in the process of checking whether the motion information of the current block, that is, mvCand, is included in the HMVP table. The process shown in FIG. 16 can take as input a motion candidate mvCand having two motion vectors mvL0 and mvL1, two reference indexes refIdxL0 and refIdxL1, and two prediction list utilization flags predFlagL0 and predFlagL1. And the process shown in FIG. 16 can output a modified HMVPCandList array. For example, when checking whether mvCand is included in the HMVP table, the encoder / decoder can perform a comparison process (or a pruning process) excluding a specific number of predefined motion information (or candidates) corresponding to an HMVPIdx value of 0 or more. Specifically, when a variable indicating the number of elements to be compared is indicated by NumPrune, the encoder / decoder can check whether mvCand matches the HMVP table elements corresponding to HMVPIdx (HMVPCandNum - NumPrune + 1) to (HMVPCandNum - 1). Thereby, compared with the embodiments described above, the number of comparisons can be significantly reduced.

[0118] In other embodiments, the encoder / decoder can compare mvCand with a specific set position other than HMVPIdx0 in the HMVP table. For example, the encoder / decoder can compare whether mvCand is motion information overlapping with the HMVP table elements corresponding to the HMVPIdx PruneStart to (HMVPCandNum - 1).

[0119] FIG. 17 is a diagram for explaining a method of updating an HMVP table according to an embodiment of the present invention. Referring to FIG. 17, in an embodiment of the present invention, in the process of checking whether mvCand is included in the HMVP table, the encoder / decoder can compare mvCand in order from the most recently added HMVP table element to the oldest HMVP table element. The process shown in FIG. 17 can take as input a motion candidate mvCand having two motion vectors mvL0 and mvL1, two reference indices refIdxL0 and refIdxL1, and two prediction list utilization flags predFlagL0 and predFlagL1. Then, the process shown in FIG. 17 can output a modified HMVPCandList array. Since mvCand is likely to be similar to the motion information of spatially adjacent blocks, it may be similar to the motion information that was relatively recently added within the HMVP table. In such a manner, when similar (or identical) motion information is searched, the pruning process can be terminated. Thus, as in the embodiment of FIG. 17, by performing the comparison from the HMVPIdx corresponding to the most recently added element, the number of comparisons can be reduced.

[0120] FIG. 18 is a diagram illustrating a pruning process according to an embodiment of the present invention. According to an embodiment of the present invention, when an encoder / decoder checks whether motion information of a current block is already included in an HMVP table, it can compare with some information in the HMVP table. This is to reduce the complexity caused by the comparison process. For example, when checking whether motion information is included in HMVPCandList[HMVPIdx], the encoder / decoder can perform a pruning process using a subset of HMVPCandList[HMVPIdx] or HMVPCandList[HMVPIdx] having a predefined index. For example, HMVPCandList[HMVPIdx] may include information regarding L0 and L1 respectively, but when both L0 and L1 utilization flags are 1, the encoder / decoder can select and compare one of L0 and L1 according to a predefined convention (or condition). For example, a pruning process for checking overlapping motion information can be performed only for the smaller reference list among the L0 reference index and the L1 reference index.

[0121] In one embodiment, the encoder / decoder can perform a pruning process only on the smaller reference list among the L0 reference index and the L1 reference index of the motion information in the HMVP table. For example, the encoder / decoder can perform a pruning process only on the smaller reference index among the L0 reference index and the L1 reference index of the motion information of the current block. In other embodiments, the encoder / decoder can compare a value based on a subset of the reference index bits and the motion information bits of the current block with a value based on a subset of the reference index bits and the motion information bits of the HMVP table element to determine whether it is already included in the HMVP table. As an example, the value based on a subset of the reference index bits and the motion information bits may be some of the reference index bits and the bits and some of the motion information bits. In other embodiments, the value based on a subset of the reference index bits and the motion information bits may be a value (or output) obtained by passing the first subset and the second subset through a hash function. In other embodiments, if the size difference between the two motion vectors in the pruning process is less than or equal to a predefined threshold, the encoder / decoder can determine that the two motion vectors are the same or similar.

[0122] FIG. 19 is a diagram illustrating a method for adding an HMVP candidate according to an embodiment of the present invention. According to an embodiment of the present invention, the encoder / decoder can add the MVs included in the HMVP table to the motion candidate list. As an example, since the motion of the current block may be similar to the most recently added motion and the most recently added motion may be useful as a candidate, the encoder / decoder can add the most recently added element in the HMVP table to the candidate list.

[0123] At this time, in one embodiment, referring to FIG. 19, the encoder / decoder can add to the candidate list from an element having a specific index value rather than an element of the most recently added HMVP table. Alternatively, the encoder / decoder may add a specific number of elements previously added to the candidate list from the element next to the most recently added HMVP table element rather than the most recently added HMVP table element.

[0124] Alternatively, in one embodiment, the encoder / decoder may add a specific number of elements previously added to the candidate list from elements excluding one or more of the most recently added HMVP table elements. At this time, the encoder / decoder can preferentially add the relatively recently input motion information among the HMVP table elements to the candidate list.

[0125] The most recently added HMVP table element may correspond to a block that is spatially close to the current block, so it is highly likely that it has already been added as a spatial candidate in the candidate list construction process. Therefore, as in this embodiment, by constructing the candidate list excluding a specific number of candidates added recently, it is possible to prevent adding unnecessary candidates to the candidate list, and when adding to the candidate list, the complexity of the pruning processor can be reduced. Although the HMVP table elements and their use have been mainly described with respect to the construction of the motion candidate list, the present invention is not limited thereto and is applicable to other parameters or inter / intra prediction-related information.

[0126] FIG. 20 is a diagram illustrating a merge sharing node according to an embodiment of the present invention. According to an embodiment of the present invention, an encoder / decoder may share the same candidate list for a plurality of blocks in order to facilitate parallel processing. The candidate list may be a motion candidate list. The plurality of blocks may be defined according to a preset convention. For example, a block below a block (or region) that satisfies a specific condition may be defined as the plurality of blocks. Or, a block included in a block that satisfies a specific condition may be defined as the plurality of blocks. In this specification, the candidate list that is identically used by the plurality of blocks may be referred to as a shared list. If applied to a merge candidate list, it may be referred to as a shared merge list.

[0127] Also, in this specification, the block that satisfies the specific condition may be referred to as a merge sharing node (i.e., the portion indicated by the dotted line in FIG. 20), a shared merge node, a merge sharing region, a shared merge region, a shared merge list node, a shared merge list region, etc. The encoder / decoder can construct a motion candidate list based on the motion information of the peripheral blocks adjacent to the merge sharing node, thereby ensuring parallel processing for the coding units within the merge sharing node. Also, a threshold value can be used as the specific condition described above. For example, the block defined as the merge sharing node may be defined (or determined) based on the threshold value. As an example, the threshold value used in the specific condition may be set to a value related to the block size, the width / height of the block.

[0128] FIG. 21 is a diagram for explaining an HMVP update method when a shared list according to an embodiment of the present invention is used. In FIG. 21, it is assumed that the shared list described above with reference to FIG. 20 is used. When constructing a shared list using HMVP candidates, the encoder and the decoder must maintain the same HMVP table. If the update rules for the HMVP table are not defined based on the motion information used for a plurality of blocks using the same shared list, the same candidate list cannot be constructed. Therefore, in FIGS. 22 to 24 described below, assuming that CU1, CU2, CU3, and CU4 use the shared list within the merge shared node shown in FIG. 21, an HMVP table update method for ensuring the same shared list within the merge shared node will be described.

[0129] FIG. 22 is a diagram illustrating a method of updating an HMVP table based on the motion information of blocks within a merge shared node according to an embodiment of the present invention. Referring to FIG. 22, the encoder / decoder can apply the same shared list to CU1, CU2, CU3, and CU4 in the sub-node coding units of the merge shared node. Specifically, the encoder / decoder can construct the same motion candidate list for CU1, CU2, CU3, and CU4 based on the motion information of the surrounding blocks adjacent to the merge shared node (S2201, S2202, S2203, S2204). The motion candidate list can include motion vector components and reference indices.

[0130] After motion information derivation or encoding / decoding for CUs CU1, CU2, CU3, and CU4 using the common motion candidate list, the encoder / decoder can update the HMVP table using the motion information for CUs CU1, CU2, CU3, and CU4 (S2205). In situations where the same candidate list should be used, that is, when multiple CUs within a merge common node update the HMVP table, it may not be possible to use the same candidate list. Therefore, according to an embodiment of the present invention, in situations where a shared list is used, the update of the HMVP table may be performed after the MV derivation or coding of all CUs belonging to the merge common node has been completed, whereby all CUs belonging to the merge common node can construct a motion candidate list based on the same HMVP table.

[0131] FIG. 23 is a diagram for explaining an HMVP update method when a shared list according to an embodiment of the present invention is used. Referring to FIG. 23, assume the case where the shared list described in FIGS. 20 to 22 above is used. At this time, when the MV derivation or encoding / decoding of each CU in FIG. 22 is completed, since there are multiple CUs within the merge common node, there may also be multiple pieces of motion information used.

[0132] According to an embodiment of the present invention, as shown in FIG. 23, the encoder / decoder can update the HMVP table using all the motion information within the merge common node. In other words, the encoder / decoder can update the HMVP table using the motion information used by all CUs that use the same shared list. At this time, the update order needs to be already set in the encoder and decoder.

[0133] In one embodiment, the encoder / decoder can refer to the regular decoding order as the update order. For example, as shown in FIG. 23, according to the regular coding order, the motion information corresponding to each CU can be sequentially used as the input of the HMVP table update process. Alternatively, as one embodiment, the encoder / decoder may determine the HMVP table update order for the motion information of a plurality of CUs by referring to the reference index order, the POC relationship between the current picture and the reference picture, etc. For example, all the motion information in the merge shared node can be updated to the HMVP table in the reference index order. Or, for example, all the motion information in the merge shared node can be updated to the HMVP table in the order of low or high POC difference between the current picture and the reference picture.

[0134] FIG. 24 is a diagram illustrating a method of updating an HMVP table based on the motion information of blocks in a merge shared node according to an embodiment of the present invention. Referring to FIG. 24, it is assumed that the shared lists described in FIGS. 20 to 22 above are used. At this time, when the MV derivation or encoding / decoding of each CU in FIG. 22 is completed, since there are a plurality of CUs in the merge shared node, there may also be a plurality of pieces of motion information used.

[0135] According to an embodiment of the present invention, an encoder / decoder can update an HMVP table using some of the motion information used in a plurality of CUs within a merge sharing node. In other words, for at least some of the CUs among the motion information used in a plurality of CUs within a merge sharing node, the encoder / decoder may not update the motion information in the HMVP table. As an example, the encoder / decoder can refer to the regular coding order as the update order. For example, the encoder / decoder can update the HMVP table using a set number of motion information with a late decoding order among the motion information corresponding to each CU according to the regular decoding order. This is because a block that is late in the regular coding order is likely to be spatially adjacent to the next coded block and may require similar motion during motion compensation.

[0136] In one embodiment, referring to FIG. 24, when there are CUs CU1 to CU4 using the same shared list, the encoder / decoder can use only the motion information for CU4, which is coded last in the regular decoding order, for updating the HMVP table. As another example, the encoder / decoder can update the HMVP table using some of the motion information among the motion information of a plurality of CUs by referring to the reference index order, the POC relationship between the current picture and the reference picture, and the like.

[0137] Further, according to another embodiment of the present invention, when adding an HMVP candidate from the HMVP table to the candidate list, the encoder / decoder can add the HMVP candidate to the candidate list with reference to the relationship between the current block and the block corresponding to the element in the HMVP table. Alternatively, when adding an HMVP candidate from the HMVP table to the candidate list, the encoder / decoder can refer to the positional relationship between the candidate blocks included in the HMVP table. For example, the encoder / decoder can add the HMVP candidate from the HMVP table to the candidate list in consideration of the decoding order of the HMVP or the block.

[0138] FIG. 25 is a diagram illustrating a method for processing a video signal based on HMVP according to an embodiment of the present invention. Referring to FIG. 25, for convenience of explanation, the decoder will be mainly described, but the present invention is not limited thereto, and the HMVP-based video signal processing method according to this embodiment is substantially similarly applicable to the encoder.

[0139] Specifically, when the current block is located in a merge sharing node including a plurality of coding blocks, the decoder can construct a merge candidate list using a spatial candidate adjacent to the merge sharing node (S2501). The decoder can add a specific HMVP in the HMVP table including at least one history-based motion vector predictor (HMVP) to the merge candidate list (S2502). Here, the HMVP represents the motion information of a block encoded before the plurality of coding blocks.

[0140] The decoder acquires index information indicating a merge candidate used for prediction of the current block within the merge candidate list (S2503), and generates a predicted block of the current block based on the motion information of the merge candidate (S2504). The decoder can generate a restored block of the current block by adding the predicted block and the residual block. As described above, according to one embodiment of the present invention, the motion information of at least one coding block among the plurality of coding blocks included in the merge sharing node may not be updated in the HMVP table. As described above, the decoder can update the HMVP table using the motion information of a predefined number of coding blocks with a relatively late decoding order among the plurality of coding blocks included in the merge sharing node.

[0141] Also, as described above, the decoder can update the HMVP table using the motion information of the coding block with the relatively latest decoding order among the plurality of coding blocks included in the merge sharing node. Further, when the current block is not located within the merge sharing node, the decoder can update the HMVP table using the motion information of the merge candidate. Also, as described above, the decoder can check whether there is motion information overlapping with the candidates in the merge candidate list using an HMVP having a specific index predefined in the HMVP table.

[0142] FIG. 26 is a diagram for explaining a multi-hypothesis prediction method according to an embodiment of the present invention. According to an embodiment of the present invention, an encoder / decoder can generate a prediction block based on a plurality of prediction methods. In this specification, a prediction method based on such a plurality of prediction modes is referred to as multi-hypothesis prediction. However, the present invention is not limited to such a name, and in this specification, the multi-hypothesis prediction may be referred to as multiple prediction, plural prediction, combined prediction, inter-intra weighted prediction, combined inter-intra prediction, combined inter-intra weighted prediction, etc. As an example, multi-hypothesis prediction can mean a block generated by any prediction method. Also, as an example, the prediction methods in multiple prediction can include methods such as intra prediction and inter prediction. Or, the prediction methods in multiple prediction may further mean, when subdivided, a merge mode, an AMVP mode, a specific mode of intra prediction, etc. Also, the encoder / decoder can generate a final prediction block by weighted-summing the prediction blocks (or prediction samples) generated based on multiple prediction.

[0143] According to an embodiment of the present invention, the maximum number of prediction methods used for multiple prediction may be set in advance. For example, the maximum number of multiple prediction may be 2. Therefore, the encoder / decoder can apply 2 predictions in the case of uni prediction, or 2 (i.e., when using multiple prediction only for prediction from one reference list) or 4 (i.e., when using multiple prediction for prediction from two reference lists) predictions in the case of bi prediction to generate a prediction block.

[0144] Alternatively, according to one embodiment of the present invention, prediction modes that can be used in multiple hypothesis prediction may be set in advance. Alternatively, combinations of prediction modes that can be used in multiple hypothesis prediction may be set in advance. For example, an encoder / decoder can perform multiple hypothesis prediction using prediction blocks (or prediction samples) generated by inter prediction and intra prediction.

[0145] According to one embodiment of the present invention, an encoder / decoder can use only some of the prediction modes among the inter prediction and / or intra prediction modes for multiple hypothesis prediction. For example, the encoder / decoder can use only the merge mode among the inter predictions for multiple hypothesis prediction. Alternatively, the encoder / decoder can use the merge mode instead of the subblock merge mode among the inter predictions for multiple hypothesis prediction. Alternatively, the encoder / decoder can use a specific intra mode among the intra prediction modes for multiple hypothesis prediction. For example, the encoder / decoder can restrictively use a prediction mode including at least one of planar, DC, vertical, and / or horizontal modes among the intra predictions for multiple hypothesis prediction. As an example, the encoder / decoder can generate a prediction block based on predictions of the merge mode and intra prediction, and at this time, only a limited prediction mode of at least one of planar, DC, vertical, and / or horizontal modes can be used for the intra prediction.

[0146] Referring to FIG. 26, the encoder / decoder can generate a prediction block using Prediction 1 and Prediction 2. Specifically, the encoder / decoder can apply Prediction 1 to generate a first temporary prediction block (or prediction sample), and apply Prediction 2 to generate a second temporary prediction block. The encoder / decoder can generate a final prediction block by performing a weighted sum of the first temporary prediction block and the second temporary prediction block. At this time, the encoder / decoder can apply a first weight value w1 to the first temporary prediction block generated by the first prediction, and apply a second weight value w2 to the second temporary prediction block generated by the second prediction to perform a weighted sum.

[0147] According to an embodiment of the present invention, when generating a prediction block based on multiple hypothesis prediction, the weight applied to the multiple hypothesis prediction may be determined based on a specific position within the block. At this time, the block may be the current block or a neighboring block. Alternatively, the weight of the multiple hypothesis prediction may be based on the mode of generating the prediction. For example, the encoder / decoder can determine the weight value based on the prediction mode when one of the modes of generating the prediction is intra prediction. Also, for example, when one of the prediction modes is intra prediction and is a directional mode, the encoder / decoder can increase the weight value for samples at positions far from the reference sample.

[0148] According to an embodiment of the present invention, when the intra prediction mode used for multi-hypothesis prediction is a directional mode and other prediction modes are inter prediction, a relatively high weighting value can be applied to the prediction sample generated based on the intra prediction on the side far from the reference sample. This is because in the case of inter prediction, motion compensation can be performed using spatial neighboring candidates, and in such a case, the probability that the motion of the current block and the spatial neighboring block referred to for motion compensation is the same or similar is high. As a result, the prediction of the region adjacent to the spatial neighboring block and the prediction of the region including an object with motion are more accurate than other parts. In this case, the residual signal adjacent to the boundary in the opposite direction of the spatial neighboring block may be left more than other regions (or parts). According to an embodiment of the present invention, this can be offset by combining and applying the samples intra-predicted in multi-hypothesis prediction. Also, in one embodiment, since the reference sample position of intra prediction can be in the vicinity of the spatial neighboring candidates of inter prediction, the encoder / decoder can apply a high weighting value to a relatively distant region therefrom.

[0149] As another example, when one of the modes used for generating multi-hypothesis prediction samples is intra prediction and is a directional mode, the encoder / decoder can apply a high weighting value to the samples located relatively close to the reference sample. More specifically, when one of the modes used for generating multi-hypothesis prediction samples is intra prediction and is a directional mode, and the modes used for generating other multi-hypothesis prediction samples are inter prediction, the encoder / decoder can apply a high weighting value to the prediction generated based on the intra prediction on the side close to the reference sample. This is because the closer the distance between the prediction sample and the reference sample in intra prediction, the higher the prediction accuracy.

[0150] As another example, when one of the modes used for generating multiple hypothesis prediction samples is intra prediction and it is not a directional mode (for example, when it is planar or DC mode), the weighting value may be set to a constant value regardless of the position within the block. Also, in one embodiment, the weighting value for prediction 2 in multiple hypothesis prediction may be determined based on the weighting value for prediction 1. The following equation represents an example of determining a prediction sample based on multiple hypothesis prediction.

[0151]

Equation

[0152] In Equation 5, pbSamples represents the (final) prediction sample (or prediction block) generated by multiple hypothesis prediction. And predSamples represents the block / sample generated by inter prediction, and predSamplesIntra represents the block / sample generated by intra prediction. In Equation 5, x and y represent the coordinates of the samples within the block, and may be in the following ranges: x = 0..nCbW - 1 and y = 0..nCbH - 1. Also, nCbW and nCbH may be the width and height of the current block, respectively. Also, in one embodiment, the weighting value w may be determined by the following process.

[0153] - If predModeIntra is INTRA_PLANAR or INTRA_DC, or nCbW < 4, or nCbH < 4, or cIdx > 0, w may be set to 4.

[0154] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and y < (nCbH / 4), w may be set to 6.

[0155] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and (nCbH / 4)<=y<(nCbH / 2), then w may be set to 5.

[0156] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and (nCbH / 2)<=y<(3*nCbH / 4), then w may be set to 4.

[0157] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and (3*nCbH / 4)<=y<nCbH, then w may be set to 3.

[0158] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and x<(nCbW / 4), then w may be set to 6.

[0159] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and (nCbW / 4)<=x<(nCbW / 2), then w may be set to 5.

[0160] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and (nCbW / 2)<=x<(3*nCbW / 4), then w may be set to 4.

[0161] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and (3*nCbW / 4)<=x<nCbW, then w may be set to 3.

[0162] The following Table 2 illustrates the multiple hypothesis prediction related syntax structure according to an embodiment of the present invention.

[0163]

Table 2

[0164] In Table 2, mh_intra_flag is a flag indicating whether to use multiple hypothesis prediction. According to an embodiment of the present invention, multiple hypothesis prediction may be applied only when a predefined specific condition for multiple hypothesis prediction (referred to as mh_conditions herein for convenience of explanation) is satisfied. If mh_conditions are not satisfied, the encoder / decoder may not parse mh_intra_flag and infer it as 0. For example, mh_conditions may include conditions related to the block size. Also, mh_conditions may include conditions related to whether to use a predefined specific mode. For example, mh_intra_flag can be parsed when merge_flag, which is a flag indicating whether the merge mode is applied, is 1 and subblock_merge_flag, which is a flag indicating whether the subblock merge mode is applied, is 0. In other words, the encoder / decoder can consider (or apply) multiple hypothesis prediction when the merge mode is applied to the current block and the subblock merge mode is not applied.

[0165] Also, according to an embodiment of the present invention, the encoder can divide candidate modes into multiple lists to determine the mode in multiple hypothesis prediction and signal to the decoder which list to use. Referring to Table 2, mh_intra_luma_mpm_flag may be a flag indicating which list among the multiple lists to use. If mh_intra_luma_mpm_flag does not exist, it can be inferred as 1 (or considered as 1). Also, as an embodiment of the present invention, the multiple lists may be an MPM list and a non-MPM list.

[0166] Also, as an example, the encoder can signal to the decoder an index (or index information) indicating which index candidate to use in the list among the plurality of lists. Referring to Table 2, mh_intra_luma_mpm_idx may be the above-mentioned index. Also, as an example, the index may be signaled only when a specific list is selected. The decoder can parse mh_intra_luma_mpm_idx only when the specific list is determined by mh_intra_luma_mpm_flag.

[0167] According to an embodiment of the present invention, as in the embodiment described with reference to FIG. 26, multiple hypothesis prediction can be performed based on the prediction generated by inter prediction and the prediction generated by intra prediction. As an example, the encoder / decoder can perform multiple hypothesis prediction only when signaled to use inter prediction. Alternatively, the encoder / decoder can perform multiple hypothesis prediction only when signaled to use a specific mode of inter prediction, such as the merge mode. In this case, signaling for inter prediction may not be required separately. Also, as an example, when the encoder / decoder generates a prediction by intra prediction, the number of candidate modes may be four in total. For example, among the four candidate modes in total, the first list and the second list can be configured using three and one candidate modes, respectively. At this time, when the second list including one prediction mode is selected, the encoder may not signal the index to the decoder. Also, when the first list is selected, the index indicating a specific candidate can be signaled to the decoder. In this case, since the number of candidates included in the first list is three, signaling can be performed with 1 bit or 2 bits in a variable length coding method.

[0168] The following Table 3 illustrates a multiple hypothesis prediction-related syntax structure according to an embodiment of the present invention.

[0169]

Table 3

[0170] In Table 3, as described in Table 2 above, signaling indicating whether to use a list among multiple lists can exist, and in Tables 2 and 3, mh_intra_luma_mpm_flag may be such a syntax element. Syntax elements overlapping with Table 2 described above in this regard are omitted from the description.

[0171] According to an embodiment of the present invention, signaling indicating which list to use may be explicitly signaled only in specific cases. If not explicitly signaled, the encoder / decoder can infer the value of the syntax element by a preset method. Referring to Table 3, when the condition of mh_mpm_infer_condition is satisfied, there is no explicit signaling, and when the condition of mh_mpm_infer_condition is not satisfied, explicit signaling may exist. Also, when the mh_mpm_infer_condition is satisfied and mh_intra_luma_mpm_flag does not exist, it may be inferred as 1 in that case. That is, in this case, the encoder / decoder can infer that the MPM list is used.

[0172] The following Table 4 exemplifies a multiple hypothesis prediction-related syntax structure according to an embodiment of the present invention.

[0173]

Table 4

[0174] As described in Tables 2 and 3 above, signaling indicating which list to use among multiple lists can exist, and when a predefined condition is satisfied, the encoder / decoder can infer the value. Syntax elements overlapping with Tables 2 and 3 described above in this regard are omitted from the description.

[0175] According to an embodiment of the present invention, the condition for inferring a signaling (or syntax element, parameter) value indicating which list among a plurality of lists to use may be determined based on the current block size. For example, the encoder / decoder may be determined based on the width and height of the current block. Specifically, the encoder / decoder can infer the signaling value when the larger one of the width and height of the current block is greater than n times the smaller one. For example, n may be set to a natural number value such as 2, 3, 4, etc. Referring to Table 4, the condition for inferring the signaling value indicating which list among a plurality of lists to use may be that the larger one of the width and height of the current block is greater than twice the smaller one. Assuming that the width and height of the current block are cbWidth and cbHeight respectively, the Abs(Log2(cbWidth / cbHeight)) value is 0 when cbWidth and cbHeight are the same, and 1 when the difference is twice. Therefore, when the difference between cbWidth and cbHeight is greater than twice, the Abs(Log2(cbWidth / cbHeight)) value is greater than 1 (i.e., can have a value of 2 or more).

[0176] FIG. 27 is a diagram showing a multiple hypothesis prediction mode determination method according to an embodiment of the present invention. As described in Tables 2 to 4 above, the determination of the prediction mode used for multiple hypothesis prediction may be performed based on a plurality of lists. As an example, the prediction mode can indicate an intra mode that generates a prediction based on intra prediction. Further, the plurality of lists may include two lists, a first list and a second list. Referring to FIG. 27, it is possible to determine whether to use the first list from list1_flag. In one embodiment, there may be a plurality of candidates belonging to the first list, and one candidate belonging to the second list.

[0177] If the list1_flag is inferred, the encoder / decoder can infer that the first list is to be used and its value (S2701). In this case, the list1_index, which is an index indicating which candidate to use in the first list, can be parsed (S2704). Also, if the list1_flag is not inferred (S2701), the encoder / decoder can parse the list1_flag (S2702). The encoder / decoder can parse the list1_index if the list1_flag is 1, and does not have to parse the index if the list1_flag is not 1. Also, if the list1_flag is 1, the mode actually used among the candidate modes of the first list can be determined based on the index (S2703). Also, the encoder / decoder can determine the mode actually used as the candidate mode of the second list without an index if the list1_flag is not 1. That is, the mode may be determined based on the flag and index in the first list, and the mode may be determined based on the flag in the second list.

[0178] FIG. 28 is a diagram showing a multiple hypothesis prediction mode determination method according to an embodiment of the present invention. According to an embodiment of the present invention, when variable-length coding an index for determining a candidate mode in a list, a method for determining the mode order included in the candidate list may be applied to increase coding efficiency. For example, there may be a method for determining the mode order included in the first list. At this time, the encoder / decoder can refer to the modes around the current block to determine the mode order. Also, the second list can be determined without referring to the modes around the current block. For example, the encoder / decoder can generate the first list by referring to the modes around the current block and include the modes not included in the first list in the second list. In one embodiment, the first list may be the MPM mode and the second list may be the non-MPM mode. Also, the total number of candidate modes may be four, with three modes included in the first list and one mode included in the second list.

[0179] Referring to FIG. 28, there may be a syntax element list1_flag indicating whether the first list is used. If the first list is used, the encoder / decoder can generate the first list (S2801, S2802) and select a specific mode from the first list. At this time, the generation of the first list and the confirmation of whether the first list is used may be performed in any order. As an example, in a situation where the first list is used, the first list may be generated before or after the confirmation of whether the first list is used. Also, when the first list is used, the encoder / decoder may not perform the process of generating the second list.

[0180] If the first list is not used, the encoder / decoder can generate the second list (S2803) and select a specific mode from the second list. At this time, the encoder / decoder can generate the first list to generate the second list. Then, among the candidate modes, the encoder / decoder can include candidates not included in the first list in the second list. Also, according to an embodiment of the present invention, the first list generation method may be the same regardless of whether the first list is used (list1_flag value), whether the first list is used, or whether inference is performed. At this time, the methods described above in Tables 2 to 4 and FIG. 27 may be applied to list signaling and mode signaling.

[0181] Hereinafter, a plurality of list configuration (or generation) methods for determining a prediction mode used in the multiple hypothesis prediction described with reference to FIGS. 27 to 28 above will be further described. As an example, as described above, the plurality of lists may be composed of two lists. That is, the plurality of lists may include a first list and a second list. Also, the plurality of lists may be used in the multiple hypothesis prediction process.

[0182] According to an embodiment of the present invention, an encoder / decoder can generate a plurality of lists by referring to the modes around the current block. Further, the encoder / decoder performs intra prediction using the mode selected from the list, and combines the prediction sample (or, prediction block) generated by the intra prediction with the inter-predicted prediction sample to generate a multiple hypothesis prediction block. In this specification, the final prediction sample (or, prediction block) generated by multiple hypothesis prediction is referred to as a multiple hypothesis prediction block, but the present invention is not limited thereto. For example, the multiple hypothesis prediction block can be referred to as a prediction block, a final prediction block, a multiple prediction block, a combined prediction block, an inter-intra weighted prediction block, a combined inter-intra prediction block, a combined inter-intra weighted prediction block, etc.

[0183] As an embodiment, the mode (candidate mode) that may be included in the list may be set to at least one of a planar mode, a DC mode, a vertical mode, and / or a horizontal mode of an intra prediction method. At this time, the vertical mode may be the mode with index (or, mode number) 50 in FIG. 6 described above, and the horizontal mode may be the mode with index 18 in FIG. 6. Further, the planar mode and the DC mode may be indices 0 and 1, respectively.

[0184] According to an embodiment of the present invention, a candidate mode list can be generated by referring to the prediction modes of the surrounding blocks of the current block. Further, the candidate mode list can be referred to as candModeList in this specification. As an example, the candidate mode list may be the first list described in the above embodiments. Further, as an embodiment, the mode of the surrounding blocks of the current block may be expressed as candIntraPredModeX. That is, candIntraPredModeX represents a variable indicating the mode of the surrounding block. Here, X represents a variable indicating a specific position around the current block such as A or B.

[0185] As an example, the encoder / decoder can generate a candidate mode list based on the presence or absence of a match among a plurality of candIntraPredModeX. For example, candIntraPredModeX can exist for two positions, and the modes of the positions may be expressed as candIntraPredModeA and candIntraPredModeB. As an example, if candIntraPredModeA and candIntraPredModeB are the same, the candidate mode list can include a planar mode and a DC mode.

[0186] As an example, when candIntraPredModeA and candIntraPredModeB are the same and their values indicate the planar mode or the DC mode, the encoder / decoder can add the modes indicated by candIntraPredModeA and candIntraPredModeB to the candidate mode list. Also, the encoder / decoder can add to the candidate mode list a mode that is not indicated by candIntraPredModeA and candIntraPredModeB among the planar mode and the DC mode. Further, the encoder / decoder can add to the candidate mode list a mode that has already been set and is not the planar mode or the DC mode. As an example, in this case, the order within the candidate mode list of the planar mode, the DC mode, and the specific mode that has already been set may already be set. For example, it may be the order of planar, DC, and the mode that has already been set. That is, candModeList[0] = planar mode, candModeList[1] = DC mode, and candModeList[2] = the mode that has already been set may be possible. Also, the mode that has already been set may be the vertical mode. As yet another example, among the planar mode, the DC mode, and the mode that has already been set, the mode indicated by candIntraPredModeA and candIntraPredModeB may be added to the candidate mode list first, a mode that is not indicated by candIntraPredModeA and candIntraPredModeB among the planar mode and the DC mode may be added as the next candidate, and the mode that has already been set may subsequently be added.

[0187] Also, if candIntraPredModeA and candIntraPredModeB are the same and their values do not indicate the planar mode and the DC mode, the encoder / decoder can add the modes indicated by candIntraPredModeA and candIntraPredModeB to the candidate mode list. Also, the planar mode and the DC mode may be added to the candidate mode list. Also, in this case, the order of the modes indicated by candIntraPredModeA and candIntraPredModeB and the planar mode and the DC mode in the candidate mode list may already be set. Also, the already set order may be the order of the modes indicated by candIntraPredModeA and candIntraPredModeB, the planar mode, and the DC mode. That is, candModeList[0]=candIntraPredModeA, candModeList[1]=planar mode, and candModeList[2]=DC mode may be possible.

[0188] Also, if candIntraPredModeA and candIntraPredModeB are different from each other, the encoder / decoder can add both candIntraPredModeA and candIntraPredModeB to the candidate mode list. Also, candIntraPredModeA and candIntraPredModeB may be included in the candidate mode list according to a predefined specific order. For example, they may be included in the candidate mode list in the order of candIntraPredModeA, candIntraPredModeB. Also, there may already be a set order among the candidate modes, and the encoder / decoder can add modes other than candIntraPredModeA and candIntraPredModeB to the candidate mode list according to the already set order. Also, the modes other than candIntraPredModeA and candIntraPredModeB may be added after candIntraPredModeA and candIntraPredModeB in the candidate mode list. Also, the already set order may be the planar mode, the DC mode, the vertical mode. Or, the already set order may be the planar mode, the DC mode, the vertical mode, the horizontal mode. That is, candModeList[0]=candIntraPredModeA, candModeList[1]=candIntraPredModeB may be fine, and candModeList[2] may be the mode that is the foremost among the modes that are not candIntraPredModeA and not candIntraPredModeB among the planar mode, the DC mode, and the vertical mode.

[0189] Also, in one embodiment, among the candidate modes, a mode not included in the candidate mode list may be defined as candIntraPredModeC. As an example, candIntraPredModeC may be included in the second list. Further, when the signaling indicating whether the aforementioned first list is used indicates not to use it, candIntraPredModeC can be determined. When using the first list, the encoder / decoder can determine the mode by the index among the candidate mode lists, and when not using the first list, can use the mode of the second list.

[0190] Also, as described above, after generating the candidate mode list, a process of modifying the candidate mode list may be added. For example, the encoder / decoder may or may not further perform the process of modification according to the current block size condition. For example, the current block size condition may be determined based on the width and height of the current block. For example, when the larger one of the width and height of the current block is larger than n times the other one, the process of modification can be further performed. For example, the n may be defined as 2.

[0191] Also, in one embodiment, the process of modification may be a process of changing a mode in the candidate mode list to another mode when any mode is included in the candidate mode list. For example, when the vertical mode is included in the candidate mode list, the encoder / decoder may insert the horizontal mode into the candidate mode list instead of the vertical mode. Or, when the vertical mode is included in the candidate mode list, the encoder / decoder may insert the candIntraPredModeC into the candidate mode list instead of the vertical mode. Or, since the planar mode and the DC mode may always be included in the candidate mode list when generating the candidate mode list described above, in this case, candIntraPredModeC may be the horizontal mode. Also, using such a modification process may be when the height of the current block is greater than n times the width. For example, n may be defined as 2. This is because when the height is greater than the width, the lower side of the block may be far from the reference samples for intra prediction, so the accuracy of the vertical mode may be low. Or, using such a modification process may be when it is inferred that the first list is used.

[0192] Also, according to an embodiment of the present invention, as another example of the candidate list modification process, when the horizontal mode is included in the candidate mode list, the encoder / decoder may add the vertical mode to the candidate mode list instead of the horizontal mode. Or, when the horizontal mode is included in the candidate mode list, the encoder / decoder may add the candIntraPredModeC to the candidate mode list instead of the horizontal mode. Or, since the planar mode and the DC mode may always be included in the candidate mode list when generating the candidate mode list described above, in this case, candIntraPredModeC may be the vertical mode. Also, using such a modification process may be when the width of the current block is greater than n times the height. For example, n may be defined as 2. This is because when the width is greater than the height, the right side of the block may be far from the reference samples for intra prediction, so the accuracy of the horizontal mode may be low. Or, using such a modification process may be when it is inferred that the first list is used.

[0193] Next, an example of the list setting method described above will be further described. In this specification, IntraPredModeY indicates a mode used for intra prediction during multiple hypothesis prediction. Also, IntraPredModeY can indicate the mode of the luminance component. As an example, in multiple hypothesis prediction, the intra prediction mode of the chrominance component may be derived from the luminance component. Also, in this specification, mh_intra_luma_mpm_flag represents a variable (or syntax element) indicating which list among a plurality of lists to use. For example, mh_intra_luma_mpm_flag may be the mh_intra_luma_mpm_flag in Tables 2 to 4, and list1_flag in FIGS. 27 and 28. Also, in this specification, mh_intra_luma_mpm_idx represents an index indicating which candidate to use in the list. For example, mh_intra_luma_mpm_idx may be the mh_intra_luma_mpm_idx in Tables 2 to 4, and list1_index in FIG. 27. Also, in this specification, xCb and yCb may be the x and y coordinates of the top-left of the current block. Also, cbWidth and cbHeight may be the width and height of the current block.

[0194] FIG. 29 is a diagram showing peripheral positions referred to in multiple hypothesis prediction according to an embodiment of the present invention. Referring to FIG. 29, as described above, the encoder / decoder can refer to peripheral positions in the process of creating a candidate list for multiple hypothesis prediction. For example, the aforementioned candIntraPredModeX may be required. At this time, the positions of A and B around the current block to be referred to may be NbA and NbB shown in FIG. 29. That is, they may be the left and upper positions adjacent to the top-left sample of the current block. If the top-left position of the current block is Cb as shown in FIG. 18 and its coordinates are (xCb, yCb), then NbA may be (xNbA, yNbA) = (xCb - 1, yCb), and NbB may be (xNbB, yNbB) = (xCb, yCb - 1).

[0195] FIG. 30 is a diagram showing a method of referring to a surrounding mode according to an embodiment of the present invention. Referring to FIG. 30, as described above, the encoder / decoder can refer to a surrounding position in the process of creating a candidate list for multi-hypothesis prediction. Also, a surrounding mode can be used as it is, or a mode based on the surrounding mode can be used to generate a candidate mode list. As an example, the mode generated by referring to the surrounding position may be candIntraPredModeX. As an embodiment, when the surrounding position is unusable, candIntraPredModeX may be a previously set mode. When it is unusable, it can include cases where the surrounding position uses inter prediction, and the mode has not been determined in the defined encoding / decoding order. Or, when the surrounding position does not use multi-hypothesis prediction, candIntraPredModeX may be a previously set mode. Or, when the surrounding position is above the CTU to which the current block belongs, candIntraPredModeX may be a previously set mode. As another example, when the surrounding position is outside the CTU to which the current block belongs, candIntraPredModeX may be a previously set mode. Also, as an embodiment, the previously set mode may be a DC mode. As yet another embodiment, the previously set mode may be a planar mode.

[0196] According to an embodiment of the present invention, candIntraPredModeX may be set depending on whether the mode of the peripheral position exceeds a threshold angle or whether the index of the mode of the peripheral position exceeds a threshold. For example, if the index of the mode of the peripheral position is greater than a specific directional mode index, the encoder / decoder may set candIntraPredModeX to the vertical mode index. Also, if the index of the mode of the peripheral position is less than or equal to the specific directional mode index and it is a directional mode, the encoder / decoder may set candIntraPredModeX to the horizontal mode index. For example, the specific directional mode index may be mode 34 in FIG. 6 described above. Also, if the mode of the peripheral position is a planar mode or a DC mode, candIntraPredModeX may be set to the planar mode or the DC mode.

[0197] Referring to FIG. 30, mh_intra_flag may be a syntax element (or variable, parameter) indicating whether multiple hypothesis prediction is used (or has been used). Also, the intra prediction mode used in the peripheral block may be X. Also, the current block can use multiple hypothesis prediction and generate a candidate list using a candidate intra prediction mode based on the mode of the peripheral block. However, since the periphery did not use multiple hypothesis prediction, regardless of the intra prediction mode of the peripheral block used or whether the peripheral block used intra prediction or not, the encoder / decoder can set the candidate intra prediction mode to the DC mode, which is the already set mode.

[0198] An example of the peripheral mode reference method described above will be further described below.

[0199] According to an embodiment of the present invention, the intra prediction mode candIntraPredModeX (X is A or B) of the peripheral block can be derived in the following manner.

[0200] 1. The availability guidance process for the block specified in the adjacent block availability confirmation process is called, and the availability guidance process can set the position (xCurr, yCurr) to (xCb, yCb), set the input (xNbX, yNbX) to the adjacent position (xNbY, yNbY), and the output can be assigned to the availability variable availableX.

[0201] 2. The candidate intra prediction mode candIntraPredModeX may be derived in the following way:

[0202] A. If one or more of the following conditions are true, candIntraPredModeX may be set to the INTRA_DC mode.

[0203] a) When the variable availableX is FALSE.

[0204] b) When mh_intra_flag[xNbX][yNbX] is not 1.

[0205] c) When X is B and yCb - 1 is less than ((yCb >> CtbLog2SizeY) << CtbLog2SizeY).

[0206] B. Otherwise, if IntraPredModeY[xNbX][yNbX] > INTRA_ANGULAR34, candIntraPredModeX may be set to INTRA_ANGULAR50.

[0207] C. Otherwise, if IntraPredModeY[xNbX][yNbX] <= INTRA_ANGULAR34 and IntraPredModeY[xNbX][yNbX] > INTRA_DC, candIntraPredModeX may be set to INTRA_ANGULAR18.

[0208] D. Otherwise, candIntraPredModeX may be set equal to IntraPredModeY[xNbX][yNbX].

[0209] As described above, in the list setting method described above, candIntraPredModeX may be determined by the peripheral mode reference method.

[0210] FIG. 31 is a diagram showing a candidate list generation method according to an embodiment of the present invention. According to the first list and second list generation methods described in Tables 2 to 4 above, the encoder / decoder generates a first list by referring to the modes around the current block, and can generate a second list using the modes not included in the first list among the candidate modes. Since there is spatial similarity within the picture, those referring to the surrounding modes may have a higher priority. That is, the first list may have a higher priority than the second list. However, according to the first list and second list signaling methods described in Tables 2 to 4, when the signaling for determining the list is not inferred, the encoder / decoder can use a flag and an index for signaling to use the modes of the first list, and can use only the flag for signaling to use the modes of the second list. That is, the encoder / decoder can use relatively fewer bits for signaling the second list. However, using relatively more bits for signaling the modes of the list with a higher priority may not be good in terms of coding efficiency. Therefore, according to an embodiment of the present invention, a method of using signaling with relatively fewer bits for the list and modes with a higher priority is proposed.

[0211] That is, according to an embodiment of the present invention, a candidate list generation method may be individually defined depending on whether only a list with a relatively high priority can be used. Whether only the first list can be used can indicate whether signaling indicating the list to be used is inferred. For example, assuming that there is a second list generated by a method already set using a candidate mode, the encoder / decoder can divide and insert the third list into the first list and the second list. For example, for the third list generated by a method already set and the generation method of the third list, the candidate mode list and its generation method described above may be applied. If signaling indicating the list to be used is inferred, only the first list can be used. In that case, the encoder / decoder can fill the first list from the front of the third list. Also, if signaling indicating the list to be used is not inferred, the first list or the second list may be used. In that case, the second list can be filled from the front of the third list, and the remainder can be filled with the first list. At this time, the encoder / decoder can also fill the first list in the order of the third list when filling the first list. That is, candIntraPredModeX can be put into the candidate list by referring to the mode around the current block, but candIntraPredModeX can be put into the first list when signaling indicating the list is inferred, and added to the second list when not inferred.

[0212] In one embodiment, the size of the second list may be 1. In this case, when the signaling indicating the list is inferred, the encoder / decoder can add candIntraPredModeA to the first list, and when it is not inferred, it can add it to the second list. candIntraPredModeA may be list3[0] which is the mode at the head of the third list. Therefore, in the present invention, in some cases, candIntraPredModeA, which is a mode based on the neighboring mode, may be added to both the first list and the second list. On the other hand, in the methods described in Tables 2 to 4, candIntraPredModeA could only be put in the first list. That is, in the present invention, the first list generation method may be individually set according to whether the signaling indicating the list to be used is inferred or not.

[0213] Referring to FIG. 31, the candidate mode may be a candidate that can be used to generate an intra prediction of multiple hypothesis prediction. That is, the candidate list generation method may be different depending on whether list1_flag, which is the signaling indicating the list to be used, is inferred (S3101, S3102). If it is inferred, it can be inferred that the first list is used and only the first list can be used, so the encoder / decoder can add from the top of the third list to the first list (S3103). When generating the third list, the encoder / decoder can add the mode based on the neighboring mode to the top. Also, if it is not inferred, both the first list and the second list may be used, so the encoder / decoder can add from the top of the third list to the second list with less signaling (S3104). And when the first list is required, for example, when it is signaled that the first list is used, the encoder / decoder can add to the first list excluding the candidates included in the second list from the third list (S3105). In the present invention, the case of generating the third list is mainly described for convenience of explanation, but it is not limited to this, and candidates can be temporarily classified and the first list and the second list can be generated based on this.

[0214] According to an embodiment of the present invention, the method of generating a candidate list by the embodiments described below and the method of generating the candidate list described in FIGS. 27 to 28 can be adaptively used depending on the case. For example, the encoder / decoder can select either one of the two methods of creating two candidate lists depending on whether signaling indicating which list to use is inferred. Also, this may be the case of multiple hypothesis prediction. Further, the following first list may include three modes, and the second list may include one mode. Also, as described in FIGS. 27 to 28 above, the encoder / decoder can signal the mode of the first list with a flag and an index, and signal the mode of the second list with a flag.

[0215] As an example, when candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is the planar mode or the DC mode, it may be determined that List2[0]=planar mode, List1[0]=DC mode, List1[1]=vertical mode, and List1[2]=horizontal mode. As yet another example, when candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is the planar mode or the DC mode, it may be determined that List2[0]=candIntraPredModeA, List1[0]=!candIntraPredModeA, List1[1]=vertical mode, and List1[2]=horizontal mode.

[0216] Also, as an example, when candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is a directional mode, it may be determined that List2[0] = candIntraPredModeA, List1[0] = the planar mode, List1[1] = the DC mode, and List1[2] = the vertical mode. Also, as an example, when candIntraPredModeA and candIntraPredModeB are different, List2[0] = candIntraPredModeA and List1[0] = candIntraPredModeB may be used. List1[1] and List1[2] may be added from those that are not candIntraPredModeA and candIntraPredModeB among the planar mode, the DC mode, the vertical mode, and the horizontal mode.

[0217] FIG. 32 is a diagram showing a candidate list generation method according to an embodiment of the present invention. In the above embodiments, a method for determining a mode based on a plurality of lists has been described. Referring to FIG. 32, according to an embodiment of the present invention, the mode may be determined based on a single list instead of a plurality of lists. Specifically, as shown in FIG. 32, one candidate list including all candidate modes for multiple hypothesis prediction may be generated. The following Table 5 illustrates a multiple hypothesis prediction related syntax structure according to an embodiment of the present invention.

[0218]

Table 5

[0219] Referring to Table 5, since there is one candidate list, there is no signaling to select the list, and there may be index signaling indicating which mode among the candidate list modes to use. Therefore, when the mh_intra_flag indicating the presence or absence of using multiple hypothesis prediction is 1, the decoder can parse the mh_intra_luma_idx which is the candidate index. According to an embodiment of the present invention, the method for generating a candidate list for multiple hypothesis prediction can be based on the MPM list generation method in existing intra prediction. Or, according to an embodiment, the candidate list for multiple hypothesis prediction may be composed of a list in a form where the first list and the second list are combined in the order of the first list and the second list described in FIG. 28 above.

[0220] That is, if the candidate list for multiple hypothesis prediction is regarded as the candidate mode list, the size of the candidate mode list may be 4 in this embodiment. If candIntraPredModeA and candIntraPredModeB are the same and are the planar mode or the DC mode, the candidate mode list may be determined according to the set order. For example, candModeList[0]=planar mode, candModeList[1]=DC mode, candModeList[2]=vertical mode, candModeList[3]=horizontal mode may be possible. As another example, if candIntraPredModeA and candIntraPredModeB are the same and are the planar mode or the DC mode, candModeList[0]=candIntraPredModeA, candModeList[1]=!candIntraPredModeA, candModeList[2]=vertical mode, candModeList[3]=horizontal mode may be determined.

[0221] Alternatively, if candIntraPredModeA and candIntraPredModeB are the same and are directional modes, candModeList[0]=candIntraPredModeA, candModeList[1]=planar mode, candModeList[2]=DC mode, candModeList[3]=candIntraPredModeA, and modes other than the planar mode and the DC mode may be determined. If candIntraPredModeA and candIntraPredModeB are different, candModeList[0]=candIntraPredModeA and candModeList[1]=candIntraPredModeB may be used. Also, candModeList[2] and candModeList[3] can sequentially add modes other than candIntraPredModeA and candIntraPredModeB according to the already set order of candidate modes. The already set order may be determined as the planar mode, the DC mode, the vertical mode, and the horizontal mode.

[0222] According to an embodiment of the present invention, the candidate list may change according to the block size condition. If, among the block width and height, the relatively larger one is greater than n times the other, the candidate list may be shorter. For example, if the width is greater than n times the height, the encoder / decoder can remove the horizontal mode from the candidate list described in FIG. 31 and move the next mode to the front to satisfy it. Also, if the height is greater than n times the width, the encoder / decoder can remove the vertical mode from the candidate list described in FIG. 32 and move the next mode to the front to satisfy it. Therefore, when the width is greater than n times the height, the candidate list size may be 3. Also, when the width is greater than n times the height, the candidate list size may be smaller than or equal to that in the case where it is not.

[0223] As an example, the candidate index of the embodiment described with reference to FIG. 32 may be variable length coded. This can increase the signaling efficiency by adding modes with relatively high probabilities of use earlier in the list. As another example, the candidate index in the embodiment described with reference to FIG. 32 may be fixed length coded. The number of modes used in multiple hypothesis prediction may be a power of 2. For example, as described above, it can be used among 4 intra prediction modes. In such a case, no values that cannot be assigned occur even with fixed length coding, and unnecessary parts in signaling do not occur. Also, when fixed length coded, the number of cases for list configuration may be 1. The number of bits is the same regardless of which index is signaled.

[0224] According to one embodiment, the candidate index may be either variable length coded or fixed length coded depending on the case. For example, as in the above embodiment, the candidate list size may vary depending on the case. As an example, the candidate index may be variable length coded or fixed length coded depending on the candidate list size. For example, when the candidate list size is a power of 2, it may be fixed length coded, and when it is not a power of 2, it may be variable length coded. That is, according to the above embodiment, the coding method may vary depending on the block size condition.

[0225] According to an embodiment of the present invention, when using multiple hypothesis prediction, when using the DC mode, since the weighting values between multiple predictions for the entire block can be the same, it is possible to achieve the same / similar results as adjusting the weighting value of the prediction block. Therefore, the DC mode can be excluded in multiple hypothesis prediction. As an example, the encoder / decoder can use only one of the planar mode, vertical mode, and horizontal mode in multiple hypothesis prediction. In such a case, as shown in FIG. 32, the encoder / decoder can signal multiple hypothesis prediction using one list. Also, the encoder / decoder can use variable length coding for index signaling. As an example, a list can be generated in a fixed order. For example, it can be in the order of planar mode, vertical mode, and horizontal mode.

[0226] As yet another example, the encoder / decoder can generate a list by referring to the modes around the current block. For example, if candIntraPredModeA and candIntraPredModeB are the same, candModeList[0] may be candIntraPredModeA. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is the planar mode, candModeList[1] and candModeList[2] can be set according to the already set order. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is not the planar mode, candModeList[1] = the planar mode, candModeList[2] may be a mode other than candIntraPredModeA and other than the planar mode. If candIntraPredModeA and candIntraPredModeB are different, candModeList[0] = candIntraPredModeA, candModeList[1] = candIntraPredModeB, and candModeList[2] may be a mode other than candIntraPredModeA and other than candIntraPredModeB.

[0227] As another example, the encoder / decoder can use only one of three modes in multiple hypothesis prediction. The three modes can include a planar mode and a DC mode. Also, the three modes can conditionally include either a vertical mode or a horizontal mode. The condition may be a condition related to the block size. For example, depending on which of the width and height of the block is larger, it can be determined whether to include the horizontal mode or the vertical mode. For example, when the width of the block is larger than the height, the vertical mode can be included. When the height of the block is larger than the width, the horizontal mode can be included. When the height and width of the block are equal, one of the vertical mode or the horizontal mode, a predefined specific mode can be included.

[0228] Also, in one embodiment, the encoder / decoder can generate a list in a fixed order. For example, it may be defined in the order of planar mode, DC mode, vertical or horizontal mode. Also, as another example, a list can be generated by referring to the modes around the current block. For example, if candIntraPredModeA and candIntraPredModeB are the same, candModeList[0] = candIntraPredModeA may be true. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is not a directional mode, the encoder / decoder can set candModeList[1] and candModeList[2] according to the already set order. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is a directional mode, candModeList[1] = planar mode, candModeList[2] = DC mode may be true. If candIntraPredModeA and candIntraPredModeB are different, candModeList[0] = candIntraPredModeA, candModeList[1] = candIntraPredModeB, and candModeList[2] may be a mode other than candIntraPredModeA and candIntraPredModeB.

[0229] According to another embodiment of the present invention, only one of two modes can be used in multiple hypothesis prediction. The two modes can include a planar mode. Also, the two modes can conditionally include either a vertical mode or a horizontal mode. The condition may be a condition related to the block size. For example, the encoder / decoder can determine whether to include the horizontal mode or the vertical mode depending on which of the width and height of the block is larger. For example, if the width of the block is larger than the height, the vertical mode can be included. If the height of the block is larger than the width, the horizontal mode can be included. If the height and width of the block are equal, either the vertical mode or the horizontal mode that is agreed upon can be included. In such a case, a flag indicating which mode to use in multiple hypothesis prediction may be signaled.

[0230] According to one embodiment, the encoder / decoder can exclude a specific mode according to the block size. For example, when the block size is small, a specific mode can be excluded. For example, when the block size is small, only the planar mode can be used in multiple hypothesis prediction. If a specific mode is excluded, mode signaling can be omitted or reduced for signaling.

[0231] According to another embodiment of the present invention, the encoder / decoder can use only one mode (i.e., the intra prediction mode) to generate samples intra-predicted in multiple hypothesis prediction. As an example, the one mode may be defined as the planar mode. As described above, multiple hypothesis prediction can use inter-predicted samples and intra-predicted samples. When generating inter-predicted samples and at the same time determining the optimal prediction mode to generate intra-predicted samples, the prediction accuracy can be improved, but it may lead to the problem of increased encoding complexity and increased signaling bits. Therefore, when performing multiple hypothesis prediction, by using only the planar mode that statistically occurs most frequently as the intra prediction mode, the encoding complexity can be improved, signaling bits can be saved, and thereby the compression performance of the video can be increased.

[0232] As yet another example, the one mode may be determined based on the block size, from among the vertical mode and the horizontal mode. For example, it may be determined as either the vertical mode or the horizontal mode depending on which of the width and height of the block is larger. For example, if the width of the block is larger than the height, it may be determined as the vertical mode, and if the height of the block is larger than the width, it may be determined as the horizontal mode. If the width and height of the block are equal, the encoder / decoder may be determined as the already set mode. If the width and height of the block are equal, the encoder / decoder can be determined as the already set mode, from among the horizontal mode or the vertical mode. If the width and height of the block are equal, the encoder / decoder can be determined as the already set mode, from among the planar mode or the DC mode.

[0233] Also, according to an embodiment of the present invention, there may be flipping signaling that flips the prediction generated by multiple hypothesis prediction. By this, even if one mode is selected in multiple hypothesis prediction, it can have the effect of eliminating the residual on the opposite side by flipping. Also, by this, it can have the effect of reducing candidate modes available for multiple hypothesis prediction. More specifically, for example, among the above embodiments, flipping can be used when only one mode is used. By this, prediction performance can be improved. The flipping can mean flipping with respect to the x-axis, flipping with respect to the y-axis, or flipping with respect to both the x and y axes. As an example, the flipping direction may be determined based on the mode selected by multiple hypothesis prediction. For example, when the mode selected by multiple hypothesis prediction is a planar mode, the encoder / decoder can determine that it is flipping with respect to both the x and y axes. Also, flipping with respect to both the x and y axes may be due to the block size or shape. For example, when the block is not square, the encoder / decoder can determine not to flip with respect to both the x and y axes. For example, when the mode selected by multiple hypothesis prediction is a horizontal mode, the encoder / decoder can determine that it is flipping with respect to the x-axis. For example, when the mode selected by multiple hypothesis prediction is a vertical mode, the encoder / decoder can determine that it is flipping with respect to the y-axis. Also, when the mode selected by multiple hypothesis prediction is a DC mode, the encoder / decoder can determine that there is no flipping and may not perform explicit signaling / parsing.

[0234] Also, in multiple hypothesis prediction, the DC mode can have an effect similar to illumination compensation. Therefore, according to an embodiment of the present invention, if either the DC mode or the illumination compensation method is used in multiple hypothesis prediction, the other does not have to be used. Also, multiple hypothesis prediction can have an effect similar to GBi (generalized bi-prediction). For example, in multiple hypothesis prediction, the DC mode can have an effect similar to GBi. The GBi prediction may be a technique for adjusting the weighting value between two reference blocks of bidirectional prediction in units of blocks or CUs. Therefore, according to an embodiment of the present invention, if either multiple hypothesis prediction (or the DC mode in multiple hypothesis prediction) or the GBi prediction method is used, the other does not have to be used. Also, this may include the case where the prediction of multiple hypothesis prediction is bidirectional prediction. For example, when the selected merge candidate of multiple hypothesis prediction is bidirectional prediction, the GBi prediction does not have to be used. The relationship between multiple hypothesis prediction and GBi prediction in these embodiments may be limited to when a specific mode of multiple hypothesis prediction, for example, the DC mode, is used. Or, if the GBi prediction-related signaling exists before the multiple hypothesis prediction-related signaling, if the GBi prediction is used, multiple hypothesis prediction or a specific mode of multiple hypothesis prediction does not have to be used. In the present invention, not using any of the methods can mean not signaling for any of the above methods and not parsing the related syntax.

[0235] FIG. 33 is a diagram showing peripheral positions referred to in multiple hypothesis prediction according to an embodiment of the present invention. As described above, the encoder / decoder can refer to the peripheral positions in the process of creating a candidate list for multiple hypothesis prediction. For example, the aforementioned candIntraPredModeX may be used. At this time, the positions of A and B around the current block to be referred to may be NbA and NbB shown in FIG. 33. If the position of the upper left corner sample of the current block is Cb as shown in FIG. 29 and its coordinates are (xCb, yCb), then NbA may be (xNbA, yNbA) = (xCb - 1, yCb + cbHeight - 1), and NbB may be (xNbB, yNbB) = (xCb + cbWidth - 1, yCb - 1). Here, cbWidth and cbHeight may be the width and height of the current block, respectively. Also, in the process of creating a candidate list for multiple hypothesis prediction, the peripheral positions may be the same as the peripheral positions referred to in the generation of the MPM list for intra prediction.

[0236] As yet another embodiment, the peripheral positions referred to in the process of creating a candidate list for multiple hypothesis prediction may be the left central position and the upper central position of the current block or positions close thereto. For example, NbA and NbB may be (xCb - 1, yCb + cbHeight / 2 - 1), (xCb + cbWidth / 2 - 1, yCb - 1). Or, NbA and NbB may be (xCb - 1, yCb + cbHeight / 2), (xCb + cbWidth / 2, yCb - 1).

[0237] FIG. 34 is a diagram showing a method of referring to a surrounding mode according to an embodiment of the present invention. As described above, the surrounding position can be referred to in the process of creating a candidate list for multiple hypothesis prediction. However, in the embodiment of FIG. 30, when the surrounding position does not use multiple hypothesis prediction, candIntraPredModeX is set to the already set mode. This is probably because when setting candIntraPredModeX, the mode of the surrounding position may not directly become candIntraPredModeX. Therefore, according to an embodiment of the present invention, even when the surrounding position does not use multiple hypothesis prediction, if the mode used by the surrounding position is a mode used for multiple hypothesis prediction, the encoder / decoder can set candIntraPredModeX to the mode used by the surrounding position. The modes used for multiple hypothesis prediction may be a planar mode, a DC mode, a vertical mode, or a horizontal mode.

[0238] Alternatively, even when the surrounding position does not use multiple hypothesis prediction, if the mode used by the surrounding position is a specific mode, the encoder / decoder can set candIntraPredModeX to the mode used by the surrounding position. Or, even when the surrounding position does not use multiple hypothesis prediction, if the mode used by the surrounding position is a vertical mode or a height mode, the encoder / decoder can set candIntraPredModeX to the mode used by the surrounding position. Or, when the surrounding position is above the current block, even when the surrounding position does not use multiple hypothesis prediction, if the mode used by the surrounding position is a vertical mode, the encoder / decoder can set candIntraPredModeX to the mode used by the surrounding position. Also, when the surrounding position is on the left side of the current block, even when the surrounding position does not use multiple hypothesis prediction, if the mode used by the surrounding position is a horizontal mode, the encoder / decoder can set candIntraPredModeX to the mode used by the surrounding position.

[0239] Referring to FIG. 34, the mh_intra_flag may be a syntax element (or variable) indicating whether multiple hypothesis prediction is used (or has been used). Also, the intra prediction mode used in the neighboring blocks may be the horizontal mode. Also, the current block can use multiple hypothesis prediction, and a candidate list can be generated using candIntraPredMode based on the modes of the neighboring blocks. However, even if the neighborhood does not use multiple hypothesis prediction, since the intra prediction mode of the neighboring blocks is a specific mode, for example, the horizontal mode, the encoder / decoder can set candIntraPredMode to the horizontal mode.

[0240] Examples of the neighboring mode reference method described above will be described again below in combination with other embodiments of FIG. 30 described above. According to an embodiment of the present invention, the intra prediction mode candIntraPredModeX (X is A or B) of the neighboring blocks can be derived in the following manner.

[0241] 1. The availability derivation process for the block specified in the neighboring block availability confirmation process is called. The availability derivation process may set the position (xCurr, yCurr) to (xCb, yCb), and can take the neighboring positions (xNbY, yNbY) as (xNbX, yNbX) as input, and the output may be assigned to the availability variable availableX.

[0242] 2. The candidate intra prediction mode candIntraPredModeX may be derived in the following manner:

[0243] A. If one or more of the following conditions are true, candIntraPredModeX may be set to the INTRA_DC mode.

[0244] a) When the variable availableX is FALSE.

[0245] b) When mh_intra_flag[xNbX][yNbX] is not 1 and IntraPredModeY[xNbX][yNbX] is not INTRA_ANGULAR50 and INTRA_ANGULAR18.

[0246] c) When X is B and yCb-1 is smaller than ((yCb >> CtbLog2SizeY) << CtbLog2SizeY).

[0247] B. Otherwise, if IntraPredModeY[xNbX][yNbX] > INTRA_ANGULAR34, candIntraPredModeX may be set to INTRA_ANGULAR50.

[0248] C. Otherwise, if IntraPredModeY[xNbX][yNbX] <= INTRA_ANGULAR34 and IntraPredModeY[xNbX][yNbX] > INTRA_DC, candIntraPredModeX may be set to INTRA_ANGULAR18.

[0249] D. Otherwise, candIntraPredModeX may be set to IntraPredModeY[xNbX][yNbX].

[0250] According to another embodiment of the present invention, the intra prediction mode candIntraPredModeX (X is A or B) of the peripheral block can be derived in the following manner.

[0251] 1. The availability derivation process for the block specified in the adjacent block availability confirmation process is called. The availability derivation process may set the position (xCurr, yCurr) to (xCb, yCb), and can take the adjacent positions (xNbY, yNbY) as the input (xNbX, yNbX), and the output may be assigned to the availability variable availableX.

[0252] 2. The candidate intra prediction mode candIntraPredModeX can be derived in the following way:

[0253] A. If one or more of the following conditions are true, candIntraPredModeX may be set to the INTRA_DC mode.

[0254] a) If the variable availableX is FALSE.

[0255] b) If mh_intra_flag[xNbX][yNbX] is not 1 and IntraPredModeY[xNbX][yNbX] is not INTRA_PLANAR, INTRA_DC, INTRA_ANGULAR50, and INTRA_ANGULAR18.

[0256] c) If X is B and yCb-1 is less than ((yCb >> CtbLog2SizeY) << CtbLog2SizeY).

[0257] B. Otherwise, if IntraPredModeY[xNbX][yNbX] > INTRA_ANGULAR34, candIntraPredModeX may be set to INTRA_ANGULAR50.

[0258] C. Otherwise, if IntraPredModeY[xNbX][yNbX] <= INTRA_ANGULAR34 and IntraPredModeY[xNbX][yNbX] > INTRA_DC, candIntraPredModeX may be set to INTRA_ANGULAR18.

[0259] D. Otherwise, candIntraPredModeX may be set to IntraPredModeY[xNbX][yNbX].

[0260] In the above-described list setting method, candIntraPredModeX may be determined by the peripheral mode reference method.

[0261] FIG. 35 is a diagram showing a method of using a peripheral reference sample according to an embodiment of the present invention. As described above, when using multiple hypothesis prediction, intra prediction can be combined with other predictions and used. Therefore, when using multiple hypothesis prediction, intra prediction can be generated using samples around the current block as reference samples. According to an embodiment of the present invention, when using multiple hypothesis prediction, a mode using restored samples can be used. Also, when not using multiple hypothesis prediction, it is not necessary to use a mode using restored samples. The restored sample may be a restored sample around the current block.

[0262] As an example of the mode using the restored sample, a template matching method can be used. That is, a restored sample at a position already set based on a certain block can be defined as a template (or a template area). The template matching may be an operation of comparing the cost of the template of the block to be compared with the template of the current block and searching for a block with a smaller cost. At this time, the cost can be defined as the sum of the absolute values of the templates, the sum of the squares of the differences, etc. For example, an encoder / decoder can search for a block expected to be similar to the current block using template matching between the current block and the block of the reference picture, and based on this, set a motion vector or refine the motion vector. Also, examples of the mode using the restored sample include motion compensation using the restored sample, motion vector refinement, and the like.

[0263] In order to use the restored samples around the current block, when decoding the current block, it is necessary to wait for the decoding of the surrounding blocks to be completed. In such a case, it may be difficult to process the current block and the surrounding blocks in parallel. Therefore, when not using multiple hypothesis prediction, the encoder / decoder may not need to use a mode that uses the restored samples around the current block in order to enable parallel processing. Or, when using multiple hypothesis prediction, since intra prediction can be generated using the restored samples around the current block, the encoder / decoder can also use other modes that use the restored samples around the current block.

[0264] Also, according to an embodiment of the present invention, even when using multiple hypothesis prediction, whether or not to use the restored samples around the current block may vary depending on the candidate index. As an example, when the candidate index is smaller than a predefined specific threshold, it is possible to use the restored samples around the current block. When the candidate index is small, the number of candidate index signaling bits is small, the accuracy of the candidate is high, and the accuracy can be further improved by using the restored samples for candidates with high coding efficiency. As another example, when the candidate index is larger than the specific threshold, it is possible to use the restored samples around the current block. When the candidate index is large, the number of candidate index signaling bits is large, and the accuracy of the candidate may be low. The accuracy can be complemented by using the restored samples around the current block for candidates with low accuracy.

[0265] According to an embodiment of the present invention, when using multiple hypothesis prediction, the encoder / decoder can generate an inter prediction using the restored samples around the current block and combine the inter prediction with the intra prediction of the multiple hypothesis prediction to generate a predicted block. Referring to FIG. 35, the value of the mh_intra_flag, which is a signaling indicating whether the current block uses multiple hypothesis prediction, is 1. Since the current block uses multiple hypothesis prediction, a mode that uses the restored samples around the current block can be used.

[0266] FIG. 36 is a diagram showing a conversion mode according to an embodiment of the present invention. According to an embodiment of the present invention, there may be a conversion mode in which conversion is performed only on a sub-part of a block. In this specification, a conversion mode in which conversion is applied only to a sub-part in this way can be called a sub-block transform (SBT) or a spatially varying transform (SVT), etc. For example, a CU or a PU may be divided into a plurality of TUs, and only a part of the plurality of TUs may be converted. Or, for example, only any one of the plurality of TUs may be converted. Among the plurality of TUs, the TUs that are not converted can be set to have a residual of 0.

[0267] Referring to FIG. 36, there may be two types, SBT-V (SBT-vertical) and SBT-H (SBT-horizontal), as types that divide one CU or PU into a plurality of TUs. SBT-V may be a type in which the height of a plurality of TUs is equal to the CU or PU height, and the width of the plurality of TUs is different from the CU or PU width. SBT-H may be a type in which the height of a plurality of TUs is different from the CU or PU height, and the width of the plurality of TUs is equal to the CU or PU width. As an example, the width and position of the TU to be converted in SBT-V may be signaled. Also, the height and position of the TU to be converted in SBT-H may be signaled. According to an example, a conversion kernel may be set in advance according to the SBT type and position, width, or height.

[0268] Thus, there is a mode of converting only a part of the CU or PU because it is possible that the residual mainly exists in a part of the CU or PU after prediction. That is, SBT has the same concept as the skip mode for a TU. At this time, the existing skip mode may be a skip mode for a CU.

[0269] Referring to FIG. 36, for each type of SBT-V ((a) and (b) in FIG. 36) and SBT-H ((c) and (d) in FIG. 36), the position to be converted as indicated by A may be defined. And the width or height may be defined as 1 / 2 or 1 / 4 of the CU width or CU height. Also, the portions other than the region indicated by A can be regarded as having a residual value of 0. Further, the conditions under which SBT can be used may be defined. For example, the conditions for SBT to be possible may be signaled as to whether it can be used in the syntax of conditions related to the block size and the high level (e.g., sequence, slice, tile, etc.).

[0270] According to an embodiment of the present invention, there can be a relationship between multiple hypothesis prediction and the transformation mode. For example, the presence or absence of the use of one of them may determine the presence or absence of the use of the other. Or, the presence or absence of the use of one mode may determine the presence or absence of the use of the other mode. Or, the presence or absence of the use of one of them may determine the presence or absence of the use of the other mode. As an example, the transformation mode may be the SBT described with reference to FIG. 36. That is, the presence or absence of the use of SBT may be determined by the presence or absence of the use of multiple hypothesis prediction. Or, the presence or absence of the use of multiple hypothesis prediction may be determined by the presence or absence of the use of SBT. The following Table 6 shows a syntax structure exemplifying the relationship between multiple hypothesis prediction and the transformation mode according to an embodiment of the present invention.

[0271]

Table 6

[0272] Referring to Table 6, the presence or absence of SBT usage may be determined by the presence or absence of using multiple hypothesis prediction. That is, when multiple hypothesis prediction is not applied to the current block (!mh_intra_flag is true), the decoder can parse the cu_sbt_flag, which is a syntax element indicating the presence or absence of SBT application. When multiple hypothesis prediction is not applied to the current block, it may not be necessary to parse the cu_sbt_flag. In this case, the value of cu_sbt_flag may be inferred to be 0 according to predefined conditions.

[0273] As described above, for SBT, when the residual exists only in a part of the block after prediction for the processing block, the compression performance improvement continues. On the other hand, in the case of multiple hypothesis prediction, while effectively reflecting the movement of the object using inter prediction, intra prediction can be efficiently used to perform prediction for the remaining area, so an improvement in the prediction performance for the entire block is expected. That is, when multiple hypothesis prediction is applied, the prediction performance for the entire block is improved, and the phenomenon that the residual is concentrated only in a part of the block occurs relatively less. Therefore, according to an embodiment of the present invention, the encoder / decoder may not apply SBT when multiple hypothesis prediction is applied. Or, the encoder / decoder may not apply multiple hypothesis prediction when SBT is applied.

[0274] According to an embodiment of the present invention, the position of the TU to be transformed by SBT may be restricted by the presence or absence of using multiple hypothesis prediction or the mode of multiple hypothesis prediction. Or, the width (SBT-V) or height (SBT-H) of the TU to be transformed by SBT may be restricted by the presence or absence of using multiple hypothesis prediction or the mode of multiple hypothesis prediction. Thereby, signaling regarding the position, width, or height can be reduced. For example, the position of the TU to be transformed by SBT does not need to be in an area where the intra prediction weight value is high in multiple hypothesis prediction. This is because the residual in the area with a large weight value can be reduced by multiple hypothesis prediction.

[0275] Therefore, when using multiple hypothesis prediction, it may not be necessary to have a mode that converts the side with a larger weight value in SBT. For example, when using the horizontal mode or the vertical mode in multiple hypothesis prediction, the SBT types in (b) and (d) of FIG. 36 may be omitted (i.e., not considered). As yet another example, when using the planar mode in multiple hypothesis prediction, the position of the TU to be converted in SBT may be restricted. For example, when using the planar mode in multiple hypothesis prediction, the SBT types in (a) and (c) of FIG. 36 may be omitted. This is because when using the planar mode in multiple hypothesis prediction, the region adjacent to the reference sample of the intra prediction can have a value similar to the reference sample value, and thus the region adjacent to the reference sample can have a relatively small residual.

[0276] In other embodiments, when using multiple hypothesis prediction, the values of the width or height of the TU to be converted in SBT may be changed. Alternatively, when using a specific mode in multiple hypothesis prediction, the values of the width or height of the TU to be converted in SBT may be changed. For example, when using multiple hypothesis prediction, since there will be no large remaining residuals in the wider part of the block, the large values of the width or height of the TU to be converted in SBT can be excluded. Alternatively, when using multiple hypothesis prediction, the values of the width or height of the TU to be converted in SBT, such as the unit for which the weight value changes in multiple hypothesis prediction, can be excluded.

[0277] Referring to Table 6, there may be a cu_sbt_flag indicating whether SBT is used and a mh_intra_flag indicating whether multiple hypothesis prediction is used. When the mh_intra_flag is 0, the cu_sbt_flag can be parsed. Also, when the cu_sbt_flag does not exist, it may be inferred as 0. Both combining intra prediction with multiple hypothesis prediction and SBT can be used to solve the problem that only a large amount of residual can remain in a part of the CU or PU when the technology is not used. Therefore, since there can be a correlation between the two technologies, it is possible to determine based on the presence or absence of the use of one technology or the presence or absence of the use of a specific mode of one technology with respect to the other technology.

[0278] Also, in Table 6, sbtBlockConditions can indicate the conditions under which SBT is possible. The conditions under which SBT is possible can include conditions related to the block size, signaling values regarding usability at a high level (e.g., sequence, slice, tile, etc.).

[0279] FIG. 37 is a diagram showing the relationship between color difference components according to an embodiment of the present invention. Referring to FIG. 37, a color format may be expressed as chroma_format_idc, Chroma format, separate_colour_plane_flag, etc. If it is Monochrome, only one sample array may exist. Also, both SubWidthC and SubHeightC may be 1. If it is 4:2:0 sampling, two chroma arrays (or color difference components, color difference blocks) may exist. Also, a chroma array can have half the width and half the height of a luma array (or luma component, luma block). Both SubWidthC and SubHeightC may be 2. SubWidthC and SubHeightC can indicate the size of the chroma array compared to the luma array. When the width or height of the chroma array is half the size of the luma array, SubWidthC or SubHeightC may be 2. When the width or height of the chroma array is the same size as the luma array, SubWidthC or SubHeightC may be 1.

[0280] If it is 4:2:2 sampling, there may be two chrominance arrays. Also, the chrominance arrays may have half the width and the same height as the luminance array. SubWidthC and SubHeightC may be 2 and 1 respectively. If it is 4:4:4 sampling, the chrominance arrays may have the same width and the same height as the luminance array. SubWidthC and SubHeightC may both be 1. At this time, it may be processed individually based on the separate_colour_plane_flag. If the separate_colour_plane_flag is 0, the chrominance arrays may have the same width and height as the luminance array. If the separate_colour_plane_flag is 1, the three color planes (luminance, Cb, Cr) may be processed respectively. Regardless of the separate_colour_plane_flag, if it is 4:4:4, SubWidthC and SubHeightC may both be 1.

[0281] If the separate_colour_plane_flag is 1, there may be only one corresponding to one color component in one slice. If the separate_colour_plane_flag is 0, there may be those corresponding to multiple color components in one slice. Referring to FIG. 38, it is possible for SubWidthC and SubHeightC to be different only when it is 4:2:2. Therefore, when it is 4:2:2, the relationship between the luminance reference width and height may be different from the relationship between the chrominance reference width and height.

[0282] For example, if the luminance sample reference width is widthL and the chrominance sample reference width is widthC, and if widthL and widthC correspond, the relationship is as shown in the following mathematical formula 6.

[0283]

Equation

[0284] Also, similarly, when the luminance sample reference height is heightL and the color difference sample reference height is heightC, and assuming that heightL and heightC correspond, the relationship is as shown in the following mathematical formula 7.

[0285]

Equation

[0286] Also, there may be a value indicating (showing) a color component. For example, cIdx can indicate a color component. For example, cIdx may be a color component index. If cIdx is 0, it can indicate a luminance component. Also, if cIdx is not 0, it can indicate a color difference component. Also, if cIdx is 1, it can indicate a color difference Cb component. Also, if cIdx is 2, it can indicate a color difference Cr component.

[0287] FIG. 38 is a diagram showing the relationship between color components according to an embodiment of the present invention. FIGS. 38(a), (b), and (c) assume the cases of 4:2:0, 4:2:2, and 4:4:4, respectively. Referring to FIG. 38(a), one color difference sample (one Cb and one Cr) may be positioned per two luminance samples in the horizontal direction. Also, one color difference sample (one Cb and one Cr) may be positioned per two luminance samples in the vertical direction. Referring to FIG. 38(b), one color difference sample (one Cb and one Cr) may be positioned per two luminance samples in the horizontal direction. Also, one color difference sample (one Cb and one Cr) may be positioned per one luminance sample in the vertical direction. Referring to FIG. 38(c), one color difference sample (one Cb and one Cr) may be positioned per one luminance sample in the horizontal direction. Also, one color difference sample (one Cb and one Cr) may be positioned per one luminance sample in the vertical direction.

[0288] As described above, SubWidthC and SubHeightC described with reference to FIG. 37 may be determined by such a relationship. Then, based on SubWidthC and SubHeightC, conversion between the luminance sample standard and the color difference sample standard can be performed.

[0289] FIG. 39 is a diagram showing a peripheral reference position according to an embodiment of the present invention. According to an embodiment of the present invention, the encoder / decoder can refer to the peripheral position during prediction. For example, as described above, when performing CIIP (combined inter-picture merge and intra-picture prediction), the peripheral position can be referred to. CIIP may be the multiple hypothesis prediction described above. That is, CIIP may be a prediction method that combines inter prediction (for example, merge mode inter prediction) and intra prediction. According to an embodiment of the present invention, the encoder / decoder can combine inter prediction and intra prediction by referring to the peripheral position. For example, the encoder / decoder can determine the ratio of inter prediction to intra prediction by referring to the peripheral position. Or, when combining inter prediction and intra prediction by referring to the peripheral position, the encoder / decoder can determine a weighting value. Or, when performing a weighted sum (or weighted average) of inter prediction and intra prediction by referring to the peripheral position, the encoder / decoder can determine a weighting value.

[0290] According to an embodiment of the present invention, the surrounding positions to be referred to may include NbA and NbB. The coordinates of NbA and NbB may be (xNbA, yNbA) and (xNbB, yNbB), respectively. Also, NbA may be the left position of the current block. Specifically, when the top-left coordinates of the current block are (xCb, yCb), and the width and height of the current block are cbWidth and cbHeight, respectively, NbA may be (xCb - 1, yCb + cbHeight - 1). The top-left coordinates (xCb, yCb) of the current block may be values based on the luminance samples. Or, the top-left coordinates (xCb, yCb) of the current block may be the position of the top-left luminance sample of the current luminance coding block with respect to the top-left luminance sample of the current picture. Also, the cbWidth and cbHeight may respectively indicate the width and height based on the color component. The coordinates described above may be for the luminance component (or luminance block). For example, cbWidth and cbHeight may indicate the width and height based on the luminance component.

[0291] Also, NbB may be the upper position of the current block. More specifically, when the top-left coordinates of the current block are (xCb, yCb), and the width and height of the current block are cbWidth and cbHeight, respectively, NbB may be (xCb + cbWidth - 1, yCb - 1). The top-left coordinates (xCb, yCb) of the current block may be values based on the luminance samples. Or, the top-left coordinates (xCb, yCb) of the current block may be the position of the top-left luminance sample of the current luminance coding block with respect to the top-left luminance sample of the current picture. Also, the cbWidth and cbHeight may be values based on the color component. The coordinates described above may be for the luminance component (luminance block). For example, cbWidth and cbHeight may be values based on the luminance component.

[0292] Referring to FIG. 39, the upper left corner, the coordinates of NbA, the coordinates of NbB, etc. for the luminance block are illustrated. Also, NbA may be the left side position of the current block. More specifically, when the upper left corner coordinates of the current block are (xCb, yCb) and the width and height of the current block are cbWidth and cbHeight respectively, NbA may be (xCb - 1, yCb + 2 * cbHeight - 1). The upper left corner coordinates (xCb, yCb) of the current block may be values based on the luminance samples. Or, the upper left corner coordinates (xCb, yCb) of the current block may be the upper left corner luminance sample position of the current luminance coding block with respect to the upper left corner luminance sample of the current picture. Also, the cbWidth and cbHeight may be values based on the color component. The coordinates described above may be for the color difference components (color difference blocks). For example, cbWidth and cbHeight may be values based on the color difference components. Also, this coordinate may apply to the case of the 4:2:0 format.

[0293] Also, NbB may be at the upper position of the current block. More specifically, when the upper left coordinate of the current block is (xCb, yCb), and the width and height of the current block are cbWidth and cbHeight respectively, NbB may be (xCb + 2 * cbWidth - 1, yCb - 1). The upper left coordinate (xCb, yCb) of the current block may be a value based on the luminance sample. Or, the upper left coordinate (xCb, yCb) of the current block may be the upper left luminance sample position of the current luminance coding block relative to the upper left luminance sample of the current picture. Also, the cbWidth and cbHeight may be values based on the color component. The coordinates described above may be for the chroma component (chroma block). For example, cbWidth and cbHeight may be values based on the chroma component. Also, this coordinate may apply to the case of 4:2:0 format or 4:2:2 format. Referring to FIG. 39, the upper left, the coordinates of NbA, the coordinates of NbB, etc. for the chroma block are illustrated.

[0294] FIG. 40 is a diagram showing a weighted sample prediction process according to an embodiment of the present invention. In the embodiment of FIG. 40, a method of combining two or more prediction signals is described. Also, the embodiment of FIG. 40 can be applied when using CIIP. Also, the embodiment of FIG. 40 can include the peripheral position reference method described in FIG. 39. Referring to FIG. 44, the variable scallFact indicating the scaling factor can be described by the following mathematical formula 8.

[0295]

Equation

[0296] In Mathematical Formula 8, the encoder / decoder can set scallFact to 0 when cIdx is 0, and set scallFact to 1 when cIdx is not 0. In an embodiment of the present invention, x?y:z may indicate a y value when x is true or x is not 0, and indicate a z value otherwise (when x is false (or when x is 0)).

[0297] Also, the encoder / decoder can set the coordinates (xNbA, yNbA) and (xNbB, yNbB) of the peripheral positions NbA and NbB referred to in the multiple hypothesis prediction. According to the embodiment described in FIG. 39, for the luminance component, (xNbA, yNbA) and (xNbB, yNbB) are (xCb - 1, yCb + cbHeight - 1) and (xCb + cbWidth - 1, yCb - 1) respectively, and for the chrominance component, (xNbA, yNbA) and (xNbB, yNbB) may be set to (xCb - 1, yCb + 2*cbHeight - 1) and (xCb + 2*cbWidth - 1, yCb - 1) respectively. Also, the operation of multiplying by 2^n may be the same as the operation of left-shifting n bits. For example, the operation of multiplying by 2 may be calculated as the value of left-shifting 1 bit. Also, left-shifting x by n bits can be expressed as "x << n". Also, the operation of dividing by 2^n may be the same as the operation of right-shifting n bits. Also, the operation of dividing by 2^n and discarding the decimal part may be calculated as the value of right-shifting n bits. For example, the operation of dividing by 2 may be calculated as the value of right-shifting 1 bit. Also, right-shifting x by n bits can be expressed as "x >> n". Therefore, (xCb - 1, yCb + 2*cbHeight - 1) and (xCb + 2*cbWidth - 1, yCb - 1) can be expressed as (xCb - 1, yCb + (cbHeight << 1) - 1) and (xCb + (cbWidth << 1) - 1, yCb - 1). Therefore, when representing both the coordinates for the luminance component and the coordinates for the chrominance component described above, it is as shown in the following Mathematical Formula 9.

[0298] [Number]

[0299] In Mathematical formula 9, scallFact may be determined as (cIdx == 0)? 0: 1 as described above. At this time, cbWidth and cbHeight represent the width and height respectively based on each color component. For example, the width and height based on the luminance component are cbWidthL and cbHeightL respectively. When performing the weighted sample prediction process for the luminance component, cbWidth and cbHeight may be cbWidthL and cbHeightL respectively. Also, the width and height based on the luminance component are cbWidthL and cbHeightL respectively. When performing the weighted sample prediction process for the color difference component, cbWidth and cbHeight may be cbWidthL / SubWidthC and cbHeightL / SubHeightC respectively.

[0300] Also, according to an embodiment of the present invention, the encoder / decoder can determine the prediction mode of the position with reference to the surrounding positions. For example, the encoder / decoder can determine whether the prediction mode is intra prediction. Also, the prediction mode may be indicated by CuPredMode. When CuPredMode is MODE_INTRA, it may be a mode that uses intra prediction. Also, the CuPredMode value may be MODE_INTRA, MODE_INTER, MODE_IBC, or MODE_PLT. When CuPredMode is MODE_INTER, inter prediction can be used. Also, when CuPredMode is MODE_IBC, intra block copy (IBC) can be used. Also, when CuPredMode is MODE_PLT, the palette mode can be used. Also, CuPredMode may be indicated by the channel type (chType) and the position. For example, it may be indicated as CuPredMode[chType][x][y], and this value may be the CuPredMode value for the channel type chType at the (x, y) position.

[0301] Also, according to an embodiment of the present invention, chType may be based on the tree type. For example, the tree type may be set to values such as SINGLE_TREE, DUAL_TREE_LUMA, DUAL_TREE_CHROMA. When it is SINGLE_TREE, there may be a part where the block partitioning of the luminance component and the chrominance component is shared. For example, when it is SINGLE_TREE, the block partitioning of the luminance component and the chrominance component may be the same. Or, when it is SINGLE_TREE, the block partitioning of the luminance component and the chrominance component may be the same or partially the same. Or, when it is SINGLE_TREE, the block partitioning of the luminance component and the chrominance component may be performed by the same syntax element value.

[0302] Also, according to an embodiment of the present invention, in the case of DUAL TREE, the block partitioning of the luminance component and the chrominance component may be independent. Or, in the case of DUAL TREE, the block partitioning of the luminance component and the chrominance component may be performed with different syntax element values. Also, in the case of DUAL TREE, the tree type value may be DUAL_TREE_LUMA or DUAL_TREE_CHROMA. If the tree type is DUAL_TREE_LUMA, it can be shown that DUAL TREE is used and it is a process for the luminance component. If the tree type is DUAL_TREE_CHROMA, it can be shown that DUAL TREE is used and it is a process for the chrominance component. Also, chType may be determined based on whether the tree type is DUAL_TREE_CHROMA. For example, chType may be set to 1 when the tree type is DUAL_TREE_CHROMA, and set to 0 when the tree type is not DUAL_TREE_CHROMA. Therefore, referring to FIG. 40, the CuPredMode[0][xNbX][yNbY] value can be determined. X can be replaced with A and B. That is, the CuPredMode value for the NbA and NbB positions can be determined.

[0303] Also, according to an embodiment of the present invention, the isIntraCodedNeighbourX value can be set based on the determination of the prediction mode for the peripheral position. For example, the isIntraCodedNeighbourX value can be set according to whether the CuPredMode for the peripheral position is MODE_INTRA. If the CuPredMode for the peripheral position is MODE_INTRA, the isIntraCodedNeighbourX value can be set to TRUE, and if the CuPredMode for the peripheral position is not MODE_INTRA, the isIntraCodedNeighbourX value can be set to FALSE. As described above and in the present invention described below, X may be replaced with A or B, etc. Also, what is written as X can indicate that it corresponds to the X position.

[0304] Also, according to an embodiment of the present invention, with reference to the peripheral position, it can be determined whether the position is available. Whether the position is available can be set by availableX. Also, isIntraCodedNeighbourX can be set based on availableX. For example, when availableX is TRUE, isIntraCodedNeighbourX can be set to TRUE, and for example, when availableX is FALSE, isIntraCodedNeighbourX can be set to FALSE. Referring to FIG. 40, whether the position is available can be determined by calling "the derivation process for neighbouring block availability". Also, whether the position is available can be determined based on whether the position is inside the current picture. If the position is (xNbY, yNbY) and xNbY or yNbY is less than 0, it is outside the current picture and availableX may be set to FALSE. Also, if xNbY is greater than or equal to the picture width, it is outside the current picture and availableX may be set to FALSE. The picture width can be indicated by pic_width_in_luma_samples. Also, if yNbY is greater than or equal to the picture height, it is outside the current picture and availableX may be set to FALSE. The picture height can be indicated by pic_height_in_luma_samples. Also, if the position is in a block or slice different from the current block, availableX may be set to FALSE. Also, if the reconstruction of the position is not completed, availableX may be set to FALSE. Whether the reconstruction is completed can be indicated by IsAvailable[cIdx][xNbY][yNbY]. In short, when any one of the following conditions is satisfied, availableX can be set to FALSE, and otherwise (when all of the following conditions are not satisfied), availableX can be set to TRUE.

[0305] - Condition 1: xNbY < 0

[0306] - Condition 2: yNbY < 0

[0307] - Condition 3: xNbY >= pic_width_in_luma_samples

[0308] - Condition 4: yNbY >= pic_height_in_luma_samples

[0309] - Condition 5: IsAvailable[cIdx][xNbY][yNbY] == FALSE

[0310] - Condition 6: When the position (the peripheral position (xNbY, yNbY) position) belongs to a block (or a different slice) different from the current block

[0311] Also, according to an embodiment of the present invention, it is possible to determine whether the current position and the position are in the same CuPredMode optionally, and set availableX. The two described conditions can be combined to set isIntraCodedNeighbourX. For example, when all of the following conditions are satisfied, isIntraCodedNeighbourX can be set to TRUE, and otherwise (when at least one of the following conditions is not satisfied), isIntraCodedNeighbourX can be set to FALSE.

[0312] - Condition 1: availableX == TRUE

[0313] - Condition 2: CuPredMode[0][xNbX][yNbX] == MODE_INTRA

[0314] Also, according to an embodiment of the present invention, the weighted value of CIIP can be determined based on a plurality of isIntraCodedNeighbourX. For example, when combining inter prediction and intra prediction based on a plurality of isIntraCodedNeighbourX, the weighted value can be determined. For example, it can be determined based on isIntraCodedNeighbourA and isIntraCodedNeighbourB. According to one embodiment, when both isIntraCodedNeighbourA and isIntraCodedNeighbourB are TRUE, w can be set to 3. For example, w may be the weighted value of CIIP or a value for determining the weighted value. Also, when both isIntraCodedNeighbourA and isIntraCodedNeighbourB are FALSE, w can be set to 1. Also, when one of isIntraCodedNeighbourA and isIntraCodedNeighbourB is FALSE (the same as when one of the two is TRUE), w can be set to 2. That is, w can be set based on whether the surrounding position is predicted by intra prediction or based on how much the surrounding position is predicted by intra prediction.

[0315] Also, according to an embodiment of the present invention, w may be a weighted value corresponding to intra prediction. Also, the weighted value corresponding to inter prediction may be determined based on w. For example, the weighted value corresponding to inter prediction may be (4 - w). When combining two or more prediction signals, the encoder / decoder can use the following mathematical formula 10.

[0316]

Equation

[0317] In Mathematical Expression 10, predSamplesIntra and predSamplesInter may be prediction signals. For example, predSamplesIntra and predSamplesInter may be prediction signals predicted by intra prediction and prediction signals predicted by inter prediction (for example, merge mode, more specifically, regular merge mode), respectively. Also, predSampleComb may be a prediction signal used in CIIP.

[0318] Also, according to an embodiment of the present invention, prior to applying Mathematical Expression 10, a process of updating the prediction signal before combination may be included. For example, the following Mathematical Expression 11 may be applied to the updating process. For example, the process of updating the prediction signal may be a process of updating the inter prediction signal of CIIP.

[0319] [Number]

[0320] FIG. 41 is a diagram showing a peripheral reference position according to an embodiment of the present invention.

[0321] Although the peripheral reference positions were described with reference to FIGS. 39 and 40, problems may occur if the described positions are used in all cases (e.g., all color difference blocks), and this problem will be described with reference to FIG. 41. The embodiment of FIG. 41 shows a color difference block. In FIGS. 39 and 40, for the color difference block, the NbA and NbB coordinates based on the luminance samples were (xCb - 1, yCb + 2*cbHeight - 1) and (xCb + 2*cbWidth - 1, yCb - 1), respectively. However, when SubWidthC or SubHeightC is 1, it may indicate a position shown in FIG. 41 different from the position shown in FIG. 39. Multiplying cbWidth and cbHeight by 2 in the above coordinates, cbWidth and cbHeight are shown based on each color component (in this embodiment, the color difference component), and since the coordinates are shown based on luminance, in the case of 4:2:0, it may be for compensating the number of luminance samples and color difference samples. That is, it may be for representing the coordinates based on luminance in the case where there is 1 color difference sample corresponding to 2 luminance samples on the x-axis of the luminance sample and 1 color difference sample corresponding to 2 luminance samples on the y-axis of the luminance sample. Therefore, when SubWidthC or SubHeightC is 1, different positions can be indicated. Therefore, when always using the positions (xCb - 1, yCb + 2*cbHeight - 1) and (xCb + 2*cbWidth - 1, yCb - 1) for the color difference block in this way, it may occur that a position far from the current color difference block is referred to. Also, in this case, the relative position used in the luminance block of the current block may not match the relative position used in the color difference block. Also, since different positions are referred to for the color difference block, it may happen that a weighting value is set by referring to a position with little relevance to the current block, or that decoding / restore is not performed in the block decoding order.

[0322] Referring to FIG. 41, in the case of 4:4:4, that is, when both SubWidthC and SubHeightC are 1, the positions of the above-described luminance reference coordinates are shown. NbA and NbB exist at positions away from the color difference block shown by the solid line.

[0323] FIG. 42 is a diagram showing a weighted sample prediction process according to an embodiment of the present invention. The embodiment of FIG. 42 may be an embodiment for solving the problems described with reference to FIGS. 39 to 41. Also, the description of the content overlapping with the above-described content will be omitted. In FIG. 40, the peripheral position is set based on scallFact. As described in FIG. 41, scallFact is a value for converting the position when SubWidthC and SubHeightC are 2. However, as described above, problems may occur depending on the color format, and the ratio of the color difference sample to the luminance sample may be different between the horizontal and vertical directions. According to an embodiment of the present invention, scallFact can be separated into horizontal (width) and vertical (height).

[0324] According to an embodiment of the present invention, scallFactWidth and scallFactHeight may exist, and the peripheral position can be set based on scallFactWidth and scallFactHeight. Also, the peripheral position may be set based on a luminance sample (luminance block). Also, scallFactWidth can be set based on cIdx and SubWidthC. For example, when cIdx is 0 or SubWidthC is 1, scallFactWidth can be set to 0, and otherwise (when cIdx is not 0 and SubWidthC is not 1 (when SubWidthC is 2)), scallFactWidth can be set to 1. At this time, scallFactWidth can be determined using the following mathematical formula 12.

[0325]

Equation

[0326] Also, scallFactHeight can be set based on cIdx and SubHeightC. For example, when cIdx is 0 or SubHeightC is 1, scallFactHeight can be set to 0, and if not (when cIdx is not 0 and SubHeightC is not 1 (when SubHeightC is 2)), scallFactHeight can be set to 1. At this time, scallFactHeight can be determined using the following mathematical formula 13.

[0327]

Number

[0328] Also, the x coordinate of the peripheral position can be indicated based on scallFactWidth, and the y coordinate of the peripheral position can be indicated based on scallFactHeight. For example, the coordinates of NbB can be set based on scallFactWidth. For example, the coordinates of NbA can be set based on scallFactHeight. Also, as described above, here, being based on scallFactWidth may mean being based on SubWidthC, and being based on scallFactHeight may mean being based on SubHeightC. For example, the peripheral position coordinates are as follows in mathematical formula 14.

[0329]

Number

[0330] At this time, xCb and yCb can be the coordinates shown based on the luminance sample standard as described above. Also, cbWidth and cbHeight can be those shown based on each color component.

[0331] Therefore, in the case of a color difference block where SubWidthC is 1, (xNbB, yNbB) may be (xCb + cbWidth - 1, yCb - 1). That is, in this case, the NbB coordinates for the luminance block and the NbB coordinates for the color difference block may be the same. Also, in the case of a color difference block where SubHeightC is 1, (xNbA, yNbA) may be (xCb - 1, yCb + cbHeight - 1). That is, in this case, the NbA coordinates for the luminance block and the NbA coordinates for the color difference block may be the same.

[0332] Therefore, in the case of the embodiment of FIG. 42, when it is in the 4:2:0 format, the same peripheral coordinates as those of the embodiments of FIGS. 39 and 40 can be set, and when it is in the 4:2:2 format or the 4:4:4 format, different peripheral coordinates from those of the embodiments of FIGS. 39 and 40 can be set.

[0333] The processes other than FIG. 42 may be the same as those described in FIG. 40. That is, based on the peripheral position coordinates described in FIG. 42, it is possible to determine the prediction mode or the availability, and to determine the weighting value of CIIP. In the present invention, the peripheral position and the coordinates of the peripheral position may be used in the same meaning.

[0334] FIG. 43 is a diagram showing a weighted sample prediction process according to an embodiment of the present invention. The embodiment of FIG. 43 represents the peripheral position coordinates described in FIG. 42 in another manner. Therefore, the content overlapping with the above description may be omitted. As described above, bit shift can be expressed by multiplication. FIG. 42 may be shown using bit shift, and FIG. 43 may be shown using multiplication.

[0335] According to one embodiment, scallFactWidth can be set based on cIdx and SubWidthC. For example, when cIdx is 0 or SubWidthC is 1, scallFactWidth can be set to 1; otherwise (when cIdx is not 0 and SubWidthC is not 1 (when SubWidthC is 2)), scallFactWidth can be set to 2. At this time, scallFactWidth can be determined using the following mathematical formula 15.

[0336]

Number

[0337] Also, scallFactHeight can be set based on cIdx and SubHeightC. For example, when cIdx is 0 or SubHeightC is 1, scallFactHeight can be set to 1; otherwise (when cIdx is not 0 and SubHeightC is not 1 (when SubHeightC is 2)), scallFactHeight can be set to 2. At this time, scallFactHeight can be determined using the following mathematical formula 16.

[0338]

Number

[0339] Also, the x coordinate of the peripheral position can be indicated based on scallFactWidth, and the y coordinate of the peripheral position can be indicated based on scallFactHeight. For example, the coordinates of NbB can be set based on scallFactWidth. For example, the coordinates of NbA can be set based on scallFactHeight. Also, as described above, here, being based on scallFactWidth may be being based on SubWidthC, and being based on scallFactHeight may be being based on SubHeightC. For example, the peripheral position coordinates are as shown in the following mathematical formula 17.

[0340]

Number

[0341] At this time, xCb and yCb may be the coordinates indicated on the luminance sample basis as described above. Also, cbWidth and cbHeight may be those indicated based on each color component.

[0342] FIG. 44 is a diagram showing a weighted sample prediction process according to an embodiment of the present invention. In the embodiments such as FIGS. 40, 42, and 43, it is determined whether the position is available with reference to the peripheral position. At this time, cIdx, which is an index indicating a color component, is set to 0 (luminance component). Also, when determining whether the position is available, cIdx can be used to determine whether the reconstruction of the position cIdx has been completed. That is, when determining whether the position is available, cIdx can be used to determine the IsAvailable[cIdx][xNbY][yNbY] value. However, when performing a weighted sample prediction process for a color difference block, referring to the IsAvailable value corresponding to cIdx0 may result in an incorrect determination. For example, if the reconstruction of the luminance component of the block including the peripheral position has not been completed but the reconstruction of the color difference component has been completed, if cIdx is not 0, IsAvailable[0][xNbY][yNbY] may be FALSE and IsAvailable[cIdx][xNbY][yNbY] may be TRUE. Therefore, a situation may occur where it is determined that the peripheral position is not available even though it is actually available. In the embodiment of FIG. 44, in order to solve such a problem, when determining whether the position is available with reference to the peripheral position, the cIdx of the current coding block can be used as an input. That is, when calling “the derivation process for neighbouring block availability”, the input cIdx can be set to the cIdx of the current coding block. In this embodiment, descriptions overlapping with those described in FIGS. 42 and 43 are omitted.

[0343] Also, when determining the prediction mode of the peripheral position described above, CuPredMode[0][xNbX][yNbY], which is the CuPredMode corresponding to chType0, was referred to. However, if the chType for the current block does not match, incorrect parameters may be referred to. Therefore, according to an embodiment of the present invention, when determining the prediction mode of the peripheral position, CuPredMode[chType][xNbX][yNbY] corresponding to the chType value corresponding to the current block can be referred to.

[0344] FIG. 45 is a diagram illustrating a video signal processing method based on multiple hypothesis prediction according to an embodiment to which the present invention is applied. Referring to FIG. 45, for convenience of explanation, the decoder will be mainly described, but the present invention is not limited thereto, and the multiple hypothesis prediction-based video signal processing method according to this embodiment can be applied to the encoder in substantially the same manner.

[0345] Specifically, when the merge mode is applied to the current block, the decoder can obtain a first syntax element indicating whether combined prediction is applied to the current block (S4501). Here, the combined prediction represents a prediction mode in which inter prediction and intra prediction are combined. As described above, the present invention is not limited to such a name, and in this specification, the multiple hypothesis prediction can be referred to as multiple prediction, plurality of predictions, combined prediction, inter-intra weighted prediction, combined inter-intra prediction, combined inter-intra weighted prediction, and the like. In one embodiment, prior to the step S4501, the decoder receives information for predicting the current block, and based on the information for prediction, can determine whether the merge mode is applied to the current block.

[0346] When the decoder indicates that the first syntax element indicates that the combined prediction is applicable to the current block, the decoder can generate an inter-prediction block and an intra-prediction block of the current block (S4502). Then, the decoder can generate a combined prediction block by weighted-summing the inter-prediction block and the intra-prediction block (S4503). Then, the decoder can decode the residual block of the current block and restore the current block using the combined prediction block and the residual block.

[0347] As described above, in one embodiment, the step of decoding the residual block may include obtaining a second syntax element indicating whether a sub-block transform is applicable to the current block when the first syntax element indicates that the combined prediction is not applicable to the current block. That is, the sub-block transform is applicable only when the combined prediction is not applicable to the current block, and syntax signaling regarding whether to apply it may be performed. Here, the sub-block transform represents a transform mode in which a transform is applied to any one of the sub-blocks of the current block divided in the horizontal or vertical direction.

[0348] As described above, in one embodiment, when the second syntax element does not exist, the value of the second syntax element may be inferred to be 0.

[0349] As described above, in one embodiment, when the first syntax element indicates that the combined prediction is applicable to the current block, the intra-prediction mode for intra-prediction of the current block may be set to the planar mode.

[0350] As described above, in one embodiment, the decoder can set the positions of the left and upper neighboring blocks referred to for the combined prediction, and can perform the combined prediction based on the intra prediction mode of the set positions. In one embodiment, the decoder can determine the weighting values used for the combined prediction based on the intra prediction mode of the set positions. Also, as described above, the positions of the left and upper neighboring blocks may be determined using a scaling factor variable determined by the color component index value of the current block.

[0351] The embodiments of the present invention described above are implemented through various means. For example, the embodiments of the present invention are implemented by hardware, firmware, software, or a combination thereof.

[0352] In the case of implementation by hardware, the method according to the embodiments of the present invention is implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSDPs (Digital Signal Processing Devices), PDLs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), processors, controllers, microcontrollers, microprocessors, etc.

[0353] In the case of implementation by firmware or software, the method according to the embodiments of the present invention is implemented in the form of modules, procedures, functions, etc. that perform the above-described functions or operations. The software code is stored in a memory and implemented by a processor. The memory is located inside or outside the processor and exchanges data with the processor by various means already known.

[0354] Some embodiments are also embodied in the form of a recording medium containing computer-executable instructions such as program modules executed by a computer. A computer-readable medium is any available medium that can be accessed by a computer, including both volatile and non-volatile media, and removable and non-removable media. Also, a computer-readable medium includes both storage media and communication media. A computer storage medium includes volatile and non-volatile media, removable and non-removable media, embodied in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. A communication medium typically includes modulated data signals such as computer-readable instructions, data structures, or program modules, as well as other data, or other transmission mechanisms, and includes any information delivery medium.

[0355] The above description of the present invention is for illustrative purposes, and those with ordinary knowledge in the technical field to which the present invention pertains should be able to understand that the present invention can be easily changed into other specific forms without changing its technical idea or essential features. Therefore, it should be understood that the above-described embodiments are exemplary in all aspects and not limiting. For example, each component described as a single type may be implemented in a distributed manner, and similarly, components described as being distributed may also be implemented in a combined form.

[0356] The scope of the present invention is indicated by the claims hereinafter rather than the above detailed description, and it should be interpreted that all changes or modified forms derived from the meaning and scope of the claims and the equivalent concept thereof are included in the scope of the present invention.

Explanation of Reference Numerals

[0357] 110 Conversion Unit 115 Quantization Unit 120 Inverse Quantization Unit 125 Inverse Conversion Unit 130 Filtering Unit 150 Prediction Unit 152 Intra Prediction Unit 154 Inter Prediction Unit 154a Motion Estimation Unit 154b Motion Compensation Unit 160 Entropy Coding Unit 210 Entropy Decoding Unit 220 Inverse Quantization Unit 225 Inverse Transformation Unit 230 Filtering Unit 250 Prediction Unit 252 Intra Prediction Unit 254 Inter Prediction Unit

Claims

1. A video signal processing method, comprising: receiving information for prediction of a current block; determining whether a merge mode is applicable to the current block based on the information for prediction; when the merge mode is applicable to the current block, obtaining a first syntax element indicating whether a combined prediction is applicable to the current block, wherein the combined prediction indicates a prediction mode combining an inter prediction and an intra prediction; when the first syntax element indicates that the combined prediction is applicable to the current block, generating an inter prediction block and an intra prediction block of the current block; and generating a combined prediction block of the current block by weighted-summing the inter prediction block and the intra prediction block.

2. decoding a residual block of the current block; and further comprising restoring the current block using the combined prediction block and the residual block, according to Claim 1.

3. The step of decoding the residual block further comprises: when the first syntax element indicates that the combined prediction is not applicable to the current block, further obtaining a second syntax element indicating whether a sub-block transform is applicable to the current block, wherein the sub-block transform indicates a transform mode that applies a transform to only one of the sub-blocks of the current block divided in a horizontal or vertical direction, according to Claim 2.

4. According to Claim 3, when the second syntax element does not exist, the value of the second syntax element is inferred to be 0.

5. According to Claim 1, when the first syntax element indicates that the combined prediction is applicable to the current block, the intra prediction mode for generating the intra prediction block of the current block is set to a planar mode.

6. further comprising setting positions of a left surrounding block and an upper surrounding block referred to for the combined prediction, The video signal processing method according to claim 1, wherein positions of the left surrounding block and the upper surrounding block are the same as positions referred to by the intra prediction.

7. The video signal processing method according to claim 6, wherein positions of the left surrounding block and the upper surrounding block are determined using a scaling factor variable determined by a color component index value of the current block.

8. A video signal processing apparatus, comprising a processor, wherein the processor, receives information for prediction of a current block, determines whether a merge mode is applicable to the current block based on the information for prediction, when the merge mode is applicable to the current block, obtains a first syntax element indicating whether a combined prediction is applicable to the current block, where the combined prediction indicates a prediction mode combining an inter prediction and an intra prediction, when the first syntax element indicates that the combined prediction is applicable to the current block, generates an inter prediction block and an intra prediction block of the current block, A video signal processing apparatus that generates a combined prediction block of the current block by weighted-summing the inter prediction block and the intra prediction block.

9. The processor, decodes a residual block of the current block, The video signal processing apparatus according to claim 8, wherein the current block is restored using the combined prediction block and the residual block.

10. The processor, when the first syntax element indicates that the combined prediction is not applicable to the current block, obtains a second syntax element indicating whether a sub-block transform is applicable to the current block, The video signal processing apparatus according to claim 9, wherein the sub-block transform indicates a transform mode in which a transform is applied to any one of sub-blocks of the current block divided in a horizontal or vertical direction.

11. The video signal processing apparatus according to claim 10, wherein when the second syntax element does not exist, the value of the second syntax element is inferred to be 0.

12. The video signal processing apparatus according to claim 8, wherein when the first syntax element indicates that the combined prediction is applied to the current block, the intra prediction mode for intra prediction of the current block is set to a planar mode.

13. The video signal processing apparatus according to claim 8, wherein the processor sets positions of left and upper peripheral blocks referred to for the combined prediction, and the positions of the left and upper peripheral blocks are the same as the positions referred to by the intra prediction.

14. The video signal processing apparatus according to claim 13, wherein the positions of the left and upper peripheral blocks are determined using a scaling factor variable determined by a color component index value of the current block.

15. A video signal processing method, comprising: determining whether a merge mode is applied to a current block; when the merge mode is applied to the current block, decoding a first syntax element indicating whether a combined prediction is applied to the current block, wherein the combined prediction indicates a prediction mode combining an inter prediction and an intra prediction; when the first syntax element indicates that the combined prediction is applied to the current block, generating an inter prediction block and an intra prediction block of the current block; and generating a combined prediction block of the current block by weighted-summing the inter prediction block and the intra prediction block.