Video signal processing method and apparatus using multiple assumption prediction
By constructing merge candidate lists with HMVPs and applying combined inter- and intra-prediction modes, the method enhances video signal processing efficiency, addressing inefficiencies in existing coding techniques.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- WILUS INSTITUTE OF STANDARDS & TECHNOLOGY INC
- Filing Date
- 2026-02-16
- Publication Date
- 2026-05-11
AI Technical Summary
Existing video signal processing methods lack efficiency in coding, particularly in handling spatial and temporal correlations, necessitating improved techniques for better compression and decoding.
The method involves constructing a merge candidate list using spatial candidates and incorporating history-based motion vector predictors (HMVP) to predict current blocks, while updating the HMVP table based on specific motion information, and applying combined inter- and intra-prediction modes to enhance coding efficiency.
This approach increases coding efficiency by optimizing the selection of conversion kernels for video blocks, leading to improved compression and decoding performance.
Smart Images

Figure 2026076325000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video signal processing method and apparatus, and more particularly, to a video signal processing method and apparatus for encoding or decoding a video signal.
Background Art
[0002] Compression encoding means a series of signal processing techniques for transmitting digitized information via a communication line or storing it in a form suitable for a storage medium. The targets of compression encoding include audio, video, characters, etc., and in particular, the technique of performing compression encoding for video is called video compression. Compression encoding for a video signal is performed by removing redundant information in consideration of spatial correlation, temporal correlation, probabilistic correlation, etc. However, due to the recent development of various media and data transmission media, a more efficient video signal processing method and apparatus are required.
Summary of the Invention
Problems to be Solved by the Invention
[0003] An object of the present invention is to improve the coding efficiency of a video signal. Specifically, the present invention has an object to improve the coding efficiency by using a conversion kernel suitable for a conversion block.
Means for Solving the Problems
[0004] In order to solve the above problems, the present invention provides a video signal processing apparatus and a video signal processing method as follows.
[0005] According to one embodiment of the present invention, a video signal processing method is provided, comprising the steps of: constructing a merge candidate list using spatial candidates; adding a specific HMVP in an HMVP table, which includes at least one history-based motion vector predictor (HMVP), to the merge candidate list, where the HMVP indicates motion information of blocks coded prior to the plurality of coding blocks; obtaining index information indicating a merge candidate to be used to predict the current block in the merge candidate list; and generating a predicted block for the current block based on the motion information of the merge candidate determined based on the index information, wherein if the current block is located in a merge sharing node that includes a plurality of coding blocks, the merge candidate list is constructed using spatial candidates adjacent to the merge sharing node, and the motion information of at least one coding block among the plurality of coding blocks included in the merge sharing node is not updated in the HMVP table.
[0006] Furthermore, according to one embodiment of the present invention, a video signal processing device is provided, which includes a processor, the processor constructs a merge candidate list using spatial candidates, adds a specific HMVP in an HMVP table containing at least one history-based motion vector predictor (HMVP) to the merge candidate list, wherein the HMVP obtains index information indicating a merge candidate used to predict the current block in the merge candidate list, which indicates motion information of blocks coded before the plurality of coding blocks, generates a predicted block for the current block based on the motion information of the merge candidate determined based on the index information, and if the current block is located in a merge sharing node containing a plurality of coding blocks, the merge candidate list is constructed using spatial candidates adjacent to the merge sharing node, and the motion information of at least one coding block among the plurality of coding blocks contained in the merge sharing node is not updated in the HMVP table.
[0007] As an embodiment, the process may further include a step of updating the HMVP table using motion information of a predefined number of coding blocks among the multiple coding blocks included in the merge shared node, the coding order of which is relatively slow.
[0008] As an embodiment, the process may further include a step of updating the HMVP table using the motion information of the coding block with the relatively slowest decoding order among the multiple coding blocks included in the merge shared node.
[0009] As an embodiment, if the current block is not located within the merge shared node, the process may further include the step of updating the HMVP table using the motion information of the merge candidate.
[0010] As an embodiment, the step of adding the HMVP to the merge candidate list may include the steps of: checking whether the HMVP having a specific index predefined in the HMVP table has motion information that overlaps with the candidates in the merge candidate list; and, if the HMVP having the specific index does not have motion information that overlaps with the candidates in the merge candidate list, adding the HMVP having the specific index to the merge candidate list.
[0011] Furthermore, according to one embodiment of the present invention, a video signal processing method is provided, comprising the steps of: receiving information for predicting the current block; determining whether or not a merge mode is applied to the current block based on the information for prediction; obtaining a first syntax element indicating whether or not a combined prediction is applied to the current block if a merge mode is applied to the current block, wherein the combined prediction indicates a prediction mode that combines an inter-prediction and an intra-prediction; generating an inter-prediction block and an intra-prediction block of the current block if the first syntax element indicates that the combined prediction is applied to the current block; and generating a combined prediction block of the current block by weighting the inter-prediction block and the intra-prediction block.
[0012] Furthermore, according to one embodiment of the present invention, a video signal processing device is provided, which includes a processor, wherein when a merge mode is applied to the current block, the processor obtains a first syntax element indicating whether or not a combined prediction is applied to the current block, where the combined prediction indicates a prediction mode that combines inter-prediction and intra-prediction, and when the first syntax element indicates that the combined prediction is applied to the current block, the processor generates an inter-prediction block and an intra-prediction block of the current block, and generates a combined prediction block of the current block by weighting the inter-prediction block and the intra-prediction block.
[0013] As an embodiment, the process may further include the steps of decoding the residual block of the current block and restoring the current block using the combination prediction block and the residual block.
[0014] As an embodiment, the step of decoding the residual block further includes, if the first syntax element indicates that the combinatorial prediction is not applied to the current block, a step of obtaining a second syntax element indicating whether or not a sub-block transform is applied to the current block, wherein the sub-block transform may indicate a transform mode that applies the transform to only one of the sub-blocks of the current block divided horizontally or vertically.
[0015] As an example, if the second syntax element is not present, the value of the second syntax element may be inferred to be 0.
[0016] As an example, if the first syntax element indicates that the combination prediction is applied to the current block, the intra-prediction mode for the intra-prediction for the current block may be set to planar mode.
[0017] As an embodiment, the process further includes the step of setting the positions of the left peripheral block and the upper peripheral block that are referenced for the combination prediction, and the positions of the left peripheral block and the upper peripheral block may be the same as the positions referenced by the intra prediction.
[0018] As an example, the positions of the left peripheral block and the upper peripheral block may be determined using a scaling factor variable determined by the color component index value of the current block. [Effects of the Invention]
[0019] According to an embodiment of the present invention, the coding efficiency of video signals can be increased. Furthermore, according to one embodiment of the present invention, a conversion kernel suitable for the current conversion block can be selected. [Brief explanation of the drawing]
[0020] [Figure 1] This is a schematic block diagram of a video signal encoding device according to one embodiment of the present invention. [Figure 2] This is a schematic block diagram of a video signal decoding device according to one embodiment of the present invention. [Figure 3] This figure shows an example in which a coding tree unit is divided into coding units within a picture. [Figure 4] This figure shows one embodiment of a method for signaling the splitting of quad trees and multi-type trees. [Figure 5] This figure provides a more detailed illustration of the intra-prediction method according to an embodiment of the present invention. [Figure 6] This figure provides a more detailed illustration of the intra-prediction method according to an embodiment of the present invention. [Figure 7] An interpretation method according to one embodiment of the present invention is illustrated. [Figure 8]A diagram specifically showing how an encoder converts a residual signal. [Figure 9] A diagram specifically showing how an encoder and a decoder obtain a residual signal by inverse-transforming a conversion coefficient. [Figure 10] A diagram illustrating a motion vector signaling method according to an embodiment of the present invention. [Figure 11] A diagram illustrating a signaling method for adaptive motion vector resolution information according to an embodiment of the present invention. [Figure 12] A diagram illustrating a history-based motion vector prediction (HMVP) method according to an embodiment of the present invention. [Figure 13] A diagram illustrating a method for updating an HMVP table according to an embodiment of the present invention. [Figure 14] A diagram for explaining a method for updating an HMVP table according to an embodiment of the present invention. [Figure 15] A diagram for explaining a method for updating an HMVP table according to an embodiment of the present invention. [Figure 16] A diagram for explaining a method for updating an HMVP table according to an embodiment of the present invention. [Figure 17] A diagram for explaining a method for updating an HMVP table according to an embodiment of the present invention. [Figure 18] A diagram illustrating a pruning process according to an embodiment of the present invention. [Figure 19] A diagram illustrating a method for adding an HMVP candidate according to an embodiment of the present invention. [Figure 20] A diagram illustrating a merge sharing node according to an embodiment of the present invention. [Figure 21]This diagram illustrates an HMVP update method when a shared list according to one embodiment of the present invention is used. [Figure 22] This figure illustrates a method for updating an HMVP table based on motion information of blocks within a merged shared node according to one embodiment of the present invention. [Figure 23] This diagram illustrates an HMVP update method when a shared list according to one embodiment of the present invention is used. [Figure 24] This figure illustrates a method for updating an HMVP table based on motion information of blocks within a merged shared node according to one embodiment of the present invention. [Figure 25] This figure illustrates a method for processing a video signal based on HMVP according to one embodiment of the present invention. [Figure 26] This is a diagram illustrating a multi-hypothesis prediction method according to one embodiment of the present invention. [Figure 27] This figure shows a method for determining multiple assumption prediction modes according to one embodiment of the present invention. [Figure 28] This figure shows a method for determining multiple assumption prediction modes according to one embodiment of the present invention. [Figure 29] This figure shows the peripheral position referenced in the multiple assumption prediction according to one embodiment of the present invention. [Figure 30] This figure shows a method for referencing peripheral modes according to one embodiment of the present invention. [Figure 31] This figure shows a method for generating a candidate list according to one embodiment of the present invention. [Figure 32] This figure shows a method for generating a candidate list according to one embodiment of the present invention. [Figure 33] This figure shows the peripheral position referenced in the multiple assumption prediction according to one embodiment of the present invention. [Figure 34] This figure shows a method for referencing peripheral modes according to one embodiment of the present invention. [Figure 35]This figure shows a method using a peripheral reference sample according to one embodiment of the present invention. [Figure 36] This figure shows a conversion mode according to one embodiment of the present invention. [Figure 37] This figure shows the relationship between color difference components according to one embodiment of the present invention. [Figure 38] This figure shows the relationship between color components according to one embodiment of the present invention. [Figure 39] This figure shows the peripheral reference position according to one embodiment of the present invention. [Figure 40] This figure shows a weighted sample prediction process according to one embodiment of the present invention. [Figure 41] This figure shows the peripheral reference position according to one embodiment of the present invention. [Figure 42] This figure shows a weighted sample prediction process according to one embodiment of the present invention. [Figure 43] This figure shows a weighted sample prediction process according to one embodiment of the present invention. [Figure 44] This figure shows a weighted sample prediction process according to one embodiment of the present invention. [Figure 45] This figure illustrates a video signal processing method based on multiple assumption prediction, according to one embodiment to which the present invention is applied. [Modes for carrying out the invention]
[0021] The terminology used herein has been selected as widely used and general terms as possible, taking into account the function of the present invention; however, this may vary depending on the intent of the articulators, conventions, or the emergence of new technologies. In addition, in certain cases, the applicant has arbitrarily selected some terms, in which case their meaning will be described in the section describing the form of implementation of the invention. Therefore, it is important to clarify that the terminology used herein is not merely a set of names, but should be interpreted based on the substantive meaning of the term and the overall content of this specification.
[0022] In this specification, some terms are interpreted as follows: Coding may be interpreted as encoding or coding in some cases. In this specification, a device that encodes a video signal to generate a video signal bitstream is referred to as an encoding device or encoder, and a device that decodes a video signal bitstream to restore a video signal is referred to as a decoding device or decoder. In this specification, the term video signal processing device is used as a conceptual term that includes both encoders and decoders. Information is a term that includes values, parameters, coefficients, elements, etc., and may be interpreted differently in some cases, so the present invention is not limited thereto. 'Unit' is used to mean a basic unit of image processing or a specific location in a picture, and refers to an image region that includes at least one of the luma component and chroma component. Also, 'block' refers to an image region that includes a specific component among the luminance component and chrominance component (i.e., Cb and Cr). However, in some embodiments, terms such as 'unit', 'block', 'partition', and 'region' may be used interchangeably. Furthermore, in this specification, the term "unit" is used as a concept that includes coding units, prediction units, and transformation units. "Picture" refers to a field or frame, and in some embodiments, these terms are used interchangeably.
[0023] Figure 1 is a schematic block diagram of a video signal encoding device 100 according to one embodiment of the present invention. Referring to Figure 1, the encoding device 100 according to this specification includes a conversion unit 110, a quantization unit 115, an inverse quantization unit 120, an inverse conversion unit 125, a filtering unit 130, a prediction unit 150, and an entropy coding unit 160.
[0024] The conversion unit 110 converts the residual signal, which is the difference between the input video signal and the predicted signal generated by the prediction unit 150, to obtain the converted coefficient values. For example, discrete cosine transform (DCT), discrete sine transform (DST), or wavelet transform may be used. Discrete cosine transform and discrete sine transform divide the input picture signal into block form and perform the transformation. In the transformation, the coding efficiency may differ depending on the distribution and characteristics of the values within the transformation domain. The quantization unit 115 quantizes the values of the conversion coefficients output in the conversion unit 110.
[0025] To improve coding efficiency, instead of directly coding the picture signal, a method is used in which the picture is predicted using a pre-coded region via the prediction unit 150, and the restored picture is obtained by adding the residual value between the original picture and the predicted picture to the predicted picture. To prevent mismatches between the encoder and decoder, the encoder should use information that is also available to the decoder when performing prediction. For this purpose, the encoder performs a further process of restoring the currently encoded block. The inverse quantization unit 120 inversely quantizes the conversion coefficient values, and the inverse transformation unit 125 restores the residual value using the inversely quantized conversion coefficient values. Meanwhile, the filtering unit 130 performs filtering operations to improve the quality of the restored picture and enhance coding efficiency. For example, this may include a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter. The filtered picture is stored in the Decoded Picture Buffer (DPB) 156 for output or use as a reference picture.
[0026] To improve coding efficiency, instead of directly coding the picture signal, a method is used in which the picture is predicted using an already coded region in the prediction unit 150, and the restored picture is obtained by adding the residual value between the original picture and the predicted picture to the predicted picture. The intra-prediction unit 152 performs in-screen prediction within the current picture, and the inter-prediction unit 154 predicts the current picture using a reference picture stored in the decoded picture buffer 156. The intra-prediction unit 152 performs in-screen prediction from the restored region within the current picture and transmits the in-screen coding information to the entropy coding unit 160. The inter-prediction unit 154 may further include a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains the motion vector value of the current region by referring to a specific restored region. The motion estimation unit 154a transmits position information of the reference region (reference frame, motion vector, etc.) to the entropy coding unit 160 so that it can be included in the bitstream. The motion compensation unit 154b performs inter-screen motion compensation using the motion vector values transmitted from the motion estimation unit 154a.
[0027] The prediction unit 150 includes an intra-prediction unit 152 and an inter-prediction unit 154. The intra-prediction unit 152 performs intra-prediction within the current picture, and the inter-prediction unit 154 performs inter-prediction, predicting the current picture using a reference buffer stored in the decoded picture buffer 156. The intra-prediction unit 152 performs intra-prediction from the restored samples in the current picture and transmits intra-coded information to the entropy coding unit 160. The intra-coded information includes at least one of the following: intra-prediction mode, MPM (Most Probable Mode) flag, and MPM index. The intra-coded information may include information about the reference sample. The inter-prediction unit 154 includes a motion estimation unit 154a and a motion compensation unit 154b. The motion estimation unit 154a obtains motion vector values for the current region by referring to a specific region of the restored reference signal picture. The motion estimation unit 154a transmits a set of motion information (reference picture index, motion vector information) for the reference region to the entropy coding unit 160. The motion compensation unit 154b performs motion compensation using the motion vector values transmitted from the motion compensation unit 154a. The inter-prediction unit 154 transmits inter-coded information, including motion information for the reference region, to the entropy coding unit 160.
[0028] In a further embodiment, the prediction unit 150 includes an intrablock copy (BC) prediction unit (not shown). The intraBC prediction unit performs intraBC prediction from the restored samples in the current picture and transmits the intraBC encoded information to the entropy coding unit 160. The intraBC prediction unit obtains a block vector value indicating a reference region used for prediction of the current region by referring to a specific region in the current picture. The intraBC prediction unit performs intraBC prediction using the obtained block vector value. The intraBC prediction unit transmits the intraBC encoded information to the entropy coding unit 160. The intraBC prediction unit includes block vector information.
[0029] Once the picture prediction described above is performed, the conversion unit 110 converts the residual values between the original picture and the predicted picture to obtain conversion coefficient values. In this case, the conversion is performed in units of specific blocks within the picture, but the size of the specific block is variable within a preset range. The quantization unit 115 quantizes the values of the conversion coefficients generated by the conversion unit 110 and transmits them to the entropy coding unit 160.
[0030] The entropy coding unit 160 generates a video signal bitstream by entropy coding information indicating quantized conversion coefficients, intra-coded information, and inter-coded information. The entropy coding unit 160 uses methods such as variable length coding (VLC) and arithmetic coding. Variable length coding (VLC) converts input symbols into a sequence of codewords, but the length of the codewords is variable. For example, frequently occurring symbols are represented by short codewords, and less frequently occurring symbols are represented by long codewords. Context-based Adaptive Variable Length Coding (CAVLC) is used as the variable length coding method. Arithmetic coding converts a sequence of data symbols into a single prime number, obtaining the optimal number of prime bits necessary to represent each symbol. Context-based Adaptive Binary Arithmetic Coding (CABAC) is used as the arithmetic coding method. For example, the entropy coding unit 160 can binary-code information indicating quantized transformation coefficients. Furthermore, the entropy coding unit 160 can arithmetically encode the binary-coded information to generate a bitstream.
[0031] The generated bitstream is encapsulated in Network Abstraction Layer (NAL) units as its basic units. Each NAL unit contains an encoded integer number of coding tree units. To decode the bitstream with a video decoder, the bitstream must first be separated into NAL units, and then each separated NAL unit must be decoded. Meanwhile, the information necessary for decoding the video signal bitstream is transmitted via Raw Byte Sequence Payloads (RBSPs) of higher-level sets such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS).
[0032] On the other hand, the block diagram in Figure 1 shows an encoding device 100 according to one embodiment of the present invention, and the separated blocks show the elements of the encoding device 100 in a logically distinguishable manner. Thus, the elements of the encoding device 100 described above are mounted on one chip or multiple chips depending on the device design. According to one embodiment, the operation of each element of the encoding device 100 described above is performed by a processor (not shown).
[0033] Figure 2 is a schematic block diagram of a video signal decoding apparatus 200 according to an embodiment of the present invention. Referring to Figure 2, the decoding apparatus 200 according to this specification includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transformation unit 225, a filtering unit 230, and a prediction unit 250.
[0034] The entropy decoding unit 210 entropy decodes the video signal bitstream and extracts conversion coefficient information, intra-encoded information, inter-encoded information, etc., for each region. For example, the entropy decoding unit 210 can obtain binary-coded conversion coefficient information for a specific region from the video signal bitstream. The entropy decoding unit 210 also performs inverse binary-coded conversion of the binary-coded conversion coefficient to obtain quantized conversion coefficients. The inverse quantization unit 220 inverse quantizes the quantized conversion coefficients, and the inverse conversion unit 225 restores the residual value using the inverse quantized conversion coefficient. The video signal processing device 200 adds the residual value obtained from the inverse conversion unit 225 with the predicted value obtained from the prediction unit 250 to restore the original pixel value.
[0035] Meanwhile, the filtering unit 230 improves image quality by filtering the picture. This includes a deblocking filter to reduce block distortion and / or an adaptive loop filter to remove distortion from the entire picture. The filtered picture is either output or stored in the decoded picture buffer (DPB) 256 to be used as a reference picture for the next picture.
[0036] The prediction unit 250 includes an intra-prediction unit 252 and an inter-prediction unit 254. The prediction unit 250 generates a prediction picture by utilizing the encoding type decoded through the aforementioned entropy decoding unit 210, the conversion coefficients for each region, intra / inter-encoded information, etc. To restore the current block from which decoding is performed, the current picture containing the current block or the decoded region of another picture can be used. A picture (or tile / slice) that uses only the current picture for restoration, i.e., performs intra-prediction or intra-BC prediction, is called an intra-picture or I-picture (or tile / slice), and a picture (or tile / slice) that can perform intra-prediction, inter-prediction, and intra-BC prediction is called an inter-picture (or tile / slice). Among interpictures (or tiles / slice), a picture (or tile / slice) that uses up to one motion vector and a reference picture index to predict the sample value of each block is called a predictive picture or P-picture (or tile / slice), and a picture (or tile / slice) that uses up to two motion vectors and a reference picture index is called a bi-predictive picture or B-picture (or tile / slice). In other words, a P-picture (or tile / slice) uses up to one motion information set to predict each block, and a B-picture (or tile / slice) uses up to two motion information sets to predict each block. Here, a motion information set includes one or more motion vectors and one reference picture index.
[0037] The intra-prediction unit 252 generates a prediction block using intra-encoded information and the recovered sample in the current picture. As described above, the intra-encoded information includes at least one of the intra-prediction mode, the MPM (MOST Probable Mode) flag, and the MPM index. The intra-prediction unit 252 predicts the sample value of the current block using the recovered sample located to the left and / or above the current block as a reference sample. In this disclosure, the recovered sample, the reference sample, and the sample of the current block represent pixels. The sample value also represents the pixel value.
[0038] In one embodiment, the reference sample is a sample included in the surrounding blocks of the current block. For example, the reference sample is a sample adjacent to the left boundary and / or the upper boundary of the current block. Alternatively, the reference sample is a sample located in the surrounding blocks of the current block that lies on a line within a predetermined distance from the left boundary and / or on a line within a predetermined distance from the upper boundary of the current block. In this case, the surrounding blocks of the current block include at least one of the following blocks adjacent to the current block: the left (L) block, the upper (A) block, the lower left (BL) block, the upper right (AR) block, or the upper left (AL) block.
[0039] The interprediction unit 254 generates a prediction block using the reference picture and intercoded information stored in the decoded picture buffer 256. The intercoded information includes a set of motion information for the current block relative to the reference block (reference picture index, motion vector, etc.). Interpretation includes L0 prediction, L1 prediction, and bi-prediction. L0 prediction is a prediction that uses one reference picture included in the L0 picture list, and L1 prediction means a prediction that uses one reference picture included in the L1 picture list. For this, one set of motion information (e.g., motion vector and reference picture index) is required. In the bi-prediction method, up to two reference regions are used, but these two reference regions may reside in the same reference picture or in different pictures. In other words, in the bi-prediction method, up to two sets of motion information (e.g., motion vector and reference picture index) are used, but the two motion vectors may correspond to the same reference picture index or to different reference picture indices. In this case, the reference picture is displayed (or output) either before or after the current picture in terms of time. In one embodiment, the two reference regions used in the dual prediction method may be regions selected from the L0 picture list and the L1 picture list, respectively.
[0040] The interpretation unit 254 obtains the current reference block using the motion vector and the reference picture index. The reference block resides within the reference picture corresponding to the reference picture index. The sample value of the block identified by the motion vector, or an interpolated value thereof, is used as the predictor for the current block. For motion prediction with sub-pel pixel accuracy, for example, an 8-tab interpolation filter is used for the luminance signal and a 4-tab interpolation filter is used for the chrominance signal. However, the interpolation filter for sub-pel motion prediction is not limited to these. In this way, the interpretation unit 254 performs motion compensation, predicting the texture of the current unit from the previously restored picture. In this process, the interpretation unit utilizes a motion information set.
[0041] In a further embodiment, the prediction unit 250 may include an intra-BC prediction unit (not shown). The intra-BC prediction unit can reconstruct the current region by referring to a specific region containing the reconstructed sample in the current picture. The intra-BC prediction unit obtains intra-BC encoded information for the current region from the entropy decoding unit 210. The intra-BC prediction unit obtains a block vector value of the current region that points to a specific region in the current picture. The intra-BC prediction unit can perform intra-BC prediction using the obtained block vector value. The intra-BC encoded information may include block vector information.
[0042] A video picture is generated by summing the predicted values output from the intra-prediction unit 252 or the inter-prediction unit 254 and the residual values output from the inverse conversion unit 225. In other words, the video signal decoding device 200 restores the current block using the predicted block generated by the prediction unit 250 and the residual obtained from the inverse conversion unit 225.
[0043] On the other hand, the block diagram in Figure 2 shows a decoding device 200 according to one embodiment of the present invention, and the separated blocks show the elements of the decoding device 200 in a logically distinguishable manner. Thus, the elements of the decoding device 200 described above are mounted on one chip or multiple chips depending on the device design. In one embodiment, the operation of each element of the decoding device 200 described above is performed by a processor (not shown).
[0044] Figure 3 shows an example in which a Coding Tree Unit (CTU) is divided into Coding Units (CUs) within a picture. In the coding process of a video signal, the picture is divided into a sequence of Coding Tree Units (CTUs). A Coding Tree Unit consists of two blocks: an NXN block of luminance samples and its corresponding chrominance samples. A Coding Tree Unit is divided into multiple Coding Units. A Coding Tree Unit may also become a leaf node without being divided. In this case, the Coding Tree Unit itself can become a Coding Unit. A Coding Unit refers to a basic unit for processing a picture in the video signal processing processes described above, i.e., intra / inter prediction, transformation, quantization, and / or entropy coding. Within a single picture, the size and pattern of the Coding Units are not constant. Coding Units have a square or rectangular pattern. A rectangular Coding Unit (or rectangular block) includes vertical Coding Units (or vertical blocks) and horizontal Coding Units (or horizontal blocks). In this specification, a vertical block is a block whose height is greater than its width, and a horizontal block is a block whose width is greater than its height. Furthermore, in this specification, a non-square block refers to a rectangular block, but the present invention is not limited to this.
[0045] Referring to Figure 3, the coding tree unit is first divided into a quad tree (QT) structure. That is, in the quad tree structure, one node with a size of 2N × 2N is divided into four nodes with a size of N × N. In this specification, the quad tree is also referred to as a quaternary tree. The quad tree division is performed recursively, and it is not necessary for all nodes to be divided to the same depth.
[0046] On the other hand, the leaf nodes of the quad tree described above are further divided into a multi-type tree (MTT) structure. According to embodiments of the present invention, in a multi-type tree structure, one node is divided into a binary or ternary tree structure with horizontal or vertical division. In other words, there are four division structures in a multi-type tree structure: vertical binary division, horizontal binary division, vertical ternary division, and horizontal ternary division. According to embodiments of the present invention, in each of the tree structures, the width and height of the node are both powers of 2. For example, in a binary tree (BT) structure, a node of size 2N×2N is divided into two N×2N nodes by vertical binary division and into two 2N×N nodes by horizontal binary division. Furthermore, in a Ternary Tree (TT) structure, a node of size 2N×2N is divided into (N / 2)×2N, N×2N, and (N / 2)×2N nodes by vertical ternary decomposition, and into 2N×(N / 2), 2N×N, and 2N×(N / 2) nodes by horizontal ternary decomposition. Such multi-type tree decomposition is performed recursively.
[0047] Leaf nodes in a multi-type tree can be coding units. If no splitting is instructed for a coding unit, or if the coding unit is not larger than the maximum transformation length, the coding unit is used as the unit for prediction and transformation without further splitting. On the other hand, in the quad tree and multi-type tree described above, at least one of the following parameters is predefined or transmitted via RBSP of a higher-level set such as PPS, SPS, VPS, etc.: 1) CTU size: the size of the root node of the quad tree, 2) MinQtSize: the minimum allowed QT leaf node size, 3) MaxBtSize: the maximum allowed BT root node size, 4) MaxTtSize: the maximum allowed TT root node size, 5) MaxMttDepth: the maximum allowed depth of MTT splitting from the leaf nodes of the QT, 6) MinBtSize: the minimum allowed BT leaf node size, 7) MinTT size: the minimum allowed TT leaf node size.
[0048] Figure 4 illustrates one embodiment of a method for signaling the splitting of quad trees and multi-type trees. The flags previously set to signal the splitting of quad trees and multi-type trees can be used. Referring to Figure 4, at least one of the following can be used: 'qt_split_flag' which indicates whether or not to split a quad tree node, 'mtt_split_flag' which indicates whether or not to split a multi-type tree node, 'mtt_split_vertical_flag' which indicates the splitting direction of a multi-type tree node, or 'mtt_split_binary_flag' which indicates the splitting form of a multi-type tree node.
[0049] According to an embodiment of the present invention, the coding tree unit is the root node of a quad tree and can be divided into a quad tree structure beforehand. In the quad tree structure, each node 'QT_node' is signaled with 'qt_split_flag'. If the value of 'qt_split_flag' is 1, the node is divided into four square nodes, and if the value of 'qt_split_flag' is 0, the node becomes a leaf node 'QT_leaf_node' of the quad tree.
[0050] Each quadtree leaf node 'QT_leaf_node' can be further divided into a multi-type tree structure. In the multi-type tree structure, each node 'MTT_node' is signaled with 'mtt_split_flag'. If the value of 'mtt_split_flag' is 1, the node is divided into multiple rectangular nodes, and if the value of 'mtt_split_flag' is 0, the node becomes a leaf node 'MTT_leaf_node' in the multi-type tree. When a multi-type tree node 'MTT_node' is divided into multiple rectangular nodes (i.e., when the value of 'mtt_split_flag' is 1), additional 'mtt_split_vertical_flag' and 'mtt_split_binary_flag' can be signaled for the node 'MTT_node'. If the value of 'mtt_split_vertical_flag' is 1, the node 'MTT_node' is instructed to be split vertically. If the value of 'mtt_split_vertical_flag' is 0, the node 'MTT_node' is instructed to be split horizontally. Also, if the value of 'mtt_split_binary_flag' is 1, the node 'MTT_node' is split into two rectangular nodes. If the value of 'mtt_split_binary_flag' is 0, the node 'MTT_node' is split into three rectangular nodes.
[0051] Picture prediction (motion compensation) for coding is performed on coding units that cannot be further divided (i.e., leaf nodes in the coding unit tree). The basic unit for performing such predictions is referred to below as a prediction unit or prediction block.
[0052] Hereinafter, the term "unit" as used herein is used as a substitute for the prediction unit, which is the basic unit for making predictions. However, the present invention is not limited thereto and is understood in a broader sense as a concept that includes the coding unit.
[0053] Figures 5 and 6 illustrate in more detail the intra-prediction method according to an embodiment of the present invention. As described above, the intra-prediction unit predicts the sample value of the current block by using the restored sample located to the left and / or above the current block as a reference sample.
[0054] First, Figure 5 shows an example of a reference sample used to predict the current block in intra-prediction mode. In this example, the reference sample is a sample adjacent to the left boundary and / or the upper boundary of the current block. As shown in Figure 5, if the size of the current block is W×H and a single reference line sample adjacent to the current block is used for intra-prediction, the reference sample is set using up to 2W+2H+1 peripheral samples located to the left and / or above the current block.
[0055] Furthermore, if at least some of the samples to be used as reference samples have not yet been recovered, the intra-prediction unit performs a reference sample padding process to acquire reference samples. The intra-prediction unit also performs a reference sample filtering process to reduce the error of the intra-prediction. That is, it filters the peripheral samples and / or the reference samples acquired by the reference sample padding process to acquire filtered reference samples. The intra-prediction unit uses the reference samples thus acquired to predict the samples of the current block. The intra-prediction unit predicts the samples of the current block using either the unfiltered or filtered reference samples. In this disclosure, peripheral samples may include samples on at least one reference line. For example, peripheral samples may include adjacent samples on lines adjacent to the boundary of the current block.
[0056] Next, Figure 6 illustrates one embodiment of a prediction mode used for intra-prediction. For intra-prediction, intra-prediction mode information indicating the direction of intra-prediction can be signaled. The intra-prediction mode information indicates one of several intra-prediction modes that constitute an intra-prediction mode set. If the current block is an intra-predicted block, the decoder receives the intra-prediction mode information for the current block from the bitstream. The intra-prediction unit of the decoder performs intra-prediction for the current block based on the extracted intra-prediction mode information.
[0057] According to an embodiment of the present invention, the intra-prediction mode set includes all intra-prediction modes used for intra-prediction (e.g., a total of 67 intra-prediction modes). More specifically, the intra-prediction mode set includes planar modes, DC modes, and a plurality of (e.g., 65) angular modes (i.e., direction modes). Each intra-prediction mode is indicated by a preset index (i.e., an intra-prediction mode index). For example, as shown in Figure 6, intra-prediction mode index 0 indicates the planar mode, and intra-prediction mode index 1 indicates the DC mode. Intra-prediction mode indices 2 through 66 each indicate different angular modes. Each angular mode indicates a different angle within a pre-set angular range. For example, an angular mode can indicate an angle within an angular range of 45° to -135° in a clockwise direction (i.e., a first angular range). The angular modes may be defined relative to the 12 o'clock direction. In this case, intra-prediction mode index 2 indicates horizontal diagonal (HDIA) mode, intra-prediction mode index 18 indicates horizontal (HOR) mode, intra-prediction mode index 34 indicates diagonal (DIA) mode, intra-prediction mode index 50 indicates vertical (VER) mode, and intra-prediction mode index 66 indicates vertical diagonal (VDIA) mode.
[0058] The interpretation method according to an embodiment of the present invention will be described below with reference to Figure 7. In the present invention, the interpretation method can include a general interpretation method optimized for translation motion and an affine model-based interpretation method. Furthermore, the motion vector can typically include at least one of a general motion vector for motion compensation and a control point motion vector for affine motion compensation, based on the interpretation method.
[0059] Figure 7 illustrates an interpretation method according to one embodiment of the present invention. As previously described, the decoder can predict the current block by referring to a restored sample of another decoded picture. Referring to Figure 7, the decoder obtains a reference block 702 in reference picture 720 based on a motion information set of the current block 701. The motion information set may include a reference picture index and a motion vector 703. The reference picture index points to a reference picture 720 containing a reference block for interpretation of the current block in the reference picture list. According to one embodiment, the reference picture list may include at least one of the L0 picture list or L1 picture list described above. The motion vector 703 indicates the offset between the coordinate values of the current block 701 in the current picture 710 and the coordinate values of the reference block 702 in reference picture 720. The decoder obtains a predictor of the current block 701 based on the sample values of the reference block 702 and uses the predictor to restore the current block 701.
[0060] Specifically, the encoder can find a reference block by searching for a block similar to the current block using the picture with the earliest restoration order. For example, the encoder can search for a reference block within a pre-configured search area that minimizes the sum of the differences between the current block and the sample values. In this process, at least one of either SAD (Sum of Absolute Difference) or SATD (Sum of Hadamard Transformed Difference) can be used to measure the similarity between the samples in the current block and the reference block. Here, SAD can be the sum of all the absolute values of the differences between the sample values contained in the two blocks. SATD can be the sum of all the absolute values of the Hadamard Transform coefficients obtained by performing a Hadamard Transform on the differences between the sample values contained in the two blocks.
[0061] On the other hand, the current block can also be predicted using one or more reference regions. As mentioned above, the current block can be interpreted using a biprediction scheme with two or more reference regions. According to one embodiment, the decoder can acquire two reference blocks based on two motion information sets of the current block. The decoder can also acquire a first predictor and a second predictor of the current block based on the sample values of each of the two acquired reference blocks. The decoder can then reconstruct the current block using the first and second predictors. For example, the decoder can reconstruct the current block based on the sample-by-sample average of the first and second predictors.
[0062] As mentioned above, one or more motion information sets can signal for motion compensation of the current block. In this case, the similarity between the motion information sets for each of the multiple blocks' motion compensations can be utilized. For example, the motion information set used for predicting the current block can be derived from the motion information set used for predicting any one of the other samples that have already been reconstructed. Through this, the encoder and decoder can reduce signaling overhead. Various embodiments in which the motion information set of the current block is signaled will be described below.
[0063] For example, there may be multiple candidate blocks that could have been predicted based on motion information sets identical or similar to the motion information set of the current block. The decoder can generate a merge candidate list based on these multiple candidate blocks. Here, the merge candidate list may include candidates corresponding to samples that were restored before the current block and could have been predicted based on motion information sets associated with the motion information set of the current block. The encoder and decoder can construct the merge candidate list of the current block based on predefined rules. In this case, the merge candidate lists constructed by the encoder and decoder may be identical to each other. For example, the encoder and decoder can construct the merge candidate list of the current block based on the position of the current block in the current picture. The method by which the encoder and decoder construct the merge candidate list of the current block will be described later in Figure 9. In this disclosure, the position of a particular block represents the relative position of the top-left sample of the particular block in the picture containing the particular block.
[0064] On the other hand, to improve coding efficiency, instead of directly coding the residual signal as described above, a method can be used in which the transformed residual signal is converted, the resulting conversion coefficient value is quantized, and the quantized conversion coefficient is coded. As mentioned above, the conversion unit can convert the residual signal to obtain the conversion coefficient value. In this case, the residual signal of a particular block may be distributed throughout the entire area of the current block. This allows for the concentration of energy in the low-frequency region using frequency domain conversion on the residual signal, thereby improving coding efficiency. The following describes in detail how the residual signal is converted or inversely converted.
[0065] Figure 8 is a diagram illustrating how the encoder converts a resistive signal. As previously mentioned, the resistive signal in the spatial domain may be converted to the frequency domain. The encoder can convert the acquired resistive signal to obtain conversion coefficients. First, the encoder can acquire at least one resistive block containing the resistive signal for the current block. The resistive block may be either the current block or a block divided from the current block. In this disclosure, a resistive block may be referred to as a resistive array or resistive matrix containing resistive samples of the current block. Also in this disclosure, a resistive block represents a block the same size as the conversion unit or conversion block.
[0066] Next, the encoder can transform the residual block using a transformation kernel. The transformation kernel used for the transformation of the residual block may be a transformation kernel having separable characteristics for vertical and horizontal transformations. In this case, the transformation of the residual block may be performed separately as a vertical and horizontal transformation. For example, the encoder can perform a vertical transformation by applying the transformation kernel to the vertical direction of the residual block. The encoder can also perform a horizontal transformation by applying the transformation kernel to the horizontal direction of the residual block. In this disclosure, the term "transformation kernel" may be used to represent a set of parameters used for the transformation of a residual signal, such as a transformation matrix, transformation array, transformation function, or transformation. In one embodiment, the transformation kernel may be one of several available kernels. Furthermore, different transformation types based on transformation kernels may be used for the vertical and horizontal transformations, respectively. A method for selecting one of several available transformation kernels will be described later in Figures 12 to 26.
[0067] The encoder can quantize a transformed block obtained from a residual block by transmitting it to a quantization unit. The transformed block may contain multiple transformation coefficients. Specifically, the transformed block may consist of multiple transformation coefficients arranged in a two-dimensional array. The size of the transformed block may be the same as that of the residual block, either the current block or a block divided from the current block. The transformation coefficients transmitted to the quantization unit may be represented by their quantized values.
[0068] Furthermore, the encoder can perform additional transformations before the transformation coefficients are quantized. As shown in Figure 8, the transformation method described above can be called a primary transform, and the additional transformation can be called a secondary transform. The secondary transform may be selective for each resistive block. In one embodiment, the encoder can improve coding efficiency by performing a secondary transform on regions where it is difficult to concentrate energy in the low-frequency region with only a primary transform. For example, a secondary transform may be added to blocks where the resistive value appears large in directions other than the horizontal or vertical direction of the resistive block. The resistive value of an intra-predicted block may have a higher probability of changing in directions other than the horizontal or vertical direction compared to the resistive value of an inter-predicted block. This allows the encoder to perform a further secondary transform on the resistive signal of an intra-predicted block. Alternatively, the encoder may omit the secondary transform on the resistive signal of an inter-predicted block.
[0069] As another example, whether or not to perform a quadratic transformation may be determined by the size of the current block or residual block. Also, different size transformation kernels may be used depending on the size of the current block or residual block. For example, an 8x8 quadratic transformation may be applied to blocks where the length of the shorter side of the width or height is shorter than a first already set length. Also, a 4x4 quadratic transformation may be applied to blocks where the length of the shorter side of the width or height is longer than a second already set length. In this case, the first already set length may be a larger value than the second already set length, but this disclosure is not limited to this. Furthermore, unlike a linear transformation, a quadratic transformation does not have to be performed by separating it into a vertical transformation and a horizontal transformation. Such a quadratic transformation can be called a Low Frequency Non-Separable Transform (LFNST).
[0070] Furthermore, in the case of video signals in a specific region, abrupt brightness changes may prevent the reduction of high-frequency bandwidth energy even after frequency conversion. This can lead to a decrease in compression performance due to quantization. Also, when performing conversion on a region where residual values rarely exist, encoding and decoding times may increase unnecessarily. For this reason, conversion of residual signals in a specific region may be omitted. Whether or not to perform conversion on residual signals in a specific region may be determined by a syntax element associated with the conversion of that region. For example, the syntax element may include transform skip information. Transform skip information may be a transform skip flag. If the transform skip information for a residual block indicates a transform skip, the conversion of that residual block is not performed. In this case, the encoder can immediately quantize the residual signal in that region that is not converted. The operation of the encoder described with reference to Figure 8 can be performed by the conversion unit in Figure 1.
[0071] The aforementioned conversion-related syntax elements may be information parsed from the video signal bitstream. A decoder can obtain the conversion-related syntax elements by entropy decoding the video signal bitstream. An encoder can generate the video signal bitstream by entropy coding the conversion-related syntax elements.
[0072] Figure 9 is a diagram illustrating in detail how an encoder and decoder obtain a residual signal by inversely transforming the conversion coefficients. For the sake of explanation, it will be assumed that the inverse transformation operation is performed in each inverse transformation section of the encoder and decoder. The inverse transformation section can obtain a residual signal by inversely transforming the inversely quantized conversion coefficients. First, the inverse transformation section can detect whether or not an inverse transformation is performed for a particular region from the conversion-related syntax elements of that region. In one embodiment, if the conversion-related syntax elements for a particular conversion block indicate a conversion skip, the transformation for that conversion block may be omitted. In this case, the aforementioned first-order and second-order inverse transformations may all be omitted for the conversion block. Furthermore, the inversely quantized conversion coefficients may be used as a residual signal. For example, the decoder can use the inversely quantized conversion coefficients as a residual signal to reconstruct the current block.
[0073] In other embodiments, the transformation-related syntax elements for a particular transformation block do not necessarily have to represent a transformation skip. In this case, the inverse transformation unit can decide whether or not to perform a quadratic inverse transformation on a quadratic transformation. For example, if the transformation block is a transformation block of an intra-predicted block, a quadratic inverse transformation may be performed on the transformation block. Alternatively, the quadratic transformation kernel used for the transformation block may be determined based on the intra-prediction mode corresponding to the transformation block. As another example, the decision to perform a quadratic inverse transformation may be made based on the size of the transformation block. The quadratic inverse transformation may be performed after the inverse quantization process and before the linear inverse transformation is performed.
[0074] The inverse transformer can perform a linear inverse transform on the inversely quantized transform coefficients or the quadratic inversely transformed transform coefficients. In the case of a linear inverse transform, it may be separated into a vertical transform and a horizontal transform, similar to the linear transform. For example, the inverse transformer can obtain a residual block by performing a vertical inverse transform and a horizontal inverse transform on the transform block. The inverse transformer can inverse the transform block based on the transform kernel used to transform the transform block. For example, the encoder can explicitly or implicitly signal information indicating which transform kernel is currently applied to the transform block from among several available transform kernels. The decoder can use the signaled transform kernel information to select the transform kernel to be used for the inverse transform of the transform block from among several available transform kernels. The inverse transformer can reconstruct the current block using the residual signal obtained by the inverse transform on the transform coefficients.
[0075] Figure 10 illustrates a motion vector signaling method according to one embodiment of the present invention. According to one embodiment of the present invention, the motion vector (MV) may be generated based on a motion vector prediction (or predictor) (MVP). For example, the MV may be determined to be the same as the MVP as shown in the following mathematical formula 1. In other words, the MV may be determined (or set, guided) to the same value as the MVP.
[0076]
number
[0077] As another example, MV may be determined based on MVP and motion vector difference (MVD), as shown in the following mathematical equation 2. The encoder can signal MVD information to the decoder to indicate a more accurate MV, and the decoder can induce MV by adding the acquired MVD to MVP.
[0078]
number
[0079] According to one embodiment of the present invention, the encoder transmits determined motion information to the decoder, and the decoder can generate (or induce) an MV from the received motion information and generate a prediction block based on it. For example, the motion information may include MVP information and MVD information. In this case, the components of the motion information may change depending on the interprediction mode. For example, in merge mode, the motion information may include MVP information but not MVD information. As another example, in AMVP (advanced motion vector prediction) mode, the motion information may include both MVP information and MVD information.
[0080] To determine, transmit, and receive information about the MVP, the encoder and decoder can generate MVP candidates (or a list of MVP candidates) in the same way. For example, the encoder and decoder can generate the same MVP candidates in the same order. The encoder then transmits an index to the decoder indicating the determined (or selected) MVP from the generated MVP candidates, and the decoder can derive the determined MVP and / or MV based on the received index.
[0081] According to one embodiment of the present invention, the MVP candidate may include a spatial candidate, a temporal candidate, and the like. When merge mode is applied, the MVP candidate may be called a merge candidate, and when AMVP mode is applied, it may be called an AMVP candidate. The spatial candidate may be the MV (or motion information) for a block at a specific position relative to the current block. For example, the spatial candidate may be the MV of a block at a position adjacent to or not adjacent to the current block. The temporal candidate may be the MV corresponding to the current picture and blocks in other pictures. Furthermore, for example, the MVP candidate may include affine MV, ATMVP, STMVP, a combination of the aforementioned MVs (or candidates), the average MV of the aforementioned MVs (or candidates), zero MV, and the like.
[0082] In one embodiment, the encoder can signal information indicating a reference picture to the decoder. In one embodiment, if the reference picture of the MVP candidate is different from the reference picture of the current block (or the current processing block), the encoder / decoder can scale the MV of the MVP candidate (motion vector scaling). In this case, the MV scaling may be performed based on the picture order count (POC) of the current picture, the POC of the reference picture of the current block, and the POC of the reference picture of the MVP candidate.
[0083] The following describes specific examples of MVD signaling methods. Table 1 below illustrates the syntax structure for MVD signaling.
[0084] [Table 1]
[0085] Referring to Table 1, according to one embodiment of the present invention, the MVD may be coded with its sign and absolute value separately. That is, the sign and absolute value of the MVD may be represented by different syntax (or syntax elements). The absolute value of the MVD may be coded directly, or it may be coded stepwise based on a flag indicating whether the absolute value is greater than N, as shown in Table 1. If the absolute value is greater than N (absolute value - N), both may be signaled. Specifically, in the example in Table 1, an abs_mvd_greater0_flag indicating whether the absolute value is greater than 0 may be transmitted. If the abs_mvd_greater0_flag indicates (or suggests) that the absolute value is not greater than 0, the absolute value of the MVD may be determined to be 0. Furthermore, if abs_mvd_greater0_flag indicates that its absolute value is greater than 0, additional syntax (or syntax element) may be present.
[0086] For example, an abs_mvd_greater1_flag indicating whether the absolute value is greater than 1 may be transmitted. If abs_mvd_greater1_flag indicates (or suggests) that the absolute value is not greater than 1, then the absolute value of the MVD may be determined to be 1. If abs_mvd_greater1_flag indicates that the absolute value is greater than 1, then additional syntax may exist. For example, abs_mvd_minus2 may exist. abs_mvd_minus2 may be the value of (absolute value - 2). Since the absolute value has been determined to be greater than 1 (i.e., 2 or greater) by the values of abs_mvd_greater0_flag and abs_mvd_greater1_flag, the value of (absolute value - 2) may be signaled. In this way, by hierarchically syntactically signaling information about absolute values, fewer bits are used compared to simply binary-coding the absolute values and signaling them directly.
[0087] In one embodiment, the syntax related to absolute values described above may be coded by applying binary evolution methods with variable lengths such as Exponential-Golomb, truncated unary, or truncated Rice. Furthermore, a flag indicating the sign of the MVD may be signaled by mvd_sign_flag.
[0088] In the above-described embodiment, a coding method for MVD was explained, but information other than MVD can also be signaled by separating the sign and absolute value. The absolute value may be coded as a flag indicating whether the absolute value is greater than a predetermined specific value, and as the value obtained by subtracting the specific value from the absolute value. In Table 1, [0] and [1] can represent component indices. For example, they can represent x-components (i.e., horizontal components) and y-components (i.e., vertical components).
[0089] Figure 11 illustrates a signaling method for adaptive motion vector resolution information according to one embodiment of the present invention. According to one embodiment of the present invention, the resolution for indicating MV or MVD can be diverse. For example, the resolution may be expressed based on pixels (or pel). For example, MV or MVD may be signaled in units such as 1 / 4 (quarter), 1 / 2 (half), 1 (integer), 2, 4 pixels, etc. The encoder can then signal the resolution information of MV or MVD to the decoder. Also, for example, 16 may be coded as 64 when it is in 1 / 4 units (1 / 4 * 64 = 16), as 16 when it is in 1 unit (1 * 16 = 16), and as 4 when it is in 4 units (4 * 0.4 = 16). That is, the MV or MVD value may be determined by the following mathematical formula 3.
[0090]
number
[0091] In mathematical formula 3, valueDetermined represents the MV or MVD value. ValuePerResolution represents the value signaled based on the determined resolution. If the value signaled by MV or MVD is not divisible by the determined resolution, a rounding process or the like may be applied. Using a higher resolution can improve accuracy, but may use more bits because the encoded value is larger. Using a lower resolution may result in lower accuracy, but may require fewer bits because the encoded value is smaller. As one example, the resolution described above may be set individually in units such as sequence, picture, slice, coding tree unit (CTU), and coding unit (CU). That is, the encoder / decoder can adaptively determine / apply the resolution according to a predefined unit from among the above units.
[0092] According to one embodiment of this specification, the resolution information described above may be signaled from the encoder to the decoder. In this case, the resolution information may be binary-coded and signaled based on the variable length described above. In such a case, signaling overhead can be reduced if the signaling is based on the index corresponding to the smallest value (i.e., the earliest value). In one embodiment, the signaling index may be mapped in order from high resolution to low resolution.
[0093] According to one embodiment of this specification, Figure 11 illustrates a signaling method assuming that three resolutions are used from among several various resolutions. In this case, the three signaling bits may be 0, 10, and 11, and the three signaling indices can represent the first resolution, second resolution, and third resolution, respectively. Since one bit is required to signal the first resolution and two bits are required to signal the remaining resolutions, the signaling overhead can be relatively reduced when signaling the first resolution. In the example in Figure 11, the first resolution, second resolution, and third resolution may be defined as 1 / 4 and 1.4 pixel resolutions, respectively. In the following embodiments, MV resolution can mean the resolution of MVD.
[0094] Figure 12 illustrates a history-based motion vector prediction (HMVP) method according to one embodiment of the present invention. As mentioned above, the encoder / decoder can use spatial candidates, temporal candidates, etc., as motion vector candidates, and in one embodiment of the present invention, a history-based motion vector, i.e., HMVP, can be further used as a motion vector candidate.
[0095] According to one embodiment of this specification, an encoder / decoder can store motion information of previously coded blocks in a table. In this specification, HMVP represents motion information of previously coded blocks. That is, an encoder / decoder can store HMVP in a table. In this specification, the table that stores the HMVP is referred to as a table or HMVP table, but the present invention is not limited to such names. For example, the table (or HMVP table) may be referred to as a buffer, HMVP buffer, HMVP candidate buffer, HMVP list, HMVP candidate list, etc.
[0096] Motion information stored in the HMVP table may include at least one of the following: MV, reference list, reference index, or utilization flag. For example, motion information may include at least one of the following: MV of reference list L0, MV of L1, L0 reference index, L1 reference index, L0 prediction list utilization flag, or L1 prediction list utilization flag. In this case, the prediction list utilization flag can indicate whether the information is available for use with that list, whether it is significant information, etc.
[0097] Furthermore, according to one embodiment of this specification, the motion information stored in the HMVP table may be generated / stored in a history-based manner. History-based motion information represents the motion information of blocks coded before the current block in the coding order. For example, an encoder / decoder can store the motion information of blocks coded before the current block in the HMVP table. In this case, the block may be a coding unit (CU) or a prediction unit (PU), etc. The motion information of the block can mean the motion information used for motion compensation of the block or candidate motion information used for motion compensation. The motion information stored in the HMVP table (i.e., HMVP) may be used for motion compensation of blocks that are later encoded / decoded. For example, the motion information stored in the HMVP table may be used for motion candidate list construction.
[0098] Referring to Figure 12, the encoder / decoder can take one or more HMVP candidates from the HMVP table (S1201). The encoder / decoder can then add the HMVP candidates to the motion candidate list. For example, the motion candidate list may be a merge candidate list (or merge list) or an AMVP candidate list (or AMVP list). The encoder / decoder can then perform motion compensation and encoding / decoding for the current block based on the motion candidate list (S1202). The encoder / decoder can then update the HMVP table using the information used for motion compensation or decoding of the current block (S1203).
[0099] According to one embodiment of the present invention, HMVP candidates may be used in the merge candidate list construction process. The most recent HMVP candidates in the HMVP table may be reviewed in order and inserted (or added) into the merge candidate list in the order of temporal motion vector prediction (or predictor, TMVP) candidates. When adding HMVP candidates, a pruning process (or pruning check) may be performed on the spatial or temporal merge candidates included in the merge list, excluding sub-block motion candidates (i.e., ATMVPs). In one embodiment, the following embodiment may be applied to reduce the number of pruning operations.
[0100] 1) For example, the number of HMPV candidates may be set as shown in the following mathematical formula 4.
[0101]
number
[0102] In formula 4, L represents the number of HMPV candidates, N represents the number of available non-sub-block merge candidates, and M represents the number of HMVP candidates available in the table. For example, N can represent the number of non-sub-block merge candidates included in the merge candidate list. If N is less than or equal to 4, the number of HMVP candidates may be determined to be M; otherwise, the number of HMVP candidates may be determined to be (8-N).
[0103] 2) Also, for example, the process of constructing the merge candidate list from the HMVP list may be terminated when the total number of available merge candidates reaches the number of signaled maximum allowable merge candidates minus 1.
[0104] 3) Also, for example, the number of candidate pairs for inducing a combined bi-predictive merge candidate may be reduced from 12 to 6.
[0105] Furthermore, according to one embodiment of the present invention, HMVP candidates may also be used in the AMVP candidate list construction process. In the table, the motion vector of the HMVP candidate having the last K index may be inserted after the TMVP candidates. In one embodiment, only HMVP candidates having the same reference picture as the AMVP target reference picture may be used to construct the AMVP candidate list. In this case as well, the pruning process described above may be applied to the HMVP candidates. For example, K may be set to 4 and the size (or length) of the AMVP list may be set to 2.
[0106] Figure 13 illustrates an example of an HMVP table update method according to one embodiment of the present invention. According to one embodiment of the present invention, the HMVP table may be maintained / managed in a FIFO (first-in, first-out) manner. That is, when there is a new input, the most recent element (or candidate) may be output first. For example, when adding motion information used in the current block to the HMVP table, if the HMVP table has filled up to the maximum number of elements, the encoder / decoder can output the most recently added motion information from the HMVP table and add the motion information used in the current block to the HMVP table. At this time, if motion information that already exists in the HMVP table is output, the encoder / decoder can move the motion information at the next position (or index) while filling the output position. For example, if the HMVP table index of the output motion information is m, and the most recent element is located at HMVP table index 0, the motion information corresponding to an index n greater than m can each be moved to the HMVP table index (n-1) position. This leaves high-index positions empty in the HMVP table, allowing the encoder / decoder to assign the highest index in the HMVP table to the motion information used in the current block. In other words, the motion information for the current block may be inserted at index (M+1) when the highest index containing a valid element is M.
[0107] Furthermore, according to one embodiment of the present invention, when the encoder / decoder updates the HMVP table based on specific motion information, a pruning process can be applied. That is, the encoder / decoder can check whether the specific motion information or information corresponding to the specific motion information is included in the HMVP table. Then, the HMVP table update method can be defined separately for cases where duplicate motion information is included and cases where it is not. This prevents the HMVP table from containing duplicate motion information and allows various candidates to be considered for motion compensation.
[0108] In one embodiment, when the encoder / decoder updates the HMVP table based on the motion information currently used in the block, if the motion information is already included in the HMVP table, it can delete the duplicate motion information already included in the HMVP table and add the motion information to the HMVP table as a new entry. At this time, the method described in FIFO can be used before deleting existing candidates and adding new ones. That is, the index of the motion information to be added to the HMVP table is determined to be m, and the oldest motion information may be output. If the motion information is not already included in the HMVP table, the encoder / decoder can delete the oldest added motion information and add the motion information to the HMVP table.
[0109] In other embodiments, when updating the HMVP table based on motion information currently used in a block, the encoder / decoder can update the HMVP table in a FIFO manner if the motion information is already included in the HMVP table, or if it is not already included.
[0110] Furthermore, according to one embodiment of the present invention, the encoder / decoder can initialize (or reset) the HMVP table at a predetermined time or position. Since the encoder and decoder must use the same motion candidate list, they must use the same HMVP table. In this case, if the HMVP table is used continuously without initialization, a dependency problem arises between coding blocks. Therefore, in order to support parallel processing, it is necessary to reduce the dependency between blocks, and the operation of initializing the HMVP table by the unit that supports parallel processing may be pre-configured. For example, the encoder / decoder can be configured to initialize the HMVP table at the slice level, CTU row level, CTU level, etc. For example, if initialization of the HMVP table is defined at the CTU row level, the encoder / decoder can perform encoding / decoding with an empty HMVP table when it starts coding for each CTU row.
[0111] Figure 14 is a diagram illustrating a method for updating an HMVP table according to one embodiment of the present invention. Referring to Figure 14, in one embodiment of the present invention, the encoder / decoder can update HMVPCandList based on mvCand. In this specification, HMVPCandList represents an HMVP table, and mvCand represents the motion information of the current block. The process shown in Figure 14 can take as input a motion candidate mvCand having two motion vectors mvL0 and mvL1, two reference indices refIdxL0 and refIdxL1, and two prediction list utilization flags predFlagL0 and predFlagL1. The process shown in Figure 14 can then output a modified HMVPCandList array.
[0112] In the first stage (Step 1), the encoder / decoder can check whether mvCand is the same as HMVPCandList[HMVPIdx] by changing the HMVPIdx value, a variable that indicates the index of an HMVP in the HMVP table, from 0 to (HMVPCandNum-1). Here, HMVPCandNum indicates the number of HMVPs included in the HMVP table, and HMVPCandList[HMVPIdx] indicates the candidates in the HMVP table that have the HMVPIdx value. If mvCand is the same as HMVPCandList[HMVPIdx], the encoder / decoder can set the variable sameCand, which indicates whether the candidates are the same, to true. In one embodiment, when checking whether mvCand is the same as HMVPCandList[HMVPIdx], the encoder / decoder can perform a comparison with motion information where the prediction list utilization flag is 1, among the MV and reference index associated with L0, or the MV and reference index associated with L1.
[0113] Then, in the second step (Step 2), the encoder / decoder can set the variable tempIdx, which indicates the temporary index, to HMVPCandNum. In the third step (Step 3), if sameCand is true or if HMVPCandNum is equal to the size (or length) of the maximum HMVP table, the encoder / decoder can copy HMVPCandList[tempIdx] to HMVPCandList[tempIdx-1] while varying tempIdx from (sameCand?HMVPIdx:1) to (HMVPCandNum-1). That is, if sameCand is true, tempIdx starts from HMVPIdx, and the encoder / decoder can set HMVPIdx to a value that is the same element index as mvCand in the HMVP table plus 1. If sameCand is false, tempIdx starts from 1, and HMVPCandList[0] may be written with the contents of HMVPCandList[1].
[0114] In the fourth step (Step 4), the encoder / decoder can copy the mvCand, which is the motion information to be updated, to HMVPCandList[tempIdx]. In the fifth step (Step 5), if HMVPCandNum is smaller than the maximum HMVP table size, the encoder / decoder can increment HMVPCandNum by 1.
[0115] Figure 15 is a diagram illustrating a method for updating an HMVP table according to one embodiment of the present invention. Referring to Figure 15, in one embodiment of the present invention, the encoder / decoder can start comparing HMVPIdx from values greater than 0 in the process of checking whether the motion information of the current block is included in the HMVP table. For example, when checking whether the motion information of the current block is included in the HMVP table, the encoder / decoder can compare except for those corresponding to HMVPIdx0. In other words, the encoder / decoder can compare mvCand starting from the candidates corresponding to HMVPIdx1. The process shown in Figure 15 can take as input a motion candidate mvCand having two motion vectors mvL0 and mvL1, two reference indices refIdxL0 and refIdxL1, and two prediction list usage flags predFlagL0 and predFlagL1. The process shown in Figure 15 can then output a modified HMVPCandList array.
[0116] To further explain the HMVP table update method described in Figure 14 above, in the first step, it is possible to check whether there is the same motion information as mvCand in HMVPIdx0. However, in the update method described above, if mvCand exists in HMVPCandList[0], or if it does not exist in HMVPCandList[HMVPIdx] corresponding to HMVPCandNum from HMVPIdx0, the motion information of HMVPCandList[0] is output. If mvCand exists in HMVPCandList[HMVPIdx] among HMVPIdx where mvCand is not 0, the contents of HMVPCandList[0] are not updated, so it is not necessary to compare mvCand with HMVPCandList[0]. Also, since the motion information of the current block is similar to the motion information of spatially adjacent blocks, it is highly likely to be similar to the motion information recently added in the HMVP table. Under these assumptions, mvCand's motion information may be more similar to that of candidates with HMVPIdx greater than 0 than that of candidates with HMVPIdx 0. If similar (or identical) motion information is found, the pruning process can be terminated. Therefore, by performing the first step in Figure 14 above starting with HMVPIdx values greater than 0, the number of comparisons can be reduced compared to the embodiment in Figure 14 described above.
[0117] Figure 16 is a diagram illustrating a method for updating an HMVP table according to one embodiment of the present invention. Referring to Figure 16, in one embodiment of the present invention, the encoder / decoder can compare HMVPIdx values from those greater than 0 in the process of checking whether the motion information of the current block, i.e., mvCand, is included in the HMVP table. The process shown in Figure 16 can take as input a motion candidate mvCand having two motion vectors mvL0 and mvL1, two reference indices refIdxL0 and refIdxL1, and two prediction list usage flags predFlagL0 and predFlagL1. The process shown in Figure 16 can then output a modified HMVPCandList array. For example, when the encoder / decoder checks whether mvCand is included in the HMVP table, it can perform the comparison process (or pruning process) while excluding a predefined number of motion information (or candidates) whose HMVPIdx values are 0 or greater. Specifically, when the variable indicating the number of elements to be compared is represented by NumPrune, the encoder / decoder can check whether mvCand matches the HMVP table element corresponding to HMVPIdx(HMVPCandNum-NumPrune+1) to (HMVPCandNum-1). This significantly reduces the number of comparisons compared to the previously described embodiment.
[0118] In other embodiments, the encoder / decoder can compare mvCand with a specific, already set position in the HMVP table that is not HMVPIdx0. For example, the encoder / decoder can compare mvCand with HMVP table elements corresponding to HMVPIdx PruneStart to (HMVPCandNum-1) to determine whether the motion information overlaps.
[0119] Figure 17 illustrates a method for updating an HMVP table according to one embodiment of the present invention. Referring to Figure 17, in one embodiment of the present invention, the encoder / decoder can compare mvCand with HMVP table elements in order from the most recently added HMVP table elements to the oldest HMVP table elements in the process of checking whether mvCand is included in the HMVP table. The process shown in Figure 17 can take as input a motion candidate mvCand having two motion vectors mvL0 and mvL1, two reference indices refIdxL0 and refIdxL1, and two prediction list utilization flags predFlagL0 and predFlagL1. The process shown in Figure 17 can then output a modified HMVPCandList array. Since mvCand is likely to be similar to the motion information of spatially adjacent blocks, it may be similar to motion information that has been added relatively recently in the HMVP table. In this manner, if similar (or identical) motion information is found, the pruning process can be terminated. Therefore, as shown in the example in Figure 17, the number of comparisons can be reduced by performing the comparison starting with the HMVPIdx corresponding to the most recently added element.
[0120] Figure 18 illustrates a pruning process according to one embodiment of the present invention. According to one embodiment of the present invention, when an encoder / decoder checks whether the motion information of the current block is already included in the HMVP table, it can compare it with information from a portion of the HMVP table. This is to reduce the complexity of the comparison process. For example, when checking whether motion information is included in HMVPCandList[HMVPIdx], the encoder / decoder can perform the pruning process using a subset of HMVPCandList[HMVPIdx] or an HMVPCandList[HMVPIdx] with a predefined index. For example, HMVPCandList[HMVPIdx] may contain information on L0 and L1, but if both the L0 and L1 utilization flags are 1, the encoder / decoder can select and compare either L0 or L1 according to a pre-established agreement (or condition). For example, a pruning process can be performed to check for duplicate motion information only for the smaller of the L0 and L1 reference index reference lists.
[0121] In one embodiment, the encoder / decoder can perform the pruning process only on the smaller of the L0 and L1 reference indices of the motion information in the HMVP table. For example, the encoder / decoder can perform the pruning process only on the smaller of the L0 and L1 reference indices of the motion information in the current block. In another embodiment, the encoder / decoder can determine whether a value based on a subset of the reference index bits and motion information bits of the current block is already included in the HMVP table by comparing it with a value based on a subset of the reference index bits and motion information bits of the HMVP table element. In one embodiment, the value based on a subset of the reference index bits and motion information bits may be the reference index bits and some of the motion information bits. In another embodiment, the value based on a subset of the reference index bits and motion information bits may be the value (or output) obtained by passing the first subset and the second subset through a hash function. In another embodiment, if the size difference between the two motion vectors is smaller than or equal to a predefined critical value during the pruning process, the encoder / decoder can determine that the two motion vectors are identical or similar.
[0122] Figure 19 illustrates a method for adding HMVP candidates according to one embodiment of the present invention. According to one embodiment of the present invention, an encoder / decoder can add MVs included in the HMVP table to the motion candidate list. In one embodiment, since the motion of the current block may be similar to a recently added motion, and the recently added motion may be useful as a candidate, the encoder / decoder can add the recently added element in the HMVP table to the candidate list.
[0123] In one embodiment, referring to Figure 19, the encoder / decoder can add elements to the candidate list starting from elements with a specific index value, rather than the most recently added element in the HMVP table. Alternatively, the encoder / decoder may add elements to the candidate list starting from the next element, rather than the most recently added element in the HMVP table, up to a specific number of elements added previously.
[0124] Alternatively, in one embodiment, the encoder / decoder may add a specific number of previously added elements to the candidate list, excluding one or more recently added HMVP table elements. In this case, the encoder / decoder can preferentially add relatively recently input motion information from the HMVP table elements to the candidate list.
[0125] The most recently added HMVP table element may correspond to a block that is spatially adjacent to the current block, and therefore is likely to have already been added as a spatial candidate during the candidate list construction process. Thus, by constructing the candidate list excluding a specific number of recently added candidates, as in this embodiment, it is possible to prevent the addition of unnecessary candidates to the candidate list and reduce the complexity of the pruning processor when adding candidates to the list. While the HMVP table element and its use have been described primarily in the context of motion candidate list construction, the present invention is not limited thereto and is applicable to other parameters or inter / intra prediction-related information.
[0126] Figure 20 illustrates a merge sharing node according to one embodiment of the present invention. According to one embodiment of the present invention, an encoder / decoder may share the same candidate list for multiple blocks in order to facilitate parallel processing. The candidate list may be a motion candidate list. The multiple blocks may be defined according to a previously established agreement. For example, the multiple blocks may be defined as blocks that are subordinate to a block (or region) that satisfies a specific condition. Alternatively, the multiple blocks may be defined as blocks that are included in a block that satisfies a specific condition. In this specification, a candidate list used identically by the multiple blocks may be called a shared list. If applied to a merge candidate list, it may be called a shared merge list.
[0127] Furthermore, in this specification, blocks that satisfy the above-mentioned specific conditions may be referred to as merge sharing nodes (i.e., the portion shown by the dotted line in Figure 20), shared merge nodes, merge sharing areas, shared merge areas, shared merge list nodes, shared merge list areas, etc. The encoder / decoder can construct a motion candidate list based on the motion information of surrounding blocks adjacent to the merge sharing node, thereby ensuring parallel processing for coding units within the merge sharing node. In addition, critical values can be used as the above-mentioned specific conditions. For example, the block defined as the merge sharing node may be defined (or determined) based on the critical value. As an example, the critical value used in the above-mentioned specific conditions may be set to a value related to the block size, block width / height.
[0128] Figure 21 is a diagram illustrating an HMVP update method when a shared list according to one embodiment of the present invention is used. Figure 21 assumes that the shared list described in Figure 20 is used. When a shared list is constructed using HMVP candidates, the encoder and decoder must maintain the same HMVP table. If multiple blocks using the same shared list do not have HMVP table update rules defined based on the motion information used, they cannot construct the same candidate list. Therefore, Figures 22 to 24, described later, assume that CU1, CU2, CU3, and CU4 use a shared list within the merged shared node shown in Figure 21, and explain an HMVP table update method to ensure the same shared list within the merged shared node.
[0129] Figure 22 illustrates a method for updating the HMVP table based on motion information of blocks in a merged shared node according to one embodiment of the present invention. Referring to Figure 22, the encoder / decoder can apply the same shared list to CU1, CU2, CU3, and CU4 in the lower node coding unit of the merged shared node. Specifically, the encoder / decoder can construct the same motion candidate list for CU1, CU2, CU3, and CU4 based on the motion information of surrounding blocks adjacent to the merged shared node (S2201, S2202, S2203, S2204). The motion candidate list may include motion vector components and reference indices.
[0130] After guiding motion information for CU1, CU2, CU3, and CU4 using a shared motion candidate list, or encoding / decoding, the encoder / decoder can update the HMVP table using the motion information for CU1, CU2, CU3, and CU4 (S2205). In situations where the same candidate list should be used, i.e., when multiple CUs in a merged shared node update the HMVP table, it may not be possible to use the same candidate list. Therefore, according to one embodiment of the present invention, in situations where a shared list is used, the update of the HMVP table may be performed after the MV guidance or coding of all CUs belonging to the merged shared node has been completed, thereby enabling all CUs belonging to the merged shared node to construct their motion candidate lists based on the same HMVP table.
[0131] Figure 23 is a diagram illustrating the HMVP update method when a shared list according to one embodiment of the present invention is used. Referring to Figure 23, we assume that the shared list described in Figures 20 to 22 above is used. In this case, when the MV induction or encoding / decoding of each CU in Figure 22 is completed, since there are multiple CUs in the merged shared node, there may also be multiple motion information used.
[0132] According to one embodiment of the present invention, as shown in Figure 23, an encoder / decoder can update the HMVP table using all motion information within a merged shared node. In other words, the encoder / decoder can update the HMVP table using motion information used by all CUs using the same shared list. The update order must be already set for the encoder and decoder.
[0133] In one embodiment, the encoder / decoder may refer to the regular decoding order as the update order. For example, as shown in Figure 23, the motion information corresponding to each CU can be sequentially used as input to the HMVP table update process according to the regular decoding order. Alternatively, in one embodiment, the encoder / decoder may determine the HMVP table update order for the motion information of multiple CUs by referring to the reference index order, the POC relationship between the current picture and the reference picture, etc. For example, all motion information within a merged shared node can be updated to the HMVP table in the reference index order. Alternatively, for example, all motion information within a merged shared node can be updated to the HMVP table in order of the lowest or highest POC difference between the current picture and the reference picture.
[0134] Figure 24 illustrates a method for updating the HMVP table based on the motion information of blocks within a merged shared node according to one embodiment of the present invention. Referring to Figure 24, we assume that the shared list described in Figures 20 to 22 above is used. In this case, once the MV induction or encoding / decoding of each CU in Figure 22 is completed, since there are multiple CUs within the merged shared node, there may also be multiple pieces of motion information used.
[0135] According to one embodiment of the present invention, an encoder / decoder can update the HMVP table using some of the motion information used by multiple CUs within a merged shared node. In other words, the encoder / decoder does not need to update the HMVP table with motion information from at least some of the CUs within a merged shared node. As an example, the encoder / decoder can refer to the regular coding order as the update order. For example, the encoder / decoder can update the HMVP table using a set number of motion information entries that are late in the decoding order among the motion information corresponding to each CU according to the regular decoding order. This is because late blocks in the regular coding order are likely to be spatially adjacent to the next block to be coded, and similar motion may be required during motion compensation.
[0136] In one embodiment, referring to Figure 24, when there are CU1 to CU4 using the same shared list, the encoder / decoder can use only the motion information for CU4, which is coded last in the normal decoding order, to update the HMVP table. In another embodiment, the encoder / decoder can update the HMVP table using some of the motion information from multiple CUs by referring to the reference index order, the POC relationship between the current picture and the reference picture, etc.
[0137] Furthermore, according to another embodiment of the present invention, when adding an HMVP candidate from an HMVP table to a candidate list, the encoder / decoder can add the HMVP candidate to the candidate list by referring to the relationship between the current block and the block corresponding to the element in the HMVP table. Alternatively, when adding an HMVP candidate from an HMVP table to a candidate list, the encoder / decoder can refer to the positional relationships between candidate blocks included in the HMVP table. For example, the encoder / decoder can add an HMVP candidate from an HMVP table to a candidate list by considering the decoding order of the HMVP or blocks.
[0138] Figure 25 illustrates a method for processing a video signal based on an HMVP according to one embodiment of the present invention. Referring to Figure 25, the explanation will focus on the decoder for convenience, but the present invention is not limited thereto, and the HMVP-based video signal processing method according to this embodiment is substantially applicable to encoders as well.
[0139] Specifically, if the current block is located within a merge sharing node containing multiple coding blocks, the decoder can construct a merge candidate list using spatial candidates adjacent to the merge sharing node (S2501). The decoder can add a specific HMVP from an HMVP table containing at least one history-based motion vector predictor (HMVP) to the merge candidate list (S2502). Here, the HMVP represents motion information of blocks coded prior to the multiple coding blocks.
[0140] The decoder obtains index information indicating the merge candidate to be used to predict the current block within the merge candidate list (S2503), and generates a predicted block for the current block based on the motion information of the merge candidate (S2504). The decoder can generate a restored block for the current block by adding the predicted block and the residual block. As described above, according to one embodiment of the present invention, the motion information of at least one coding block among the multiple coding blocks included in the merge shared node does not need to be updated in the HMVP table. As described above, the decoder can update the HMVP table using the motion information of a predefined number of coding blocks among the multiple coding blocks included in the merge shared node that have a relatively slow decoding order.
[0141] Furthermore, as described above, the decoder can update the HMVP table using the motion information of the coding block with the relatively latest decoding order among the multiple coding blocks included in the merge shared node. Also, if the current block is not located within the merge shared node, the decoder can update the HMVP table using the motion information of the merge candidate. Furthermore, as described above, the decoder can use an HMVP with a predefined specific index in the HMVP table to check whether it has motion information that overlaps with a candidate in the merge candidate list.
[0142] Figure 26 is a diagram illustrating a multi-hypothesis prediction method according to one embodiment of the present invention. According to one embodiment of the present invention, an encoder / decoder can generate prediction blocks based on multiple prediction methods. In this specification, a prediction method based on such multiple prediction modes is called multi-hypothesis prediction. However, the present invention is not limited to this name, and in this specification, the multi-hypothesis prediction may be called multiple prediction, multiple prediction, combined prediction, inter-intra weighted prediction, combined inter-intra prediction, combined inter-intra weighted prediction, etc. In one embodiment, multi-hypothesis prediction can mean blocks generated by any prediction method. Also, in one embodiment, the prediction method in multi-hypothesis prediction may include methods such as intra prediction and inter prediction. Alternatively, the prediction method in multi-hypothesis prediction may be further subdivided to mean merge mode, AMVP mode, specific mode of intra prediction, etc. Furthermore, the encoder / decoder can generate a final prediction block by weighted summing the prediction blocks (or prediction samples) generated based on multi-hypothesis prediction.
[0143] According to one embodiment of the present invention, the maximum number of prediction methods used in multiple prediction may be set in advance. For example, the maximum number of multiple predictions may be 2. Therefore, the encoder / decoder can generate a prediction block by applying two predictions in the case of uni-prediction, or two predictions (i.e., when multiple prediction is used only for predictions from one reference list) or four predictions (i.e., when multiple prediction is used for predictions from two reference lists) in the case of bi-prediction.
[0144] Alternatively, according to one embodiment of the present invention, prediction modes usable in multiple assumption prediction may be pre-set. Or, combinations of prediction modes usable in multiple assumption prediction may be pre-set. For example, the encoder / decoder can perform multiple assumption prediction using prediction blocks (or prediction samples) generated by inter-prediction and intra-prediction.
[0145] According to one embodiment of the present invention, an encoder / decoder can use only some of the prediction modes among the inter-prediction and / or intra-prediction modes for multiple assumption prediction. For example, the encoder / decoder can use only the merge mode among the inter-prediction modes for multiple assumption prediction. Alternatively, the encoder / decoder can use the merge mode, rather than the subblock merge mode, among the inter-prediction modes for multiple assumption prediction. Alternatively, the encoder / decoder can use a specific intra-prediction mode among the intra-prediction modes for multiple assumption prediction. For example, the encoder / decoder can restrictively use prediction modes among the intra-prediction modes that include at least one of the planar, DC, vertical, and / or horizontal modes for multiple assumption prediction. In one embodiment, the encoder / decoder can generate prediction blocks based on the predictions of the merge mode and intra-prediction, in which case only at least one restricted prediction mode among the planar, DC, vertical, and / or horizontal modes can be used for intra-prediction.
[0146] Referring to Figure 26, the encoder / decoder can generate prediction blocks using the first prediction (prediction 1) and the second prediction (prediction 2). Specifically, the encoder / decoder can generate a first temporary prediction block (or prediction sample) by applying the first prediction, and generate a second temporary prediction block by applying the second prediction. The encoder / decoder can generate the final prediction block by weighting the first temporary prediction block and the second temporary prediction block. At this time, the encoder / decoder can perform a weighted sum by applying the first weight value w1 to the first temporary prediction block generated by the first prediction, and the second weight value w2 to the second temporary prediction block generated by the second prediction.
[0147] According to one embodiment of the present invention, when generating a prediction block based on multiple assumption predictions, the weights applied to the multiple assumption predictions may be determined based on a specific position within the block. In this case, the block may be the current block or a surrounding block. Alternatively, the weights of the multiple assumption predictions may be based on the mode in which the predictions are generated. For example, if one of the modes in which the predictions are generated is intra-prediction, the encoder / decoder can determine the weights based on the prediction mode. Also, for example, if one of the prediction modes is intra-prediction and is a directional mode, the encoder / decoder can increase the weights for samples located far from the reference sample.
[0148] According to one embodiment of the present invention, when the intra-prediction mode used in multiple assumption prediction is a directional mode and the other prediction mode is inter-prediction, the encoder / decoder can apply a relatively high weight to the predicted sample generated based on the intra-prediction on the side farther from the reference sample. This is because, in the case of inter-prediction, motion compensation can be performed using spatial neighbor candidates, and in such cases, there is a high probability that the motion of the current block and the spatial neighbor block referenced for motion compensation are the same or similar, and as a result, the prediction of the region adjacent to the spatial neighbor block and the prediction of the region containing the object with motion are more likely to be accurate than other parts. In this case, the residual signal adjacent to the opposite directional boundary of the spatial neighbor block may be greater than in other regions (or parts). According to one embodiment of the present invention, this can be canceled out by combining and applying the intra-predicted samples in multiple assumption prediction. Also, in one embodiment, since the reference sample position of the intra-prediction may be near the spatial neighbor candidate of the inter-prediction, the encoder / decoder can apply a high weight to the region relatively farther from it.
[0149] In another embodiment, if one of the modes used to generate multiple-assumption prediction samples is intra-prediction and is in directional mode, the encoder / decoder can apply a higher weight to samples that are relatively close to the reference sample. More specifically, if one of the modes used to generate multiple-assumption prediction samples is intra-prediction and is in directional mode, and the other mode used to generate multiple-assumption prediction samples is inter-prediction, the encoder / decoder can apply a higher weight to the prediction generated based on the intra-prediction that is closer to the reference sample. This is because, in intra-prediction, the closer the distance between the prediction sample and the reference sample, the higher the accuracy of the prediction.
[0150] In another embodiment, if one of the modes used to generate the multiple assumption prediction sample is intra-prediction and not a directional mode (e.g., planar, DC mode), the weighting value may be set to a constant value regardless of the position within the block. Also, in one embodiment, the weighting value for prediction 2 in multiple assumption prediction may be determined based on the weighting value for prediction 1. The following equation represents an example of determining a prediction sample based on multiple assumption prediction.
[0151]
number
[0152] In equation 5, pbSamples represents the (final) predicted samples (or predicted blocks) generated by multiple assumption prediction. predSamples represents the blocks / samples generated by interpretation, and predSamplesIntra represents the blocks / samples generated by intrapretation. In equation 5, x and y represent the coordinates of the sample within the block and may be within the following ranges: x = 0..nCbW-1 and y = 0..nCbH-1. nCbW and nCbH may be the current width and height of the block, respectively. In one embodiment, the weighted value w may be determined by the following process.
[0153] - If predModeIntra is INTRA_PLANAR or INTRA_DC, nCbW<4, nCbH<4, or cIdx>0, then w may be set to 4.
[0154] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and y < (nCbH / 4), then w may be set to 6.
[0155] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and (nCbH / 4) <= y < (nCbH / 2), w may be set to 5.
[0156] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and (nCbH / 2) <= y < (3*nCbH / 4), w may be set to 4.
[0157] - Otherwise, if predModeIntra is INTRA_ANGULAR50 and (3*nCbH / 4) <= y < nCbH, w may be set to 3.
[0158] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and x < (nCbW / 4), w may be set to 6.
[0159] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and (nCbW / 4) <= x < (nCbW / 2), w may be set to 5.
[0160] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and (nCbW / 2) <= x < (3*nCbW / 4), w may be set to 4.
[0161] - Otherwise, if predModeIntra is INTRA_ANGULAR18 and (3*nCbW / 4) <= x < nCbW, w may be set to 3.
[0162] The following Table 2 illustrates a multiple hypothesis prediction related syntax structure according to an embodiment of the present invention.
[0163]
Table 2
[0164] In Table 2, mh_intra_flag is a flag indicating whether or not to use multiple assumption prediction. According to one embodiment of the present invention, multiple assumption prediction may be applied only if certain predefined conditions for multiple assumption prediction (referred to herein as mh_conditions for convenience of explanation) are met. If mh_conditions are not met, the encoder / decoder may not parse mh_intra_flag and infer it to 0. For example, mh_conditions may include conditions relating to block size. Also, mh_conditions may include conditions relating to whether or not to use a certain predefined mode. For example, mh_intra_flag may be parsed if merge_flag, a flag indicating whether or not merge mode is applied, is 1, and subblock_merge_flag, a flag indicating whether or not subblock merge mode is applied, is 0. In other words, the encoder / decoder may consider (or apply) multiple assumption prediction if merge mode is currently applied to the block and subblock merge mode is not applied.
[0165] Furthermore, according to one embodiment of the present invention, the encoder can configure candidate modes by dividing them into multiple lists in order to determine the mode in multiple assumption prediction, and can signal to the decoder which list to use. Referring to Table 2, mh_intra_luma_mpm_flag may be a flag indicating which of the multiple lists to use. If mh_intra_luma_mpm_flag does not exist, it can be inferred (or considered) to be 1. Also, in one embodiment of the present invention, the multiple lists may be an MPM list and a non-MPM list.
[0166] In one embodiment, the encoder can signal to the decoder an index (or index information) indicating which index candidate to use from among the multiple lists. Referring to Table 2, mh_intra_luma_mpm_idx may be the index described above. In another embodiment, the index may be signaled only when a specific list is selected. The decoder can parse mh_intra_luma_mpm_idx only when a specific list is determined by mh_intra_luma_mpm_flag.
[0167] According to one embodiment of the present invention, as shown in the embodiment described in Figure 26, multiple assumption predictions can be made based on predictions generated by inter-prediction and predictions generated by intra-prediction. For example, the encoder / decoder can perform multiple assumption predictions only when signaled to use inter-prediction. Alternatively, the encoder / decoder can perform multiple assumption predictions only when signaled to use a specific mode of inter-prediction, such as merge mode. In this case, separate signaling for inter-prediction is not required. Also, in one embodiment, when the encoder / decoder generates predictions with intra-prediction, there may be a total of four candidate modes. For example, a first list and a second list can be constructed using three and one candidate modes respectively from the total of four candidate modes. In this case, if the second list containing one prediction mode is selected, the encoder does not need to signal an index to the decoder. Also, if the first list is selected, an index indicating a specific candidate can be signaled to the decoder. In this case, since there are three candidates in the first list, signaling can be performed using variable length coding with 1 or 2 bits.
[0168] Table 3 illustrates a multiple-assumption prediction-related syntax structure according to one embodiment of the present invention.
[0169] [Table 3]
[0170] In Table 3, as explained in Table 2 above, there can be signaling to indicate which of several lists to use, and in Tables 2 and 3, mh_intra_luma_mpm_flag may be such a syntax element. Syntax elements that overlap with those in Table 2 mentioned above will not be explained.
[0171] According to one embodiment of the present invention, signaling indicating which list to use may be explicitly signaled only in specific cases. If explicit signaling is not provided, the encoder / decoder can infer the value of the syntax element in a pre-configured manner. Referring to Table 3, if the condition mh_mpm_infer_condition is met, explicit signaling may not exist, and if the condition mh_mpm_infer_condition is not met, explicit signaling may exist. Also, if the condition mh_mpm_infer_condition is met, mh_intra_luma_mpm_flag may not exist, in which case it may be inferred to be 1. That is, in this case, the encoder / decoder can infer that the MPM list is to be used.
[0172] Table 4 illustrates a multiple-assumption prediction-related syntax structure according to one embodiment of the present invention.
[0173] [Table 4]
[0174] As explained above in Tables 2 and 3, a signaling mechanism can exist to indicate which of several lists to use, and if predefined conditions are met, the encoder / decoder can infer the value. Syntax elements that overlap with those mentioned in Tables 2 and 3 above will not be explained further.
[0175] According to one embodiment of the present invention, the conditions for inferring a signaling (or syntax element, parameter) value indicating which of a plurality of lists to use may be determined based on the current block size. For example, the encoder / decoder may determine this based on the current block width and height. Specifically, the encoder / decoder can infer a signaling value if the larger of the current block's width and height is greater than n times the smaller of the smaller of the two. For example, n may be set to a natural number such as 2, 3, or 4. Referring to Table 4, the condition for inferring a signaling value indicating which of a plurality of lists to use may be the condition that the larger of the current block's width and height is greater than twice the smaller of the two. If the current block's width and height are cbWidth and cbHeight, respectively, then the Abs(Log2(cbWidth / cbHeight)) value is 0 when cbWidth and cbHeight are the same, and 1 when they differ by a factor of two. Therefore, if the difference between cbWidth and cbHeight is greater than 2, the Abs(Log2(cbWidth / cbHeight)) value will be greater than 1 (i.e., it can have a value of 2 or greater).
[0176] Figure 27 shows a method for determining a multiple assumption prediction mode according to one embodiment of the present invention. As explained in Tables 2 to 4 above, the determination of the prediction mode used for multiple assumption prediction may be based on a plurality of lists. As an example, the prediction mode may represent an intra-mode that generates predictions based on intra-prediction. The plurality of lists may include two lists, a first list and a second list. Referring to Figure 27, it is possible to determine whether or not to use the first list from list1_flag. In one embodiment, there may be multiple candidates that belong to the first list, and there may be one candidate that belongs to the second list.
[0177] If list1_flag is inferred, the encoder / decoder can infer that the first list will be used and its value (S2701). In this case, it can parse list1_index, which is the index indicating which candidate in the first list to use (S2704). Also, if list1_flag is not inferred (S2701), the encoder / decoder can parse list1_flag (S2702). If list1_flag is 1, the encoder / decoder can parse list1_index, but if list1_flag is not 1, it does not need to parse the index. Also, if list1_flag is 1, the encoder / decoder can determine which of the candidate modes in the first list will actually be used based on the index (S2703). Also, if list1_flag is not 1, the encoder / decoder can determine which of the candidate modes in the second list will actually be used without an index. In other words, the mode may be determined based on the flags and index in the first list, and the mode may be determined based on the flags in the second list.
[0178] Figure 28 shows a method for determining multiple hypothetical prediction modes according to one embodiment of the present invention. According to one embodiment of the present invention, when the index for determining candidate modes in a list is coded with variable length, a method for determining the order of modes included in the candidate list may be applied to improve coding efficiency. For example, there may be a method for determining the order of modes included in the first list. In this case, the encoder / decoder can refer to the modes around the current block to determine the mode order. The second list can be determined without referring to the modes around the current block. For example, the encoder / decoder can generate the first list by referring to the modes around the current block and include modes not included in the first list in the second list. In one embodiment, the first list may be MPM modes and the second list may be non-MPM modes. Also, there are four candidate modes in total, and the first list may contain three modes and the second list may contain one mode.
[0179] Referring to Figure 28, there may be a syntax element list1_flag indicating whether or not the first list is used. If the first list is used, the encoder / decoder generates the first list (S2801, S2802) and can select a specific mode from the first list. In this case, the generation of the first list and the confirmation of whether or not the first list is used may be performed in any order. For example, in a situation where the first list is used, the first list may be generated before or after the confirmation of whether or not the first list is used. Also, if the first list is used, the encoder / decoder does not need to perform the process of generating the second list.
[0180] If the first list is not used, the encoder / decoder generates a second list (S2803) and can select a specific mode from the second list. In this case, the encoder / decoder can generate the first list in order to generate the second list. The encoder / decoder can then include candidates that are not included in the first list from among the candidate modes in the second list. Furthermore, according to one embodiment of the present invention, the method for generating the first list may be the same regardless of whether the first list is used or not (list1_flag value), whether the first list is used or not, whether inference is performed or not, etc. In this case, the methods described above in Tables 2 to 4 and Figure 27 may be applied to list signaling and mode signaling.
[0181] The following further describes the method for constructing (or generating) multiple lists for determining the prediction mode used in the multiple assumption prediction described in Figures 27 and 28 above. As an example, as mentioned above, the multiple lists may consist of two lists. That is, the multiple lists may include a first list and a second list. Furthermore, the multiple lists may be used in the multiple assumption prediction process.
[0182] According to one embodiment of the present invention, an encoder / decoder can generate multiple lists by referring to the modes around the current block. The encoder / decoder can also perform intraprediction using the modes selected from the lists and combine the predicted samples (or predicted blocks) generated by the intraprediction with the inter-predicted predicted samples to generate a multiple-hypothesis predicted block. In this specification, the final predicted samples (or predicted blocks) generated by multiple-hypothesis prediction are referred to as a multiple-hypothesis predicted block, but the present invention is not limited thereto. For example, a multiple-hypothesis predicted block may be called a predicted block, a final predicted block, a multiple-prediction block, a combined predicted block, an inter-intra weighted predicted block, a combined inter-intra predicted block, a combined inter-intra weighted predicted block, and so on.
[0183] As one embodiment, the modes (candidate modes) that may be included in the above list may be set to at least one of the planar mode, DC mode, vertical mode, and / or horizontal mode of the intra prediction method. In this case, the vertical mode may be the mode with index (or mode number) 50 in Figure 6 described above, and the horizontal mode may be the mode with index 18 in Figure 6. Also, the planar mode and DC mode may have indices 0 and 1, respectively.
[0184] According to one embodiment of the present invention, a candidate mode list can be generated by referring to the predicted modes of the surrounding blocks of the current block. The candidate mode list may be referred to as candModeList in this specification. For example, the candidate mode list may be the first list described in the above embodiment. In another embodiment, the mode of the surrounding blocks of the current block may be expressed as candIntraPredModeX. That is, candIntraPredModeX represents a variable that indicates the mode of the surrounding block. Here, X represents a variable that indicates a specific location around the current block, such as A, B, etc.
[0185] In one embodiment, an encoder / decoder can generate a candidate mode list based on whether or not there are matches between multiple candIntraPredModeX values. For example, candIntraPredModeX can exist for two positions, and the modes at those positions may be represented as candIntraPredModeA and candIntraPredModeB. As an example, if candIntraPredModeA and candIntraPredModeB are the same, the candidate mode list may include planar mode and DC mode.
[0186] In one embodiment, if candIntraPredModeA and candIntraPredModeB are the same and their values indicate either planar mode or DC mode, the encoder / decoder can add the modes indicated by candIntraPredModeA and candIntraPredModeB to the candidate mode list. The encoder / decoder can also add any mode from planar mode and DC mode that is not indicated by candIntraPredModeA and candIntraPredModeB to the candidate mode list. Furthermore, the encoder / decoder can add a mode that has already been set, rather than planar mode or DC mode, to the candidate mode list. In one embodiment, in this case, the order of planar mode, DC mode, and the previously set specific mode within the candidate mode list may already be set. For example, it may be planar, DC, and the previously set mode order. That is, candModeList[0] = planar mode, candModeList[1] = DC mode, candModeList[2] = the previously set mode. Also, the previously set mode may be vertical mode. In yet another embodiment, the planar mode, DC mode, and the modes indicated by candIntraPredModeA and candIntraPredModeB among the already set modes are initially added to the candidate mode list, the planar mode and DC mode not indicated by candIntraPredModeA and candIntraPredModeB are added as the next candidates, and the already set modes may be added thereafter.
[0187] Furthermore, if candIntraPredModeA and candIntraPredModeB are the same and their values do not indicate planar mode or DC mode, the encoder / decoder can add the modes indicated by candIntraPredModeA and candIntraPredModeB to the candidate mode list. Planar mode and DC mode may also be added to the candidate mode list. In this case, the order of the modes indicated by candIntraPredModeA and candIntraPredModeB, planar mode, and DC mode in the candidate mode list may already be set. The already set order may be the mode indicated by candIntraPredModeA and candIntraPredModeB, planar mode, and DC mode. That is, candModeList[0]=candIntraPredModeA, candModeList[1]=planar mode, and candModeList[2]=DC mode.
[0188] Furthermore, if candIntraPredModeA and candIntraPredModeB are different, the encoder / decoder can add both candIntraPredModeA and candIntraPredModeB to the candidate mode list. Also, candIntraPredModeA and candIntraPredModeB may be included in the candidate mode list according to a predefined specific order. For example, they may be included in the candidate mode list in the order candIntraPredModeA, candIntraPredModeB. There may also be a pre-set order between the candidate modes, and the encoder / decoder can add modes other than candIntraPredModeA and candIntraPredModeB from the modes in the pre-set order to the candidate mode list. Also, the modes other than candIntraPredModeA and candIntraPredModeB may be added in the candidate mode list after candIntraPredModeA and candIntraPredModeB. Furthermore, the pre-set order may be planar mode, DC mode, and vertical mode. Alternatively, the previously set order may be planar mode, DC mode, vertical mode, horizontal mode. That is, candModeList[0]=candIntraPredModeA, candModeList[1]=candIntraPredModeB, and candModeList[2] may be the earliest mode among the planar mode, DC mode, and vertical mode that is neither candIntraPredModeA nor candIntraPredModeB.
[0189] Furthermore, in one embodiment, a mode among the candidate modes that is not included in the candidate mode list may be defined as candIntraPredModeC. For example, candIntraPredModeC may be included in the second list. Also, if the signaling indicating whether or not the aforementioned first list is used indicates that it is not used, candIntraPredModeC can be determined. When the first list is used, the encoder / decoder determines the mode from the candidate mode list by index, and when the first list is not used, it can use the modes from the second list.
[0190] Furthermore, as mentioned above, a process of modifying the candidate mode list may be added after the candidate mode list has been generated. For example, the encoder / decoder may or may not perform the modification process depending on the current block size condition. For example, the current block size condition may be determined based on the width and height of the current block. For example, if the larger of the width and height of the current block is greater than n times the other, the modification process may be performed further. For example, n may be defined as 2.
[0191] Furthermore, in one embodiment, the modification process may be a process of changing a mode to another mode when any mode is included in the candidate mode list. For example, if the candidate mode list includes the vertical mode, the encoder / decoder may insert the horizontal mode into the candidate mode list instead of the vertical mode. Alternatively, if the candidate mode list includes the vertical mode, the encoder / decoder may insert candIntraPredModeC into the candidate mode list instead of the vertical mode. Alternatively, since the planar mode and DC mode can always be included in the candidate mode list when the candidate mode list is generated as described above, in this case candIntraPredModeC may be the horizontal mode. Furthermore, such a modification process may be used when the current block height is greater than n times the width. For example, n may be defined as 2. This is because when the height is greater than the width, the lower side of the block is far from the reference sample for intra-prediction, which can result in lower accuracy for the vertical mode. Alternatively, such a modification process may be used when it is inferred that the first list is being used.
[0192] Furthermore, according to embodiments of the present invention, as another example of the candidate list modification process, the encoder / decoder may, if the candidate mode list includes a horizontal mode, add a vertical mode to the candidate mode list instead of a horizontal mode. Alternatively, if the candidate mode list includes a horizontal mode, candIntraPredModeC may be added to the candidate mode list instead of a horizontal mode. Alternatively, since the planar mode and DC mode may always be included in the candidate mode list when the candidate mode list is generated as described above, in this case candIntraPredModeC may be a vertical mode. Moreover, such a modification process may be used when the width of the block is greater than n times the height. For example, n may be defined as 2. This is because, when the width is greater than the height, the right side of the block is farther from the reference sample for intra-prediction, which can result in lower accuracy for the horizontal mode. Alternatively, such a modification process may be used when it is inferred that the first list is being used.
[0193] The following describes an example of the list setting method described above. In this specification, IntraPredModeY indicates the mode used for intraprediction during multiple assumption prediction. IntraPredModeY can also indicate the mode of the luminance component. In one embodiment, the intraprediction mode of the color difference component in multiple assumption prediction may be derived from the luminance component. In this specification, mh_intra_luma_mpm_flag represents a variable (or syntax element) that indicates which of the multiple lists to use. For example, mh_intra_luma_mpm_flag may be mh_intra_luma_mpm_flag in Tables 2 to 4, or list1_flag in Figures 27 and 28. In this specification, mh_intra_luma_mpm_idx represents an index that indicates which candidate in the list to use. For example, mh_intra_luma_mpm_idx may be the mh_intra_luma_mpm_idx in Tables 2 to 4, or list1_index in Figure 27. Also, in this specification, xCb and yCb may be the x and y coordinates of the top-left corner of the current block. Also, cbWidth and cbHeight may be the width and height of the current block.
[0194] Figure 29 shows a peripheral position referenced in a multiple assumption prediction according to one embodiment of the present invention. Referring to Figure 29, as mentioned above, the encoder / decoder can reference a peripheral position in the process of creating a candidate list for multiple assumption predictions. For example, the aforementioned candIntraPredModeX may be required. In this case, the positions A and B around the current block that are referenced may be NbA and NbB shown in Figure 29. That is, they may be the left and upper positions adjacent to the upper left end sample of the current block. For example, if the upper left end position of the current block is Cb as shown in Figure 18, and its coordinates are (xCb, yCb), then NbA may be (xNbA, yNbA) = (xCb-1, yCb) and NbB may be (xNbB, yNbB) = (xCb, yCb-1).
[0195] Figure 30 shows a method for referencing a peripheral mode according to one embodiment of the present invention. Referring to Figure 30, as described above, the encoder / decoder can reference a peripheral position in the process of creating a candidate list of multiple assumption predictions. Furthermore, the peripheral mode can be used as is, or a candidate mode list can be generated using a mode based on the peripheral mode. As an example, the mode generated by referencing a peripheral position may be candIntraPredModeX. In one embodiment, if the peripheral position is unavailable, candIntraPredModeX may be a mode that has already been set. If it is unavailable, this may include cases where the peripheral position uses interpretation, or where the mode has not been determined in the defined encoding / decoding order. Alternatively, if the peripheral position does not use multiple assumption prediction, candIntraPredModeX may be a mode that has already been set. Alternatively, if the peripheral position is above the CTU to which the current block belongs, candIntraPredModeX may be a mode that has already been set. As another example, if the peripheral position is outside the CTU to which the current block belongs, candIntraPredModeX may be a mode that has already been set. Also, in one embodiment, the aforementioned mode that has already been set may be a DC mode. In yet another embodiment, the already set mode may be the planar mode.
[0196] According to one embodiment of the present invention, candIntraPredModeX may be set depending on whether the mode of the peripheral position exceeds the threshold angle, or whether the index of the mode of the peripheral position exceeds the threshold. For example, if the index of the mode of the peripheral position is greater than the specific directional mode index, the encoder / decoder can set candIntraPredModeX to the vertical mode index. Also, if the index of the mode of the peripheral position is less than or equal to the specific directional mode index and is a directional mode, the encoder / decoder can set candIntraPredModeX to the horizontal mode index. For example, the specific directional mode index may be mode 34 in Figure 6 described above. Furthermore, if the mode of the peripheral position is a planar mode or a DC mode, candIntraPredModeX may be set to a planar mode or a DC mode.
[0197] Referring to Figure 30, mh_intra_flag may be a syntax element (or variable, parameter) indicating whether or not multiple assumption prediction is used. Also, the intra prediction mode used by the surrounding block may be X. Furthermore, although the current block can use multiple assumption prediction and can generate a candidate list using a candidate intra prediction mode based on the mode of the surrounding block, since the surrounding did not use multiple assumption prediction, the encoder / decoder can set the candidate intra prediction mode to the already set mode, DC mode, regardless of the intra prediction mode of the surrounding block used, or regardless of whether or not the surrounding block used intra prediction.
[0198] An example of the peripheral mode referencing method described above is further described below.
[0199] According to one embodiment of the present invention, the intra-predictive mode candIntraPredModeX (where X is A or B) of a peripheral block can be induced in the following manner.
[0200] 1. The availability induction process for the block specified in the adjacent block availability confirmation process is called. The availability induction process can set the position (xCurr, yCurr) to (xCb, yCb), set the input (xNbX, yNbX) to the adjacent position (xNbY, yNbY), and the output can be assigned to the availability variable availableX.
[0201] 2. The candidate intra prediction mode candIntraPredModeX may be induced in the following way:
[0202] A. If one or more of the following conditions are true, candIntraPredModeX may be set to the INTRA_DC mode.
[0203] a) If the variable availableX is FALSE.
[0204] b) If mh_intra_flag[xNbX][yNbX] is not 1.
[0205] c) If X is B and yCb - 1 is less than ((yCb >> CtbLog2SizeY) << CtbLog2SizeY).
[0206] B. Otherwise, if IntraPredModeY[xNbX][yNbX] > INTRA_ANGULAR34, candIntraPredModeX may be set to INTRA_ANGULAR50.
[0207] C. Otherwise, if IntraPredModeY[xNbX][yNbX] <= INTRA_ANGULAR34 and IntraPredModeY[xNbX][yNbX] > INTRA_DC, candIntraPredModeX may be set to INTRA_ANGULAR18.
[0208] D. Otherwise, candIntraPredModeX may be set equal to IntraPredModeY[xNbX][yNbX].
[0209] As mentioned above, in the list setting method described above, candIntraPredModeX may be determined by the peripheral mode reference method.
[0210] Figure 31 shows a candidate list generation method according to one embodiment of the present invention. According to the first and second list generation methods described in Tables 2 to 4 above, the encoder / decoder can generate a first list by referring to the modes around the current block, and generate a second list using the candidate modes that are not included in the first list. Since there is spatial similarity within the picture, modes that refer to the surrounding area may have a higher priority. That is, the first list may have a higher priority than the second list. However, according to the first and second list signaling methods described in Tables 2 to 4, if the signaling to determine the list is not inferred, the encoder / decoder can signal using flags and indices to use the modes in the first list, and use only flags to use the modes in the second list. That is, the encoder / decoder can use relatively fewer bits for signaling the second list. However, using relatively more bits for signaling the modes in a higher priority list may be detrimental to coding efficiency. Therefore, according to one embodiment of the present invention, a method is proposed in which relatively fewer bits of signaling are used for high-priority lists and modes.
[0211] In other words, according to one embodiment of the present invention, the candidate list generation method may be individually defined depending on whether or not only lists with relatively high priority are usable. Whether or not only the first list is usable can indicate whether or not signaling indicating the list to be used is inferred. For example, assuming that there is a second list that is generated by a method already set using candidate modes, the encoder / decoder can insert the third list by dividing it into the first list and the second list. For example, the third list and the method for generating the third list that are generated by a method already set may be the candidate mode list and its generation method described above. If signaling indicating the list to be used is inferred, only the first list can be used, in which case the encoder / decoder can fill the first list from the beginning of the third list. Also, if signaling indicating the list to be used is not inferred, either the first or second list may be used, in which case the second list can be filled from the beginning of the third list, and the remainder can be filled in the first list. At this time, the encoder / decoder can also fill the third list in order when filling the first list. In other words, we can now refer to the mode around the current block and add candIntraPredModeX to the candidate list. CandIntraPredModeX can be added to the first list if a signaling indicating the list is inferred, and to the second list if it is not inferred.
[0212] In one embodiment, the size of the second list may be 1, in which case the encoder / decoder can add candIntraPredModeA to the first list if a signaling indicating the list is inferred, and to the second list if it is not inferred. candIntraPredModeA may be list 3[0], which is the first mode in the third list. Therefore, in the present invention, candIntraPredModeA, a mode based on peripheral modes, may be added to both the first and second lists as needed. On the other hand, in the methods described in Tables 2 to 4, candIntraPredModeA could only be included in the first list. That is, in the present invention, the first list generation method may be set individually depending on whether or not a signaling indicating the list to be used is inferred.
[0213] Referring to Figure 31, the candidate modes may be candidates that can be used to generate intra-predictions for multiple assumption predictions. That is, the method of generating the candidate list may differ depending on whether or not the signaling list1_flag, which indicates the list to be used, is inferred (S3101, S3102). If it is inferred, it can be inferred using the first list, and only the first list can be used, so the encoder / decoder can add the top-ranking elements from the third list to the first list (S3103). When generating the third list, the encoder / decoder can add modes based on surrounding modes to the top. Also, if it is not inferred, both the first and second lists may be used, so the encoder / decoder can add the top-ranking elements from the third list to the second list, which has fewer signaling elements (S3104). And if the first list is necessary, for example, if it is signaled to use the first list, the encoder / decoder can add the candidates from the third list to the first list, excluding those included in the second list (S3105). Although this invention primarily describes the case of generating a third list for the sake of explanation, it is not limited to this, and candidates can be temporarily classified, and the first and second lists can be generated based on this classification.
[0214] According to one embodiment of the present invention, the method for generating a candidate list described below and the method for generating a candidate list described in Figures 27-28 can be used adaptively depending on the circumstances. For example, the encoder / decoder can choose one of the two methods for generating candidate lists depending on whether or not signaling indicating which list to use is inferred. This may also be the case for multiple assumption prediction. Furthermore, the first list below may contain three modes, and the second list may contain one mode. Also, as described in Figures 27-28 above, the encoder / decoder can signal the modes in the first list with flags and indices, and the modes in the second list with flags.
[0215] In one embodiment, when candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is in planar mode or DC mode, it may be determined that List2[0] = planar mode, List1[0] = DC mode, List1[1] = vertical mode, and List1[2] = horizontal mode. In yet another embodiment, when candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is in planar mode or DC mode, it may be determined that List2[0] = candIntraPredModeA, List1[0] = !candIntraPredModeA, List1[1] = vertical mode, and List1[2] = horizontal mode.
[0216] Furthermore, in one embodiment, when candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is a directional mode, List2[0]=candIntraPredModeA, List1[0]=planar mode, List1[1]=DC mode, and List1[2]=vertical mode may be determined. Also, in one embodiment, when candIntraPredModeA and candIntraPredModeB are different, List2[0]=candIntraPredModeA and List1[0]=candIntraPredModeB may be determined. List1[1] and List1[2] may be added from among the planar mode, DC mode, vertical mode, and horizontal mode that are not candIntraPredModeA and candIntraPredModeB.
[0217] Figure 32 shows a method for generating a candidate list according to one embodiment of the present invention. In the above embodiments, a method for determining a mode based on multiple lists has been described. Referring to Figure 32, according to one embodiment of the present invention, the mode may be determined based on a single list instead of multiple lists. Specifically, as shown in Figure 32, a single candidate list containing all candidate modes for multiple assumption prediction may be generated. The following Table 5 illustrates a syntax structure related to multiple assumption prediction according to one embodiment of the present invention.
[0218] [Table 5]
[0219] Referring to Table 5, since there is only one candidate list, there is no signaling to select the list, and there may be an index signaling to indicate which mode of the candidate list to use. Therefore, the decoder can parse the candidate index mh_intra_luma_idx if mh_intra_flag, which indicates whether or not multiple assumption prediction is used, is 1. According to one embodiment of the present invention, the method for generating a candidate list for multiple assumption prediction can be based on the MPM list generation method in existing intra prediction. Alternatively, according to one embodiment, the candidate list for multiple assumption prediction may consist of a list in which the first list and the second list are combined in the order described in the first list and second list generation method described in Figure 28 above.
[0220] In other words, if the candidate list for multiple assumption prediction is the candidate mode list, then in this embodiment the size of the candidate mode list can be 4. If candIntraPredModeA and candIntraPredModeB are the same and are either planar mode or DC mode, the candidate mode list may be determined by the already set order. For example, candModeList[0]=planar mode, candModeList[1]=DC mode, candModeList[2]=vertical mode, candModeList[3]=horizontal mode. As yet another example, if candIntraPredModeA and candIntraPredModeB are the same and are either planar mode or DC mode, candModeList[0]=candIntraPredModeA, candModeList[1]=!candIntraPredModeA, candModeList[2]=vertical mode, candModeList[3]=horizontal mode.
[0221] Alternatively, if candIntraPredModeA and candIntraPredModeB are the same and are directional modes, candModeList[0]=candIntraPredModeA, candModeList[1]=planar mode, candModeList[2]=DC mode, and candModeList[3]=candIntraPredModeA and modes other than planar mode and DC mode may be determined. If candIntraPredModeA and candIntraPredModeB are different, candModeList[0]=candIntraPredModeA and candModeList[1]=candIntraPredModeB may be determined. Furthermore, candModeList[2] and candModeList[3] can sequentially add modes other than candIntraPredModeA and candIntraPredModeB according to the already set order of candidate modes. The already set order may be determined to be planar mode, DC mode, vertical mode, and horizontal mode.
[0222] According to one embodiment of the present invention, the candidate list may vary depending on the block size condition. For example, if one of the block width and height is more than n times the other, the candidate list may be shorter. For example, if the width is more than n times the height, the encoder / decoder can satisfy the condition by removing the horizontal mode from the candidate list as shown in Figure 31 and moving the next mode forward. Also, if the height is more than n times the width, the encoder / decoder can satisfy the condition by removing the vertical mode from the candidate list as shown in Figure 32 and moving the next mode forward. Therefore, if the width is more than n times the height, the candidate list size may be 3. Furthermore, if the width is more than n times the height, the size of the candidate list may be smaller than or equal to the size of the list in the case where the width is not greater.
[0223] As one embodiment, the candidate index in the embodiment described in Figure 32 may be coded with variable length. This can improve signaling efficiency by adding modes with a relatively high probability of being used to the beginning of the list. As another embodiment, the candidate index in the embodiment described in Figure 32 may be coded with fixed length. The number of modes used in multiple assumption prediction may be a number raised to the power of 2. For example, as mentioned above, it can be used in four intra-prediction modes. In such a case, even with fixed-length coding, no values that cannot be assigned will occur, and no unnecessary parts will be generated in the signaling. Also, when coded with fixed length, the number of items in the list configuration is only one. The number of bits will be the same regardless of which index is signaled.
[0224] In one embodiment, the candidate index may be coded with variable length or fixed length depending on the circumstances. For example, as in the above embodiment, the candidate list size may vary depending on the circumstances. In one embodiment, the candidate index may be coded with variable length or fixed length depending on the candidate list size. For example, if the candidate list size is a power of 2, it may be coded with fixed length, and if it is not a power of 2, it may be coded with variable length. In other words, according to the above embodiment, the coding method may change depending on the block size condition.
[0225] According to one embodiment of the present invention, when using multiple assumption prediction, if DC mode is used, the weighting values between multiple predictions may be the same for the entire block, thus leading to the same / similar results as adjusting the weighting values of the prediction block. Therefore, DC mode can be excluded in multiple assumption prediction. In one embodiment, the encoder / decoder can use only one of the following in multiple assumption prediction: planar mode, vertical mode, or horizontal mode. In such a case, the encoder / decoder can signal multiple assumption predictions using a single list, as shown in Figure 32. The encoder / decoder can also use variable-length coding for index signaling. In one embodiment, the list can be generated in a fixed order. For example, the order may be planar mode, vertical mode, and horizontal mode.
[0226] In another embodiment, the encoder / decoder can generate a list by referring to the mode around the current block. For example, if candIntraPredModeA and candIntraPredModeB are the same, candModeList[0] = candIntraPredModeA. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is in planar mode, candModeList[1] and candModeList[2] can be set according to the already set order. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is not in planar mode, candModeList[1] = planar mode and candModeList[2] = a mode other than candIntraPredModeA. If candIntraPredModeA and candIntraPredModeB are different, then candModeList[0]=candIntraPredModeA, candModeList[1]=candIntraPredModeB, and candModeList[2]=candIntraPredModeA or any mode other than candIntraPredModeB.
[0227] In another embodiment, the encoder / decoder may use only one of three modes in multiple assumption prediction. The three modes may include planar mode and DC mode. The three modes may also include either vertical mode or horizontal mode depending on the condition. The condition may be related to the block size. For example, whether the horizontal mode or vertical mode is included can be determined by which of the block width or height is greater. For example, if the block width is greater than the height, the vertical mode may be included. If the block height is greater than the width, the horizontal mode may be included. If the block height and width are equal, a predefined specific mode from either the vertical mode or the horizontal mode may be included.
[0228] In one embodiment, the encoder / decoder can generate a list in a fixed order. For example, this may be defined as the order of planar mode, DC mode, and vertical or horizontal mode. In another embodiment, the list can be generated by referring to the mode around the current block. For example, if candIntraPredModeA and candIntraPredModeB are the same, then candModeList[0] = candIntraPredModeA. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is not a directional mode, the encoder / decoder can set candModeList[1] and candModeList[2] according to the previously set order. If candIntraPredModeA and candIntraPredModeB are the same and candIntraPredModeA is a directional mode, then candModeList[1] = planar mode and candModeList[2] = DC mode. If candIntraPredModeA and candIntraPredModeB are different, then candModeList[0]=candIntraPredModeA, candModeList[1]=candIntraPredModeB, and candModeList[2]=candIntraPredModeA is not the case; any mode other than candIntraPredModeB is acceptable.
[0229] According to another embodiment of the present invention, only one of two modes may be used in multiple assumption prediction. The two modes may include a planar mode. The two modes may also include either a vertical mode or a horizontal mode, depending on the conditions. The conditions may be related to the block size. For example, an encoder / decoder may include either a horizontal mode or a vertical mode depending on which of the block width or height is greater. For example, if the block width is greater than the height, it may include a vertical mode. If the block height is greater than the width, it may include a horizontal mode. If the block height and width are equal, it may include the promised mode, either a vertical mode or a horizontal mode. In such cases, a flag may be signaled to indicate which mode to use in multiple assumption prediction.
[0230] In one embodiment, an encoder / decoder can exclude certain modes depending on the block size. For example, if the block size is small, certain modes can be excluded. For instance, if the block size is small, only the planar mode can be used in multiple assumption prediction. If certain modes are excluded, mode signaling can be omitted or reduced.
[0231] According to another embodiment of the present invention, the encoder / decoder can use only one mode (i.e., an intra-prediction mode) to generate intra-predicted samples in multiple assumption prediction. In one embodiment, the one mode may be defined as a planar mode. As previously mentioned, multiple assumption prediction can use inter-predicted samples and intra-predicted samples. Generating inter-predicted samples while simultaneously determining the optimal prediction mode to generate intra-predicted samples increases prediction accuracy but can lead to increased encoding complexity and an increase in signaling bits. Therefore, when performing multiple assumption prediction, using only the planar mode, which statistically occurs most frequently, as the intra-prediction mode can improve encoding complexity, save signaling bits, and thereby improve video compression performance.
[0232] In yet another embodiment, the mode may be determined based on the block size, between the vertical mode and the horizontal mode. For example, the mode may be determined by which of the block's width or height is greater. For example, if the block's width is greater than its height, it may be determined as the vertical mode, and if the block's height is greater than its width, it may be determined as the horizontal mode. If the block's width and height are equal, the encoder / decoder may be determined as the previously set mode. If the block's width and height are equal, the encoder / decoder can be determined as the previously set mode, between the horizontal mode and the vertical mode. If the block's width and height are equal, the encoder / decoder can be determined as the previously set mode, between the planar mode and the DC mode.
[0233] Furthermore, according to embodiments of the present invention, there may be a flipping signaling that flips the predictions generated by the multiple assumption prediction. This allows for the effect of eliminating the residual on the opposite side by flipping, even if one mode is selected in the multiple assumption prediction. This also has the effect of reducing the number of candidate modes available for use in the multiple assumption prediction. More specifically, for example, in the above embodiments, flipping can be used when using only one mode. This can improve prediction performance. The flipping can mean flipping with respect to the x axis, flipping with respect to the y axis, or flipping with respect to both the x and y axes. In one embodiment, the flipping direction may be determined based on the mode selected in the multiple assumption prediction. For example, if the mode selected in the multiple assumption prediction is the planar mode, the encoder / decoder can determine that it is flipping with respect to both the x and y axes. Also, flipping with respect to both the x and y axes may depend on the block size or shape. For example, if the block is not square, the encoder / decoder can determine that it is not flipping with respect to both the x and y axes. For example, if the mode selected by the multiple assumption prediction is the horizontal mode, the encoder / decoder can determine that it is flipping along the x-axis. For example, if the mode selected by the multiple assumption prediction is the vertical mode, the encoder / decoder can determine that it is flipping along the y-axis. Also, if the mode selected by the multiple assumption prediction is the DC mode, the encoder / decoder can determine that there is no flipping and does not need to perform explicit signaling / parsing.
[0234] Furthermore, in multiple-assumption prediction, the DC mode can have an effect similar to illumination compensation. Therefore, according to one embodiment of the present invention, if either the DC mode or the illumination compensation method is used in multiple-assumption prediction, the other may not be used. Also, multiple-assumption prediction can have an effect similar to GBi (generalized bi-prediction). For example, in multiple-assumption prediction, the DC mode can have an effect similar to GBi. GBi prediction may be a technique for adjusting the weighting values between two reference blocks of bidirectional prediction at the block level or CU level. Therefore, according to one embodiment of the present invention, if either multiple-assumption prediction (or the DC mode in multiple-assumption prediction) or the GBi prediction method is used, the other may not be used. Furthermore, this may include cases where the prediction in multiple-assumption prediction is a bidirectional prediction. For example, if the selected merge candidate in multiple-assumption prediction is a bidirectional prediction, GBi prediction may not be used. In these embodiments, the relationship between multiple-assumption prediction and GBi prediction may be limited to when a specific mode of multiple-assumption prediction, such as the DC mode, is used. Alternatively, if GBi prediction-related signaling exists before multiple assumption prediction-related signaling, using GBi prediction eliminates the need to use multiple assumption prediction or a specific mode of multiple assumption prediction. In the present invention, not using either method can mean not signaling to either of the aforementioned methods and not parsing the related syntax.
[0235] Figure 33 shows a peripheral position referenced in a multiple assumption prediction according to one embodiment of the present invention. As mentioned above, the encoder / decoder can reference a peripheral position in the process of creating a candidate list for multiple assumption predictions. For example, the candIntraPredModeX mentioned above may be used. In this case, the positions A and B around the current block that are referenced may be NbA and NbB shown in Figure 33. For example, if the position of the upper left corner sample of the current block is Cb, as shown in Figure 29, and its coordinates are (xCb, yCb), then NbA may be (xNbA, yNbA) = (xCb-1, yCb + cbHeight-1) and NbB may be (xNbB, yNbB) = (xCb + cbWidth-1, yCb-1). Here, cbWidth and cbHeight may be the width and height of the current block, respectively. Furthermore, in the process of creating the candidate list for multiple assumption predictions, the peripheral positions may be the same as the peripheral positions referenced in the generation of the MPM list for intra-prediction.
[0236] In another embodiment, the peripheral positions referenced in the process of creating the candidate list for multiple assumption predictions may be the left-center and top-center positions of the current block, or positions close to them. For example, NbA and NbB may be (xCb-1, yCb+cbHeight / 2-1) and (xCb+cbWidth / 2-1, yCb-1). Alternatively, NbA and NbB may be (xCb-1, yCb+cbHeight / 2) and (xCb+cbWidth / 2, yCb-1).
[0237] Figure 34 shows a method for referencing the surrounding mode according to one embodiment of the present invention. As mentioned above, the surrounding position can be referenced in the process of creating a candidate list for multiple assumption prediction. However, in the embodiment of Figure 30, if the surrounding position does not use multiple assumption prediction, candIntraPredModeX is set to an already set mode. This is because when setting candIntraPredModeX, the mode of the surrounding position may not be the same as candIntraPredModeX. Therefore, according to one embodiment of the present invention, even if the surrounding position does not use multiple assumption prediction, if the mode used by the surrounding position is a mode used for multiple assumption prediction, the encoder / decoder can set candIntraPredModeX to the mode used by the surrounding position. The modes used for multiple assumption prediction may be planar mode, DC mode, vertical mode, or horizontal mode.
[0238] Alternatively, even if the surrounding position does not use multiple assumption prediction, if the mode used by the surrounding position is a specific mode, the encoder / decoder can set candIntraPredModeX to the mode used by the surrounding position. Alternatively, even if the surrounding position does not use multiple assumption prediction, if the mode used by the surrounding position is vertical mode or height mode, the encoder / decoder can set candIntraPredModeX to the mode used by the surrounding position. Alternatively, if the surrounding position is above the current block, even if the surrounding position does not use multiple assumption prediction, if the mode used by the surrounding position is vertical mode, the encoder / decoder can set candIntraPredModeX to the mode used by the surrounding position. Furthermore, if the surrounding position is to the left of the current block, even if the surrounding position does not use multiple assumption prediction, if the mode used by the surrounding position is horizontal mode, the encoder / decoder can set candIntraPredModeX to the mode used by the surrounding position.
[0239] Referring to Figure 34, mh_intra_flag may be a syntax element (or variable) indicating whether or not multiple assumption prediction is used. Also, the intra prediction mode used by the surrounding block may be horizontal mode. Furthermore, even if the current block can use multiple assumption prediction and a candidate list can be generated using candIntraPredMode based on the mode of the surrounding block, if the surrounding block does not use multiple assumption prediction, the encoder / decoder can set candIntraPredMode to horizontal mode because the intra prediction mode of the surrounding block is a specific mode, for example, horizontal mode.
[0240] The example of the peripheral mode referencing method described above will be described again below in combination with the other embodiment shown in Figure 30 above. According to one embodiment of the present invention, the intra-prediction mode candIntraPredModeX (where X is A or B) of a peripheral block can be induced in the following manner.
[0241] 1. An availability induction process is called for the block specified in the adjacent block availability check process. The availability induction process sets the position (xCurr, yCurr) to (xCb, yCb) and takes the adjacent position (xNbY, yNbY) as input (xNbX, yNbX). The output may be assigned to the availability variable availableX.
[0242] 2. The candidate intra-prediction mode candIntraPredModeX may be derived in the following way:
[0243] A. If one or more of the following conditions are true, candIntraPredModeX may be set to INTRA_DC mode.
[0244] a) When the variable availableX is FALSE.
[0245] b) When mh_intra_flag[xNbX][yNbX] is not 1 and IntraPredModeY[xNbX][yNbX] is not INTRA_ANGULAR50 and INTRA_ANGULAR18.
[0246] c) When X is B and yCb-1 is smaller than ((yCb>>CtbLog2SizeY)<<CtbLog2SizeY).
[0247] B. Otherwise, if IntraPredModeY[xNbX][yNbX]>INTRA_ANGULAR34, candIntraPredModeX may be set to INTRA_ANGULAR50.
[0248] C. Otherwise, if IntraPredModeY[xNbX][yNbX]<=INTRA_ANGULAR34 and IntraPredModeY[xNbX][yNbX]>INTRA_DC, candIntraPredModeX may be set to INTRA_ANGULAR18.
[0249] D. Otherwise, candIntraPredModeX may be set to IntraPredModeY[xNbX][yNbX].
[0250] According to another embodiment of the present invention, the intra prediction mode candIntraPredModeX (X is A or B) of the peripheral block can be derived in the following manner.
[0251] 1. The availability derivation process for the block specified in the adjacent block availability confirmation process is called. The availability derivation process may set the position (xCurr, yCurr) to (xCb, yCb), and the adjacent positions (xNbY, yNbY) can take (xNbX, yNbX) as input, and the output may be assigned to the availability variable availableX.
[0252] 2. The candidate intra prediction mode candIntraPredModeX can be derived in the following way:
[0253] A. If one or more of the following conditions are true, candIntraPredModeX may be set to the INTRA_DC mode.
[0254] a) If the variable availableX is FALSE.
[0255] b) If mh_intra_flag[xNbX][yNbX] is not 1 and IntraPredModeY[xNbX][yNbX] is not INTRA_PLANAR, INTRA_DC, INTRA_ANGULAR50, or INTRA_ANGULAR18.
[0256] c) If X is B and yCb-1 is less than ((yCb >> CtbLog2SizeY) << CtbLog2SizeY).
[0257] B. Otherwise, if IntraPredModeY[xNbX][yNbX] > INTRA_ANGULAR34, candIntraPredModeX may be set to INTRA_ANGULAR50.
[0258] C. Otherwise, if IntraPredModeY[xNbX][yNbX] <= INTRA_ANGULAR34 and IntraPredModeY[xNbX][yNbX] > INTRA_DC, candIntraPredModeX may be set to INTRA_ANGULAR18.
[0259] D. Otherwise, candIntraPredModeX may be set to IntraPredModeY[xNbX][yNbX].
[0260] In the list setting method described above, candIntraPredModeX may be determined by the peripheral mode reference method described above.
[0261] Figure 35 shows a method using peripheral reference samples according to one embodiment of the present invention. As described above, when using multiple assumption prediction, intra predictions can be used in combination with other predictions. Therefore, when using multiple assumption prediction, intra predictions can be generated by using samples around the current block as reference samples. According to one embodiment of the present invention, when using multiple assumption prediction, a mode using reconstructed samples can be used. Also, when not using multiple assumption prediction, the mode using reconstructed samples does not need to be used. The reconstructed samples may be reconstructed samples around the current block.
[0262] One example of a mode using the aforementioned restored samples is the template matching method. That is, a restored sample at a position already set based on a certain block can be defined as a template (or template region). The template matching may be an operation that compares the cost of the template of the block to be compared with the template of the current block and searches for the block with the smaller cost. In this case, the cost can be defined as the sum of the absolute values of the templates, the sum of the squares of the differences, etc. For example, an encoder / decoder can use template matching between the current block and the reference picture block to search for a block that is expected to be similar to the current block, and based on this, it can set or refine the motion vector. Other examples of modes using the aforementioned restored samples include motion compensation and motion vector improvement using restored samples.
[0263] To use the reconstructed samples around the current block, decoding the current block must wait until the decoding of the surrounding blocks is complete. In such cases, parallel processing of the current block and surrounding blocks may be difficult. Therefore, if multiple assumption prediction is not used, the encoder / decoder does not need to use a mode that uses reconstructed samples around the current block to enable parallel processing. Alternatively, if multiple assumption prediction is used, the encoder / decoder can also use other modes that use reconstructed samples around the current block, since intra-predictions can be generated using the reconstructed samples around the current block.
[0264] Furthermore, according to one embodiment of the present invention, even when using multiple assumption predictions, the availability of reconstructed samples around the current block may differ depending on the candidate index. In one embodiment, when the candidate index is smaller than a predetermined specific threshold, reconstructed samples around the current block can be used. When the candidate index is small, the number of candidate index signaling bits is small, the accuracy of the candidate can be high, and the accuracy can be further improved by using reconstructed samples for candidates with high coding efficiency. In another embodiment, when the candidate index is larger than a specific threshold, reconstructed samples around the current block can be used. When the candidate index is large, the number of candidate index signaling bits is large, the accuracy of the candidate may be low, and the accuracy can be compensated for by using reconstructed samples around the current block for candidates with low accuracy.
[0265] According to one embodiment of the present invention, when using multiple assumption prediction, the encoder / decoder can generate an inter-prediction using reconstructed samples around the current block, and combine the inter-prediction with the intra-prediction of the multiple assumption prediction to generate a prediction block. Referring to Figure 35, the mh_intra_flag value, which is a signaling that the current block uses multiple assumption prediction, is 1. Since the current block uses multiple assumption prediction, the mode using reconstructed samples around the current block can be used.
[0266] Figure 36 shows a transformation mode according to one embodiment of the present invention. According to one embodiment of the present invention, there may be a transformation mode that performs a transformation only on the sub-parts of a block. In this specification, a transformation mode to which a transformation is applied only to sub-parts may be called a sub-block transform (SBT) or a spatially varying transform (SVT). For example, a CU or PU may be divided into multiple TUs, and only a part of the multiple TUs may be transformed. Or, for example, only one of the multiple TUs may be transformed. The TUs that are not transformed among the multiple TUs may be set to have a residual of 0.
[0267] Referring to Figure 36, there are two types of SBT-V (SBT-vertical) and SBT-H (SBT-horizontal) for dividing one CU or PU into multiple TUs. SBT-V may be a type in which the height of the multiple TUs is equal to the height of the CU or PU, and the width of the multiple TUs is different from the width of the CU or PU. SBT-H may be a type in which the height of the multiple TUs is different from the height of the CU or PU, and the width of the multiple TUs is equal to the width of the CU or PU. In one embodiment, the width and position of the TUs to be converted in SBT-V may be signaled. Also, the height and position of the TUs to be converted in SBT-H may be signaled. In one embodiment, a conversion kernel based on the SBT type and position, width or height may be pre-set.
[0268] Thus, the existence of a mode that transforms only a portion of the CU or PU is possible because, after prediction, the residuals can primarily reside in a portion of the CU or PU. In other words, SBT has the same concept as the skip mode for TU. In this case, the existing skip mode may be the skip mode for CU.
[0269] Referring to FIG. 36, for each type of SBT-V ((a) and (b) in FIG. 36) and SBT-H ((c) and (d) in FIG. 36), the position for A to be displayed and converted may be defined. And the width or height may be defined as 1 / 2 or 1 / 4 of the CU width or CU height. Also, the parts other than the area where A is displayed can be regarded as having a residual value of 0. Also, the conditions under which SBT can be used may be defined. For example, the conditions for SBT to be possible may be signaled as to whether it can be used in the syntax of conditions related to the block size and the high level (e.g., sequence, slice, tile, etc.).
[0270] According to an embodiment of the present invention, there can be a relationship between multiple hypothesis prediction and the conversion mode. For example, the presence or absence of the use of one of them may determine the presence or absence of the use of the other. Or, the presence or absence of the use of one mode may determine the presence or absence of the use of the other mode. Or, the presence or absence of the use of one of them may determine the presence or absence of the use of the other mode. As an example, the conversion mode may be the SBT described in FIG. 36. That is, the presence or absence of the use of SBT may be determined by the presence or absence of the use of multiple hypothesis prediction. Or, the presence or absence of the use of multiple hypothesis prediction may be determined by the presence or absence of the use of SBT. The following Table 6 shows a syntax structure exemplifying the relationship between multiple hypothesis prediction and the conversion mode according to an embodiment of the present invention.
[0271]
Table 6
[0272] Referring to Table 6, the use of SBT may be determined by whether or not multiple assumption prediction is used. That is, if multiple assumption prediction is not applied to the current block (!mh_intra_flag is true), the decoder can parse cu_sbt_flag, which is a syntax element indicating whether or not SBT is applied. If multiple assumption prediction is not applied to the current block, cu_sbt_flag does not need to be parsed. In this case, the value of cu_sbt_flag may be inferred to be 0 by a predefined condition.
[0273] As mentioned above, SBT continues to improve compression performance when residuals after prediction for a processing block exist only in a portion of the block. On the other hand, in the case of multiple assumption prediction, inter-prediction effectively reflects the movement of objects, while intra-prediction efficiently makes predictions for the remaining area, thus improving prediction performance for the entire block. In other words, when multiple assumption prediction is applied, prediction performance for the entire block improves, and the phenomenon of residuals being concentrated in only a portion of the block occurs relatively less frequently. Therefore, according to one embodiment of the present invention, if multiple assumption prediction is applied, the encoder / decoder does not need to apply SBT. Or, if SBT is applied, the encoder / decoder does not need to apply multiple assumption prediction.
[0274] According to one embodiment of the present invention, the position of the TU transformed by SBT may be restricted depending on whether or not multiple assumption prediction is used or the mode of multiple assumption prediction. Alternatively, the width (SBT-V) or height (SBT-H) of the TU transformed by SBT may be restricted depending on whether or not multiple assumption prediction is used or the mode of multiple assumption prediction. This reduces signaling related to position, width, or height. For example, the position of the TU transformed by SBT does not have to be in a region where the weighting value of the intra prediction is high in the multiple assumption prediction. This is because the residual in the region with a large weighting value can be reduced by the multiple assumption prediction.
[0275] Therefore, when using multiple assumption prediction, there is no need for a mode in SBT that transforms the side with the larger weight value. For example, when using horizontal or vertical mode in multiple assumption prediction, the SBT types in Figure 36(b) and (d) may be omitted (i.e., not considered). In another embodiment, when using planar mode in multiple assumption prediction, the position of the TU transformed by SBT may be restricted. For example, when using planar mode in multiple assumption prediction, the SBT types in Figure 36(a) and (c) may be omitted. This is because when using planar mode in multiple assumption prediction, the region adjacent to the reference sample in the intra prediction may have a value similar to the reference sample value, and as a result, the residual in the region adjacent to the reference sample may be relatively small.
[0276] In other embodiments, when using multiple assumption prediction, the possible width or height values of TUs to be SBT converted may change. Alternatively, when using a specific mode in multiple assumption prediction, the possible width or height values of TUs to be SBT converted may change. For example, when using multiple assumption prediction, large values for the width or height of TUs to be SBT converted can be excluded because a large amount of residual does not need to remain in a large part of the block. Alternatively, when using multiple assumption prediction, values for the width or height of TUs to be SBT converted that are such as units where the weighting value changes in multiple assumption prediction can be excluded.
[0277] Referring to Table 6, there may be a cu_sbt_flag indicating whether SBT is used and a mh_intra_flag indicating whether multiple assumption prediction is used. If mh_intra_flag is 0, cu_sbt_flag can be parsed. Also, if cu_sbt_flag does not exist, it may be inferred to be 0. Both combining intra prediction with multiple assumption prediction and SBT can be used to solve the problem where a large amount of residuals may remain in only a part of the CU or PU when the respective techniques are not used. Therefore, since there may be a correlation between the two techniques, the use of one technique or the use of a specific mode of one technique can be determined based on its relationship to the other technique.
[0278] Furthermore, as shown in Table 6, sbtBlockConditions can indicate the conditions under which SBT is possible. These conditions may include conditions related to block size, and signaling values indicating whether or not it is possible at a higher level (e.g., sequence, slice, tile).
[0279] Figure 37 shows the relationship between color difference components according to one embodiment of the present invention. Referring to Figure 37, the color format may be expressed as chroma_format_idc, Chroma format, separate_colour_plane_flag, etc. If it is monochrome, there may be only one sample array. Also, SubWidthC and SubHeightC may both be 1. If it is 4:2:0 sampling, there may be two color difference arrays (or color difference components, color difference blocks). Also, the color difference array can have half the width and half the height of the lumen array (or lumen component, lumen block). SubWidthC and SubHeightC may both be 2. SubWidthC and SubHeightC can indicate the size of the chrominance array compared to the luminance array. If the width or height of the chrominance array is half the size of the luminance array, SubWidthC or SubHeightC can be 2. If the width or height of the chrominance array is the same size as the luminance array, SubWidthC or SubHeightC can be 1.
[0280] If the sampling is 4:2:2, there may be two color difference arrays. The color difference array may have half the width and the same height as the luminance array. SubWidthC and SubHeightC may be 2 and 1, respectively. If the sampling is 4:4:4, the color difference array may have the same width and height as the luminance array. SubWidthC and SubHeightC may both be 1. In this case, they may be processed individually based on separate_colour_plane_flag. If separate_colour_plane_flag is 0, the color difference array may have the same width and height as the luminance array. If separate_colour_plane_flag is 1, the three color planes (luminance, Cb, Cr) may be processed individually. Regardless of separate_colour_plane_flag, if the sampling is 4:4:4, SubWidthC and SubHeightC may both be 1.
[0281] If separate_colour_plane_flag is 1, then only one color component may exist in a single slice. If separate_colour_plane_flag is 0, then multiple color components may exist in a single slice. Referring to Figure 38, SubWidthC and SubHeightC can be different only in the case of 4:2:2. Therefore, in the case of 4:2:2, the relationship between luminance reference width and height, and the relationship between chrominance reference width and height, may be different.
[0282] For example, if the luminance sample reference width is widthL and the chrominance sample reference width is widthC, and if widthL and widthC correspond, then their relationship is as shown in mathematical equation 6 below.
[0283]
number
[0284] Similarly, if the luminance sample reference height is heightL and the chromatic difference sample reference height is heightC, and heightL and heightC correspond, then their relationship is as shown in the following mathematical equation 7.
[0285]
number
[0286] Furthermore, there may be values that indicate (indicate) a color component. For example, cIdx can indicate a color component. For example, cIdx may be a color component index. If cIdx is 0, it can indicate the luminance component. If cIdx is not 0, it can indicate the chrominance component. If cIdx is 1, it can indicate the chrominance Cb component. If cIdx is 2, it can indicate the chrominance Cr component.
[0287] Figure 38 shows the relationship between color components according to one embodiment of the present invention. Figures 38(a), (b), and (c) assume the cases of 4:2:0, 4:2:2, and 4:4:4, respectively. Referring to Figure 38(a), one color difference sample (one Cb and one Cr) may be located for every two luminance samples in the horizontal direction. Also, one color difference sample (one Cb and one Cr) may be located for every two luminance samples in the vertical direction. Referring to Figure 38(b), one color difference sample (one Cb and one Cr) may be located for every two luminance samples in the horizontal direction. Also, one color difference sample (one Cb and one Cr) may be located for every one luminance sample in the vertical direction. Referring to Figure 38(c), one color difference sample (one Cb and one Cr) may be located for every one luminance sample in the horizontal direction. Furthermore, one color difference sample (one Cb and one Cr) may be positioned for each luminance sample in the vertical direction.
[0288] As mentioned above, the SubWidthC and SubHeightC values explained in Figure 37 can be determined by this relationship. Based on SubWidthC and SubHeightC, conversions between luminance sample criteria and chrominance sample criteria can then be performed.
[0289] Figure 39 shows a peripheral reference position according to one embodiment of the present invention. According to an embodiment of the present invention, the encoder / decoder can reference a peripheral position during prediction. For example, as described above, when performing CIIP (combined inter-picture merge and intra-picture prediction), the peripheral position can be referenced. CIIP may be the multiple-assumption prediction described above. That is, CIIP may be a prediction method that combines inter-prediction (e.g., merge-mode inter-prediction) and intra-prediction. According to an embodiment of the present invention, the encoder / decoder can combine inter-prediction and intra-prediction by referencing a peripheral position. For example, the encoder / decoder can determine the ratio of inter-prediction to intra-prediction by referencing a peripheral position. Alternatively, the encoder / decoder can determine the weighting when combining inter-prediction and intra-prediction by referencing a peripheral position. Alternatively, the encoder / decoder can determine the weighting when performing a weighted sum (or weighted average) of inter-prediction and intra-prediction by referencing a peripheral position.
[0290] According to one embodiment of the present invention, the referenced peripheral position may include NbA and NbB. The coordinates of NbA and NbB may be (xNbA, yNbA) and (xNbB, yNbB), respectively. NbA may be the left position of the current block. Specifically, if the top-left coordinates of the current block are (xCb, yCb) and the width and height of the current block are cbWidth and cbHeight, respectively, then NbA may be (xCb-1, yCb + cbHeight-1). The top-left coordinates of the current block (xCb, yCb) may be a value based on the luminance sample. Alternatively, the top-left coordinates of the current block (xCb, yCb) may be the position of the top-left luminance sample of the current luminance coding block relative to the top-left luminance sample of the current picture. Furthermore, cbWidth and cbHeight can represent the width and height relative to the color component, respectively. The coordinates described above may be relative to the lumen component (or lumen block). For example, cbWidth and cbHeight can represent the width and height relative to the lumen component.
[0291] Furthermore, NbB may be the upper position of the current block. More specifically, if the top-left coordinates of the current block are (xCb, yCb) and the width and height of the current block are cbWidth and cbHeight respectively, then NbB may be (xCb + cbWidth-1, yCb-1). The top-left coordinates (xCb, yCb) of the current block may be a value based on a luminance sample. Alternatively, the top-left coordinates (xCb, yCb) of the current block may be the top-left luminance sample position of the current luminance coding block relative to the top-left luminance sample of the current picture. Furthermore, cbWidth and cbHeight may be values based on the corresponding color component. The coordinates described above may be relative to a luminance component (luminance block). For example, cbWidth and cbHeight may be values based on a luminance component.
[0292] Referring to Figure 39, the coordinates of the top-left corner, NbA, and NbB are illustrated relative to the luminance block. NbA may be the left-side position of the current block. More specifically, if the coordinates of the top-left corner of the current block are (xCb, yCb), and the width and height of the current block are cbWidth and cbHeight, respectively, then NbA may be (xCb-1, yCb+2*cbHeight-1). The coordinates of the top-left corner of the current block (xCb, yCb) may be values based on the luminance sample. Alternatively, the coordinates of the top-left corner of the current block (xCb, yCb) may be the position of the top-left luminance sample of the current luminance coding block relative to the top-left luminance sample of the current picture. Furthermore, cbWidth and cbHeight may be values based on the corresponding color component. The coordinates described above may be relative to the color difference component (color difference block). For example, cbWidth and cbHeight may be values based on the color difference component. Furthermore, these coordinates may be in a 4:2:0 format.
[0293] Furthermore, NbB may be the upper position of the current block. More specifically, if the top-left coordinates of the current block are (xCb, yCb) and the width and height of the current block are cbWidth and cbHeight respectively, then NbB may be (xCb + 2 * cbWidth - 1, yCb - 1). The top-left coordinates (xCb, yCb) of the current block may be values based on a luminance sample. Alternatively, the top-left coordinates (xCb, yCb) of the current block may be the position of the top-left luminance sample of the current luminance coding block relative to the top-left luminance sample of the current picture. Furthermore, cbWidth and cbHeight may be values based on the corresponding color component. The coordinates described above may be relative to a chroma block. For example, cbWidth and cbHeight may be values based on a chroma block. Furthermore, these coordinates may correspond to cases where the format is 4:2:0 or 4:2:2. Referring to Figure 39, the coordinates of the top left corner, NbA, NbB, etc., are illustrated for the color difference block.
[0294] Figure 40 shows a weighted sample prediction process according to one embodiment of the present invention. The embodiment in Figure 40 describes a method for combining two or more prediction signals. The embodiment in Figure 40 can also be applied when using CIIP. Furthermore, the embodiment in Figure 40 can include the peripheral position referencing method described in Figure 39. Referring to Figure 44, the variable scallFact, which represents the scaling factor, can be described by the following mathematical formula 8.
[0295]
number
[0296] In Mathematical Formula 8, the encoder / decoder can set scallFact to 0 when cIdx is 0, and set scallFact to 1 when cIdx is not 0. In an embodiment of the present invention, x?y:z may indicate a y value when x is true or x is not 0, and indicate a z value otherwise (when x is false (or when x is 0)).
[0297] Also, the encoder / decoder can set the coordinates (xNbA, yNbA) and (xNbB, yNbB) of the peripheral positions NbA and NbB referred to for multiple hypothesis prediction. According to the embodiment described in FIG. 39, for the luminance component, (xNbA, yNbA) and (xNbB, yNbB) are respectively (xCb - 1, yCb + cbHeight - 1) and (xCb + cbWidth - 1, yCb - 1), and for the color difference component, (xNbA, yNbA) and (xNbB, yNbB) may be set to (xCb - 1, yCb + 2*cbHeight - 1) and (xCb + 2*cbWidth - 1, yCb - 1) respectively. Also, the operation of multiplying by 2^n may be the same as the operation of left-shifting n bits. For example, the operation of multiplying by 2 may be calculated as the value of left-shifting 1 bit. Also, left-shifting x by n bits can be expressed as "x << n". Also, the operation of dividing by 2^n may be the same as the operation of right-shifting n bits. Also, the operation of dividing by 2^n and discarding the fractional part may be calculated as the value of right-shifting n bits. For example, the operation of dividing by 2 may be calculated as the value of right-shifting 1 bit. Also, right-shifting x by n bits can be expressed as "x >> n". Therefore, (xCb - 1, yCb + 2*cbHeight - 1) and (xCb + 2*cbWidth - 1, yCb - 1) can be expressed as (xCb - 1, yCb + (cbHeight << 1) - 1) and (xCb + (cbWidth << 1) - 1). Therefore, when representing both the coordinates for the luminance component and the coordinates for the color difference component described above, it is as shown in the following Mathematical Formula 9.
[0298]
number
[0299] In mathematical formula 9, scallFact may be determined as (cIdx == 0) ? 0:1, as mentioned above. In this case, cbWidth and cbHeight represent the width and height, respectively, relative to each color component. For example, if the width and height relative to the luminance component are cbWidthL and cbHeightL, respectively, and a weighted sample prediction process is performed on the luminance component, then cbWidth and cbHeight may be cbWidthL and cbHeightL, respectively. Also, if the width and height relative to the luminance component are cbWidthL and cbHeightL, respectively, and a weighted sample prediction process is performed on the chrominance component, then cbWidth and cbHeight may be cbWidthL / SubWidthC and cbHeightL / SubHeightC, respectively.
[0300] Furthermore, according to one embodiment of the present invention, an encoder / decoder can determine the prediction mode for a given position by referring to the surrounding position. For example, the encoder / decoder can determine whether the prediction mode is intra-prediction. The prediction mode may be indicated by CuPredMode. If CuPredMode is MODE_INTRA, it may be a mode that uses intra-prediction. The CuPredMode value may also be MODE_INTRA, MODE_INTER, MODE_IBC, or MODE_PLT. If CuPredMode is MODE_INTER, inter-prediction can be used. If CuPredMode is MODE_IBC, intra-block copy (IBC) can be used. If CuPredMode is MODE_PLT, palette mode can be used. The CuPredMode may also be indicated by the channel type (chType) and position. For example, it may be indicated as CuPredMode[chType][x][y], where this value may be the CuPredMode value for the channel type chType at position (x,y).
[0301] Furthermore, according to one embodiment of the present invention, chType may be based on a tree type. For example, the tree type may be set to values such as SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA. When it is SINGLE_TREE, there may be a portion where the block partitioning of the luminance component and the chrominance component are shared. For example, when it is SINGLE_TREE, the block partitioning of the luminance component and the chrominance component may be the same. Or, when it is SINGLE_TREE, the block partitioning of the luminance component and the chrominance component may be the same or partially the same. Or, when it is SINGLE_TREE, the block partitioning of the luminance component and the chrominance component may be performed using the same syntax element value.
[0302] Furthermore, according to one embodiment of the present invention, when it is a DUAL TREE, the block partitioning of the luminance component and the chrominance component may be independent. Alternatively, when it is a DUAL TREE, the block partitioning of the luminance component and the chrominance component may be performed by different syntax element values. Also, when it is a DUAL TREE, the tree type value may be DUAL_TREE_LUMA or DUAL_TREE_CHROMA. If the tree type is DUAL_TREE_LUMA, DUAL TREE can be used to indicate that it is a process for the luminance component. If the tree type is DUAL_TREE_CHROMA, DUAL TREE can be used to indicate that it is a process for the chrominance component. Also, chType may be determined based on whether or not the tree type is DUAL_TREE_CHROMA. For example, chType may be set to 1 if the tree type is DUAL_TREE_CHROMA, and to 0 if the tree type is not DUAL_TREE_CHROMA. Therefore, referring to Figure 40, the CuPredMode[0][xNbX][yNbY] values can be determined. X can be replaced with A and B. That is, the CuPredMode values for the NbA and NbB positions can be determined.
[0303] Furthermore, according to one embodiment of the present invention, the isIntraCodedNeighbourX value can be set based on the determination of the prediction mode for the surrounding position. For example, the isIntraCodedNeighbourX value can be set depending on whether the CuPredMode for the surrounding position is MODE_INTRA or not. If the CuPredMode for the surrounding position is MODE_INTRA, the isIntraCodedNeighbourX value can be set to TRUE, and if the CuPredMode for the surrounding position is not MODE_INTRA, the isIntraCodedNeighbourX value can be set to FALSE. In the present invention described above and below, X may be replaced with A or B, etc. Also, X can indicate that it corresponds to position X.
[0304] Furthermore, according to one embodiment of the present invention, it is possible to determine whether a location is available by referring to its surrounding locations. Whether a location is available can be set with availableX. In addition, isIntraCodedNeighbourX can be set based on availableX. For example, if availableX is TRUE, isIntraCodedNeighbourX can be set to TRUE, and if availableX is FALSE, isIntraCodedNeighbourX can be set to FALSE. Referring to Figure 40, whether a location is available can be determined by calling "the derivation process for neighboring block availability". In addition, whether a location is available can be determined based on whether the location is currently inside the picture. If the location is (xNbY, yNbY) and xNbY or yNbY is less than 0, it is currently outside the picture, and availableX may be set to FALSE. Furthermore, if xNbY is greater than or equal to the picture width, the object is currently outside the picture, and availableX may be set to FALSE. The picture width can be indicated by pic_width_in_luma_samples. Also, if yNbY is greater than or equal to the picture height, the object is currently outside the picture, and availableX may be set to FALSE. The picture height can be indicated by pic_height_in_luma_samples. Furthermore, if the location is in a different block or slice than the current block, availableX may be set to FALSE. Also, if the reconstruction of the location is not yet complete, availableX may be set to FALSE. Whether the reconstruction is complete or not can be indicated by IsAvailable[cIdx][xNbY][yNbY]. In short, if any one of the following conditions is met, availableX can be set to FALSE, and if none of the following conditions are met, availableX can be set to TRUE.
[0305] - Condition 1: xNbY<0
[0306] - Condition 2: yNbY < 0
[0307] - Condition 3: xNbY>=pic_width_in_luma_samples
[0308] - Condition 4:yNbY>=pic_height_in_luma_samples
[0309] - Condition 5: IsAvailable[cIdx][xNbY][yNbY]==FALSE
[0310] - Condition 6: If the location in question (the surrounding location (xNbY, yNbY)) belongs to a different block (or slice) than the current block.
[0311] Furthermore, according to one embodiment of the present invention, it is possible to determine whether the current position and the position in question are in the same CuPredMode by option and set availableX. isIntraCodedNeighbourX can be set by combining the two conditions described above. For example, isIntraCodedNeighbourX can be set to TRUE if all of the following conditions are met, and isIntraCodedNeighbourX can be set to FALSE if not (if at least one of the following conditions is not met).
[0312] - Condition 1: availableX==TRUE
[0313] - Condition 2: CuPredMode[0][xNbX][yNbX]==MODE_INTRA
[0314] Furthermore, according to one embodiment of the present invention, the weights of CIIP can be determined based on a plurality of isIntraCodedNeighbourX. For example, when combining interprediction and intraprediction based on a plurality of isIntraCodedNeighbourX, the weights can be determined. For example, they can be determined based on isIntraCodedNeighbourA and isIntraCodedNeighbourB. According to one embodiment, if both isIntraCodedNeighbourA and isIntraCodedNeighbourB are TRUE, w can be set to 3. For example, w may be the weight of CIIP or a value that determines the weights. Also, if both isIntraCodedNeighbourA and isIntraCodedNeighbourB are FALSE, w can be set to 1. Also, if one of isIntraCodedNeighbourA and isIntraCodedNeighbourB is FALSE (the same as if one of both is TRUE), w can be set to 2. In other words, w can be set based on whether or not the surrounding position was predicted by intra-prediction, or to what extent the surrounding position was predicted by intra-prediction.
[0315] Furthermore, according to one embodiment of the present invention, w may be a weighting value corresponding to intra-prediction. Also, the weighting value corresponding to inter-prediction may be determined based on w. For example, the weighting value corresponding to inter-prediction may be (4-w). When the encoder / decoder combines two or more prediction signals, the following mathematical formula 10 can be used.
[0316]
number
[0317] In mathematical formula 10, predSamplesIntra and predSamplesInter may be prediction signals. For example, predSamplesIntra and predSamplesInter may be prediction signals predicted by intra-prediction and inter-prediction (e.g., merge mode, more specifically regular merge mode), respectively. Also, predSampleComb may be a prediction signal used in CIIP.
[0318] Furthermore, according to one embodiment of the present invention, a process of updating the pre-combination prediction signal may be included prior to applying mathematical formula 10. For example, the following mathematical formula 11 may be applied to the update process. For example, the process of updating the prediction signal may be a process of updating the interpretation prediction signal of the CIIP.
[0319]
number
[0320] Figure 41 is a diagram showing the peripheral reference position according to one embodiment of the present invention.
[0321] While peripheral reference positions were explained in Figures 39 and 40, problems can arise if the explained positions are used in all cases (for example, all color difference blocks), and this problem is explained in Figure 41. The embodiment in Figure 41 shows a color difference block. In Figures 39 and 40, the NbA and NbB coordinates relative to the luminance sample for the color difference block were (xCb-1, yCb+2*cbHeight-1) and (xCb+2*cbWidth-1, yCb-1), respectively. However, when SubWidthC or SubHeightC is 1, it may indicate a position different from that shown in Figure 39, as shown in Figure 41. Multiplying cbWidth and cbHeight by 2 in the above coordinates means that cbWidth and cbHeight are shown relative to each color component (color difference component in this embodiment), and the coordinates are shown relative to luminance, so in the case of 4:2:0, this may be to compensate for the number of luminance samples versus color difference samples. In other words, it is sufficient to represent the coordinates of a chrominance-based case where one chrominance sample corresponds to two luminance samples on the x-axis and one chrominance sample corresponds to two luminance samples on the y-axis. Therefore, if SubWidthC or SubHeightC is 1, it can indicate a different position. Consequently, if the (xCb-1, yCb+2*cbHeight-1) and (xCb+2*cbWidth-1, yCb-1) positions are always used for chrominance blocks in this way, it may occur to reference a position far from the current chrominance block. In this case, the relative position used for the luminance block of the current block and the relative position used for the chrominance block may not match. Also, because a different position is referenced for the chrominance block, it may be possible to set weight values by referencing a position with little correlation to the current block, or decoding / reconstruction may not occur in the block decoding order.
[0322] Referring to Figure 41, the position of the luminance reference coordinates described above is shown for the 4:4:4 case, that is, when both SubWidthC and SubHeightC are 1. NbA and NbB are located away from the color difference block shown by the solid line.
[0323] Figure 42 shows a weighted sample prediction process according to one embodiment of the present invention. The embodiment in Figure 42 may be an embodiment for solving the problems described in Figures 39 to 41. Furthermore, explanations that overlap with the above will be omitted. In Figure 40, the peripheral position is set based on scallFact, and as explained in Figure 41, scallFact is a value for converting the position when SubWidthC and SubHeightC are 2. However, as mentioned above, problems can occur depending on the color format, and the ratio of color difference samples to luminance samples may differ horizontally and vertically. Therefore, according to one embodiment of the present invention, scallFact can be separated into horizontal (width) and vertical (height).
[0324] According to one embodiment of the present invention, scallFactWidth and scallFactHeight may exist, and the peripheral position can be set based on scallFactWidth and scallFactHeight. The peripheral position may be set based on luminance samples (luminance blocks). Furthermore, scallFactWidth can be set based on cIdx and SubWidthC. For example, if cIdx is 0 or SubWidthC is 1, scallFactWidth can be set to 0, and otherwise (if cIdx is not 0 and SubWidthC is not 1 (if SubWidthC is 2)), scallFactWidth can be set to 1. In this case, scallFactWidth can be determined using the following mathematical formula 12.
[0325]
number
[0326] Furthermore, scallFactHeight can be set based on cIdx and SubHeightC. For example, if cIdx is 0 or SubHeightC is 1, scallFactHeight can be set to 0; otherwise (if cIdx is not 0 and SubHeightC is not 1 (i.e., if SubHeightC is 2)), scallFactHeight can be set to 1. In this case, scallFactHeight can be determined using the following mathematical formula 13.
[0327]
number
[0328] Furthermore, the x-coordinate of the surrounding position can be shown based on scallFactWidth, and the y-coordinate of the surrounding position can be shown based on scallFactHeight. For example, the coordinates of NbB can be set based on scallFactWidth. For example, the coordinates of NbA can be set based on scallFactHeight. Also, as mentioned above, basing on scallFactWidth here can be done by basing on SubWidthC, and basing on scallFactHeight can be done by basing on SubHeightC. For example, the coordinates of the surrounding position are as shown in the following mathematical formula 14.
[0329]
number
[0330] In this case, xCb and yCb may be the coordinates shown based on the luminance sample, as mentioned above. Also, cbWidth and cbHeight may be those shown based on each color component.
[0331] Therefore, if it is a color difference block and SubWidthC is 1, then (xNbB, yNbB) = (xCb + cbWidth - 1, yCb - 1). In other words, in this case, the NbB coordinate for the luminance block and the NbB coordinate for the color difference block may be the same. Also, if it is a color difference block and SubHeightC is 1, then (xNbA, yNbA) = (xCb - 1, yCb + cbHeight - 1). In other words, in this case, the NbA coordinate for the luminance block and the NbA coordinate for the color difference block may be the same.
[0332] Therefore, in the embodiment shown in Figure 42, if the format is 4:2:0, the same peripheral coordinates as in the embodiments shown in Figures 39 and 40 can be set, and if the format is 4:2:2 or 4:4:4, different peripheral coordinates from those in the embodiments shown in Figures 39 and 40 can be set.
[0333] The processes other than those shown in Figure 42 may be the same as those described in Figure 40. That is, the prediction mode or availability can be determined based on the peripheral position coordinates described in Figure 42, and the weighting value of CIIP can be determined. In this invention, peripheral position and peripheral position coordinates may be used interchangeably.
[0334] Figure 43 shows a weighted sample prediction process according to one embodiment of the present invention. The embodiment in Figure 43 represents the peripheral position coordinates described in Figure 42 in a different way. Therefore, content that overlaps with what has been said above may be omitted. As mentioned above, bit shifts can be represented by multiplication. Figure 42 may be shown using bit shifts, and Figure 43 may be shown using multiplication.
[0335] In one embodiment, scallFactWidth can be set based on cIdx and SubWidthC. For example, if cIdx is 0 or SubWidthC is 1, scallFactWidth can be set to 1; otherwise (if cIdx is not 0 and SubWidthC is not 1 (i.e., if SubWidthC is 2)), scallFactWidth can be set to 2. In this case, scallFactWidth can be determined using the following mathematical formula 15.
[0336]
number
[0337] Furthermore, scallFactHeight can be set based on cIdx and SubHeightC. For example, if cIdx is 0 or SubHeightC is 1, scallFactHeight can be set to 1; otherwise (if cIdx is not 0 and SubHeightC is not 1 (i.e., if SubHeightC is 2)), scallFactHeight can be set to 2. In this case, scallFactHeight can be determined using the following mathematical formula 16.
[0338]
number
[0339] Furthermore, the x-coordinate of the surrounding position can be shown based on scallFactWidth, and the y-coordinate of the surrounding position can be shown based on scallFactHeight. For example, the coordinates of NbB can be set based on scallFactWidth. For example, the coordinates of NbA can be set based on scallFactHeight. Also, as mentioned above, here, being based on scallFactWidth may be based on SubWidthC, and being based on scallFactHeight may be based on SubHeightC. For example, the coordinates of the surrounding position are as shown in the following mathematical formula 17.
[0340]
number
[0341] In this case, xCb and yCb may be the coordinates shown based on the luminance sample, as mentioned above. Also, cbWidth and cbHeight may be those shown based on each color component.
[0342] Figure 44 shows a weighted sample prediction process according to one embodiment of the present invention. In embodiments such as Figures 40, 42, and 43, the availability of a location was determined by referring to a surrounding location. At this time, cIdx, an index indicating a color component, was set to 0 (luminance component). Furthermore, when determining whether a location is available, cIdx can be used to determine whether the reconstruction of the location cIdx has been completed. That is, when determining whether a location is available, cIdx can be used to determine the IsAvailable[cIdx][xNbY][yNbY] value. However, when performing a weighted sample prediction process for a color difference block, referring to the IsAvailable value corresponding to cIdx0 may lead to an incorrect judgment. For example, if the restoration of the luminance component of a block including the neighboring position is not completed, but the restoration of the color difference component is completed, then if cIdx is not 0, IsAvailable[0][xNbY][yNbY] may be FALSE and IsAvailable[cIdx][xNbY][yNbY] may be TRUE. Therefore, it is possible that a neighboring position may be judged as unavailable even though it is actually available. In the embodiment of Figure 44, in order to solve this problem, when determining whether a position is available by referring to the neighboring position, the cIdx of the currently coded block can be used as input. That is, when calling "the derivation process for neighboring block availability", the input cIdx can be the cIdx of the currently coded block. In this embodiment, explanations that overlap with those described in Figures 42 and 43 are omitted.
[0343] Furthermore, when determining the prediction mode for the surrounding position as described above, CuPredMode[0][xNbX][yNbY], which corresponds to chType0, was referenced. However, if the chType for the current block does not match, an incorrect parameter may be referenced. Therefore, according to one embodiment of the present invention, when determining the prediction mode for the surrounding position, CuPredMode[chType][xNbX][yNbY] corresponding to the chType value for the current block can be referenced.
[0344] Figure 45 illustrates a video signal processing method based on multiple assumption prediction according to one embodiment to which the present invention is applied. Referring to Figure 45, the explanation will focus on the decoder for convenience, but the present invention is not limited thereto, and the multiple assumption prediction-based video signal processing method according to this embodiment can be applied to an encoder in substantially the same manner.
[0345] Specifically, the decoder can obtain a first syntax element indicating whether a combined prediction is applied to the current block if a merge mode is applied to the current block (S4501). Here, the combined prediction represents a prediction mode that combines inter-prediction and intra-prediction. As stated above, the present invention is not limited to these names, and in this specification, the multiple assumption prediction may be called multiple prediction, multiple prediction, combined prediction, inter-intra weighted prediction, combined inter-intra prediction, combined inter-intra weighted prediction, etc. In one embodiment, prior to step S4501, the decoder receives information for prediction of the current block and can determine whether a merge mode is applied to the current block based on the information for prediction.
[0346] The decoder can generate an inter-prediction block and an intra-prediction block for the current block if the first syntax element indicates that the combined prediction is applied to the current block (S4502). The decoder can then generate a combined prediction block by weighting the inter-prediction block and the intra-prediction block (S4503). The decoder can then decode the residual block of the current block and reconstruct the current block using the combined prediction block and the residual block.
[0347] As described above, in one embodiment, the step of decoding the residual block may include a step of obtaining a second syntax element indicating whether or not a sub-block transform is applied to the current block if the first syntax element indicates that the combinatorial prediction is not applied to the current block. That is, the sub-block transform is applicable only when the combinatorial prediction is not applied to the current block, and syntax signaling regarding whether or not to apply it may be performed. Here, the sub-block transform represents a transform mode that applies a transform to any one of the sub-blocks of the current block divided horizontally or vertically.
[0348] As mentioned above, in one embodiment, if the second syntax element is not present, the value of the second syntax element may be inferred to be 0.
[0349] As described above, in one embodiment, if the first syntax element indicates that the combination prediction is applied to the current block, the intra-prediction mode for the intra-prediction for the current block may be set to planar mode.
[0350] As described above, in one embodiment, the decoder can set the positions of the left and upper peripheral blocks referenced for the combination prediction, and can perform combination prediction based on the intra-prediction mode of the set positions. In one embodiment, the decoder can determine the weight values used for combination prediction based on the intra-prediction mode of the set positions. Also, as described above, the positions of the left and upper peripheral blocks may be determined using a scaling factor variable determined by the color component index value of the current block.
[0351] The embodiments of the present invention described above can be embodied through a variety of means. For example, embodiments of the present invention can be embodied through hardware, firmware, software, or a combination thereof.
[0352] In the case of hardware implementation, the method according to the embodiment of the present invention is implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSDPs (Digital Signal Processing Devices), PDLs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), processors, controllers, microcontrollers, microprocessors, etc.
[0353] In the case of implementation by firmware or software, the method according to the embodiment of the present invention is implemented in the form of a module, procedure, or function that performs the functions or operations described above. The software code is stored in memory and implemented by a processor. The memory is located inside or outside the processor and exchanges data with the processor by various already known means.
[0354] Some embodiments also embody the form of recording media containing computer-executable instructions, such as program modules executed by a computer. Computer-readable media are any available media that can be accessed by a computer, and include both volatile and non-volatile media, and isolated and non-isolated media. Computer-readable media also include both storage media and communication media. Computer storage media include both volatile and non-volatile media, isolated and non-isolated media, and embodied in any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Communication media typically include computer-readable instructions, data structures, or other data such as modulated data signals or program modules, or other transmission mechanisms, and include any information transmission media.
[0355] The above description of the present invention is illustrative, and a person with ordinary skill in the art to which the present invention pertains should understand that it can be easily modified into other specific forms without altering the technical idea or essential features of the present invention. Therefore, the above embodiments should be understood to be illustrative and not limiting in all respects. For example, each component described as a single type may be implemented in a distributed manner, and similarly, components described as distributed may be implemented in a combined manner.
[0356] The scope of the present invention is indicated by the claims described below rather than by the detailed description above, and all modifications or altered forms derived from the meaning and scope of the claims and the concept of equivalents thereto should be interpreted as being included within the scope of the present invention. [Explanation of Symbols]
[0357] 110 Conversion Unit 115 Quantization section 120 Inverse quantization section 125 Inverse Transform Section 130 Filtering section 150 Prediction Section 152 Intra Prediction Unit 154 Interpretation Unit 154a Motion Estimation Unit 154b Motion compensation unit 160 Entropy coding section 210 Entropy Decoding Unit 220 Inverse quantization section 225 Inverse Transform Section 230 Filtering section 250 Prediction Section 252 Intra Prediction Unit 254 Interpretation Unit
Claims
1. A video signal processing method, Currently, we are in the stage of receiving information for block prediction; Based on the information for the prediction, the step of determining whether or not a merge mode is applied to the current block; If a merge mode is applied to the current block, the step is to obtain a first syntax element indicating whether or not a combined prediction is applied to the current block, wherein the combined prediction indicates a prediction mode that combines inter-prediction and intra-prediction; If the first syntax element indicates that the combination prediction is applied to the current block, the steps are to generate an inter-prediction block and an intra-prediction block for the current block; and A video signal processing method comprising the step of generating a combined prediction block of the current block by weighting the inter prediction block and the intra prediction block.
2. The step of decoding the residual block of the current block; and The video signal processing method according to claim 1, further comprising the step of restoring the current block using the combination prediction block and the residual block.
3. The step of decoding the residual block is as follows: If the first syntax element indicates that the combination prediction is not applied to the current block, the process further includes obtaining a second syntax element indicating whether or not a sub-block transformation is applied to the current block. The video signal processing method according to claim 2, wherein the subblock conversion indicates a conversion mode in which the conversion is applied to only one subblock of the current block which is divided horizontally or vertically.
4. The video signal processing method according to claim 3, wherein if the second syntax element is not present, the value of the second syntax element is inferred to be 0.
5. The video signal processing method according to claim 1, wherein if the first syntax element indicates that the combination prediction is applied to the current block, the intra-prediction mode for generating the intra-prediction block of the current block is set to planar mode.
6. The process further includes the step of setting the positions of the left peripheral block and the upper peripheral block that are referenced for the aforementioned combination prediction, The video signal processing method according to claim 1, wherein the positions of the left peripheral block and the upper peripheral block are the same as the positions referenced by the intra prediction.
7. The video signal processing method according to claim 6, wherein the positions of the left peripheral block and the upper peripheral block are determined using a scaling factor variable determined by the color component index value of the current block.
8. A video signal processing device, Including the processor, The aforementioned processor, We are currently receiving information for block prediction. Based on the information for the prediction, it is determined whether or not the merge mode is applied to the current block. If the merge mode is applied to the current block, a first syntax element is obtained indicating whether or not combined prediction is applied to the current block, where the combined prediction indicates a prediction mode that combines interpretation and intraprediction. If the first syntax element indicates that the combination prediction is applied to the current block, then the inter-prediction block and intra-prediction block of the current block are generated. A video signal processing device that generates a combined prediction block of the current block by weighting the inter prediction block and the intra prediction block.
9. The aforementioned processor, The residual block of the current block is decoded, The video signal processing apparatus according to claim 8, wherein the current block is restored using the combination prediction block and the residual block.
10. The aforementioned processor, If the first syntax element indicates that the combination prediction is not applied to the current block, a second syntax element is obtained indicating whether or not a sub-block transformation is applied to the current block. The video signal processing apparatus according to claim 9, wherein the subblock conversion indicates a conversion mode in which the conversion is applied to any one of the subblocks of the current block which is divided horizontally or vertically.
11. The video signal processing apparatus according to claim 10, wherein if the second syntax element is not present, the value of the second syntax element is inferred to be 0.
12. The video signal processing apparatus according to claim 8, wherein if the first syntax element indicates that the combination prediction is applied to the current block, the intra-prediction mode for the intra-prediction for the current block is set to planar mode.
13. The video signal processing apparatus according to claim 8, wherein the processor sets the positions of the left and upper peripheral blocks referenced for the combination prediction, and the positions of the left peripheral block and the upper peripheral block are the same as the positions referenced by the intra prediction.
14. The video signal processing apparatus according to claim 13, wherein the positions of the left peripheral block and the upper peripheral block are determined using a scaling factor variable determined by the color component index value of the current block.
15. A video signal processing method, Currently, the block is being determined as to whether merge mode is applied or not; If a merge mode is applied to the current block, the step is to decode a first syntax element indicating whether or not a combined prediction is applied to the current block, wherein the combined prediction indicates a prediction mode that combines inter-prediction and intra-prediction; If the first syntax element indicates that the combination prediction is applied to the current block, the steps are to generate an inter-prediction block and an intra-prediction block for the current block; and A video signal processing method comprising the step of generating a combined prediction block of the current block by weighting the inter prediction block and the intra prediction block.