Video decoding device and video coding device
By generating two prediction images for a target block using geometric position-based weights and distinct parameter options, the video decoding and encoding device addresses complexity issues in MMVD prediction, enhancing encoding efficiency and reducing processing demands.
Patent Information
- Application Number
- PCT/JP2024/041585
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2024-11-25
- Publication Date
- 2025-09-04
AI Technical Summary
Existing video coding methods, such as those using GPM prediction, face increased complexity in MMVD prediction processing, leading to inefficiencies in encoding and decoding processes.
A video decoding and encoding device that employs a synthesis mode to generate two first prediction images for a target block, using weights determined by the geometric position of the block, with different parameter options for MMVD and intra prediction, reducing processing complexity and improving encoding efficiency.
The proposed solution reduces the complexity of MMVD prediction processing, thereby improving encoding efficiency while minimizing increases in processing volume and memory bandwidth.
Smart Images

Figure JP2024041585_04092025_PF_FP_ABST
Abstract
Description
Video decoding device and video encoding device
[0001] An embodiment of the present invention relates to a video decoding device and a video encoding device.
[0002] In order to efficiently transmit or record moving images, a moving image encoding device is used that encodes moving images to generate coded data, and a moving image decoding device is used that decodes the coded data to generate decoded images.
[0003] Specific video encoding methods include, for example, H.265 / HEVC (High-Efficiency Video Coding) and H.266 / VVC (Versatile Video Coding).
[0004] In such a video coding method, images (pictures) constituting a video are managed in a hierarchical structure consisting of slices obtained by dividing images, coding tree units (CTUs) obtained by dividing slices, coding units (sometimes called coding units (CUs)) obtained by dividing coding tree units, and transform units (TUs) obtained by dividing coding units, and are coded / decoded for each CU.
[0005] In such video coding methods, a predicted image is typically generated based on a locally decoded image obtained by encoding / decoding an input image, and the predicted image is subtracted from the input image (original image) to obtain a prediction error (sometimes called a "difference image" or "residual image"), which is then coded. Methods for generating predicted images include inter-frame prediction (inter-prediction) and intra-frame prediction (intra-prediction).
[0006] Furthermore, Non-Patent Document 1 discloses a GPM (Geometry Partition Mode) technique in which the area of a target block is divided into two partitions based on a selected partition mode, and a predicted image of the target block is derived by synthesizing predicted images generated for each partition.
[0007] "Algorithm description of Enhanced Compression Model 10 (ECM 10)", JVET-AF2025, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 2023-10-02
[0008] However, the method described in Non-Patent Document 1 has a problem in that MMVD prediction using GPM prediction is more complex to process than MMVD prediction without GPM prediction. Also, GPM mode using intra mode has a problem in that it is more complex to process than normal intra prediction.
[0009] A video decoding device according to one embodiment of the present invention is a video decoding device that decodes encoded data, and includes a motion information derivation unit that derives motion information of a target block and a prediction image generation unit that generates a prediction image of the target block by referring to the motion information. In a synthesis mode in which the prediction image generation unit generates two first prediction images for the target block and synthesizes a second prediction image of the target block from the two first prediction images using weights determined according to the geometric position of the target block, when the prediction image generation unit generates the first prediction image for at least one of the first prediction images using MMVD prediction, which shifts a value of a prediction motion vector derived in a merge mode using a prediction parameter that is a motion vector difference derived using two parameters, the options for a specific parameter that specifies candidates for the prediction parameter derived by the motion information derivation unit are different from the options for the specific parameter derived in an MMVD mode in which a prediction image of the target block is generated without using the synthesis mode.
[0010] A video decoding device according to one aspect of the present invention is a video decoding device that decodes encoded data, and is equipped with a motion information derivation unit that derives motion information of a target block and a prediction image generation unit that generates a prediction image of the target block by referring to the motion information, and is characterized in that in a synthesis mode in which the prediction image generation unit generates two first prediction images for the target block and synthesizes a second prediction image of the target block from the two first prediction images using weights determined according to the geometric position of the target block, when the prediction image generation unit generates at least one of the first prediction images using intra prediction, the options for the type of intra prediction mode derived by the motion information derivation unit are different from the options for the type of intra prediction mode derived in a mode in which a prediction image of the target block is generated without using the synthesis mode.
[0011] A video decoding device according to one aspect of the present invention is a video decoding device that decodes encoded data, and includes a motion information derivation unit that derives motion information of a target block, and a prediction image generation unit that generates a prediction image of the target block by referring to the motion information, wherein in a synthesis mode in which the prediction image generation unit generates two first prediction images for the target block and synthesizes a second prediction image of the target block from the two first prediction images using weights determined according to the geometric position of the target block, the filter coefficient options of a filter used to generate the prediction image derived by the motion information derivation unit are different from the filter coefficient options of a filter used to generate the prediction image derived in a mode in which a prediction image of the target block is generated without using the synthesis mode.
[0012] A video coding device according to one aspect of the present invention is a video coding device that codes a residual between a predicted image and an image to be coded, and includes a motion information derivation unit that derives motion information of a target block and a predicted image generation unit that generates a predicted image of the target block by referring to the motion information. In a synthesis mode in which the predicted image generation unit generates two first predicted images for the target block and synthesizes a second predicted image of the target block from the two first predicted images using weights determined according to the geometric position of the target block, when the predicted image generation unit generates the first predicted image for at least one of the first predicted images using MMVD prediction, which shifts a value of a predicted motion vector derived in a merge mode using a prediction parameter that is a motion vector difference derived using two parameters, the options for a specific parameter that specifies candidates for the prediction parameter derived by the motion information derivation unit are different from the options for the specific parameter derived in an MMVD mode in which a predicted image of the target block is generated without using the synthesis mode.
[0013] A video encoding device according to one aspect of the present invention is a video decoding device that decodes encoded data, and includes a motion information derivation unit that derives motion information of a target block, and a prediction image generation unit that generates a prediction image of the target block by referring to the motion information, wherein the prediction image generation unit generates two first prediction images for the target block, and in a synthesis mode in which a second prediction image of the target block is synthesized from the two first prediction images using weights determined according to the geometric position of the target block, when at least one of the first prediction images is generated using intra prediction, the options for the type of intra prediction mode are different from the options for the type of intra prediction mode derived in a mode in which a prediction image of the target block is generated without using the synthesis mode.
[0014] A video coding device according to one aspect of the present invention is a video coding device that codes a residual between a predicted image and an image to be coded, and is equipped with a motion information derivation unit that derives motion information of a target block and a predicted image generation unit that generates a predicted image of the target block by referring to the motion information, and is characterized in that in a synthesis mode in which the predicted image generation unit generates two first predicted images for the target block and synthesizes a second predicted image of the target block from the two first predicted images using weights determined according to the geometric position of the target block, the filter coefficient options of a filter used to generate the predicted image derived by the motion information derivation unit are different from the filter coefficient options of a filter used to generate the predicted image derived in a mode in which a predicted image of the target block is generated without using the synthesis mode.
[0015] A video encoding device according to one aspect of the present invention is a video encoding device that encodes a residual between a predicted image and an image to be encoded, and is equipped with a motion information derivation unit that derives motion information of a target block, and a predicted image generation unit that generates a predicted image of the target block by referring to the motion information, wherein the predicted image generation unit generates two first predicted images for the target block, and in a synthesis mode in which a second predicted image of the target block is synthesized from the two first predicted images using weights determined according to the geometric position of the target block, when at least one of the first predicted images is generated using intra prediction, the options for the type of intra prediction mode are different from the options for the type of intra prediction mode derived in a mode in which a predicted image of the target block is generated without using the synthesis mode.
[0016] A video coding device according to one aspect of the present invention is a video coding device that codes a residual between a predicted image and an image to be coded, and includes a motion information derivation unit that derives motion information of a target block, and a predicted image generation unit that generates a predicted image of the target block by referring to the motion information, wherein the predicted image generation unit generates two first predicted images for the target block, and in a synthesis mode in which a second predicted image of the target block is synthesized from the two first predicted images using weights determined according to a geometric position in the target block, the filter coefficient options of a filter used to generate the predicted images are different from the filter coefficient options of a filter used to generate the predicted image derived in a mode in which a predicted image of the target block is generated without using the synthesis mode.
[0017] According to one aspect of the present invention, in video encoding and decoding processes, by reducing the complexity of MMVD prediction processing using GPM prediction, it is possible to improve encoding efficiency while suppressing increases in processing volume and memory bandwidth.
[0018] FIG. 1 is a schematic diagram showing the configuration of an image transmission system according to this embodiment. FIG. 2 is a diagram showing a hierarchical structure of data of an encoded stream. FIG. 3 is a schematic diagram showing the configuration of a video decoding device. FIG. 4 is a block diagram showing the configuration of a video encoding device. FIG. 4 is a schematic diagram showing types of intra prediction modes (mode numbers). FIG. 5 is a schematic diagram showing the configuration of an inter prediction parameter derivation unit. FIG. 6 is a schematic diagram showing the configuration of an intra prediction parameter derivation unit. FIG. 7 is a schematic diagram showing the configuration of an inter prediction image generation unit. FIG. 8 is a diagram explaining merging and MMVD. FIG. 9 is a diagram explaining GPM prediction. FIG. 10 is a syntax diagram explaining coding parameters for GPM prediction. FIG. 11 is a syntax diagram explaining coding parameters for GPM prediction. FIG. 12 is a syntax diagram explaining coding parameters for GPM prediction. FIG. 13 is a diagram showing an example of limiting the number of MMVD candidates in MMVD prediction using GPM. FIG. 14 is a diagram showing another example of limiting the number of MMVD candidates in MMVD prediction using GPM. FIG. 15 is a diagram showing yet another example of limiting the number of MMVD candidates in MMVD prediction using GPM. FIG. 16 is a diagram showing yet another example of limiting the number of MMVD candidates in MMVD prediction using GPM. 1 is a diagram showing an example of limiting the number of intra-predictions when prediction is performed by a combination of MMVD prediction and intra-prediction using GPM; FIG. 2 is a diagram showing a table used for GPM prediction; FIG. 3 is a diagram showing an example of 8-directional candidates in MMVD candidates; and FIG. 4 is a diagram showing an example of 16-directional candidates in MMVD candidates.
[0019] First Embodiment Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0020] FIG. 1 is a schematic diagram showing the configuration of an image transmission system 1 according to this embodiment.
[0021] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding a target image, decodes the transmitted encoded stream, and displays the image. The image transmission system 1 includes a video encoding device (image encoding device) 11, a network 21, a video decoding device (image decoding device) 31, and a video display device (image display device) 41.
[0022] An image T is input to the video encoding device 11 .
[0023] The network 21 transmits the coded stream Te generated by the video coding device 11 to the video decoding device 31. The network 21 is the Internet, a wide area network (WAN, including 4G / 5G / 6G, etc.), a local area network (LAN), or a combination thereof. The network 21 is not necessarily limited to a bidirectional communication network, but may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Furthermore, the network 21 may be replaced by a storage medium on which the coded stream Te is recorded, such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark).
[0024] The video decoding device 31 decodes each of the coded streams Te transmitted over the network 21, and generates one or more decoded images Td.
[0025] The video display device 41 displays all or part of one or more decoded images Td generated by the video decoding device 31. The video display device 41 includes a display device such as a liquid crystal display or an organic EL (Electro-Luminescence) display. The display may be in the form of a stationary display, a mobile display, an HMD, or the like. Furthermore, if the video decoding device 31 has high processing power, it displays high-quality images, and if it has only low processing power, it displays images that do not require high processing power or display power.
[0026] <Operators> The operators used in this specification are listed below.
[0027] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, | is a bitwise OR, |= is the OR assignment operator, and || indicates logical sum.
[0028] x?y:z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0).
[0029] Clip3(a, b, c) is a function that clips c to a value between a and b. It returns a if c < a, b if c > b, and c otherwise (where a <= b).
[0030] ClipH(o, W, x) is a function that returns x if x < 0, x - o if x > W - 1, and x otherwise.
[0031] Clip1(x) is Clip3(0, (1 << BitDepth) - 1, x).
[0032] sign(a) is a function that returns 1 if a > 0, 1 if a == 0, and -1 if a < 0.
[0033] abs(a) is a function that returns the absolute value of a.
[0034] int(a) is a function that returns the integer value of a.
[0035] floor(a) is a function that returns the largest integer less than or equal to a.
[0036] ceil(a) is a function that returns the smallest integer greater than or equal to a. <000007
[0042] 2 is a diagram showing a hierarchical structure of data in an encoded stream Te. The encoded stream Te illustratively includes a sequence SEQ and multiple pictures constituting the sequence. Fig. 2 shows an encoded video sequence that defines the sequence SEQ, an encoded picture that defines a picture PICT, an encoded slice that defines a slice S, encoded slice data that defines slice data, a coding tree unit included in the coded slice data, and an encoding unit included in the coding tree unit.
[0043] (Encoded Video Sequence) An encoded video sequence defines a set of data that the video decoding device 31 refers to in order to decode a sequence SEQ to be processed. As shown in Fig. 2, the sequence SEQ includes a video parameter set, a sequence parameter set SPS (Sequence Parameter Set), a picture parameter set PPS (Picture Parameter Set), an adaptation parameter set (APS), a picture PICT, and supplemental enhancement information SEI (Supplemental Enhancement Information).
[0044] The video parameter set VPS specifies a set of coding parameters common to multiple videos composed of multiple layers, as well as a set of coding parameters related to multiple layers included in the video and each individual layer.
[0045] The sequence parameter set SPS defines a set of coding parameters that the video decoding device 31 references to decode the target sequence. For example, the width and height of a picture are defined. Note that there may be multiple SPSs. In this case, one of the multiple SPSs is selected from the PPS.
[0046] The picture parameter set PPS defines a set of coding parameters that the video decoding device 31 references to decode each picture in the target sequence. For example, the PPS includes a reference value for the quantization width used in decoding the picture (pic_init_qp_minus26) and a flag indicating the application of weighted prediction (weighted_pred_flag). Note that there may be multiple PPSs. In this case, one of the multiple PPSs is selected for each picture in the target sequence.
[0047] (Coded Picture) A coded picture defines a set of data that the video decoding device 31 references in order to decode a picture PICT to be processed. As shown in Fig. 2, the picture PICT includes slices 0 to NS-1 (NS is the total number of slices included in the picture PICT).
[0048] In the following description, when there is no need to distinguish between slices 0 to NS-1, the subscripts of the symbols may be omitted. This also applies to other data that are included in the coded stream Te described below and have subscripts.
[0049] (Encoded Slice) An encoded slice defines a set of data to be referenced by the video decoding device 31 in order to decode a slice S to be processed. As shown in Fig. 2, a slice includes a slice header and slice data.
[0050] The slice header includes a group of coding parameters that the video decoding device 31 refers to in order to determine a decoding method for the current slice. Slice type designation information (slice_type) that designates the slice type is an example of a coding parameter included in the slice header.
[0051] Slice types that can be specified by the slice type specification information include (1) an I slice that uses only intra prediction during encoding, (2) a P slice that uses unidirectional prediction or intra prediction during encoding, and (3) a B slice that uses unidirectional prediction, bidirectional prediction, or intra prediction during encoding. Note that inter prediction is not limited to uni-prediction or bi-prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referred to as a P or B slice, it refers to a slice including a block that can use inter prediction.
[0052] Note that the slice header may include a reference to the picture parameter set PPS (pic_parameter_set_id).
[0053] (Encoded Slice Data) The encoded slice data defines a set of data that the video decoding device 31 references in order to decode the slice data to be processed. As shown in Fig. 2, the slice data includes a CTU. A CTU is a block of a fixed size (e.g., 64x64, 128x128) that constitutes a slice.
[0054] (Coding Trees and Coding Units) A CTU is divided into coding units (CUs), which are the basic units of the coding process, by recursive quad-tree division (QT (Quad Tree) division), binary tree division (BT (Binary Tree) division), or ternary tree division (TT (Ternary Tree) division). BT division and TT division are collectively called multi-tree division (MT (Multi Tree) division). A node in the tree structure obtained by recursive quad-tree division is called a coding tree (CT). The intermediate nodes of a quad-tree, binary tree, or ternary tree are coding trees, and the CTU itself is defined as the top-level coding node. The lowest-level coding tree is defined as a coding unit (CU).
[0055] Different trees (separate trees or dual trees) may be used for luminance and chrominance. The tree type is indicated by treeType. For example, when using a common tree for luminance (Y, cIdx=0) and chrominance (Cb / Cr, cIdx=1,2), the common single tree is indicated by treeType=SINGLE_TREE. When using two different trees (DUAL trees) for luminance and chrominance, the luminance tree is indicated by treeType=DUAL_TREE_LUMA and the chrominance tree is indicated by treeType=DUAL_TREE_CHROMA.
[0056] A CU is composed of a CU header CUH, prediction parameters, transformation parameters, quantized transformation coefficients, etc. The CU header specifies a prediction mode, etc.
[0057] The prediction process may be performed in units of CUs, or in units of sub-CUs obtained by further dividing a CU. When the sizes of a CU and a sub-CU are the same, there is one sub-CU in the CU. When the size of a CU is larger than the size of a sub-CU, the CU is divided into sub-CUs. For example, when a CU is 8x8 and a sub-CU is 4x4, the CU is divided into four sub-CUs, divided horizontally in half and vertically in half.
[0058] Prediction types (prediction modes) include intra prediction (MODE_INTRA), inter prediction (MODE_INTER), and intra block copy (MODE_IBC). Intra prediction is prediction within the same picture, while inter prediction refers to prediction processing performed between different pictures (for example, between display times or between layer images).
[0059] The transformation and quantization processes are performed in units of CUs, but the quantized transformation coefficients may be entropy coded in units of sub-blocks such as 4x4.
[0060] (Prediction Parameters) A predicted image is derived from prediction parameters associated with a block. Prediction parameters include intra-prediction and inter-prediction parameters.
[0061] (Prediction Parameters for Inter Prediction) The prediction parameters for inter prediction will be described. The inter prediction parameters are composed of prediction list usage flags predFlagL0 and predFlagL1, reference picture indices refIdxL0 and refIdxL1, and motion vectors mvL0 and mvL1, and are decoded from encoded data in units of CU. predFlagL0 and predFlagL1 are flags indicating whether a reference picture list (L0 list, L1 list) is used, and when the value is 1, the corresponding reference picture list is used. Note that in this specification, when the term "flag indicating whether XX is true" is used, a flag other than 0 (e.g., 1) indicates XX, and 0 indicates non-XX, and in logical negation, logical product, etc., 1 is treated as true and 0 is treated as false (the same applies below). However, in actual devices and methods, other values may be used as true and false values.
[0062] To derive inter prediction parameters, for example, the following syntax elements are decoded on a CU-by-CU basis: skip flag skip_flag, merge flag merge_flag (general_merge_flag), merge index merge_idx, merge_subblock_flag, regular_merge_flag, ciip_flag, merge_gpm_partition_idx, merge_gpm_idx0, merge_gpm_idx1, inter_pred_idc, reference picture index refIdxLX, mvp_LX_idx, difference vector mvdLX, motion vector precision flag amvr_flag, and motion vector precision index amvr_precision_idx. merge_subblock_flag is a flag indicating whether to use subblock-based inter prediction. regular_merge_flag is a flag indicating whether to use normal merge mode or MMVD. ciip_flag is a flag indicating whether to use CIIP (Combined Inter-picture merge and Intra-picture Prediction) mode. merge_gpm_partition_idx is an index indicating the partition shape in GPM mode. merge_gpm_idx0 and merge_gpm_idx1 are indices indicating the merge index in GPM mode. inter_pred_idc is an inter prediction identifier for selecting a reference picture to be used in AMVP mode. mvp_LX_idx is a predicted vector index for deriving a motion vector.
[0063] (Reference Picture List) The reference picture list is a list of reference pictures stored in the reference picture memory 306. In each CU, refIdxLX specifies which picture in the reference picture list RefPicList[X] (X = 0 or 1) to actually reference. Note that LX is a notation method used when there is no distinction between L0 prediction and L1 prediction; hereinafter, parameters for the L0 list and parameters for the L1 list will be distinguished by replacing LX with L0 or L1.
[0064] (Merge Prediction and AMVP Prediction) Prediction parameter decoding (encoding) methods include merge prediction mode (merge mode) and AMVP (Advanced Motion Vector Prediction, adaptive motion vector prediction) mode, and general_merge_flag is a flag for distinguishing between them. Merge mode is a prediction mode that omits some or all of the motion vector difference, and derives the prediction list usage flag predFlagLX, reference picture index refIdxLX, and motion vector mvLX from the encoded data, instead of including them in the encoded data, and instead derives them from prediction parameters of already processed neighboring blocks, etc. AMVP mode is a mode that includes inter_pred_idc, refIdxLX, and mvLX in the encoded data. Note that mvLX is encoded as mvp_LX_idx, which identifies the prediction vector mvpLX, and the difference vector mvdLX. The general term for prediction modes that omit or simplify motion vector differences is called the general merge mode, and general_merge_flag can be used to select between general merge mode and AMVP prediction. general_merge_flag is decoded from the coded data if the current block is not in skip mode, and is set to 1 if the current block is in skip mode.
[0065] If general_merge_flag is 1, regular_merge_flag may be transmitted separately. If regular_merge_flag is 1, normal merge mode or MMVD may be selected, and otherwise CIIP mode or GPM mode may be selected. CIIP mode generates a predicted image by weighted sum of inter-predicted image and intra-predicted image. GPM mode generates a predicted image by combining two regions separated by a line segment within the target CU.
[0066] inter_pred_idc is a value indicating the type and number of reference pictures, and takes one of the values PRED_L0, PRED_L1, or PRED_BI. PRED_L0 and PRED_L1 indicate uni-prediction using one reference picture managed in the L0 list and L1 list, respectively. PRED_BI indicates bi-prediction using two reference pictures managed in the L0 list and L1 list.
[0067] The merge_idx is an index indicating which prediction parameter from among prediction parameter candidates (merge candidates) derived from blocks for which processing has been completed is to be used as the prediction parameter for the current block.
[0068] (Motion Vector) mvLX indicates the amount of shift between blocks on two different pictures. The predicted vector and differential vector related to mvLX are called mvpLX and mvdLX, respectively.
[0069] (Inter prediction identifier inter_pred_idc and prediction list use flag predFlagLX) The relationship between inter_pred_idc, predFlagL0, and predFlagL1 is as follows, and they can be converted into each other.
[0070] inter_pred_idc = (predFlagL1<<1)+predFlagL0 predFlagL0 = inter_pred_idc & 1 predFlagL1 = inter_pred_idc >> 1 Note that, as the inter prediction parameters, a prediction list usage flag or an inter prediction identifier may be used. Furthermore, a determination using the prediction list usage flag may be replaced with a determination using the inter prediction identifier. Conversely, a determination using the inter prediction identifier may be replaced with a determination using the prediction list usage flag.
[0071] (Determination of Bi-Prediction biPred) The flag biPred indicating whether bi-prediction is performed can be derived based on whether two prediction list usage flags are both 1. For example, it can be derived using the following formula.
[0072] biPred = (predFlagL0==1 && predFlagL1==1) Alternatively, biPred can also be derived based on whether the inter prediction identifier is a value indicating the use of two prediction lists (reference pictures). For example, it can be derived using the following formula:
[0073] biPred = (inter_pred_idc==PRED_BI) ? 1 : 0 (Configuration of Video Decoding Apparatus) The configuration of the video decoding apparatus 31 according to this embodiment will be described. FIG.
[0074] The video decoding device 31 includes an entropy decoding unit 301, a parameter decoding unit (prediction parameter decoding device) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a prediction image generation unit (prediction image generation device) 308, an inverse quantization and inverse transform unit 311, an adder 312, and a prediction parameter derivation unit 320. Note that the video decoding device 31 may also be configured without including the loop filter 305, in accordance with the video coding device 11 described below.
[0075] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit), and the CU decoding unit 3022 further includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, PPS, and APS, and slice headers (slice information) from the coded data. The CT information decoding unit 3021 decodes the CT from the coded data. The CU decoding unit 3022 decodes the CU from the coded data. The TU decoding unit 3024 decodes the CU from the coded data.
[0076] When a TU includes a prediction error, the TU decoding unit 3024 decodes QP update information (quantization correction value) and quantization prediction error (residual_coding) from the coded data. The QP update information is a difference value from a quantization parameter predicted value qPpred, which is a predicted value of the quantization parameter QP.
[0077] The TU decoding unit 3024 decodes the QP update information and the quantized prediction error from the coded data when the mode is other than the skip mode (skip_mode==0). More specifically, when skip_mode==0, the TU decoding unit 3024 decodes the flag cu_cbp indicating whether or not the current block includes a quantized prediction error, and decodes the quantized prediction error when cu_cbp is 1. When cu_cbp does not exist in the coded data, it is derived as 0.
[0078] The TU decoding unit 3024 decodes an index mts_idx indicating a transform base from the coded data. The TU decoding unit 3024 also decodes an index stIdx indicating the use of a secondary transform and the transform base from the coded data. stIdx indicates no application of a secondary transform when it is 0, indicates one transform of a set (pair) of secondary transform bases when it is 1, and indicates the other transform of the pair when it is 2.
[0079] The TU decoding unit 3024 may also decode a sub-block transform flag cu_sbt_flag. When cu_sbt_flag is 1, the CU is divided into multiple sub-blocks and the residual of only one specific sub-block is decoded. The TU decoding unit 3024 may also decode a flag cu_sbt_quad_flag indicating whether the number of sub-blocks is 4 or 2, cu_sbt_horizontal_flag indicating the division direction, and cu_sbt_pos_flag indicating a sub-block that includes a non-zero transform coefficient.
[0080] The predicted image generating unit 308 includes an inter predicted image generating unit 309 and an intra predicted image generating unit 310 .
[0081] The prediction parameter derivation unit 320 includes an inter prediction parameter derivation unit 303 and an intra prediction parameter derivation unit.
[0082] In the following, an example will be described in which CTUs and CUs are used as processing units, but this is not limiting and processing may be performed in sub-CU units. Alternatively, CTUs and CUs may be read as blocks, and sub-CUs as sub-blocks, and processing may be performed in block or sub-block units.
[0083] The entropy decoding unit 301 performs entropy decoding on the externally input coded stream Te to decode individual codes (syntax elements). Entropy coding can be performed in two ways: one is to perform variable-length coding of syntax elements using a context (probability model) adaptively selected according to the type of syntax element and the surrounding circumstances, and the other is to perform variable-length coding of syntax elements using a predetermined table or formula.
[0084] The entropy decoding unit 301 performs entropy decoding on the externally input coded stream Te to decode individual codes (syntax elements). Entropy coding can be divided into two types: variable-length coding of syntax elements using a context (probability model) adaptively selected according to the type of syntax element and surrounding conditions, and variable-length coding of syntax elements using a predefined table or formula. The former, CABAC (Context Adaptive Binary Arithmetic Coding), stores the CABAC state of the context (the type of most probable symbol (0 or 1) and a probability state index pStateIdx that specifies the probability) in memory. The entropy decoding unit 301 initializes all CABAC states at the beginning of a segment (tile, CTU row, slice). The entropy decoding unit 301 converts the syntax elements into a binary string (bin string) and decodes each bit of the bin string. When a context is used, a context index (ctxInc) is derived for each bit of the syntax element, the bit is decoded using the context, and the CABAC state of the used context is updated. Bits that do not use a context are decoded with equal probability (EP, bypass), and the ctxInc derivation and CABAC state are omitted. The decoded syntax elements include prediction information for generating a predicted image and a prediction error for generating a difference image.
[0085] The entropy decoding unit 301 outputs the decoded code to the parameter decoding unit 302. The decoded code includes, for example, a prediction mode predMode, general_merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX, amvr_flag, amvr_precision_idx, ciip_flag, merge_gpm_partition_idx, gpm_mmvd_flag0, gpm_mmvd_distance_idx0, gpm_mmvd_direction_idx0, gpm_mmvd_enable_flag1, gpm_mmvd_distance_idx1, gpm_mmvd_direction_idx1, merge_gpm_idx0, merge_gpm_idx1, etc. Control of which code to decode is performed based on an instruction from the parameter decoding unit 302.
[0086] The loop filter 305 is a filter provided in the encoding loop that removes block distortion and ringing distortion to improve image quality. The loop filter 305 applies filters such as a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF) to the decoded image of the CU generated by the adder 312.
[0087] The reference picture memory 306 stores the decoded image of the CU generated by the adder 312 in a predetermined location for each current picture and current CU.
[0088] The prediction parameter memory 307 stores prediction parameters at a predetermined location for each CTU or CU to be decoded. Specifically, the prediction parameter memory 307 stores the parameters decoded by the parameter decoding unit 302 and the prediction mode predMode decoded by the entropy decoding unit 301.
[0089] The predicted image generation unit 308 receives input of predMode, prediction parameters, etc. The predicted image generation unit 308 also reads a reference picture from the reference picture memory 306. The predicted image generation unit 308 generates a predicted image of a block or sub-block using the prediction parameters and the read reference picture (reference block) in the prediction mode indicated by predMode. Here, the reference block is a set of pixels on the reference picture (usually rectangular, and therefore called a block), and is an area referenced to generate a predicted image.
[0090] (Configuration of Inter Prediction Parameter Derivation Unit) FIG. 6 is a schematic diagram showing an example of the configuration of an inter prediction parameter derivation unit. The inter prediction parameter derivation unit 303 derives inter prediction parameters by referring to prediction parameters stored in the prediction parameter memory 307 based on the syntax elements decoded by the parameter decoding unit 302. The inter prediction parameter derivation unit 303 also outputs the inter prediction parameters to the inter prediction image generation unit 309 and the prediction parameter memory 307. The inter prediction parameter derivation unit 303 and its internal elements, namely the AMVP prediction parameter derivation unit 3032, the merge prediction parameter derivation unit 3036, the MMVD prediction unit 30376, the GPM prediction unit 30377, the affine prediction unit 30372, the DMVR unit 30537, and the MV addition unit 3038, are means common to the video encoding device and the video decoding device, and may therefore be collectively referred to as a motion vector derivation unit (motion vector derivation device).
[0091] If general_merge_flag is 1, that is, if it indicates merge prediction mode, merge_idx is decoded and output to the merge prediction parameter derivation unit 3036 .
[0092] When general_merge_flag is 0, that is, when it indicates the AMVP prediction mode, the AMVP prediction parameter derivation unit 3032 derives mvpLX from inter_pred_idc, refIdxLX, or mvp_LX_idx.
[0093] If sym_mvd_flag is 0, that is, if it does not indicate SMVD (Symmetric Motion Vector Difference) mode, the AMVP prediction parameter derivation unit 3032 uses mvdLX corresponding to L0 and L1 decoded by the parameter decoding unit 302.
[0094] When sym_mvd_flag is 1, i.e., when SMVD mode is indicated, the parameter decoding unit 302 derives refIdxLX and mvdLX only for L0, and the AMVP prediction parameter derivation unit 3032 derives refIdxL1 and MvdL1 so that the MVD is temporally and spatially symmetric with the MVD of L0. MvdL1[0] = -MvdL0[0] MvdL1[1] = -MvdL0[1] (Merge Prediction) The merge prediction parameter derivation unit 3036 includes a merge candidate derivation unit 30361 and a merge candidate selection unit 30362. Note that merge candidates include prediction parameters (predFlagLX, mvLX, refIdxLX) and are stored in a merge candidate list. Merge candidates stored in the merge candidate list are assigned indices according to a predetermined rule.
[0095] The merge candidate derivation unit 30361 derives merge candidates by directly using the motion vectors and refIdxLX of the decoded adjacent blocks. Alternatively, the merge candidate derivation unit 30361 may apply a spatial merge candidate derivation process, a temporal merge candidate derivation process, or the like, which will be described later.
[0096] As a spatial merge candidate derivation process, the merge candidate derivation unit 30361 reads prediction parameters stored in the prediction parameter memory 307 according to a predetermined rule and sets them as merge candidates. Figure 9 is a diagram explaining merging and MMVD. For example, prediction parameters for the following positions A1, B1, B0, A0, and B2 shown in Figure 9 are read.
[0097] A1: (xCb-1, yCb+cbHeight-1) B1: (xCb+cbWidth-1, yCb-1) B0: (xCb+cbWidth, yCb-1) A0: (xCb-1, yCb+cbHeight) B2: (xCb-1, yCb-1) The top left coordinates of the target block are (xCb, yCb), the width is cbWidth, and the height is cbHeight.
[0098] As a temporal merge derivation process, the merge candidate derivation unit 30361 may read the prediction parameters of the lower right CBR of the target block or the block C in the reference image including the center coordinates from the prediction parameter memory 307, set it as a merge candidate Col, and store it in the merge candidate list mergeCandList[].
[0099] The order in which mergeCandList[] is stored is, for example, spatial merge candidates (B1, A1, B0, A0, B2), followed by temporal merge candidate Col. Reference blocks that are unavailable (e.g., blocks that are intra-predicted) are not stored in the merge candidate list. i = 0 if(availableFlagB1) mergeCandList[i++] = B1 if(availableFlagA1) mergeCandList[i++] = A1 if(availableFlagB0) mergeCandList[i++] = B0 if(availableFlagA0) mergeCandList[i++] = A0 if(availableFlagB2) mergeCandList[i++] = B2 if(availableFlagCol) mergeCandList[i++] = Col. Furthermore, history merge candidate HmvpCand, pairwise average candidate avgCand, and zero merge candidate zeroCandm may be added to mergeCandList[] and used. for (j=1, j <= numHmvpCand; j++) if( i < MaxNumMergeCand && numHmvpCand > 0) mergeCandList[ i++ ] = HmvpCand[numHmvpCand - j] if( i < MaxNumMergeCand && i > 1 ) mergeCandList[ i++ ] = avgCand if( i < MaxNumMergeCand ) mergeCandList[ i++ ] = zeroCand Furthermore, non-adjacent spatial merge candidates may be added to mergeCandList[] and used. Unlike the spatial merge candidates (B1, A1, B0, A0, B2), non-adjacent spatial merge candidates are merge candidates that use prediction parameters from positions that are not adjacent to the target block.
[0100] The merge candidate selection unit 30362 selects a merge candidate N indicated by merge_idx from among the merge candidates included in the merge candidate list using the following formula.
[0101] N = mergeCandList[merge_idx] where N is a label indicating a merge candidate, such as A1, B1, B0, A0, B2, Col, etc. The motion information of the merge candidate indicated by label N is indicated by (mvLXN[0], mvLXN[1]), predFlagLXN, and refIdxLXN.
[0102] The merge candidate selection unit 30362 stores the inter prediction parameters of the selected merge candidate in the prediction parameter memory 307 and outputs them to the inter prediction image generation unit 309.
[0103] (MMVD prediction unit 30376) When mmvd_flag is 1, the MMVD prediction unit 30376 decodes the further restricted difference vector and performs MMVD (Merge with Motion Vector Difference) processing. The MMVD processing adds a difference vector restricted to a predetermined distance and a predetermined direction to the motion vector of the merge candidate. In MMVD mode, the value range of the difference vector is restricted to a predetermined distance (e.g., 6 ways, 8 ways, etc.) and a predetermined direction (e.g., 4 directions, 8 directions, 16 directions, etc.), thereby efficiently deriving a motion vector.
[0104] The MMVD prediction unit 30376 selects the central vector mvLXN[ ] using base_candidate_idx.
[0105] N = mergeCandList[base_candidate_idx] The MMVD prediction unit 30376 derives the base distance (mvdUnit[0], mvdUnit[1]) from mmvd_direction_idx and a table (DirectionTable), and derives the distance MmvdDistance from distance_idx and a table (DistanceTable). Note that the DirectionTable may use a table DirectionTableN with N elements.
[0106] DirectionTable4x[] = { 1, -1, 0, 0} DirectionTable4y[] = { 0, 0, 1, -1} mvdUnit[0] = DirectionTable4x[[direction_idx] mvdUnit[1] = DirectionTable4y[direction_idx] MmvdDistance = DistanceTable[distance_idx] Here, DistanceTable may be {1, 2, 4, 8, 16, 32, 64, 128, 256}. It may also be derived by a left shift operation without referencing the table.
[0107] MmvdDistance = 1<< distance_idx Figure 21 is a diagram showing an example of eight direction candidates in MMVD candidates. When there are eight directions as shown in Figure 21A, they may be derived as follows.
[0108] DirectionTable8x[] = { 1, -1, 0, 0, 1, -1, 1,-1} DirectionTable8y[] = { 0, 0, 1, -1, 1, -1,-1, 1} mvdUnit[0] = DirectionTable8x[direction_idx] mvdUnit[1] = DirectionTable8y[direction_idx] Figure 22 is a diagram showing an example of 16 direction candidates in MMVD candidates. When there are 16 directions as shown in Figure 21B, they may be derived as follows.
[0109] DirectionTable16x[] = { 1, -1, 0, 0, 1, -1, 1, -1, 2, -2, 2, -2, 1, 1, -1, -1}; DirectionTable16y[] = { 0, 0, 1, -1, 1, -1, -1, 1, 1, 1, -1, -1, 2, -2, 2, -2}; mvdUnit[0] = DirectionTable16x[direction_idx]; mvdUnit[1] = DirectionTable16y[direction_idx]; Note that in MMVD and GPM-MMVD, it is derived by replacing direction_idx with mmvd_direction_idx and gpm_mmvd_direction_idxKK, and distance_idx with mmvd_distance_idx and gpm_mmvd_distance_idxKK.
[0110] The MMVD prediction unit 30376 derives the differential vector refineMvdLX[] using the product of (mvdUnit[0], mvdUnit[1]) and MmvdDistance.
[0111] refineMvdL0[0] = (MmvdDistance << shiftMMVD) * mvdUnit[0]; refineMvdL0[1] = (MmvdDistance << shiftMMVD) * mvdUnit[1]; refineMvdL1[0] = -(MmvdDistance << shiftMMVD) * mvdUnit[0]; refineMvdL1[1] = -(MmvdDistance << shiftMMVD) * mvdUnit[1]; Here, shiftMMVD is a value that adjusts the magnitude of the differential vector so as to match the accuracy MVPREC of the motion vector in the motion compensation unit 3091 (interpolation unit). For example, shiftMMVD = 2 may be used. Also, in the case of 1 / 16 accuracy, shiftMMVD = 1, in the case of 1 / 4 pel accuracy, shiftMMVD = 2, and in the case of 1 pei accuracy, shiftMMVD = 4 may be used.
[0112] Scaling may be performed according to the POC distance between the reference picture and the target picture.
[0113] MmvdOffset[0] = (MmvdDistance << shiftMMVD) * mvdUnit[0] MmvdOffset[1] = (MmvdDistance << shiftMMVD) * mvdUnit[1] If predFlagL0 == 1 and predFlagL1 == 1, apply the following. currPocDiffL0 = DiffPicOrderCnt( currPic, RefPicList[ 0 ][ refIdxL0 ] ) currPocDiffL1 = DiffPicOrderCnt( currPic, RefPicList[ 1 ][ refIdxL1 ] ) Here, DiffPicOrderCnt(Pic1,Pic2) is a function that returns the difference in the temporal information (e.g., POC, PictureOrderCount) between Pic1 and Pic2. If currPocDiffL0 == currPocDiffL1, apply the following. refineMvdL0[ 0 ] = MmvdOffset[ 0 ] refineMvdL0[ 1 ] = MmvdOffset[ 1 ] refineMvdL1[ 0 ] = MmvdOffset[ 0 ] refineMvdL1[ 1 ] = MmvdOffset[ 1 ] Otherwise, if Abs( currPocDiffL0 ) >= Abs( currPocDiffL1 ), apply the following. refineMvdL0[ 0 ] = MmvdOffset[ 0 ] refineMvdL0[ 1 ] = MmvdOffset[ 1 ] If RefPicList[ 0 ][ refIdxL0 ] is marked as "used for long-term reference" and not marked as such for RefPicList[ 1 ][ refIdxL1 ], and not marked as "used for long-term reference", apply the following.td = Clip3( -128, 127, currPocDiffL0 ) tb = Clip3( -128, 127, currPocDiffL1 ) tx = ( 16384 + ( Abs( td ) >> 1 ) ) / td distScaleFactor = Clip3( -4096, 4095, ( tb * tx + 32 ) >> 6 ) refineMvdL1[ 0 ] = Clip3( -217, 217 - 1, (distScaleFactor * refineMvdL0[ 0 ] + 128 - ( distScaleFactor * refineMvdL0[ 0 ] >= 0 ) ) >> 8 ) refineMvdL1[ 1 ] = Clip3( -217, 217 - 1, (distScaleFactor * refineMvdL0[ 1 ] + 128 - ( distScaleFactor * refineMvdL0[ 1 ] >= 0 ) ) >> 8 ) Otherwise apply the following: refineMvdL1[0] = Sign(currPocDiffL0) == Sign(currPocDiffL1) ? refineMvdL0[0]:-refineMvdL0[0] refineMvdL1[1] = Sign(currPocDiffL0) == Sign(currPocDiffL1) ? refineMvdL0[1]:-refineMvdL0[1] Otherwise (Abs( currPocDiffL0 ) < Abs( currPocDiffL1 )), apply the following: refineMvdL1[ 0 ] = MmvdOffset[ 0 ] refineMvdL1[ 1 ] = MmvdOffset[ 1 ] If RefPicList[ 0 ][ refIdxL0 ] is not marked "used for long-term reference" and RefPicList[ 1 ][ refIdxL1 ] is not marked "used for long-term reference", then the following applies.td = Clip3( -128, 127, currPocDiffL1 ) tb = Clip3( -128, 127, currPocDiffL0 ) tx = ( 16384 + ( Abs( td ) >> 1 ) ) / td distScaleFactor = Clip3( -4096, 4095, ( tb * tx + 32 ) >> 6 ) refineMvdL0[ 0 ] = Clip3( -217, 217 - 1, ( distScaleFactor * refineMvdL1[ 0 ] +128 - ( distScaleFactor * refineMvdL1[ 0 ] >= 0 ) ) >> 8 ) refineMvdL0[ 1 ] = Clip3( -217, 217 - 1, ( distScaleFactor * refineMvdL1[ 1 ] +128 - ( distScaleFactor * refineMvdL1[ 1 ] >= 0 ) ) >> 8 ) Else apply the following: refineMvdL0[0] = Sign(currPocDiffL0) == Sign(currPocDiffL1) ? refineMvdL1[0]:-refineMvdL1[0] refineMvdL0[1] = Sign(currPocDiffL0) == Sign(currPocDiffL1) ? refineMvdL1[1]:-refineMvdL1[1] Else if ( predFlagL0 == 1 or predFlagL1 == ), apply the following. refineMvdLX[ 0 ] = ( predFlagLX == 1 ) ? MmvdOffset[ 0 ] : 0 refineMvdLX[ 1 ] = ( predFlagLX == 1 ) ? MmvdOffset[ 1 ] : 0 Finally, the MMVD prediction unit 30376 derives the motion vector of the MMVD merge candidate from refineMvdLX and the center vector mvLXN as follows:mvL0[ 0 ] = mvL0N[ 0 ] + refineMvdL0[0] mvL0[ 1 ] = mvL0N[ 1 ] + refineMvdL0[1] mvL1[ 0 ] = mvL1N[ 0 ] + refineMvdL1[0] mvL1[ 1 ] = mvL1N[ 1 ] + refineMvdL1[1] When the flag tm_merge_flag used to indicate whether template matching is used is true, the MMVD prediction unit 30376 may calculate a cost tempCost (template matching cost) for each MMVD candidate and sort the MMVD candidates in ascending order of cost to derive the final MMVD candidate list mmvdLUT. All elements of the mmvdLUT may be derived in advance first, and the mmvdLUT may be derive by swapping the mmvdLUT candidates using the costs. Alternatively, the MMVD prediction unit 30376 may derive the MMVD candidate list mmvdLUT by starting with an empty MMVD candidate list mmvdLUT and adding MMVD candidates to the list in ascending order of cost. Furthermore, when adding an MMVD candidate to the list, the MMVD prediction unit 30376 may skip adding the MMVD candidate if the addition position exceeds the maximum number of candidates maxNumMmvdLUT (e.g., 12) in the list.
[0114] (DMVR) This section describes the DMVR (Decoder-side Motion Vector Refinement) process performed by the DMVR unit 30375. The DMVR process is a process in which, when the target CU is in merge mode (when merge_flag is 1 or when the skip flag skip_flag is 1), the motion vectors mvL0 and mvL1 of the target CU derived by the merge prediction unit 30374 and the MMVD prediction unit 30376 are corrected using predicted images derived from motion vectors corresponding to two reference pictures.
[0115] DMVR divides a target block into sub-blocks of 16 pixels each. However, if the height or width is less than 16 pixels, it does not divide the block. For each sub-block, mvL0 and mvL1 are displaced by the motion vector displacements dMvL0 and dMvL1, respectively, to derive an L0 predicted image (mvL0 + dMvL0) and an L1 predicted image (mvL1 + dMvL1). Then, dMvL0 and dMvL1 are calculated to minimize the sum of absolute differences (SAD) between these predicted images. Note that when the components of the motion vector displacements dMvL0 and dMvL1 are (dMvL0[0], dMvL0[1]) and (dMvL1[0], dMvL1[1]), respectively, dMvL1[0] = -dMvL0[0] and dMvL1[1] = -dMvL0[1].
[0116] dMvL1[0] = -dMvL0[0] dMvL1[1] = -dMvL0[1] When dmvrFlag is 1, the DMVR unit 30375 adds the derived differential vector dmvLX to the motion vector mvLX input from the merge prediction unit 30374 to derive the interpolation motion vector refMvLX (X = 0..1).
[0117] The interpolation motion vector refMvLX is used in the motion compensation unit 3091 (interpolation unit). The DMVR unit 30375 outputs mvLX to the predicted image generation unit 308 and the prediction parameter memory 307.
[0118] refMvLX[0] = mvLX[0] + dMvLX[0] refMvLX[1] = mvLX[1] + dMvLX[1] When dmvrFlag=0, the DMVR unit 30375 sets the motion vector mvLX input from the merge prediction unit 30374 to refMvLx and outputs it to the motion compensation unit 3091 (interpolation unit) (X=0..1).
[0119] refMvLX[0] = mvLX[0] refMvLX[1] = mvLX[1] (TM prediction) The TM prediction unit performs processing in TM (Template Matching) mode. In TM prediction, the prediction parameters (motion information) of the current block are corrected based on the matching cost (error value) of the template region. The upper and left adjacent regions of the current block are used as the template region, and the position where the error value is minimum is searched for within the periphery of the initial MV (for example, within a range of ±8 pixels), and the motion information is updated to that position.
[0120] In the case of the AMVP prediction mode, the MVP candidate with the smallest template matching error value at that position is selected and used as the initial MV.
[0121] In the case of merge prediction mode, a merge candidate indicated by the merge index merge_idx is used as the initial MV, and similar correction is performed.
[0122] When TM prediction is applied to a bi-predictive block, an error value is derived for each of L0 and L1, and the other is further corrected with the MV with the smaller cost.
[0123] (AMVP Prediction) The AMVP prediction parameter derivation unit 3032 includes a vector candidate derivation unit 3033 and a vector candidate selection unit 3034. The vector candidate derivation unit 3033 derives predictor vector candidates from the motion vectors of decoded adjacent blocks stored in the prediction parameter memory 307 based on refIdxLX, and stores the candidates in a predictor vector candidate list mvpListLX[ ].
[0124] The vector candidate selection unit 3034 selects, as mvpLX, the motion vector mvpListLX[mvp_LX_idx] indicated by mvp_LX_idx from among the predicted vector candidates in mvpListLX[ ]. The vector candidate selection unit 3034 outputs the selected mvpLX to the MV addition unit 3038.
[0125] (MV Addition Unit) The MV addition unit 3038 calculates mvLX by adding the mvpLX input from the AMVP prediction parameter derivation unit 3032 and the decoded mvdLX. The MV addition unit 3038 outputs the calculated mvLX to the inter predicted image generation unit 309 and the prediction parameter memory 307.
[0126] mvLX[0] = mvpLX[0] + mvdLX[0] mvLX[1] = mvpLX[1] + mvdLX[1] The MV adder derives a shift value amvrShift based on the values of amvr_flag and amvr_precision_idx, and may use it to change the precision of the derived mvLX. amvr_precision_idx is a syntax element that switches the precision of the vector together with amvr_flag. For example, in AMVP mode, when amvr_flag==0, amvr_flag==&&amvr_precision_idx=0, amvr_flag==&&amvr_precision_idx=1, amvr_flag==1&& amvr_precision_idx=2, amvrShift=2, 3, 4, 6 are set, respectively, to switch between 1 / 4 pixel, 1 / 2 pixel, 1 pixel, and 4 pixel precision.
[0127] (Inter-prediction image generation unit 309) When predMode indicates inter-prediction, the inter-prediction image generation unit 309 generates a prediction image of a block or sub-block by inter-prediction using the inter-prediction parameters and reference picture input from the inter-prediction parameter derivation unit 303.
[0128] 8 is a schematic diagram showing the configuration of an inter-prediction image generation unit 309 included in the prediction image generation unit 308 according to this embodiment. The inter-prediction image generation unit 309 includes a motion compensation unit (prediction image generation device) 3091 and a synthesis unit 3095. The synthesis unit 3095 includes an intra-inter synthesis unit 30951, a GPM synthesis unit 30952, a weighted prediction unit 30953, a BDOF processing unit 30954, and a BCW processing unit 30955. It may also include an LIC processing unit 30956 and an OBMC processing unit 30957, which are not shown.
[0129] (Motion Compensation) The motion compensation unit 3091 (interpolated image generation unit 3091) generates an interpolated image (motion-compensated image) by reading a reference block from the reference picture memory 306 based on the inter-prediction parameters (predFlagLX, refIdxLX, mvLX) input from the inter-prediction parameter derivation unit 303. The reference block is a block located at a position shifted by mvLX from the position of the current block on the reference picture RefPicLX[X][refIdxLX] specified by refIdxLX. Here, if mvLX does not have integer precision, a filter for generating pixels at decimal positions called a motion compensation filter is applied to generate an interpolated image.
[0130] The motion compensation unit 3091 derives the integer position (xInt, yInt) and phase (xFrac, yFrac) corresponding to the top left coordinates (xPb, yPb) of a block of size bW*bH, the coordinates within the prediction block (xL, yL), and the motion vector (mvLX[0], mvLX[1]) using the following formula (MC-P1).
[0131] xInt = xPb+(mvLX[0]>>(log2MVPREC))+xL xFrac = mvLX[0]&(MVPREC-1) yInt = yPb+(mvLX[1]>>(log2MVPREC))+yL yFrac = mvLX[1]&(MVPREC-1) Here, MVPREC indicates the accuracy of mvLX (1 / MVPREC pixel accuracy), log2MVPREC=log2(MVPREC), x=0...bW-1, y=0...bH-1. For example, MVPREC=16.
[0132] The motion compensation unit 3091 derives the temporary image temp[][] by performing horizontal interpolation on the reference picture refImg using an interpolation filter. In the following, Σ is the sum over k, k=0..NTAP-1, mcFilter[Frac][k] is the kth interpolation filter coefficient in phase Frac, shift1 is a normalization parameter that adjusts the value range, and offset1=1<<(shift1-1).
[0133] temp[x][y] = (ΣmcFilter[xFrac][k]*refImg[xInt+k-NTAP / 2+1][yInt]+offset1)>>shift1 Next, the motion compensation unit 3091 derives the interpolated image Pred[][] by vertically interpolating the temporary image temp[][]. In the following, Σ is the sum for k=0..NTAP-1, shift2 is a normalization parameter that adjusts the value range, and offset2=1<<(shift2-1).
[0134] Pred[x][y] = (ΣmcFilter[yFrac][k]*temp[x][y+k-NTAP / 2+1]+offset2)>>shift2 In the case of bi-prediction, the above Pred[][] is used to derive interpolated images PredL0[][] and PredL1[][] for each L0 list and L1 list, and the interpolated image Pred[][] is generated from PredL0[][] and PredL1[][].
[0135] Here, shift1 = Min (4, BitDepth - 8) shift2 = 6 shift3 = Max (2, 14 - BitDepth) may be used.
[0136] (Configuration of intra prediction parameter derivation unit 304) The intra prediction parameter derivation unit 304 derives intra prediction parameters, for example, an intra prediction mode IntraPredMode, by referring to prediction parameters stored in the prediction parameter memory 307, based on the syntax elements decoded by the parameter decoding unit 302 (or syntax elements derived by the parameter encoding unit 111). The intra prediction parameter derivation unit 304 outputs the intra prediction parameters to the predicted image generation unit 308, and also stores them in the prediction parameter memory 307. The intra prediction parameter derivation unit 304 may derive different intra prediction modes for luma and chroma.
[0137] 7 is a schematic diagram showing the configuration of the intra prediction parameter derivation unit 304 of the prediction parameter derivation unit 320. As shown in the figure, the intra prediction parameter derivation unit 304 is configured to include a luma intra prediction parameter derivation unit 3042 and a chroma intra prediction parameter derivation unit 3043.
[0138] The luma intra prediction parameter derivation unit 3042 includes an MPM candidate list derivation unit 30421, an MPM parameter derivation unit 30422, and a non-MPM parameter derivation unit 30423 (decoding unit, derivation unit).
[0139] The MPM parameter derivation unit 30422 derives IntraPredModeY by referencing the mpmCandList[ ] and intra_luma_mpm_idx derived by the MPM candidate list derivation unit 30421 , and outputs the IntraPredModeY to the intra-predicted image generation unit 310 .
[0140] The non-MPM parameter derivation unit 30423 derives RemIntraPredMode from mpmCandList[ ] and intra_luma_mpm_reminder, and outputs IntraPredModeY to the intra-predicted image generation unit 310.
[0141] The chrominance intra-prediction parameter derivation unit 3043 derives IntraPredModeC from the syntax elements of the chrominance intra-prediction parameters, and outputs it to the intra-prediction image generation unit 310.
[0142] (IntraInter synthesis processing) When ciip_mode is 1, the IntraInter synthesis unit 30951 generates a predicted image in CIIP mode by weighting the inter predicted image and the intra predicted image. For example, the inter predicted image is generated using merge mode, and the intra predicted image is generated using planar prediction for the current block. The weights for the inter predicted image and the intra predicted image are derived based on, for example, whether the adjacent blocks above and to the left of the current block are intra blocks.
[0143] In the case of CIIP prediction (when ciip_flag is 1), the intra-prediction image generation unit 310 generates the predicted image PredIntra[][] using planar prediction (IntraPredModeY=INTRA_PLANAR).
[0144] When ciip_flag is 1, the inter predicted image generating unit 309 performs motion compensation using the motion vector obtained by merge prediction to generate a predicted image PredInter[][].
[0145] When ciip_flag is 1, the IntraInter synthesis unit 30951 generates a predicted image PredComb[][] by weighting the inter-predicted image PredInter[][] and the intra-predicted image PredIntra[][], and outputs the generated image to the addition unit 312.
[0146] PredComb[x][y] = (w * PredIntra[x][y] + (4 - w) * PredInter[x][y] + 2) >> 2 where w is set to 3 if both the upper and left neighboring blocks of the target CU are in intra mode, 1 if both are not in intra mode, and 2 otherwise.
[0147] (GPM Prediction) Next, GPM prediction will be described. Fig. 10 is a diagram illustrating GPM prediction. As shown in Fig. 10(a), in GPM prediction, a target CU is divided into two prediction units (hereinafter also referred to as partitions or regions) with a line segment as the boundary.
[0148] As shown in Fig. 10b, a line segment spanning a target CU is specified by an angle index angleIdx and a distance index distanceIdx. angleIdx indicates the angle φ between a vertical line and the line segment. distanceIdx indicates the distance ρ from the center of the target CU to the line segment. For angles, as shown in Fig. 10c, one angle mode (angle index) is assigned every 15 degrees. In the figure, 24 angle modes numbered 0 to 23 are used. For distance, for example, distanceIdx = 0 to 3 shown in Fig. 10d is used as the distance mode (distance index).
[0149] The GPM predicted image is derived by deriving two "rectangular" predicted images that include the prediction unit, rather than deriving a "non-rectangular" predicted image corresponding to each prediction unit, and performing weighting according to the shape of the prediction unit. In other words, the motion compensation unit 3091 or the intra-predicted image generation unit 310 derives two temporary predicted images for the target CU, and the GPM synthesis unit 30952 derives a predicted image by performing weighting processing on each pixel of the two temporary predicted images according to the pixel position.
[0150] The GPM prediction unit 30377 derives prediction parameters for the two regions used in GPM prediction and supplies them to the inter predicted image generation unit 309. The derivation of the two predicted images and synthesis using the predicted images are performed by the motion compensation unit 3091 and GPM synthesis unit 30952.
[0151] (Decoding Syntax in GPM Prediction) The upper parts of Figures 11 and 12 show part of the syntax of merge_data that is decoded when merge prediction is on (general_merge_flag==1) for the current block. The syntax shown in the upper part of Figure 12 is a continuation of the syntax in Figure 11. The middle and lower parts of Figure 12 show syntax that indicates the GPM partition configuration, merge candidates, or prediction mode, and the syntax shown in the lower part of Figure 12 is used when at least one of the partitions is intra-predicted. The parameter decoding unit 302 decodes syntax elements in the encoded data, and the GPM prediction unit 30377 (inter-prediction parameter derivation unit 303) derives GPM prediction parameters in accordance with the following rules.
[0152] In the example shown in the figure, GPM prediction is determined to be on when ciip_flag is 0. At this time, the syntax element gpm_adaptive_blending_idx for GPM prediction is decoded. gpm_adaptive_blending_idx is a parameter that controls the degree of weighting of the predicted image at the boundary.
[0153] <Syntax structure of GPM prediction intra MMVD> gpm_mmvd_flag0 is decoded, and if gpm_mmvd_flag0 is 1, it is determined that GPM-MMVD is on in area A (area label KK=0). In this case, the GPM-MMVD syntax elements gpm_mmvd_distance_idx0 and gpm_mmvd_direction_idx0 are decoded. gpm_mmvd_distance_idx0 indicates the magnitude of the differential motion vector MVD that corrects the motion vector of area A, and gpm_mmvd_direction_idx0 indicates the direction of the MVD of area A. Otherwise, when GPM-MMVD is off (gpm_mmvd_flag0==0), gpm_mmvd_distance_idx0 and gpm_mmvd_direction_idx0 may not be decoded, and the syntax element gpm_intra_flag0 indicating whether or not region A is in intra prediction mode may be decoded. Note that when gpm_intra_flag0 is not decoded and gpm_intra_flag0 does not appear, it is estimated that gpm_intra_flag0=0.
[0154] For region B (region label KK=1), the same syntax element gpm_mmvd_flag1 as for region A is decoded, and if gpm_mmvd_flag1==1, gpm_mmvd_distance_idx1 and gpm_mmvd_direction_idx1 are decoded. Otherwise (region B is other than MMVD, gpm_mmvd_flag1==0) and region A is other than intra (gpm_intra_flag0==0), the flag gpm_intra_flag1 indicating whether region B is intra is decoded. Note that if gpm_intra_flag1 is not decoded or if gpm_intra_flag1 does not appear, it is assumed that gpm_intra_flag1=0. In the above example, intra prediction and MMVD are exclusive in the GPM prediction of regions A and B, eliminating the need for redundant syntax decoding.
[0155] <Exclusion Between GPM-Intra and GPM-MMVD> Alternatively, instead of decoding flags gpm_intra_flag0 and gpm_intra_flag1 indicating whether region A and region B are intra, only the flag gpm_intra_flag0 indicating whether region A is intra may be decoded. In this case, region B is always inter prediction, and gpm_intra_flag1 is always 0. In other words, the determination of if(!gpm_intra_flag0[x0][y0]) in the figure and the decoding of gpm_intra_flag1[x0][y0] may be omitted.
[0156] <GPM-MMVD Restrictions for Area B> Furthermore, gpm_mmvd_flag1 may be decoded unless area A is intra (gpm_intra_flag0==1). If gpm_mmvd_flag1 is not decoded, or if gpm_mmvd_flag1 does not appear, it is assumed that gpm_mmvd_flag1=0. In other words, GPM-MMVD may be selected in area B only if area A is intra. Furthermore, gpm_mmvd_flag1 may be decoded unless area A is intra (gpm_intra_flag0==1), the slice is a P slice (sh_slice_type == P), or there is one merge candidate (sps_max_num_geo_cand == 1). That is, if region A is not intra, or if the slice is a P slice (sh_slice_type == P) or there is one merge candidate (sps_max_num_geo_cand == 1), gpm_mmvd_flag1 may not be decoded and may be estimated to be 0. Note that if the slice is not a P slice (sh_slice_type == P) or there is one merge candidate (sps_max_num_geo_cand == 1), it may be determined that the slice is a B slice (sh_slice_type == B) and there are two or more merge candidates (sps_max_num_geo_cand > 1).
[0157] <Another example of exclusion between GPM-intra and GPM-MMVD> Note that gpm_intra_flagKK may be decoded before gpm_mmvd_flagKK, and when gpm_intra_flagKK==0, gpm_mmvd_flagKK may be further decoded, and when gpm_mmvd_flagKK==1, gpm_mmvd_distance_idxKK and gpm_mmvd_direction_idxKK may be decoded. KK=0 or 1. In the above example, intra prediction and MMVD are exclusive in GPM prediction for regions A and B, eliminating the need for redundant syntax decoding.
[0158] <GPM Mode Syntax Decoding> Furthermore, the GPM prediction syntax elements merge_gpm_partition_idx, merge_gpm_idx0, and merge_gpm_idx1 indicated by gpm_merge_idx() or gpm_merge_idx1() are decoded. merge_gpm_partition_idx is an index (partition index) indicating the division pattern of the GPM prediction mode. Specifically, it indicates a combination of angleIdx and distanceIdx that identify a line segment spanning the target block to divide the target block into two regions. The number of partition index options is sometimes referred to as the number of division patterns in this specification.
[0159] If neither region A nor region B is in GPM-MMVD mode (gpm_mmvd_flag0==0 and gpm_mmvd_flag0==1), a flag tm_merge_flag indicating whether the current block uses template matching may be further decoded. Here, if both region A and region B are in a mode other than intra prediction mode (gpm_intra_flag0==0 and gpm_intra_flag0==1, that is, if (gpm_intra_flag0||gpm_intra_flag0) is 0), tm_merge_flag may be decoded. If tm_merge_flag is not decoded, or does not appear in the encoded data, it may be assumed that tm_merge_flag==0. If tm_merge_flag is true (==1), both region A and region B are inter (both are other than intra), and gpm_merge_idx is decoded. If tm_merge_flag is false and one of area A and area B is intra (gpm_intra_flag0||gpm_intra_flag0), decode gpm_merge_idx1(). Otherwise, both area A and area B are inter (both are not intra), and decode gpm_merge_idx().
[0160] In addition to the above, if both area A and area B are in GPM-MMVD mode (gpm_mmvd_flag0==1 and gpm_mmvd_flag0==1), both area A and area B are inter (both are not intra) and gpm_merge_idx() is decoded.
[0161] In addition to the above, if at least one of area A and area B is in GPM-MMVD mode (gpm_mmvd_flag0==1 or gpm_mmvd_flag0==1), decode gpm_merge_idx1().
[0162] Here, in the above, gpm_merge_idx() is decoded when both region A and region B are in inter mode. gpm_merge_idx1() is decoded when either region A or region B is in intra mode. When decoding gpm_merge_idx1(), if region A is in intra mode (gpm_intra_flag0==1), the syntax element gpm_intra_idx0 (equivalent to intra_luma_mpm_idx) that specifies the intra mode of region A is decoded, and if region B is in intra mode (gpm_intra_flag1==1), the syntax element gpm_intra_idx1 (equivalent to intra_luma_mpm_idx) that specifies the intra mode of region B is decoded. The intra prediction mode IntraPredModeY0 of region A and the intra prediction mode IntraPredModeY1 of region B may be derived from the following.
[0163] IntraPredModeY0 = mpmCandList[gpm_intra_idx0] IntraPredModeY1 = mpmCandList[gpm_intra_idx1] Also, different intra prediction mode lists may be used for the regions A and B. For example, different intra mode candidate lists mpmCandList0 and mpmCandList1 may be derived as follows.
[0164] IntraPredModeY0 = mpmCandList0[gpm_intra_idx0] IntraPredModeY1 = mpmCandList1[gpm_intra_idx1] (Another configuration) As shown in the syntax shown in the upper and middle sections of Figure 13, regardless of tm_merge_flag, if one of area A and area B is intra (gpm_intra_flag0||gpm_intra_flag0), gpm_merge_idx1() is decoded; otherwise, both area A and area B may be inter (both other than intra) and gpm_merge_idx() may be decoded.
[0165] When sps_gpm_enabled_flag is 1, the number of candidates for partition index options is NumGPMFull (for example, 82, 54, 40, 26, etc.), and merge_gpm_partition_idx takes an integer value between 0 and NumGPMFull-1.
[0166] MergeGpmFlag is a flag that indicates whether or not to perform GPM prediction on the target block. If all of the following conditions (GPM determination conditions) are met, the GPM prediction unit 30377 sets MergeGpmFlag=1; otherwise, the GPM prediction unit 30377 sets MergeGpmFlag=0. sps_gpm_enabled_flag=1 (GPM prediction is available in the target SPS) slice_type is B slice general_merge_flag=1 (merge prediction is on, and inter prediction parameters for the target block are estimated from neighboring inter prediction blocks) cbWidth>=8 and cbHeight>=8 regular_merge_flag=0 (basic merge prediction or MMVD prediction is off) merge_subblock_flag=0 (subblock-by-subblock inter prediction is off) ciip_flag=0 (combining process of intra-predicted image and inter-predicted image is off) When MergeGpmFlag=1, the GPM prediction unit 30377 derives parameters necessary for generating a predicted image in the following procedure. Also, when gpm_mmvd_flag0==1 or gpm_mmvd_flag1==1, it uses GPM-MMVD (GPM with Merge mode with Motion Vector Difference) to generate a differential motion vector MVD (Motion Vector Difference). The GPM synthesis unit 30952 calculates the motion vector difference (Vector Difference) and adds the motion vector to the motion vector.
[0167] When a region is inter-predicted, the GPM prediction unit 30377 derives merge indexes m and n from merge_gpm_idx0 and merge_gpm_idx1 as syntax indicating motion information of the two regions. If syntax gpm_intra_flag0 is 0, region A is inter-predicted, and if it is 1, it is intra-predicted. Similarly, if gpm_intra_flag1 is 0, region B is inter-predicted, and if it is 1, it is intra-predicted.
[0168] m = merge_gpm_idx0 n = merge_gpm_idx1 + (merge_gpm_idx1 >= m) ? 1 : 0 Derive M, the merge candidate pointed to by merge index m, and N, the merge candidate pointed to by merge index n, from mergeCandList.
[0169] M = mergeCandList[m] N = mergeCandList[n] Using the mergCandList derived by the merge prediction parameter derivation unit 3036, the motion information of merge candidates M and N (mvLXM, refIdxLXM, predFlagLXM, mvLXN, refIdxLXN, predFlagLXN, bcwIdx) derived using merge_gpm_idx0 and merge_gpm_idx1, and refineMvdLXN and refineMvdLXM derived by MPG-MMVD processing, the GPM prediction unit 30377 derives motion vectors mvA and mvB of predicted images, reference indices refIdxA and refIdxB, and prediction list flags predListFlagA and predListFlagB for the target area 2 as follows:
[0170] mvA[0] = mvLXM[0] + refineMvdLXM[0] mvA[1] = mvLXM[1] + refineMvdLXM[1] refIdxA = refIdxLXM predListFlagA = X Here, the GPM prediction unit 30377 sets the lowest 1 bit of m to X (X = m & 0x01). Note that if predFlagLXM is 0, the GPM prediction unit 30377 sets X to (1-X).
[0171] mvB[0] = mvLXN[0] + refineMvdLXM [0] mvB[1] = mvLXN[1] + refineMvdLXM [1] refIdxB = refIdxLXN predListFlagB = X Here, the GPM prediction unit 30377 sets the lowest 1 bit of n to X (X = n & 0x01). Note that if predFlagLXN is 0, the GPM prediction unit 30377 sets X to (1-X).
[0172] (GPM-MMVD processing) The GPM prediction unit 30377 derives the predicted motion vector index gpmMmvdIdx0 or gpmMmvdIdx1 for each of region A and region B when the flags gpm_mmvd_flag0 and gpm_mmvd_flag1 are 1 (true). The following is an example of how mvdL0 is derived for region A. In the case of bi-prediction, mvdL1 uses a vector in the opposite direction to mvdL0. The derivation of mvdLX for region B is similar.
[0173] (1) Derive the distance and direction of MVD in GPM-MMVD mode based on the distance (step) index gpm_mmvd_distance_idx0 and the direction index gpm_mmvd_direction_idx0.
[0174] (1.1) Derivation of the distance MmvdDistance: Decode gpm_mmvd_distance_idx0 and gpm_mmvd_distance_idx1 according to the number of options of the distance index of MMVD. For example, when the number of options is NDist, gpm_mmvd_distance_idx0 and gpm_mmvd_distance_idx can be decoded using the binaryization of the TR code with cMax = NDist - 1 as the maximum value. Here, the Rice parameter cRiceParam of the TR code (Truncated Rice) can be 0. The TR code is a code composed of a prefix bin determined by prefixVal and a suffix bin with an FL length of cMax = (1 << cRiceParam) - 1 determined by suffixVal. The relationship between the symbol value symbolVal, prefixVal, and suffixVal satisfies suffixVal = symbolVal - (prefixVal << cRiceParam). Also, the binaryization of the FL code of an equal-length code with cMax = NDist - 1 as the maximum value can be used. At this time, the bit length of the FL code is log2(N). The same applies when decoding the syntax mmvd_distance_idx of MMVD other than GPM prediction.
[0175] For example, based on the indexes gpm_mmvd_distance_idx0, gpm_mmvd_distance_idx1 and the conversion table DistanceTableNDist, MmvdDistance0 and MmvdDistance1 are derived by the following formula.
[0176] MmvdDistance0 = DistanceTable[gpm_mmvd_distance_idx0] MmvdDistance1 = DistanceTable[gpm_mmvd_distance_idx1] DistanceTable = {1,2,3,4,5,0,6,7,8} Here, if the flag ph_gpm_ext_mmvd_flag is 1 (true), the number of directions NDist that can be selected in GPM-MMVD may be set to 9, and if the ph_gpm_ext_mmvd_flag flag is 0 (false), it may be set to 8. In other words, NDist indicates that the value range of gpm_mmvd_distance_idx0 is from 0 to NDist-1.
[0177] (1.2) Derivation of basic MVD: gpm_mmvd_direction_idx0 and gpm_mmvd_direction_idx1 are decoded according to the number of choices for the MMVD direction index. For example, if the number of choices is NDir, gpm_mmvd_direction_idx0 and gpm_mmvd_direction_idx1 may be decoded using the binarization of the TR code with cMax = NDir-1 as the maximum value. Alternatively, binarization of the FL (Fixed Length) code, which is an equal-length code with cMax = NDir-1 as the maximum value, may be used. In this case, the bit length of the FL code is log2(N). The same applies when decoding the mmvd_direction_idx syntax of MMVD other than GPM prediction. Based on the indexes gpm_mmvd_distance_idx0 and gpm_mmvd_distance_idx1 and the conversion table DirectionTableNDir, mvdUnit0 and mvdUnit1 are derived using the following formula:
[0178] mvdUnit0[0] = DirectionTableNDirx[gpm_mmvd_direction_idx0] mvdUnit0[1] = DirectionTableNDiry[gpm_mmvd_direction_idx0] mvdUnit1[0] = DirectionTableNDirx[gpm_mmvd_direction_idx1] mvdUnit1[1] = DirectionTableNDiry[gpm_mmvd_direction_idx1] Note that if the flag ph_gpm_ext_mmvd_flag is 1 (true), the number of directions selectable in GPM-MMVD, numGpmMMVDDir, can be set to 8, and if the ph_gpm_ext_mmvd_flag flag is 0 (false), it can be set to 4.
[0179] (1.3) (Example of MMVD Candidates in GPM-MMVD Processing) Here, an example of the number of MMVD candidates in GPM-MMVD processing will be described.
[0180] The number of MMVD candidates in the GPM-MMVD process is determined according to the number of distances or directions indicated by the index. The number of MMVD candidates may be changed depending on the combination of prediction modes used for two partitions in GPM prediction. That is, the maximum value of the distance index (merge_gpm_mmvd_distance_idxKK) and direction index (merge_gpm_mmvd_direction_idxKK) included in the coded data may be changed depending on the combination of prediction modes used for two partitions in GPM prediction.
[0181] For example, in GPM-MMVD processing, when the inter-predicted image generation unit 309 generates at least one of two temporary predicted images (first predicted images) using MMVD prediction, the specific parameters for identifying MMVD candidates may be restricted as follows: The options indicated by the distance index (distance_idx) and direction index (direction_idx) for identifying MMVD candidates, derived by the GPM prediction unit 30377, may be different numbers for MMVD prediction using GPM prediction and MMVD prediction without GPM prediction. For example, the distance index gpm_mmvd_distance_idxKK for MMVD prediction using GPM prediction may use a different value range than the distance index mmvd_distance_idx for MMVD prediction without GPM prediction. The distance index gpm_direction_idxKK for MMVD prediction using GPM prediction may use a different value range than the distance index mmvd_direction_idx for MMVD prediction without GPM prediction. For example, the maximum value cMax of mmvd_direction_idx is less than the maximum value cMax of gpm_mmvd_direction_idxKK.
[0182] That is, when a first predicted image is generated by MMVD prediction in which a prediction parameter is used to shift the value of a predicted motion vector derived in a merge mode, the number of options for a specific parameter that identifies candidates for the prediction parameter may be as follows: The number of options for the specific parameter may be less than the number of options for the specific parameter derived in an MMVD mode in which a predicted image for the target block is generated without using a synthesis mode.
[0183] In this specification, GPM prediction may be referred to as a synthesis mode in which a second predicted image of a current block is synthesized from two first predicted images using weights derived according to boundaries in the current block. This synthesis mode uses weights determined according to positions within the current block (weights determined according to geometric positions). In addition, in this specification, MMVD candidates, which are parameter candidates used to shift values of predicted motion vectors derived in merge mode in MMVD prediction, may be referred to as prediction parameter candidates. In addition, in this specification, there may be two parameters identifying an MMVD candidate. In particular, a distance index (mmvd_distance_idx, gpm_mmvd_distance_idxkk) and a direction index (mmvd_direction_idx, gpm_mmvd_direction_idxKK) may be referred to as specific parameters.
[0184] "A mode for generating a predicted image for a target block without using a synthesis mode" can also be expressed as "a mode for generating a predicted image for a target block using only one predicted image, or using weighting of the predicted image that is not determined according to the position within the target block."
[0185] A specific example of a restriction on the number of MMVD candidates in the GPM-MMVD process will be described below.
[0186] (1.3.1: Example 1 of Limiting the Number of MMVD Candidates in GPM-MMVD Processing) In this example, an example of limiting the number of MMVD candidates in MMVD prediction performed for a partition related to GPM prediction will be described.
[0187] For example, when inter prediction using MMVD prediction without GPM prediction is performed on a current block, the total number of prediction candidates is the number of MMVD candidates. Note that, hereinafter, "inter prediction using MMVD prediction" will also be simply referred to as "MMVD prediction."
[0188] On the other hand, when predicting a target block using GPM prediction by combining MMVD prediction and intra prediction, the total number of prediction candidates is the product of the number of partitions related to GPM prediction, the number of MMVD candidates, and the number of intra prediction candidates. In other words, the total number of prediction candidates is "number of partitions × number of MMVD candidates × number of intra prediction candidates."
[0189] Furthermore, when MMVD prediction is performed for each partition of a target block using GPM prediction, the total number of prediction candidates is the product of the number of partitions and the number of MMVD candidates. That is, the total number of prediction candidates is "number of partitions × number of MMVD candidates."
[0190] Therefore, MMVD prediction using GPM prediction is more complex to process than MMVD prediction without GPM prediction. This example describes an example of reducing the complexity of processing in prediction that combines GPM prediction and MMVD prediction.
[0191] In this example, the number of MMVD candidates in MMVD prediction using GPM prediction is limited to be less than the number of MMVD candidates in MMVD prediction without GPM prediction.
[0192] 14 is a diagram showing an example of limiting the number of MMVD candidates in MMVD prediction using GPM prediction. The diagram shows the number of distance options indicated by the distance index (gpm_mmvd_distance_idxKK) and the number of direction options indicated by the direction index (merge_gpm_mmvd_direction_idxKK) derived by the GPM prediction unit 30377 in MMVD prediction (inter).
[0193] As shown in the second row of Figure 14, in MMVD prediction (inter) that does not use GPM prediction, the number of distance options indicated by the distance index is NDistL (NDistL = 8, mmvd_distance_idx = 0..NDistL-1, cMax = NDistL-1), and the number of direction options indicated by the direction index is NDirL (NDirL = 8, mmvd_direction_idx = NDirL-1, cMax = NDirL-1).
[0194] The example shown in the third row of FIG. 14 illustrates an example in which MMVD prediction is performed on each of partitions 1 and 2, which are two partitions related to GPM prediction. As shown in the figure, in MMVD prediction (inter) using GPM prediction, the number of distance options indicated by the distance index is NDistS4 (NDistS=4, gpm_mmvd_distance_idx=0..NDistS-1, cMax=NDistS-1). Also, the number of direction options indicated by the direction index is NDirS4 (NDirS=4, gpm_mmvd_direction_idx=0..NDirS-1, cMax=NDirS-1).
[0195] The example shown in the fourth row of FIG. 14 illustrates an example in which MMVD prediction is performed on partition 2, which is one of two partitions related to GPM prediction, and intra prediction is performed on the other partition 1 (gpm_intra_flag0==true or gpm_intra_flag1==true). In MMVD prediction (inter) on partition 2, the number of distance options indicated by the distance index is NDistS4. The number of direction options indicated by the direction index is also NDirS4. The value range of the syntax and the maximum value cMax of the binarization of the syntax during decoding may be gpm_mmvd_distance_idx=0..NDistS-1, cMax=NDistS-1, gpm_mmvd_direction_idx=0..NDirS-1, and cMax=NDirS-1.
[0196] Here, when GPM prediction is not used, MVD may be derived using DistanceTableNDistL, in which the number of options is NDistL, and when GPM prediction is used, DistanceTableNDistS, in which the number of options is NDistS.
[0197] MmvdDistance = DistanceTableNDistL[mmvd_distance_idx] MmvdDistance = DistanceTableNDistS[gpm_mmvd_distance_idxKK] Here, if GPM prediction is not used, MVD may be derived using DirectionTableNDirL, where the number of options is NDirL, and if GPM prediction is used, MVD may be derived using DirectionTableNDirS, where the number of options is NDirS. When GPM prediction is not used, mvdUnit[0] = DirectionTableNDirLx[mmvd_direction_idx] mvdUnit[1] = DirectionTableNDirLy[mmvd_direction_idx] When GPM prediction is used, for KK=0..1, mvdUnit[0] = DirectionTableNDirSx[gpm_mmvd_direction_idxKK] mvdUnit[1] = DirectionTableNDirSy[gpm_mmvd_direction_idxKK] In other words, in this configuration, the number of distance candidates for which GPM prediction is not used, NDistL, is set to the number of distance candidates for which GPM prediction is used, NDistS (NDistL > NDistS may be true, for example, NDistL = 8, NDirS = 4). Alternatively, the number of direction candidates for which GPM prediction is not used, NDirL, is set to the number of direction candidates for which GPM prediction is used, NDirS (NDirL > NDirS may be true, for example, NDirL = 8, NDirS = 4). As already explained, each syntax is decoded with cMax=Ndir-1.
[0198] For example, in the example shown in Fig. 14, in MMVD prediction using GPM prediction, the number of distance options indicated by the distance index is limited from 8 to 4 compared to MMVD prediction not using GPM prediction. Also, in MMVD prediction using GPM prediction, the number of direction options indicated by the direction index is limited from 8 to 4 compared to MMVD prediction not using GPM prediction.
[0199] The above example can also be expressed as follows:
[0200] The number of options for the parameters indicating direction and distance, which are parameters specifying candidates for prediction parameters derived by the inter prediction parameter derivation unit 303, is limited as follows: That is, compared to a mode in which a predicted image is generated using a synthesis mode, the number of options for the parameters (the syntax element cMax and the number of elements in DirectionTable) is smaller in a mode in which a predicted image is generated without using a synthesis mode.
[0201] As another example, in MMVD prediction using GPM prediction, at least one of the number of distance options indicated by the distance index and the number of direction options indicated by the direction index may be limited to be smaller than that in MMVD prediction without GPM prediction.
[0202] (1.3.2: Example 2 of Limiting the Number of MMVD Candidates in GPM-MMVD Processing) This example describes another example of limiting the number of MMVD candidates in MMVD prediction performed on partitions related to GPM prediction. In this example, the number of MMVD candidates is limited differently when MMVD prediction is performed on both of two partitions related to GPM prediction and when MMVD prediction is performed on one of the partitions and intra prediction is performed on the other (gpm_intra_flag0==true or gpm_intra_flag1==true).
[0203] For example, when prediction is performed using a combination of MMVD prediction and intra prediction using GPM prediction, the total number of prediction candidates may be the product of the number of partitions related to GPM prediction, the number of MMVD candidates, and the number of intra prediction candidates. In other words, the total number of prediction candidates may be "the number of partitions × the number of MMVD candidates × the number of intra prediction candidates."
[0204] Furthermore, when MMVD prediction is performed on each partition of a current block using GPM prediction, the total number of prediction candidates may be the product of the number of partitions and the number of MMVD candidates. That is, the total number of prediction candidates may be "the number of partitions × the number of MMVD candidates."
[0205] Therefore, in MMVD prediction using GPM prediction, the processing may be more complicated when prediction is performed by a combination of MMVD prediction and intra prediction compared to when only MMVD prediction is performed. This example shows an example of reducing the processing complexity in prediction by a combination of GPM prediction, MMVD prediction, and intra prediction.
[0206] In this example, in MMVD prediction using GPM prediction, the number of MMVD candidates when prediction is performed by a combination of MMVD prediction and intra prediction is set to be different from the number of MMVD candidates when only MMVD prediction is performed. Here, the number of MMVD candidates when prediction is performed by a combination of MMVD prediction and intra prediction is limited to a smaller number. For example, when intra is included (gpm_intra_flag0=1 or gpm_intra_flag1=1), NDir=NDirL is set, and otherwise NDir=NDirS is set, and gpm_mmvd_direction_idxKK may be decoded using binarization with cMax=NDir-1. Also, different Direction Tables may be selected when intra is included (gpm_intra_flag0=1 or gpm_intra_flag1=1) and when not. For example, a Direction Table may be selected depending on NDir (DirectionTableNDir is selected). Here, NDirL > NDirS. NDirL = 8, NDirS = 4 may also be used.
[0207] 15 is a diagram showing another example of limiting the number of MMVD candidates in MMVD prediction using GPM prediction. The diagram shows the number of distance options indicated by the distance index and the number of direction options indicated by the direction index derived by the GPM prediction unit 30377 in MMVD prediction.
[0208] As shown in the second row of Fig. 15, in MMVD prediction (inter) that does not use GPM prediction, the number of distance options indicated by the distance index is 8. In addition, the number of direction options indicated by the direction index is also 8.
[0209] The example shown in the third row of Fig. 15 shows an example in which MMVD prediction is performed on each of partitions 1 and 2, which are two partitions related to GPM prediction. As shown in the third row of Fig. 15, in MMVD prediction (inter) using GPM prediction, the number of distance options indicated by the distance index is 4. Also, the number of direction options indicated by the direction index is 8.
[0210] The example shown in the fourth row of FIG. 15 shows a case where MMVD prediction is performed on one of two partitions related to GPM prediction, partition 2, and intra prediction is performed on the other partition, partition 1 (in the case of intra GPM, gpm_intra_flag0=1 or gpm_intra_flag1=1). In the MMVD prediction (inter) for partition 2, the number of distance options indicated by the distance index is 4. Also, the number of direction options indicated by the direction index is 4.
[0211] The above example can also be expressed as follows:
[0212] The number of options for a direction parameter, which is a parameter that specifies candidates for a prediction parameter derived by the GPM prediction unit 30377, is limited as follows: In a mode in which the inter-predicted image generation unit 309 generates a predicted image using a synthesis mode, the number of options (the syntax elements cMax and the number of elements in DirectionTable) is smaller when one of the first predicted images is generated using intra prediction (gpm_intra_flag0==true or gpm_intra_flag1==true) than when two first predicted images are generated using an MMVD mode. For example, with respect to cMax of gpm_mmvd_direction_idxKK, cMax = (is intra GPM)? NDirS:NDirL may be used. With respect to the DirectionTable from which mvdUnit is derived, DirectionTable = (is intra GPM)? DirectionTableNDirS: DirectionTableNDirL may be used. Here, NDirS < NDirL may be used. For example, (NDirS, NDirL) = (4, 8), (8, 16), etc.
[0213] As another example, the number of MMVD candidates may be limited as follows: In a prediction in which inter prediction is performed on one of the partitions related to GPM prediction and intra prediction is performed on the other, the use of MMVD may be prohibited in the inter prediction. In other words, the inter prediction may be merge prediction that does not use MMVD.
[0214] (1.3.3: Example 3 of Limiting the Number of MMVD Candidates in GPM-MMVD Processing) This example describes another example of limiting the number of MMVD candidates in MMVD prediction performed on partitions related to GPM prediction. In this example, the number of MMVD candidates is limited differently when MMVD prediction is performed on both of two partitions related to GPM prediction and when MMVD prediction is performed on one of the partitions and intra prediction is performed on the other.
[0215] For example, when GPM prediction is used to perform prediction by a combination of MMVD prediction and intra prediction, the calculation of the predicted image may be a calculation obtained by adding an intra prediction calculation and an L0 prediction calculation, or a calculation obtained by adding an intra prediction calculation and an L1 prediction calculation. In other words, the prediction calculation may be "intra prediction + L0 inter prediction, or intra prediction + L1 inter prediction."
[0216] Also, when GPM prediction is used and MMVD prediction is performed for each partition, the prediction calculation may be a calculation that adds the L1 prediction calculation and the L0 prediction calculation. That is, the prediction calculation may be "L0 inter prediction + L1 inter prediction".
[0217] Therefore, in MMVD prediction using GPM prediction, the processing may be more complicated when only MMVD prediction is performed compared to when prediction is performed using a combination of MMVD prediction and intra prediction. This example shows an example of reducing the processing complexity in prediction using a combination of GPM prediction and MMVD prediction without intra prediction.
[0218] 16 is a diagram showing another example of limiting the number of MMVD candidates in MMVD prediction using GPM prediction. The diagram shows the number of distance options indicated by the distance index and the number of direction options indicated by the direction index derived by the GPM prediction unit 30377 in MMVD prediction (inter).
[0219] As shown in the second row of Fig. 16, in MMVD prediction (inter) that does not use GPM prediction, the number of distance options indicated by the distance index is 8. In addition, the number of direction options indicated by the direction index is also 8.
[0220] The example shown in the third row of Fig. 16 shows an example in which MMVD prediction is performed on each of partitions 1 and 2, which are two partitions related to GPM prediction. As shown in the third row of Fig. 16, in MMVD prediction using GPM prediction (GPM inter), the number of distance options indicated by the distance index is 4. Also, the number of direction options indicated by the direction index is 4.
[0221] The example shown in the fourth row of FIG. 16 illustrates an example in which MMVD prediction is performed on one of two partitions related to GPM prediction, partition 2, and intra prediction is performed on the other, partition 1. In the MMVD prediction (GPM inter) for partition 2, the number of distance options indicated by the distance index is 4. Also, the number of direction options indicated by the direction index is 8.
[0222] That is, in this example, compared to a prediction in which MMVD prediction is performed on one of the partitions and intra prediction is performed on the other, a prediction in which MMVD prediction is performed on each of the two partitions limits the number of direction options indicated by the direction index to a smaller number.
[0223] The above example can also be expressed as follows:
[0224] The number of options for a direction parameter, which is a parameter specifying prediction parameter candidates and which is derived by the GPM prediction unit 30377, is limited as follows: In a mode in which the inter-predicted image generation unit 309 generates a predicted image using a synthesis mode, the number of options (the syntax element cMax and the number of elements in DirectionTable) is smaller when two first predicted images are generated using MMVD mode than when one of the first predicted images is generated using intra-prediction. For example, with regard to cMax of gpm_mmvd_direction_idxKK, cMax = (is intra GPM) ? NDirL : NDirS may be satisfied. With regard to the DirectionTable from which mvdUnit is derived, DirectionTable = (is intra GPM) ? DirectionTableNDirL : DirectionTableNDirS may be satisfied. Here, NDirS < NDirL may be satisfied. For example, (NDirS, NDirL) = (4, 8), (8, 16), etc. may be satisfied.
[0225] (1.3.4: Example 4 of Limiting the Number of MMVD Candidates in GPM-MMVD Processing) This example describes another example of limiting the number of MMVD candidates in MMVD prediction performed on partitions related to GPM prediction. In this example, when MMVD prediction is performed on one of the partitions related to GPM prediction and intra prediction is performed on the other, and the intra prediction is intra block copy (IBC) prediction, the number of MMVD candidates is limited to a small number. Note that intra block copy, which is one of the prediction modes, is a prediction mode included in intra prediction, which obtains a predicted image by copying a predicted image in units of prediction blocks from decoded surrounding areas of the same picture.
[0226] For example, if MMVD prediction is performed on each partition of a current block using GPM prediction, the total number of prediction candidates may be the number of partitions multiplied by the number of MMVD candidates. That is, the total number of prediction candidates may be "number of partitions × number of MMVD candidates."
[0227] Furthermore, when GPM prediction is used to perform prediction by combining MMVD prediction and intra prediction other than IBC, the total number of prediction candidates may be the product of the number of partitions related to GPM prediction, the number of MMVD candidates, and the number of intra prediction candidates. In other words, the total number of prediction candidates may be "the number of partitions × the number of MMVD candidates × the number of intra prediction candidates."
[0228] Furthermore, when GPM prediction is used to perform prediction by combining MMVD prediction and IBC prediction, the total number of prediction candidates may be a value obtained by multiplying the number of partitions related to GPM prediction by the number of MMVD candidates and the number of IBC candidates. In other words, the total number of prediction candidates may be "the number of partitions × the number of MMVD candidates × the number of IBC candidates."
[0229] Therefore, in MMVD prediction using GPM prediction, the prediction process for the combination of MMVD prediction and IBC prediction may be more complex than the combination of MMVD prediction and intra prediction other than IBC prediction. This example shows an example of reducing the processing complexity in the prediction process for the combination of GPM prediction, MMVD prediction, and IBC prediction.
[0230] 17 is a diagram showing another example of limiting the number of MMVD candidates in MMVD prediction using GPM prediction. The diagram shows the number of distance options indicated by the distance index and the number of direction options indicated by the direction index derived by the GPM prediction unit 30377 in MMVD prediction (inter).
[0231] As shown in the second row of Fig. 17, in MMVD prediction (inter) that does not use GPM prediction, the number of distance options indicated by the distance index is 8. In addition, the number of direction options indicated by the direction index is also 8.
[0232] The example shown in the third row of Fig. 17 shows an example in which MMVD prediction is performed on each of partitions 1 and 2, which are two partitions related to GPM prediction. As shown in the third row of Fig. 17, in MMVD prediction (inter) using GPM prediction, the number of distance options indicated by the distance index is 4. Also, the number of direction options indicated by the direction index is 8.
[0233] The example shown in the fourth row of Fig. 17 shows an example in which MMVD prediction is performed on one of two partitions related to GPM prediction, partition 2, and intra prediction other than IBC is performed on the other partition, partition 1. In the MMVD prediction (inter) on partition 2, the number of distance options indicated by the distance index is 4. Also, the number of direction options indicated by the direction index is 8.
[0234] The example shown in the fifth row of Fig. 17 shows an example in which MMVD prediction (inter) is performed on one of two partitions related to GPM prediction, partition 2, and IBC prediction is performed on the other, partition 1. In the MMVD prediction (GPM inter) on partition 2, the number of distance options indicated by the distance index is 4. Also, the number of direction options indicated by the direction index is 4.
[0235] The above example can also be expressed as follows:
[0236] The number of options for a direction parameter, which is a parameter that specifies candidates for a prediction parameter derived by the GPM prediction unit 30377, is derived as follows. In a mode in which the inter-predicted image generation unit 309 generates a predicted image using a synthesis mode, the number of options (the syntax elements cMax and the number of elements in DirectionTable) is smaller when one of the first predicted images is generated using IBC prediction than when one of the first predicted images is generated using intra prediction other than IBC prediction. For example, with respect to cMax of gpm_mmvd_direction_idxKK, cMax = (IBC-GPM) ? NDirS : NDirL may be satisfied. With respect to the DirectionTable referenced by gpm_mmvd_direction_idxKK, DirectionTable = (IBC-GPM) ? DirectionTableNDirS : DirectionTableNDirL may be satisfied. Here, NDirS < NDirL may be satisfied. For example, (NDirS, NDirL) = (4, 8), (8, 16), etc. may be satisfied.
[0237] (1.3.5: Example 4 of Limiting the Number of MMVD Candidates in GPM-MMVD Processing) This example describes another example of limiting the number of MMVD candidates in MMVD prediction performed on partitions related to GPM prediction. In this example, the number of MMVD candidates is limited differently when the direction of the boundary between two partitions related to GPM prediction is diagonal with respect to the target block and when the direction of the boundary is horizontal or vertical with respect to the target block. Note that a diagonal with respect to the target block is any direction other than horizontal or vertical with respect to the target block.
[0238] More specifically, in MMVD prediction using GPM prediction, the number of MMVD candidates when the boundary direction between two partitions for GPM prediction is diagonal is limited to be less than the number of MMVD candidates when the boundary direction is horizontal or vertical relative to the target block.
[0239] 18 is a diagram showing another example of limiting the number of MMVD candidates in MMVD prediction using GPM prediction. The diagram shows the number of distance options indicated by the distance index and the number of direction options indicated by the direction index derived by the GPM prediction unit 30377 in MMVD prediction (inter).
[0240] As shown in the second row of Fig. 18, in MMVD prediction (inter) that does not use GPM prediction, the number of distance options indicated by the distance index is 8. In addition, the number of direction options indicated by the direction index is also 8.
[0241] The example shown in the third row of Figure 18 shows an example of the number of MMVD candidates when the boundary direction between two partitions for GPM prediction is horizontal or vertical with respect to the target block. As shown in the figure, in MMVD prediction (inter) when the boundary direction is horizontal or vertical with respect to the target block, the number of distance options indicated by the distance index is 8. Also, the number of direction options indicated by the direction index is 8.
[0242] The example shown in the fourth row of Figure 18 shows an example of the number of MMVD candidates when the boundary direction between two partitions for GPM prediction is diagonal to the target block. As shown in the figure, in MMVD prediction (GPM inter) when the boundary direction is diagonal to the target block, the number of distance options indicated by the distance index is 4. Also, the number of direction options indicated by the direction index is 8.
[0243] That is, in this example, when the direction of the boundary between two partitions related to GPM prediction is diagonal to the target block, the following is derived: The number of MMVD candidates is restricted to be smaller when the direction of the boundary is diagonal to the target block than when the direction of the boundary is horizontal or vertical to the target block.
[0244] The above example can also be expressed as follows. The number of options for a distance parameter, which is a parameter specifying candidates for prediction parameters derived by the GPM prediction unit 30377, is derived as follows. In a mode in which the inter-predicted image generation unit 309 generates a predicted image using a synthesis mode, the number of options (the syntax elements cMax and the number of elements in DirectionTable) is smaller when the boundary of the target block is diagonal to the target block than when the boundary is horizontal or vertical to the target block. For example, with respect to cMax of gpm_mmvd_direction_idxKK, cMax = (angleIdx is in a predetermined range) ? NDirS: NDirL may be used. With respect to DirectionTable from which mvdUnit is derived, DirectionTable = (angleIdx is in a predetermined range) ? DirectionTableNDirS: DirectionTableNDirL may be used. Here, NDirS < NDirL. For example, (NDirS, NDirL) = (4, 8), (8, 16), etc. may be used.
[0245] (1.4.1: Example of Limiting the Number of Intra-Predictions in GPM Prediction) In this example, an example of limiting the number of intra-predictions performed on partitions related to GPM prediction will be described.
[0246] For example, when a current block is predicted using a combination of MMVD prediction and intra prediction using GPM prediction, the total number of prediction candidates may be the product of the number of partitions related to GPM prediction, the number of MMVD candidates, and the number of intra prediction candidates. In other words, the total number of prediction candidates may be "the number of partitions × the number of MMVD candidates × the number of intra prediction candidates."
[0247] Therefore, when GPM prediction is used to perform prediction by combining MMVD prediction and intra prediction, the processing becomes complicated depending on the number of intra predictions. This example describes an example of reducing the complexity when GPM prediction is used.
[0248] In particular, in this example, when the inter-prediction image generation unit 309 generates at least one of two temporary predicted images (first predicted images) using intra prediction in GPM prediction, the type of intra prediction mode is derived as follows: That is, the options for the type of intra prediction mode derived by the GPM prediction unit 30377 are different from the options for the type of intra prediction mode derived in a mode in which a predicted image of the current block is generated without using GPM prediction.
[0249] For example, the number of options for the types of intra prediction modes derived by the GPM prediction unit 30377 is fewer than the number of options for the types of intra prediction modes derived in a mode for generating a predicted image for a current block without using GPM prediction. For example, cMax = NIntraS-1, which indicates the maximum value of the range of the syntax element gpm_intra_idxKK that selects the intra prediction mode for GPM intra prediction, may be smaller than the number of elements in the candidate list mpmCandMode of intra prediction modes used for intra prediction other than GPM intra prediction, minus 1.
[0250] More specifically, the number of intra predictions in a combination of intra predictions using GPM prediction is limited to be less than the number of intra predictions in a combination of intra predictions without GPM prediction.
[0251] 19 is a diagram showing an example of limiting the number of intra-prediction modes when GPM prediction and intra-prediction are combined. The diagram shows the number or types of intra-prediction modes (intra-modes) used in GPM prediction. In the diagram, the number of distance options indicated by the distance index and direction index is 8, but it may be 4, 16, or the like.
[0252] 19, in intra prediction without GPM prediction for a current block, the intra prediction mode options derived by the intra prediction parameter derivation unit 304 are NIntraL, which includes, for example, prediction in 65 directions, planar prediction, and DC prediction.
[0253] As shown in the third row of Figure 19, when intra prediction using GPM prediction is performed on a current block, the choice of intra prediction mode is NIntraS. Here, it is preferable that NIntraL > NIntraS. For example, (NIntraL, NIntraS) = (67, 3), (67, 4), (131, 3), (131, 4), etc. may be used.
[0254] For example, the limited choices of intra prediction modes used in the GPM may be the first NIntraS in mpmCandList[ ].
[0255] For example, the limited options for intra prediction modes used in GPM may be any three of the intra modes used for blocks adjacent to the current block. The adjacent blocks may be adjacent blocks located above, to the left, and above left of the current block. That is, mpmCandList = {IntraPredA, IntraPredB, IntraPredB2} may be used. Here, IntraPredA, IntraPredB, and IntraPredB2 may be values of IntraPredModeY at positions A1, B1, and B2, respectively. Alternatively, the limited options may be any two of the intra modes used for blocks adjacent to the current block and planar prediction. mpmCandList = {IntraPredA, IntraPredB, Planar} may be used, or mpmCandList[] = {Planar, IntraPredA, IntraPredB} may be used.
[0256] The above example can also be expressed as follows: In GPM prediction, when the inter-predicted image generation unit 309 generates at least one of the first predicted images using intra prediction, the type of intra prediction mode derived by the GPM prediction unit 30377 is at least one of Planar and the intra mode used for the block adjacent to the current block.
[0257] As another example of the limited options, the options for the intra prediction mode may be limited to, for example, three options including planar prediction, horizontal prediction (IntraPredModeH), and vertical prediction (IntraPredModeV). mpmCandList[] = {Planar, IntraPredModeV, IntraPredModeH} may be used. IntraPredModeV = 50, and IntraPredModeH = 18 may be used.
[0258] The above example can also be expressed as follows: In GPM prediction, when the inter-prediction image generation unit 309 generates at least one of the first predicted images using intra prediction, the type of intra prediction mode derived by the GPM prediction unit 30377 is at least one of planar prediction, horizontal prediction, and vertical prediction.
[0259] Furthermore, in the above example, the number of options for the intra prediction mode is limited to three, but the number of options for the intra prediction mode to be limited may be changed as appropriate.
[0260] (Example of limiting the number of intra predictions when performing MPM in GPM prediction) In GPM prediction, when intra prediction is used for one partition and inter prediction is performed for the other partition, the number of intra prediction mode options when MMVD is used in GPM inter prediction may be smaller than the number of intra prediction mode options when MMVD is not used in GPM inter prediction.
[0261] That is, when decoding a syntax element indicating the intra mode of the GPM intra mode, the maximum value cMax of the syntax element gpm_intra_idxKK may be cMax = gpm_intra_flag ? NIntraMMVD : NIntraS, where NIntraMMVD < NIntraS. For example, NIntraMMVD = 2 and NIntraS = 3.
[0262] (1.4.2: Example 2 of Restriction on the Number of Intra-Predictions in GPM Prediction) As another example of restricting intra-predictions in GPM prediction, a restriction may be made to prohibit prediction modes other than those included in the MPM candidate list mpmCandList[] (MPM: Most Probable Mode). In other words, in this restriction example, the use of non-MPMs is prohibited. For example, as a syntax element for intra-prediction not used in GPM, a syntax element intra_luma_mpm_idx restricted to cMax = NMpmL - 1 may be decoded. As a syntax element indicating the intra-prediction mode for intra-prediction used in GPM, a syntax element gpm_intra_idxKK restricted to cMax = NIntraS - 1 may be decoded. Here, NIntraS < NMpmL < NIntraL may be satisfied. For example, NIntraS = 3 and NMpmL = 5 may be satisfied.
[0263] For example, in GPM prediction, when the inter-prediction image generation unit 309 generates at least one of the first predicted images using intra prediction, the type of intra prediction mode derived by the GPM prediction unit 30377 may be limited to a primary most probable mode (MPM). Here, the primary MPM is a list of intra modes derived from intra prediction modes included in blocks adjacent to the current block, and includes at least an intra prediction mode (IntraPredModeA) included in a block adjacent to the left of the current block and an intra prediction mode (IntraPredModeA) included in a block adjacent above. Note that when deriving a secondary most probable mode (an MPM different from the primary most probable mode) whose elements do not overlap with the primary most probable mode, the secondary most probable mode is not used in the intra prediction mode predicted by GPM prediction. That is, the primary most probable mode and secondary primary most probable mode may be used to derive intra prediction modes other than GPM, and the primary most probable mode may be used to derive intra prediction modes for GPM.
[0264] Furthermore, in GPM prediction, when the inter-predicted image generation unit 309 generates at least one of the first predicted images using intra prediction, the types of intra prediction modes derived by the GPM prediction unit 30377 may be predictions from the first in the MPM candidate list up to a predetermined order. Note that the number of predictions from the first in the predetermined order is less than the number of MPMs used in intra prediction without GPM prediction. For example, as a syntax element for intra prediction not used in GPM, a syntax element intra_luma_mpm_idx limited to cMax = NMpmL - 1 may be decoded. As a syntax element indicating the intra prediction mode for intra prediction used in GPM, a syntax element gpm_intra_idxKK limited to cMax = NIntraS - 1 may be decoded. Here, NIntraS < NMpmL < NIntraL may be satisfied. For example, NIntraS = 3 and NMpmL = 5 may be satisfied.
[0265] (1.4.3: Example 3 of limiting the number of intra predictions in GPM prediction) In GPM prediction, when the inter prediction image generation unit 309 generates at least one of the first predicted images using intra prediction, the number of types of intra prediction modes derived by the GPM prediction unit 30377 may be a number corresponding to the block size of the target block.
[0266] For example, if the block size of the target block is smaller than a predetermined block size, the number of types of intra prediction modes may be reduced, and if the block size of the target block is equal to or larger than the predetermined block size, the number of types of intra prediction modes may be increased. As an example, if the block size of the target block is 8x8 or larger, the number of types of intra prediction modes may be set to three, and if the block size of the target block is smaller than 8x8, the number of types of intra prediction modes may be set to two.
[0267] Alternatively, if the block size of the current block is smaller than a predetermined block size, the number of types of intra prediction modes may be increased, and if the block size of the current block is equal to or larger than the predetermined block size, the number of types of intra prediction modes may be decreased. For example, if the block size of the current block is 32x32 or larger, the number of types of intra prediction modes may be set to three, and if the block size of the current block is smaller than 32x32, the number of types of intra prediction modes may be set to six.
[0268] Alternatively, if the block size of the current block is within a predetermined range, the number of types of intra prediction modes may be increased, and if the block size of the current block is outside the predetermined range, the number of types of intra prediction modes may be decreased. For example, if the block size of the current block is equal to or greater than 8x8 and less than 32x32, the number of types of intra prediction modes may be set to 6, and if the block size of the current block is less than 8x8 or equal to or greater than 32x32, the number of types of intra prediction modes may be set to 3.
[0269] (1.5: Example of Restriction on Derivation of Predicted Image in Combination of MMVD Prediction and Intra Prediction in GPM Prediction) In this example, the method of deriving a predicted image in combination of MMVD prediction and intra prediction in GPM prediction is restricted.
[0270] For example, in GPM prediction, the choice of filter coefficients of the filter used to generate a predicted image derived by the GPM prediction unit 30377 is different from the choice of filter coefficients of the filter used to generate a predicted image derived in a mode in which a predicted image of a current block is generated without using GPM prediction. As an example, in GPM prediction, the number of choice of filter coefficients of the filter used to generate a predicted image derived by the GPM prediction unit 30377 may be fewer than the number of choice of filter coefficients of the filter used to generate a predicted image derived in a mode in which a predicted image of a current block is generated without using GPM prediction.
[0271] For example, the filter may be the loop filter 305. As an example, in GPM prediction, the number of filter coefficient options for the filter used to generate the predicted image derived by the GPM prediction unit 30377 may be one. (2) The GPM prediction unit 30377 corrects the motion vector mvLXN using the method described in (MMVD prediction unit 30376) with distance_idx=gpmMMVDStep0 and direction_idx=gpmMMVDDir0.
[0272] These motion information are referenced to generate predicted images for the two regions.
[0273] For each region, when intra prediction is used for the region, the GPM prediction unit 30377 derives the intra prediction mode using the syntax elements merge_gpm_idx0 and merge_gpm_idx1. merge_gpm_idx0 and merge_gpm_idx1 are indices that select an intra prediction mode to apply to the GPM partition from a candidate list of intra prediction modes. The number of choices for intra prediction modes is, for example, three. This is not necessarily the same as the number of merge candidates used in inter prediction, so merge_gpm_idx0 and merge_gpm_idx1 may be decoded using a method different from that for the inter prediction mode. The GPM prediction unit 30377 uses the intra prediction image generation unit 310 to generate a predicted image corresponding to the intra prediction partition.
[0274] (Weighting Coefficient Derivation Process in GPM Prediction) The GPM synthesis unit 30952 derives the weighting prediction coefficient wValue to be applied to two regions using the following procedure. Here, nCbW = cbWidth, and nCbH = cbHeight. First, we will explain the weighting coefficient derivation process in GPM prediction and preprocessing for motion vector storage process, which will be described later.
[0275] 20 is a diagram showing an example of a table used for GPM prediction. angleIdx and distanceIdx are derived according to the table shown in the upper part of FIG.
[0276] The GPM synthesis unit 30952 derives variables as follows. nW = ( cIdx = = 0 ) ? nCbW : nCbW * SubWidthC nH = ( cIdx = = 0 ) ? nCbH : nCbH * SubHeightC shift1 = Max( 5, 17 - BitDepth ) offset1 = 1 << ( shift1 - 1 ) displacementX = angleIdx displacementY = ( angleIdx + 8 ) % 3 partFlip = ( angleIdx >= 13 && angleIdx <= 27 ) ? 0 : 1 shiftHor = ( angleIdx % 16 = = 8 | | ( angleIdx % 16 != 0 && nH >= nW )) ? 0 : 1 The GPM synthesis unit 30952 derives the variables offsetX and offsetY as follows. If shiftHor == 0 offsetX = ( -nW ) >> 1 offsetY = ( ( -nH ) >> 1 ) + ( angleIdx < 16 ? ( distanceIdx * nH ) >> 3 : - ( ( distanceIdx * nH ) >> 3 ) ) Otherwise (shiftHor == 1 ) offsetX = ( ( -nW ) >> 1 ) + ( angleIdx < 16 ? ( distanceIdx * nW ) >> 3 : - ( ( distanceIdx * nW ) >> 3 ) ) offsetY = ( - nH ) >> 1 The GPM synthesis unit 30952 derives weightIdx and wValue for each x and y using the disLut table.xL = ( cIdx == 0 ) ? x : x * SubWidthC yL = ( cIdx == 0 ) ? y : y * SubHeightC weightIdx = ( ( ( xL + offsetX ) << 1 ) + 1 ) * disLut[ displacementX ] + ( ( ( yL + offsetY ) << 1 ) + 1 ) * disLut[ displacementY ] weightIdxL = partFlip ? 32 + weightIdx : 32 - weightIdx wValue = Clip3( 0, 8, ( weightIdxL + 4 ) >> 3 ) The GPM prediction unit 30377 generates Pred using wValue as follows: A predicted image Pred is derived for x = 0..nCbW - 1, y = 0..nCbH - 1.
[0277] Pred[x][y] = Clip3(0, (1 << bitDepth) - 1, (PredLA[x][y] * (wValue) + PredLB[x][y] * (8 - wValue) + offset1) >> shift1) Here, Pred is a prediction block of size cbWidth*cbHeight. PredLA and PredLB are prediction images generated by the motion compensation unit 3091 using the motion information of areas A and B.
[0278] (Motion Vector Storing Process in GPM Prediction) The GPM prediction unit 30377 stores the motion vectors of areas A and B in memory in units of 4*4 sub-blocks in the following procedure so that they can be referenced in subsequent processes.
[0279] The GPM prediction unit 30377 sets numSbX = cbWidth >> 2 and numSbY = cbHeight >> 2. numSbX and numSbY are the numbers of 4*4 sub-blocks in the horizontal and vertical directions of the current block, respectively.
[0280] The GPM prediction unit 30377 derives displacementX, displacementY, isFlip, and shiftHor. displacementX = angleIdx displacementY = ( angleIdx + 8 ) % 32 isFlip = ( angleIdx >= 13 && angleIdx <= 27 ) ? 1 : 0 shiftHor = ( angleIdx % 16 = = 8 | | ( angleIdx % 16 != 0 && cbHeight >= cbWidth ) ) ? 0 : 1 The GPM prediction unit 30377 derives offsetX and offsetY according to shiftHor. If shiftHor == 1 offsetX = ( -cbWidth ) >> 1 offsetY = ( -cbHeight ) >> 1 ) +( angleIdx < 16 ? ( distanceIdx * cbHeight ) >> 3 : -( ( distanceIdx * cbHeight ) >> 3 ) ) Otherwise (shiftHor == 0) offsetX = ( -cbWidth ) >> 1 ) + ( angleIdx < 16 ? ( distanceIdx * cbWidth ) >> 3 : -( ( distanceIdx * cbWidth ) >> 3 ) ) offsetY = ( -cbHeight ) >> 1 The GPM prediction unit 30377 calculates xSbIdx = 0..numSbX-1, ySbIdx = For each position (xSbIdx, ySbIdx) of the 4*4 sub-blocks from 0 to numSbY-1, the following process is performed.
[0281] The GPM prediction unit 30377 uses disLut, offsetX, and offsetY to calculate motionIdx of xSbIdx and ySbIdx as follows:
[0282] motionIdx = ( ( ( 4 * xSbIdx + offsetX ) << 1 ) + 5 ) * disLut[ displacementX ] + ( ( ( 4 * ySbIdx + offsetY ) << 1 ) + 5 ) * disLut[ displacementY ] The GPM prediction unit 30377 derives sType as follows:
[0283] sType = (abs(motionIdx) < 32) ? 2 : ((motionIdx <= 0) ? partIdx : 1 - partIdx) If sType is 0, the GPM prediction unit 30377 performs the following.
[0284] When the prediction list flag for area A is 0 (predListFlagA==0), the GPM predictor 30377 stores the motion vector of A in L0 as unidirectional prediction. When the prediction list flag for A is not 0 (predListFlagA!=0), the GPM predictor 30377 stores the motion vector of A in L1 as unidirectional prediction.
[0285] predFlagL0 = (predListFlagA == 0) ? 1 : 0 predFlagL1 = (predListFlagA == 0) ? 0 : 1 refIdxL0 = (predListFlagA == 0) ? refIdxA : -1 refIdxL1 = (predListFlagA == 0) ? -1 : refIdxA mvL0[0] = (predListFlagA == 0) ? mvA[0] : 0 mvL0[1] = (predListFlagA == 0) ? mvA[1] : 0 mvL1[0] = (predListFlagA == 0) ? 0 : mvA[0] mvL1[1] = (predListFlagA == 0) ? 0 : mvA[1] Otherwise, if sType is 1, or if sType is 2 and predListFlagA+predListFlagB is not 1, then the GPM predictor 30377 performs the following: where predListFlagA+predListFlagB is not 1 indicates that the reference picture lists of A and B are the same.
[0286] If the prediction list flag for B is 0 (predListFlagB==0), the GPM predictor 30377 stores the motion vector of B in L0 as unidirectional prediction. If the prediction list flag for B is not 0 (predListFlagB!=0), the GPM predictor 30377 stores the motion vector of B in L1 as unidirectional prediction.
[0287] predFlagL0 = (predListFlagB == 0) ? 1 : 0 predFlagL1 = (predListFlagB == 0) ? 0 : 1 refIdxL0 = (predListFlagB == 0) ? refIdxB : -1 refIdxL1 = (predListFlagB == 0) ? -1 : refIdxB mvL0[0] = (predListFlagB == 0) ? mvB[0] : 0 mvL0[1] = (predListFlagB == 0) ? mvB[1] : 0 mvL1[0] = (predListFlagB == 0) ? 0 : mvB[0] mvL1[1] = (predListFlagB == 0) ? 0 : mvB[1] Otherwise (if sType is 2 and predListFlagA+predListFlagB is 1), the GPM predictor 30377 performs the following: where predListFlagA+predListFlagB being 1 indicates that the reference picture lists of A and B are different.
[0288] If the prediction list flag for A is 0 (predListFlagA==0), the GPM predictor 30377 stores the motion vector of A in L0 and sets bidirectional prediction with the motion vector of B in L1. If the prediction list flag for A is not 0 (predListFlagA!=0), the GPM predictor 30377 stores the motion vector of B in L0 and sets bidirectional prediction with the motion vector of A in L1.
[0289] predFlagL0 = 1 predFlagL1 = 1 refIdxL0 = (predListFlagA == 0) ? refIdxA : refIdxB refIdxL1 = (predListFlagA == 0) ? refIdxB : refIdxA mvL0[0] = (predListFlagA == 0) ? mvA[0] : mvB[0] mvL0[1] = (predListFlagA == 0) ? mvA[1] : mvB[1] mvL1[0] = (predListFlagA == 0) ? mvB[0] : mvA[0] mvL1[1] = (predListFlagA == 0) ? mvB[1] : mvA[1] (weight prediction) The weighted prediction unit 30953 generates a predicted image for the block by multiplying the interpolated image PredLX by a weighting factor. The weighting factor is set for each reference image and each luma and chroma channel, and is coded in the slice header or picture header.
[0290] The weighted prediction unit 30953 generates a predicted image of the block by multiplying the interpolated image PredLX by a weighting coefficient. When one of the prediction list usage flags (predFlagL0 or predFlagL1) is 1 (uni-prediction) and weighted prediction is not used, the weighted prediction unit 30953 performs the following equation processing to adjust PredLX (LX is L0 or L1) to the pixel bit depth bitDepth.
[0291] Pred[x][y] = Clip3(0,(1<<bitDepth)-1,(PredLX[x][y]+offset1)> >shift1) where shift1=14-bitDepth, offset1=1<<(shift1-1). Furthermore, when both prediction list usage flags (predFlagL0 and predFlagL1) are 1 (bi-prediction PRED_BI) and weighted prediction is not used, PredL0 and PredL1 are averaged and adjusted to the number of pixel bits according to the following equation.
[0292] Pred[x][y] = Clip3(0,(1<<bitDepth)-1,(PredL0[x][y]+PredL1[x][y]+offset2)> >shift2) where shift2=15-bitDepth and offset2=1<<(shift2-1).
[0293] Furthermore, when uni-prediction and weighted prediction are performed, the weighted prediction unit 30953 derives a weighted prediction coefficient w0 and an offset o0 from the coded data, and performs processing according to the following equations.
[0294] Pred[x][y] = Clip3(0,(1<<bitDepth)-1,((PredLX[x][y]*w0+2^(log2WD-1))> >log2WD)+o0) where log2WD is a variable indicating a predetermined shift amount.
[0295] Furthermore, when performing bi-prediction PRED_BI and weighted prediction, the weighted prediction unit 30953 derives weighted prediction coefficients w0, w1, o0, and o1 from the coded data, and performs processing according to the following equations.
[0296] Pred[x][y] = Clip3(0,(1< <bitDepth)-1,(PredL0[x][y]*w0+PredL1[x][y]*w1+((o0+o1+1)<<log2WD))> >(log2WD+1)) The inter predicted image generation unit 309 outputs the generated predicted image of the block to the addition unit 312.
[0297] (Intra Prediction Parameters) The following describes the prediction parameters for intra prediction. The intra prediction parameters consist of a luminance prediction mode, IntraPredModeY, and a chrominance prediction mode, IntraPredModeC. FIG. 5 is a schematic diagram showing the types of intra prediction modes (mode numbers). For example, there are planar prediction (0), DC prediction (1), and angular prediction (others). Furthermore, CCLM modes (81 to 83) may be added for chrominance. Template-based intra mode derivation (TIMD) prediction and decoder-side intra mode derivation (DIMD) prediction may also be used.
[0298] To derive intra prediction parameters, for example, intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_remainder, etc. are decoded on a CU-by-CU basis. Also, a flag intra_luma_not_planar_flag indicating planar prediction and an index intra_luma_ref_idx for selecting a reference line may be included.
[0299] (MPM) intra_luma_mpm_flag is a flag indicating whether the IntraPredModeY of the current block matches the MPM (Most Probable Mode). The MPM is a prediction mode included in the MPM candidate list mpmCandList[]. The MPM candidate list is a list that stores candidates that are estimated to have a high probability of being applied to the current block based on the intra prediction modes of neighboring blocks and predetermined intra prediction modes. When intra_luma_mpm_flag is 1, the IntraPredModeY of the current block is derived using the MPM candidate list and intra_luma_mpm_idx.
[0300] IntraPredModeY = mpmCandList[intra_luma_mpm_idx] (REM) When intra_luma_mpm_flag is 0, IntraPredModeY is derived using intra_luma_mpm_remainder. Specifically, IntraPredModeY is selected from the remaining modes RemIntraPredMode obtained by excluding the intra prediction modes included in the MPM candidate list from all intra prediction modes.
[0301] (LM Prediction) The LM prediction unit 31044 predicts chrominance pixel values based on luminance pixel values. Specifically, this is a method of generating a predicted image of a chrominance image (Cb, Cr) using a linear model based on the decoded luminance image. LM prediction includes CCLM (Cross-Component Linear Model prediction) prediction and MMLM (Multiple Model ccLM) prediction. CCLM prediction is a prediction method that uses one linear model to predict chrominance from luminance for one block. MMLM prediction is a prediction method that uses two or more linear models to predict chrominance from luminance for one block.
[0302] (Matrix Intra Prediction) The MIP unit 31045 generates a provisional predicted image Pred[x][y] by multiplying and accumulating the reference sample rec[x][y] derived from the adjacent region and a weighting matrix, and outputs the generated image to the predicted image correction unit 3105.
[0303] (Intra-prediction image generation unit 310) When predMode indicates an intra-prediction mode, the intra-prediction image generation unit 310 performs intra-prediction using the intra-prediction parameters input from the intra-prediction parameter derivation unit 304 and reference pixels read from the reference picture memory 306.
[0304] Specifically, the intra-prediction image generation unit 310 reads neighboring blocks within a predetermined range from the current block in the current picture from the reference picture memory 306. The predetermined range refers to neighboring blocks to the left, upper left, lower left, upper, and upper right of the current block, and the area to be referenced differs depending on the intra-prediction mode.
[0305] The intra-predicted image generation unit 310 generates a predicted image of the current block by referring to the read decoded pixel values and the prediction mode indicated by IntraPredMode. The intra-predicted image generation unit 310 outputs the generated predicted image of the block to the adder 312.
[0306] Generation of a predicted image based on an intra prediction mode will be described below. In planar prediction, DC prediction, and angular prediction, a decoded surrounding region adjacent (close to) the block to be predicted is set as a reference region R. Then, a predicted image is generated by extrapolating pixels in the reference region R in a specific direction. For example, the reference region R may be set as an L-shaped region including the left and top of the block to be predicted (or further including the top left, top right, and bottom left).
[0307] (Details of Prediction Image Generation Unit) Next, we will explain the details of the configuration of the intra-prediction image generation unit 310. The intra-prediction image generation unit 310 includes a reference sample filter unit, a prediction unit, and a prediction image correction unit (prediction image correction unit, filter switching unit, and weighting coefficient change unit).
[0308] The prediction unit generates a temporary predicted image (pre-corrected predicted image) of the block to be predicted based on each reference pixel (reference image) in the reference region R, a filtered reference image generated by applying a reference sample filter (first filter), and the intra prediction mode, and outputs the generated image to the predicted image correction unit. The predicted image correction unit corrects the temporary predicted image according to the intra prediction mode, and generates and outputs a predicted image (corrected predicted image).
[0309] Each unit included in the intra-predicted image generation unit 310 will be described below.
[0310] (Reference Sample Filter Unit) The reference sample filter unit derives a reference sample s[x][y] at each position (x, y) on the reference region R by referring to the reference image. Furthermore, the reference sample filter unit applies a reference sample filter (first filter) to the reference sample rec[][] according to the intra prediction mode to update the reference sample rec[x][y] at each position (x, y) on the reference region R (derives a filtered reference image rec[x][y]). Specifically, a low-pass filter is applied to the reference image at and around the position (x, y) to derive a filtered reference image. Note that it is not necessarily necessary to apply a low-pass filter to all intra prediction modes, and a low-pass filter may be applied to some intra prediction modes. Note that the filter applied to the reference image on the reference region R in the reference sample filter unit is referred to as a "reference pixel filter (first filter)," whereas the filter that corrects the tentative predicted image in the predicted image correction unit (described later) is referred to as a "position-dependent filter (second filter)."
[0311] (Configuration of Intra Prediction Unit) The intra prediction unit generates a provisional predicted image (provisional predicted pixel value, uncorrected predicted image) of the block to be predicted based on the intra prediction mode, the reference image, and the filtered reference pixel value, and outputs it to the predicted image correction unit. The prediction unit internally includes a planar prediction unit, a DC prediction unit, an angular prediction unit, and an LM prediction unit. The prediction unit selects a specific prediction unit depending on the intra prediction mode, and inputs the reference image and the filtered reference image. The relationship between the intra prediction mode and the corresponding prediction unit is as follows:・Planar prediction...Planar prediction unit ・DC prediction...DC prediction unit ・Angular prediction...Angular prediction unit ・LM prediction...LM prediction unit ・Matrix intra prediction...MIP unit ・DIMD prediction...DIMD prediction unit ・TIMD prediction...TIMD prediction unit (Planar prediction) The planar prediction unit generates a provisional predicted image by linearly adding the reference samples rec[x][y] according to the distance between the pixel position to be predicted and the reference pixel position, and outputs this to the predicted image correction unit.
[0312] (DC Prediction) The DC prediction unit derives a DC predicted value equivalent to the average value of the reference samples rec[x][y], and outputs a provisional predicted image Pred[x][y] whose pixel values are the DC predicted values.
[0313] (Angular prediction) The angular prediction unit generates a temporary predicted image Pred[x][y] using reference samples rec[x][y] in the prediction direction (reference direction) indicated by the intra prediction mode, and outputs it to the predicted image correction unit. When IntraPredMode >= DIR, ref[x] = rec[-1-refIdx+x][-1-refIdx] (x=0..bW+refIdx+1) Furthermore, the following is performed for x=0..bW-1, y = 0..bH-1.
[0314] iIdx = (((y+1+refIdx) * intraPredAngle) >> 5) + refIdx iFact = ((y + 1 + refIdx) * intraPredAngle) & 31 Pred[x][y] = Clip1(((Σ(fT[i]*ref[x+iIdx+i])) + 32) >> 6) Otherwise (IntraPredMode < DIR) ref[x] = rec[-1-refIdx][-1-refIdx+x](x=0..bH+refIdx+1) Furthermore, do the following for x=0..bW-1, y = 0..bH-1.
[0315] iIdx = (((x+1+refIdx) * intraPredAngle) >> 5) + refIdx iFact = ((x + 1 + refIdx) * intraPredAngle) & 31 Pred[x][y] = Clip1(((Σ(fT[i]*ref[y+iIdx+i])) + 32) >> 6) DIR and refIdx are predetermined constants, for example, DIR=34, 66, etc. refIdx=0, 1, 2, etc. In the case of normal angular prediction, refIdx may be set by decoding the syntax of the coded data. fT is an interpolation filter coefficient for intra-predicted images. In addition, in the case of generating predicted images for TIMD prediction, which will be described later, refIdx may be fixed at 0. In addition, when used to generate a template predicted image, refIdx=2 or 4 may be used.
[0316] (Configuration of predicted image correction unit) The predicted image correction unit corrects the temporary predicted image output from the prediction unit according to the intra prediction mode. Specifically, the predicted image correction unit derives a position-dependent weighting coefficient for each pixel of the temporary predicted image according to the reference region R and the position of the target predicted pixel. Then, the predicted image correction unit performs weighted addition (weighted averaging) of the reference sample s[][] and the temporary predicted image to derive a predicted image (corrected predicted image) Pred[][] obtained by correcting the temporary predicted image. Note that in some intra prediction modes, the predicted image correction unit 3105 may not correct the temporary predicted image, and the output of the prediction unit may be used as the predicted image as is.
[0317] The inverse quantization and inverse transform unit 311 inverse quantizes the quantized transform coefficients input from the parameter decoding unit 302 to obtain transform coefficients.
[0318] The inverse quantization and inverse transform unit 311 includes a scaling unit (inverse quantization unit), a secondary transform unit, and a core transform unit.
[0319] The scaling unit derives the scaling factor ls[x][y] using the quantization matrix m[x][y] or the uniform matrix m[x][y]=16 decoded from the encoded data, the quantization parameter qP, and the rectNonTsFlag derived from the TU size.
[0320] When dependency quantization is used (dep_quant_enabled_flag==1), ls[x][y] may be derived using the following formula:
[0321] ls[x][y] = (m[x][y] * levelScale[rectNonTsFlag][(qP+1)%6]) << (qP / 6) In other cases, it may be derived using the following formula.
[0322] ls[x][y] = (m[x][y] * levelScale[rectNonTsFlag][qP%6]) << (qP / 6) where levelScale[] = { { 40, 45, 51, 57, 64, 72}, { 57, 64, 72, 81, 91, 102}}.
[0323] rectNonTsFlag = (((Log2(nTbW) + Log2(nTbH)) & 1) == 1 && transform_skip_flag == 0) The scaling unit 31112 derives dnc[][] from the product of ls[][] and the transform coefficient TransCoeffLevel, and performs inverse quantization. It then derives d[][] by clipping.
[0324] ) >> 1 The secondary transform unit restores the modified transform coefficients d[ ][ ] (the transform coefficients after transformation by the second transform unit) by applying a transform using a transform matrix to some or all of the transform coefficients d[ ][ ] received from the scaling unit. The secondary transform unit applies the secondary transform to the transform coefficients d[ ][ ] of a specified unit for each TU. The secondary transform is applied only to the intra CU, and the transform base is determined by referring to stIdx and IntraPredMode. The secondary transform unit outputs the restored modified transform coefficients d[ ][ ] to the core transform unit.
[0325] The core transform unit transforms the transform coefficients d[ ][ ] or the modified transform coefficients d[ ][ ] using the selected transform matrix to derive prediction errors r[ ][ ]. The core transform unit outputs the prediction errors r[ ][ ] (resSamples[ ][ ]) to the adder 312. Note that the inverse quantization and inverse transform unit 311 sets all prediction errors of the current block to 0 when skip_flag is 1 or cu_cbp is 0. The transform matrix may be selected from multiple transform matrices using mts_idx.
[0326] The prediction error d[][] after the core transformation may be further shifted to have the same accuracy as the predicted image Pred[][] to derive resSamples[][].
[0327] resSamples[x][y] = (r[x][y] + (1<<(bdShift-1))) >> bdShift bdShift = Max( 20 - bitDepth, 0 ) When joint coding is performed, the prediction error resSamples[][] (e.g., resSamplesCb[][]) of the first color component (cIdx=cIdx0) may be used to derive the prediction error resSamples[][] of the second color component (e.g., cIdx=cIdx1, resSamplesCr[][]). cIdx0 and cIdx1 may be 1, 2 (deriving Cr from Cb) or 2, 1 (deriving Cb from Cr). For example, the following processing may be performed.
[0328] resSamplesCr[x][y] = -(resSamplesCb[x][y])>>1 Furthermore, for stereo encoding, resSamplesCb[][] and resSamplesCr[][] may be derived by adding or subtracting resSamplesCb[][] and resSamplesCr[][].
[0329] resSamples[x][y] = resSamplesCb[x][y]+resSamplesCr[x][y] (cIdx=cIdx0) resSamples[x][y] = resSamplesCb[x][y]-resSamplesCr[x][y] (cIdx=cIdx1) The adder 312 adds the predicted image Pred of the block input from the predicted image generation unit 308 and the prediction error resSamples input from the inverse quantization and inverse transform unit 311 for each pixel to generate a decoded image rec of the block.
[0330] rec[x][y]=Pred[x][y]+resSamples[x][y] The adder 312 stores the decoded image of the block in the reference picture memory 306 and also outputs it to the loop filter 305.
[0331] The inter-prediction image generation unit 309 outputs the generated prediction image of the block to the addition unit 312 .
[0332] (Configuration of Video Encoding Device) Next, the configuration of the video encoding device 11 according to this embodiment will be described. Fig. 4 is a block diagram showing the configuration of the video encoding device 11 according to this embodiment. The video encoding device 11 includes a prediction image generation unit 101, a subtraction unit 102, a transformation / quantization unit 103, an inverse quantization / inverse transformation unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, a coding parameter determination unit 110, a parameter coding unit 111, a prediction parameter derivation unit 120, and an entropy coding unit 104.
[0333] The predicted image generation unit 101 generates a predicted image for each CU. The predicted image generation unit 101 includes the inter predicted image generation unit 309 and the intra predicted image generation unit 310, which have already been described, and therefore further description thereof will be omitted.
[0334] The subtraction unit 102 generates a prediction error by subtracting the pixel values of the predicted image of the block input from the predicted image generation unit 101 from the pixel values of the image T. The subtraction unit 102 outputs the prediction error to the transformation / quantization unit 103.
[0335] The transform / quantization unit 103 calculates transform coefficients by frequency transforming the prediction errors input from the subtraction unit 102, and derives quantized transform coefficients by quantizing them. The transform / quantization unit 103 outputs the quantized transform coefficients to the parameter coding unit 111 and the inverse quantization / inverse transform unit 105.
[0336] The transform / quantization unit 103 includes a separation transform unit (first transform unit), a non-separation transform unit (second transform unit), and a scaling unit.
[0337] The separate transform unit applies a separate transform to the prediction error, and the scaling unit scales the transform coefficients with a quantization matrix.
[0338] The inverse quantization and inverse transform unit 105 is the same as the inverse quantization and inverse transform unit 311 in the video decoding device 31, and a description thereof will be omitted. The calculated prediction error is output to the adder .
[0339] The parameter coding unit 111 includes a header coding unit 1110, a CT information coding unit 1111, and a CU coding unit 1112 (prediction mode coding unit). The CU coding unit 1112 further includes a TU coding unit 1114. The following describes an outline of the operation of each module.
[0340] The header encoding unit 1110 performs encoding processing of parameters such as header information, division information, prediction information, and quantized transform coefficients.
[0341] The CT information encoding unit 1111 encodes QT, MT (BT, TT) division information and the like.
[0342] The CU encoding unit 1112 encodes CU information, prediction information, division information, and the like.
[0343] When a prediction error is included in a TU, the TU encoding unit 1114 encodes the QP update information and the quantized prediction error.
[0344] The CT information encoding unit 1111 and the CU encoding unit 1112 supply syntax elements such as inter prediction parameters (predMode, general_merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX, amvr_flag, amvr_precision_idx), intra prediction parameters, and quantized transform coefficients to the parameter encoding unit 111.
[0345] The entropy coding unit 104 receives the quantized transform coefficients and coding parameters (division information, prediction parameters) from the parameter coding unit 111. The entropy coding unit 104 entropy codes these to generate and output a coded stream Te.
[0346] The prediction parameter derivation unit 120 is a means including the inter prediction parameter coding unit 112 and an intra prediction parameter coding unit, and derives intra prediction parameters and inter prediction parameters from the parameters input from the coding parameter determination unit 110. The derived intra prediction parameters and inter prediction parameters are output to the parameter coding unit 111.
[0347] (Configuration of Inter Prediction Parameter Encoding Unit) The inter prediction parameter encoding unit 112 includes a parameter encoding control unit 1121 and an inter prediction parameter derivation unit 303. The inter prediction parameter derivation unit 303 has the same configuration as the video decoding device. The parameter encoding control unit 1121 includes a merge index derivation unit 11211 and a vector candidate index derivation unit 11212.
[0348] The merge index derivation unit 11211 derives merge candidates and the like, and outputs them to the inter prediction parameter derivation unit 303. The vector candidate index derivation unit 11212 derives predicted vector candidates and the like, and outputs them to the inter prediction parameter derivation unit 303 and the parameter coding unit 111.
[0349] (Configuration of Intra Prediction Parameter Encoding Unit) The intra prediction parameter encoding unit includes a parameter encoding control unit and an intra prediction parameter derivation unit. The intra prediction parameter derivation unit has the same configuration as the video decoding device.
[0350] However, unlike the video decoding device, the inputs to the inter prediction parameter derivation unit 303 and the intra prediction parameter derivation unit are the coding parameter determination unit 110 and the prediction parameter memory 108 , and output to the parameter coding unit 111 .
[0351] The adder 106 generates a decoded image by adding, for each pixel, the pixel values of the predicted block input from the predicted image generation unit 101 and the prediction errors input from the inverse quantization and inverse transform unit 105. The adder 106 stores the generated decoded image in a reference picture memory 109.
[0352] The loop filter 107 performs deblocking filtering, SAO, and ALF on the decoded image generated by the adder 106. Note that the loop filter 107 does not necessarily have to include the above three types of filters, and may be configured with only a deblocking filter, for example.
[0353] The prediction parameter memory 108 stores the prediction parameters generated by the coding parameter determination unit 110 in a predetermined location for each current picture and CU.
[0354] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at a predetermined location for each current picture and CU.
[0355] The coding parameter determination unit 110 selects one set of coding parameters from among multiple sets of coding parameters. The coding parameters are the above-mentioned QT, BT, or TT division information, prediction parameters, or parameters to be coded that are generated in relation to these. The predicted image generation unit 101 generates a predicted image using these coding parameters.
[0356] The coding parameter determination unit 110 calculates an RD cost value indicating the magnitude of the information amount and the coding error for each of the multiple sets. The RD cost value is, for example, the sum of the code amount and the value obtained by multiplying the squared error by a coefficient λ. The code amount is the information amount of the coded stream Te obtained by entropy coding the quantization error and the coding parameters. The squared error is the sum of the squares of the prediction errors calculated by the subtraction unit 102. The coefficient λ is a preset real number greater than zero. The coding parameter determination unit 110 selects the set of coding parameters that minimizes the calculated cost value. The coding parameter determination unit 110 outputs the determined coding parameters to the prediction parameter derivation unit 120.
[0357] Note that a portion of the video encoding device 11 and the video decoding device 31 in the above-described embodiments, such as the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generation unit 308, the inverse quantization / inverse transform unit 311, the adder 312, the prediction parameter derivation unit 320, the predicted image generation unit 101, the subtractor 102, the transform / quantization unit 103, the entropy encoding unit 104, the inverse quantization / inverse transform unit 105, the loop filter 107, the encoding parameter determination unit 110, the parameter encoding unit 111, and the prediction parameter derivation unit 120, may be implemented by a computer. In this case, a program for implementing this control function may be recorded on a computer-readable recording medium, and the program may be read and executed by a computer system. Note that the term "computer system" used here refers to a computer system built into either the video encoding device 11 or the video decoding device 31, and includes hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. Furthermore, "computer-readable recording media" may also include devices that dynamically store programs for a short period of time, such as communication lines used when transmitting programs over networks like the Internet or communication lines like telephone lines, or devices that store programs for a fixed period of time, such as volatile memory within computer systems that serve as servers or clients in such cases. Furthermore, the programs may be programs that realize some of the aforementioned functions, or may be programs that can realize the aforementioned functions in combination with programs already stored in the computer system.
[0358] Furthermore, part or all of the video encoding device 11 and video decoding device 31 in the above-described embodiments may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the video encoding device 11 and video decoding device 31 may be individually implemented as a processor, or part or all of them may be integrated into a processor. Furthermore, the integrated circuit implementation method is not limited to LSI, and may be implemented using a dedicated circuit or a general-purpose processor. Furthermore, if an integrated circuit implementation technology that can replace LSI emerges due to advances in semiconductor technology, an integrated circuit based on that technology may be used.
[0359] A video encoding device according to one aspect of the present invention is a video encoding device that encodes the residual between a predicted image and an image to be encoded, and is equipped with a motion information derivation unit that derives motion information of a target block and a predicted image generation unit that generates a predicted image of the target block by referring to the motion information, and is characterized in that in a synthesis mode in which the predicted image generation unit generates two first predicted images for the target block and synthesizes a second predicted image of the target block from the two first predicted images using weights determined according to the geometric position of the target block, when the predicted image generation unit generates at least one of the first predicted images using intra prediction, the options for the type of intra prediction mode derived by the motion information derivation unit are different from the options for the type of intra prediction mode derived in a mode in which a predicted image of the target block is generated without using the synthesis mode.
[0360] A video coding device according to one aspect of the present invention is a video coding device that codes a residual between a predicted image and an image to be coded, and is equipped with a motion information derivation unit that derives motion information of a target block and a predicted image generation unit that generates a predicted image of the target block by referring to the motion information, and is characterized in that in a synthesis mode in which the predicted image generation unit generates two first predicted images for the target block and synthesizes a second predicted image of the target block from the two first predicted images using weights determined according to the geometric position of the target block, the filter coefficient options of a filter used to generate the predicted image derived by the motion information derivation unit are different from the filter coefficient options of a filter used to generate the predicted image derived in a mode in which a predicted image of the target block is generated without using the synthesis mode.
[0361] One embodiment of the present invention has been described in detail above with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes and the like are possible within the scope that does not deviate from the gist of the present invention.
[0362] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. In other words, embodiments obtained by combining technical means modified appropriately within the scope of the claims are also included in the technical scope of the present invention.
[0363] The embodiments of the present invention can be suitably applied to a video decoding device that decodes coded data obtained by coding image data, and a video coding device that generates coded data obtained by coding image data, and can also be suitably applied to the data structure of coded data generated by a video coding device and referenced by the video decoding device.
[0364] 31 Video decoding device 301 Entropy decoding unit 302 Parameter decoding unit 3022 CU decoding unit 3024 TU decoding unit 303 Inter prediction parameter derivation unit 305, 107 Loop filter 306, 109 Reference picture memory 307, 108 Prediction parameter memory 308, 101 Prediction image generation unit 309 Inter prediction image generation unit 3092 OOB processing unit 311, 105 Inverse quantization / inverse transform unit 312, 106 Addition unit 320 Prediction parameter derivation unit 11 Video encoding device 102 Subtraction unit 103 Transform / quantization unit 104 Entropy encoding unit 110 Encoding parameter determination unit 111 Parameter encoding unit 112 Inter prediction parameter encoding unit 120 Prediction parameter derivation unit
Claims
1. A video decoding device that decodes encoded data, comprising: a motion information derivation unit that derives motion information of a current block; and a predicted image generation unit that generates a predicted image of the current block by referring to the motion information, wherein the predicted image generation unit generates two first predicted images for the current block, and in a synthesis mode that synthesizes a second predicted image of the current block from the two first predicted images using weights determined according to the geometric position of the current block, when at least one of the first predicted images is generated using intra prediction, the options for the type of intra prediction mode are different from the options for the type of intra prediction mode derived in a mode that generates a predicted image of the current block without using the synthesis mode.
2. The video decoding device described in claim 1, characterized in that when at least one of the first predicted images is generated using intra prediction, the number of options for the type of intra prediction mode is less than the number of options for the type of intra prediction mode derived in a mode that generates a predicted image of the target block without using the synthesis mode.
3. The video decoding device according to claim 2, characterized in that when at least one of the first predicted images is generated using intra prediction, the type of intra prediction mode is at least one of Planar and the intra mode used for a block adjacent to the current block.
4. The video decoding device of claim 2, characterized in that when at least one of the first predicted images is generated using intra prediction, the type of intra prediction mode is at least one of planar prediction, horizontal prediction, and vertical prediction.
5. The video decoding device according to claim 2, characterized in that when at least one of the first predicted images is generated using intra prediction, the only type of intra prediction mode is Primary Most Probable Mode (MPM).
6. The video decoding device according to claim 2, characterized in that when at least one of the first predicted images is generated using intra prediction, the types of intra prediction modes are predictions from the first to a predetermined order in a Most Probable Mode (MPM) candidate list.
7. The video decoding device of claim 2, characterized in that when at least one of the first predicted images is generated using intra prediction, the number of options for the type of intra prediction mode corresponds to the block size of the target block.
8. A video decoding device that decodes encoded data, comprising: a motion information derivation unit that derives motion information of a current block; and a predicted image generation unit that generates a predicted image of the current block by referring to the motion information, wherein the predicted image generation unit generates two first predicted images for the current block, and in a synthesis mode that synthesizes a second predicted image of the current block from the two first predicted images using weights determined according to a geometric position in the current block, the filter coefficient options of a filter used to generate the predicted image are different from the filter coefficient options of a filter used to generate the predicted image derived in a mode in which a predicted image of the current block is generated without using the synthesis mode.
9. A video encoding device that encodes a residual between a predicted image and an image to be encoded, comprising: a motion information derivation unit that derives motion information of a current block; and a predicted image generation unit that generates a predicted image of the current block by referring to the motion information, wherein the predicted image generation unit generates two first predicted images for the current block, and in a synthesis mode that synthesizes a second predicted image of the current block from the two first predicted images using weights determined according to a geometric position in the current block, when at least one of the first predicted images is generated using intra prediction, the options for the type of intra prediction mode are different from the options for the type of intra prediction mode derived in a mode that generates a predicted image of the current block without using the synthesis mode.
10. A video coding device that codes the residual between a predicted image and a current image to be coded, comprising: a motion information derivation unit that derives motion information of a current block; and a predicted image generation unit that generates a predicted image of the current block by referring to the motion information, wherein the predicted image generation unit generates two first predicted images for the current block, and in a synthesis mode that synthesizes a second predicted image of the current block from the two first predicted images using weights determined according to the geometric position of the current block, the filter coefficient options of a filter used to generate the predicted image are different from the filter coefficient options of a filter used to generate the predicted image derived in a mode in which a predicted image of the current block is generated without using the synthesis mode.
Citation Information
Patent Citations
Video decoding device and video encoding device
JP2022150312A