Moving image decoding device, moving image encoding device, and computer-readable recording medium
Patent Information
- Application Number
- JP2022101647
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-06-12
AI Technical Summary
【0011】 本発明の態様によれば、動画像符号化·復号処理においてMMVD予測にかかる計算量を削減することができる。
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] An embodiment of the present invention relates to a video decoding device and a video encoding device. [Background technology]
[0002] In order to efficiently transmit or record moving images, a moving image encoding device is used that generates encoded data by encoding moving images, and a moving image decoding device is used that generates a decoded image by decoding the encoded data.
[0003] Specific examples of video encoding methods include H.264 / AVC and High-Efficiency Video Coding (HEVC).
[0004] In such a video coding method, images (pictures) constituting a video are managed in a hierarchical structure consisting of slices obtained by dividing images, coding tree units (CTUs) obtained by dividing slices, coding units (sometimes called coding units: CUs) obtained by dividing coding tree units, and transform units (TUs) obtained by dividing coding units, and are coded / decoded for each CU.
[0005] In such a video coding method, a predicted image is usually generated based on a locally decoded image obtained by encoding / decoding an input image, and a prediction error (sometimes called a "difference image" or "residual image") obtained by subtracting the predicted image from the input image (original image) is coded. Methods for generating a predicted image include inter-prediction and intra-prediction.
[0006] Furthermore, VVC / H.266 discloses MMVD prediction that obtains a motion vector by adding a difference vector of a predetermined distance and a predetermined direction to the motion vector.
[0007] Furthermore, Non-Patent Document 1 discloses a technique for deriving a difference vector with a small amount of coding by calculating template matching costs for all MMVD difference vector candidates and sorting them according to the costs. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] “Non-EE2: Template Matching-based Reordering for Extended MMVD Design”, JVET-X0085, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 24th Meeting, by teleconference Summary of the Invention [Problem to be solved by the invention]
[0009] The method described in Non-Patent Document 1 has a problem in that the amount of calculation is large because the template matching cost is derived for all MMVD difference vector candidates. [Means for solving the problem]
[0010] In order to solve the above problem, a video decoding device according to one aspect of the present invention includes an MMVD prediction unit that obtains a motion vector by adding a difference vector at a predetermined distance and in a predetermined direction to a predicted motion vector of a current block, and an index function that specifies the difference vector from an MMVD candidate list. A video decoding device including a parameter decoding unit that decodes a parameter from encoded data, The MMVD prediction unit derives a difference vector of a specific distance and direction from the predetermined distance and predetermined direction, performs a search by deriving a template matching cost for the difference vector, and derives an MMVD candidate list by inserting difference vector candidates according to the cost; In the search, only a first directional subset (DIR1) is searched at a predetermined distance, and only a second directional subset (DIR2) is searched at distances other than the predetermined distance. Effect of the Invention
[0011] According to an aspect of the present invention, it is possible to reduce the amount of calculation required for MMVD prediction in video encoding / decoding processing. [Brief description of the drawings]
[0012] [Figure 1] 1 is a schematic diagram showing a configuration of an image transmission system according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram showing a hierarchical structure of data in an encoded stream. [Diagram 3] FIG. 1 is a schematic diagram showing a configuration of a video decoding device. [Figure 4] FIG. 13 is a schematic diagram showing a configuration of an inter-prediction image generating unit. [Diagram 5] FIG. 13 is a schematic diagram showing a configuration of an inter-prediction parameter derivation unit. [Figure 6] 11 is a flowchart illustrating a schematic operation of the video decoding device. [Figure 7] Fig. 11 is a diagram showing an example of an index used in the MMVD mode. (a) is a diagram showing an example of an index mmvd_cand_idx indicating an MMVD candidate in an MMVD candidate list. (b) is a diagram showing an example of a block position adjacent to a target block. (c) is a diagram showing an example of mmvd_distance_idx. (d) is a diagram showing an example of mmvd_direction_idx. [Figure 8]11A and 11B are diagrams illustrating an example of the number of search distance candidates and the number of derivation direction candidates in the MMVD mode. [Figure 9] FIG. 13 is a syntax diagram for merge prediction and MMVD prediction. [Figure 10] FIG. 13 is a syntax diagram for another configuration of MMVD prediction. [Figure 11] 13 is a flowchart showing a process flow in another configuration of MMVD prediction. [Figure 12] 13 is a flowchart showing the flow of processing in another configuration of the MMVD candidate list derivation processing. [Figure 13] 13 is another flowchart showing the process flow in another configuration of the MMVD candidate list derivation process. [Figure 14] FIG. 13 is a diagram showing examples of 16 direction candidates in MMVD candidates. [Figure 15] This figure shows four distances for 16 direction candidates in MMVD candidates. [Figure 16] FIG. 1 is a block diagram showing a configuration of a video encoding device. [Figure 17] 13 is a schematic diagram showing a configuration of an inter-prediction parameter encoding unit. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] (First embodiment) Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0014] FIG. 1 is a schematic diagram showing the configuration of an image transmission system 1 according to this embodiment.
[0015] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding a target image, decodes the transmitted encoded stream, and displays an image. The image transmission system 1 includes a video encoding device (image encoding device) 11, a network 21, a video decoding device (image decoding device) 31, and a video display device (image display device) 41.
[0016] An image T is input to the video encoding device 11 .
[0017] The network 21 transmits the encoded stream Te generated by the video encoding device 11 to the video decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 21 is not necessarily limited to a bidirectional communication network, and may be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. The network 21 may also be replaced by a storage medium on which the encoded stream Te is recorded, such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark).
[0018] The video decoding device 31 decodes each of the coded streams Te transmitted by the network 21, and generates one or more decoded images Td.
[0019] The video display device 41 displays all or part of one or more decoded images Td generated by the video decoding device 31. The video display device 41 includes a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. The display may be in the form of a stationary display, a mobile display, an HMD, or the like. When the video decoding device 31 has high processing power, it displays high quality images, and when it has only low processing power, it displays images that do not require high processing power or display power.
[0020] <operator> The operators used in this specification are listed below.
[0021] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, | is a bitwise OR, |= is the OR assignment operator, and || indicates logical OR.
[0022] x?y:z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0).
[0023] Clip3(a, b, c) is a function that clips c to a value between a and b (inclusive). It returns a if c < a, b if c > b, and c otherwise (assuming a <= b).
[0024] abs(a) is a function that returns the absolute value of a.
[0025] Int(a) is a function that returns the integer value of a.
[0026] floor(a) is a function that returns the largest integer less than or equal to a.
[0027] ceil(a) is a function that returns the smallest integer greater than or equal to a.
[0028] a / d represents the division of a by d, rounded down to the nearest integer.
[0029] <Structure of the Encoded Stream Te> Prior to the detailed description of the moving image encoding apparatus 11 and the moving image decoding apparatus 31 according to the present embodiment, the data structure of the encoded stream Te generated by the moving image encoding apparatus 11 and decoded by the moving image decoding apparatus 31 will be described.
[0030] FIG. 2 is a diagram showing the hierarchical structure of the data in the encoded stream Te. The encoded stream Te illustratively includes a sequence and a plurality of pictures that make up the sequence. FIGS. 2(a) to 2(f) respectively show an encoded video sequence that defines the sequence SEQ, an encoded picture that defines the picture PICT, an encoded slice that defines the slice S, encoded slice data that defines the slice data, an encoded tree unit included in the encoded slice data, and an encoded unit included in the encoded tree unit.
[0031] (Coded Video Sequence) The coded video sequence defines a set of data that the video decoding device 31 refers to in order to decode the sequence SEQ to be processed. As shown in Fig. 2, the sequence SEQ includes a video parameter set, a sequence parameter set SPS (Sequence Parameter Set), a picture parameter set PPS (Picture Parameter Set), an adaptation parameter set (APS), a picture PICT, and supplemental enhancement information SEI (Supplemental Enhancement Information).
[0032] The video parameter set VPS specifies a set of coding parameters common to multiple videos composed of multiple layers, as well as a set of coding parameters related to multiple layers and each individual layer included in the video.
[0033] The sequence parameter set SPS specifies a set of coding parameters that the video decoding device 31 refers to in order to decode the target sequence. For example, the width and height of a picture are specified. Note that there may be multiple SPSs. In that case, one of the multiple SPSs is selected from the PPS.
[0034] The picture parameter set PPS specifies a set of coding parameters that the video decoding device 31 refers to in order to decode each picture in the target sequence. For example, the picture parameter set PPS includes a reference value of the quantization width (pic_init_qp_minus26) used in decoding the picture and a flag (weighted_pred_flag) indicating the application of weighted prediction. Note that there may be multiple PPSs. In that case, one of the multiple PPSs is selected for each picture in the target sequence.
[0035] (Encoded Picture) A coded picture defines a set of data to be referenced by the video decoding device 31 in order to decode a picture PICT to be processed. As shown in FIG. 2, the picture PICT includes slices 0 to NS-1 (NS is the total number of slices included in the picture PICT).
[0036] In the following description, when there is no need to distinguish between slices 0 to NS-1, the subscripts of the symbols may be omitted. The same applies to other data that are included in the coded stream Te and that are to be described below and that are to be given subscripts.
[0037] (Coded Slice) A coded slice defines a set of data to be referenced by the video decoding device 31 in order to decode a target slice S. As shown in Fig. 2, a slice includes a slice header and slice data.
[0038] The slice header includes a group of coding parameters to be referred to by the video decoding device 31 in order to determine a decoding method for the current slice. Slice type designation information (slice_type) that designates the slice type is an example of a coding parameter included in the slice header.
[0039] Slice types that can be specified by the slice type specification information include (1) an I slice that uses only intra prediction when encoding, (2) a P slice that uses unidirectional prediction or intra prediction when encoding, and (3) a B slice that uses unidirectional prediction, bidirectional prediction, or intra prediction when encoding. Note that inter prediction is not limited to uni-prediction or bi-prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referring to P or B slice, it refers to a slice including a block that can use inter prediction.
[0040] In addition, the slice header may include a reference to a picture parameter set PPS (pic_parameter_set_id).
[0041] (Encoded slice data) The coded slice data specifies a set of data to be referenced by the video decoding device 31 in order to decode the slice data to be processed. As shown in Fig. 2(d), the slice data includes a CTU. A CTU is a block of a fixed size (e.g., 64x64) that constitutes a slice, and is also called a Largest Coding Unit (LCU).
[0042] (coding tree unit) 2 specifies a set of data that the video decoding device 31 refers to in order to decode a CTU to be processed. The CTU is divided into coding units CU, which are basic units of coding processing, by recursive quad tree division (QT (Quad Tree) division), binary tree division (BT (Binary Tree) division), or ternary tree division (TT (Ternary Tree) division). BT division and TT division are collectively called multi tree division (MT (Multi Tree) division). A node of a tree structure obtained by recursive quad tree division is called a coding node. Intermediate nodes of the quad tree, binary tree, and ternary tree are coding nodes, and the CTU itself is specified as the top coding node. The lowest coding node is specified as a coding unit.
[0043] (Encoding Unit) 2 specifies a set of data to be referenced by the video decoding device 31 in order to decode a coding unit to be processed. Specifically, a CU is composed of a CU header CUH, prediction parameters, transformation parameters, quantization transformation coefficients, etc. The CU header specifies a prediction mode, etc.
[0044] The prediction process may be performed on a CU basis, or on a sub-CU basis by further dividing a CU. If the size of a CU and a sub-CU are the same, there is one sub-CU in the CU. If the size of a CU is larger than that of a sub-CU, the CU is divided into sub-CUs. For example, if the CU is 8x8 and the sub-CU is 4x4, the CU is divided into 2 parts horizontally and 2 parts vertically, into 4 sub-CUs.
[0045] Prediction types (prediction modes) include intra prediction (MODE_INTRA), inter prediction (MODE_INTER), and intra block copy (MODE_IBC). Intra prediction is a prediction within the same picture, and inter prediction refers to a prediction process performed between different pictures (for example, between display times or between layer images).
[0046] The transform and quantization processes are performed in units of CUs, but the quantized transform coefficients may be entropy coded in units of sub-blocks such as 4x4.
[0047] (Prediction parameters) The predicted image is derived from prediction parameters associated with the block, which include intra-prediction and inter-prediction parameters.
[0048] (Inter prediction parameters) The prediction parameters of inter prediction will be described. The inter prediction parameters are composed of prediction list use flags predFlagL0 and predFlagL1, reference picture indexes refIdxL0 and refIdxL1, and motion vectors mvL0 and mvL1. predFlagL0 and predFlagL1 are flags indicating whether or not a reference picture list (L0 list, L1 list) is used, and when the value is 1, the corresponding reference picture list is used. Note that in this specification, when the "flag indicating whether or not XX" is written, a flag other than 0 (for example, 1) is XX, and 0 is not XX, and in logical negation, logical product, etc., 1 is treated as true and 0 is treated as false (similar below). However, in an actual device or method, other values can also be used as true and false values.
[0049] Syntax elements for deriving inter prediction parameters include, for example, a merge flag merge_flag (general_merge_flag), a merge index merge_idx, merge_subblock_flag, regulare_merge_flag, ciip_flag, merge_gpm_partition_idx, merge_gpm_idx0, merge_gpm_idx1, inter_pred_idc, a reference picture index refIdxLX, mvp_LX_idx, a difference vector mvdLX, and a motion vector precision mode amvr_mode. merge_subblock_flag is a flag indicating whether to use inter prediction in subblock units. regulare_merge_flag is a flag indicating whether to use a normal merge mode or MMVD. ciip_flag is a flag indicating whether to use a combined inter-picture merge and intra-picture prediction (CIIP) mode or a geometric partitioning merge mode (GPM) mode. merge_gpm_partition_idx is an index indicating the partition shape in the GPM mode. merge_gpm_idx0 and merge_gpm_idx1 are indexes indicating the merge index in the GPM mode. inter_pred_idc is an inter prediction identifier for selecting a reference picture to be used in the AMVP mode. mvp_LX_idx is a predicted vector index for deriving a motion vector.
[0050] (Reference picture list) The reference picture list is a list of reference pictures stored in the reference picture memory 306. In each CU, refIdxLX is used to specify which picture in the reference picture list RefPicListX (X=0 or 1) is actually referenced. Note that LX is a description method used when there is no distinction between L0 prediction and L1 prediction, and hereinafter, parameters for the L0 list and parameters for the L1 list are distinguished by replacing LX with L0 or L1.
[0051] (Merge prediction and AMVP prediction) There are two methods of decoding (encoding) prediction parameters: merge prediction mode (merge mode) and AMVP (Advanced Motion Vector Prediction, adaptive motion vector prediction) mode, and general_merge_flag is a flag for identifying these. The merge mode is a prediction mode in which some or all of the motion vector difference is omitted, and the prediction list usage flag predFlagLX, the reference picture index refIdxLX, and the motion vector mvLX are not included in the encoded data, but are derived from the prediction parameters of the neighboring blocks that have already been processed. The AMVP mode is a mode in which inter_pred_idc, refIdxLX, and mvLX are included in the encoded data. Note that mvLX is encoded as mvp_LX_idx, which identifies the prediction vector mvpLX, and the difference vector mvdLX. The general name for the prediction mode in which the motion vector difference is omitted or simplified is called the general merge mode, and the general merge mode and the AMVP prediction may be selected by the general_merge_flag.
[0052] When general_merge_flag is 1, merge_data() shown in FIG. 10 may be transmitted, in which regular_merge_flag may be transmitted. When regular_merge_flag is 1, normal merge mode or MMVD may be selected, and otherwise CIIP mode or GPM mode may be selected. In CIIP mode, a predicted image is generated by a weighted sum of an inter predicted image and an intra predicted image. In GPM mode, a predicted image is generated by dividing a target CU into two non-rectangular prediction units by a line segment.
[0053] inter_pred_idc is a value indicating the type and number of reference pictures, and takes one of the values PRED_L0, PRED_L1, and PRED_BI. PRED_L0 and PRED_L1 indicate uni-prediction using one reference picture managed in the L0 list and L1 list, respectively. PRED_BI indicates bi-prediction using two reference pictures managed in the L0 list and L1 list.
[0054] merge_idx is used to determine which prediction parameter of the candidate prediction parameters (merge candidates) derived from the blocks for which processing has been completed will be used as the prediction parameter of the target block. This is an index showing:
[0055] (Motion Vector) mvLX indicates the amount of shift between blocks on two different pictures. A prediction vector and a difference vector related to mvLX are called mvpLX and mvdLX, respectively.
[0056] (Inter prediction identifier inter_pred_idc and prediction list usage flag predFlagLX) The relationship between inter_pred_idc, predFlagL0, and predFlagL1 is as follows, and they are mutually convertible.
[0057] inter_pred_idc = (predFlagL1<<1)+predFlagL0 predFlagL0 = inter_pred_idc & 1 predFlagL1 = inter_pred_idc >> 1 In addition, the inter prediction parameter may use a prediction list usage flag or an inter prediction identifier. Furthermore, the determination using the prediction list usage flag may be replaced with a determination using the inter prediction identifier. Conversely, the determination using the inter prediction identifier may be replaced with a determination using the prediction list usage flag.
[0058] (Bi-predictive biPred decision) A flag biPred indicating whether or not the prediction is bi-predictive can be derived based on whether or not two prediction list usage flags are both 1.
[0059] Alternatively, biPred can be derived based on whether the inter-prediction identifier is a value indicating the use of two prediction lists (reference pictures).
[0060] (Configuration of a video decoding device) The configuration of a video decoding device 31 (FIG. 3) according to this embodiment will be described.
[0061] The video decoding device 31 includes an entropy decoding unit 301, a parameter decoding unit (prediction image decoding device) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a prediction image generating unit (prediction image generating device) 308, an inverse quantization and inverse transform unit 311, an adder 312, and a prediction parameter derivation unit 320. Note that, in accordance with the video encoding device 11 described below, the video decoding device 31 may also be configured not to include the loop filter 305.
[0062] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit), and the CU decoding unit 3022 further includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, PPS, and APS, and slice header (slice information) from the encoded data. The CT information decoding unit 3021 decodes the CT from the encoded data. The CU decoding unit 3022 decodes the CU from the encoded data. The TU decoding unit 3024 decodes the CU from the encoded data.
[0063] When a prediction error is included in the TU, the TU decoding unit 3024 decodes the QP update information and the quantized transform coefficient from the encoded data. The quantized transform coefficient may be derived in a plurality of modes (e.g., RRC mode and TSRC mode). RRC (Regular Residual Coding) is a decoding mode of a prediction error using a transform, and TSRC (Transform Skip Residual Coding) is a decoding mode of a prediction error in a transform skip mode in which a transform is not performed. The TU decoding unit 3024 decodes the last position of the transform coefficient in the RRC mode, and may not decode the last position in the TSRC mode. The QP update information is a difference value from a quantization parameter predicted value qPpred, which is a predicted value of a quantization parameter QP.
[0064] The predicted image generating unit 308 generates an inter-predicted image and an intra-predicted image. The device 300 includes a unit 310.
[0065] The prediction parameter derivation unit 320 includes the inter prediction parameter derivation unit 303 (FIG. 5) and an intra prediction parameter derivation unit.
[0066] In addition, although an example in which CTU and CU are used as processing units will be described below, the present invention is not limited to this example, and processing may be performed in sub-CU units. Alternatively, CTU and CU may be read as blocks, and sub-CU as sub-blocks, and processing may be performed in block or sub-block units.
[0067] The entropy decoding unit 301 performs entropy decoding on the externally input coded stream Te to decode individual codes (syntax elements). Entropy coding includes a method of variable-length coding the syntax elements using a context (probability model) adaptively selected according to the type of syntax element or surrounding circumstances, and a method of variable-length coding the syntax elements using a predetermined table or formula.
[0068] The entropy decoding unit 301 outputs the decoded code to the parameter decoding unit 302. The decoded code is, for example, a prediction mode predMode, general_merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX, amvr_mode, etc. Control of which code to decode is performed based on an instruction from the parameter decoding unit 302.
[0069] (Basic flow) FIG. 6 is a flowchart illustrating a schematic operation of the video decoding device 31.
[0070] (S1100: Decode Parameter Set Information) The header decoding unit 3020 decodes parameter set information such as the VPS, SPS, and PPS from the encoded data.
[0071] (S1200: Decode slice information) The header decoding unit 3020 decodes the slice header (slice information) from the encoded data.
[0072] Thereafter, the video decoding device 31 repeats the processes from S1300 to S5000 for each CTU included in the target picture to derive a decoded image of each CTU.
[0073] (S1300: Decode CTU Information) The CT information decoding unit 3021 decodes the CTU from the encoded data.
[0074] (S1400: Decode CT Information) The CT information decoding unit 3021 decodes the CT from the encoded data.
[0075] (S1500: CU Decoding) The CU decoding unit 3022 performs S1510 and S1520 to decode the CU from the encoded data.
[0076] (S1510: Decode CU Information) The CU decoding unit 3022 decodes the CU information, prediction information, TU division flag, CU residual flag, and the like from the encoded data.
[0077] (S1520: Decode TU information) When a prediction error is included in a TU, the TU decoding unit 3024 decodes the quantized prediction error and the like from the encoded data.
[0078] (S2000: Generation of predicted image) The predicted image generation unit 308 generates a predicted image for each block included in the current CU based on the prediction information.
[0079] (S3000: Inverse Quantization and Inverse Transformation) The inverse quantization and inverse transform unit 311 executes inverse quantization and inverse transform processing for each TU included in the target CU.
[0080] (S4000: Generate decoded image) The adder 312 adds the predicted image supplied from the predicted image generation unit 308 and the prediction error supplied from the inverse quantization and inverse transform unit 311 to generate a decoded image of the current CU.
[0081] (S5000: Loop Filter) The loop filter 305 applies a loop filter such as a deblocking filter, SAO, or ALF to the decoded image to generate a decoded image.
[0082] The loop filter 305 is a filter provided in the encoding loop, which removes block distortion and ringing distortion to improve image quality. The loop filter 305 applies a filter such as a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF) to the decoded image of the CU generated by the adder 312.
[0083] The reference picture memory 306 stores the decoded image of the CU generated by the adder 312 in a location that is determined in advance for each current picture and current CU.
[0084] The prediction parameter memory 307 stores prediction parameters at a predetermined position for each CTU or CU to be decoded. Specifically, the prediction parameter memory 307 stores the parameters decoded by the parameter decoding unit 302 and the prediction mode predMode separated by the entropy decoding unit 301.
[0085] The prediction image generating unit 308 receives a prediction mode predMode, prediction parameters, and the like. The prediction image generating unit 308 also reads a reference picture from the reference picture memory 306. The prediction image generating unit 308 generates a prediction image of a block or sub-block using the prediction parameters and the read reference picture (reference picture block) in the prediction mode indicated by the prediction mode predMode. Here, the reference picture block is a set of pixels on the reference picture (usually rectangular, so called a block), and is an area to be referenced for generating a prediction image.
[0086] (Configuration of inter-prediction parameter derivation unit) 5, the inter prediction parameter derivation unit 303 derives inter prediction parameters by referring to prediction parameters stored in the prediction parameter memory 307 based on the syntax elements input from the parameter decoding unit 302. The inter prediction parameter derivation unit 303 also outputs the inter prediction parameters to the inter prediction image generation unit 309 and the prediction parameter memory 307. The inter prediction parameter derivation unit 303 and its internal elements, the AMVP prediction parameter derivation unit 3032, the merge prediction parameter derivation unit 3036, the GPM prediction unit 30377, and the MV addition unit 3038, are means common to the video encoding device and the video decoding device, and therefore may be collectively referred to as a motion vector derivation unit (motion vector derivation device).
[0087] When general_merge_flag is 1, that is, when it indicates the merge prediction mode, it derives merge_idx and outputs it to the merge prediction parameter derivation unit 3036 .
[0088] When general_merge_flag is 0, that is, when it indicates the AMVP prediction mode, the AMVP prediction parameter derivation unit 3032 derives mvpLX from inter_pred_idc, refIdxLX, or mvp_LX_idx.
[0089] (MV addition section) The MV adder 3038 adds the derived mvpLX and mvdLX to derive mvLX.
[0090] (Merge prediction) The merge prediction parameter derivation unit 3036 includes a merge candidate derivation unit 30361 and a merge candidate selection unit 30362. Note that the merge candidate includes prediction parameters (predFlagLX, mvLX, refIdxLX). The merge candidates stored in the merge candidate list are assigned indices according to a predetermined rule.
[0091] The merging candidate derivation unit 30361 derives merging candidates by directly using the motion vectors and refIdxLX of the decoded adjacent blocks. Alternatively, the merging candidate derivation unit 30361 may apply a spatial merging candidate derivation process, a temporal merging candidate derivation process, or the like, which will be described later.
[0092] As a spatial merge candidate derivation process, the merge candidate derivation unit 30361 reads out prediction parameters stored in the prediction parameter memory 307 according to a predetermined rule, and sets them as merge candidates. For example, the prediction parameters for the positions A1, B1, B0, A0, and B2 shown in FIG. 7B are read out.
[0093] A1: (xCb-1, yCb+cbHeight-1) B1: (xCb+cbWidth-1, yCb-1) B0: (xCb+cbWidth, yCb-1) A0: (xCb-1, yCb+cbHeight) B2: (xCb-1, yCb-1) The upper left coordinates of the target block are (xCb, yCb), its width is cbWidth, and its height is cbHeight.
[0094] As a temporal merge derivation process, the merge candidate derivation unit 30361 may read prediction parameters of a block C in a reference image including the lower right CBR or center coordinates of the target block from the prediction parameter memory 307, set it as a merge candidate Col, and store it in the merge candidate list mergeCandList[ ].
[0095] The order of storing the mergeCandList[] is, for example, spatial merge candidates (B1, A1, B0, A0, B2), then temporal merge candidates Col. Note that reference blocks that are unavailable (blocks that are intra-predicted, etc.) are not stored in the merge candidate list. i = 0 if(availableFlagB1) mergeCandList[i++] = B1 if(availableFlagA1) mergeCandList[i++] = A1 if(availableFlagB0) mergeCandList[i++] = B0 if(availableFlagA0) mergeCandList[i++] = A0 if(availableFlagB2) mergeCandList[i++] = B2 if(availableFlagCol) mergeCandList[i++] = Col Furthermore, the history merge candidate HmvpCand, the pairwise average candidate avgCand, and the zero merge candidate zeroCandm may be added to mergeCandList[] for use.
[0096] The merging candidate selection unit 30362 selects a merging candidate N indicated by merge_idx from among the merging candidates included in the merging candidate list, using the following formula.
[0097] N = mergeCandList[merge_idx] Here, N is a label indicating a merging candidate, and can be A1, B1, B0, A0, B2, Col, etc. The motion information of the merging candidate indicated by the label N is indicated by (mvLXN[0], mvLXN[1]), predFlagLXN, and refIdxLXN.
[0098] The merge candidate selection unit 30362 selects (mvLXN[0], mvLXN[0]), predFlagLXN, and refIdxLXN as inter prediction parameters of the current block using merge_idx. The merge candidate selection unit 30362 stores the inter prediction parameters of the selected merge candidate in the prediction parameter memory 307, and also outputs the inter prediction parameters to the inter predicted image generation unit 309.
[0099] (MMVD Prediction Section 30376) The MMVD prediction unit 30376 performs processing in MMVD (Merge with Motion Vector Difference) mode. In merge prediction, a motion vector obtained from an adjacent block is used as the motion vector of a merge candidate. The MMVD mode is a mode in which a precise motion vector is obtained by adding a difference vector of a predetermined distance and a predetermined direction to the motion vector of the merge candidate. In the MMVD mode, the MMVD prediction unit 30376 uses the merge candidate and efficiently derives a motion vector by limiting the value range of the difference vector to a predetermined distance (e.g., 6 ways, 8 ways, etc.) and a predetermined direction (e.g., 4 directions, 8 directions, 16 directions, etc.). An example of four directions is shown in FIG. 8.
[0100] The MMVD prediction unit 30376 derives a motion vector mvLX[] using the merge candidates mergeCandList[] and the syntax elements mmvd_cand_flag, mmvd_direction_idx, and mmvd_distance_idx. These are syntax elements decoded from the encoded data or coded into the encoded data. Furthermore, the MMVD prediction unit 30376 may code or decode the syntax element distance_list_idx that selects a distance table and use it.
[0101] The MMVD prediction unit 30376 decodes the MMVD flag (mmvd_merge_flag) when regular_merge_flag is 1 (indicating that the normal merge mode or the MMVD mode is applied) for the target CU. Furthermore, when the MMVD flag indicates that the MMVD mode is applied (mmvd_merge_flag=1), the MMVD prediction unit 30376 applies the MMVD mode and encodes or decodes mmvd_cand_flag, mmvd_distance_idx, and mmvd_direction_idx.
[0102] The MMVD prediction unit 30376 generates an MMVD candidate list using one of the first two predicted vectors in the merge candidate list mergeCandList[] and a difference vector (MVD: motion vector difference) between the predicted vector and the MVD candidate list, and derives a motion vector. The difference vector is coded or decoded separately for direction and distance. Furthermore, the MMVD prediction unit 30376 derives a motion vector from the predicted vector and the difference vector.
[0103] 8 shows candidates for the difference vector refineMvLX derived in the MMVD prediction unit 30376. In the example shown in the figure, the black circle in the center is the position indicated by the predicted vector mvLXN (center vector).
[0104] Figure 7(a) shows the relationship between the index mmvd_cand_flag of mergeCandList[] and mvLXN. The motion vector of mergeCandList[mmvd_cand_flag] is set in mvLXN. The difference between the position indicated by this center vector (black circle in Figure 8) and the actual motion vector is the difference vector refineMvLX.
[0105] FIG. 7B is a diagram showing an example of blocks adjacent to the target block. For example, in the case of mergeCandList[]={A1, B1, B0, A0, B2}, when the decoded mmvd_cand_flag indicates 0, the MMVD prediction unit 30376 selects the motion vector of block A1 shown in FIG. 7B as the central vector mvLXN. When the decoded mmvd_cand_flag indicates 1, the MMVD prediction unit 30376 selects the motion vector of block B1 shown in FIG. 7B as the central vector mvLXN. When mmvd_cand_flag is not notified in the encoded data, it may be estimated that mmvd_cand_flag=0.
[0106] Furthermore, the MMVD prediction unit 30376 derives refineMvLX using an index mmvd_distance_idx indicating the length of the difference vector refineMvLX and an index mmvd_direction_idx indicating the direction of refineMvLX.
[0107] Fig. 7(c) is a diagram showing an example of mmvd_distance_idx. As shown in Fig. 7(c), in mmvd_distance_idx, values 0, 1, 2, 3, 4, 5, 6, and 7 correspond to eight distances (lengths), 1 / 4pel, 1 / 2pel, 1pel, 2pel, 4pel, 8pel, 16pel, and 32pel, respectively.
[0108] (d) of Fig. 7 is a diagram showing an example of mmvd_direction_idx. As shown in (d) of Fig. 7, in mmvd_direction_idx, the values 0, 1, 2, and 3 correspond to the positive direction of the x-axis, the negative direction of the x-axis, the positive direction of the y-axis, and the negative direction of the y-axis, respectively. The MMVD prediction unit 30376 derives a basic motion vector (mvdUnit[0], mvdUnit[1]) by referring to the direction table DirectionTable from mmvd_direction_idx. (mvdUnit[0], mvdUnit[1]) may be written as (sign[0], sign[1]).
[0109] Furthermore, the MMVD prediction unit 30376 derives the magnitude of the difference vector DistFromBaseMV (=MmvdDistance) from the distance DistanceTable[mmvd_distance_idx] indicated by mmvd_distance_idx in the distance table DistanceTable using the following formula.
[0110] DistFromBaseMV = DistanceTable[mmvd_distance_idx] Also, depending on the flag, the DistanceTable may be selected from the following two:
[0111] DistanceTable[] = {1, 2, 4, 8, 16, 32, 64, 128} DistanceTable[] = {4, 8, 16, 32, 64, 128, 256, 512} Also, taking into consideration the precision of the motion vector (for example, 1 / 16), the magnitudes of the center vector and the difference vector may be made the same by shifting left.
[0112] DistFromBaseMV = DistFromBaseMV << 2 (Other than the four directions) In the above, the case where the basic motion vector (mvdUnit[0], mvdUnit[1]) has four directions, up, down, left, and right, has been described, but it is not limited to four directions and may have eight directions. An example of the x component dir_table_x[] and the y component dir_table_y[] of the direction table DirectionTable when the basic motion vector has eight directions is shown below.
[0113] dir_table_x[] = { 8, -8, 0, 0, 6, -6, -6, 6} dir_table_y[] = { 0, 0, 8, -8, 6, -6, 6, -6} The size and order of the direction table may be other than those described above.
[0114] The MMVD prediction unit 30376 derives a basic motion vector (mvdUnit[0], mvdUnit[1]) by referencing the DirectionTable from mmvd_direction_idx.
[0115] mvdUnit[0] = dir_table_x[mmvd_direction_idx] mvdUnit[1] = dir_table_y[mmvd_direction_idx] Also, for example, by using a direction table such as the one below, the number of directions may be 4, 6, 12, or 16. - 6 directions (numDir=6) dir_table_x[] = { 8, -8, 2, -2, -2, 2} dir_table_y[] = { 0, 0, 4, -4, 4, -4} or dir_table_x[] = { 8, -8, 3, -3, -3, 3} dir_table_y[] = { 0, 0, 6, -6, 6, -6} - 12 directions (numDir=12) dir_table_x[] = { 8, -8, 0, 0, 4, 2, -4, -2, -2, -4, 2, 4} dir_table_y[] = { 0, 0, 8, -8, 2, 4, -2, -4, 4, 2, -4, -2} or dir_table_x[] = { 8, -8, 0, 0, 6, 3, -6, -3, -3, -6, 3, 6} dir_table_y[] = { 0, 0, 8, -8, 3, 6, -3, -6, 6, 3, -6, -3} - 16 directions (numDir=16) dir_table_x[] = {8, -8, 0, 0, 4, -4, -4, 4, 6, 2, -6, -2, -2, -6, 2, 6} dir_table_y[] = {0, 0, 8, -8, 4, -4, 4, -4, 2, 6, -2, -6, 6, 2, -6, -2} or dir_table_x[] = {1, -1, 0, 0, 1, -1, 1, -1, 2, -2, 2, -2, 1, 1, -1, -1} dir_table_y[] = {0, 0, 1, -1, 1, -1, -1, 1, 1, 1, -2, -1, 2, -2, 2, -2} for 4 directions (numDir=4) dir_table_x[] = { 1, -1, 0, 0} dir_table_y[] = { 0, 0, 1, -1} The size and order of the direction table may be other than those described above.
[0116] (multiple distance tables) Also, the number of distance tables is not limited to one, and may be multiple. For example, the MMVD prediction unit 30376 derives the length DistFromBaseMV of the difference vector using DistanceTable[] indicated by distance_list_idx decoded or derived from the encoded data. By referring to these, DistFromBaseMV may be derived from the first distance table DistanceTable1[] and the second distance table DistanceTable2[] as follows.
[0117] DistanceTable1 [] = {1, 2, 3, 5} DistanceTable2 [] = {4, 8, 16, 32} DistanceTable = DistanceTable1 (distance_list_idx == 0) DistanceTable = DistanceTable2 (distance_list_idx == 1) DistFromBaseMV = DistanceTable[mmvd_distance_idx] Furthermore, the MMVD prediction unit 30376 may switch between two distance tables using a two-dimensional table DistanceTable2d.
[0118] DistanceTable2d [] = {{1, 2, 3, 5},{4, 8, 16, 32}} DistFromBaseMV = DistanceTable2d[distance_list_idx][mmvd_distance_idx] (Derivation of the difference vector) The MMVD predictor 30376 derives a difference vector refineMvLX from the base motion vector and the magnitude of the difference vector DistFromBaseMV. When the merge candidate N for the central vector is uni-prediction from the L0 reference picture (predFlagL0N=1, predFlagL1N=0), the MMVD predictor 30376 derives an L0 difference vector refineMvL0 from the base motion vector and the magnitude of the difference vector DistFromBaseMV.
[0119] refineMvL0[0] = (DistFromBaseMV< <shiftMMVD) * mvdUnit[0] refineMvL0[1] = (DistFromBaseMV< <shiftMMVD) * mvdUnit[1] refineMvL1[0] = 0 refineMvL1[1] = 0 Here, shiftMMVD is a value that adjusts the magnitude of the difference vector so that it matches the accuracy MVPREC of the motion vector in the motion compensation unit 3091 (interpolation unit). For example, when MVPREC is 16, that is, the motion vector accuracy is 1 / 16 pixels, and there are four directions, that is, when mvdUnit[0] and mvdUnit[1] are 0 or 1, it is appropriate to use 2. In addition, the shift direction of shiftMMVD is not limited to a left shift. For example, when mvdUnit[0] and mvdUnit[1] use values other than 0 or 1 (for example, 8), such as 6, 8, 12, and 16 directions, the MMVD prediction unit 30376 may normalize by shifting to the right. For example, the MMVD prediction unit 30376 may perform a right shift after multiplying by a basic motion vector (mvdUnit[0], mvdUnit[1]) as follows:
[0120] refineMvL0[0] = (DistFromBaseMV * mvdUnit[0]) >> shiftMMVD refineMvL0[1] = (DistFromBaseMV * mvdUnit[1]) >> shiftMMVD Also, the MMVD prediction unit 30376 may calculate the magnitude and sign of the motion vector separately. The same applies to other methods of deriving the difference vector hereinafter.
[0121] refineMvL0[0] = ((DistFromBaseMV * abs(mvdUnit[0])) >> shiftMMVD) * sign(mvdUnit[0]) refineMvL0[1] = ((DistFromBaseMV * abs(mvdUnit[1])) >> shiftMMVD) * sign(mvdUnit[1]) Otherwise, when the merge candidate N for the central vector is uni-prediction from the L1 reference picture (predFlagL0N=0, predFlagL1N=1), the MMVD predictor 30376 derives the L1 difference vector refineMvL1 from the base motion vector and the magnitude of the difference vector DistFromBaseMV.
[0122] refineMvL0[0] = 0 refineMvL0[1] = 0 refineMvL1[0] = (DistFromBaseMV< <shiftMMVD) * mvdUnit[0] refineMvL1[1] = (DistFromBaseMV< <shiftMMVD) * mvdUnit[1] or refineMvL1[0] = (DistFromBaseMV * mvdUnit[0]) >> shiftMMVD refineMvL1[1] = (DistFromBaseMV * mvdUnit[1]) >> shiftMMVD Otherwise, when the merge candidate N for the central vector is bi-predictive (predFlagL0N=1, predFlagL1N=1), the MMVD predictor 30376 derives the first difference vector firstMv from the base motion vector and the magnitude of the difference vector DistFromBaseMV.
[0123] firstMv[0] = (DistFromBaseMV< <shiftMMVD) * mvdUnit[0] firstMv[1] = (DistFromBaseMV< <shiftMMVD) * mvdUnit[1] or firstMv = (DistFromBaseMV * mvdUnit[0]) >> shiftMMVD firstMv = (DistFromBaseMV * mvdUnit[1]) >> shiftMMVD Here, firstMv corresponds to a difference vector between the reference picture with the larger POC distance (POC difference) between the target picture and the reference picture. That is, among the reference pictures in the reference picture list L0 and the reference picture list L1, the reference picture with the larger POC distance (POC difference) between the target picture and the reference picture is set as the reference picture in the reference picture list LX. firstMv is a difference vector between the reference block of the reference picture in the list LX with the larger POC distance (POC difference) and the target block on the target picture.
[0124] Next, the MMVD prediction unit 30376 may derive a second motion vector secondMv of the other reference picture (reference list LY (Y=1-X)) by scaling firstMv. secondMv is a difference vector for the reference picture in list LY with a smaller POC distance.
[0125] For example, if the distance between the current picture currPic and the L0 picture RefPicList0[refIdxLN0] is equal to or greater than the distance between the current picture and the L1 picture RefPicList1[refIdxLN1], then firstMv corresponds to the L0 difference vector refineMvL0. Furthermore, the MMVD prediction unit 30376 may scale firstMv to derive the L1 difference vector refineMvL1.
[0126] refineMvL0[0] = firstMv[0] refineMvL0[1] = firstMv[1] refineMvL1[0] = Clip3(-32768, 32767, Sign(distScaleFactor * firstMv[0]) * ((Abs(distScaleFactor * refineMvL0[0]) + 127) >> 8)) refineMvL1[1] = Clip3(-32768, 32767, Sign(distScaleFactor * firstMv[1]) * ((Abs(distScaleFactor * refineMvL0[1]) + 127) >> 8)) Here, the MMVD prediction unit 30376 derives the scaling value distScaleFactor from the POC difference between currPic and the L0 reference picture and the POC difference between currPic and the L1 reference picture as follows: distScaleFactor = Clip3(-4096, 4095, ( tb * tx + 32 ) >> 6) (formula scale-1) tx = (16384 + (Abs(td) >> 1)) / td td = Clip3(-128, 127, DiffPicOrderCnt(currPic, RefPicList0[refIdxLN0])) tb = Clip3(-128, 127, DiffPicOrderCnt(currPic, RefPicList1[refIdxLN1])) DiffPicOrderCnt(currPic, RefPicList0[refIdxLN0]) is the POC difference between currPic and the L0 reference picture, and DiffPicOrderCnt(currPic, RefPicList1[refIdxLN1]) is the POC difference between currPic and the L1 reference picture.
[0127] Otherwise, if the distance between the current picture currPic and the L0 picture RefPicList0[refIdxLN0] is less than the distance between the current picture and the L1 picture RefPicList1[refIdxLN1], the first vector firstMv corresponds to the L1 difference vector refineMvL1. In this case, the MMVD prediction unit 30376 may scale the first vector firstMv to derive the L0 difference vector refineMvL0.
[0128] refineMvL0[0] = Clip3(-32768, 32767, Sign(distScaleFactor * firstMv[0]) * ((Abs(distScaleFactor * firstMv[0]) + 127) >> 8)) refineMvL0[1] = Clip3(-32768, 32767, Sign(distScaleFactor * firstMv[1]) * ((Abs(distScaleFactor * firstMv[1]) + 127) >> 8)) refineMvL1[0] = firstMv[0] refineMvL1[1] = firstMv[1] distScaleFactor is derived using the above formula (scale-1), but is calculated using td and tb. td = Clip3(-128, 127, DiffPicOrderCnt(currPic, RefPicList1[refIdxLN1])) tb = Clip3(-128, 127, DiffPicOrderCnt(currPic, RefPicList0[refIdxLN0])) In addition, the above branching between "If equal to or greater than the distance" and "Other than the above, if less than the distance" may be changed to "If greater than the distance" and "Other than the above, if less than the distance". In addition, if the distance between the target picture currPic and the L0 picture RefPicList0[refIdxLN0] is equal to the distance between the target picture and the L1 picture RefPicList1[refIdxLN1], the MMVD prediction unit 30376 may set refineMvLX[] from the following processing (processing A or processing B) without scaling firstMv[]. Action A: refineMvL0[0] = firstMv[0] refineMvL0[1] = firstMv[1] refineMvL1[0] = -firstMv[0] refineMvL1[1] = -firstMv[1] Process B: refineMvL0[0] = firstMv[0] refineMvL0[1] = firstMv[1] refineMvL1[0] = firstMv[0] refineMvL1[1] = firstMv[1] More specifically, the MMVD prediction unit 30376 derives refineMvLX[ ] by process A when the L0 reference picture, the current picture currPic, and the L1 current picture are arranged in time order, and derives refineMvLX[ ] by process B otherwise.
[0129] Note that the case where the signals are arranged in chronological order is when (POC_L0-POC_curr) * (POC_L1-POC_curr) < 0, that is, the following formula. DiffPicOrderCnt(RefPicList0[refIdxLN0], currPic) * DiffPicOrderCnt(currPic, RefPicList1[refIdxLN1]) > 0 Here, POC_L0, POC_L1, and POC_curr indicate the Picture Order Count (POC) of the L0 reference picture, the L1 reference picture, and the target picture, respectively.
[0130] The opposite case (reverse chronological order) is when (POC_L0-POC_curr) * (POC_L1-POC_curr)>0, that is, the following formula. DiffPicOrderCnt(RefPicList0[refIdxLN0], currPic) * DiffPicOrderCnt(currPic, RefPicList1[refIdxLN1]) < 0 In addition, even if the distance between POCs is different, the MMVD prediction unit 30376 may derive refineMvLX[] by processing A or processing B, and then scale refineMvLX[] according to the POC distance between the reference picture and the target picture to derive the final refineMvLX[].
[0131] (Addition of center vector and difference vector) Finally, the MMVD prediction unit 30376 derives a motion vector of the MMVD merging candidate from the difference vector refineMv[] and the central vector mvLXN[] (mvpLX[]) as follows. mvL0[0] = mvL0N[0] + refineMvL0[0] mvL0[1] = mvL0N[1] + refineMvL0[1] mvL1[0] = mvL1N[0] + refineMvL1[0] mvL1[1] = mvL1N[1] + refineMvL1[1] When deriving a temporary motion vector, the above mvLX[] is written as tempMvLX[].
[0132] (summary) In this way, even if the central vector is bi-predictive, the MMVD prediction unit 30376 notifies only one motion vector information (mmvd_direction_idx, mmvd_distance_idx). Then, two motion vectors are derived from this information. The MMVD prediction unit 30376 scales the motion vector as necessary based on the difference between the POC of each of the two reference pictures and the POC of the target picture. The difference vector between the reference block of the reference picture with the larger POC distance (POC difference) and the target block on the target picture is the first difference vector (firstMv).
[0133] firstMv[0] = (DistFromBaseMV< <shiftMMVD) * mvdUnit[0] firstMv[1] = (DistFromBaseMV< <shiftMMVD) * mvdUnit[1] The MMVD prediction unit 30376 derives the motion vector mvdLY(secondMv), LY(Y=1-X) of the reference picture with the smaller POC distance by scaling it by the ratio of the POC distances between the pictures (POCS / POCL).
[0134] secondMv[0] = (DistFromBaseMV< <shiftMMVD) * mvdUnit[0] *POCS / POCL secondMv[1] = (DistFromBaseMV< <shiftMMVD) * mvdUnit[1] *POCS / POCL The reference picture with the smaller POC distance corresponds to the picture with the smaller POC distance (POC difference) between the target picture and the reference picture. Here, POCS is the difference value of the POC difference between the target picture and the reference picture closer to the target picture, and POCL is the difference value of the POC difference between the target picture and the reference picture farther from the target picture.
[0135] As described above, the MMVD prediction unit 30376 derives mvpLX[] (mvLXN[]) and refineMvLX[], and uses these to derive the motion vector mvLX[] of the current block.
[0136] mvLX[0] = mvpLX[0]+ refineMvLX[0] mvLX[1] = mvpLX[1]+ refineMvLX[1] (Integer rounding of motion vectors) When the magnitude DistFromBaseMV of the difference vector to be added to the central vector is greater than a predetermined threshold, the MMVD prediction unit 30376 may correct the motion vector mvLX of the current block so as to indicate an integer pixel position. For example, the MMVD prediction unit 30376 may round off the motion vector mvLX to an integer when DistFromBaseMV is equal to or greater than a predetermined threshold of 16.
[0137] Furthermore, the MMVD prediction unit 30376 may convert mvLX to an integer when distance_list_idx is a specific distance table (e.g., DistanceTable2) and mmvd_distance_idx is in a specific range (e.g., mmvd_distance_idx is 2 or 3). distance_list_idx is an index that selects a distance table, and mmvd_distance_idx is an index that selects an element of the distance table (selects a distance coefficient). For example, when distance_list_idx==1 and mmvd_distance_idx>=2, the MMVD prediction unit 30376 may modify mvLX using the following formula.
[0138] mvLX[0] = (mvLX[0] / MVPREC) * MVPREC mvLX[1] = (mvLX[1] / MVPREC) * MVPREC The MMVD prediction unit 30376 may also derive mvLX using a shift.
[0139] mvLX[0] = (mvLX[0] >> MVBIT) << MVBIT mvLX[1] = (mvLX[1] >> MVBIT) << MVBIT Here, MVBIT = log2(MVPREC). For example, 4. You can also derive it below, taking into account positive and negative values.
[0140] mvLX[0] = mvLX[0]>=0 ? (mvLX[0] >> MVBIT) << MVBIT : -((-mvLX[0] >> MVBIT) << MVBIT) mvLX[1] = mvLX[1]>=0 ? (mvLX[1] >> MVBIT) << MVBIT : -((-mvLX[1] >> MVBIT) << MVBIT) In this way, by rounding the motion vector to an integer, it is possible to reduce the amount of calculation required for generating a predicted image.
[0141] (Syntax) FIG. 9 shows an example of the syntax of merge_data() notified when merge prediction is on (general_merge_flag==1) in the current block. general_merge_flag is a flag notified when the current block is not in skip mode, indicating whether or not prediction parameters of the current block are derived from adjacent blocks. In other words, it is a flag indicating whether or not merge mode is used. In the case of skip mode, the inter prediction parameter derivation unit 303 sets general_merge_flag=1. merge_data() is a syntax structure for notifying parameters of merge prediction.
[0142] The merge_subblock_flag is a flag indicating whether or not parameters of subblock-based inter prediction of the current block are derived from adjacent blocks. When the merge_subblock_flag is 1, it indicates that subblock-based inter prediction is used. When the merge_subblock_flag is 0, it indicates that subblock-based inter prediction is not used.
[0143] The regular_merge_flag is a flag indicating whether the normal merge mode or the merge mode using differential motion vectors (MMVD) is used for the current block. When the regular_merge_flag is 1, it indicates that the normal merge mode or MMVD is used. When the regular_merge_flag is 0, it indicates that the normal merge mode and MMVD are not used.
[0144] The mmvd_merge_flag is a flag indicating whether or not the MMVD is used in the current block. When the mmvd_merge_flag is 1, it indicates that the MMVD is used in the current block. At this time, the parameter decoding unit 302 uses the mmvd_cand_flag, mmvd_distance_idx, and mmvd_dir The MMVD candidate list is a list that stores MMVD candidates. ...
[0145] mmvd_cand_idx = mmvd_distance_idx * numDir + mmvd_direction_idx Here, mmvd_distance_idx = 0..numDist-1, mmvd_direction_idx = 0..numDir-1. numDir represents the number of predetermined directions of difference vectors in MMVD (for example, if there are 16 directions, numDir = 16). numDist represents the number of predetermined distances of difference vectors in MMVD (for example, if there are 6 distances, numDist = 6). mmvd_distance_idx and mmvd_direction_idx may be abbreviated to dist and dir below.
[0146] Fig. 10 is another example of a syntax table of the MMVD prediction unit 30376. When it is indicated that MMVD is used in the target block (mmvd_merge_flag==1), the parameter decoding unit 302 decodes mmvd_cand_flag and mmvd_cand_idx. The value of mmvd_cand_idx may be 0 to maxNumMmvdLUT-1 (for example, 11). Fig. 10 shows an example in which mmvd_cand_idx is signaled instead of mmvd_distance_idx and mmvd_direction_idx in Fig. 9. The parameter decoding unit 302 may encode and decode mmvd_cand_idx using TR (Truncated Rice) binary with cMax=11 and a Rice parameter of 1, or TB (Truncated Binary). cMax is the upper limit of the value that the syntax element can take.
[0147] (Another configuration of the MMVD prediction unit 30376) FIG. 11 is a flowchart showing the process of another configuration of the MMVD prediction unit.
[0148] Another configuration of the MMVD prediction unit 30376 will be described. In this embodiment, when MMVD prediction is used (mmvd_merge_flag is 1), the MMVD prediction unit 30376 derives an MMVD candidate list mmvdLUT (S301) and derives a difference vector (MVD) using an index mmvd_cand_idx indicating an MMVD candidate (S302). The MMVD prediction unit 30376 selects an MMVD candidate using the mmvdLUT and mmvd_cand_idx, and derives a motion vector of a target block. For example, the MMVD prediction unit 30376 may select an MMVD candidate by mmvdLUT[mmvd_cand_idx].
[0149] The MMVD candidates may be the motion vectors of L0 and L1 (MvL0[0], MvL0[1], MvL0[0], MvL0[1]). In this case, the mmvdLUT becomes as follows:
[0150] mmvdLUT[][] = { (MvL0[0], MvL0[1], MvL0[0], MvL0[1]), (MvL0[0], MvL0[1], MvL0[0], MvL0[1]), …} The MMVD prediction unit 30376 uses the mmvdLUT and mmvd_cand_idx (=tempIdx) to select MMVD candidates and derive motion vectors. mvL0[0] (= tempMvL0[0]) = mvL0[0] of mmvdLUT[mmvd_cand_idx] (Formula MMVD-1) mvL0[1] (= tempMvL0[1]) = mvL0[1] of mmvdLUT[mmvd_cand_idx] mvL1[0] (= tempMvL1[0]) = mvL1[0] of mmvdLUT[mmvd_cand_idx] mvL1[1] (= tempMvL1[1]) = mvL1[1] of mmvdLUT[mmvd_cand_idx] The above can also be expressed as follows:
[0151] mvL0[0] (= tempMvL0[0]) = mmvdLUT[mmvd_cand_idx][0] mvL0[1] (= tempMvL0[1]) = mmvdLUT[mmvd_cand_idx][1] mvL1[0] (= tempMvL1[0]) = mmvdLUT[mmvd_cand_idx][2] mvL1[1] (= tempMvL1[1]) = mmvdLUT[mmvd_cand_idx][3] Also, the MMVD candidate may be the difference vector (firstMv[0], firstMv[1]).
[0152] mmvdLUT[][] = { (fistMv[0], fisetMv[1]), (fistMv[0], fisetMv[1]), (fistMv[0], fisetMv[1]), …} The MMVD prediction unit 30376 selects an MMVD candidate using the mmvdLUT and mmvd_cand_idx (=tempIdx), derives firstMv, and derives a motion vector mvLX (=tempMvLX) from firstMv.
[0153] firstMv[0] = firstMv[0] of mmvdLUT[mmvd_cand_idx] firstMv[1] = firstMv[1] of mmvdLUT[mmvd_cand_idx] The above can also be expressed as follows:
[0154] firstMv[0] = mmvdLUT[mmvd_cand_idx][0] firstMv[1] = mmvdLUT[mmvd_cand_idx][1] The method of deriving the motion vector tempMvLX from firstMv has already been explained.
[0155] Alternatively, the MMVD candidate may be the position of the difference vector (the direction and distance from the central vector).
[0156] mmvdLUT[][] = { (mmvd_direction_idx, mmvd_distance_idx), (mmvd_direction_idx, mmvd_distance_idx), (mmvd_direction_idx, mmvd_distance_idx), …} The MMVD prediction unit 30376 may select an MMVD candidate (mmvd_direction_idx, mmvd_distance_idx) using the mmvdLUT and mmvd_cand_idx (=tempIdx).
[0157] mmvd_direction_idx = mmvd_direction_idx in mmvdLUT[mmvd_cand_idx] mmvd_distance_idx = mmvd_distance_idx in mmvdLUT[mmvd_cand_idx] The above can also be expressed as follows:
[0158] mmvd_direction_idx = mmvdLUT[mmvd_cand_idx][0] mmvd_distance_idx = mmvdLUT[mmvd_cand_idx][1] mmvd_cand_idx = mmvd_distance_idx * numDir + mmvd_direction_idx The motion vector mvLX (=tempMvLX) may be derived from mmvd_cand_idx by (equation MMVD-1).
[0159] Also, the MMVD candidate may simply be the number of the MMVD candidate.
[0160] mmvdLUT[][] = { 0, 1, 2, 3, …} In this case, the motion vector may be derived from the optimal number mmvdLUT[mmvd_cand_idx] or mmvdLUT[tempIdx] using the following firstMv[].
[0161] dist = tempIdx / numDir; dir = tempIdx - dist * numDir; firstMv[0] = (dist< <shiftMMVD) * dir_table_x[dir] firstMv[1] = (dist< <shiftMMVD) * dir_table_y[dir] In this embodiment, by using an mmvd candidate list arranged in ascending order of cost, the amount of coding of information required to derive a difference vector (mmvd_distance_idx, mmvd_direction_idx, or mmvd_cand_idx) is reduced, thereby achieving the effect of improving coding efficiency.
[0162] (MMVD candidate list derivation process) The MMVD prediction unit 30376 calculates a cost tempCost (template matching cost) for each MMVD candidate, and sorts the MMVD candidates in ascending order of cost to derive the final MMVD candidate list mmvdLUT. All elements of the mmvdLUT may be derived in advance first, and the mmvdLUT1 may be derived by replacing the candidates in the mmvdLUT using the cost. The MMVD prediction unit 30376 may also derive the MMVD candidate list mmvdLUT by adding MMVD candidates to the list in ascending order of cost, with the MMVD candidate list mmvdLUT being set as an initial state in an empty state. When adding an MMVD candidate to the list, the MMVD prediction unit 30376 may skip adding the MMVD candidate if the addition position exceeds the maximum number of candidates in the list maxNumMmvdLUT (for example, 12).
[0163] Here, the flow chart of FIG. 11 will be described in detail.
[0164] First, the MMVD prediction unit 30376 initializes the MMVD candidate list mmvdLUT (S3011). All candidate motion vectors may be derived as elements of the mmvdLUT. For example, they are derived for tempIdx=0..numDir*numDist-1.
[0165] dist = tempIdx / numDir; dir = tempIdx - dist * numDir; Using the above formula, an mmvdLUT with {dir, dist} as elements may be used. An mmvdLUT with {firstMv[0], firstMv[1]} derived from {dir, dist} as elements may be used. Furthermore, a table with {mvLX[0], mvLX[1]} derived from firstMv as elements may be used. Here, the MMVD prediction unit 30376 may obtain the direction (dir) and distance (dist) from the integer (index) tempIdx indicating each mmvdMergeCand.
[0166] Derive firstMv using dist and dir.
[0167] firstMv[0] = (dist< <shiftMMVD) * dir_table_x[dir] firstMv[1] = (dist< <shiftMMVD) * dir_table_y[dir] Then, in a similar manner to (deriving the difference vector), refineMvL0 and refineMvL1 are derived from firstMv, and tempMvLX[] (X = 0, 1) is derived.
[0168] tempMvLX[0] = mvpLX[0] + refineMvLX[0] tempMvLX[1] = mvpLX[1] + refineMvLX[1] The derived tempMvLX is stored as element mmvdLUT[tempIdx] of mmvdLUT.
[0169] Next, the MMVD prediction unit 30376 loops the number of times required for the number of candidates (for example, NumMmvdCand = numDir * numDist) and derives the cost tempCost of the MMVD candidates (S3012).
[0170] The MMVD prediction unit 30376 calculates a direction (dir) and a distance (dist) for each MMVD candidate (mmvdMergeCand). The MMVD prediction unit 30376 then calculates a motion vector tempMvLX for the MMVD candidate (S3013).
[0171] Next, the MMVD prediction unit 30376 performs template matching using tempMvLX and calculates the cost (S3014). The left and top pixels adjacent to the target block are set as a template (rec template) and a pixel of a reference block indicated by tempMvLx (ref template). The MMVD prediction unit 30376 may perform template matching to derive a ref template (and a corresponding difference vector tempMvLx) that minimizes the difference (cost, tempCost) between the rec template and the ref template. In addition, the sum of absolute difference (SAD) or the sum of squared difference (SSD) may be used as tempCost.
[0172] For example, tempCost may be the sum of the absolute difference between the left region of the reference block and the left region of the target block, and the sum of the absolute difference between the upper region of the reference block refSamplesLX and the upper region of the target block.
[0173] tempCost = Σabs(refSamplesLX[xC+iL+tempMvLX[0]][yC+jL+tempMvLX[1]] - recSamples[xC+i][xC+j])+ Σabs(refSamplesLX[xC+iT+tempMvLX[0]][yC+jT+tempMvLX[1]] - recSamples[xC+i][xC+jT]) where X=0 or 1.
[0174] Here, refSamples is the reference picture, recSamples is the current picture, tempMvLX[2] is the motion vector of the MMVD candidate, xC, yC are the top left coordinates of the current block, nW, nH are the width and height of the block, and Σ is the sum over iL=-1, jL=0..nH-1, iT=0..nW-1, and jT=-1. The above processing may be performed after converting the decimal precision to integer precision.
[0175] tempMvLX[0] = (abs(tempMvLX[0] >> shiftMMVD) * sign(tempMvLX[0]) tempMvLX[1] = (abs(tempMvLX[1] >> shiftMMVD) * sign(tempMvLX[1]) The above is for 1 / 16 accuracy, shiftMMVD=4. To derive firstMv with integer precision, firstMv may be derived in advance as follows.
[0176] firstMv[0] = dist * dir_table_x[dir] firstMv[1] = dist * dir_table_y[dir] Then, the MMVD prediction unit 30376 updates the MMVD candidate list according to the cost calculated in S3014 (S3015). The MMVD prediction unit 30376 may sort the MMVD candidate list in ascending order of cost. Note that, as shown below, a position insertPos at which the mmvdLUT is inserted may be derived, and the table below insertPos may be updated. insertPos = 0 while (insertPos < maxNumMmvdLUT && tempCost < candCostList[endIdx-1-insertPos]){ insertPos++; } if (insertPos != 0) { for (i = 1; i < insertPos; i++) { mmvdLUT[endIdx - i] = mmvdLUT[endIdx - 1 - i] candCostList[endIdx - i] = candCostList[endIdx - 1 - i] } mmvdLUT[endIdx - insertPos] = tempIdx candCostList[endIdx - insertPos] = tempCost } Here, while(x){process} is a loop process that repeats process while x is true. mmvdLUT is a list that stores tempIdx. endIdx is the last index of the group, for example, endIdx=(tempIdx / grpSize) * grpSize + grpSize, where grpSize=1, 2, 4, or 8. candCostList[idx] is a list of tempCosts of MMVD candidates indicated by idx.
[0177] The MMVD prediction unit 30376 performs the processes of S3013 to S3015 the number of times corresponding to the number of MMVD candidates. Through the above processes, the MMVD prediction unit 30376 derives an MMVD candidate list.
[0178] According to this embodiment, the MMVD prediction unit 30376 performs template matching for all MMVD candidates and calculates the cost. For example, if there are 16 MMVD directions and 6 distances, template matching (cost derivation and sorting processing) is performed 96 times.
[0179] (Another example of the MMVD candidate list derivation process) In this embodiment, MMVD candidates for cost calculation are adaptively selected, which has the effect of reducing the amount of calculation without reducing the coding efficiency.
[0180] FIG. 12 is a flow chart showing the process of determining the direction and distance of the difference vector in this embodiment. 11. Note that the process in which the MMVD prediction unit 30376 selects MMVD candidates using the mmvdLUT and mmvd_cand_idx (S302) is the same as that described in FIG.
[0181] First, the MMVD prediction unit 30376 initializes an MMVD candidate list (S4011). An MMVD candidate indicates the position of a difference vector (direction and distance from the central vector). The MMVD candidate list is a list that stores MMVD candidates. Next, the MMVD prediction unit 30376 loops the cost derivation process for candidates that can become difference vectors as many times as the number of candidates (numDir * numDist) (S4012).
[0182] The MMVD prediction unit 30376 derives the direction (dir) and distance (dist) corresponding to each MMVD candidate (mmvdMergeCand) (S4013).
[0183] The MMVD prediction unit 30376 determines the search direction (dir) of mmvdMergeCand at the current distance (dist) (S4014). For each mmvdMergeCand, the MMVD prediction unit 30376 determines whether the search direction is the derived one (S4015), and if it is the search direction (Y in S4015), derives the cost. Hereinafter, searching XX means deriving tempCost with XX as a search candidate.
[0184] If it is determined that it is not the search direction (N in S4015), the MMVD prediction unit 30376 does not derive a cost for mmvdMergeCand, and proceeds to deriving the cost of the next mmvdMergeCand.
[0185] If it is the search direction (Y in S4015), the MMVD prediction unit 30376 obtains the current temporary difference vector (firstMv) from the direction and distance of the search candidate (S4016).
[0186] Next, the MMVD prediction unit 30376 performs template matching using tempMvLX derived in the same manner as in the MMVD candidate list derivation process, and calculates the cost tempCost (S4017). Then, the MMVD prediction unit 30376 updates the MMVD candidate list in ascending order of cost according to the cost calculated in S4017 (S4018).
[0187] The MMVD prediction unit 30376 performs the processes of S4013 to S4018 the number of times corresponding to the number of MMVD candidates. Through the above processes, the MMVD prediction unit 30376 derives an MMVD candidate list.
[0188] 13 is another flowchart showing the flow of processing for determining the direction and distance of a difference vector in this embodiment. Note that the processing in which the MMVD prediction unit 30376 selects an MMVD candidate (S302) using the mmvdLUT and mmvd_cand_idx is the same as that in FIG.
[0189] First, the MMVD prediction unit 30376 initializes an MMVD candidate list (S5011). Next, the MMVD prediction unit 30376 loops through distance candidates that can be difference vectors the number of times (numDist) (S5012). Here, the MMVD prediction unit 30376 determines the direction (dir) of the difference vector to search for at the current distance (dist) (S5013). A detailed determination method will be described later. Next, the MMVD prediction unit 30376 loops through candidates in the direction determined to be searched in S5013 the number of times (S5014).
[0190] The MMVD prediction unit 30376 derives the motion vector of the MMVD candidate from the direction and distance of the difference vector (S5016). Next, the MMVD prediction unit 30376 performs template matching using the motion vector of the MMVD candidate and calculates the cost (S5017). Then, the MMVD prediction unit 30376 updates the MMVD candidate list in ascending order of cost according to the cost calculated in S5017 (S4018).
[0191] The MMVD prediction unit 30376 performs the processes of S5016 to S5018 for the number of candidates in the direction to search for the MMVD candidates. The MMVD prediction unit 30376 also performs the processes of S5013 to S5018 for the number of candidates in the distance to the MMVD candidates. Through the above processes, the MMVD prediction unit 30376 derives an MMVD candidate list.
[0192] (Learn more about how the search direction is determined) A method of determining the search direction according to the distance (dist) of mmvdMergeCand, which corresponds to processes S4014 and S5013, will be described.
[0193] <Search direction determination method 1> The MMVD prediction unit 30376 may determine the direction to search based on the following: The following describes the case where numDir=16 (16 directions). 1: If the distance is 0 (dist==0), only directions included in the direction group DIR1 are searched. 2: If the distance is any other than this (dist!=0), only directions included in the direction group DIR2 are searched.
[0194] Here, the direction groups DIR1 and DIR2 are each a subset (part) of the directions of the MMVD candidates. DIR1 may include 1 / 2 of all directions. For example, it may be a direction dir where dir=0...numDir / 2-1. For example, DIR1 may include 8 directions, including horizontal, vertical, and 45-degree diagonal directions. The directions may be indicated by integers representing the directions as shown below, or may be indicated by labels.
[0195] DIR1 = {0, 1, 2, 3, 4, 5, 6, 7} DIR1 = {Top, Top Right, Right, Bottom Right, Bottom, Bottom Left, Left, Top Left} The direction group DIR2 may include a total of three directions, the direction that gives the smallest cost among the template matching costs for the number of searches found in the previous distance dist and its adjacent directions. Alternatively, the direction group DIR2 may include two directions that give the smallest cost among the template matching costs for the number of searches found in the previous distance dist and their adjacent directions, up to six directions. The directions included in the direction group DIR2 may be updated for each dist.
[0196] Hereinafter, an example of limiting the search direction according to the distances shown in Figs. 12 and 13 will be described with reference to Figs. 14 and 15. It is assumed that there are a total of 96 candidates for difference vectors in MMVD, with 6 distances and 16 directions. The relationship between the directions and indexes is as shown in Figs. 14(a) and (b). Fig. 15 shows search candidates with a distance dist of up to 3 in this case. The following shows how the MMVD prediction unit 30376 searches for MMVD candidates. In the example below, the MMVD prediction unit 30376 searches for 23 (8+3*5) candidates.
[0197] Hereinafter, the operation of "adding the searched candidates to the MMVD candidate list and updating the MMVD candidate list in ascending order of cost" will be described as "updating the mmvdLUT for the searched candidates." <step1>Behavior when :dist=0 Search 8 directions, horizontal, vertical, and 45 degree diagonal directions, using the direction group DIR1 = {0, 1, 2, 3, 4, 5, 6, 7}. Update the mmvdLUT for the searched candidates. As a result of the search, dir=5 was found to have the smallest cost. <step2>Behavior when :dist=1 In step 1, the three directions dir=5 and adjacent dir=11, 15 are selected, and a direction group DIR2 = {5, 11, 15} for step 2 is derived. Three directions are searched using DIR2 = {5, 11, 15}. The mmvdLUT is updated for the searched candidates. As a result of the search, dir=5 was found to have the smallest cost. <step3>Behavior when :dist=2 From the result of step 2, derive DIR2 = {5, 11, 15} for step 3. Search three directions using DIR2 = {5, 11, 15}. Update the mmvdLUT for the searched candidates. As a result of the search, dir=15 was found to have the smallest cost. <step4>Behavior when :dist=3 From the result of step 3, derive DIR2 = {3, 5, 15} for step 4. Search three directions using DIR2 = {3, 5, 15}. Update the mmvdLUT for the searched candidates. As a result of the search, dir=3 was found to have the smallest cost. <step5>Behavior when :dist=4 From the result of step 4, DIR2 = {3, 13, 15} is derived for step 5. Three directions are searched using DIR2 = {3, 13, 15}. The mmvdLUT is updated for the searched candidates. As a result of the search, dir=3 was found to have the smallest cost. <step6>Behavior when :dist=5 From the result of step 5, derive DIR2 = {3, 13, 15} for step 6. Search three directions using DIR2 = {3, 13, 15}. Update the mmvdLUT for the searched candidates.
[0198] From the above steps, the MMVD prediction unit 30376 derives an MMVD candidate list.
[0199] An example in which the direction group DIR2 configures search candidates from a total of six directions, including the two directions that provide the smallest cost and their adjacent directions, is shown below. <step1>Behavior when :dist=0 Using the direction group DIR1 = {0, 1, 2, 3, 4, 5, 6, 7}, search for 8 directions: horizontal, vertical, and 45 degree diagonal. Update the mmvdLUT for the searched candidates. As a result of the search, dir=5 has the smallest cost, and dir=0 has the second smallest cost. <step2>Behavior when :dist=1 From the result of step 1, search six directions using the direction group DIR2 = {5, 11, 15, 0, 8, 10}. Update the mmvdLUT for the searched candidates. As a result of the search, dir=10 has the smallest cost, and dir=11 has the second smallest cost. <step3>Behavior when :dist=2 From the result of step 2, search six directions using the direction group DIR2 = {0, 6, 10, 1, 5, 11}. Update the mmvdLUT for the searched candidates. As a result of the search, dir=10 was the smallest cost, and dir=0 was the second smallest cost. <step4>Behavior when :dist=3 From the result of step 3, a four-direction search is performed using the direction group DIR2 = {0, 6, 8, 10}. If the directions that give the smallest cost are adjacent to each other, the six directions may not be the same as in this example. The mmvdLUT is updated for the searched candidates. The subsequent processing is omitted, but the MMVD prediction unit 30376 <step6>The same process is carried out up to
[0200] From the above steps, the MMVD prediction unit 30376 derives an MMVD candidate list.
[0201] In this embodiment, selection is made for each distance of mmvdMergeCand for which cost calculation is performed in the MMVD candidate search. In particular, the search is characterized in that only the first direction subset (DIR1) is searched for a predetermined distance, and only the second direction subset (DIR2) is searched for distances other than the predetermined distance. The search is characterized in that directions included in the second direction subset (DIR2) are searched for each distance candidate, and the directions included in the second direction subset (DIR2) are updated for each distance candidate.
[0202] This has the effect of reducing the amount of calculation without reducing the coding efficiency.
[0203] <Search direction determination method 2> The MMVD predictor 30376 has multiple distance sets. The MMVD predictor 30376 may then determine the direction to search based on the following: 1: For a distance included in the distance set DIST1, only directions included in the direction group DIR1 are searched. 2: For distances included in other distance sets, only directions included in direction group DIR2 are searched.
[0204] Here, the direction group DIR1 may be as in <Search direction determination method 1>.
[0205] The direction group DIR2 may include a total of three directions, the direction that gives the smallest cost among the template matching costs for the number of searches found in the previous distance set, and its adjacent directions. Alternatively, the direction group DIR2 may include two directions that give the smallest cost among the template matching costs for the number of searches found in the previous distance set, and up to four directions that are adjacent to the direction with the smallest cost. Alternatively, the direction group DIR2 may include two directions that give the smallest cost among the template matching costs for the number of searches found in the previous distance set, and up to six directions that are adjacent to each of the two directions.
[0206] The number of directions included in the direction group DIR2 may be updated for each distance set. Also, the method of selecting the directions included in the direction group DIR2 may differ for each distance set. For example, the number of directions to be searched in the distance set DIST2 may be six, and the number of directions to be searched in the distance set DIST3 may be four or three.
[0207] The distance set DIST1 may include only dist = 0. In this case, the distance set DIST2 may include dist = 1, 2, and the distance set DIST3 may include dist = 3, 4, ....
[0208] Alternatively, the distance set DIST1 may include dist=0, 1. In this case, the distance set DIST2 may include dist=2, 3, and the distance set DIST3 may include dist=4, 5, . . .
[0209] The following shows how the MMVD prediction unit 30376 searches for MMVD candidates. <step1>: Behavior when distance set DIST1={0} Using the direction group DIR1 = {0, 1, 2, 3, 4, 5, 6, 7}, search for 8 directions: horizontal, vertical, and 45 degree diagonal. Update the mmvdLUT for the searched candidates. Also, as a result of the search, dir=5 was found to have the smallest cost. <step2>: Behavior when distance set DIST2={1,2} From the result of step 1, the direction group DIR2 = {5, 11, 15} is used to search for the three directions dir=5, which has the smallest cost, and the adjacent dir=11, 15. The mmvdLUT is updated for the searched candidates. Furthermore, as a result of the search, dir=5 was found to have the smallest cost. <step3>: Behavior when distance set DIST3={3,4,5} From the result of step 2, search three directions using the direction group DIR2 = {5, 11, 15}. Update the mmvdLUT for the searched candidates.
[0210] From the above steps, the MMVD prediction unit 30376 derives an MMVD candidate list.
[0211] In addition, an example will be shown in which the method of selecting directions included in the direction group DIR2 differs for each distance set. The following is an example in which the number of directions to be searched in the distance set DIST2 is six, and the number of directions to be searched in the distance set DIST3 is three. <step1>: Behavior when distance set DIST1={0} Using the direction group DIR1 = {0, 1, 2, 3, 4, 5, 6, 7}, search for 8 directions: horizontal, vertical, and 45 degree diagonal. Update the mmvdLUT for the searched candidates. Also, as a result of the search, dir=5 was the smallest cost, and dir=0 was the second smallest cost. <step2>: Behavior when distance set DIST2={1,2} From the result of step 1, 6 directions are searched using the direction group DIR2 = {5, 11, 15, 0, 8, 10}. The mmvdLUT is updated for the searched candidates. In addition, as a result of the search, dir=10 was found to have the smallest cost. <step3>: Behavior when distance set DIST3={3,4,5} From the result of step 2, search for three directions using the direction group DIR2 = {0, 6, 10}. Update the mmvdLUT for the searched candidates.
[0212] From the above steps, the MMVD prediction unit 30376 derives an MMVD candidate list.
[0213] In this embodiment, mmvdMergeCand, which performs cost calculation in searching for MMVD candidates, is selected for each distance. In particular, only the first direction subset (DIR1) is searched for distances included in a predetermined distance set, and only the second direction subset (DIR2) is searched for distances included in a distance set other than the predetermined distance set. Also, for a distance set including multiple distance candidates, directions included in the second direction subset (DIR2) are searched, and the directions included in the second direction subset (DIR2) are updated for each distance set including multiple distance candidates. Also, in the search, the number of directions included in the second direction subset (DIR2) varies depending on the distance set.
[0214] Although the above describes the case of 16 directions, in the case of 8 directions (numDir=8), the search direction may be determined according to the distance by setting dir=0...numDir-1, DIR1=0...numDir / 2-1. In the case of 4 directions (numDir=4), the search direction may be determined according to the distance by setting dir=0...numDir-1, DIR1=0...numDir / 2-1.
[0215] This provides the effect of reducing the amount of calculation without reducing the coding efficiency. In addition, by selecting the direction to search for each of a plurality of distances, it is possible to reduce dependency and improve parallelism, thereby providing the effect of enabling efficient calculation.
[0216] (Early termination of search) The MMVD prediction unit 30376 calculates the template matching cost for each MMVD candidate mmvdMergeCand. At this time, the MMVD prediction unit 30376 may end the search if the above cost satisfies the following condition.
[0217] <Condition 1 for early termination of search> The cost is smaller than a predetermined threshold TH1. Here, TH1 may be set as follows using the width and height of the target block: TH1 = C * width * height C is a constant (for example, 5). That is, when the cost per pixel achieves a value less than C, the MMVD prediction unit 30376 ends the search.
[0218] <Condition 2 for early termination of search> The cost is greater than the minimum cost by a predetermined constant TH2, where TH2 is, for example, 1.05 or 1.1. That is, if sufficient cost reduction cannot be achieved for candidates with large distances, the MMVD prediction unit 30376 does not search for candidates with larger distances.
[0219] The above-mentioned early search termination has the effect of reducing the amount of calculation without reducing the coding efficiency.
[0220] (Inter-prediction image generation unit 309) When predMode indicates inter prediction, the inter prediction image generation unit 309 generates a prediction image of a block or sub-block by inter prediction using the inter prediction parameters and reference picture input from the inter prediction parameter derivation unit 303.
[0221] 4 is a schematic diagram showing a configuration of an inter-prediction image generation unit 309 included in the prediction image generation unit 308 according to this embodiment. The inter-prediction image generation unit 309 includes a motion compensation unit (prediction image generation device) 3091 and a synthesis unit 3095. The synthesis unit 3095 includes an IntraInter synthesis unit 30951, a GPM synthesis unit 30952, a BIO unit 30954, and a weight prediction unit 3094.
[0222] (Motion Compensation) The motion compensation unit 3091 (interpolated image generation unit 3091) generates an interpolated image (motion compensated image) by reading a reference block from the reference picture memory 306 based on the inter prediction parameters (predFlagLX, refIdxLX, mvLX) input from the inter prediction parameter derivation unit 303. The reference block is a block at a position shifted by mvLX from the position of the target block on the reference picture RefPicLX specified by refIdxLX. Here, if mvLX is not integer precision, a filter for generating pixels at decimal positions called a motion compensation filter is applied to generate an interpolated image.
[0223] The motion compensation unit 3091 first derives an integer position (xInt, yInt) and a phase (xFrac, yFrac) corresponding to coordinates (x, y) in the prediction block using the following formula.
[0224] xInt = xPb+(mvLX[0]>>(log2(MVPREC)))+x xFrac = mvLX[0]&(MVPREC-1) yInt = yPb+(mvLX[1]>>(log2(MVPREC)))+y yFrac = mvLX[1]&(MVPREC-1) Here, (xPb, yPb) are the upper left coordinates of a block of size bW*bH, where x=0...bW-1, y=0...bH-1, and MVPREC indicates the accuracy of mvLX (1 / MVPREC pixel accuracy), e.g., MVPREC=16.
[0225] The motion compensation unit 3091 derives a temporary image temp[][] by performing horizontal interpolation processing on the reference picture refImg using an interpolation filter. In the following, Σ is the sum for k=0..NTAP-1, shift1 is a normalization parameter that adjusts the value range, and offset1=1<<(shift1-1).
[0226] temp[x][y] = (ΣmcFilter[xFrac][k]*refImg[xInt+k-NTAP / 2+1][yInt]+offset1)>>shift1 Next, the motion compensation unit 3091 derives the interpolated image Pred[][] by vertically interpolating the temporary image temp[][]. In the following, Σ is the sum for k=0..NTAP-1, shift2 is a normalization parameter that adjusts the value range, and offset2=1<<(shift2-1).
[0227] Pred[x][y] = (ΣmcFilter[yFrac][k]*temp[x][y+k-NTAP / 2+1]+offset2)>>shift2 In the case of bi-prediction, the above Pred[][] is used to derive interpolated images PredL0[][] and PredL1[][] for each of the L0 and L1 lists, and an interpolated image Pred[][] is generated from PredL0[][] and PredL1[][].
[0228] (GPM synthesis processing) When ciip_mode is 0, the GPM synthesis unit 30952 generates a predicted image in GPM mode by a weighted sum of multiple inter predicted images.
[0229] (IntraInter synthesis processing) When ciip_mode is 1, the IntraInter synthesis unit 30951 generates a predicted image in CIIP mode by using a weighted sum of an inter predicted image and an intra predicted image.
[0230] (BIO forecast) In bi-prediction mode, the BIO unit 30954 generates a predicted image by referring to two predicted images (a first predicted image and a second predicted image) and a gradient correction term.
[0231] (Weighted prediction) The weighted prediction unit 3094 generates a predicted image of the block by multiplying the interpolated image PredLX by a weighting coefficient.
[0232] The inter-prediction image generation unit 309 outputs the generated prediction image of the block to the addition unit 312 .
[0233] (Configuration of a video encoding device) Next, the configuration of the video encoding device 11 according to this embodiment will be described. Fig. 16 is a block diagram showing the configuration of the video encoding device 11 according to this embodiment. The video encoding device 11 includes a prediction image generating unit 101, a subtraction unit 102, a transformation and quantization unit 103, an inverse quantization and inverse transformation unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, an encoding parameter determining unit 110, a parameter encoding unit 111, a prediction parameter derivation unit 120, and an entropy encoding unit 104.
[0234] The predicted image generation unit 101 generates a predicted image for each CU. The predicted image generation unit 101 includes the inter predicted image generation unit 309 and the intra predicted image generation unit already described, and therefore a description thereof will be omitted.
[0235] The subtraction unit 102 generates a prediction error by subtracting the pixel values of the predicted image of the block input from the predicted image generation unit 101 from the pixel values of the image T. The subtraction unit 102 outputs the prediction error to the transformation and quantization unit 103.
[0236] The transform / quantization unit 103 calculates transform coefficients by frequency transforming the prediction errors input from the subtraction unit 102, and derives quantized transform coefficients by quantizing the prediction errors. The transform / quantization unit 103 outputs the quantized transform coefficients to the parameter coding unit 111 and the inverse quantization / inverse transform unit 105.
[0237] The transform / quantization unit 103 includes a separation transform unit (first transform unit), a non-separation transform unit (second transform unit), and a scaling unit.
[0238] The separate transform unit applies a separate transform to the prediction error, and the scaling unit scales the transform coefficients with a quantization matrix.
[0239] The inverse quantization and inverse transform unit 105 is the same as the inverse quantization and inverse transform unit 311 in the video decoding device 31, and a description thereof will be omitted.
[0240] The parameter coding unit 111 includes a header coding unit 1110, a CT information coding unit 1111, and a CU coding unit 1112 (prediction mode coding unit). The CU coding unit 1112 further includes a TU coding unit 1114. The following describes an outline of the operation of each module.
[0241] The header encoding unit 1110 performs encoding processing of parameters such as header information, division information, prediction information, and quantized transform coefficients.
[0242] The CT information encoding unit 1111 encodes the QT, MT (BT, TT) division information and the like.
[0243] The CU encoding unit 1112 encodes the CU information, prediction information, division information, and so on.
[0244] When a prediction error is included in a TU, the TU encoding unit 1114 encodes the QP update information and the quantized prediction error.
[0245] The CT information encoding unit 1111 and the CU encoding unit 1112 supply syntax elements such as inter prediction parameters (predMode, general_merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX), intra prediction parameters, and quantized transform coefficients to the parameter encoding unit 111.
[0246] The entropy coding unit 104 receives the quantized transform coefficients and the coding parameters from the parameter coding unit 111. The entropy coding unit 104 entropy codes these parameters (division information, prediction parameters) to generate and output a coded stream Te.
[0247] The prediction parameter derivation unit 120 is a means including the inter prediction parameter encoding unit 112 and an intra prediction parameter encoding unit, and derives intra prediction parameters and inter prediction parameters from the parameters input from the encoding parameter determination unit 110. The derived intra prediction parameters and inter prediction parameters are output to the parameter encoding unit 111.
[0248] (Configuration of the inter-prediction parameter encoding unit) 17, the inter prediction parameter encoding unit 112 includes a parameter encoding control unit 1121 and an inter prediction parameter derivation unit 303. The inter prediction parameter derivation unit 303 has the same configuration as the video decoding device. The parameter encoding control unit 1121 includes a merge index derivation unit 11211 and a vector candidate index derivation unit 11212.
[0249] The merge index derivation unit 11211 derives merge candidates and the like, and outputs them to the inter prediction parameter derivation unit 303. The vector candidate index derivation unit 11212 derives predictive vector candidates and the like, and outputs them to the inter prediction parameter derivation unit 303 and the parameter coding unit 111.
[0250] (Configuration of intra-prediction parameter encoding unit) The intra-prediction parameter coding unit includes a parameter coding control unit and an intra-prediction parameter derivation unit. The intra-prediction parameter derivation unit has a common configuration with the video decoding device.
[0251] However, unlike the video decoding device, the inputs to the inter prediction parameter derivation unit 303 and the intra prediction parameter derivation unit are the coding parameter determination unit 110 and the prediction parameter memory 108 , and they are output to the parameter coding unit 111 .
[0252] The adder 106 generates a decoded image by adding, for each pixel, the pixel value of the predicted block input from the predicted image generation unit 101 and the prediction error input from the inverse quantization and inverse transform unit 105. The adder 106 stores the generated decoded image in a reference picture memory 109.
[0253] The loop filter 107 performs deblocking filtering, SAO, and ALF on the decoded image generated by the adder 106. Note that the loop filter 107 does not necessarily have to include the above three types of filters, and may be configured, for example, as only a deblocking filter.
[0254] The prediction parameter memory 108 stores the prediction parameters generated by the coding parameter determination unit 110 in a predetermined location for each current picture and CU.
[0255] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at a predetermined position for each current picture and CU.
[0256] The coding parameter determination unit 110 selects one set from among a plurality of sets of coding parameters. The coding parameters are the above-mentioned QT, BT or TT division information, prediction parameters, or parameters to be coded that are generated in relation to these. The predicted image generation unit 101 generates a predicted image using these coding parameters.
[0257] The coding parameter determination unit 110 calculates an RD cost value indicating the amount of information and the coding error for each of the multiple sets. The RD cost value is, for example, the sum of the code amount and a value obtained by multiplying the square error by a coefficient λ. The code amount is calculated by entropy coding the quantization error and the coding parameters. The squared error is the sum of squares of the prediction errors calculated in the subtraction unit 102. The coefficient λ is a preset real number greater than zero. The coding parameter determination unit 110 selects a set of coding parameters that minimizes the calculated cost value. The coding parameter determination unit 110 outputs the determined coding parameters to the parameter coding unit 111 and the prediction parameter derivation unit 120.
[0258] In addition, a part of the video encoding device 11 and the video decoding device 31 in the above-mentioned embodiment, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generating unit 308, the inverse quantization and inverse transform unit 311, the addition unit 312, the prediction parameter derivation unit 320, the predicted image generating unit 101, the subtraction unit 102, the transform and quantization unit 103, the entropy coding unit 104, the inverse quantization and inverse transform unit 105, the loop filter 107, the coding parameter determination unit 110, the parameter coding unit 111, and the prediction parameter derivation unit 120 may be realized by a computer. In this case, a program for realizing this control function may be recorded in a computer-readable recording medium, and the program recorded in the recording medium may be read into and executed by a computer system. In addition, the "computer system" referred to here is a computer system built into either the video encoding device 11 or the video decoding device 31, and includes hardware such as an OS and peripheral devices. In addition, "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. Furthermore, "computer-readable recording medium" may also include devices that dynamically hold a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and devices that hold a program for a certain period of time, such as volatile memory inside a computer system that serves as a server or client in such cases. Furthermore, the above program may be one that realizes part of the above-mentioned functions, or may be one that can realize the above-mentioned functions in combination with a program already recorded in the computer system.
[0259] In addition, a part or the whole of the video encoding device 11 and the video decoding device 31 in the above-mentioned embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the video encoding device 11 and the video decoding device 31 may be individually made into a processor, or a part or the whole may be integrated into a processor. The integrated circuit method is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. Furthermore, when an integrated circuit technology that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit based on that technology may be used.
[0260] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes, etc. are possible within the scope that does not deviate from the gist of the present invention.
[0261] The present invention is not limited to the above-described embodiment, and various modifications are possible within the scope of the claims. In other words, the technical scope of the present invention also includes embodiments obtained by combining technical means that are appropriately modified within the scope of the claims. [Industrial Applicability]
[0262] The embodiments of the present invention can be suitably applied to a video decoding device that decodes coded data in which image data is coded, and a video coding device that generates coded data in which image data is coded, and can also be suitably applied to the data structure of coded data that is generated by a video coding device and referenced by the video decoding device. [Explanation of symbols]
[0263] 31 Video Decoding Device 301 Entropy Decoding Unit 302 Parameter Decoding Unit 3022 CU Decoding Unit 3024 TU Decoding Unit 303 Inter-prediction parameter derivation unit 30376 MMVD Prediction Department 305, 107 Loop Filter 306, 109 Reference Picture Memory 307, 108 Prediction parameter memory 308, 101 Prediction image generation unit 309 Inter-prediction image generation unit 311, 105 Inverse quantization and inverse transformation unit 312, 106 Addition section 320 Prediction Parameter Derivation Unit 11 Video Encoding Device 102 Subtraction section 103 Transformation and Quantization Section 104 Entropy coding unit 110 Encoding parameter determination unit 111 Parameter Encoding Unit 112 Inter-prediction parameter coding unit 120 Prediction parameter derivation part
Claims
1. An MPEG motion vector difference (MMVD) prediction unit that obtains a motion vector by adding a difference vector in a predetermined distance and a predetermined direction to a predicted motion vector of a target block, and a parameter decoding unit that decodes an index for specifying the difference vector from an MMVD candidate list from encoded data, the moving image decoding apparatus comprising: The MMVD prediction unit is characterized by deriving a difference vector in a specific distance and direction from the predetermined distance and direction, performing a search by deriving a template matching cost for the difference vector, and deriving an MMVD candidate list by inserting a difference vector candidate according to the cost. In the search, only a subset (DIR1) of the first direction is searched at a predetermined distance, and only a subset (DIR2) of the second direction is searched at a distance other than the predetermined distance. A moving image decoding apparatus characterized by the above.
2. The moving image decoding apparatus according to claim 1, wherein the subset (DIR1) of the first direction includes eight directions of horizontal, vertical, and diagonal 45 degrees.
3. The moving image decoding apparatus according to claim 1, wherein the subset (DIR2) of the second direction includes three directions of the direction giving the minimum cost among the template matching costs for the number of searches obtained at the previous distance and the directions adjacent thereto.
4. The moving image decoding apparatus according to claim 1, wherein the subset (DIR2) of the second direction includes two directions giving the smallest cost among the template matching costs for the number of searches obtained at the previous distance and up to six directions adjacent to each of them.
5. In the search, for each distance candidate, a direction included in a subset (DIR2) of the second direction is searched, and the direction included in the subset (DIR2) of the second direction is updated for each distance candidate. A moving image decoding apparatus according to claim 1, characterized by the above.
6. An MPEG motion vector difference (MMVD) prediction unit that obtains a motion vector by adding a difference vector in a predetermined distance and a predetermined direction to a predicted motion vector of a target block, and a parameter decoding unit that decodes an index for specifying the difference vector from an MMVD candidate list from encoded data, the moving image decoding apparatus comprising: The MMVD prediction unit derives a differential vector of a specific distance and direction from the predetermined distance and direction, performs a search by deriving a template matching cost for the differential vector, and derives an MMVD candidate list by inserting a differential vector candidate according to the cost, and is characterized in that, In the search, only a subset of the first direction (DIR1) is searched for distances included in a predetermined distance set, and only a subset of the second direction (DIR2) is searched for distances included in a distance set other than the predetermined distance set. A moving image decoding apparatus characterized by that.
7. In the search, for a distance set including a plurality of distance candidates, a direction included in a subset of the second direction (DIR2) is searched, and for each distance set including a plurality of distance candidates, a subset of the second direction (DIR2) is searched. The moving image decoding apparatus according to claim 6, characterized in that the direction included in (DIR2) is updated.
8. In the search, the moving image decoding apparatus according to claim 7, characterized in that the number of directions included in a subset of the second direction (DIR2) is different according to the distance set.
9. A moving image encoding apparatus including an MMVD prediction unit that obtains a motion vector by adding a differential vector of a predetermined distance and a predetermined direction to a predicted motion vector of a target block, and a parameter encoding unit that encodes an index for specifying a differential vector from an MMVD candidate list, The MMVD prediction unit derives a differential vector of a specific distance and direction from the predetermined distance and direction, performs a search by deriving a template matching cost for the differential vector, and derives an MMVD candidate list by inserting a differential vector candidate according to the cost, and is characterized in that, In the search, only a subset of the first direction (DIR1) is searched for a predetermined distance, and only a subset of the second direction (DIR2) is searched for a distance other than the predetermined distance. A moving image encoding apparatus characterized by that.
10. A computer-readable recording medium for recording a program, The program causes a computer to, decode an index for specifying a differential vector from an MMVD candidate list from encoded data, derive a differential vector of a specific distance and direction from a predetermined distance and direction, A step of performing search by deriving a template matching cost for the differential vector; A step of deriving an MMVD candidate list by inserting differential vector candidates according to the cost; are executed, In the search, only a subset (DIR1) in a first direction is searched at a predetermined distance, and only a subset (DIR2) in a second direction is searched at a distance other than the predetermined distance. A computer-readable recording medium characterized by that.