Moving image decoding apparatus and moving image encoding apparatus
By optimizing DMVR processing with weighted prediction and specific conditions, the complexity of bidirectional prediction methods is reduced, improving the efficiency of moving image decoding and encoding without compromising image quality.
Patent Information
- Application Number
- JP2024156934
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-03-08
- Filing Date
- 2024-09-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-03-06
AI Technical Summary
The high processing complexity of bidirectional prediction methods, such as DMVR and BDOF, in moving image encoding and decoding processes, particularly when correcting motion vectors using two prediction images, hinders efficient image quality improvement.
Implementing a moving image decoding and encoding apparatus that utilizes DMVR processing with weighted prediction and specific conditions for setting dmvrFlag based on the presence of weight coefficients and offsets in reference pictures, reducing complexity by optimizing the DMVR process.
The proposed solution reduces the complexity of high-quality image processing while maintaining or improving prediction image quality, thereby enhancing the efficiency of moving image decoding and encoding operations.
Smart Images

Figure 0007706621000003 
Figure 0007706621000004 
Figure 0007706621000005
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a moving image decoding apparatus and a moving image encoding apparatus.
Background Art
[0002] In order to efficiently transmit or record a moving image, a moving image encoding apparatus that generates encoded data by encoding a moving image, and a moving image decoding apparatus that generates a decoded image by decoding the encoded data are used.
[0003] Specific moving image encoding methods include, for example, the H.264 / AVC and HEVC (High-Efficiency Video Coding) methods.
[0004] In such a moving image encoding method, an image (picture) constituting a moving image is managed by a hierarchical structure including a slice obtained by dividing the image, a coding tree unit (CTU) obtained by dividing the slice, a coding unit (sometimes called a coding unit (CU)) obtained by dividing the coding tree unit, and a transform unit (TU) obtained by dividing the coding unit, and is encoded / decoded for each CU.
[0005] Also, in such a moving image encoding method, usually, a prediction image is generated based on a local decoded image obtained by encoding / decoding an input image, and a prediction error (sometimes called a "difference image" or a "residual image") obtained by subtracting the prediction image from the input image (original image) is encoded. Examples of the method for generating a prediction image include inter-picture prediction (inter prediction) and intra-picture prediction (intra prediction).
[0006] Also, Non-Patent Document 1 is cited as a technique for recent moving image encoding and decoding.
Prior Art Documents
Non-Patent Literature
[0007]
Non-Patent Literature 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] When deriving the bidirectional prediction image of Non-Patent Literature 1, the prediction (BDOF prediction) using the DMVR process of correcting the motion vector using two prediction images to improve the quality of the prediction image or the BDOF process of improving the quality of the prediction image using the gradient image has a problem of high processing complexity.
[0009] An embodiment of the present invention aims to realize an image decoding device and an image encoding device that reduce the complexity of these high-quality processing.
Means for Solving the Problems
[0010] To solve the above problems, an image decoding device according to an aspect of the present invention is A moving image decoding device that performs DMVR (Decoder side Motion Vector Refinement) processing using two reference pictures and motion vectors mvL0 and mvL1, A DMVR unit that executes the DMVR process when a dmvrFlag indicating whether the DMVR process is to be performed is TRUE, A weighted prediction unit that executes weighted prediction using a first weight coefficient, a first offset, a second weight coefficient, and a second offset. The DMVR unit sets TRUE to the dmvrFlag based on a predetermined condition. Here, the predetermined condition for setting the dmvrFlag to TRUE includes that both luma_weight_l0_flag[refIdxL0] and luma_weight_l1_flag[refIdxL1] are FALSE. The luma_weight_l0_flag[refIdxL0] indicates whether the first weight coefficient and the first offset of the luminance corresponding to the L0 reference picture indicated by the reference picture index refIdxL0 exist. The luma_weight_l1_flag[refIdxL1] indicates whether the second weight coefficient and the second offset of the luminance corresponding to the L1 reference picture indicated by the reference picture index refIdxL1 exist.
[0011] Also, an image encoding apparatus according to an aspect of the present invention is A moving image encoding apparatus that performs DMVR (Decoder side Motion Vector Refinement) processing using two reference pictures and motion vectors mvL0 and mvL1, A DMVR unit that executes the DMVR process when a dmvrFlag indicating whether the DMVR process is to be performed is TRUE, A weighted prediction unit that performs weighted prediction using a first weight coefficient, a first offset, a second weight coefficient, and a second offset, and The DMVR unit sets TRUE to the dmvrFlag based on a predetermined condition. Here, the predetermined condition for setting the dmvrFlag to TRUE includes that both luma_weight_l0_flag[refIdxL0] and luma_weight_l1_flag[refIdxL1] are FALSE. The luma_weight_l0_flag[ refIdxL0 ] indicates whether or not there exist the first weight coefficient and the first offset of the luminance corresponding to the L0 reference picture indicated by the reference picture index refIdxL0. The luma_weight_l1_flag[ refIdxL1 ] indicates whether or not there exist the second weight coefficient and the second offset of the luminance corresponding to the L1 reference picture indicated by the reference picture index refIdxL1.
Advantages of the Invention
[0012] According to the above configuration, it is possible to realize an image decoding apparatus and an image encoding apparatus that reduce the complexity of high image quality processing.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
[0014] (First embodiment) Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0015] FIG. 1 is a schematic diagram showing the configuration of an image transmission system 1 according to this embodiment.
[0016] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding an image to be encoded and decodes the transmitted encoded stream to display an image. The image transmission system 1 includes a moving image encoding device (image encoding device) 11, a network 21, a moving image decoding device (image decoding device) 31, and a moving image display device (image display device) 41.
[0017] An image T is input to the moving image encoding device 11.
[0018] The network 21 transmits the encoded stream Te generated by the moving image encoding device 11 to the moving image decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network 21 is not necessarily limited to a two-way communication network and may be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Further, the network 21 may be replaced by a storage medium on which the encoded stream Te such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark) is recorded.
[0019] The moving image decoding device 31 decodes each of the encoded streams Te transmitted by the network 21 and generates one or more decoded images Td obtained by the decoding.
[0020] The moving image display device 41 displays all or part of the one or more decoded images Td generated by the moving image decoding device 31. The moving image display device 41 includes, for example, a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Examples of the form of the display include a stationary type, a mobile type, and an HMD (Head Mount Display). Further, when the moving image decoding device 31 has high processing power, the moving image display device 41 displays an image with high image quality, and when it has only lower processing power, the moving image display device 41 displays an image that does not require high processing power and display ability.
[0021] <Operator> The operators used in this specification are described below.
[0022] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, | is a bitwise OR, |= is an OR assignment operator, and || represents a logical OR.
[0023] x?y:z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0).
[0024] Clip3(a, b, c) is a function that clips c to a value between a and b (inclusive), returning a if c < a, b if c > b, and c otherwise (where a <= b).
[0025] abs(a) is a function that returns the absolute value of a.
[0026] Int(a) is a function that returns the integer value of a.
[0027] floor(a) is a function that returns the largest integer less than or equal to a.
[0028] ceil(a) is a function that returns the smallest integer greater than or equal to a.
[0029] a / d represents the division of a by d (truncating the fractional part).
[0030] sign(a) is a function that returns the sign of a.
[0031] a^b represents a to the power of b.
[0032] <Structure of the Encoded Stream Te> Prior to the detailed description of the moving image encoding apparatus 11 and the moving image decoding apparatus 31 according to this embodiment, the data structure of the encoded stream Te generated by the moving image encoding apparatus 11 and decoded by the moving image decoding apparatus 31 will be described.
[0033] FIG. 4 is a diagram showing the hierarchical structure of data in the encoded stream Te. The encoded stream Te illustratively includes a sequence and a plurality of pictures constituting the sequence. FIGS. 4(a) to 4(f) respectively show an encoded video sequence that defines the sequence SEQ, an encoded picture that defines the picture PICT, an encoded slice that defines the slice S, encoded slice data that defines the slice data, an encoded tree unit included in the encoded slice data, and an encoded unit included in the encoded tree unit.
[0034] (Encoded video sequence) In the encoded video sequence, a set of data that the moving image decoding apparatus 31 refers to in order to decode the sequence SEQ to be processed is defined. As shown in FIG. 4(a), the sequence SEQ includes a video parameter set, a sequence parameter set SPS (Sequence Parameter Set), a picture parameter set PPS (Picture Parameter Set), a picture PICT, and supplemental enhancement information SEI (Supplemental Enhancement Information).
[0035] The video parameter set VPS defines a set of encoding parameters common to a plurality of moving images and a set of encoding parameters related to the plurality of layers and individual layers included in the moving image in a moving image composed of a plurality of layers.
[0036] In the sequence parameter set SPS, a set of encoding parameters that the moving image decoding device 31 refers to for decoding the target sequence is defined. For example, the width and height of a picture are defined. Note that there may be multiple SPSs. In that case, one of the multiple SPSs is selected from the PPS.
[0037] In the picture parameter set PPS, a set of encoding parameters that the moving image decoding device 31 refers to for decoding each picture in the target sequence is defined. For example, it includes a reference value of the quantization width (pic_init_qp_minus26) used for decoding a picture and a flag (weighted_pred_flag) indicating the application of weighted prediction. Note that there may be multiple PPSs. In that case, one of the multiple PPSs is selected from each picture in the target sequence.
[0038] (Encoded Picture) In the encoded picture, a set of data that the moving image decoding device 31 refers to for decoding the picture PICT to be processed is defined. As shown in FIG. 4(b), the picture PICT includes slices 0 to NS - 1 (NS is the total number of slices included in the picture PICT).
[0039] Note that hereinafter, when it is not necessary to distinguish each of slices 0 to NS - 1, the subscript of the symbol may be omitted in the description. The same applies to the data included in the encoding stream Te described below and other data with subscripts.
[0040] (Encoded Slice) In the encoded slice, a set of data that the moving image decoding device 31 refers to for decoding the slice S to be processed is defined. As shown in FIG. 4(c), the slice includes a slice header and slice data.
[0041] The slice header includes a group of coding parameters that the moving image decoding apparatus 31 refers to in order to determine the decoding method of the target slice. The slice type designation information (slice_type) for designating the slice type is an example of the coding parameters included in the slice header.
[0042] Examples of slice types that can be specified by the slice type designation information include: (1) an I slice that uses only intra prediction during encoding; (2) a P slice that uses uni-directional prediction or intra prediction during encoding; (3) a B slice that uses uni-directional prediction, bi-directional prediction, or intra prediction during encoding, etc. Note that inter prediction is not limited to uni-prediction and bi-prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referring to P and B slices, it refers to a slice including a block that can use inter prediction.
[0043] Note that the slice header may include a reference (pic_parameter_set_id) to the picture parameter set PPS.
[0044] (Encoded slice data) In the encoded slice data, a set of data that the moving image decoding apparatus 31 refers to in order to decode the slice data to be processed is defined. As shown in FIG. 4(d), the slice data includes CTUs. A CTU is a block of a fixed size (for example, 64x64) that constitutes a slice, and is sometimes also referred to as the largest coding unit (LCU).
[0045] (Coding tree unit) FIG. 4(e) defines a set of data that the moving image decoding apparatus 31 refers to in order to decode the CTU to be processed. The CTU is divided into coding units (CUs), which are the basic units of the encoding process, by recursive quadtree (QT) partitioning, binary tree (BT) partitioning, or ternary tree (TT) partitioning. The combination of BT partitioning and TT partitioning is called multi-tree (MT) partitioning. The nodes of the tree structure obtained by recursive quadtree partitioning are called coding nodes. The intermediate nodes of the quadtree, binary tree, and ternary tree are coding nodes, and the CTU itself is defined as the topmost coding node.
[0046] The CT includes, as CT information, a QT partitioning flag (qt_split_cu_flag) indicating whether QT partitioning is to be performed, an MT partitioning flag (mtt_split_cu_flag) indicating the presence or absence of MT partitioning, an MT partitioning direction (mtt_split_cu_vertical_flag) indicating the partitioning direction of MT partitioning, and an MT partitioning type (mtt_split_cu_binary_flage) indicating the partitioning type of MT partitioning. The qt_split_cu_flag, mtt_split_cu_flag, mtt_split_cu_vertical_flag, and mtt_split_cu_binary_flag are transmitted for each coding node.
[0047] FIG. 5 is a diagram showing an example of CTU partitioning. When the qt_split_cu_flag is 1, the coding node is divided into four coding nodes (FIG. 5(b)).
[0048] When the qt_split_cu_flag is 0 and the mtt_split_cu_flag is 0, the coding node is not divided and has one CU as a node (FIG. 5(a)). The CU is the terminal node of the coding node and is not further divided. The CU is the basic unit of the encoding process.
[0049] When mtt_split_cu_flag is 1, the coding node is MT-split as follows. When mtt_split_cu_vertical_flag is 0 and mtt_split_cu_binary_flag is 1, the coding node is horizontally split into two coding nodes (Figure 5(d)). When mtt_split_cu_vertical_flag is 1 and mtt_split_cu_binary_flag is 1, the coding node is vertically split into two coding nodes (Figure 5(c)). Also, when mtt_split_cu_vertical_flag is 0 and mtt_split_cu_binary_flag is 0, the coding node is horizontally split into three coding nodes (Figure 5(f)). When mtt_split_cu_vertical_flag is 1 and mtt_split_cu_binary_flag is 0, the coding node is vertically split into three coding nodes (Figure 5(e)). These are shown in Figure 5(g).
[0050] Also, when the size of the CTU is 64x64 pixels, the size of the CU can be any of 64x64 pixels, 64x32 pixels, 32x64 pixels, 32x32 pixels, 64x16 pixels, 16x64 pixels, 32x16 pixels, 16x32 pixels, 16x16 pixels, 64x8 pixels, 8x64 pixels, 32x8 pixels, 8x32 pixels, 16x8 pixels, 8x16 pixels, 8x8 pixels, 64x4 pixels, 4x64 pixels, 32x4 pixels, 4x32 pixels, 16x4 pixels, 4x16 pixels, 8x4 pixels, 4x8 pixels, and 4x4 pixels.
[0051] (Coding Unit) As shown in Figure 4(f), a set of data that the moving image decoding device 31 refers to for decoding the coding unit to be processed is defined in the CU. Specifically, the CU is composed of a CU header CUH, prediction parameters, transformation parameters, quantized transform coefficients, etc. The prediction mode, etc. are defined in the CU header.
[0052] The prediction process may be performed in units of CUs or in units of sub-CUs obtained by further dividing CUs. When the sizes of the CU and the sub-CU are equal, there is one sub-CU in the CU. When the CU is larger than the size of the sub-CU, the CU is divided into sub-CUs. For example, when the CU is 8x8 and the sub-CU is 4x4, the CU is divided into four sub-CUs consisting of two horizontal divisions and two vertical divisions.
[0053] There are two types of prediction (prediction modes), intra prediction and inter prediction. Intra prediction is prediction within the same picture, and inter prediction refers to the prediction process performed between different pictures (for example, between display times).
[0054] The transformation and quantization processes are performed in units of CUs, but the quantized transform coefficients may be entropy-coded in units of sub-blocks such as 4x4.
[0055] (Prediction parameters) The predicted image is derived from the prediction parameters associated with the block. The prediction parameters include intra prediction and inter prediction parameters.
[0056] Hereinafter, the inter prediction parameters will be described. The inter prediction parameters are composed of the prediction list usage flags predFlagL0 and predFlagL1, the reference picture indices refIdxL0 and refIdxL1, and the motion vectors mvL0 and mvL1. The prediction list usage flags predFlagL0 and predFlagL1 are flags indicating whether reference picture lists called the L0 list and the L1 list are used for inter prediction, respectively. When the value is 1, the corresponding reference picture list is used for inter prediction. In this specification, when it is described as "a flag indicating whether XX", when the flag is other than 0 (for example, 1), it is considered that XX is the case, and when the flag is 0, it is considered that XX is not the case. In logical negation, logical product, etc., 1 is treated as true and 0 is treated as false (the same applies hereinafter). However, in actual devices and methods, other values can also be used as true values and false values.
[0057] Syntax elements for deriving inter prediction parameters include, for example, an affine flag, a merge flag, a merge index, an inter prediction identifier, a reference picture index, a prediction vector index, a differential vector, and a motion vector precision mode.
[0058] (Reference Picture List) The reference picture list is a list consisting of reference pictures stored in the reference picture memory 306. FIG. 6 is a conceptual diagram showing an example of a reference picture and a reference picture list. In FIG. 6(a), rectangles represent pictures, arrows represent the reference relationships of the pictures, the horizontal axis represents time, I, P, and B in the rectangles represent an intra picture, a single prediction picture, and a dual prediction picture, respectively, and the numbers in the rectangles indicate the decoding order. As shown in the figure, the decoding order of the pictures is I0, P1, B2, B3, B4, and the display order is I0, B3, B2, B4, P1. FIG. 6(b) shows an example of the reference picture list of the picture B3 (target picture). The reference picture list is a list representing candidates for reference pictures, and one picture (slice) may have one or more reference picture lists. In the example of the figure, the target picture B3 has two reference picture lists, the L0 list RefPicList0 and the L1 list RefPicList1. For each individual CU, the reference picture index refIdxLX specifies which picture in the reference picture list RefPicListX (X = 0 or 1) is actually referenced. The figure shows an example where refIdxL0 = 2 and refIdxL1 = 0. Note that LX is a description method used when distinguishing between L0 prediction and L1 prediction is not required. Hereinafter, by replacing LX with L0 and L1, the parameters for the L0 list and the parameters for the L1 list are distinguished.
[0059] (Merge Prediction and AMVP Prediction) The decoding (encoding) method of prediction parameters includes a merge prediction mode and an AMVP (Advanced Motion Vector Prediction) mode, and the merge flag merge_flag is a flag for identifying these. The merge prediction mode is a mode that derives from the prediction parameters of already processed neighboring blocks without including the prediction list utilization flag predFlagLX (or inter-prediction identifier inter_pred_idc), reference picture index refIdxLX, and motion vector mvLX in the encoded data. The AMVP mode is a mode that includes the inter-prediction identifier inter_pred_idc, reference picture index refIdxLX, and motion vector mvLX in the encoded data. Note that the motion vector mvLX is encoded as a motion vector index mvp_LX_idx for identifying the prediction vector mvpLX, a differential vector mvdLX, and a motion vector accuracy mode amvr_mode. The merge prediction mode is a mode that selects a merge candidate derived from the motion information of adjacent blocks and obtains the motion vector mvLX (motion vector information). In addition to the merge prediction mode, there may be an affine prediction mode identified by the affine flag affine_flag. As a form of the merge prediction mode, there may be a skip mode identified by the skip flag skip_flag. Note that the skip mode is a mode used to derive prediction parameters in the same way as the merge mode and does not include the prediction error (residual image, residual information) in the encoded data. In other words, when the skip flag skip_flag is 1, for the target CU, only the syntax related to the merge mode such as the skip flag skip_flag and the merge index merge_idx is included, and the motion vector, residual information, etc. are not included in the encoded data.
[0060] (Motion Vector) The motion vector mvLX indicates the shift amount between blocks on two different pictures. The prediction vector and differential vector related to the motion vector mvLX are called the prediction vector mvpLX and the differential vector mvdLX, respectively.
[0061] (Inter-prediction identifier inter_pred_idc and prediction list usage flag predFlagLX) The inter-prediction identifier inter_pred_idc is a value indicating the type and number of reference pictures, and takes any one of the values of PRED_L0, PRED_L1, and PRED_BI. PRED_L0 and PRED_L1 indicate single prediction using one reference picture managed in the L0 list and the L1 list, respectively. PRED_BI indicates bi-prediction BiPred using two reference pictures managed in the L0 list and the L1 list.
[0062] The merge index merge_idx is an index indicating which prediction parameter among the prediction parameter candidates (merge candidates) derived from the processed blocks is used as the prediction parameter of the target block.
[0063] The relationship between the inter-prediction identifier inter_pred_idc and the prediction list usage flags predFlagL0 and predFlagL1 is as follows and they are mutually convertible.
[0064] inter_pred_idc = (predFlagL1 << 1) + predFlagL0 predFlagL0 = inter_pred_idc & 1 predFlagL1 = inter_pred_idc >> 1 (Determination of bi-prediction biPred) The flag biPred indicating whether it is bi-prediction BiPred can be derived depending on whether both of the two prediction list usage flags are 1. For example, it can be derived by the following formula.
[0065] biPred = (predFlagL0 == 1 && predFlagL1 == 1) Alternatively, the flag biPred can also be derived depending on whether the inter-prediction identifier indicates that two prediction lists (reference pictures) are used. For example, it can be derived using the following formula.
[0066] biPred = (inter_pred_idc == PRED_BI)? 1 : 0 (Configuration of Video Decoding Apparatus) The configuration of the video decoding apparatus 31 (FIG. 7) according to this embodiment will be described.
[0067] The video decoding apparatus 31 includes an entropy decoding unit 301, a parameter decoding unit 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a prediction image generation unit (prediction image generation apparatus) 308, an inverse quantization / inverse transform unit 311, and an addition unit 312. Note that, in accordance with the video encoding apparatus 11 described later, there is also a configuration in which the video decoding apparatus 31 does not include the loop filter 305.
[0068] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit), not shown in the figure. The CU decoding unit 3022 further includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS, and slice headers (slice information) from the encoded data. The CT information decoding unit 3021 decodes CT from the encoded data. The CU decoding unit 3022 decodes CU from the encoded data. The TU decoding unit 3024 decodes QP update information (quantization correction value) and quantized prediction error (residual_coding) from the encoded data when the TU contains a prediction error.
[0069] When it is not in the skip mode (skip_mode == 0), the TU decoding unit 3024 decodes QP update information (quantization correction value) and quantization prediction error (residual_coding) from the encoded data. More specifically, when skip_mode == 0, the TU decoding unit 3024 decodes a flag cu_cbp indicating whether the quantization prediction error is included in the target block from the encoded data, and decodes the quantization prediction error when cu_cbp is 1. When cu_cbp does not exist in the encoded data, the TU decoding unit 3024 derives cu_cbp as 0.
[0070] In addition, the parameter decoding unit 302 includes an inter-prediction parameter decoding unit 303 and an intra-prediction parameter decoding unit 304 (not shown). The prediction image generation unit 308 includes an inter-prediction image generation unit 309 and an intra-prediction image generation unit 310.
[0071] In the following, examples using CTUs and CUs as processing units are described, but the present invention is not limited to this example, and processing may be performed in units of sub-CUs. Alternatively, CTUs and CUs may be read as blocks, and sub-CUs may be read as sub-blocks, and processing may be performed in units of blocks or sub-blocks.
[0072] The entropy decoding unit 301 performs entropy decoding on the encoded stream Te input from the outside to decode individual codes (syntax elements). The decoded codes include prediction information for generating a prediction image, a prediction error for generating a difference image, and the like.
[0073] The entropy decoding unit 301 outputs the decoded codes to the parameter decoding unit 302. The decoded codes are, for example, predMode, merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX, amvr_mode, etc. The control of which code to decode is performed based on the instruction of the parameter decoding unit 302.
[0074] (Configuration of Inter-Prediction Parameter Decoding Unit) Based on the code input from the entropy decoding unit 301, the inter-prediction parameter decoding unit 303 decodes the inter-prediction parameters by referring to the prediction parameters stored in the prediction parameter memory 307. Further, the inter-prediction parameter decoding unit 303 outputs the decoded inter-prediction parameters to the predicted image generation unit 308 and stores them in the prediction parameter memory 307.
[0075] FIG. 8 is a schematic diagram showing the configuration of the inter-prediction parameter decoding unit 303 according to the present embodiment. The inter-prediction parameter decoding unit 303 includes a merge prediction unit 30374, a DMVR unit 30375, a sub-block prediction unit (affine prediction unit) 30372, an MMVD prediction unit 30376, a Triangle prediction unit 30377, an AMVP prediction parameter derivation unit 3032, and an addition unit 3038. The merge prediction unit 30374 includes a merge prediction parameter derivation unit 3036. Since the AMVP prediction parameter derivation unit 3032, the merge prediction parameter derivation unit 3036, and the affine prediction unit 30372 are common means in the moving image encoding device and the moving image decoding device, they may be collectively referred to as a motion vector derivation unit (motion vector derivation device).
[0076] (Affine prediction unit) The affine prediction unit 30372 derives the affine prediction parameters of the target block. In the present embodiment, as the affine prediction parameters, the motion vectors (mv0_x, mv0_y) and (mv1_x, mv1_y) of two control points (V0, V1) of the target block are derived. Specifically, the motion vectors of each control point may be derived by predicting from the motion vectors of adjacent blocks of the target block, or the motion vectors of each control point may be derived by the sum of the predicted vectors derived as the motion vectors of the control points and the difference vectors derived from the encoded data.
[0077] (Merge prediction) FIG. 9(a) is a schematic diagram showing the configuration of a merge prediction parameter derivation unit 3036 included in a merge prediction unit 30374. The merge prediction parameter derivation unit 3036 includes a merge candidate derivation unit 30361 and a merge candidate selection unit 30362. Note that a merge candidate is configured to include a prediction list usage flag predFlagLX, a motion vector mvLX, and a reference picture index refIdxLX, and is stored in a merge candidate list. An index is assigned to the merge candidate stored in the merge candidate list according to a predetermined rule.
[0078] The merge candidate derivation unit 30361 derives a merge candidate by directly using the motion vector and the reference picture index refIdxLX of the decoded adjacent block.
[0079] The order of storing in the merge candidate list mergeCandList[] is, for example, spatial merge candidates A1, B1, B0, A0, B2, temporal merge candidate Col, pairwise merge candidate avgK, zero merge candidate ZK. Note that a reference block that is not available (such as when the block is intra predicted) is not stored in the merge candidate list.
[0080] The merge candidate selection unit 30362 selects a merge candidate N indicated by a merge index merge_idx from among the merge candidates included in the merge candidate list by the following formula.
[0081] N = mergeCandList[merge_idx] Here, N is a label indicating a merge candidate, and takes values such as A1, B1, B0, A0, B2, Col, avgK, ZK. The motion information of the merge candidate indicated by the label N is indicated by (mvLXN[0], mvLXN[1]), predFlagLXN, and refIdxLXN.
[0082] The merge candidate selection unit 30362 selects the motion information (mvLXN[0], mvLXN[1]), predFlagLXN, and refIdxLXN of the selected merge candidate as the inter-prediction parameters of the target block. The merge candidate selection unit 30362 stores the inter-prediction parameters of the selected merge candidate in the prediction parameter memory 307 and outputs them to the prediction image generation unit 308.
[0083] (AMVP Prediction) FIG. 9(b) is a schematic diagram showing the configuration of the AMVP prediction parameter derivation unit 3032 according to the present embodiment. The AMVP prediction parameter derivation unit 3032 includes a vector candidate derivation unit 3033 and a vector candidate selection unit 3034. The vector candidate derivation unit 3033 derives prediction vector candidates from the decoded motion vectors mvLX of adjacent blocks stored in the prediction parameter memory 307 based on the reference picture index refIdxLX, and stores them in the prediction vector candidate list mvpListLX[].
[0084] The vector candidate selection unit 3034 selects the motion vector mvpListLX[mvp_LX_idx] indicated by the prediction vector index mvp_LX_idx among the prediction vector candidates in the prediction vector candidate list mvpListLX[] as the prediction vector mvpLX. The vector candidate selection unit 3034 outputs the selected prediction vector mvpLX to the addition unit 3038.
[0085] Note that the prediction vector candidates are derived by scaling the motion vectors of decoded adjacent blocks within a predetermined range from the target block. Note that the adjacent blocks include blocks that are spatially adjacent to the target block, such as the left block and the upper block, and in addition, regions that are temporally adjacent to the target block, such as regions obtained from the prediction parameters of blocks that include the same position as the target block but have different display times.
[0086] The addition unit 3038 adds the predicted vector mvpLX input from the AMVP prediction parameter derivation unit 3032 and the decoded difference vector mvdLX to calculate the motion vector mvLX. The addition unit 3038 outputs the calculated motion vector mvLX to the prediction image generation unit 308 and the prediction parameter memory 307.
[0087] mvLX[0] = mvpLX[0]+mvdLX[0] mvLX[1] = mvpLX[1]+mvdLX[1] The motion vector accuracy mode amvr_mode is a syntax for switching the accuracy of the motion vector derived in the AMVP mode. For example, at amvr_mode = 0, 1, 2, the accuracy is switched between 1 / 4 pixel, 1 pixel, and 4 pixels.
[0088] When the accuracy of the motion vector is set to 1 / 16 accuracy (MVPREC = 16), in order to change the motion vector differences with 1 / 4, 1, and 4 pixel accuracies to motion vector differences with 1 / 16 pixel accuracy, the parameter decoding unit 302 may perform inverse quantization using MvShift (= 1 << amvr_mode) derived from amvr_mode as follows.
[0089] mvdLX[0] = mvdLX[0] << (MvShift + 2) mvdLX[1] = mvdLX[1] << (MvShift + 2) Note that the parameter decoding unit 302 may also derive the mvdLX[] before shifting by the above MvShift by decoding the following syntax. ·abs_mvd_greater0_flag ·abs_mvd_minus2 ·mvd_sign_flag Then, the parameter decoding unit 302 decodes the difference vector lMvd[] from the syntax by using the following formula.
[0090] lMvd[compIdx] = abs_mvd_greater0_flag[compIdx] * (abs_mvd_minus2[compIdx]+2) * (1-2*mvd_sign_flag[compIdx]) Furthermore, the parameter decoding unit 302 sets the decoded differential vector lMvd[] to MvdLX in the case of translational MVD (MotionModelIdc[x][y] == 0), and sets it to MvdCpLX in the case of control point MVD (MotionModelIdc[x][y] != 0).
[0091] if (MotionModelIdc[x][y] == 0) mvdLX[x0][y0][compIdx] = lMvd[compIdx] else mvdCpLX[x0][y0][compIdx] = lMvd[compIdx]<<2 (Motion vector scaling) The method for deriving the motion vector scaling will be described. Assuming the motion vector Mv (reference motion vector), the picture PicMv including the block having Mv, the reference picture PicMvRef of Mv, the scaled motion vector sMv, the picture CurPic including the block having sMv, and the reference picture CurPicRef to which sMv refers, the derivation function MvScale(Mv, PicMv, PicMvRef, CurPic, CurPicRef) of sMv is expressed by the following formula.
[0092] sMv = MvScale(Mv,PicMv,PicMvRef,CurPic,CurPicRef) = Clip3(-R1,R1-1,sign(distScaleFactor*Mv)*((abs(distScaleFactor*Mv)+round1-1)>>shift1)) distScaleFactor = Clip3(-R2,R2-1,(tb*tx+round2)>>shift2) tx = (16384+abs(td)>>1) / td td = DiffPicOrderCnt(PicMv, PicMvRef) tb = DiffPicOrderCnt(CurPic, CurPicRef) Here, round1, round2, shift1, and shift2 are round values and shift values for performing division using reciprocals. For example, round1 = 1 << (shift1 - 1), round2 = 1 << (shift2 - 1), shift1 = 8, shift2 = 6, etc. DiffPicOrderCnt(Pic1, Pic2) is a function that returns the difference in the time information (e.g., POC) between Pic1 and Pic2. R1 and R2 are used to limit the value range for performing the process with limited precision. For example, R1 = 32768, R2 = 4096, etc.
[0093] Also, the scaling function MvScale(Mv, PicMv, PicMvRef, CurPic, CurPicRef) may be expressed by the following formula.
[0094] MvScale(Mv, PicMv, PicMvRef, CurPic, CurPicRef) = Mv * DiffPicOrderCnt(CurPic, CurPicRef) / DiffPicOrderCnt(PicMv, PicMvRef) That is, Mv may be scaled according to the ratio of the difference in the time information between CurPic and CurPicRef to the difference in the time information between PicMv and PicMvRef.
[0095] (DMVR Unit 30375) Subsequently, the DMVR (Decoder side Motion Vector Refinement) process performed by the DMVR unit 30375 will be described. The DMVR process is a process of correcting the motion vectors mvL0 and mvL1 using two reference pictures.
[0096] FIG. 10 is a schematic diagram showing the configuration of the DMVR unit 30375. With reference to FIG. 10, the content of the processing performed by the specific DMVR unit 30375 will be described. The DMVR unit 30375 includes a prediction image generation unit 303751 for modified motion vector search, an initial error generation unit 303752, a motion vector search unit 303753, and a modified vector derivation unit 303754.
[0097] The DMVR unit 30375 refers to · The upper left position (xCb, yCb) of the target block · The width bW of the target block · The height bH of the target block · Motion vectors mvL0 and mvL1 with 1 / 16 pixel accuracy · Reference pictures refPicL0L and refPicL1L to derive the amounts of change dmvL0 and dmvL1 of the motion vectors for modifying mvL0 and mvL1, and output them to the inter prediction image generation unit 309.
[0098] First, the prediction image generation unit 303751 for modified motion vector search · The upper left position (xSb, ySb) of the target sub-block · The width sbW of the target sub-block of luminance · The height sbH of the target sub-block of luminance · Motion vector mvLX (X = 0, 1) · Reference picture refPicLXL (X = 0, 1) to derive a prediction image predSamplesLXL having a size of (sbW)*(sbH).
[0099] In the prediction image generation unit 303751 for modified motion vector search, the motion vector MvLsX (X = 0, 1) is derived by the following formula.
[0100] MvLsX[0] = MvLX[0] - 32 MvLsX[1] = MvLX[1] - 32 Further, the DMVR unit 30375 sets the values of the variables srRange, offsetH[0], offsetV[0], offsetH[1], and offsetV[1] to 2 respectively.
[0101] Let the position of the pixel in the reference block in integer pixel units corresponding to the pixel position (xL, yL) in the target block be (xIntL, yIntL). Also, let the offset in 1 / 16 pixel units from (xIntL, yIntL) be (xFracL, yFracL). These coordinates are derived from the integer components (mvLX[0]>>4, mvLX[1]>>4) and the fractional components (mvLX[0]&15, mvLX[1]&15) of the motion vector (mvLX[0], mvLX[1]), and indicate the position of the pixel with fractional precision within the reference picture refPicLXL. For the pixel whose position in predSamplesLXL is (xL, yL) (xL = 0, ..., sbW-1, yL = 0, ..., sbH-1), the DMVR unit 30375 derives xIntL, yIntL, xFracL, and yFracL using the following equations.
[0102] xIntL = xSb + (mvLX[0]>>4) + xL yIntL = ySb + (mvLX[1]>>4) + yL xFracL = mvLX[0]&15 yFracL = mvLX[1]&15 Subsequently, the DMVR unit 30375 ·(xIntL, yIntL) ·(xFracL, yFracL) ·refPicLXL to derive predSamplesLXL.
[0103] First, the prediction image generation unit 303751 for modified motion vector search derives the variables shift1, shift2, shift3, and shift4 using the following equations.
[0104] shift1 = BitDepthY - 6 offset1 = 1 << (shift1 - 1) shift2 = 4 offset2 = 8 shift3 = 10 - BitDepthY offset3 = 1 << (shift3 - 1) shift4 = BitDepthY - 10 Note that in the above formula, BitDepthY is the number of pixel bits.
[0105] Next, the prediction image generation unit 303751 for modified motion vector search sets picW equal to the value of the picture width pic_width_in_luma_samples. Also, the prediction image generation unit 303751 for modified motion vector search sets picH equal to the value of the picture height pic_height_in_luma_samples.
[0106] After that, in the prediction image generation unit 303751 for modified motion vector search, predSamplesLXL is derived as follows. In the following description, fb L [p] represents the filter coefficient for deriving the pixel value at 1 / 16 pixel accuracy. fb L The value of fb L [p] depends on the position p (p = 1, 2,..., 15) at 1 / 16 pixel accuracy. The position p is equal to xFracL or yFracL. As the value of p increases, fb L [p][0] monotonically decreases, and fb
[0107] First, the prediction image generation unit 303751 for modified motion vector search determines whether xFracL and yFracL are both 0. If both xFracL and yFracL are 0, the DMVR unit 30375 derives predSamplesLXL according to one of the following formulas depending on the value of BitDepthY.
[0108] predSamplesLXL = (BitDepthY <= 10)? (refPicLXL[xIntL][yIntL] << shift3) : ((refPicLXL[xIntL][yIntL]+offset3) >> shift4) When xFracL is not 0 and yFracL is 0, the prediction image generation unit 303751 for the modified motion vector search derives predSamplesLXL according to the following formula.
[0109] predSamplesLXL = (fb L [xFracL][0] * refPicLXL[Clip3(0,picW-1,xIntL)][yIntL] + fb L [xFracL][1] * refPicLXL[Clip3(0,picW-1,xIntL+1)][yIntL] + offset1)>>shift1 When xFracL is 0 and yFracL is not 0, the DMVR unit 30375 derives predSamplesLXL according to the following formula.
[0110] predSamplesLXL = (fb L [yFracL][0] * refPicLXL[xIntL][Clip3(0,picH-1,yIntL)] + fb L [yFracL][1] * refPicLXL[xIntL][Clip3(0,picH-1,yIntL+1)]+offset1)>>shift1 When both xFracL and yFracL are not 0, the prediction image generation unit 303751 for the modified motion vector search derives predSamplesLXL as follows. First, the DMVR unit 30375 derives temp[n] according to the following formula. The derivation process of temp[] is performed n times by changing the reference position. n = 0 represents the first derivation process, and n = 1 represents the second derivation process.
[0111] yPosL = Clip3(0, PicH-1, yIntL+n-3) temp[n] = (fb L [xFracL][0] * refPicLXL[Clip3(0,picW-1,xIntL)][yPosL] + fb L [xFracL][1] * refPicLXL[Clip3(0,picW-1,xIntL+1)][yPosL]+offset1)>>shift1 After that, the DMVR unit 30375 derives predSamplesLXL according to the following formula.
[0112] predSamplesLXL = (fb L [yFracL][0] * temp[0] + fb L [yFracL][1] * temp[1])>>shift2 Next, the initial error generation unit 303752 · the width nCbW of the target block · the height nCbH of the target block · two predicted images predSampleL1 and predSampleL2 having a size of (nCbW + 4) x (nCbH + 4), and variables offsetH[0], offsetH[1], offsetV[0], and offsetV[1] to derive a list Sad1 of the sum of absolute differences of the pixel values included in predSampleL1 and predSampleL2 and a variable centerSad.
[0113] The DMVR unit 30375 sets the value of each element of the 2×9 array bC as follows.
[0114] bC[0][0] = -1 bC[1][0] = -1 bC[0][1] = -1 bC[1][1] = 0 bC[0][2] = -1 bC[1][2] = 1 bC[0][3] = 0 bC[1][3] = -1 bC[0][4] = 0 bC[1][4] = 0 bC[0][5] = 0, bC[1][5] = 1 bC[0][6] = 1, bC[1][6] = -1 bC[0][7] = 1, bC[1][7] = 0 bC[0][8] = 1, bC[1][8] = 1 The initial error generation unit 303752 derives the elements sadList[i] (i = 0,..., 8) of Sad1 according to the following formula.
[0115]
Equation
[0116] Furthermore, the initial error generation unit 303752 derives centerSad according to the following formula.
[0117]
Equation
[0118] The initial error generation unit 303752 determines whether centerSad is greater than or equal to (bH>>1)*(bW)*4. dmvrFlag is a flag indicating whether to perform DMVR processing when it is TRUE and not to perform DMVR processing when it is FALSE. When centerSad is smaller than (bH>>1)*(bW)*4, since the error is small, the initial error generation unit 303752 determines that there is no need to perform DMVR processing, sets dmvrFlag to FALSE, and proceeds to the inter-prediction image generation unit 309 without modifying the motion vectors.
[0119] When centerSad is greater than or equal to (bH>>1)*(bW)*4, the initial error generation unit 303752 sets dmvrFlag to TRUE, and the motion vector search unit 303753 · The number of search points n · The list sadList of the absolute difference sums of the search points, which is an element of Sad1 to derive the index bestIdx with reference to. n is a positive integer.
[0120] Hereinafter, the case where n = 9 will be described. Note that the value of the search point number n may be other than 9, and in addition to the method described in this embodiment, for example, the minimum value of the sadList value may be simply selected when n = 25.
[0121] The motion vector search unit 303753 determines whether sadList[1] < sadList[7] and whether sadList[3] < sadList[5].
[0122] When sadList[1] < sadList[7] and sadList[3] < sadList[5], the DMVR unit 30375 sets the value of idx to 0. Then, the motion vector search unit 303753 determines whether sadList[1] < sadList[3]. The motion vector search unit 303753 sets the value of bestIdx to 1 when sadList[1] < sadList[3], and sets it to 3 when sadList[1] < sadList[3] is not satisfied.
[0123] Otherwise, when sadList[1] >= sadList[7] and sadList[3] < sadList[5], the motion vector search unit 303753 sets the value of idx to 6. Then, the DMVR unit 30375 determines whether sadList[7] < sadList[3]. The motion vector search unit 303753 sets the value of bestIdx to 7 when sadList[7] < sadList[3], and sets it to 3 when sadList[7] < sadList[3] is not satisfied.
[0124] Otherwise, if sadList[1] < sadList[7] and sadList[3] >= sadList[5], the motion vector search unit 303753 sets the value of idx to 2. Thereafter, the DMVR unit 30375 determines whether sadList[1] < sadList[5]. The DMVR unit 30375 sets the value of bestIdx to 1 if sadList[1] < sadList[5], and sets the value of bestIdx to 5 if sadList[1] < sadList[5] is not true.
[0125] Otherwise, if sadList[1] >= sadList[7] and sadList[3] >= sadList[5], the motion vector search unit 303753 sets the value of idx to 8. Thereafter, the DMVR unit 30375 determines whether sadList[7] < sadList[5]. The DMVR unit 30375 sets the value of bestIdx to 7 if sadList[7] < sadList[5], and sets the value of bestIdx to 5 if sadList[7] < sadList[5] is not true.
[0126] Furthermore, the motion vector search unit 303753 determines whether sadList[4] <= sadList[bestIdx]. If sadList[4] <= sadList[bestIdx], the DMVR unit 30375 updates the value of bestIdx to 4. On the other hand, if sadList[4] <= sadList[bestIdx] is not true, the DMVR unit 30375 does not update the value of bestIdx.
[0127] Furthermore, the motion vector search unit 303753 determines whether sadList[idx] < sadList[bestIdx]. If sadList[idx] < sadList[bestIdx], the DMVR unit 30375 sets bestIdx to idx. On the other hand, if sadList[idx] < sadList[bestIdx] is not true, the motion vector search unit 303753 does not update the value of bestIdx.
[0128] The motion vector search unit 303753 determines whether the value of bestIdx is 4. If the value of bestIdx is 4, the motion vector search unit 303753 sets halfPelAppliedflag to true.
[0129] If the value of bestIdx is not 4, the motion vector search unit 303753 calculates the values of variables dmvx and dmvy according to the following formula.
[0130] dmvx = (bestIdx / 3 - 1) dmvy = (bestIdx%3 - 1) Furthermore, the motion vector search unit 303753 updates offsetH and offsetV according to the following formula.
[0131] offsetH[0] = offsetH[0] + dmvx, offsetV[0] = offsetV[0] + dmvy offsetH[1] = offsetH[1] - dmvx, offsetV[1] = offsetV[1] - dmvy The motion vector search unit 303753 derives Sad2 through a process similar to the process of deriving Sad1 described above using the updated offsetH and offsetV. Furthermore, the motion vector search unit 303753 derives bestIdx again using Sad2 instead of Sad1.
[0132] The motion vector search unit 303753 determines whether the value of the re-derived bestIdx is 4. If the value of bestIdx is 4, the motion vector search unit 303753 sets halfPelAppliedflag to true.
[0133] If the value of bestIdx is not 4, the motion vector search unit 303753 calculates dmvx and dmvy according to the following formula.
[0134] dmvx = (bestIdx / 3 - 1), dmvy = (bestIdx%3 - 1) Furthermore, the DMVR unit 30375 calculates dmvL0 and dmvL1 according to the following formula.
[0135] dmvL0[0] = 16*dmvx, dmvL0[1] = 16*dmvy dmvL1[0] = -16*dmvx, dmvL1[1] = -16*dmvy If halfPelAppliedflag is true, the motion vector search unit 303753 derives the corrected dmvL0 and dmvL1 as follows. The following sadList is an element of Sad2 if Sad2 exists, and an element of Sad1 if Sad2 does not exist.
[0136] First, the motion vector search unit 303753 determines whether sadList[1] + sadList[7] == sadList[4]. If sadList[1] + sadList[7] == sadList[4], and mrSadT + mrSadB - (mrSadC<<1) == 0, the motion vector search unit 303753 sets dmv[0]=0. If sadList[1] + sadList[7]!= sadList[4], the motion vector search unit 303753 calculates dmv[0] according to the following formula.
[0137] dmv[0] = ((sadList[1] - sadList[7])<<3) / (sadList[1] + sadList[7] - (sadList[4]<<1)) Next, the correction vector derivation unit 303754 determines whether sadList[3] + sadList[5] == sadList[4]. If sadList[3] + sadList[5] == sadList[4] and mrSadL + mrSadR - (mrSadC<<1) == 0, the correction vector derivation unit 303754 sets dmv[1]=0. If sadList[3] + sadList[5] != sadList[4], the correction vector derivation unit 303754 calculates dmv[1] using the following formula.
[0138] dmv[1] = ((sadList[3] - sadList[5])<<3) / (sadList[3] + sadList[5] - (sadList[4]<<1)) Furthermore, the correction vector derivation unit 303754 corrects the motion vectors mvL0 and mvL1 using the following formula.
[0139] dmvL0[0] = dmvL0[0] + dmv[0] dmvL0[1] = dmvL0[1] + dmv[1] dmvL1[0] = dmvL1[0] - dmv[0] dmvL1[1] = dmvL1[1] - dmv[1] The DMVR unit 30375 calculates the motion vector mvLX by adding the differential vector dmvLX derived from the prediction vector mvpLX input from the merge prediction unit 30374. The DMVR unit 30375 outputs mvLX to the inter-prediction image generation unit 309.
[0140] mvLX[0] = mvpLX[0]+dmvLX[0] mvLX[1] = mvpLX[1]+dmvLX[1] Note that the values of dmvLX[0] and dmvLX[1] are restricted to be between -8 and 8 regardless of the number of bits of sadList.
[0141] (DMVR Judgment Criteria) dmvrFlag is a flag indicating that DMVR processing is performed when it is TRUE and DMVR processing is not performed when it is FALSE.
[0142] When the flag of SPS, which indicates that DMVR processing is possible, is On, the initial error generation unit 303752 sets dmvrFlag to TRUE. Otherwise, the initial error generation unit 303752 sets dmvrFlag to FALSE.
[0143] Also, when the merge_flag of the block is TRUE, the initial error generation unit 303752 may set dmvrFlag to TRUE. Otherwise, the initial error generation unit 303752 sets dmvrFlag to FALSE.
[0144] Also, when both predFlagL0 and predFlagL1 are TRUE, that is, in the case of bidirectional prediction, the initial error generation unit 303752 may set dmvrFlag to TRUE. Otherwise, the initial error generation unit 303752 sets dmvrFlag to FALSE.
[0145] When the mmvd_flag of the block is FALSE, that is, in the non-MMVD mode, the initial error generation unit 303752 may set dmvrFlag to TRUE. Otherwise, in the MMVD mode, the initial error generation unit 303752 sets dmvrFlag to FALSE.
[0146] If DiffPicOrderCnt( currPic, RefPicList
[0000] [ refIdxL0 ]) is equal to DiffPicOrderCnt( RefPicList
[0001] [ refIdxL1 ], currPic ), that is, when the current picture currPic is in a positional relationship such that the L0 reference picture RefPicList
[0000] [ refIdxL0 ] and the L1 reference picture RefPicList
[0001] [ refIdxL1 ] are interpolated at equal distances, the initial error generation unit 303752 may set dmvrFlag to TRUE. Otherwise, the initial error generation unit 303752 sets dmvrFlag to FALSE. Here, DiffPicOrderCnt() is a function that derives the difference in POC (Picture Order Count) of two images as follows.
[0147] DiffPicOrderCnt(picA,picB) = PicOrderCnt(picA)-PicOrderCnt(picB) In addition, when DiffPicOrderCnt( currPic, RefPicList
[0000] [ refIdxL0 ])*DiffPicOrderCnt( currPic, RefPicList
[0001] [ refIdxL1 ])<0, that is, when they are simply in a positional relationship for interpolation, the initial error generation unit 303752 may set dmvrFlag to TRUE, and otherwise, dmvrFlag may be set to FALSE.
[0148] Also, when the size of the processing block is less than a certain specific value, the initial error generation unit 303752 may set dmvrFlag to FALSE. For example, when bH is 8 or more and bH*bW is 64, the initial error generation unit 303752 may set dmvrFlag to TRUE. Otherwise, the initial error generation unit 303752 sets dmvrFlag to FALSE.
[0149] FIG. 11 is a flowchart showing the processing flow in the DMVR unit 30375. In the present embodiment, in addition to the above determination criteria, as shown in FIG. 11, conditions are added such that the DMVR process is applied only when the GBI process described later is not applied.
[0150] Specifically, first, the DMVR unit 30375 executes the determination process (S1101) of the dmvrFlag described above. Next, the DMVR unit 30375 determines (S1102) whether gbiIdx is 0. As will be described later, when gbiIdex has a non-zero value, non-uniform weighted prediction is performed based on the table gbiWLut. When gbiIdx is 0, in addition to the condition that dmvrFlag is set to TRUE, when gbiIdx is non-zero, dmvrFlag is set to FALSE (S1103).
[0151] Furthermore, the DMVR unit 30375 determines (S1104) whether dmvrFlag is TRUE. If it is TRUE, the DMVR process (S1105) is executed, and if it is FALSE, it is not executed.
[0152] When applying GBI prediction, since weighted prediction is applied, considering that the error cannot be correctly evaluated, the overall processing amount can be reduced by restricting the application conditions.
[0153] Similarly, in the weighted prediction described later, when either the L0 prediction or the L1 prediction that applies the DMVR process performs weighted prediction, dmvrFlag is set to FALSE. Specifically, when both luma_weight_l0_flag[refIdxL0], which indicates whether there is a luminance weight coefficient w0 and an offset o0 in the L0 prediction picture, and luma_weight_l1_flag[refIdxL1], which indicates whether there is a luminance weight coefficient w1 and an offset o1 in the L1 prediction picture, are both FALSE, in addition to the condition that dmvrFlag is set to TRUE, otherwise, dmvrFlag is set to FALSE.
[0154] (Determination of BDOF by Error Threshold Processing in DMVR) In DMVR, a process of obtaining the error between the L0 predicted image and the L1 predicted image is performed. Based on the error value at this time, it is determined in advance whether to execute the BDOF process described later that is performed in the subsequent stage.
[0155] FIG. 12 is a flowchart for explaining the process of determining BDOF by error threshold processing in DMVR.
[0156] First, the initial error generation unit 303752 sets bdofFlag to TRUE in advance (S1201). Next, the initial error generation unit 303752 derives centerSad (S1202) and determines whether the value of centerSad is greater than or equal to the value of the threshold (bH>>1)*bW*4 (S1203). If the value of centerSad is smaller than the threshold, the initial error generation unit 303752 determines that the error is small and sets bdofFlag, which indicates whether to perform the BDOF process, to FALSE (S1204) to prevent the BDOF process from being performed in advance. Since this determination is the same as the determination by the above-described initial error generation unit 30752, the motion vector search unit 303753 and the correction vector derivation unit 303754 are skipped and the DMVR process is not performed either. If the value of centerSad is greater than or equal to the threshold, the motion vector search unit 303753 performs a corrected motion vector search (S1205). As a result, the correction vector derivation unit 303754 determines whether the value of sadList[bestIdx], which is the SAD value of the minimum bestIdx, is smaller than the value of the threshold (bH>>1)*bW*8 (S1206). If the value of sadList[bestIdx] is smaller than the threshold, the correction vector derivation unit 303754 determines that the error is small and sets bdofFlag, which indicates whether to perform the BDOF process, to FALSE (S1207) to prevent the BDOF process from being performed in advance.
[0157] Note that the threshold in (S1206) is the same as or larger than the threshold in (S1203).
[0158] In the DMVR process, in order to search for the modified motion vector, it is necessary to calculate the error between the L0 predicted image and the L1 predicted image. On the other hand, in the BDOF process, when the error is small, there is no effect. Therefore, by adding such a process, it becomes possible to determine whether to perform BDOF without adding additional error calculation.
[0159] In the BDOF section described later, it is determined (S1208) whether bdofFlag is TRUE. If Yes, the inter-prediction parameter decoding section 303 performs the BDOF process (S1209). If No, it is determined that the BDOF process is not performed in the block.
[0160] (Triangle Prediction) Subsequently, the Triangle prediction will be described. In the Triangle prediction, the target CU is divided into two triangle prediction units with the diagonal or anti-diagonal as the boundary. The predicted image in each triangle prediction unit is derived by performing a weighted mask process according to the pixel position on each pixel of the predicted image of the target CU (rectangular block including the triangle prediction unit). For example, by multiplying a mask with 1 for the pixels in the triangle region within the rectangular region and 0 for the regions outside the triangle, the triangle image can be derived from the rectangular image. The adaptive weighting process of the predicted image is applied to both regions across the diagonal, and one predicted image of the target CU (rectangular block) is derived by the adaptive weighting process using the two predicted images. This process is called the Triangle synthesis process. The transformation (inverse transformation) and quantization (inverse quantization) processes are applied to the entire target CU. Note that the Triangle prediction is applied only in the case of the merge prediction mode or the skip mode.
[0161] The Triangle prediction unit 30377 derives prediction parameters corresponding to two triangle regions used for Triangle prediction and outputs them to the inter-prediction image generation unit 309. In Triangle prediction, for simplicity of processing, a configuration that does not use dual prediction may be used. In this case, the inter-prediction parameters for uni-directional prediction are derived in one triangle region. Note that the derivation of two prediction images and the synthesis using the prediction images are performed by the motion compensation unit 3091 and the Triangle synthesis unit 30952.
[0162] (MMVD prediction unit 30376) The MMVD prediction unit 30376 performs processing in the MMVD (Merge with Motion Vector Difference) mode. The MMVD mode is a mode in which a motion vector is obtained by adding a difference vector in a predetermined distance and a predetermined direction to a motion vector derived from merge candidates (a motion vector derived from a motion vector of an adjacent block or the like). In the MMVD mode, the MMVD prediction unit 30376 uses merge candidates and efficiently derives a motion vector by restricting the value range of the difference vector to a predetermined distance (for example, eight values) and a predetermined direction (for example, four directions, eight directions, etc.).
[0163] The loop filter 305 is a filter provided within the encoding loop, which removes block distortion and ringing distortion to improve the image quality. The loop filter 305 applies filters such as a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF) to the decoded image of the CU generated by the addition unit 312.
[0164] The reference picture memory 306 stores the decoded image of the CU generated by the addition unit 312 at a predetermined position for each target picture and each target CU.
[0165] The prediction parameter memory 307 stores prediction parameters at predetermined positions for each CTU or CU to be decoded. Specifically, the prediction parameter memory 307 stores the parameters decoded by the parameter decoding unit 302 and the prediction mode predMode decoded by the entropy decoding unit 301, etc.
[0166] The prediction image generation unit 308 is input with the prediction mode predMode, prediction parameters, etc. Also, the prediction image generation unit 308 reads a reference picture from the reference picture memory 306. The prediction image generation unit 308 generates a prediction image of a block or sub-block using the prediction mode indicated by the prediction mode predMode, the prediction parameters, and the read reference picture (reference picture block). Here, the reference picture block is a set of pixels on the reference picture (usually a rectangle, so it is called a block), and is an area referred to for generating the prediction image.
[0167] (Inter prediction image generation unit 309) When the prediction mode predMode indicates the inter prediction mode, the inter prediction image generation unit 309 generates a prediction image of a block or sub-block by inter prediction using the inter prediction parameters input from the inter prediction parameter decoding unit 303 and the read reference picture.
[0168] FIG. 13 is a schematic diagram showing the configuration of the inter prediction image generation unit 309 included in the prediction image generation unit 308 according to the present embodiment. The inter prediction image generation unit 309 includes a motion compensation unit (prediction image generation device) 3091 and a synthesis unit 3095.
[0169] (Motion compensation) The motion compensation unit 3091 (interpolation image generation unit) generates an interpolation image (motion-compensated image) by reading, from the reference picture memory 306, a block located at a position shifted by the motion vector mvLX from the position of the target block in the reference picture RefPicLX specified by the reference picture index refIdxLX, based on the inter-prediction parameter (prediction list utilization flag predFlagLX, reference picture index refIdxLX, motion vector mvLX) input from the inter-prediction parameter decoding unit 303. Here, when the accuracy of the motion vector mvLX is not integer accuracy, a filter for generating pixels at fractional positions, called a motion compensation filter, is applied to generate the interpolation image.
[0170] The motion compensation unit 3091 first derives the integer position (xInt, yInt) and phase (xFrac, yFrac) corresponding to the coordinates (x, y) within the prediction block using the following equations.
[0171] xInt = xPb+(mvLX[0]>>(log2(MVPREC)))+x xFrac = mvLX[0]&(MVPREC-1) yInt = yPb+(mvLX[1]>>(log2(MVPREC)))+y yFrac = mvLX[1]&(MVPREC-1) Here, (xPb, yPb) is the upper left coordinate of a block of size bW*bH, x = 0…bW-1, y = 0…bH-1, and MVPREC indicates the accuracy of the motion vector mvLX (1 / MVPREC pixel accuracy). For example, MVPREC may be 16.
[0172] The motion compensation unit 3091 performs horizontal interpolation processing on the reference picture refImg using an interpolation filter to derive a temporary image temp[][]. The following Σ is the sum over k from k = 0..NTAP-1, shift1 is a normalization parameter for adjusting the value range, and offset1 = 1<<(shift1-1).
[0173] temp[x][y] = (ΣmcFilter[xFrac][k]*refImg[xInt+k-NTAP / 2+1][yInt]+offset1)>>shift1 Subsequently, the motion compensation unit 3091 derives an interpolated image Pred[][] from the temporary image temp[][] through vertical interpolation processing. The following Σ is the sum over k from k = 0..NTAP - 1, shift2 is a normalization parameter for adjusting the value range, and offset2 = 1<<(shift2 - 1).
[0174] Pred[x][y] = (ΣmcFilter[yFrac][k]*temp[x][y+k-NTAP / 2+1]+offset2)>>shift2 (Synthesis unit) The synthesis unit 3095 generates a predicted image with reference to the interpolated image input from the motion compensation unit 3091, the inter prediction parameters input from the inter prediction parameter decoding unit 303, and the intra image input from the intra prediction image generation unit 310, and outputs the generated predicted image to the addition unit 312.
[0175] The synthesis unit 3095 includes a Combined intra / inter synthesis unit 30951, a Triangle synthesis unit 30952, an OBMC unit 30953, and a BDOF unit 30956.
[0176] (Combined intra / inter synthesis processing) The Combined intra / inter synthesis unit 30951 generates a predicted image by comprehensively using unidirectional prediction, skip mode, merge mode, and intra prediction in AMVP.
[0177] (Triangle synthesis processing) The Triangle synthesis unit 30952 generates a predicted image using the above-described Triangle prediction.
[0178] (OBMC processing) The OBMC unit 30953 generates a predicted image using OBMC (Overlapped block motion compensation) processing. The OBMC processing includes the following processes. · Using the interpolation image (PU interpolation image) generated using the inter-prediction parameter added to the target sub-block and the interpolation image (OBMC interpolation image) generated using the motion parameter of the adjacent sub-block of the target sub-block, generate the interpolation image (motion compensation image) of the target sub-block. · Generate a predicted image by weighted averaging the OBMC interpolation image and the PU interpolation image.
[0179] (Weighted Prediction Unit 30954) The weighted prediction unit 309454 generates a predicted image of a block by multiplying the motion compensation images PredL0 and PredL1 by weight coefficients. When one of the prediction list utilization flags (predFlagL0 or predFlagL1) is 1 (single prediction) and weighted prediction is not used, perform the following processing on the motion compensation image PredLX (LX is L0 or L1) to match the pixel bit depth bitDepth.
[0180] Pred[x][y] = Clip3(0,(1<<bitDepth)-1,(PredLX[x][y]+offset1)>>shift1) Here, shift1 = Max(2,14 - bitDepth) and offset1 = 1<<(shift1 - 1).
[0181] (Bidirectional Prediction Processing) Also, when both of the prediction list utilization flags (predFlagL0 and predFlagL1) are 1 (bidirectional prediction BiPred) and weighted prediction is not used, average the motion compensation images PredL0 and PredL1 and perform the following processing to match the pixel bit depth.
[0182] Pred[x][y] = Clip3(0,(1<<bitDepth)-1,(PredL0[x][y]+PredL1[x][y]+offset2)>>shift2) Here, shift2 = Max(3, 15 - bitDepth), offset2 = 1<<(shift2 - 1). Hereinafter, this process is also called normal bidirectional prediction.
[0183] Furthermore, when a flag (luma_weight_l0_flag for luminance and chroma_weight_l0_flag for chrominance) indicating whether there are a weight prediction coefficient w0 and an offset o0 in the reference picture of L0 in single prediction is on, the weighted prediction unit 30954 derives the weight prediction coefficient w0 and the offset o0 from the encoded data in the case of L0 prediction and performs the following formula processing.
[0184] Pred[x][y] = Clip3(0,(1<<bitDepth)-1,((PredL0[x][y]*w0+(1<<(log2WD - 1)))>>log2WD)+o0) In the case of L1 prediction, when a flag (luma_weight_l1_flag for luminance and chroma_weight_l1_flag for chrominance) indicating whether there are a weight prediction coefficient w1 and an offset o1 in the reference picture of L1 is on, the weighted prediction unit 30954 derives the weight prediction coefficient w1 and the offset o1 from the encoded data and performs the following formula processing.
[0185] Pred[x][y] = Clip3(0,(1<<bitDepth)-1,((PredL1[x][y]*w1+(1<<(log2WD - 1)))>>log2WD)+o1) Here, log2WD is a variable obtained by summing up the values of Log2WeightDenom + shift1 explicitly sent in the slice header for luminance and chrominance separately.
[0186] (Weighted bidirectional prediction process) Furthermore, when double prediction BiPred is used and there are flags (luma_weight_l0_flag, luma_weight_l1_flag for luminance, chroma_weight_l0_flag, chroma_weight_l1_flag for color difference) indicating the presence or absence of weight prediction coefficients and offsets, if weight prediction is performed, the weighted prediction unit 30954 derives the weight prediction coefficients w0, w1, o0, o1 from the encoded data and performs the processing of the following formula.
[0187] Pred[x][y] = Clip3(0,(1<<bitDepth)-1,(PredL0[x][y]*w0+PredL1[x][y]*w1+((o0+o1+1)<<log2WD))>>(log2WD+1)) (GBI unit 30955) In the above weighted prediction, an example of generating a predicted image by multiplying the interpolation image by weight coefficients has been described. Here, another example of generating a predicted image by multiplying the interpolation image by weight coefficients will be described. Specifically, the process of generating a predicted image using generalized bi-prediction (hereinafter referred to as GBI prediction) will be described. In GBI prediction, the predicted image Pred is generated by multiplying the L0 predicted image PredL0 and the L1 predicted image PredL1 in double prediction by weight coefficients (w0, w1).
[0188] Also, when generating a predicted image using GBI prediction, the GBI unit 30955 switches the weight coefficients (w0, w1) in units of encoding units. That is, the GBI unit 30954 of the inter prediction image generation unit 309 sets the weight coefficients for each encoding unit. In GBI prediction, a plurality of weight coefficient candidates are defined in advance, and gbiIdx is an index indicating the weight coefficient to be used for the target block among the plurality of weight coefficient candidates included in the table gbiWLut.
[0189] The GBI unit 30955 checks the flag gbiAppliedFlag indicating whether GBI prediction is used. If it is FALSE, the motion compensation unit 3091 generates a predicted image using the following formula.
[0190] Pred[x][y]=Clip3(0,(1<<bitDepth)-1, (PredL0[x][y]+ PredL1[x][y]+offset2)>>shift2 ) Here, the initial state of gbiAppliedFlag is FALSE. The GBI unit 30955 sets gbiAppliedFlag to TRUE when the flag indicating that GBI processing is possible in the SPS is On and it is bidirectional prediction. Further, as an additional (AND) condition, gbiAppliedFlag may be set to TRUE when gbiIdx, which is the index of the GBI prediction weight coefficient table gbiWLut, is not 0 (the index value when the weights of the L0 prediction image and the L1 prediction image are equal). Further, as an additional (AND) condition, gbiAppliedFlag may be set to TRUE when the block size of the CU is a certain value or more.
[0191] When gbiAppliedFlag is true, the GBI unit 30955 derives the predicted image Pred from the weights w0, w1 and PredL0, PredL1 by the following formula.
[0192] Pred[x][y]=Clip3(0,(1<<bitDepth)-1, (w0*PredL0[x][y]+w1*PredL1[x][y]+offset3)>>(shift2+3)) Here, the weight coefficient w1 is a coefficient derived from the table iWLut[] = {4, 5, 3, 10, -2} by gbiIdx explicitly shown in the syntax. The weight coefficient w0 is set to (8 - w1). When gbiIdx = 0, w0 = w1 = 4, which is equivalent to normal bidirectional prediction.
[0193] shift1, shift2, offset1, offset2 are derived by the following formula.
[0194] shift1=Max(2,14-bitDepth) shift2 = Max(3, 15 - bitDepth) = shift1 + 1 offset1 = 1 << (shift1 - 1) offset2 = 1 << (shift2 - 1) offset3 = 1 << (shift2 + 2) Note that there are multiple tables gbiWLut with different combinations of weight coefficients, and the GBI unit 30955 may switch the table used for selecting the weight coefficients according to whether the picture structure is LowDelay (LB).
[0195] When GBI prediction is used in the AMVP prediction mode, the inter - prediction parameter decoding unit 303 decodes gbiIdx and sends it to the GBI unit 30955. Also, when GBI prediction is used in the merge prediction mode, the inter - prediction parameter decoding unit 303 decodes the merge index merge_idx, and the merge candidate derivation unit 30361 derives the gbiIdx of each merge candidate. Specifically, the merge candidate derivation unit 30361 uses the weight coefficient of the adjacent block used for deriving the merge candidate as the weight coefficient of the merge candidate to be used for the target block. That is, in the merge mode, the weight coefficient used in the past is inherited as the weight coefficient of the target block.
[0196] (Selection of Prediction Mode Using GBI Prediction) Next, with reference to FIG. 14, the selection process of the prediction mode using GBI prediction in the moving image decoding apparatus 31 will be described. FIG. 14 is a flowchart showing an example of the flow of the prediction mode selection process in the moving image decoding apparatus 31.
[0197] As shown in FIG. 14, the inter-prediction parameter decoder 303 first decodes the skip flag (S1401). When the skip flag indicates the skip mode (YES in S1402), the prediction mode becomes the merge mode (S1403), and the inter-prediction parameter decoder 303 decodes the merge index (S14031). When GBI prediction is used, the GBI unit 30955 derives the weight coefficient derived from the merge candidate as the weight coefficient of the GBI prediction.
[0198] When the skip flag does not indicate the skip mode (NO in S1402), the inter-prediction parameter decoder 303 decodes the merge flag (S1407). When the merge flag indicates the merge mode (YES in S1408), the prediction mode becomes the merge mode (S1403), and the inter-prediction parameter decoder 303 decodes the merge index (S14031). When GBI prediction is used, the GBI unit 30955 derives the weight coefficient derived from the merge candidate as the weight coefficient of the GBI prediction.
[0199] When the merge flag does not indicate the merge mode (NO in S1408), the prediction mode is the AMVP mode (S1409).
[0200] In the AMVP mode, the inter-prediction parameter decoder 303 decodes the inter-prediction identifier inter_pred_idc (S14090). Subsequently, the inter-prediction parameter decoder 303 decodes the differential vector mvdLX (S14091). Subsequently, the inter-prediction parameter decoder 303 decodes gbiIdx (S14092), and when GBI prediction is used, the GBI unit 30955 selects the weight coefficient w1 of the GBI prediction from the weight coefficient candidates in the gbiWLut table.
[0201] (BDOF Prediction) Next, the details of the prediction using the BDOF process performed by the BDOF unit 30956 (BDOF prediction) will be described. The BDOF unit 30956 generates a predicted image with reference to two predicted images (the first predicted image and the second predicted image) and a gradient correction term in the bi-prediction mode.
[0202] FIG. 15 is a flowchart for explaining the flow of the process for deriving a predicted image.
[0203] When the inter-prediction parameter decoding unit 303 determines that it is a unidirectional prediction of L0 (inter_pred_idc is 0 in S1501), the motion compensation unit 3091 generates an L0 predicted image PredL0[x][y] (S1502). When the inter-prediction parameter decoding unit 303 determines that it is a unidirectional prediction of L1 (inter_pred_idc is 1 in S1501), the motion compensation unit 3091 generates an L1 predicted image PredL1[x][y] (S1503). On the other hand, when the inter-prediction parameter decoding unit 303 determines that it is in the bi-prediction mode (inter_pred_idc is 2 in S1501), the following processing of S1504 follows. In S1504, the synthesis unit 3095 refers to the bioAvailableFlag indicating whether to perform the BDOF process and determines the necessity of the BDOF process. When the bioAvailableFlag indicates TRUE, the BDOF unit 30956 executes the BDOF process to generate a bi-directional predicted image (S1506). When the bioAvailableFlag indicates FALSE, the synthesis unit 3095 generates a predicted image by normal bi-directional predicted image generation (S1505).
[0204] The inter-prediction parameter decoding unit 303 may derive TRUE for the bioAvailableFlag when the L0 reference image refImgL0 and the L1 reference image refImgL1 are different reference images and two pictures are in opposite directions with respect to the target picture. Specifically, assuming the target image is currPic, when the condition DiffPicOrderCnt(currPic, refImgL0)*DiffPicOrderCnt(currPic, refImgL1)<0 is satisfied, the bioAvailableFlag indicates TRUE. Here, DiffPicOrderCnt() is a function that derives the difference in POC (Picture Order Count: display order of pictures) between two images as follows.
[0205] DiffPicOrderCnt(picA,picB) = PicOrderCnt(picA)-PicOrderCnt(picB) As a condition for bioAvailableFlag to indicate TRUE, a condition that the motion vector of the target block is not a motion vector in sub-block units may be added.
[0206] Also, as a condition for bioAvailableFlag to indicate TRUE, a condition that the motion vector of the target picture is not a motion vector in sub-block units may be added.
[0207] Also, as a condition for bioAvailableFlag to indicate TRUE, a condition that the sum of absolute differences between the L0 predicted image and the L1 predicted image of two prediction blocks is equal to or greater than a predetermined value may be added.
[0208] Also, as a condition for bioAvailableFlag to indicate TRUE, a condition that the prediction image creation mode is a prediction image creation mode in block units may be added.
[0209] Also, as a condition for bioAvailableFlag to indicate TRUE, in weighted prediction, a condition that neither L0 prediction nor L1 prediction performs weighted prediction may be added. Specifically, when both luma_weight_l0_flag[ refIdxL0 ], which indicates whether there is a luminance weight coefficient w0 and an offset o0 in the L0 predicted picture, and luma_weight_l1_flag[ refIdxL1 ], which indicates whether there is a luminance weight coefficient w1 and an offset o1 in the L1 predicted picture, are FALSE, it is a condition for bioAvailableFlag to indicate TRUE.
[0210] Figure 16 is a schematic diagram showing the configuration of the BDOF unit 30956. Using Figure 16, the content of the processing performed by the specific BDOF unit 30956 will be described. The BDOF processing unit 30956 includes an L0, L1 prediction image generation unit 309561, a gradient image generation unit 309562, a correlation parameter calculation unit 309563, a motion compensation correction value derivation unit 309564, and a BDOF prediction image generation unit 309565. The BDOF unit 30956 generates a prediction image from the interpolated image received from the motion compensation unit 3091 and the inter prediction parameter received from the inter prediction parameter decoding unit 303, and outputs the generated prediction image to the addition unit 312. Note that the process of deriving the motion compensation correction value modBIO (motion compensation correction image) from the gradient image and correcting and deriving the prediction images of PredL0 and PredL1 is called the bidirectional gradient change process.
[0211] Figure 17 is a diagram showing an example of an area where padding is executed. First, the L0, L1 prediction image generation unit 309561 generates L0 and L1 prediction images used for BDOF processing. In the BDOF unit 30956, BDOF processing is performed based on the L0 and L1 prediction images for each CU unit or sub-CU unit shown in Figure 17. However, in order to obtain the gradient, additional interpolation image information for two pixels around the target CU or sub-CU is required. This interpolation image information for this part is generated using adjacent integer pixels instead of a normal interpolation filter for the purpose of generating the gradient image described later. In other cases, this part is used as a padding area, and the surrounding pixels are copied and used in the same way as outside the picture. Also, the unit of BDOF processing is a CU unit or NxN pixels of sub-CU unit or less, and the processing itself is performed using (N + 2) x (N + 2) pixels with one pixel added to the surroundings.
[0212] The gradient image generation unit 309562 generates a gradient image. In the gradient change (Optical Flow), it is assumed that the pixel value of each point does not change and only its position changes. This can be expressed as follows using the change in the pixel value I in the horizontal direction (horizontal gradient value lx) and its position change Vx, the change in the pixel value I in the vertical direction (vertical gradient value ly) and its position change Vy, and the temporal change lt of the pixel value I.
[0213] lx * Vx + ly * Vy + lt = 0 Hereinafter, the change in position (Vx, Vy) is referred to as the correction weight vector (u, v).
[0214] Specifically, the gradient image generation unit 309562 derives gradient images lx0, ly0, lx1, and ly1 from the following equations. lx0 and lx1 indicate gradients along the horizontal direction, and ly0 and ly1 indicate gradients along the vertical direction.
[0215] lx0[x][y] = (PredL0[x+1][y]-PredL0[x-1][y])>>shift1 ly0[x][y] = (PredL0[x][y+1]-PredL0[x][y-1])>>shift1 lx1[x][y] = (PredL1[x+1][y]-PredL1[x-1][y])>>shift1 ly1[x][y] = (PredL1[x][y+1]-PredL1[x][y-1])>>shift1 Here, shift1 = Max(2, 14 - bitDepth).
[0216] Next, the correlation parameter calculation unit 309563 derives gradient products s1, s2, s3, s5, and s6 of (N + 2) x (N + 2) pixels using the surrounding 1 pixel for each block of N x N pixels within each CU.
[0217] s1 = sum(phiX[x][y]* phiX[x][y]) s2 = sum(phiX[x][y]* phiY[x][y]) s3 = sum(-theta[x][y]* phiX[x][y]) s5 = sum(phiY[x][y]* phiY[x][y]) s6 = sum(-theta[x][y]* phiY[x][y]) Here, sum(a) represents the sum of a with respect to the coordinates (x, y) within a block of (N + 2) x (N + 2) pixels. Also, theta[x][y]= -(PredL1[x][y]>>shift4)+(PredL0[x][y]>>shift4) phiX[x][y] = (lx1[x][y] + lx0[x][y])>>shift5 phiY[x][y] = (ly1[x][y] + ly0[x][y])>>shift5 Here, shift4=Min(8,bitDepth-4) shift5=Min(5,bitDepth-7) is set as follows.
[0218] Next, the motion compensation correction value derivation unit 309564 derives a correction weight vector (u, v) in units of N x N pixels using the derived sum of gradient products s1, s2, s3, s5, s6.
[0219] u = (s3<<3)>>log2(s1) v = ((s6<<3)-((((u*s2m)<<12)+u*s2s)>>1))>>log2(s5) Here, s2m = s2>>12 and s2s = s2&((1<<12)-1).
[0220] Note that the ranges of u and v may be further restricted using a clip as follows.
[0221] u = s1>0?Clip3(-th,th,-(s3<<3)>>floor(log2(s1))):0 v = s5>0?Clip3(-th,th,((s6<<3)-((((u*s2m)<<12)+u*s2s)>>1))>>floor(log2(s5))):0 Here, let th = Max(2, 1<<(13 - bitDepth)). Since the value of th needs to be calculated in conjunction with shift1, consider the case where the pixel bit length bitDepth is greater than 12 bits.
[0222] The motion compensation correction value derivation unit 309564 derives the modBIO[x][y] of the motion compensation correction value for NxN pixels using the correction weight vector (u, v) in units of NxN pixels and the gradient images lx0, ly0, lx1, and ly1.
[0223] modBIO[x][y] = ((lx1[x][y] - lx0[x][y]) * u + (ly1[x][y] - ly0[x][y]) * v + 1) >> 1 (Equation A3) Alternatively, modBIO may be derived as follows using a rounding function.
[0224] modBIO[x][y] = Round(((lx1[x][y] - lx0[x][y]) * u) >> 1) + Round(((ly1[x][y] - ly0[x][y]) * v) >> 1) The BDOF predicted image generation unit 309565 derives the pixel value Pred of the predicted image for NxN pixels according to the following formula using the above parameters.
[0225] At this time, the BDOF predicted image generation unit 309565 derives the pixel value Pred of the predicted image for NxN pixels according to the following formula using the above parameters.
[0226] Pred[x][y] = Clip3(0, (1 << bitDepth) - 1, (PredL0[x][y] + PredL1[x][y] + modBIO[x][y] + offset2) >> shift2) Here, shift2 = Max(3, 15 - bitDepth) and offset2 = 1 << (shift2 - 1).
[0227] Then, the BDOF predicted image generation unit 309565 outputs the generated predicted image of the block to the addition unit 312.
[0228] The inverse quantization and inverse transformation unit 311 inverse quantizes the quantization conversion coefficients input from the entropy decoding unit 301 to obtain conversion coefficients. These quantization conversion coefficients are coefficients obtained by performing frequency conversion such as DCT (Discrete Cosine Transform) and DST (Discrete Sine Transform) on the prediction error and then quantizing it in the encoding process. The inverse quantization and inverse transformation unit 311 performs inverse frequency conversion such as inverse DCT and inverse DST on the obtained conversion coefficients to calculate the prediction error. The inverse quantization and inverse transformation unit 311 outputs the prediction error to the addition unit 312. The inverse quantization and inverse transformation unit 311 sets all prediction errors to 0 when the skip_flag is 1 or when the cu_cbp is 0.
[0229] The addition unit 312 adds the predicted image of the block input from the predicted image generation unit 308 and the prediction error input from the inverse quantization and inverse transformation unit 311 for each pixel to generate the decoded image of the block. The addition unit 312 stores the decoded image of the block in the reference picture memory 306 and also outputs it to the loop filter 305.
[0230] (Configuration of the moving image encoding device) Next, the configuration of the moving image encoding device 11 according to the present embodiment will be described. FIG. 18 is a schematic diagram showing the configuration of the moving image encoding device 11 according to the present embodiment. The moving image encoding device 11 includes a predicted image generation unit 101, a subtraction unit 102, a conversion and quantization unit 103, an inverse quantization and inverse transformation unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, an encoding parameter determination unit 110, a parameter encoding unit 111, and an entropy encoding unit 104.
[0231] The predicted image generation unit 101 generates a predicted image for each CU, which is a region obtained by dividing each picture of the image T. The predicted image generation unit 101 performs the same operation as the predicted image generation unit 308 described above, and the description thereof is omitted.
[0232] The subtraction unit 102 subtracts the pixel value of the predicted image of the block input from the prediction image generation unit 101 from the pixel value of the image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the transform and quantization unit 103.
[0233] The transform and quantization unit 103 calculates transform coefficients for the prediction error input from the subtraction unit 102 by frequency conversion, and derives quantized transform coefficients by quantization. The transform and quantization unit 103 outputs the quantized transform coefficients to the entropy encoding unit 104 and the inverse quantization and inverse transform unit 105.
[0234] The inverse quantization and inverse transform unit 105 is the same as the inverse quantization and inverse transform unit 311 (FIG. 7) in the moving image decoder 31, and the description thereof is omitted. The calculated prediction error is output to the addition unit 106.
[0235] The entropy encoding unit 104 receives the quantized transform coefficients from the transform and quantization unit 103 and the encoding parameters from the parameter encoding unit 111. The encoding parameters include, for example, codes such as a reference picture index refIdxLX, a prediction vector index mvp_LX_idx, a differential vector mvdLX, a motion vector accuracy mode amvr_mode, a prediction mode predMode, and a merge index merge_idx.
[0236] The entropy encoding unit 104 entropy-encodes split information, prediction parameters, quantized transform coefficients, etc. to generate and output an encoded stream Te.
[0237] The parameter encoding unit 111 includes a header encoding unit 1110, a CT information encoding unit 1111, a CU encoding unit 1112 (prediction mode encoding unit), and a parameter encoding unit 112, which are not shown. The CU encoding unit 1112 further includes a TU encoding unit 1114.
[0238] Hereinafter, the schematic operations of each module will be described. The parameter encoding unit 111 performs encoding processing of parameters such as header information, split information, prediction information, and quantized transform coefficients.
[0239] The CT information encoding unit 1111 encodes QT, MT (BT, TT) split information, etc. from the encoded data.
[0240] The CU encoding unit 1112 encodes CU information, prediction information, TU split flag split_transform_flag, CU residual flags cbf_cb, cbf_cr, cbf_luma, etc.
[0241] When the TU contains prediction errors, the TU encoding unit 1114 encodes QP update information (quantization correction value) and quantized prediction error (residual_coding).
[0242] The CT information encoding unit 1111 and the CU encoding unit 1112 output syntax elements such as inter prediction parameters (prediction mode predMode, merge flag merge_flag, merge index merge_idx, inter prediction identifier inter_pred_idc, reference picture index refIdxLX, prediction vector index mvp_LX_idx, difference vector mvdLX), intra prediction parameters (prev_intra_luma_pred_flag, mpm_idx, rem_selected_mode_flag, rem_selected_mode, rem_non_selected_mode), and quantized transform coefficients to the entropy encoding unit 104.
[0243] (Configuration of the parameter encoding unit) Based on the prediction parameters input from the encoding parameter determination unit 110, the parameter encoding unit 112 derives inter prediction parameters. The parameter encoding unit 112 includes a configuration that is partially the same as the configuration in which the inter prediction parameter decoding unit 303 derives inter prediction parameters.
[0244] FIG. 19 is a schematic diagram showing the configuration of the parameter encoding unit 112. The configuration of the parameter encoding unit 112 will be described. As shown in FIG. 19, the parameter encoding unit 112 includes a parameter encoding control unit 1121, a merge prediction unit 30374, a sub-block prediction unit (affine prediction unit) 30372, a DMVR unit 30375, an MMVD prediction unit 30376, a Triangle prediction unit 30377, an AMVP prediction parameter derivation unit 3032, and a subtraction unit 1123. The merge prediction unit 30374 includes a merge prediction parameter derivation unit 3036. The parameter encoding control unit 1121 includes a merge index derivation unit 11211 and a vector candidate index derivation unit 11212. Further, the parameter encoding control unit 1121 derives merge_idx, affine_flag, base_candidate_idx, distance_idx, direction_idx, etc. in the merge index derivation unit 11211, and derives mvpLX, etc. in the vector candidate index derivation unit 11212. The merge prediction parameter derivation unit 3036, the AMVP prediction parameter derivation unit 3032, the affine prediction unit 30372, the MMVD prediction unit 30376, and the Triangle prediction unit 30377 may be collectively referred to as a motion vector derivation unit (motion vector derivation device). The parameter encoding unit 112 outputs the motion vector mvLX, the reference picture index refIdxLX, the inter prediction identifier inter_pred_idc, or information indicating these to the predicted image generation unit 101. Further, the parameter encoding unit 112 outputs merge_flag, skip_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_lX_idx, mvdLX, amvr_mode, and affine_flag to the entropy encoding unit 104.
[0245] FIG. 20 is a diagram showing an example of the number of candidates for the search distance and the number of candidates for the derivation direction in the moving image encoding apparatus 11. The parameter encoding control unit 1121 derives parameters (such as base_candidate_idx, distance_idx, direction_idx, etc.) representing the differential vector and outputs them to the MMVD prediction unit 30376. The derivation of the differential vector in the parameter encoding control unit 1121 will be described with reference to FIG. 20. The black circle in the center of the figure is the position pointed to by the prediction vector mvpLX. Centering on this position, eight search distances are searched in each of the four (up, down, left, right) directions. mvpLX is the motion vector of the first and second candidates at the head of the merge candidate list, and searches are performed for each of them. Since there are two prediction vectors (the first and second in the list) in the merge candidate list, the search distance is 8, and the search direction is 4, there are 64 candidates for mvdLX. The mvdLX with the smallest cost among the searched ones is represented by base_candidate_idx, distance_idx, and direction_idx.
[0246] In this way, the MMVD mode is a mode that searches for limited candidate points centered on the prediction vector and derives an appropriate motion vector.
[0247] The merge index derivation unit 11211 derives the merge index merge_idx and outputs it to the merge prediction parameter derivation unit 3036 (merge prediction unit). In the MMVD mode, the merge index derivation unit 11211 sets the value of the merge index merge_idx to the same value as the value of base_candidate_idx. The vector candidate index derivation unit 11212 derives the prediction vector index mvp_lX_idx.
[0248] The merge prediction parameter derivation unit 3036 derives the inter prediction parameter based on the merge index merge_idx.
[0249] The AMVP prediction parameter derivation unit 3032 derives a prediction vector mvpLX based on the motion vector mvLX. The AMVP prediction parameter derivation unit 3032 outputs the prediction vector mvpLX to the subtraction unit 1123. Note that the reference picture index refIdxLX and the prediction vector index mvp_lX_idx are output to the entropy encoding unit 104.
[0250] The affine prediction unit 30372 derives an inter prediction parameter (affine prediction parameter) of the sub-block.
[0251] The subtraction unit 1123 subtracts the prediction vector mvpLX, which is the output of the AMVP prediction parameter derivation unit 3032, from the motion vector mvLX input from the encoding parameter determination unit 110 to generate a differential vector mvdLX. The differential vector mvdLX is output to the entropy encoding unit 104.
[0252] The addition unit 106 adds the pixel value of the predicted image of the block input from the prediction image generation unit 101 and the prediction error input from the inverse quantization / inverse transformation unit 105 for each pixel to generate a decoded image. The addition unit 106 stores the generated decoded image in the reference picture memory 109.
[0253] The loop filter 107 performs a deblocking filter, SAO, and ALF on the decoded image generated by the addition unit 106. Note that the loop filter 107 does not necessarily include the above three types of filters, and for example, it may have a configuration including only the deblocking filter.
[0254] The prediction parameter memory 108 stores the prediction parameters generated by the encoding parameter determination unit 110 at predetermined positions for each target picture and CU.
[0255] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at predetermined positions for each target picture and CU.
[0256] The encoding parameter determination unit 110 selects one set from among a plurality of sets of encoding parameters. The encoding parameters are the QT, BT, or TT splitting information, prediction parameters, or parameters to be encoded generated in relation to these, as described above. The predicted image generation unit 101 generates a predicted image using these encoding parameters.
[0257] The encoding parameter determination unit 110 calculates an RD cost value indicating the amount of information and the encoding error for each of the plurality of sets. The encoding parameter determination unit 110 selects the set of encoding parameters for which the calculated cost value is minimized. As a result, the entropy encoding unit 104 outputs the selected set of encoding parameters as an encoded stream Te. The encoding parameter determination unit 110 stores the determined encoding parameters in the prediction parameter memory 108.
[0258] Note that, a part of the moving image encoding device 11 and the moving image decoding device 31 in the above-described embodiment, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generation unit 308, the inverse quantization / inverse transform unit 311, the addition unit 312, the predicted image generation unit 101, the subtraction unit 102, the transform / quantization unit 103, the entropy encoding unit 104, the inverse quantization / inverse transform unit 105, the loop filter 107, the encoding parameter determination unit 110, and the parameter encoding unit 111 may be realized by a computer. In that case, a program for realizing this control function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to realize it. Here, the "computer system" refers to a computer system built in either the moving image encoding device 11 or the moving image decoding device 31, and includes hardware such as an OS and peripheral devices. Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, or a storage device such as a hard disk built in a computer system. Furthermore, the "computer-readable recording medium" also includes, like a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, something that holds a program dynamically for a short time, and something that holds a program for a certain time, like a volatile memory inside a computer system that becomes a server or a client in that case. Also, the above program may be for realizing a part of the aforementioned functions, and may further be something that can be realized in combination with a program already recorded in a computer system for the aforementioned functions.
[0259] Further, part or all of the moving image encoding device 11 and the moving image decoding device 31 in the above-described embodiments may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the moving image encoding device 11 and the moving image decoding device 31 may be individually processed by a processor, or part or all of them may be integrated and processed by a processor. Further, the method of integrating into an integrated circuit is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. Also, when a technology for integrating into an integrated circuit that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit using such technology may be used.
[0260] As described above, one embodiment of the present invention has been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist of the present invention.
[0261] 〔Application Example〕 The above-described moving image encoding device 11 and moving image decoding device 31 can be mounted and used in various devices that transmit, receive, record, and play back moving images. The moving image may be a natural moving image captured by a camera or the like, or an artificial moving image (including CG and GUI) generated by a computer or the like.
[0262] First, the fact that the above-described moving image encoding device 11 and moving image decoding device 31 can be used for transmission and reception of moving images will be described with reference to FIG. 2.
[0263] FIG. 2(a) is a block diagram showing the configuration of a transmission device PROD_A equipped with the moving image encoding device 11. As shown in the figure, the transmission device PROD_A includes an encoding unit PROD_A1 that obtains encoded data by encoding a moving image, a modulation unit PROD_A2 that obtains a modulation signal by modulating a carrier wave with the encoded data obtained by the encoding unit PROD_A1, and a transmission unit PROD_A3 that transmits the modulation signal obtained by the modulation unit PROD_A2. The above-described moving image encoding device 11 is used as this encoding unit PROD_A1.
[0264] The transmission device PROD_A may further include a camera PROD_A4 for capturing a moving image, a recording medium PROD_A5 for recording a moving image, an input terminal PROD_A6 for inputting a moving image from the outside, and an image processing unit A7 for generating or processing an image, as an input source of the moving image input to the encoding unit PROD_A1. In the figure, a configuration in which the transmission device PROD_A includes all of these is illustrated, but a part of them may be omitted.
[0265] Note that the recording medium PROD_A5 may record an unencoded moving image, or may record a moving image encoded by a recording encoding method different from the transmission encoding method. In the latter case, a decoding unit (not shown) for decoding the encoded data read from the recording medium PROD_A5 according to the recording encoding method may be interposed between the recording medium PROD_A5 and the encoding unit PROD_A1.
[0266] Figure 2(b) is a block diagram showing the configuration of the receiving device PROD_B equipped with the moving image decoding device 31. As shown in the figure, the receiving device PROD_B includes a receiving unit PROD_B1 for receiving a modulated signal, a demodulating unit PROD_B2 for obtaining encoded data by demodulating the modulated signal received by the receiving unit PROD_B1, and a decoding unit PROD_B3 for obtaining a moving image by decoding the encoded data obtained by the demodulating unit PROD_B2. The above-described moving image decoding device 31 is used as this decoding unit PROD_B3.
[0267] The receiving device PROD_B may further include a display PROD_B4 for displaying a moving image, a recording medium PROD_B5 for recording a moving image, and an output terminal PROD_B6 for outputting a moving image to the outside, as an output destination of the moving image output by the decoding unit PROD_B3. In the figure, a configuration in which the receiving device PROD_B includes all of these is illustrated, but a part of them may be omitted.
[0268] Note that the recording medium PROD_B5 may be for recording unencoded moving images, or may be encoded using an encoding method for recording different from the encoding method for transmission. In the latter case, an encoding unit (not shown) for encoding the moving image acquired from the decoding unit PROD_B3 according to the encoding method for recording may be interposed between the decoding unit PROD_B3 and the recording medium PROD_B5.
[0269] Note that the transmission medium for transmitting the modulation signal may be wireless or wired. Also, the transmission mode for transmitting the modulation signal may be broadcasting (here, referring to a transmission mode where the transmission destination is not specified in advance), or may be communication (here, referring to a transmission mode where the transmission destination is specified in advance). That is, the transmission of the modulation signal may be realized by any of wireless broadcasting, wired broadcasting, wireless communication, and wired communication.
[0270] For example, a broadcasting station (broadcasting facilities, etc.) / receiving station (television receiver, etc.) for terrestrial digital broadcasting is an example of the transmitting device PROD_A / receiving device PROD_B that transmits and receives the modulation signal by wireless broadcasting. Also, a broadcasting station (broadcasting facilities, etc.) / receiving station (television receiver, etc.) for cable television broadcasting is an example of the transmitting device PROD_A / receiving device PROD_B that transmits and receives the modulation signal by wired broadcasting.
[0271] Also, a server (workstation, etc.) / client (television receiver, personal computer, smartphone, etc.) for VOD (Video On Demand) services or video sharing services using the Internet is an example of the transmitting device PROD_A / receiving device PROD_B that transmits and receives the modulation signal by communication (usually, either wireless or wired is used as the transmission medium in a LAN, and wired is used as the transmission medium in a WAN). Here, personal computers include desktop PCs, laptop PCs, and tablet PCs. Also, smartphones include multifunctional mobile phone terminals.
[0272] In addition, the client of the video sharing service has a function of decoding the encoded data downloaded from the server and displaying it on the display, and also has a function of encoding the moving images captured by the camera and uploading them to the server. That is, the client of the video sharing service functions as both the transmission device PROD_A and the reception device PROD_B.
[0273] Next, with reference to FIG. 3, it will be described that the above-described moving image encoding device 11 and moving image decoding device 31 can be used for recording and playing back moving images.
[0274] FIG. 3(a) is a block diagram showing the configuration of the recording device PROD_C equipped with the above-described moving image encoding device 11. As shown in the figure, the recording device PROD_C includes an encoding unit PROD_C1 that obtains encoded data by encoding moving images, and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 to the recording medium PROD_M. The above-described moving image encoding device 11 is used as this encoding unit PROD_C1.
[0275] Note that the recording medium PROD_M may be of a type built into the recording device PROD_C, such as (1) an HDD (Hard Disk Drive) or an SSD (Solid State Drive), or may be of a type connected to the recording device PROD_C, such as (2) an SD memory card or a USB (Universal Serial Bus) flash memory, or may be loaded into a drive device (not shown) built into the recording device PROD_C, such as (3) a DVD (Digital Versatile Disc: registered trademark) or a BD (Blu-ray Disc: registered trademark).
[0276] Further, the recording device PROD_C may further include a camera PROD_C3 for capturing a moving image, an input terminal PROD_C4 for inputting a moving image from the outside, a receiving unit PROD_C5 for receiving a moving image, and an image processing unit PROD_C6 for generating or processing an image, as an input source of the moving image input to the encoding unit PROD_C1. In the figure, a configuration in which the recording device PROD_C includes all of these is illustrated, but a part of them may be omitted.
[0277] Note that the receiving unit PROD_C5 may receive an unencoded moving image, or may receive encoded data encoded by a transmission encoding method different from the recording encoding method. In the latter case, a transmission decoding unit (not shown) for decoding the encoded data encoded by the transmission encoding method may be interposed between the receiving unit PROD_C5 and the encoding unit PROD_C1.
[0278] Examples of such a recording device PROD_C include a DVD recorder, a BD recorder, an HDD (Hard Disk Drive) recorder, etc. (in this case, the input terminal PROD_C4 or the receiving unit PROD_C5 serves as the main input source of the moving image). Also, a camcorder (in this case, the camera PROD_C3 serves as the main input source of the moving image), a personal computer (in this case, the receiving unit PROD_C5 or the image processing unit C6 serves as the main input source of the moving image), a smartphone (in this case, the camera PROD_C3 or the receiving unit PROD_C5 serves as the main input source of the moving image), etc. are also examples of such a recording device PROD_C.
[0279] Figure 3(b) is a block diagram showing the configuration of the playback device PROD_D equipped with the above-described moving image decoding device 31. As shown in the figure, the playback device PROD_D includes a reading unit PROD_D1 for reading the encoded data written in the recording medium PROD_M, and a decoding unit PROD_D2 for obtaining a moving image by decoding the encoded data read by the reading unit PROD_D1. The above-described moving image decoding device 31 is used as this decoding unit PROD_D2.
[0280] Note that the recording medium PROD_M may be of a type built into the playback device PROD_D, such as an HDD or an SSD, or may be of a type connected to the playback device PROD_D, such as an SD memory card or a USB flash memory, or may be loaded into a drive device (not shown) built into the playback device PROD_D, such as a DVD or a BD.
[0281] The playback device PROD_D may further include a display PROD_D3 for displaying a moving image, an output terminal PROD_D4 for outputting the moving image externally, and a transmission unit PROD_D5 for transmitting the moving image as output destinations of the moving image output by the decoding unit PROD_D2. In the figure, a configuration in which the playback device PROD_D includes all of these is illustrated, but a part of them may be omitted.
[0282] Note that the transmission unit PROD_D5 may transmit an unencoded moving image or may transmit encoded data encoded by a transmission encoding method different from the recording encoding method. In the latter case, an encoding unit (not shown) for encoding the moving image by the transmission encoding method may be interposed between the decoding unit PROD_D2 and the transmission unit PROD_D5.
[0283] Examples of such a playback device PROD_D include, for example, a DVD player, a BD player, an HDD player, etc. (in this case, the output terminal PROD_D4 to which a television receiver or the like is connected becomes the main output destination for moving images). Also, a television receiver (in this case, the display PROD_D3 becomes the main output destination for moving images), digital signage (also referred to as an electronic signboard or electronic bulletin board, etc., and the display PROD_D3 or the transmission unit PROD_D5 becomes the main output destination for moving images), a desktop PC (in this case, the output terminal PROD_D4 or the transmission unit PROD_D5 becomes the main output destination for moving images), a laptop or tablet PC (in this case, the display PROD_D3 or the transmission unit PROD_D5 becomes the main output destination for moving images), a smartphone (in this case, the display PROD_D3 or the transmission unit PROD_D5 becomes the main output destination for moving images), etc. are also examples of such a playback device PROD_D.
[0284] (Hardware implementation and software implementation) In addition, each block of the above-described moving image decoding device 31 and moving image encoding device 11 may be implemented hardware-wise by a logic circuit formed on an integrated circuit (IC chip), or may be implemented software-wise using a CPU (Central Processing Unit).
[0285] In the latter case, each of the above devices includes a CPU that executes instructions of a program for realizing each function, a ROM (Read Only Memory) that stores the above program, a RAM (Random Access Memory) that expands the above program, a storage device (recording medium) such as a memory that stores the above program and various data, etc. And the object of the embodiment of the present invention can also be achieved by supplying a recording medium in which program codes (executable format program, intermediate code program, source program) of control programs of each of the above devices, which are software for realizing the above-described functions, are recorded in a computer-readable manner to each of the above devices, and having the computer (or CPU or MPU) read and execute the program codes recorded in the recording medium.
[0286] Examples of the recording medium include tapes such as magnetic tapes and cassette tapes, magnetic disks such as floppy (registered trademark) disks / hard disks, and disks including optical disks such as CD-ROM (Compact Disc Read-Only Memory) / MO disks (Magneto-Optical disc) / MD (Mini Disc) / DVD (Digital Versatile Disc: registered trademark) / CD-R (CD Recordable) / Blu-ray Disc (Blu-ray Disc: registered trademark), cards such as IC cards (including memory cards) / optical cards, semiconductor memories such as mask ROM / EPROM (Erasable Programmable Read-Only Memory) / EEPROM (Electrically Erasable and Programmable Read-Only Memory: registered trademark) / flash ROM, or logic circuits such as PLD (Programmable logic device) and FPGA (Field Programmable Gate Array).
[0287] Alternatively, each of the above devices may be configured to be connectable to a communication network, and the above program code may be supplied via the communication network. This communication network only needs to be capable of transmitting the program code and is not particularly limited. For example, the Internet, intranet, extranet, LAN (Local Area Network), ISDN (Integrated Services Digital Network), VAN (Value-Added Network), CATV (Community Antenna television / Cable Television) communication network, virtual private network, telephone line network, mobile communication network, satellite communication network, etc. can be used. Also, the transmission medium constituting this communication network only needs to be a medium capable of transmitting the program code and is not limited to a specific configuration or type. For example, it can be wired such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line carrier, cable TV line, telephone line, ADSL (Asymmetric Digital Subscriber Line) line, etc., or wireless such as infrared rays like IrDA (Infrared Data Association) and remote controls, Bluetooth (registered trademark), IEEE802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance: registered trademark), mobile phone network, satellite line, terrestrial digital broadcast network, etc. Note that the embodiments of the present invention can also be realized in the form of a computer data signal embedded in a carrier wave, in which the above program code is embodied by electronic transmission.
[0288] The embodiments of the present invention are not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. That is, embodiments obtained by appropriately combining technical means modified within the scope shown in the claims are also included in the technical scope of the present invention.
[0289] (Cross - reference to related applications) This application claims the benefit of priority to Japanese Patent Application: Japanese Patent Application No. 2019 - 043097, filed on March 8, 2018, and by reference thereto, the entire contents thereof are incorporated herein.
[0290] (Summary) The present invention can also be expressed as follows.
[0291] An image decoding apparatus according to an aspect of the present invention includes an inter - prediction parameter decoding unit that has a process of correcting two motion vectors from the errors between two prediction images, and performs a process of correcting the two motion vectors when neither of the two prediction images is a case of weighted prediction.
[0292] Also, an image encoding apparatus according to an aspect of the present invention includes an inter - prediction parameter encoding unit that has a process of correcting two motion vectors from the errors between two prediction images, and performs a process of correcting the two motion vectors when neither of the two prediction images is a case of weighted prediction.
[0293] By adopting such a configuration, when weighted prediction is applied, accurate error evaluation cannot be performed, so an effect cannot be obtained. Therefore, by restricting the application conditions, the overall processing amount can be reduced.
[0294] Also, an image decoding apparatus according to an aspect of the present invention includes an inter - prediction parameter decoding unit that has a process of correcting two motion vectors from the error values between two prediction images, and a bidirectional gradient change processing unit that generates a prediction image using a gradient image derived from two interpolated images generated using the parameters decoded by the inter - prediction parameter decoding unit. Determine whether to apply the processing by the bidirectional gradient change processing unit using the error values of the two predicted images.
[0295] Also, an image encoding apparatus according to an aspect of the present invention has an inter prediction parameter encoding unit that corrects two motion vectors from the error values of two predicted images and has a process of correcting the two motion vectors, has a bidirectional gradient change processing unit that generates a predicted image using a gradient image derived from two generated interpolation images using the parameters decoded by the inter prediction parameter decoding unit, Determine whether to apply the processing by the bidirectional gradient change processing unit using the error values of the two predicted images.
[0296] By adopting such a configuration, it is necessary to obtain an error value in order to correct the motion vector. On the other hand, in the bidirectional gradient change processing unit, when the error is small, there is no effect. Therefore, by adding such processing, it is possible to determine whether to apply the processing by the bidirectional gradient change processing unit without adding an additional error value, and the overall processing amount can be reduced.
Industrial Applicability
[0297] Embodiments of the present invention can be suitably applied to a moving image decoding apparatus that decodes encoded data in which image data is encoded, and a moving image encoding apparatus that generates encoded data in which image data is encoded. Further, it can be suitably applied to the data structure of the encoded data generated by the moving image encoding apparatus and referred to by the moving image decoding apparatus.
Explanation of Signs
[0298] 31 Image decoding apparatus 301 Entropy decoding unit 302 Parameter decoding unit 3020 Header decoding unit 303 Inter prediction parameter decoding unit 304 Intra prediction parameter decoding unit 308 Prediction Image Generation Unit 309 Inter-Prediction Image Generation Unit 310 Intra-Prediction Image Generation Unit 311 Inverse Quantization / Inverse Transformation Unit 312 Addition Unit 11 Image Encoding Device 101 Prediction Image Generation Unit 102 Subtraction Unit 103 Transformation / Quantization Unit 104 Entropy Encoding Unit 105 Inverse Quantization / Inverse Transformation Unit 107 Loop Filter 110 Encoding Parameter Determination Unit 111 Parameter Encoding Unit 112 Parameter Encoding Unit 1110 Header Encoding Unit 1111 CT Information Encoding Unit 1112 CU Encoding Unit (Prediction Mode Encoding Unit) 1114 TU Encoding Unit 3091 Motion Compensation Unit 3095 Synthesis Unit 30951 Combined Intra / inter Synthesis Unit 30952 Triangle Synthesis Unit 30953 OBMC Unit 30954 Weighted Prediction Unit 30955 GBI Unit 30956 BDOF Unit 309561 L0,L1 Prediction Image Generation Unit 309562 Gradient Image Generation Unit 309563 Correlation Parameter Calculation Unit 309564 Motion Compensation Correction Value Derivation Unit 309565 BDOF Prediction Image Generation Unit
Claims
1. A moving image decoding apparatus that performs DMVR (Decoder side Motion Vector Refinement) processing using two reference pictures and motion vectors mvL0 and mvL1, a DMVR unit that executes the DMVR processing when a dmvrFlag indicating whether the DMVR processing is to be performed is TRUE, and a weighted prediction unit that performs weighted prediction using a first weight coefficient, a first offset, a second weight coefficient, and a second offset. The DMVR unit sets the dmvrFlag to TRUE based on a predetermined condition, where the predetermined condition for setting the dmvrFlag to TRUE includes that both luma_weight_l0_flag[refIdxL0] and luma_weight_l1_flag[refIdxL1] are FALSE. The luma_weight_l0_flag[refIdxL0] indicates whether the first weight coefficient and the first offset of the luminance corresponding to the L0 reference picture indicated by the reference picture index refIdxL0 exist. The moving image decoding apparatus is characterized in that the luma_weight_l1_flag[refIdxL1] indicates whether the second weight coefficient and the second offset of the luminance corresponding to the L1 reference picture indicated by the reference picture index refIdxL1 exist.
2. A moving image encoding apparatus that performs DMVR (Decoder side Motion Vector Refinement) processing using two reference pictures and motion vectors mvL0 and mvL1, a DMVR unit that executes the DMVR processing when a dmvrFlag indicating whether the DMVR processing is to be performed is TRUE, and a weighted prediction unit that performs weighted prediction using a first weight coefficient, a first offset, a second weight coefficient, and a second offset. The DMVR unit sets the dmvrFlag to TRUE based on a predetermined condition, where the predetermined condition for setting the dmvrFlag to TRUE includes that both luma_weight_l0_flag[refIdxL0] and luma_weight_l1_flag[refIdxL1] are FALSE. The luma_weight_l0_flag[ refIdxL0 ] indicates whether or not the first weight coefficient and the first offset of the luminance corresponding to the L0 reference picture indicated by the reference picture index refIdxL0 exist. The moving image encoding apparatus is characterized in that the luma_weight_l1_flag[ refIdxL1 ] indicates whether or not the second weight coefficient and the second offset of the luminance corresponding to the L1 reference picture indicated by the reference picture index refIdxL1 exist.
Citation Information
Patent Citations
Decoder Side Motion Vector Refinement in Video Coding
US20190020895A1
Block size restrictions for dmvr
WO2020008343A1
Apparatus and method for conditional decoder-side motion vector refinement in video coding
WO2020052654A1
Methods and devices for selectively applying bi-directional optical flow and decoder-side motion vector refinement for video coding
WO2020163837A1
DMVR-based inter-prediction method and device
WO2020166897A1