Moving image decoding device

The moving image decoding apparatus addresses inefficiencies in managing reference pictures and weights by using specific flags and syntax elements, improving the decoding process through clearer definitions and reduced redundancy.

JP7714079B2Active Publication Date: 2025-07-28SHARP KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024064079
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-07-28
Estimated Expiration
2040-03-17

AI Technical Summary

Technical Problem

Existing moving image encoding and decoding technologies, such as those described in Non-Patent Document 1, face issues with indefinite reference picture lists and redundant weight definitions, leading to inefficiencies in managing reference pictures and weights for prediction.

Method used

A moving image decoding apparatus is designed with specific flags and syntax elements to manage reference picture lists and weights more effectively, including a first flag for weight prediction in the picture header, a second flag for B slices, and syntax elements to derive weight coefficients, ensuring clear definitions and reducing redundancy.

Benefits of technology

This approach resolves the issues of indefinite reference pictures and redundant weight definitions, enhancing the efficiency and clarity of the moving image decoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007714079000001
    Figure 0007714079000001
  • Figure 0007714079000002
    Figure 0007714079000002
  • Figure 0007714079000003
    Figure 0007714079000003
Patent Text Reader

Abstract

To solve a problem in which, in managing a reference picture list, a reference picture list structure in which the number of reference pictures is 0 can be defined, which causes the reference pictures to become undefined.SOLUTION: A video decoding device includes an inter-prediction unit that decodes a plurality of reference picture list structures and selects one reference picture list structure from the plurality of reference picture list structures on a picture-by-picture or slice-by-slice basis, and the number of reference pictures in each of the plurality of reference picture list structures is 1 or greater.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a predicted image generation device, a moving image decoding device, and a moving image encoding device.

Background Art

[0002] In order to efficiently transmit or record a moving image, a moving image encoding device that generates encoded data by encoding the moving image and a moving image decoding device that generates a decoded image by decoding the encoded data are used.

[0003] Specific moving image encoding methods include, for example, the H.264 / AVC and H.265 / HEVC (High-Efficiency Video Coding) methods.

[0004] In such a moving image encoding method, an image (picture) constituting the moving image is managed by a hierarchical structure consisting of a slice obtained by dividing the image, a coding tree unit (CTU) obtained by dividing the slice, a coding unit (which may also be called a coding unit (CU)) obtained by dividing the coding tree unit, and a transform unit (TU) obtained by dividing the coding unit, and is encoded / decoded for each CU.

[0005] Also, in such a moving image encoding method, usually, a predicted image is generated based on a local decoded image obtained by encoding / decoding an input image, and a prediction error (which may also be called a "difference image" or a "residual image") obtained by subtracting the predicted image from the input image (original image) is encoded. Examples of the method for generating the predicted image include inter-picture prediction (inter prediction) and intra-picture prediction (intra prediction).

[0006] Non-Patent Document 1 is cited as a recent technology for moving image encoding and decoding.

[0007] In Non-Patent Document 1, in the management of the reference picture list for inter prediction, a mechanism is adopted in which a plurality of reference picture lists are defined and used with reference thereto. Also, in weighted prediction, a method of explicitly defining the number of weights is adopted.

Prior Art Documents

Non-Patent Documents

[0008]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0009] However, in Non-Patent Document 1, in the management of the reference picture list, there is a problem that since a reference picture list structure in which the number of reference pictures is 0 can be defined, the reference pictures become indefinite.

[0010] Also, among the defined reference picture list structures, there is a problem that the number of reference pictures actually used for prediction can be defined in the slice header but not in the picture header.

[0011] Also, in weighted prediction, although the number of weights is explicitly defined, since the number of reference pictures that can be used is already known, there is a problem that it is redundant description.

Means for Solving the Problems

[0012] A moving image decoding apparatus according to an aspect of the present invention includes a first flag indicating whether weight prediction information is in a picture header, a second flag indicating whether weight prediction is applied to a B slice, a first syntax element for deriving a weight coefficient, and a second syntax element indicating the number of entries included in at least reference picture list 1, a parameter decoding unit that decodes the first syntax element and the second syntax element, a prediction parameter derivation unit that derives an inter prediction parameter, a motion compensation unit that generates an interpolated image based on the inter prediction parameter and a reference picture, and a weight prediction unit that derives the weight coefficient using the first syntax element and generates a predicted image using the interpolated image and the weight coefficient. The parameter decoding unit decodes a third syntax element indicating the number of weights signaled for the entries included in reference picture list 1 based on the first flag and the second flag. When the third syntax element is not decoded, the weight prediction unit sets a variable NumWeightL1 indicating the number of weights signaled for the entries included in reference picture list 1 to be equal to 0. When the value of the variable NumWeightL1 is equal to 0, the weight prediction unit sets the value of the weight coefficient for reference picture list 1 to the power of 2 of the first syntax element.

Advantages of the Invention

[0013] According to an aspect of the present invention, the above problems can be solved.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Embodiments for Carrying Out the Invention

[0015] (First Embodiment) Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0016] FIG. 1 is a schematic diagram showing the configuration of an image transmission system 1 according to the present embodiment.

[0017] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding images of different resolutions with their resolutions converted, decodes the transmitted encoded stream, and inversely converts the image to the original resolution for display. The image transmission system 1 includes a resolution conversion device (resolution conversion unit) 51, a moving image encoding device (image encoding device) 11, a network 21, a moving image decoding device (image decoding device) 31, a resolution inverse conversion device (resolution inverse conversion unit) 61, and a moving image display device (image display device) 41.

[0018] The resolution conversion device 51 converts the resolution of the image T included in the moving image, and supplies a variable resolution moving image signal including images of different resolutions to the image encoding device 11. Further, the resolution conversion device 51 supplies information indicating whether or not the resolution of the image is converted to the image encoding device 11. When the information indicates resolution conversion, the image encoding device sets the resolution conversion information ref_pic_resampling_enabled_flag, which will be described later, to 1, and includes it in the sequence parameter set SPS (Sequence Parameter Set) of the encoded data for encoding.

[0019] An image T with its resolution converted is input to the moving image encoding device 11.

[0020] The network 21 transmits the encoded stream Te generated by the moving image encoding device 11 to the moving image decoding device 31. The network 21 is the Internet, a wide area network (WAN: Wide Area Network), a local area network (LAN: Local Area Network), or a combination thereof. The network 21 is not necessarily limited to a bidirectional communication network, and may be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Further, the network 21 may be replaced by a storage medium that records the encoded stream Te such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark).

[0021] The moving image decoding device 31 decodes each of the encoded streams Te transmitted by the network 21, generates a variable resolution decoded image signal, and supplies it to the resolution inverse conversion device 61.

[0022] When the resolution conversion information included in the variable resolution decoded image signal indicates resolution conversion, the resolution inverse conversion device 61 generates a decoded image signal of the original size by inversely converting the resolution-converted image.

[0023] The moving image display device 41 displays all or part of one or more decoded images Td indicated by the decoded image signal input from the resolution inverse conversion unit. The moving image display device 41 includes, for example, a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Examples of the form of the display include a stationary type, a mobile type, and an HMD. Also, when the moving image decoding device 31 has high processing power, it displays an image with high image quality, and when it has only lower processing power, it displays an image that does not require high processing power and display capabilities.

[0024] <Operator> The operators used in this specification are described below.

[0025] >>>> is a right bit shift, <<< is a left bit shift, & is a bitwise AND, | is a bitwise OR, |= is an OR assignment operator, and || indicates a logical OR.

[0026] x? y : z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0).

[0027] Clip3(a, b, c) is a function that clips c to a value between a and b (inclusive). If c < a, it returns a; if c > b, it returns b; otherwise, it returns c (where a <= b).

[0028] abs(a) is a function that returns the absolute value of a.

[0029] Int(a) is a function that returns the integer value of a.

[0030] floor(a) is a function that returns the largest integer less than or equal to a.

[0031] ceil(a) is a function that returns the smallest integer greater than or equal to a.

[0032] a / d represents the division of a by d (truncating the decimal part).

[0033] min(a,b) represents the smaller value of a and b.

[0034] <Structure of the Encoded Stream Te> Prior to the detailed description of the moving image encoding device 11 and the moving image decoding device 31 according to the present embodiment, the data structure of the encoded stream Te generated by the moving image encoding device 11 and decoded by the moving image decoding device 31 will be described.

[0035] FIG. 4 is a diagram showing the hierarchical structure of data in the encoded stream Te. The encoded stream Te illustratively includes a sequence and a plurality of pictures constituting the sequence. FIG. 4 shows an encoded video sequence that defines the sequence SEQ, an encoded picture that defines the picture PICT, an encoded slice that defines the slice S, encoded slice data that defines the slice data, an encoded tree unit included in the encoded slice data, and an encoded unit included in the encoded tree unit.

[0036] (Encoded Video Sequence) In a symbolized video sequence, a set of data that the moving image decoding device 31 refers to in order to decode the sequence SEQ to be processed is defined. As shown in FIG. 4, the sequence SEQ includes a video parameter set VPS (Video Parameter Set), a sequence parameter set SPS (Sequence Parameter Set), a picture parameter set PPS (Picture Parameter Set), an Adaptation Parameter Set (APS), a picture PICT, and supplemental enhancement information SEI (Supplemental Enhancement Information).

[0037] The video parameter set VPS defines a set of encoding parameters common to a plurality of moving images and a set of encoding parameters related to a plurality of layers and individual layers included in the moving images in a moving image composed of a plurality of layers.

[0038] In the sequence parameter set SPS, a set of encoding parameters that the moving image decoding device 31 refers to in order to decode the target sequence is defined. For example, the width and height of a picture are defined. Note that there may be a plurality of SPSs. In that case, one of the plurality of SPSs is selected from the PPS.

[0039] Here, the sequence parameter set SPS includes the following syntax. ·ref_pic_resampling_enabled_flag: A flag that defines whether to use a function (resampling) that makes the resolution variable when decoding each image included in a single sequence that refers to the target SPS. From another aspect, the flag indicates that the size of the reference picture referred to in the generation of the predicted image changes among the images indicated by a single sequence. When the value of the flag is 1, the above resampling is applied; when it is 0, it is not applied. ·pic_width_max_in_luma_samples: It is the syntax that specifies the width of the image with the maximum width among the images in a single sequence in units of luminance blocks. Also, the value of this syntax is required to be not 0 and an integer multiple of Max(8, MinCbSizeY). Here, MinCbSizeY is a value determined by the minimum size of the luminance block. ·pic_height_max_in_luma_samples: It is the syntax that specifies the height of the image with the maximum height among the images in a single sequence in units of luminance blocks. Also, the value of this syntax is required to be not 0 and an integer multiple of Max(8, MinCbSizeY). ·sps_temporal_mvp_enabled_flag: It is a flag that specifies whether to use temporal motion vector prediction when decoding the target sequence. If the value of this flag is 1, temporal motion vector prediction is used; if the value is 0, temporal motion vector prediction is not used. Also, by specifying this flag, it is possible to prevent the reference coordinate positions from shifting when referring to reference pictures with different resolutions, etc.

[0040] In the picture parameter set PPS, a set of encoding parameters that the moving image decoding device 31 refers to in order to decode each picture in the target sequence is defined. For example, it includes the reference value of the quantization width (pic_init_qp_minus26) used for picture decoding and the flag (weighted_pred_flag) indicating the application of weighted prediction. Note that there may be multiple PPSs. In that case, one of the multiple PPSs is selected from each picture in the target sequence.

[0041] Here, the picture parameter set PPS includes the following syntax. ·pic_width_in_luma_samples: A syntax that specifies the width of the target picture. The value of this syntax is required to be a non-zero integer multiple of Max(8, MinCbSizeY) and less than or equal to pic_width_max_in_luma_samples. ·pic_height_in_luma_samples: A syntax that specifies the height of the target picture. The value of this syntax is required to be a non-zero integer multiple of Max(8, MinCbSizeY) and less than or equal to pic_height_max_in_luma_samples. ·conformance_window_flag: A flag indicating whether the conformance (cropping) window offset parameter will be subsequently signaled, and also a flag indicating the location where the conformance window is to be displayed. If this flag is 1, the parameter is signaled; if it is 0, it indicates that there is no conformance window offset parameter. ·conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, conf_win_bottom_offset: Offset values for specifying the left, right, top, and bottom positions of the picture output in the decoding process with respect to a rectangular area specified in the output picture coordinates. Also, when the value of conformance_window_flag is 0, the values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are assumed to be 0. ·scaling_window_flag: A flag indicating whether the scaling window offset parameter exists in the target PPS, which is a flag related to the definition of the output image size. If this flag is 1, it indicates that the parameter exists in the PPS; if this flag is 0, it indicates that the parameter does not exist in the PPS. Also, when the value of ref_pic_resampling_enabled_flag is 0, the value of scaling_window_flag is also required to be 0. ·scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, scaling_win_bottom_offset: A syntax for specifying, in luminance sample units, the offsets applied to the image size for calculating the scaling ratio for the left, right, top, and bottom positions of the target picture, respectively. Also, when the value of scaling_window_flag is 0, the values of scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset are presumed to be 0. Also, the value of scaling_win_left_offset + scaling_win_right_offset is required to be less than pic_width_in_luma_samples, and the value of scaling_win_top_offset + scaling_win_bottom_offset is required to be less than pic_height_in_luma_samples.

[0042] The width PicOutputWidthL and height PicOutputHeightL of the output picture are derived as follows.

[0043] PicOutputWidthL = pic_width_in_luma_samples - (scaling_win_right_offset + scaling_win_left_offset) PicOutputHeightL = pic_height_in_pic_size_units - (scaling_win_bottom_offset + scaling_win_top_offset) · pps_collocated_from_l0_idc: The syntax indicating whether collocated_from_l0_flag exists in the slice header of the slice referring to the PPS. When the value of this syntax is 0, it indicates that collocated_from_l0_flag exists in the slice header, and when it is 1 or 2, it indicates that it does not exist in the slice header.

[0044] (Encoded Picture) In the encoded picture, a set of data that the moving image decoding device 31 refers to in order to decode the picture PICT to be processed is defined. As shown in FIG. 4, the picture PICT includes a picture header PH and slices 0 to NS-1 (NS is the total number of slices included in the picture PICT).

[0045] Hereinafter, when it is not necessary to distinguish each of slices 0 to NS-1, the subscript of the symbol may be omitted in the description. The same applies to the data included in the encoded stream Te described below and other data with subscripts.

[0046] The picture header includes the following syntax. · pic_temporal_mvp_enabled_flag: A flag that specifies whether to use temporal motion vector prediction for inter prediction of slices associated with the picture header. When the value of this flag is 0, the syntax elements of the slice associated with the picture header are restricted so that temporal motion vector prediction is not used in decoding the slice. When the value of this flag is 1, it indicates that temporal motion vector prediction is used in decoding the slice associated with the picture header. Also, when this flag is not specified, it is assumed to have a value of 0.

[0047] (Encoded slice) In the encoded slice, a set of data that the moving image decoder 31 refers to for decoding the slice S to be processed is defined. As shown in FIG. 4, the slice includes a slice header and slice data.

[0048] The slice header includes a group of encoding parameters that the moving image decoder 31 refers to for determining the decoding method of the target slice. The slice type specification information (slice_type) that specifies the slice type is an example of the encoding parameters included in the slice header.

[0049] Slice types that can be specified by the slice type specification information include: (1) I slice that uses only intra prediction during encoding; (2) P slice that uses either single prediction (L0 prediction) or intra prediction during encoding; (3) B slice that uses single prediction (L0 prediction using only reference picture list 0 or L1 prediction using only reference picture list 1), bi-prediction, or intra prediction during encoding. Note that inter prediction is not limited to single prediction and bi-prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referring to P and B slices, it means a slice including a block that can use inter prediction.

[0050] Note that the slice header may include a reference (pic_parameter_set_id) to the picture parameter set PPS.

[0051] (Encoded slice data) In the encoded slice data, a set of data that the moving image decoding device 31 refers to in order to decode the slice data to be processed is defined. As shown in the encoded slice header of FIG. 4, the slice data includes CTUs. A CTU is a block of a fixed size (for example, 64x64) that constitutes a slice, and is sometimes called the largest coding unit (LCU).

[0052] (Coding tree unit) FIG. 4 defines a set of data that the moving image decoding device 31 refers to in order to decode the CTU to be processed. The CTU is divided into coding units (CUs), which are the basic units of the encoding process, by recursive quadtree partitioning (QT (Quad Tree) partitioning), binary tree partitioning (BT (Binary Tree) partitioning), or ternary tree partitioning (TT (Ternary Tree) partitioning). The combination of BT partitioning and TT partitioning is called multi-tree partitioning (MT (Multi Tree) partitioning). A node of the tree structure obtained by recursive quadtree partitioning is called a coding node. Intermediate nodes of the quadtree, binary tree, and ternary tree are coding nodes, and the CTU itself is defined as the topmost coding node.

[0053] CT includes, as CT information, a CU split flag (split_cu_flag) indicating whether to perform CU splitting, a QT split flag (qt_split_cu_flag) indicating whether to perform QT splitting, an MT split direction (mtt_split_cu_vertical_flag) indicating the split direction of MT splitting, and an MT split type (mtt_split_cu_binary_flag) indicating the split type of MT splitting. split_cu_flag, qt_split_cu_flag, mtt_split_cu_vertical_flag, and mtt_split_cu_binary_flag are transmitted for each encoding node.

[0054] Trees that differ in luminance and color difference may be used. The type of tree is indicated by treeType. For example, when using a common tree for luminance (Y, cIdx = 0) and color difference (Cb / Cr, cIdx = 1, 2), the common single tree is indicated by treeType = SINGLE_TREE. When using two different trees (DUAL tree) for luminance and color difference, the luminance tree is indicated by treeType = DUAL_TREE_LUMA and the color difference tree is indicated by treeType = DUAL_TREE_CHROMA.

[0055] (Coded unit) FIG. 4 defines a set of data that the moving image decoding apparatus 31 refers to in order to decode the coded unit to be processed. Specifically, the CU is composed of a CU header CUH, prediction parameters, transformation parameters, quantized transform coefficients, etc. The prediction mode, etc. are defined in the CU header.

[0056] The prediction process may be performed in units of CU or in units of sub-CUs obtained by further splitting the CU. When the sizes of the CU and the sub-CU are equal, there is one sub-CU in the CU. When the CU is larger than the size of the sub-CU, the CU is split into sub-CUs. For example, when the CU is 8x8 and the sub-CU is 4x4, the CU is split into four sub-CUs consisting of two horizontal splits and two vertical splits.

[0057] There are two types of prediction (prediction modes), namely intra prediction and inter prediction. Intra prediction is prediction within the same picture, and inter prediction refers to prediction processing performed between different pictures (for example, between display times, between layer images).

[0058] The transformation and quantization processing is performed in units of CUs, but the quantized transform coefficients may be entropy coded in units of sub-blocks such as 4x4.

[0059] (Prediction parameters) The predicted image is derived by prediction parameters associated with the block. Prediction parameters include intra prediction parameters and inter prediction parameters.

[0060] Hereinafter, the prediction parameters of inter prediction will be described. The inter prediction parameters are composed of a prediction list usage flag predFlagL0 and predFlagL1, a reference picture index refIdxL0 and refIdxL1, and a motion vector mvL0 and mvL1. predFlagL0 and predFlagL1 are flags indicating whether the reference picture list (L0 list, L1 list) is used. When the value is 1, the corresponding reference picture list is used. In this specification, when it is described as "a flag indicating whether XX", if the flag is other than 0 (for example, 1), it is the case where XX, and if 0, it is the case where XX is not. In logical negation, logical product, etc., 1 is treated as true and 0 as false (the same applies hereinafter). However, in an actual device or method, other values can also be used as true values and false values.

[0061] Syntax elements for deriving inter prediction parameters include, for example, the affine flag affine_flag used in merge mode, the merge flag merge_flag, the merge index merge_idx, the MMVD flag mmvd_flag, the inter prediction identifier inter_pred_idc for selecting a reference picture used in AMVP mode, the reference picture index refIdxLX, the prediction vector index mvp_LX_idx for deriving a motion vector, the differential vector mvdLX, and the motion vector accuracy mode amvr_mode.

[0062] (Reference Picture List) The reference picture list is a list consisting of reference pictures stored in the reference picture memory 306. FIG. 5 is a conceptual diagram showing an example of a reference picture and a reference picture list. In the conceptual diagram showing an example of the reference picture in FIG. 5, a rectangle represents a picture, an arrow represents a reference relationship between pictures, the horizontal axis represents time, I, P, and B in the rectangle are an intra picture, a single prediction picture, and a dual prediction picture, respectively, and the numbers in the rectangle indicate the decoding order. As shown in the figure, the decoding order of the pictures is I0, P1, B2, B3, B4, and the display order is I0, B3, B2, B4, P1. FIG. 5 shows an example of the reference picture list of the picture B3 (target picture). The reference picture list is a list representing candidates for reference pictures, and one picture (slice) may have one or more reference picture lists. In the example of the figure, the target picture B3 has two reference picture lists, the L0 list RefPicList0 and the L1 list RefPicList1. For each individual CU, refIdxLX specifies which picture in the reference picture list RefPicListX (X = 0 or 1) is actually referenced. The figure is an example with refIdxL0 = 2 and refIdxL1 = 0. Note that LX is a description method used when distinguishing between L0 prediction and L1 prediction is not required. Hereinafter, by replacing LX with L0 and L1, the parameters for the L0 list and the parameters for the L1 list are distinguished.

[0063] (Merge Prediction and AMVP Prediction) The decoding (encoding) method of prediction parameters includes a merge prediction mode and an AMVP (Advanced Motion Vector Prediction) mode, and merge_flag is a flag for identifying them. The merge prediction mode is a mode that derives from the prediction parameters of neighboring blocks that have already been processed without including the prediction list utilization flag predFlagLX, the reference picture index refIdxLX, and the motion vector mvLX in the encoded data. The AMVP mode is a mode that includes inter_pred_idc, refIdxLX, and mvLX in the encoded data. Note that mvLX is encoded as an mvp_LX_idx that identifies the prediction vector mvpLX and a differential vector mvdLX. In addition to the merge prediction mode, there may be an affine prediction mode and an MMVD prediction mode.

[0064] inter_pred_idc is a value indicating the type and number of reference pictures, and takes any one of the values PRED_L0, PRED_L1, and PRED_BI. PRED_L0 and PRED_L1 indicate single prediction using one reference picture managed in the L0 list and the L1 list, respectively. PRED_BI indicates bi-prediction using two reference pictures managed in the L0 list and the L1 list.

[0065] merge_idx is an index indicating which prediction parameter among the prediction parameter candidates (merge candidates) derived from the blocks for which the processing has been completed is to be used as the prediction parameter of the target block.

[0066] (Motion Vector) mvLX indicates the shift amount between blocks on two different pictures. The prediction vector and the differential vector related to mvLX are called mvpLX and mvdLX, respectively.

[0067] (Inter-prediction identifier inter_pred_idc and prediction list utilization flag predFlagLX) The relationship between inter_pred_idc, predFlagL0, and predFlagL1 is as follows and they are mutually convertible.

[0068] inter_pred_idc = (predFlagL1<<1)+predFlagL0 predFlagL0 = inter_pred_idc & 1 predFlagL1 = inter_pred_idc >> 1 Note that the inter-prediction parameter may use the prediction list utilization flag or the inter-prediction identifier. Also, the determination using the prediction list utilization flag may be replaced with the determination using the inter-prediction identifier. Conversely, the determination using the inter-prediction identifier may be replaced with the determination using the prediction list utilization flag.

[0069] (Determination of bi-prediction biPred) The flag biPred indicating whether it is a bi-prediction can be derived depending on whether both of the two prediction list utilization flags are 1. For example, it can be derived by the following formula.

[0070] biPred = (predFlagL0==1 && predFlagL1==1) Alternatively, biPred can also be derived depending on whether the inter-prediction identifier indicates a value that uses two prediction lists (reference pictures). For example, it can be derived by the following formula.

[0071] biPred = (inter_pred_idc==PRED_BI)? 1 : 0 (Configuration of the moving image decoding device) The configuration of the moving image decoding device 31 (Fig. 6) according to this embodiment will be described.

[0072] The moving image decoding device 31 includes an entropy decoding unit 301, a parameter decoding unit (predicted image decoding device) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a predicted image generation unit (predicted image generation device) 308, an inverse quantization / inverse transformation unit 311, an addition unit 312, and a prediction parameter derivation unit 320. Note that, in accordance with the moving image encoding device 11 described later, there is also a configuration in which the moving image decoding device 31 does not include the loop filter 305.

[0073] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit), and the CU decoding unit 3022 further includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, PPS, APS, and slice header (slice information) from the encoded data. The CT information decoding unit 3021 decodes CT from the encoded data. The CU decoding unit 3022 decodes CU from the encoded data. The TU decoding unit 3024 decodes QP update information (quantization correction value) and quantized prediction error (residual_coding) from the encoded data when the TU contains a prediction error.

[0074] The TU decoding unit 3024 decodes QP update information and quantized prediction error from the encoded data when it is not in the skip mode (skip_mode == 0). More specifically, when skip_mode == 0, the TU decoding unit 3024 decodes a flag cu_cbp indicating whether the target block contains a quantized prediction error, and decodes the quantized prediction error when cu_cbp is 1. When cu_cbp does not exist in the encoded data, it is derived as 0.

[0075] The TU decoder 3024 decodes the index mts_idx indicating the conversion basis from the encoded data. Also, the TU decoder 3024 decodes the index stIdx indicating the use of the secondary conversion and the conversion basis from the encoded data. When stIdx is 0, it indicates non - application of the secondary conversion. When stIdx is 1, it indicates one of the conversions in the set (pair) of the secondary conversion bases. When stIdx is 2, it indicates the other conversion in the above pair.

[0076] Also, the TU decoder 3024 may decode the sub - block conversion flag cu_sbt_flag. When cu_sbt_flag is 1, the CU is divided into a plurality of sub - blocks, and only a specific one of the sub - blocks decodes the residual. Further, the TU decoder 3024 may decode the flag cu_sbt_quad_flag indicating whether the number of sub - blocks is 4 or 2, the cu_sbt_horizontal_flag indicating the division direction, and the cu_sbt_pos_flag indicating the sub - block including non - zero conversion coefficients.

[0077] The predicted image generation unit 308 is configured to include an inter - predicted image generation unit 309 and an intra - predicted image generation unit 310.

[0078] The prediction parameter derivation unit 320 is configured to include an inter - prediction parameter derivation unit 303 and an intra - prediction parameter derivation unit 304.

[0079] Also, hereinafter, examples using CTUs and CUs as the processing units are described, but the present invention is not limited to this example, and processing may be performed in units of sub - CUs. Alternatively, CTUs and CUs may be read as blocks, sub - CUs may be read as sub - blocks, and processing may be performed in units of blocks or sub - blocks.

[0080] The entropy decoding unit 301 performs entropy decoding on the encoded stream Te input from the outside to decode individual codes (syntax elements). For entropy encoding, there are a method of performing variable-length encoding of syntax elements using a context (probability model) adaptively selected according to the type of syntax element and the surrounding situation, and a method of performing variable-length encoding of syntax elements using a predefined table or calculation formula. The former, CABAC (Context Adaptive Binary Arithmetic Coding), stores the CABAC state of the context (the type of the dominant symbol (0 or 1) and the probability state index pStateIdx specifying the probability) in the memory. The entropy decoding unit 301 initializes all the CABAC states at the head of a segment (tile, CTU row, slice). The entropy decoding unit 301 converts the syntax element into a binary string (Bin String) and decodes each bit of the Bin String. When using a context, a context index ctxInc is derived for each bit of the syntax element, the bit is decoded using the context, and the CABAC state of the used context is updated. Bits not using a context are decoded with equal probability (EP, bypass), and the derivation of ctxInc and the CABAC state are omitted. The decoded syntax elements include prediction information for generating a predicted image, prediction errors for generating a differential image, and the like.

[0081] The entropy decoding unit 301 outputs the decoded code to the parameter decoding unit 302. The decoded code is, for example, the prediction mode predMode, merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX, amvr_mode, etc. The control of which code to decode is performed based on the instruction of the parameter decoding unit 302.

[0082] (Basic flow) FIG. 7 is a flowchart for explaining the schematic operation of the moving image decoding apparatus 31.

[0083] (S1100: Parameter set information decoding) The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS from the encoded data.

[0084] (S1200: Slice information decoding) The header decoding unit 3020 decodes a slice header (slice information) from the encoded data.

[0085] Hereinafter, the moving image decoding apparatus 31 derives a decoded image of each CTU by repeating the processes of S1300 to S5000 for each CTU included in the target picture.

[0086] (S1300: CTU information decoding) The CT information decoding unit 3021 decodes a CTU from the encoded data.

[0087] (S1400: CT information decoding) The CT information decoding unit 3021 decodes a CT from the encoded data.

[0088] (S1500: CU decoding) The CU decoding unit 3022 performs S1510 and S1520 to decode a CU from the encoded data.

[0089] (S1510: CU information decoding) The CU decoding unit 3022 decodes CU information, prediction information, TU split flag split_transform_flag, CU residual flags cbf_cb, cbf_cr, cbf_luma, etc. from the encoded data.

[0090] (S1520: TU information decoding) When the TU contains a prediction error, the TU decoding unit 3024 decodes QP update information, quantized prediction error, and transform index mts_idx from the encoded data. Note that the QP update information is a difference value from the quantized parameter prediction value qPpred, which is a predicted value of the quantization parameter QP.

[0091] (S2000: Predicted image generation) The predicted image generation unit 308 generates a predicted image based on the prediction information for each block included in the target CU.

[0092] (S3000: Inverse Quantization and Inverse Transformation) The inverse quantization and inverse transformation unit 311 performs inverse quantization and inverse transformation processing for each TU included in the target CU.

[0093] (S4000: Decoded Image Generation) The addition unit 312 generates a decoded image of the target CU by adding the predicted image supplied from the predicted image generation unit 308 and the prediction error supplied from the inverse quantization and inverse transformation unit 311.

[0094] (S5000: Loop Filter) The loop filter 305 applies loop filters such as a deblocking filter, SAO, and ALF to the decoded image to generate a decoded image.

[0095] (Configuration of Inter-Prediction Parameter Derivation Unit) FIG. 9 shows a schematic diagram of the configuration of the inter-prediction parameter derivation unit 303 according to the present embodiment. The inter-prediction parameter derivation unit 303 derives an inter-prediction parameter by referring to the prediction parameters stored in the prediction parameter memory 307 based on the syntax elements input from the parameter decoding unit 302. Further, the inter-prediction parameter is output to the inter-prediction image generation unit 309 and the prediction parameter memory 307. Since the inter-prediction parameter derivation unit 303 and its internal elements, i.e., the AMVP prediction parameter derivation unit 3032, the merge prediction parameter derivation unit 3036, the affine prediction unit 30372, the MMVD prediction unit 30373, the GPM prediction unit 30377, the DMVR unit 30537, and the MV addition unit 3038, are common means in the moving image encoding device and the moving image decoding device, they may be collectively referred to as a motion vector derivation unit (motion vector derivation device).

[0096] The scale parameter derivation unit 30378 derives the horizontal scaling ratio RefPicScale[i][j][0] of the reference picture, the vertical scaling ratio RefPicScale[i][j][1] of the reference picture, and RefPicIsScaled[i][j] indicating whether the reference picture is scaled. Here, i indicates whether the reference picture list is the L0 list or the L1 list, and j is the value of the L0 reference picture list or the L1 reference picture list, and is derived as follows.

[0097] RefPicScale[i][j][0] = ((fRefWidth << 14)+(PicOutputWidthL >> 1)) / PicOutputWidthL RefPicScale[ i ][ j ]

[0001] = ((fRefHeight << 14)+(PicOutputHeightL >> 1)) / PicOutputHeightL RefPicIsScaled[i][j] = (RefPicScale[i][j][0] != (1<<14)) || (RefPicScale[i][j][1] != (1<<14)) Here, the variable PicOutputWidthL is the value when calculating the horizontal scaling ratio when the coded picture is referenced, and the value obtained by subtracting the left and right offset values from the number of horizontal pixels of the luminance of the coded picture is used. The variable PicOutputHeightL is the value when calculating the vertical scaling ratio when the coded picture is referenced, and the value obtained by subtracting the upper and lower offset values from the number of vertical pixels of the luminance of the coded picture is used. The variable fRefWidth is the value of PicOutputWidthL of the reference picture of the reference list value j of list i, and the variable fRefHeight is the value of PicOutputHeightL of the reference picture of the reference picture list value j of list i.

[0098] When the affine_flag is 1, that is, when indicating the affine prediction mode, the affine prediction unit 30372 derives the inter-prediction parameters in units of sub-blocks.

[0099] When the mmvd_flag is 1, that is, when indicating the MMVD prediction mode, the MMVD prediction unit 30373 derives the inter-prediction parameters from the merge candidate derived by the merge prediction parameter derivation unit 3036 and the difference vector.

[0100] When the GPM Flag is 1, that is, when indicating the GPM (Geometric Partitioning Mode) prediction mode, the GPM prediction unit 30377 derives the GPM prediction parameters.

[0101] When the merge_flag is 1, that is, when indicating the merge prediction mode, the merge_idx is derived and output to the merge prediction parameter derivation unit 3036.

[0102] When the merge_flag is 0, that is, when indicating the AMVP prediction mode, the AMVP prediction parameter derivation unit 3032 derives the mvpLX from the inter_pred_idc, refIdxLX or mvp_LX_idx.

[0103] (MV Addition Unit) In the MV addition unit 3038, the derived mvpLX and mvdLX are added to derive the mvLX.

[0104] (Affine Prediction Unit) The affine prediction unit 30372: 1) derives the motion vectors of two control points CP0, CP1 or three control points CP0, CP1, CP2 of the target block, 2) derives the affine prediction parameters of the target block, and 3) derives the motion vectors of each sub-block from the affine prediction parameters.

[0105] In the case of merge affine prediction, the motion vectors cpMvLX[] of each control point CP0, CP1, CP2 are derived from the motion vectors of the adjacent blocks of the target block. In the case of inter-affine prediction, the cpMvLX[] of each control point is derived from the sum of the prediction vectors of each control point CP0, CP1, CP2 and the difference vector mvdCpLX[] derived from the encoded data.

[0106] (Merge prediction) FIG. 10 shows a schematic diagram showing the configuration of the merge prediction parameter derivation unit 3036 according to the present embodiment. The merge prediction parameter derivation unit 3036 includes a merge candidate derivation unit 30361 and a merge candidate selection unit 30362. Note that the merge candidates are configured to include prediction parameters (predFlagLX, mvLX, refIdxLX) and are stored in a merge candidate list. An index is assigned to the merge candidates stored in the merge candidate list according to a predetermined rule.

[0107] The merge candidate derivation unit 30361 derives merge candidates using the motion vectors and refIdxLX of the decoded adjacent blocks as they are. In addition, the merge candidate derivation unit 30361 may apply a spatial merge candidate derivation process, a temporal merge candidate derivation process, a pairwise merge candidate derivation process, and a zero merge candidate derivation process, which will be described later.

[0108] As the spatial merge candidate derivation process, the merge candidate derivation unit 30361 reads out the prediction parameters stored in the prediction parameter memory 307 according to a predetermined rule and sets them as merge candidates. The method of specifying the reference picture is, for example, the prediction parameters related to each of the adjacent blocks (for example, all or part of the blocks adjacent to the left A1, right B1, upper right B0, lower left A0, and upper left B2 of the target block) within a predetermined range from the target block. Each merge candidate is called A1, B1, B0, A0, B2. Here, A1, B1, B0, A0, B2 are motion information derived from blocks including the following coordinates, respectively. The positions of A1, B1, B0, A0, B2 in the arrangement of the merge candidates are shown in the target picture of FIG. 8.

[0109] A1: (xCb - 1, yCb + cbHeight - 1) B1: (xCb + cbWidth - 1, yCb - 1) B0: (xCb + cbWidth, yCb - 1) A0: (xCb - 1, yCb + cbHeight) B2: (xCb - 1, yCb - 1) Let the upper left coordinates of the target block be (xCb, yCb), the width be cbWidth, and the height be cbHeight.

[0110] As a time merge derivation process, as shown in the collocated picture of FIG. 8, the merge candidate derivation unit 30361 reads out the prediction parameters of block C in the reference image including the lower right CBR of the target block or the central coordinates from the prediction parameter memory 307 as the merge candidate Col, and stores it in the merge candidate list mergeCandList[].

[0111] Generally, give priority to block CBR and add it to mergeCandList[]. When CBR does not have a motion vector (for example, an intra prediction block), or when CBR is located outside the picture, add the motion vector of block C to the prediction vector candidates. By adding the motion vectors of collocated blocks with high possibility of different motions as prediction candidates, the options for prediction vectors increase and the coding efficiency improves.

[0112] When ph_temporal_mvp_enabled_flag is 0, or when cbWidth*cbHeight is 32 or less, set the collocated motion vector mvLXCol of the target block to 0, and set the availability flag availableFlagLXCol of the collocated block to 0.

[0113] In other cases (SliceTemporalMvpEnabledFlag is 1), perform the following.

[0114] For example, the merge candidate derivation unit 30361 may derive the position (xColCtr, yColCtr) of C and the position (xColCBr, yColCBr) of CBR using the following equations.

[0115] xColCtr = xCb+(cbWidth>>1) yColCtr = yCb+(cbHeight>>1) xColCBr = xCb+cbWidth yColCBr = yCb+ cbHeight If CBR is available, the merge candidate COL is derived using the motion vector of CBR. If CBR is not available, COL is derived using C. Then, availableFlagLXCol is set to 1. Note that the reference picture may be collocated_ref_idx notified in the slice header.

[0116] The pairwise candidate derivation unit derives the pairwise candidate avgK from the average of the two merge candidates (p0Cand, p1Cand) stored in mergeCandList and stores it in mergeCandList[].

[0117] mvLXavgK[0] = (mvLXp0Cand[0]+mvLXp1Cand[0]) / 2 mvLXavgK[1] = (mvLXp0Cand[1]+mvLXp1Cand[1]) / 2 The merge candidate derivation unit 30361 derives zero merge candidates Z0,..., ZM where refIdxLX is 0...M and both the X and Y components of mvLX are 0, and stores them in the merge candidate list.

[0118] The order of storage in mergeCandList[] is, for example, spatial merge candidates (A1, B1, B0, A0, B2), temporal merge candidate Col, pairwise candidate avgK, zero merge candidate ZK. Note that non-available (blocks are intra-predicted, etc.) reference blocks are not stored in the merge candidate list. i = 0 if( availableFlagA1 ) mergeCandList[ i++ ] = A1 if( availableFlagB1 ) mergeCandList[ i++ ] = B1 if( availableFlagB0 ) mergeCandList[ i++ ] = B0 if( availableFlagA0 ) mergeCandList[ i++ ] = A0 if( availableFlagB2 ) mergeCandList[ i++ ] = B2 if( availableFlagCol ) mergeCandList[ i++ ] = Col if( availableFlagAvgK ) mergeCandList[ i++ ] = avgK if( i < MaxNumMergeCand ) mergeCandList[ i++ ] = ZK The merge candidate selection unit 30362 selects the merge candidate N indicated by merge_idx from among the merge candidates included in the merge candidate list by the following formula.

[0119] N = mergeCandList[merge_idx] Here, N is a label indicating a merge candidate, and takes values such as A1, B1, B0, A0, B2, Col, avgK, ZK. The movement information of the merge candidate indicated by the label N is indicated by (mvLXN[0], mvLXN[0]), predFlagLXN, and refIdxLXN.

[0120] Select (mvLXN[0], mvLXN[0]) that has been selected, and use predFlagLXN and refIdxLXN as the inter-prediction parameters for the target block. The merge candidate selection unit 30362 stores the inter-prediction parameters of the selected merge candidate in the prediction parameter memory 307 and outputs them to the inter-prediction image generation unit 309.

[0121] (DMVR) Next, the DMVR (Decoder side Motion Vector Refinement) process performed by the DMVR unit 30375 will be described. When the merge_flag is 1 or the skip_flag is 1 for the target CU, the DMVR unit 30375 corrects the mvLX of the target CU derived by the merge prediction unit 30374 using the reference image. Specifically, when the prediction parameters derived by the merge prediction unit 30374 are dual prediction, the motion vector is corrected using the prediction images derived from the motion vectors corresponding to the two reference pictures. The corrected mvLX is supplied to the inter-prediction image generation unit 309.

[0122] Also, in the derivation of the flag dmvrFlag that defines whether to perform the DMVR process, one of the multiple conditions for setting dmvrFlag to 1 includes that the value of the above-mentioned RefPicIsScaled[0][refIdxL0] is 0 and the value of RefPicIsScaled[1][refIdxL1] is 0. When the value of dmvrFlag is set to 1, the DMVR process by the DMVR unit 30375 is executed.

[0123] Also, in the derivation of the flag dmvrFlag that defines whether to perform the DMVR process, one of the multiple conditions for setting dmvrFlag to 1 includes that ciip_flag is 0, that is, the IntraInter synthesis process is not applied.

[0124] Also, in deriving a flag dmvrFlag that defines whether to perform DMVR processing, as one of a plurality of conditions for setting dmvrFlag to 1, it includes that a flag luma_weight_l0_flag[i] indicating whether coefficient information of weight prediction for L0 prediction of luminance described later exists is 0, and a value of a flag luma_weight_l1_flag[i] indicating whether coefficient information of weight prediction for L1 prediction of luminance exists is 0. When the value of dmvrFlag is set to 1, DMVR processing by the DMVR unit 30375 is executed.

[0125] Note that, in deriving a flag dmvrFlag that defines whether to perform DMVR processing, as one of a plurality of conditions for setting dmvrFlag to 1, it may include that luma_weight_l0_flag[i] is 0, and the value of luma_weight_l1_flag[i] is 0, and a flag chroma_weight_l0_flag[i] indicating whether coefficient information of weight prediction for L0 prediction of color difference described later exists is 0, and a value of a flag chroma_weight_l1_flag[i] indicating whether coefficient information of weight prediction for L1 prediction of color difference exists is 0. When the value of dmvrFlag is set to 1, DMVR processing by the DMVR unit 30375 is executed.

[0126] (Prof) Also, if the value of RefPicIsScaled[0][refIdxLX] is 1 or the value of RefPicIsScaled[1][refIdxLX] is 1, the value of cbProfFlagLX is set to FALSE. Here, cbProfFlagLX is a flag that defines whether to perform Prediction refinement (PROF) of affine prediction.

[0127] (AMVP prediction) FIG. 10 shows a schematic diagram of the configuration of the AMVP prediction parameter derivation unit 3032 according to the present embodiment. The AMVP prediction parameter derivation unit 3032 includes a vector candidate derivation unit 3033 and a vector candidate selection unit 3034. The vector candidate derivation unit 3033 derives prediction vector candidates from the decoded motion vectors of adjacent blocks stored in the prediction parameter memory 307 based on refIdxLX, and stores them in the prediction vector candidate list mvpListLX[].

[0128] The vector candidate selection unit 3034 selects the motion vector mvpListLX[mvp_LX_idx] indicated by mvp_LX_idx from the prediction vector candidates in mvpListLX[] as mvpLX. The vector candidate selection unit 3034 outputs the selected mvpLX to the MV addition unit 3038.

[0129] (MV addition unit) The MV addition unit 3038 adds mvpLX input from the AMVP prediction parameter derivation unit 3032 and the decoded mvdLX to calculate mvLX. The addition unit 3038 outputs the calculated mvLX to the inter-prediction image generation unit 309 and the prediction parameter memory 307.

[0130] mvLX[0] = mvpLX[0]+mvdLX[0] mvLX[1] = mvpLX[1]+mvdLX[1] (Details classification of sub-block merge) Summarize the types of prediction processes related to sub-block merge. As described above, it is roughly classified into merge prediction and AMVP prediction.

[0131] Merge prediction is further classified as follows.

[0132] · Normal merge prediction (block-based merge prediction) · Sub-block merge prediction Sub-block merge prediction is further classified as follows.

[0133] · Sub-block prediction (ATMVP) · Affine prediction · Inferred affine prediction · Constructed affine prediction On the other hand, AMVP prediction is classified as follows.

[0134] · AMVP (translation) · MVD affine prediction MVD affine prediction is further classified as follows.

[0135] · 4-parameter MVD affine prediction · 6-parameter MVD affine prediction Note that MVD affine prediction refers to affine prediction that decodes and uses a difference vector.

[0136] In sub-block prediction, similar to the temporal merge derivation process, the availability availableFlagSbCol of the collocated sub-block COL of the target sub-block is determined, and if available, prediction parameters are derived. At least when the above-mentioned SliceTemporalMvpEnabledFlag is 0, availableFlagSbCol is set to 0.

[0137] MMVD prediction (Merge with Motion Vector Difference) may be classified as merge prediction or AMVP prediction. In the former case, when merge_flag = 1, mmvd_flag and MMVD-related syntax elements are decoded, and in the latter case, when merge_flag = 0, mmvd_flag and MMVD-related syntax elements are decoded.

[0138] The loop filter 305 is a filter provided within the encoding loop, which removes block distortion and ringing distortion and improves the image quality. The loop filter 305 applies filters such as a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF) to the decoded image of the CU generated by the addition unit 312.

[0139] The reference picture memory 306 stores the decoded image of the CU at a predetermined position for each target picture and each target CU.

[0140] The prediction parameter memory 307 stores prediction parameters at a predetermined position for each CTU or CU. Specifically, the prediction parameter memory 307 stores parameters decoded by the parameter decoding unit 302, parameters derived by the prediction parameter derivation unit 320, and the like.

[0141] The parameters derived by the prediction parameter derivation unit 320 are input to the prediction image generation unit 308. Also, the prediction image generation unit 308 reads a reference picture from the reference picture memory 306. The prediction image generation unit 308 generates a prediction image of a block or sub-block using the parameters and the reference picture (reference picture block) in the prediction mode indicated by predMode. Here, the reference picture block is a set of pixels on the reference picture (usually a rectangle, so it is called a block), and is an area referred to for generating the prediction image.

[0142] (Inter prediction image generation unit 309) When predMode indicates the inter prediction mode, the inter prediction image generation unit 309 generates a prediction image of a block or sub-block by inter prediction using the inter prediction parameters input from the inter prediction parameter derivation unit 303 and the reference picture.

[0143] FIG. 11 is a schematic diagram showing the configuration of the inter-prediction image generation unit 309 included in the prediction image generation unit 308 according to the present embodiment. The inter-prediction image generation unit 309 includes a motion compensation unit (prediction image generation device) 3091 and a synthesis unit 3095. The synthesis unit 3095 includes an IntraInter synthesis unit 30951, a GPM synthesis unit 30952, a BDOF unit 30954, and a weight prediction unit 3094.

[0144] (Motion Compensation) The motion compensation unit 3091 (interpolation image generation unit 3091) generates an interpolation image (motion compensation image) by reading a reference block from the reference picture memory 306 based on the inter-prediction parameters (predFlagLX, refIdxLX, mvLX) input from the inter-prediction parameter derivation unit 303. The reference block is a block at a position shifted by mvLX from the position of the target block on the reference picture RefPicLX specified by refIdxLX. Here, when mvLX is not of integer precision, a filter for generating pixels at fractional positions called a motion compensation filter is applied to generate the interpolation image.

[0145] The motion compensation unit 3091 first derives an integer position (xInt, yInt) and a phase (xFrac, yFrac) corresponding to the coordinates (x, y) within the prediction block using the following equations.

[0146] xInt = xPb+(mvLX[0]>>(log2(MVPREC)))+x xFrac = mvLX[0]&(MVPREC-1) yInt = yPb+(mvLX[1]>>(log2(MVPREC)))+y yFrac = mvLX[1]&(MVPREC-1) Here, (xPb, yPb) is the upper left coordinate of a block of size bW*bH, where x = 0…bW-1 and y = 0…bH-1, and MVPREC indicates the precision of mvLX (1 / MVPREC pixel precision). For example, MVPREC = 16.

[0147] The motion compensation unit 3091 derives a temporary image temp[][] by performing horizontal interpolation processing on the reference picture refImg using an interpolation filter. The following Σ is the sum over k from k = 0..NTAP - 1, shift1 is a normalization parameter for adjusting the value range, and offset1 = 1 << (shift1 - 1).

[0148] temp[x][y] = (ΣmcFilter[xFrac][k]*refImg[xInt + k - NTAP / 2 + 1][yInt]+offset1)>>shift1 Subsequently, the motion compensation unit 3091 derives an interpolated image Pred[][] from the temporary image temp[][] through vertical interpolation processing. The following Σ is the sum over k from k = 0..NTAP - 1, shift2 is a normalization parameter for adjusting the value range, and offset2 = 1 << (shift2 - 1).

[0149] Pred[x][y] = (ΣmcFilter[yFrac][k]*temp[x][y + k - NTAP / 2 + 1]+offset2)>>shift2. In the case of dual prediction, the above Pred[][] is derived for each of the L0 list and L1 list (referred to as the interpolated images PredL0[][] and PredL1[]), and the interpolated image Pred[][] is generated from PredL0[][] and PredL1[]

[0150] Note that the motion compensation unit 3091 has a function of scaling the interpolated image according to the horizontal scaling ratio RefPicScale[i][j][0] of the reference picture derived by the scale parameter derivation unit 30378 and the vertical scaling ratio RefPicScale[i][j][1] of the reference picture.

[0151] The synthesis unit 3095 includes an IntraInter synthesis unit 30951, a GPM synthesis unit 30952, a weight prediction unit 3094, and a BDOF unit 30954.

[0152] (interpolation filter processing) The following describes the interpolation filter process executed by the prediction image generation unit 308, which is the interpolation filter process in the case where the resampling described above is applied and the size of the reference picture changes within a single sequence. Note that this process may be executed by, for example, the motion compensation unit 3091.

[0153] When the value of RefPicIsScaled[i][j] input from the inter-prediction parameter derivation unit 303 indicates that the reference picture is scaled, the prediction image generation unit 308 switches a plurality of filter coefficients and executes the interpolation filter process.

[0154] (IntraInter Synthesis Process) The IntraInter synthesis unit 30951 generates a prediction image by weighted summation of the inter-prediction image and the intra-prediction image.

[0155] The pixel value predSamplesComb[x][y] of the prediction image is derived as follows if the flag ciip_flag indicating whether to apply the IntraInter synthesis process is 1.

[0156] predSamplesComb[x][y] =(w * predSamplesIntra[x][y] +(4 - w)*predSamplesInter[x][y] + 2)>> 2 Here, predSamplesIntra[x][y] is the intra-prediction image and is limited to planar prediction. predSamplesInter[x][y] is the reconstructed inter-prediction image.

[0157] The weight w is derived as follows.

[0158] If both the lowermost block adjacent to the left and the rightmost block adjacent to the top of the target coding block are intra, w is set to 3.

[0159] Otherwise, if both the bottom - most block adjacent to the left and the right - most block adjacent to the top of the target - coded block are not intra, w is set to 1.

[0160] Otherwise, w is set to 2.

[0161] (GPM Synthesis Processing) The GPM synthesis unit 30952 generates a predicted image using the above - described GPM prediction.

[0162] (BDOF Prediction) Next, the details of the BDOF prediction (Bi - Directional Optical Flow, BDOF processing) performed by the BDOF unit 30954 will be described. In the bi - prediction mode, the BDOF unit 30954 generates a predicted image with reference to two predicted images (the first predicted image and the second predicted image) and a gradient correction term.

[0163] (Weight Prediction) The weight prediction unit 3094 generates a predicted image pbSamples of a block from the interpolated image predSamplesLX. First, a variable weightedPredFlag indicating whether to perform weight prediction processing is derived as follows. When slice_type is equal to P, weightedPredFlag is set equal to pps_weighted_pred_flag defined in PPS. Otherwise, when slice_type is equal to B, weightedPredFlag is set equal to pps_weighted_bipred_flag && (!dmvrFlag) defined in PPS.

[0164] Hereafter, bcw_idx is the weight index for bi - prediction with CU - unit weights. If bcw_idx is not notified, set bcw_idx = 0. bcwIdx sets bcwIdxN of neighboring blocks in the merge prediction mode and sets bcw_idx of the target block in the AMVP prediction mode.

[0165] If the value of the variable weightedPredFlag is equal to 0 or the value of the variable bcwIdx is 0, then as normal prediction image processing, the prediction image pbSamples is derived as follows.

[0166] When one of the prediction list usage flags (predFlagL0 or predFlagL1) is 1 (single prediction) (without using weighted prediction), the following processing of the formula for adjusting predSamplesLX (LX is L0 or L1) to the pixel bit depth bitDepth is performed.

[0167] pbSamples[x][y] = Clip3(0,(1<<bitDepth)-1,(predSamplesLX[x][y]+offset1)>>shift1) Here, shift1 = 14 - bitDepth, offset1 = 1<<(shift1 - 1). PredLX is the interpolated image of L0 or L1 prediction.

[0168] Also, when both of the prediction list usage flags (predFlagL0 and predFlagL1) are 1 (dual prediction PRED_BI) and weighted prediction is not used, the following processing of the formula for averaging predSamplesL0 and predSamplesL1 and adjusting to the pixel bit depth is performed.

[0169] pbSamples[x][y] = Clip3(0,(1<<bitDepth)-1,(predSamplesL0[x][y]+predSamplesL1[x][y]+offset2)>>shift2) Here, shift2 = 15 - bitDepth, offset2 = 1<<(shift2 - 1).

[0170] If the value of the variable weightedPredFlag is equal to 1 and the value of the variable bcwIdx is equal to 0, then as weighted prediction processing, the prediction image pbSamples is derived as follows.

[0171] Set the variable shift1 equal to Max(2, 14 - bitDepth). Derive the variables log2Wd, o0, o1, w0, and w1 as follows.

[0172] If cIdx is 0 and it is for luminance, the following applies.

[0173] log2Wd = luma_log2_weight_denom + shift1 w0 = LumaWeightL0[refIdxL0] w1 = LumaWeightL1[refIdxL1] o0 = luma_offset_l0[refIdxL0] <<(bitDepth - 8) o1 = luma_offset_l1[refIdxL1] <<(bitDepth - 8) Otherwise (cIdx is not equal to 0, i.e., for chrominance difference), the following applies.

[0174] log2Wd = ChromaLog2WeightDenom + shift1 w0 = ChromaWeightL0[refIdxL0][cIdx - 1] w1 = ChromaWeightL1[refIdxL1][cIdx - 1] o0 = ChromaOffsetL0[refIdxL0][cIdx - 1] <<(bitDepth - 8) o1 = ChromaOffsetL1[refIdxL1][cIdx - 1] <<(bitDepth - 8) The pixel value pbSamples[x][y] of the predicted image for x = 0..nCbW - 1 and y = 0..nCbH - 1 is derived as follows.

[0175] Next, if predFlagL0 is equal to 1 and predFlagL1 is equal to 0, the pixel value pbSamples[x][y] of the predicted image is derived as follows.

[0176] if(log2Wd >= 1) pbSamples[x][y] = Clip3(0,(1 << bitDepth)- 1, ((predSamplesL0[x][y] * w0 + 2^(log2Wd - 1))>> log2Wd)+ o0) else pbSamples[x][y] = Clip3(0,(1<<bitDepth)-1, predSamplesL0[x][y]*w0 + o0) Otherwise, if predFlagL0 is 0 and predFlagL1 is 1, the pixel value pbSamples[x][y] of the predicted image is derived as follows.

[0177] if(log2Wd >= 1) pbSamples[x][y] = Clip3(0,(1 << bitDepth)- 1, ((predSamplesL1[x][y] * w1 + 2^(log2Wd - 1))>> log2Wd)+ o1) else pbSamples[x][y] = Clip3(0,(1<<bitDepth)-1, predSamplesL1[x][y]*w1 + o1) Otherwise, if predFlagL0 is equal to 1 and predFlagL1 is equal to 1, the pixel value pbSamples[x][y] of the predicted image is derived as follows.

[0178] pbSamples[x][y] = Clip3(0,(1 << bitDepth)- 1, (predSamplesL0[x][y] * w0 + predSamplesL1[x][y] * w1 + ((o0 + o1 + 1)<< log2Wd))>>(log2Wd + 1)) (BCW prediction) BCW (Bi-prediction with CU-level Weights) prediction is a prediction method that can switch pre-determined weight coefficients at the CU level. Input two variables nCbW and nCbH that specify the width and height of the current encoded block, two arrays predSamplesL0 and predSamplesL1 of (nCbW)x(nCbH), flags predFlagL0 and predFlagL1 indicating whether to use the prediction list, reference indices refIdxL0 and refIdxL1, BCW prediction index bcw_idx, and variable cIdx that specifies the indices of the luminance and chrominance components, perform BCW prediction processing, and output the pixel values of the predicted image in the array pbSamples of (nCbW)x(nCbH).

[0179] When the sps_bcw_enabled_flag indicating whether to use this prediction at the SPS level is TRUE, the variable weightedPredFlag is 0, there are no weight prediction coefficients in the reference pictures indicated by the two reference indices refIdxL0 and refIdxL1, and the encoded block size is below a certain value, explicitly notify the bcw_idx of the CU-level syntax and assign its value to the variable bcwIdx. If bcw_idx does not exist, 0 is assigned to the variable bcwIdx.

[0180] When the variable bcwIdx is 0, the pixel values of the predicted image are derived as follows.

[0181] pbSamples[x][y] = Clip3(0,(1 << bitDepth)- 1, (predSamplesL0[x][y] + predSamplesL1[x][y] + offset2)>> shift2) Otherwise (when bcwIdx is not equal to 0), the following applies.

[0182] The variable w1 is set equal to bcwWLut[bcwIdx]. bcwWLut[k] = {4, 5, 3, 10, -2}.

[0183] The variable w0 is set to (8 - w1). Also, the pixel value of the predicted image is derived as follows.

[0184] pbSamples[x][y] = Clip3(0, (1 << bitDepth) - 1, (w0 * predSamplesL0[x][y] + w1 * predSamplesL1[x][y] + offset3)>>(shift2 + 3)) When BCW prediction is used in the AMVP prediction mode, the inter prediction parameter decoding unit 303 decodes bcw_idx and sends it to the BCW unit 30955. Also, when BCW prediction is used in the merge prediction mode, the inter prediction parameter decoding unit 303 decodes the merge index merge_idx, and the merge candidate derivation unit 30361 derives the bcwIdx of each merge candidate. Specifically, the merge candidate derivation unit 30361 uses the weight coefficient of the adjacent block used for deriving the merge candidate as the weight coefficient of the merge candidate used for the target block. That is, in the merge mode, the weight coefficient used in the past is inherited as the weight coefficient of the target block.

[0185] (Intra prediction image generation unit 310) When predMode indicates the intra prediction mode, the intra prediction image generation unit 310 performs intra prediction using the intra prediction parameters input from the intra prediction parameter derivation unit 304 and the reference pixels read from the reference picture memory 306.

[0186] The inverse quantization and inverse transformation unit 311 inverse quantizes the quantized transformation coefficients input from the parameter decoding unit 302 to obtain the transformation coefficients.

[0187] The adder 312 adds the predicted image of the block input from the predicted image generation unit 308 and the prediction error input from the inverse quantization / inverse transformation unit 311 for each pixel to generate a decoded image of the block. The adder 312 stores the decoded image of the block in the reference picture memory 306 and also outputs it to the loop filter 305.

[0188] The inverse quantization / inverse transformation unit 311 inverse quantizes the quantized transform coefficients input from the parameter decoding unit 302 to obtain transform coefficients.

[0189] The adder 312 adds the predicted image of the block input from the predicted image generation unit 308 and the prediction error input from the inverse quantization / inverse transformation unit 311 for each pixel to generate a decoded image of the block. The adder 312 stores the decoded image of the block in the reference picture memory 306 and also outputs it to the loop filter 305.

[0190] (Configuration of the moving image encoding device) Next, the configuration of the moving image encoding device 11 according to the present embodiment will be described. FIG. 12 is a block diagram showing the configuration of the moving image encoding device 11 according to the present embodiment. The moving image encoding device 11 includes a predicted image generation unit 101, a subtraction unit 102, a transform / quantization unit 103, an inverse quantization / inverse transformation unit 105, an adder 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, an encoding parameter determination unit 110, a parameter encoding unit 111, a prediction parameter derivation unit 120, and an entropy encoding unit 104.

[0191] The predicted image generation unit 101 generates a predicted image for each CU. The predicted image generation unit 101 includes the inter-predicted image generation unit 309 and the intra-predicted image generation unit 310 that have already been described, and the description thereof is omitted.

[0192] The subtraction unit 102 subtracts the pixel value of the predicted image of the block input from the predicted image generation unit 101 from the pixel value of the image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the transform / quantization unit 103.

[0193] The conversion / quantization unit 103 calculates conversion coefficients for the prediction error input from the subtraction unit 102 by frequency conversion, and derives quantized conversion coefficients by quantization. The conversion / quantization unit 103 outputs the quantized conversion coefficients to the parameter encoding unit 111 and the inverse quantization / inverse conversion unit 105.

[0194] The inverse quantization / inverse conversion unit 105 is the same as the inverse quantization / inverse conversion unit 311 (Fig. 6) in the moving image decoding device 31, and the description thereof is omitted. The calculated prediction error is output to the addition unit 106.

[0195] The parameter encoding unit 111 includes a header encoding unit 1110, a CT information encoding unit 1111, and a CU encoding unit 1112 (prediction mode encoding unit). The CU encoding unit 1112 further includes a TU encoding unit 1114. Hereinafter, the schematic operations of each module will be described.

[0196] The header encoding unit 1110 performs encoding processing on parameters such as header information, segmentation information, prediction information, and quantized conversion coefficients.

[0197] The CT information encoding unit 1111 encodes QT, MT (BT, TT) segmentation information, etc.

[0198] The CU encoding unit 1112 encodes CU information, prediction information, segmentation information, etc.

[0199] When the TU contains a prediction error, the TU encoding unit 1114 encodes QP update information and the quantized prediction error.

[0200] The CT information encoding unit 1111 and the CU encoding unit 1112 supply syntax elements such as inter-prediction parameters (predMode, merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX), intra-prediction parameters (intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_reminder, intra_chroma_pred_mode), and quantized transform coefficients to the parameter encoding unit 111.

[0201] The entropy encoding unit 104 receives the quantized transform coefficients and encoding parameters (partition information, prediction parameters) from the parameter encoding unit 111. The entropy encoding unit 104 entropy-encodes these to generate and output an encoded stream Te.

[0202] The prediction parameter derivation unit 120 is a means including an inter-prediction parameter encoding unit 112 and an intra-prediction parameter encoding unit 113, and derives intra-prediction parameters and intra-prediction parameters from the parameters input from the encoding parameter determination unit 110. The derived intra-prediction parameters and intra-prediction parameters are output to the parameter encoding unit 111.

[0203] (Configuration of the inter-prediction parameter encoding unit) As shown in FIG. 13, the inter-prediction parameter encoding unit 112 includes a parameter encoding control unit 1121 and an inter-prediction parameter derivation unit 303. The inter-prediction parameter derivation unit 303 has a configuration common to the moving image decoding device. The parameter encoding control unit 1121 includes a merge index derivation unit 11211 and a vector candidate index derivation unit 11212.

[0204] The merge index derivation unit 11211 derives merge candidates and outputs them to the inter prediction parameter derivation unit 303. The vector candidate index derivation unit 11212 derives prediction vector candidates and outputs them to the inter prediction parameter derivation unit 303 and the parameter encoding unit 111.

[0205] (Configuration of Intra Prediction Parameter Encoding Unit 113) As shown in FIG. 14, the intra prediction parameter encoding unit 113 includes a parameter encoding control unit 1131 and an intra prediction parameter derivation unit 304. The intra prediction parameter derivation unit 304 has the same configuration as that of the moving image decoding device.

[0206] The parameter encoding control unit 1131 derives IntraPredModeY and IntraPredModeC. Further, it determines intra_luma_mpm_flag with reference to mpmCandList[]. These prediction parameters are output to the intra prediction parameter derivation unit 304 and the parameter encoding unit 111.

[0207] However, different from the moving image decoding device, the inputs to the inter prediction parameter derivation unit 303 and the intra prediction parameter derivation unit 304 are the encoding parameter determination unit 110 and the prediction parameter memory 108, and are output to the parameter encoding unit 111.

[0208] The addition unit 106 adds the pixel values of the prediction block input from the prediction image generation unit 101 and the prediction error input from the inverse quantization / inverse transformation unit 105 for each pixel to generate a decoded image. The addition unit 106 stores the generated decoded image in the reference picture memory 109.

[0209] The loop filter 107 performs a deblocking filter, SAO, and ALF on the decoded image generated by the addition unit 106. Note that the loop filter 107 does not necessarily include the above three types of filters, and may have a configuration including only the deblocking filter, for example.

[0210] The prediction parameter memory 108 stores the prediction parameters generated by the encoding parameter determination unit 110 at predetermined positions for each target picture and CU.

[0211] The reference picture memory 109 stores the decoded images generated by the loop filter 107 at predetermined positions for each target picture and CU.

[0212] The encoding parameter determination unit 110 selects one set from a plurality of sets of encoding parameters. The encoding parameters are the QT, BT, or TT segmentation information, prediction parameters, or parameters to be encoded generated in relation to these as described above. The prediction image generation unit 101 generates a prediction image using these encoding parameters.

[0213] The encoding parameter determination unit 110 calculates an RD cost value indicating the amount of information and the encoding error for each of the plurality of sets. The RD cost value is, for example, the sum of the amount of code and the value obtained by multiplying the mean squared error by a coefficient λ. The amount of code is the amount of information of the encoded stream Te obtained by entropy encoding the quantization error and the encoding parameters. The mean squared error is the sum of the squares of the prediction errors calculated in the subtraction unit 102. The coefficient λ is a real number greater than zero set in advance. The encoding parameter determination unit 110 selects the set of encoding parameters for which the calculated cost value is the minimum. The encoding parameter determination unit 110 outputs the determined encoding parameters to the parameter encoding unit 111 and the prediction parameter derivation unit 120.

[0214] Note that, a part of the moving image encoding device 11 and the moving image decoding device 31 in the above-described embodiment, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generation unit 308, the inverse quantization / inverse transform unit 311, the addition unit 312, the predicted parameter derivation unit 320, the predicted image generation unit 101, the subtraction unit 102, the transform / quantization unit 103, the entropy encoding unit 104, the inverse quantization / inverse transform unit 105, the loop filter 107, the encoding parameter determination unit 110, the parameter encoding unit 111, and the predicted parameter derivation unit 120 may be realized by a computer. In that case, a program for realizing this control function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to realize it. Here, the "computer system" refers to a computer system built in either the moving image encoding device 11 or the moving image decoding device 31, and includes hardware such as an OS and peripheral devices. Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, etc., and a storage device such as a hard disk built in a computer system. Furthermore, the "computer-readable recording medium" also includes something that holds a program dynamically for a short time, like a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and something that holds a program for a certain time, like a volatile memory inside a computer system serving as a server or a client in that case. Also, the above program may be for realizing a part of the aforementioned functions, and may further be something that can be realized in combination with a program already recorded in a computer system for realizing the aforementioned functions.

[0215] Further, part or all of the moving image encoding device 11 and the moving image decoding device 31 in the above-described embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the moving image encoding device 11 and the moving image decoding device 31 may be individually made into a processor, or part or all of them may be integrated and made into a processor. Further, the method of integrating into a circuit is not limited to an LSI, and it may be realized by a dedicated circuit or a general-purpose processor. Also, when a technology for integrating into a circuit that replaces an LSI appears due to the progress of semiconductor technology, an integrated circuit using such technology may be used.

[0216] As described above, an embodiment of the present invention has been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist of the present invention.

[0217] (Syntax) FIG. 15(a) shows a part of the syntax of the Sequence Paramenter Set (SPS) of Non-Patent Document 1.

[0218] sps_weighted_pred_flag is a flag indicating whether weighted prediction may be applied to a P slice that refers to the SPS. That sps_weighted_pred_flag is equal to 0 indicates that weighted prediction is applied to a P slice that refers to the SPS. That sps_weighted_pred_flag is equal to 0 indicates that weighted prediction is not applied to a P slice that refers to the SPS.

[0219] sps_weighted_bipred_flag is a flag indicating whether weighted prediction may be applied to a B slice that refers to the SPS. That sps_weighted_bipred_flag is equal to 0 indicates that weighted prediction is applied to a B slice that refers to the SPS. That sps_weighted_bipred_flag is equal to 0 indicates that weighted prediction is not applied to a B slice that refers to the SPS.

[0220] The long_term_ref_pics_flag is a flag indicating whether long-term pictures are used. The inter_layer_ref_pics_present_flag is a flag indicating whether inter-layer prediction is used. The sps_idr_rpl_present_flag is a flag indicating whether syntax elements of the reference picture list are present in the slice header of an IDR picture. The sps_idr_rpl_present_flag is a flag indicating whether syntax elements of the reference picture list are present in the slice header of an IDR picture. When the rpl1_same_as_rpl0_flag is 1, it indicates that there is no information for reference picture list 1 and it is the same as num_ref_pic_lists_in_sps[0] and ref_pic_list_struct(0, rplsIdx).

[0221] Figure 15(b) shows a part of the syntax of the Picture Parameter Set (PPS) in Non-Patent Document 1.

[0222] num_ref_idx_default_active_minus1[i] + 1 indicates the value of the variable NumRefIdxActive[0] for a P or B slice when i is 0 and num_ref_idx_active_override_flag is 0. When i is 1, it indicates the value of the variable NumRefIdxActive[1] for a B slice when num_ref_idx_active_override_flag is equal to 0. The value of num_ref_idx_default_active_minus1[i] must be within the range of values from 0 to 14.

[0223] The pps_weighted_pred_flag is a flag indicating whether weighted prediction is applied to a P slice that refers to the PPS. That the pps_weighted_pred_flag is equal to 0 indicates that weighted prediction is not applied to the P slice that refers to the PPS. That the pps_weighted_pred_flag is equal to 1 indicates that weighted prediction is applied to the P slice that refers to the PPS. When the sps_weighted_pred_flag is equal to 0, the weighted prediction unit 3094 sets the value of the pps_weighted_pred_flag to 0. If the pps_weighted_pred_flag does not exist, the value is set to 0.

[0224] The pps_weighted_bipred_flag is a flag indicating whether weighted prediction is applied to a B slice that refers to the PPS. That the pps_weighted_bipred_flag is equal to 0 indicates that weighted prediction is not applied to the B slice that refers to the PPS. That the pps_weighted_bipred_flag is equal to 1 indicates that weighted prediction is applied to the B slice that refers to the PPS. When the sps_weighted_bipred_flag is equal to 0, the weighted prediction unit 3094 sets the value of the pps_weighted_bipred_flag to 0. If the pps_weighted_bipred_flag does not exist, the value is set to 0.

[0225] That the rpl_info_in_ph_flag is equal to 1 indicates that the reference picture list information exists in the picture header. That the rpl_info_in_ph_flag is equal to 0 indicates that the reference picture list information does not exist in the picture header and there may be a slice header.

[0226] When pps_weighted_pred_flag is equal to 1, or pps_weighted_bipred_flag is equal to 1, or rpl_info_in_ph_flag is equal to 1, wp_info_in_ph_flag exists. That wp_info_in_ph_flag is equal to 1 indicates that the prediction weight information pred_weight_table exists in the picture header but does not exist in the slice header. That wp_info_in_ph_flag is equal to 0 indicates that the prediction weight information pred_weight_table does not exist in the picture header and may exist in the slice header. If wp_info_in_ph_flag does not exist, the value of wp_info_in_ph_flag shall be equal to 0.

[0227] Figure 16 shows a part of the syntax of the picture header PH of Non-Patent Document 1.

[0228] When ph_inter_slice_allowed_flag is 0, it indicates that the slice_type of all slices of the picture is 2 (I Slice). When ph_inter_slice_allowed_flag is 1, it indicates that the slice_type of at least one or more slices included in the picture is 0 (B Slice) or 1 (P Slice).

[0229] The ph_temporal_mvp_enabled_flag is a flag indicating whether to use temporal motion vector prediction for inter-prediction of slices associated with PH. When ph_temporal_mvp_enabled_flag is 0, temporal motion vector prediction cannot be used for slices associated with PH. Otherwise (when ph_temporal_mvp_enabled_flag is equal to 1), temporal motion vector prediction can be used for slices associated with PH. If it does not exist, it is assumed that the value of ph_temporal_mvp_enabled_flag is equal to 0. When the reference picture in the DPB does not have the same spatial resolution as the current picture, the value of ph_temporal_mvp_enabled_flag becomes 0. When ph_collocated_from_l0_flag is 1, it indicates that the reference picture used for temporal motion vector prediction is specified using reference picture list 0. When ph_collocated_from_l0_flag is 0, it indicates that the reference picture used for temporal motion vector prediction is specified using reference picture list 1. ph_collocated_ref_idx indicates the index value of the reference picture used for temporal motion vector prediction. When ph_collocated_from_l0_flag is 1, ph_collocated_ref_idx refers to reference picture list 0, and the value of ph_collocated_ref_idx must be within the range from 0 to num_ref_entries[0][RplsIdx[0]] - 1. Also, when ph_collocated_from_l0_flag is 0, ph_collocated_ref_idx refers to reference picture list 1, and the value of ph_collocated_ref_idx must be within the range from 0 to num_ref_entries[1][RplsIdx[1]] - 1. If it does not exist, it is assumed that the value of ph_collocated_ref_idx is equal to 0.

[0230] When ph_inter_slice_allowed_flag is not 0 and pps_weighted_pred_flag is equal to 1, or pps_weighted_bipred_flag is equal to 1, or wp_info_in_ph_flag is equal to 1, the weight prediction information pred_weight_table exists.

[0231] Figure 17(a) shows a part of the syntax of the slice header in Non-Patent Document 1.

[0232] When num_ref_idx_active_override_flag is 1, it indicates that the syntax element num_ref_idx_active_minus1[0] exists in P and B slices, and the syntax element num_ref_idx_active_minus1[1] exists in B slices. When num_ref_idx_active_override_flag is 0, it indicates that the syntax element num_ref_idx_active_minus1[0] does not exist in P and B slices. If it does not exist, it is presumed that the value of num_ref_idx_active_override_flag is equal to 1.

[0233] num_ref_idx_active_minus1[i] is used to derive the number of actually used reference pictures in the reference picture list i. The variable NumRefIdxActive[i], which is the number of actually used reference pictures, is derived in the manner shown in Figure 17(b). The value of num_ref_idx_active_minus1[i] must be a value between 0 and 14 inclusive. When the slice is a B slice, num_ref_idx_active_override_flag is 1, and num_ref_idx_active_minus1[i] does not exist, it is presumed that num_ref_idx_active_minus1[i] is equal to 0.

[0234] When the value of ph_temporal_mvp_enabled_flag is 1 and rpl_info_in_ph_flag is not 1, information regarding temporal motion vector prediction exists in the slice header. At this time, when the slice_type of the slice is equal to B, slice_collocated_from_l0_flag is specified. rpl_info_in_ph_flag is a flag indicating that information regarding the reference picture list exists in the picture header.

[0235] When slice_collocated_from_l0_flag is 1, it indicates that the reference picture used for temporal motion vector prediction is derived from reference picture list 0. When slice_collocated_from_l0_flag is 0, it indicates that the reference picture used for temporal motion vector prediction is derived from reference picture list 1. When slice_type is equal to B or P, ph_temporal_mvp_enabled_flag is equal to 1, and slice_collocated_from_l0_flag does not exist, the following applies. When rpl_info_in_ph_flag is not 1, slice_collocated_from_l0_flag is presumed to be equal to ph_collocated_from_l0_flag. Otherwise (when rpl_info_in_ph_flag is 0 and slice_type is equal to P), the value of slice_collocated_from_l0_flag is presumed to be equal to 1.

[0236] The slice_collocated_ref_idx indicates an index that specifies the reference picture used for temporal motion vector prediction. When slice_type is P, or when slice_type is B and slice_collocated_from_l0_flag is 1, slice_collocated_ref_idx refers to reference picture list 0, and the value of slice_collocated_ref_idx must be no less than 0 and no greater than NumRefIdxActive[0] - 1. When slice_type is B and slice_collocated_from_l0_flag is 0, slice_collocated_ref_idx refers to reference picture list 1, and the value of slice_collocated_ref_idx must be no less than 0 and no greater than NumRefIdxActive[1] - 1. When slice_collocated_ref_idx does not exist, the following applies. If rpl_info_in_ph_flag is 1, it is assumed that the value of slice_collocated_ref_idx is equal to ph_collocated_ref_idx. Otherwise (when rpl_info_in_ph_flag is equal to 0), it is assumed that the value of slice_collocated_ref_idx is equal to 0. Also, the reference picture indicated by slice_collocated_ref_idx must be the same for all slices within the picture. The values of pic_width_in_luma_samples and pic_height_in_luma_samples of the reference picture indicated by slice_collocated_ref_idx are equal to the values of pic_width_in_luma_samples and pic_height_in_luma_samples of the current picture, and RprConstraintsActive[slice_collocated_from_l0_flag? 0:1][slice_collocated_ref_idx] must be equal to 0.

[0237] When wp_info_in_ph_flag is not 1, and pps_weighted_pred_flag is equal to 1 and slice_type is 1 (P Slice), or when pps_weighted_bipred_flag is equal to 1 and slice_type is 0 (B Slice), pred_weight_table is called.

[0238] Figure 17(b) shows the derivation method of the variable NumRefIdxActive[i] in Non-Patent Document 1. For the reference picture list i (=0,1), in the case of a B slice or a P slice and when the reference picture list is 0, if num_ref_idx_active_override_flag is equal to 1, the value obtained by adding 1 to the value of num_ref_idx_active_minus1[i] is assigned to the variable NumRefIdxActive[i]. Otherwise (in the case of a B slice or a P slice with the reference picture list 0 and num_ref_idx_active_override_flag equal to 0), if the value of num_ref_entries[i][RplsIdx[i]] is greater than or equal to the value obtained by adding 1 to num_ref_idx_default_active_minus1[i], the value obtained by adding 1 to num_ref_idx_default_active_minus1[i] is assigned to the variable NumRefIdxActive[i]; otherwise, the value of num_ref_entries[i][RplsIdx[i]] is assigned to the variable NumRefIdxActive[i]. num_ref_idx_default_active_minus1[i] is the default value of the variable NumRefIdxActive[i] defined in the PPS. In the case of an I slice or a P slice with the reference picture list 1, 0 is assigned to the variable NumRefIdxActive[i].

[0239] Figure 18 shows the syntax of the prediction weight information pred_weight_table in Non-Patent Document 1.

[0240] Here, num_l0_weights indicates the number of weights signaled for the entries of reference picture list 0 when wp_info_in_ph_flag is equal to 1. The value of num_l0_weights ranges from 0 to min(15, num_ref_entries[0][RplsIdx[0]]). When wp_info_in_ph_flag is equal to 1, the variable NumWeightsL0 is set equal to num_l0_weights. Otherwise (when wp_info_in_ph_flag is equal to 0), NumWeightsL0 is set to NumRefIdxActive[0]. Here, num_ref_entries[i][RplsIdx[i]] indicates the number of reference pictures in reference picture list i. The variable RplsIdx[i] is an index value indicating a list in which there are multiple reference picture lists i.

[0241] num_l1_weights specifies the number of weights signaled for the entries of reference picture list 1 when both pps_weighted_bipred_flag and wp_info_in_ph_flag are equal to 1. Assume that the value of num_l1_weights ranges from 0 to min(15, num_ref_entries[1][RplsIdx[1]]).

[0242] If pps_weighted_bipred_flag is 0, set the variable NumWeightsL1 to 0. Otherwise, if wp_info_in_ph_flag is 1, assign the value of num_l1_weights to the variable NumWeightsL1. Otherwise, assign NumRefIdxActive[1] to the variable NumWeightsL1.

[0243] luma_log2_weight_denom is the logarithm to the base 2 of the denominator of all luminance weight coefficients. The value of luma_log2_weight_denom must be within the range from 0 to 7. delta_chroma_log2_weight_denom is the difference of the logarithms to the base 2 of the denominators of all chrominance weight coefficients. If delta_chroma_log2_weight_denom does not exist, it is assumed to be equal to 0. The variable ChromaLog2WeightDenom is derived to be equal to luma_log2_weight_denom + delta_chroma_log2_weight_denom, and the value must be within the range from 0 to 7.

[0244] If luma_weight_l0_flag[i] is 1, it indicates that there exists a weight coefficient for the luminance component of L0 prediction. If luma_weight_l0_flag[i] is 0, it indicates that there does not exist a weight coefficient for the luminance component of L0 prediction. If luma_weight_l0_flag[i] does not exist, the weight prediction unit 3094 assumes it to be equal to 0. If chroma_weight_l0_flag[i] is 1, it indicates that there exists a weight coefficient for the chrominance prediction value of L0 prediction. If chroma_weight_l0_flag[i] is 0, it indicates that there does not exist a weight coefficient for the chrominance prediction value of L0 prediction. If chroma_weight_l0_flag[i] does not exist, the weight prediction unit 3094 assumes it to be equal to 0.

[0245] delta_luma_weight_l0[i] is the difference in weight coefficients applied to the luminance prediction value of L0 prediction using RefPicList[0][i]. The variable LumaWeightL0[i] is derived to be equal to (1<<luma_log2_weight_denom)+delta_luma_weight_l0[i]. When luma_weight_l0_flag[i] is equal to 1, the value of delta_luma_weight_l0[i] must be within the range of -128 to 127. When luma_weight_l0_flag[i] is equal to 0, the weight prediction unit 3094 assumes that LumaWeightL0[i] is equal to the power value of 2 to the luma_log2_weight_denom (2^luma_log2_weight_denom).

[0246] luma_offset_l0[i] is the offset value applied to the luminance prediction value of L0 prediction using RefPicList[0][i]. The value of luma_offset_l0[i] must be within the range of -128 to 127. When luma_weight_l0_flag[i] is equal to 0, the weight prediction unit 3094 assumes that luma_offset_l0[i] is equal to 0.

[0247] delta_chroma_weight_l0[i][j] is the difference in the weight coefficients applied to the predicted value of the color difference of L0 prediction using RefPicList0[i] with j = 0 for Cb and j = 1 for Cr. The variable ChromaWeightL0[i][j] is derived to be equal to (1<<ChromaLog2WeightDenom)+delta_chroma_weight_l0[i][j]. When chroma_weight_l0_flag[i] is equal to 1, the value of delta_chroma_weight_l0[i][j] must be in the range from -128 to 127. When chroma_weight_l0_flag[i] is 0, the weight prediction unit 3094 assumes that ChromaWeightL0[i][j] is equal to the power of 2 to the ChromaLog2WeightDenom value (2^ChromaLog2WeightDenom). delta_chroma_offset_l0[i][j] is the difference in the offset values applied to the predicted value of the color difference of L0 prediction using RefPicList0[i] with j = 0 for Cb and j = 1 for Cr. The variable ChromaOffsetL0[i][j] is derived as follows.

[0248] ChromaOffsetL0[i][j] = Clip3(-128,127, (128 + delta_chroma_offset_l0[i][j] - ((128 * ChromaWeightL0[i][j])>> ChromaLog2WeightDenom))) The value of delta_chroma_offset_l0[i][j] must be within the range of -4*128 to 4*127. When chroma_weight_l0_flag[i] is equal to 0, the weight prediction unit 3094 assumes that ChromaOffsetL0[i][j] is equal to 0.

[0249] Interpret by replacing luma_weight_l1_flag[i], chroma_weight_l1_flag[i], delta_luma_weight_l1[i], luma_offset_l1[i], delta_chroma_weight_l1[i][j], and delta_chroma_offset_l1[i][j] with luma_weight_l0_flag[i], chroma_weight_l0_flag[i], delta_luma_weight_l0[i], luma_offset_l0[i], delta_chroma_weight_l0[i][j], and delta_chroma_offset_l0[i][j] respectively, and interpret l0, L0, list0, and List0 by replacing them with l1, l1, list1, and List1 respectively.

[0250] Figure 19(a) shows the syntax of ref_pic_lists() that defines the reference picture list of Non-Patent Document 1. ref_pic_lists() may exist in the picture header or slice header. If rpl_sps_flag[i] is 1, it indicates that the reference picture list i of ref_pic_lists() is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) of the SPS. Here, listIdx is equal to i.

[0251] If rpl_sps_flag[i] is 0, it indicates that the reference picture list i is derived based on ref_pic_list_struct(listIdx, rplsIdx). Here, listIdx is equal to i directly included in ref_pic_lists(). If rpl_sps_flag[i] does not exist, the following applies. If num_ref_pic_lists_in_sps[i] is 0, the value of rpl_sps_flag[i] is presumed to be 0. If num_ref_pic_lists_in_sps[i] is greater than 0, and rpl1_idx_present_flag is 0 and i is equal to 1, the value of rpl_sps_flag[1] is presumed to be equal to rpl_sps_flag[0].

[0252] rpl_idx[i] indicates the index of ref_pic_list_struct(listIdx, rplsIdx). ref_pic_list_struct(listIdx, rplsIdx) is used for the derivation of reference picture i. Here, listIdx is equal to i. If it does not exist, the value of rpl_idx[i] is presumed to be 0. The value of rpl_idx[i] is within the range of 0 or more and num_ref_pic_lists_in_sps[i] - 1 or less. If rpl_sps_flag[i] is 1 and num_ref_pic_lists_in_sps[i] is 1, the value of rpl_idx[i] is presumed to be 0. If rpl_sps_flag[i] is 1 and rpl1_idx_present_flag is 0, the value of rpl_idx[1] is presumed to be equal to rpl_idx[0]. The variable RplsIdx[i] is derived as follows.

[0253] RplsIdx[i]=(rpl_sps_flag[i])? rpl_idx[i]:num_ref_pic_lists_in_sps[i] Figure 19(b) shows the syntax defining the reference picture list structure ref_pic_list_struct(listIdx, rplsIdx) of Non-Patent Document 1.

[0254] The ref_pic_list_struct (listIdx, rplsIdx) may exist in the SPS, picture header, or slice header. Depending on whether the syntax is included in the SPS, picture header, or slice header, the following applies. If it exists in the picture or slice header, the ref_pic_list_struct (listIdx, rplsIdx) indicates the reference picture list listIdx of the current picture (the picture containing the slice). If it exists in the SPS, the ref_pic_list_struct (listIdx, rplsIdx) indicates candidates for the reference picture list listIdx. And the current picture can be referenced by index value from the picture header or slice header to the list of ref_pic_list_struct (listIdx, rplsIdx) included in the SPS.

[0255] Here, num_ref_entries[listIdx][rplsIdx] indicates the number of ref_pic_list_struct (listIdx, rplsIdx). The value of num_ref_entries[listIdx][rplsIdx] takes a value of 0 or more and 13 or less than MaxDpbSize + 13. MaxDpbSize is the number of decoded pictures determined by the profile level.

[0256] The ltrp_in_header_flag[listIdx][rplsIdx] is a flag indicating whether there is a long-term reference picture in the ref_pic_list_struct (listIdx, rplsIdx).

[0257] The inter_layer_ref_pic_flag[listIdx][rplsIdx][i] is a flag indicating whether the i-th of the reference picture list of the ref_pic_list_struct (listIdx, rplsIdx) is inter-layer prediction.

[0258] st_ref_pic_flag[listIdx][rplsIdx][i] is a flag indicating whether the i-th picture in the reference picture list of ref_pic_list_struct(listIdx, rplsIdx) is a short-term reference picture.

[0259] abs_delta_poc_st[listIdx][rplsIdx][i] is a syntax element for deriving the absolute value of the difference in POC of the short-term reference picture.

[0260] strp_entry_sign_flag[listIdx][rplsIdx][i] is a flag for deriving the positive / negative sign.

[0261] rpls_poc_lsb_lt[listIdx][rplsIdx][i] is a syntax element for deriving the POC of the long-term reference picture of the i-th in the reference picture list of ref_pic_list_struct(listIdx, rplsIdx).

[0262] ilrp_idx[listIdx][rplsIdx][i] is a syntax element for deriving the layer information of the reference picture for inter-layer prediction of the i-th in the reference picture list of ref_pic_list_struct(listIdx, rplsIdx).

[0263] As a problem with the method described in Non-Patent Document 1, as shown in Fig. 19(b), there is a point that 0 can be specified as the value of num_ref_entries[listIdx][rplsIdx] in the reference picture list structure ref_pic_list_struct(listIdx, rplsIdx). 0 indicates that the number of reference pictures in the reference picture list listIdx of the pic_list_truct indicated by rplsIdx is 0. num_ref_entries can be specified regardless of the slice_type. In the case of the reference picture list 0 of the P slice or in the case of the B slice, the number of reference pictures existing in the reference picture list is assumed to be at least 1 or more. If 0 is specified, since there is no reference picture, the reference picture becomes indefinite.

[0264] Therefore, in this embodiment, as shown in Fig. 20, the syntax element to be notified is not num_ref_entries[listIdx][rplsIdx] but num_ref_entries_minus1[listIdx][rplsIdx], and num_ref_entries_minus1[listIdx][rplsIdx] takes a value of 0 or more and MaxDpbSize + 14 or less. By doing so, it is possible to prevent the reference picture from becoming indefinite by arbitrarily prohibiting a reference picture number of 0.

[0265] Another problem with the method described in Non-Patent Document 1 is that in the pred_weight_table of FIG. 18, by explicitly describing num_l0_weights and num_l1_weights as syntax, the number of weights of reference picture list 0 and reference picture list 1 is described. At the time when pred_weight_table is called in the picture header, the number of reference pictures of reference picture list i has already been defined by ref_pic_list_struct(listIdx, rplsIdx), and this syntax element is redundant. Also, at the time when pred_weight_table is called in the slice header, the number of reference pictures of reference picture list i has already been defined by NumRefIdxActive[i], and this syntax element is redundant. Therefore, in this embodiment, as shown in FIG. 21(a), before pred_weight_table is called in the picture header, the value of num_ref_entries_minus1[0][RplsIdx[0]] + 1 is assigned to the variable NumWeightsL0. And the variable NumWeightsL1 is assigned the value of num_ref_entries_minus1[1][RplsIdx[1]] + 1 if pps_weighted_bipred_flag is 1, and 0 otherwise. This is because if pps_weighted_bipred_flag is 0, there is no weight prediction for bidirectional prediction. Also, as shown in FIG. 21(b), before pred_weight_table is called in the slice header, the value of the variable NumRefIdxActive[0] is assigned to the variable NumWeightsL0, and the value of the variable NumRefIdxActive[1] is assigned to the variable NumWeightsL1. Incidentally, the value of the variable NumRefIdxActive[1] at the time of a P slice is 0. And as shown in FIG. 22, pred_weight_table can eliminate redundancy by not explicitly describing num_l0_weights and num_l1_weights as syntax and using the variable NumWeightsL0 as the number of weights of reference picture list 0 and the variable NumWeightsL1 as reference picture list 1.

[0266] Figure 23 is an example of another embodiment of the present embodiment. In this example, in pred_wight_table, the variables NumWeightsL0 and NumWeightsL1 are defined. When wp_info_in_ph_flag is equal to 1, the value of num_ref_entries[0][PicRplsIdx[0]] is assigned to the variable NumWeightsL0, and when it is not, the value of the variable NumRefIdxActive[0] is assigned. wp_info_in_ph_flag is a flag indicating that the weight prediction information exists in the picture header.

[0267] Also, when wp_info_in_ph_flag is equal to 1 and pps_weighted_bipred_flag is equal to 1, the value of num_ref_entries[1][PicRplsIdx[1]] is assigned to the variable NumWeightsL1. pps_weighted_bipred_flag is a flag indicating that bidirectional weight prediction is performed. When wp_info_in_ph_flag is equal to 1 and pps_weighted_bipred_flag is 0, 0 is assigned to the variable NumWeightsL1. When wp_info_in_ph_flag is equal to 0, the value of the variable NumRefIdxActive[1] is assigned to NumWeightsL1. By doing so, without explicitly describing num_l0_weights and num_l1_weights in syntax, and using the variable NumWeightsL0 as the number of weights of the reference picture list 0 and the variable NumWeightsL1 as the number of weights of the reference picture list 1, redundancy can be eliminated.

[0268] Another problem with the method described in Non-Patent Document 1 is that the number of active reference pictures is defined in the slice header but not in the picture header.

[0269] Therefore, in another embodiment of the present implementation, as shown in FIG. 24(a), it is possible to define the number of active reference pictures even in the picture header. When ph_inter_slice_allowed_flag is equal to 1 and rpl_info_in_ph_flag is equal to 1, the number of active reference pictures is defined. When ph_inter_slice_allowed_flag is 1, it indicates that a P slice or a B slice may exist within the picture. When rpl_info_in_ph_flag is equal to 1, it indicates that the reference picture list information exists in the picture header.

[0270] ph_num_ref_idx_active_override_flag is a flag indicating whether ph_num_ref_idx_active_minus1[0] and ph_num_ref_idx_active_minus1[1] exist.

[0271] ph_num_ref_idx_active_minus1[i] is a syntax element used to derive the variable NumRefIdxActive[i] for reference picture list i, and its value is between 0 and 14 inclusive.

[0272] ph_collocated_ref_idx indicates the index of the reference picture used for temporal motion vector prediction. When ph_collocated_from_l0_flag is 1, ph_collocated_ref_idx refers to reference picture list 0, and the value of ph_collocated_ref_idx is between 0 and NumRefIdxActive[0] - 1 inclusive. When ph_collocated_from_l0_flag is 0, ph_collocated_ref_idx refers to an entry in reference picture list 1, and the value of ph_collocated_ref_idx is between 0 and NumRefIdxActive[1] - 1 inclusive. If it does not exist, it is assumed that the value of ph_collocated_ref_idx is equal to 0.

[0273] Figure 24(b) shows the method for deriving the variable NumRefIdxActive[i]. For the reference picture list i (=0,1), when ph_num_ref_idx_active_override_flag is 1, if num_ref_entries_minus1[i][RplsIdx[i]] is greater than 0, the variable NumRefIdxActive[i] is assigned the value obtained by adding 1 to the value of ph_num_ref_idx_active_minus1[i]. Otherwise, 1 is assigned. On the other hand, when ph_num_ref_idx_active_override_flag is not 1, if the value of num_ref_entries_minus1[i][RplsIdx[i]] is greater than or equal to the value obtained by adding 1 to num_ref_idx_default_active_minus1[i], the variable NumRefIdxActive[i] is assigned the value obtained by adding 1 to num_ref_idx_default_active_minus1[i]. Otherwise, the variable NumRefIdxActive[i] is assigned the value obtained by adding 1 to num_ref_entries_minus1[i][RplsIdx[i]]. num_ref_idx_default_active_minus1[i] is the default value of the variable NumRefIdxActive[i] defined in the PPS.

[0274] Figure 25(a) shows the syntax of the slice header. In the slice header, when rpl_info_in_ph_flag, which indicates that the reference picture list information exists in the picture header, is not 1, and in the case of a P slice or a B slice, the number of active reference pictures is defined.

[0275] Figure 25(b) shows the method for deriving the variable NumRefIdxActive[i] in this case. For the reference picture list i (=0,1), when rpl_info_in_ph_flag is not 1, and when i is 0 in a B slice or a P slice, the variable NumRefIdxActive[i] is rewritten. If num_ref_idx_active_override_flag is 1, if num_ref_entries_minus1[i][RplsIdx[i]] is greater than 0, the value obtained by adding 1 to the value of num_ref_idx_active_minus1[i] is assigned to the variable NumRefIdxActive[i], otherwise 1 is assigned. When num_ref_idx_active_override_flag is not 1, if the value of num_ref_entries_minus1[i][RplsIdx[i]] is greater than or equal to the value obtained by adding 1 to num_ref_idx_default_active_minus1[i], the value obtained by adding 1 to num_ref_idx_default_active_minus1[i] is assigned to the variable NumRefIdxActive[i], otherwise, the value obtained by adding 1 to num_ref_entries_minus1[i][RplsIdx[i]] is assigned to the variable NumRefIdxActive[i]. When i is 1 in an I slice or a P slice, regardless of the value of rpl_info_in_ph_flag, 0 is assigned to the variable NumRefIdxActive[i]. rpl_info_in_ph_flag is a flag indicating that the reference picture list information exists in the picture header. num_ref_idx_default_active_minus1[i] is the default value of the variable NumRefIdxActive[i] defined in the PPS.

[0276] As shown in FIG. 26, the pred_weight_table can eliminate redundancy by explicitly not describing num_l0_weights and num_l1_weights in syntax, and by setting the variable NumRefIdxActive[0] as the number of weights of reference picture list 0 and the variable NumRefIdxActive[1] as the number of weights of reference picture list 1.

[0277] 〔Application Example〕 The above-described moving image encoding apparatus 11 and moving image decoding apparatus 31 can be mounted and used in various apparatuses that perform transmission, reception, recording, and reproduction of moving images. Note that the moving image may be a natural moving image captured by a camera or the like, or an artificial moving image (including CG and GUI) generated by a computer or the like.

[0278] First, the fact that the above-described moving image encoding apparatus 11 and moving image decoding apparatus 31 can be used for transmission and reception of moving images will be described with reference to FIG. 2.

[0279] PROD_A in FIG. 2 is a block diagram showing the configuration of a transmission apparatus PROD_A equipped with the moving image encoding apparatus 11. As shown in the figure, the transmission apparatus PROD_A includes an encoding unit PROD_A1 that obtains encoded data by encoding a moving image, a modulation unit PROD_A2 that obtains a modulation signal by modulating a carrier wave with the encoded data obtained by the encoding unit PROD_A1, and a transmission unit PROD_A3 that transmits the modulation signal obtained by the modulation unit PROD_A2. The above-described moving image encoding apparatus 11 is used as this encoding unit PROD_A1.

[0280] The transmission apparatus PROD_A may further include a camera PROD_A4 that captures a moving image, a recording medium PROD_A5 that records a moving image, an input terminal PROD_A6 for inputting a moving image from the outside, and an image processing unit A7 that generates or processes an image as a supply source of the moving image input to the encoding unit PROD_A1. In the figure, a configuration in which the transmission apparatus PROD_A includes all of these is illustrated, but a part of them may be omitted.

[0281] Note that the recording medium PROD_A5 may record unencoded moving images, or may record moving images encoded by an encoding method for recording different from the encoding method for transmission. In the latter case, a decoding unit (not shown) for decoding the encoded data read from the recording medium PROD_A5 according to the encoding method for recording may be interposed between the recording medium PROD_A5 and the encoding unit PROD_A1.

[0282] PROD_B in FIG. 2 is a block diagram showing the configuration of a receiving device PROD_B equipped with a moving image decoding device 31. As shown in the figure, the receiving device PROD_B includes a receiving unit PROD_B1 that receives a modulated signal, a demodulating unit PROD_B2 that obtains encoded data by demodulating the modulated signal received by the receiving unit PROD_B1, and a decoding unit PROD_B3 that obtains a moving image by decoding the encoded data obtained by the demodulating unit PROD_B2. The above-described moving image decoding device 31 is used as this decoding unit PROD_B3.

[0283] The receiving device PROD_B may further include a display PROD_B4 that displays a moving image, a recording medium PROD_B5 for recording a moving image, and an output terminal PROD_B6 for outputting a moving image to the outside as a supply destination of the moving image output by the decoding unit PROD_B3. In the figure, a configuration in which the receiving device PROD_B includes all of these is illustrated, but a part of them may be omitted.

[0284] Note that the recording medium PROD_B5 may be for recording unencoded moving images, or may be encoded by an encoding method for recording different from the encoding method for transmission. In the latter case, an encoding unit (not shown) for encoding the moving image obtained from the decoding unit PROD_B3 according to the encoding method for recording may be interposed between the decoding unit PROD_B3 and the recording medium PROD_B5.

[0285] Note that the transmission medium for transmitting the modulation signal may be wireless or wired. Also, the transmission mode for transmitting the modulation signal may be broadcast (here, it refers to a transmission mode where the transmission destination is not specified in advance) or communication (here, it refers to a transmission mode where the transmission destination is specified in advance). That is, the transmission of the modulation signal may be realized by any of wireless broadcast, wired broadcast, wireless communication, and wired communication.

[0286] For example, a broadcasting station (such as broadcasting equipment) / receiving station (such as a television receiver) for terrestrial digital broadcasting is an example of the transmitting device PROD_A / receiving device PROD_B that transmits and receives the modulation signal by wireless broadcast. Also, a broadcasting station (such as broadcasting equipment) / receiving station (such as a television receiver) for cable television broadcasting is an example of the transmitting device PROD_A / receiving device PROD_B that transmits and receives the modulation signal by wired broadcast.

[0287] Also, a server (such as a workstation) / client (such as a television receiver, personal computer, smartphone, etc.) for a VOD (Video On Demand) service or video sharing service using the Internet is an example of the transmitting device PROD_A / receiving device PROD_B that transmits and receives the modulation signal by communication (usually, either wireless or wired is used as the transmission medium in a LAN, and wired is used as the transmission medium in a WAN). Here, personal computers include desktop PCs, laptop PCs, and tablet PCs. Also, smartphones include multifunctional mobile phone terminals.

[0288] Note that in addition to the function of decoding the encoded data downloaded from the server and displaying it on the display, the client of the video sharing service has a function of encoding the moving image captured by the camera and uploading it to the server. That is, the client of the video sharing service functions as both the transmitting device PROD_A and the receiving device PROD_B.

[0289] Next, with reference to FIG. 3, it will be described that the above-described moving image encoding apparatus 11 and moving image decoding apparatus 31 can be used for recording and playing back moving images.

[0290] PROD_C in FIG. 3 is a block diagram showing the configuration of the recording apparatus PROD_C equipped with the above-described moving image encoding apparatus 11. As shown in the figure, the recording apparatus PROD_C includes an encoding unit PROD_C1 that obtains encoded data by encoding a moving image, and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 to the recording medium PROD_M. The above-described moving image encoding apparatus 11 is used as this encoding unit PROD_C1.

[0291] Note that the recording medium PROD_M may be of a type (1) built into the recording apparatus PROD_C, such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), or (2) of a type connected to the recording apparatus PROD_C, such as an SD memory card or a USB (Universal Serial Bus) flash memory, or (3) of a type loaded into a drive device (not shown) built into the recording apparatus PROD_C, such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blu-ray Disc: registered trademark).

[0292] Further, the recording apparatus PROD_C may further include a camera PROD_C3 that captures a moving image, an input terminal PROD_C4 for inputting a moving image from the outside, a receiving unit PROD_C5 for receiving a moving image, and an image processing unit PROD_C6 for generating or processing an image, as a supply source of the moving image input to the encoding unit PROD_C1. In the figure, a configuration in which the recording apparatus PROD_C includes all of these is illustrated, but a part thereof may be omitted.

[0293] Note that the receiving unit PROD_C5 may receive an unencoded moving image, or may receive encoded data encoded by an encoding method for transmission different from the encoding method for recording. In the latter case, a transmission decoder unit (not shown) for decoding the encoded data encoded by the encoding method for transmission may be interposed between the receiving unit PROD_C5 and the encoding unit PROD_C1.

[0294] Examples of such a recording device PROD_C include a DVD recorder, a BD recorder, an HDD (Hard Disk Drive) recorder, etc. (in this case, the input terminal PROD_C4 or the receiving unit PROD_C5 serves as the main source of the moving image). Also, a camcorder (in this case, the camera PROD_C3 serves as the main source of the moving image), a personal computer (in this case, the receiving unit PROD_C5 or the image processing unit C6 serves as the main source of the moving image), a smartphone (in this case, the camera PROD_C3 or the receiving unit PROD_C5 serves as the main source of the moving image), etc. are also examples of such a recording device PROD_C.

[0295] Figure 3 PROD_D is a block diagram showing the configuration of a playback device PROD_D equipped with the above-described moving image decoder device 31. As shown in the figure, the playback device PROD_D includes a reading unit PROD_D1 that reads the encoded data written on the recording medium PROD_M, and a decoding unit PROD_D2 that obtains a moving image by decoding the encoded data read by the reading unit PROD_D1. The above-described moving image decoder device 31 is used as this decoding unit PROD_D2.

[0296] Note that the recording medium PROD_M may be of a type built into the playback device PROD_D, such as an HDD or an SSD, (1), may be of a type connected to the playback device PROD_D, such as an SD memory card or a USB flash memory, (2), or may be of a type loaded into a drive device (not shown) built into the playback device PROD_D, such as a DVD or a BD, (3).

[0297] In addition, the playback device PROD_D may further include a display PROD_D3 for displaying a moving image, an output terminal PROD_D4 for outputting the moving image externally, and a transmission unit PROD_D5 for transmitting the moving image, as destinations for the moving image output by the decoding unit PROD_D2. In the figure, a configuration in which the playback device PROD_D includes all of these is illustrated, but a part of them may be omitted.

[0298] Note that the transmission unit PROD_D5 may transmit an unencoded moving image, or may transmit encoded data encoded by a transmission encoding method different from the recording encoding method. In the latter case, an encoding unit (not shown) for encoding the moving image by the transmission encoding method may be interposed between the decoding unit PROD_D2 and the transmission unit PROD_D5.

[0299] Examples of such a playback device PROD_D include a DVD player, a BD player, an HDD player, etc. (in this case, the output terminal PROD_D4 to which a television receiver or the like is connected becomes the main destination for the moving image). Also, a television receiver (in this case, the display PROD_D3 becomes the main destination for the moving image), digital signage (also referred to as an electronic billboard or an electronic bulletin board, etc., and the display PROD_D3 or the transmission unit PROD_D5 becomes the main destination for the moving image), a desktop PC (in this case, the output terminal PROD_D4 or the transmission unit PROD_D5 becomes the main destination for the moving image), a laptop or tablet PC (in this case, the display PROD_D3 or the transmission unit PROD_D5 becomes the main destination for the moving image), a smartphone (in this case, the display PROD_D3 or the transmission unit PROD_D5 becomes the main destination for the moving image), etc. are also examples of such a playback device PROD_D.

[0300] (Hardware implementation and software implementation) Further, each block of the above-described moving image decoding apparatus 31 and moving image encoding apparatus 11 may be realized hardware-wise by a logic circuit formed on an integrated circuit (IC chip), or may be realized software-wise using a CPU (Central Processing Unit).

[0301] In the latter case, each of the above apparatuses includes a CPU that executes instructions of a program for realizing each function, a ROM (Read Only Memory) that stores the above program, a RAM (Random Access Memory) that expands the above program, a storage device (recording medium) such as a memory that stores the above program and various data, and the like. And the object of the embodiment of the present invention can also be achieved by supplying a recording medium in which program codes (executable format program, intermediate code program, source program) of control programs of each of the above apparatuses, which are software for realizing the above-described functions, are recorded in a computer-readable manner to each of the above apparatuses, and causing the computer (or CPU or MPU) to read and execute the program codes recorded in the recording medium.

[0302] Examples of the recording medium include tapes such as magnetic tapes and cassette tapes, magnetic disks such as floppy (registered trademark) disks / hard disks, and disks including optical disks such as CD-ROM (Compact Disc Read-Only Memory) / MO disks (Magneto-Optical disc) / MD (Mini Disc) / DVD (Digital Versatile Disc: registered trademark) / CD-R (CD Recordable) / Blu-ray Disc (Blu-ray Disc: registered trademark); cards such as IC cards (including memory cards) / optical cards; semiconductor memories such as mask ROM / EPROM (Erasable Programmable Read-Only Memory) / EEPROM (Electrically Erasable and Programmable Read-Only Memory: registered trademark) / flash ROM; or logic circuits such as PLD (Programmable logic device) and FPGA (Field Programmable Gate Array).

[0303] Also, each of the above devices may be configured to be connectable to a communication network, and the above program code may be supplied via the communication network. This communication network only needs to be capable of transmitting the program code and is not particularly limited. For example, the Internet, intranet, extranet, LAN (Local Area Network), ISDN (Integrated Services Digital Network), VAN (Value-Added Network), CATV (Community Antenna television / Cable Television) communication network, virtual private network, telephone line network, mobile communication network, satellite communication network, etc. can be used. Also, the transmission medium constituting this communication network only needs to be a medium capable of transmitting the program code and is not limited to a specific configuration or type. For example, it can be wired such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line carrier, cable TV line, telephone line, ADSL (Asymmetric Digital Subscriber Line) line, etc., or wireless such as IrDA (Infrared Data Association) and infrared rays like a remote control, Bluetooth (registered trademark), IEEE802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance: registered trademark), mobile phone network, satellite line, terrestrial digital broadcast network, etc. Note that the embodiments of the present invention can also be realized in the form of a computer data signal embedded in a carrier wave, in which the above program code is embodied by electronic transmission.

[0304] The embodiments of the present invention are not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. That is, embodiments obtained by combining technical means appropriately modified within the scope shown in the claims are also included in the technical scope of the present invention.

[0305] 〔Summary〕 A moving image decoding apparatus according to an aspect of the present invention has an inter prediction unit that decodes a plurality of reference picture list structures and selects one reference picture list structure from the plurality of reference picture list structures in units of pictures or slices, wherein in the plurality of reference picture list structures, all the number of reference pictures is 1 or more.

[0306] With such a configuration, the reference picture does not become indefinite.

[0307] A moving image decoding apparatus according to an aspect of the present invention When selecting one reference picture list structure from a plurality of reference picture list structures in units of pictures, a flag indicating whether to rewrite the number of reference pictures actually used for prediction is decoded in units of pictures, If the flag is true, a value obtained by subtracting 1 from the number of reference pictures actually used for prediction is decoded, and the number of reference pictures actually used is set, If the flag is false and no rewriting is performed, the number of pictures in the reference picture list structure is compared with the default number of reference pictures actually used for prediction, and the smaller value is used as the number of reference pictures actually used.

[0308] With such a configuration, the number of reference pictures actually used for prediction can be defined even in the picture header.

[0309] A moving image decoding apparatus according to an aspect of the present invention When selecting one reference picture list structure from a plurality of reference picture list structures in units of slices, If it is a B slice or a P slice and the reference picture list is 0, a flag indicating whether to rewrite the number of reference pictures actually used for prediction is decoded in units of slices, If the flag is true, a value obtained by subtracting 1 from the number of reference pictures actually used for prediction is decoded, and the number of reference pictures actually used is set, If the flag is false and no rewriting is to be done, compare the number of pictures in the reference picture list structure with the default number of reference pictures actually used for prediction, and use the smaller value as the number of reference pictures actually used, If it is an I slice or a P slice in reference picture list 1, the number of reference pictures shall be 0.

[0310] With such a configuration, the number of reference pictures actually used for prediction can be correctly defined in the slice header.

[0311] A moving image decoding apparatus according to an aspect of the present invention When selecting one reference picture list structure from a plurality of reference picture list structures on a picture-by-picture basis, determine the number of reference pictures actually used for prediction on a picture-by-picture basis, and decode the weighted prediction information based on that number. When selecting one reference picture list structure from a plurality of reference picture list structures on a slice-by-slice basis, it is characterized by determining the number of reference pictures actually used for prediction on a slice-by-slice basis and decoding the weighted prediction information based on that number.

[0312] With such a configuration, redundancy in the number of weights can be improved.

Industrial Applicability

[0313] Embodiments of the present invention can be suitably applied to a moving image decoding apparatus that decodes encoded data in which image data is encoded, and a moving image encoding apparatus that generates encoded data in which image data is encoded. Further, it can be suitably applied to the data structure of the encoded data generated by the moving image encoding apparatus and referred to by the moving image decoding apparatus.

Explanation of Signs

[0314] 31 Image decoding apparatus 301 Entropy decoding unit 302 Parameter decoding unit 303 Inter-prediction parameter derivation unit 304 Intra Prediction Parameter Derivation Unit 305, 107 Loop Filter 306, 109 Reference Picture Memory 307, 108 Prediction Parameter Memory 308, 101 Predicted Image Generation Unit 309 Inter Prediction Image Generation Unit 310 Intra Prediction Image Generation Unit 311, 105 Inverse Quantization and Inverse Transformation Unit 312, 106 Addition Unit 320 Prediction Parameter Derivation Unit 11 Image Encoding Device 102 Subtraction Unit 103 Transformation and Quantization Unit 104 Entropy Encoding Unit 110 Encoding Parameter Determination Unit 111 Parameter Encoding Unit 112 Inter Prediction Parameter Encoding Unit 113 Intra Prediction Parameter Encoding Unit 120 Prediction Parameter Derivation Unit

Claims

1. A moving image decoding apparatus, comprising: a parameter decoding unit that decodes a first flag indicating whether weight prediction information is in a picture header, a second flag indicating whether weight prediction is applied to a B slice, a first syntax element for deriving a weight coefficient, and a second syntax element indicating the number of entries included in at least reference picture list 1, where a value of the first flag equal to 1 indicates that the weight prediction information exists in the picture header, a value of the first flag equal to 0 indicates that the weight prediction information exists in a slice header, a value of the second flag equal to 1 indicates that the weight prediction is applied to the B slice, and a value of the second flag equal to 0 indicates that the weight prediction is not applied to the B slice; a prediction parameter derivation unit that derives an inter prediction parameter; a motion compensation unit that generates an interpolated image based on the inter prediction parameter and a reference picture; and a weight prediction unit that derives the weight coefficient using the first syntax element and generates a predicted image using the interpolated image and the weight coefficient, wherein the parameter decoding unit does not decode a third syntax element indicating the number of weights; wherein the weight prediction unit sets a variable NumWeightsL1 indicating the number of weights signaled for entries included in reference picture list 1 to be equal to 0 when the value of the first flag is equal to 1 and the value of the second flag is equal to 0; and the weight prediction unit sets a value of the weight coefficient for reference picture list 1 to a power value of 2 of the first syntax element when the value of the variable NumWeightsL1 is equal to 0. A moving image decoding apparatus characterized by the above is provided.

2. The moving image decoding apparatus according to claim 1, wherein the weight prediction unit sets the variable NumWeightsL1 to be equal to 0 when the value of the first flag is equal to 1, the value of the second flag is equal to 1, and the value of the second syntax element is equal to 0.

Citation Information

Patent Citations

  • Signaling of prediction weights in general constraint information of a bitstream

    WO2021167758A1

  • Methods and apparatuses for signaling of syntax elements in video coding

    WO2021207423A1