Moving image decoding device, moving image encoding device, moving image decoding method, and moving image encoding method

By decoding or encoding the marker information of the sequence and image parameter set in the motion picture decoding and encoding device, the structure of the reference image list is derived, which solves the problem of uncertain reference image quantity and achieves deterministic prediction of image generation.

CN121985118APending Publication Date: 2026-05-05SHARP KK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHARP KK
Filing Date
2021-03-26
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In Non-Patent Literature 1, there is an issue of uncertainty in the number of reference images in the management of the reference image list, which leads to uncertainty in the generation of the predicted image.

Method used

By decoding or encoding the marker information in the sequence parameter set and the image parameter set in the motion picture decoding device and the encoding device, the reference image list structure is derived, ensuring that the number of reference images is assumed to be 0 when the first reference image list is not derived, thus avoiding uncertainty about the reference images.

Benefits of technology

This solves the problem of an uncertain list of reference images, ensuring the determinism and accuracy of the generated predicted images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985118A_ABST
    Figure CN121985118A_ABST
Patent Text Reader

Abstract

A video decoding device, a video encoding device, a video decoding method, and a video encoding method, the video decoding device comprising: a parameter decoding unit that decodes (i) one or more reference picture list structures included in a sequence parameter set, (ii) a first flag included in the sequence parameter set, and (ii) a second flag included in the picture parameter set; wherein the first mark indicates whether first reference picture list information exists in a slice header of the picture with the nalunittype being the IDR, and the second mark indicates whether second reference picture list information exists in a picture header; and a prediction parameter derivation unit (i) deriving first reference picture list information using the nalunittype or the first flag and the second flag in the slice header, or deriving second reference picture list information using the second flag in the picture header, and (ii) deriving a reference picture list on the basis of a reference picture list structure.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application "Motion Picture Decoding Apparatus, Motion Picture Encoding Apparatus, Motion Picture Decoding Method and Motion Picture Encoding Method" (application number: 202180025620.9), filed on March 26, 2021. Technical Field

[0002] Embodiments of the present invention relate to a predictive image generation apparatus, a motion picture decoding apparatus, and a motion picture encoding apparatus. Background Technology

[0003] In order to efficiently transmit or record moving images, a moving image encoding device is used to generate encoded data by encoding the moving images, and a moving image decoding device is used to generate decoded images by decoding the encoded data.

[0004] Specific motion picture coding methods include, for example, H.264 / AVC and H.265 / HEVC (High-Efficiency Video Coding).

[0005] In this motion picture coding method, the images (pictures) that constitute the motion picture are managed through a hierarchical structure and encoded / decoded by each CU. The hierarchical structure includes slices obtained by segmenting the image, coding tree units (CTUs) obtained by segmenting the slices, coding units (sometimes also called coding units (CUs)) obtained by segmenting the coding tree units, and transformation units (TUs) obtained by segmenting the coding units.

[0006] Furthermore, in such moving image coding methods, a prediction image is typically generated based on a locally decoded image obtained by encoding / decoding the input image, and the prediction error (sometimes called a "difference image" or "residual image") obtained by subtracting the prediction image from the input image (original image) is encoded. Methods for generating prediction images include inter-frame prediction and intra-frame prediction.

[0007] In addition, non-patent literature 1 can be cited as a technology for motion image encoding and decoding in recent years.

[0008] In Non-Patent Literature 1, the following mechanism was used for managing the reference image list: defining multiple reference image lists and referring to and using these multiple reference image lists. Furthermore, in weighted prediction, the weights were explicitly defined.

[0009] Existing technical documents

[0010] Non-patent literature

[0011] Non-patent document 1: "Versatile Video Coding (Draft 8)", JVET-Q2001-vE, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG11, 2020-03-12 Summary of the Invention

[0012] The problem the invention aims to solve

[0013] However, in Non-Patent Literature 1, there is a problem: in the management of the reference picture list, it is possible to define a reference picture list with a reference picture quantity of 0, therefore, the reference pictures are uncertain.

[0014] Technical solution

[0015] One aspect of the present invention is a motion picture decoding apparatus that derives a predicted image using reference images included in a list of reference images, characterized by comprising:

[0016] The parameter decoding unit decodes (i) one or more reference image list structures and (ii) a first flag and a second flag included in the image parameter set, wherein the first flag indicates whether the first reference image list information exists in the slice header of the image with nal_unit_type as IDR, and the second flag indicates whether the second reference image list information exists in the image header; and

[0017] Prediction parameter derivation section,

[0018] (i) The prediction parameter derivation unit uses nal_unit_type or the first flag and the second flag in the slice header to derive the first reference image list information, or uses the second flag in the image header to derive the second reference image list information.

[0019] (ii) The prediction parameter derivation unit derives the reference image list based on the reference image list structure.

[0020] In the absence of deriving the first list of reference images, the number of entries in the reference image list structure is presumed to be 0.

[0021] One aspect of the present invention is a motion picture coding apparatus that derives a predicted image using reference images included in a list of reference images, characterized by comprising:

[0022] The parameter encoding unit encodes (i) one or more reference image list structures and (ii) a first flag and a second flag included in the image parameter set, wherein the first flag indicates whether the first reference image list information exists in the slice header of the image with nal_unit_type as IDR, and the second flag indicates whether the second reference image list information exists in the image header; and

[0023] Prediction parameter derivation section,

[0024] (i) The prediction parameter derivation unit uses nal_unit_type or the first flag and the second flag in the slice header to derive the first reference image list information, or uses the second flag in the image header to derive the second reference image list information.

[0025] (ii) The prediction parameter derivation unit derives the reference image list based on the reference image list structure.

[0026] In the absence of deriving the first list of reference images, the number of entries in the reference image list structure is presumed to be 0.

[0027] One aspect of the present invention is a motion picture decoding method that uses reference images included in a list of reference images to derive a predicted image. This method is characterized by comprising at least the following steps:

[0028] Decode (i) one or more reference image list structures and (ii) a first flag and a second flag included in the image parameter set, wherein the first flag indicates whether the first reference image list information exists in the slice header of the image with nal_unit_type of IDR, and the second flag indicates whether the second reference image list information exists in the image header; and

[0029] (i) Use nal_unit_type or the first flag and the second flag in the slice header to deduce the first reference image list information, or use the second flag in the image header to deduce the second reference image list information.

[0030] (ii) Derive the reference image list based on the structure of the reference image list.

[0031] In the absence of deriving the first list of reference images, the number of entries in the reference image list structure is presumed to be 0.

[0032] One aspect of the present invention is a motion picture coding method for deriving a predicted image using reference images included in a list of reference images, characterized by comprising at least the following steps:

[0033] Encode (i) one or more reference image list structures and (ii) a first flag and a second flag included in the image parameter set, wherein the first flag indicates whether the first reference image list information exists in the slice header of the image with nal_unit_type of IDR, and the second flag indicates whether the second reference image list information exists in the image header; and

[0034] (i) Use nal_unit_type or the first flag and the second flag in the slice header to deduce the first reference image list information, or use the second flag in the image header to deduce the second reference image list information.

[0035] (ii) Derive the reference image list based on the structure of the reference image list.

[0036] In the absence of deriving the first list of reference images, the number of entries in the reference image list structure is presumed to be 0.

[0037] By adopting this configuration, the reference image will not become uncertain.

[0038] Beneficial effects

[0039] According to one aspect of the present invention, the above-mentioned problems can be solved. Attached Figure Description

[0040] Figure 1 This is a schematic diagram showing the configuration of the image transmission system of this embodiment.

[0041] Figure 2 This diagram illustrates the configuration of a transmitting device equipped with a motion picture encoding apparatus according to this embodiment and a receiving device equipped with a motion picture decoding apparatus. PROD_A represents the transmitting device equipped with the motion picture encoding apparatus, and PROD_B represents the receiving device equipped with the motion picture decoding apparatus.

[0042] Figure 3This diagram illustrates the configuration of a recording apparatus equipped with a motion picture encoding device according to this embodiment and a playback apparatus equipped with a motion picture decoding device. PROD_C represents the recording apparatus equipped with the motion picture encoding device, and PROD_D represents the playback apparatus equipped with the motion picture decoding device.

[0043] Figure 4 It is a diagram representing the hierarchical structure of the encoded stream data.

[0044] Figure 5 This is a concept diagram representing an example of a reference image and a list of reference images.

[0045] Figure 6 This is a schematic diagram showing the configuration of a motion picture decoding device.

[0046] Figure 7 This is a flowchart illustrating the general operation of a motion picture decoding device.

[0047] Figure 8 This is a diagram illustrating the configuration of the merge candidates.

[0048] Figure 9 This is a schematic diagram showing the structure of the inter-frame prediction parameter derivation section.

[0049] Figure 10 This is a schematic diagram showing the structure of the combined prediction parameter derivation unit and the AMVP prediction parameter derivation unit.

[0050] Figure 11 This is a schematic diagram showing the structure of the inter-frame prediction image generation unit.

[0051] Figure 12 This is a block diagram illustrating the structure of a motion picture encoding device.

[0052] Figure 13 This is a schematic diagram showing the structure of the inter-frame prediction parameter coding unit.

[0053] Figure 14 This is a schematic diagram showing the structure of the intra-frame prediction parameter coding unit.

[0054] Figure 15 It is a diagram representing part of the syntax for Sequence Parameter Set (SPS) and Picture Parameter Set (PPS).

[0055] Figure 16 This is a diagram representing part of the syntax for the image header PH.

[0056] Figure 17 This is a diagram representing part of the syntax of the slice header.

[0057] Figure 18This is a graph representing the syntax of the weighted prediction information pred_weight_table.

[0058] Figure 19 This is a diagram representing the syntax of ref_pic_lists() for defining a list of reference images and ref_pic_list_struct(listIdx, rplsIdx) for defining a list of reference images.

[0059] Figure 20 This is a diagram illustrating the syntax of the reference image list structure ref_pic_list_struct(listIdx, rplsIdx) that defines this implementation.

[0060] Figure 21 This is a diagram illustrating the syntax of the image header PH and slice header in this embodiment.

[0061] Figure 22 This is a diagram representing the syntax of the weighted prediction information pred_weight_table in this embodiment.

[0062] Figure 23 This is a diagram representing the syntax of the pred_weight_table, which is another weighted prediction information in this embodiment.

[0063] Figure 24 This is a diagram illustrating the syntax of the image header and the derivation method of the variable NumRefIdxActive[i] in this embodiment.

[0064] Figure 25 This is a diagram illustrating the syntax of the slice header and the derivation of the variable NumRefIdxActive[i] in this embodiment.

[0065] Figure 26 This is a diagram representing the syntax of the pred_weight_table, which is another weighted prediction information in this embodiment.

[0066] Figure 27 This is a diagram illustrating the syntax of the image header and the derivation method of the variable NumRefIdxActive[i] in this embodiment.

[0067] Figure 28 This is a diagram illustrating the syntax of the slice header and the derivation of the variable NumRefIdxActive[i] in this embodiment. Detailed Implementation

[0068] (First Implementation)

[0069] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.

[0070] Figure 1 This is a schematic diagram showing the configuration of the image transmission system 1 of this embodiment.

[0071] Image transmission system 1 is a system that transmits encoded streams obtained by encoding images of different resolutions that have been converted, decodes the transmitted encoded streams, reverse-converts the images to their original resolution, and displays them. Image transmission system 1 is configured to include: a resolution conversion device (resolution conversion unit) 51, a moving image encoding device (image encoding device) 11, a network 21, a moving image decoding device (image decoding device) 31, a resolution inverse conversion device (resolution inverse conversion unit) 61, and a moving image display device (image display device) 41.

[0072] The resolution conversion device 51 converts the resolution of the image T included in the moving image, and supplies a variable resolution moving image signal including images of different resolutions to the image encoding device 11. Furthermore, the resolution conversion device 51 supplies information indicating whether the image resolution conversion is present to the moving image encoding device 11. If the information indicates a resolution conversion, the moving image encoding device sets the resolution conversion information ref_pic_resampling_enabled_flag (described later) to 1 and includes it in the Sequence Parameter Set (SPS) of the encoded data for encoding.

[0073] An image T with its resolution converted is input into the motion picture encoding device 11.

[0074] Network 21 transmits the encoded stream Te generated by the motion picture encoding device 11 to the motion picture decoding device 31. Network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. Network 21 is not necessarily limited to a two-way communication network; it can also be a one-way communication network transmitting broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Furthermore, network 21 can also be replaced by a storage medium containing the encoded stream Te, such as a DVD (Digital Versatile Disc) or a Blu-ray Disc (Blu-ray Disc).

[0075] The motion image decoding device 31 decodes the encoded stream Te transmitted by the network 21, generates a variable resolution decoded image signal, and supplies it to the resolution inverse conversion device 61.

[0076] When the resolution conversion information included in the variable-resolution decoded image signal indicates resolution conversion, the resolution inverse conversion device 61 generates a decoded image signal of the original size by performing inverse conversion on the resolution-converted image.

[0077] The moving image display device 41 displays all or part of one or more decoded images Td shown in the decoded image signal input from the resolution inverse conversion unit. The moving image display device 41 includes, for example, a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. As the form of the display, a fixed type, a mobile type, an HMD (Head Mounted Display), etc. can be cited. In addition, when the moving image decoding device 31 has high processing power, an image with high display quality is displayed, and when it only has low processing power, an image that does not require high processing power and high display ability is displayed. <000^0188>

[0078] <Operator>

[0079] The following describes the operators used in this specification.

[0080] >> is a right shift, << is a left shift, & is a bitwise AND, | is a bitwise OR, |= is an OR assignment operator, and || represents a logical OR.

[0081] x?y:z is a ternary operator that takes y when x is true (other than 0) and takes z when x is false (0).

[0082] Clip3(a, b, c) is a function that limits c to a value between a and b. It returns a when c < a, returns b when c > b, and returns c in other cases (where a <= b). <00^0199>abs(a) is a function that returns the absolute value of a.

[0084] Int(a) is a function that returns the integer value of a.

[0085] floor(a) is a function that returns the largest integer less than or equal to a. }

[0086] [[ID=зо]]ceil(a) is a function that returns the smallest integer greater than or equal to a.

[0087] a / d means a divided by d (discarding the decimal part). ]|END]]

[0088] min(a, b) represents the smaller value of a and b.

[0089] <Structure of the encoded stream Te>

[0090] Before providing a detailed description of the motion picture encoding device 11 and the motion picture decoding device 31 of this embodiment, the data structure of the encoded stream Te generated by the motion picture encoding device 11 and decoded by the motion picture decoding device 31 will be described.

[0091] Figure 4 This is a diagram representing the hierarchical structure of the data in the encoded stream Te. The encoded stream Te exemplarily includes a sequence and multiple images constituting the sequence. Figure 4 The diagram shows an encoded video sequence representing a given sequence SEQ, an encoded picture representing a given picture PICT, an encoded slice representing a given slice S, encoded slice data representing given slice data, coding tree units included in the encoded slice data, and coding units included in the coding tree units.

[0092] (Encoded video sequence)

[0093] In the encoded video sequence, a set of data is specified for the motion picture decoding device 31 to refer to in order to decode the sequence SEQ of the object being processed. For example... Figure 4 As shown, the sequence SEQ includes the video parameter set (VPS), the sequence parameter set (SPS), the picture parameter set (PPS), the adaptation parameter set (APS), the picture (PICT), and the supplemental enhancement information (SEI).

[0094] The Video Parameter Set (VPS) defines a set of common coding parameters for multiple motion pictures in a motion picture composed of multiple layers, as well as a set of coding parameters for the multiple layers included in the motion picture and associated with each layer.

[0095] The Sequence Parameter Set (SPS) specifies a set of encoding parameters that the moving image decoding device 31 refers to for decoding the object sequence. For example, it specifies the width and height of the image. It should be noted that multiple SPSs can exist. In this case, any one of the multiple SPSs is selected from the PPS.

[0096] Here, the Sequence Parameter Set (SPS) includes the following syntax.

[0097] • `ref_pic_resampling_enabled_flag`: This flag specifies whether to use the variable resolution resampling feature when decoding the images included in a single sequence of the reference object SPS. In other words, this flag indicates whether the size of the reference image used in generating the predicted image varies between the images shown in the single sequence. A value of 1 indicates that resampling is applied, while a value of 0 indicates that resampling is not applied.

[0098] • pic_width_max_in_luma_samples: This is a syntax for specifying the width of an image in a single sequence, in units of luminance blocks. Furthermore, the value of this syntax must be non-zero and an integer multiple of Max(8, MinCbSizeY).

[0099] Here, MinCbSizeY is a value determined based on the minimum size of the luminance block.

[0100] • pic_height_max_in_luma_samples: This is a syntax for specifying the height of the image with the maximum height in a single sequence of images, in units of luminance blocks. Furthermore, the value of this syntax must be non-zero and an integer multiple of Max(8, MinCbSizeY).

[0101] • sps_temporal_mvp_enabled_flag: This flag specifies whether temporal motion vector prediction is used when decoding object sequences. A value of 1 indicates that temporal motion vector prediction is used, while a value of 0 indicates that it is not used. Furthermore, specifying this flag can prevent shifts in the referenced coordinate positions, such as when referencing reference images of different resolutions.

[0102] The Picture Parameter Set (PPS) specifies a set of encoding parameters that the moving image decoding device 31 refers to for decoding each picture in the object sequence. These parameters include, for example, a reference value for the quantization width used for picture decoding (pic_init_qp_minus26) and a flag indicating the application of weighted prediction (weighted_pred_flag). It should be noted that multiple PPSs can exist. In this case, any one of the multiple PPSs is selected from the pictures in the object sequence.

[0103] (Encoded slice)

[0104] In the encoded slice, a set of data is specified for the motion picture decoding device 31 to refer to in order to decode the slice S of the object being processed. For example... Figure 4 As shown, a slice includes a slice header and slice data.

[0105] The slice header includes a set of encoded parameters for the moving image decoding device 31 to refer to in order to determine the decoding method for the object slice. The slice type specification information (slice_type) is an example of the encoded parameters included in the slice header.

[0106] Examples of slice types that can be specified by the slice type specification information include: (1) I slices that use only intra-frame prediction during encoding; (2) P slices that use unidirectional prediction (L0 prediction) or intra-frame prediction during encoding; and (3) B slices that use unidirectional prediction (using only L0 prediction of reference image list 0 or only L1 prediction of reference image list 1), bidirectional prediction, or intra-frame prediction during encoding. It should be noted that inter-frame prediction is not limited to unidirectional or bidirectional prediction, and more reference images can be used to generate the prediction image. Hereinafter, the cases referred to as P and B slices refer to slices that include blocks that can use inter-frame prediction.

[0107] It should be noted that the slice header may also include a reference to the image parameter set (PPS) (pic_parameter_set_id).

[0108] (Encoded slice data)

[0109] In the encoded slice data, a set of data is specified for the motion picture decoding device 31 to refer to in order to decode the slice data of the object being processed. For example... Figure 4 As shown in the encoded slice header, the slice data includes CTUs. A CTU is a fixed-size (e.g., 64×64) block that makes up a slice, also known as the Largest Coding Unit (LCU).

[0110] (Coding Tree Unit)

[0111] exist Figure 4 The specification defines a set of data for the motion picture decoding device 31 to reference in decoding the CTU of the processing object. The CTU is divided into coding units (CUs) as the basic unit of encoding processing through recursive quadtree (QT) segmentation, binary tree (BT) segmentation, or ternary tree (TT) segmentation. BT segmentation and TT segmentation are collectively referred to as multi-tree (MT) segmentation. The nodes of the tree structure obtained through recursive quadtree segmentation are called coding nodes. The intermediate nodes of quadtrees, binary trees, and ternary trees are coding nodes, and the CTU itself is defined as the top-level coding node.

[0112] CT includes the following information as CT information: a CU segmentation flag (split_cu_flag) indicating whether CT segmentation is performed, a QT segmentation flag (qt_split_cu_flag) indicating whether QT segmentation is performed, an MT segmentation direction flag (mtt_split_cu_vertical_flag) indicating the segmentation direction of MT segmentation, and an MT segmentation type flag (mtt_split_cu_binary_flag) indicating the segmentation type of MT segmentation. `split_cu_flag`, `qt_split_cu_flag`, `mtt_split_cu_vertical_flag`, and `mtt_split_cu_binary_flag` are transmitted per encoding node.

[0113] Different trees can be used based on luminance and chromatic difference. The `treeType` parameter represents the tree type. For example, when using a generic tree based on luminance (Y, cIdx=0) and chromatic difference (Cb / Cr, cIdx=1, 2), `treeType=SINGLE_TREE` represents a generic single tree. When using two different trees (DUAL trees) based on luminance and chromatic difference, `treeType=DUAL_TREE_LUMA` represents the luminance tree, and `treeType=DUAL_TREE_CHROMA` represents the chromatic difference tree.

[0114] (Encoding unit)

[0115] exist Figure 4 The document specifies a set of data for the motion picture decoding device 31 to refer to in order to decode the encoding unit of the object being processed. Specifically, the CU consists of a CU header CUH, prediction parameters, transform parameters, quantization transform coefficients, etc. The prediction mode, etc., are specified in the CU header.

[0116] Predictive processing can be performed on a per-unit basis (CU) or on a per-unit basis, based on further subdivisions of the CU into sub-CUs. When the size of the CU and its sub-CUs are equal, there is one sub-CU within the CU. When the size of the CU is larger than the size of its sub-CUs, the CU is divided into sub-CUs. For example, if the CU is 8×8 and the sub-CUs are 4×4, the CU is divided into four sub-CUs, comprising two horizontally divided parts and two vertically divided parts.

[0117] There are two types of prediction (prediction modes): intra-frame prediction and inter-frame prediction. Intra-frame prediction is prediction within the same image, while inter-frame prediction refers to prediction processing performed between different images (such as between display times or between layers).

[0118] Transformation / quantization processing is performed in units of CUs, but quantization transform coefficients can also be entropy encoded in units of 4×4 sub-blocks.

[0119] (Prediction parameters)

[0120] The predicted image is derived from the prediction parameters appended to the block. These prediction parameters include those for intra-frame and inter-frame prediction.

[0121] The prediction parameters for inter-frame prediction are explained below. The inter-frame prediction parameters consist of the prediction list using flags predFlagL0 and predFlagL1, reference image indices refIdxL0 and refIdxL1, and motion vectors mvL0 and mvL1. predFlagL0 and predFlagL1 are flags indicating whether the reference image list (L0 list, L1 list) is used; a value of 1 indicates that the corresponding reference image list is used. It should be noted that in this specification, when a flag is referred to as "the flag indicating whether it is ××", a flag other than 0 (e.g., 1) is considered to be ××, and a flag of 0 is considered not to be ××. In logical NOT, logical product, etc., 1 is considered true, and 0 is considered false (the same applies below). However, in actual devices and methods, other values ​​can also be used as true and false values.

[0122] Among the syntax elements used to derive inter-frame prediction parameters are, for example, the affine flag affine_flag, merge flag merge_flag, merge index merge_idx, MMVD flag mmvd_flag used in merge mode, the inter-frame prediction identifier inter_pred_idc used in AMVP mode for selecting a reference image, the reference image index refIdxLX, the prediction vector index mvp_LX_idx used to derive motion vectors, the difference vector mvdLX, and the motion vector precision mode amvr_mode.

[0123] (Refer to the image list)

[0124] The reference image list is a list of reference images stored in the reference image memory 306. Figure 5 This is a concept diagram representing an example of a reference image and a list of reference images. Figure 5 In a conceptual diagram representing a reference image, rectangles represent images, arrows indicate the reference relationships between images, the horizontal axis represents time, and the I, P, and B symbols within the rectangles represent intra-frame images, one-way prediction images, and two-way prediction images, respectively. The numbers within the rectangles indicate the decoding order. Figure 5 As shown, the decoding order of the image is I0, P1, B2, B3, B4, and the display order is I0, B3, B2, B4, P1. Figure 5The image shows an example of a list of reference images for image B3 (the object image). A list of reference images is a list of candidates for reference images; an image (slice) can have more than one list of reference images. Figure 5 In the example, object image B3 has a list of reference images in L0 list RefPicList0 and L1 list RefPicList1. In each CU, refIdxLX specifies which image in the list of reference images RefPicListX (X=0 or 1) is actually referenced. Figure 5 This is an example where refIdxL0=2 and refIdxL1=0. It should be noted that LX is a notation used without distinguishing between L0 and L1 predictions. Below, we will distinguish between parameters for the L0 list and parameters for the L1 list by replacing LX with L0 and L1.

[0125] (Merged forecast and AMVP forecast)

[0126] There are two main methods for decoding (encoding) prediction parameters: merge prediction mode and AMVP (Advanced Motion Vector Prediction) mode. The merge_flag is a flag used to identify them. The merge prediction mode derives the prediction parameters from processed neighboring blocks without including the prediction list in the encoded data using the predFlagLX flag, the reference image index refIdxLX, and the motion vector mvLX. The AMVP mode includes inter_pred_idc, refIdxLX, and mvLX in the encoded data. It should be noted that mvLX is encoded as mvp_LX_idx, which identifies the prediction vector mvpLX, and the difference vector mvdLX. In addition to the merge prediction mode, there are also affine prediction mode and MMVD prediction mode.

[0127] `inter_pred_idc` is a value representing the type and number of reference images, taking any one of `PRED_L0`, `PRED_L1`, or `PRED_BI`. `PRED_L0` and `PRED_L1` represent one-way prediction using a single reference image managed in the L0 and L1 lists, respectively. `PRED_BI` represents two-way prediction using two reference images managed in the L0 and L1 lists.

[0128] merge_idx is an index indicating whether to use any of the prediction parameter candidates (merge candidates) derived from the processed block as the prediction parameter for the object block.

[0129] (Motion vector)

[0130] mvLX represents the shift amount between blocks on two different images. The prediction vector and difference vector related to mvLX are called mvpLX and mvdLX, respectively.

[0131] (The inter-frame prediction identifier inter_pred_idc and the prediction list utilize the predFlagLX)

[0132] The relationships between inter_pred_idc, predFlagL0, and predFlagL1 can be converted to each other as follows.

[0133] inter_pred_idc=(predFlagL1<<1)+predFlagL0

[0134] predFlagL0=inter_pred_idc&1

[0135] predFlagL1=inter_pred_idc>>1

[0136] It should be noted that inter-frame prediction parameters can use either the prediction list utilization flag or the inter-frame prediction identifier. Furthermore, a decision using the prediction list utilization flag can be replaced by a decision using the inter-frame prediction identifier, and vice versa.

[0137] (Determination of bidirectional prediction biPred)

[0138] The flag indicating whether a prediction is bidirectional (biPred) can be derived from both prediction lists by checking if both flags are 1. For example, it can be derived using the following formula.

[0139] biPred=(predFlagL0==1&&predFlagL1==1)

[0140] Alternatively, biPred can also be derived based on whether the inter-frame prediction identifier indicates the use of two prediction lists (see image). For example, it can be derived using the following formula.

[0141] biPred=(inter_pred_idc==PRED_BI)?1:0

[0142] (Composition of a motion picture decoding device)

[0143] The motion image decoding device 31 of this embodiment ( Figure 6 The composition of ) will be explained.

[0144] The motion picture decoding device 31 is configured to include: an entropy decoding unit 301, a parameter decoding unit (predictive image decoding device) 302, a loop filter 305, a reference image memory 306, a prediction parameter memory 307, a prediction image generation unit (predictive image generation device) 308, an inverse quantization / inverse transform unit 311, an adder unit 312, and a prediction parameter derivation unit 320. It should be noted that, according to the motion picture encoding device 11 described later, there is also a configuration in which the motion picture decoding device 31 does not include the loop filter 305.

[0145] The parameter decoding unit 302 also includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit). The CU decoding unit 3022 also includes a TU decoding unit 3024. These can be collectively referred to as decoding modules. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, PPS, and APS, as well as the slice header (slice information), from the encoded data. The CT information decoding unit 3021 decodes the CT from the encoded data. The CU decoding unit 3022 decodes the CU from the encoded data. In cases where the TU includes prediction errors, the TU decoding unit 3024 decodes QP (Quantization Parameter) update information (quantization correction value) and quantization prediction error (residual_coding) from the encoded data.

[0146] When skip mode is not active (skip_mode==0), the TU decoding unit 3024 decodes QP update information and quantization prediction error from the encoded data. More specifically, when skip_mode==0, the TU decoding unit 3024 decodes the flag cu_cbp, which indicates whether quantization prediction error is included in the object block; when cu_cbp is 1, it decodes the quantization prediction error. When cu_cbp is not present in the encoded data, it is derived to be 0.

[0147] The TU decoding unit 3024 decodes the index mts_idx of the transform basis from the encoded data. Furthermore, the TU decoding unit 3024 decodes the index stIdx of the transform basis from the encoded data, representing the use of a quadratic transform. stIdx is 0 if no quadratic transform is applied, 1 if a transform of one of the sets (pairs) of quadratic transform bases is applied, and 2 if a transform of the other of the aforementioned pairs is applied.

[0148] Furthermore, the TU decoding unit 3024 can also decode the sub-block transform flag cu_sbt_flag. When cu_sbt_flag is 1, the CU is divided into multiple sub-blocks, and the residual of only a specific sub-block is decoded. Moreover, the TU decoding unit 3024 can also decode the flags cu_sbt_quad_flag indicating whether the number of sub-blocks is 4 or 2, cu_sbt_horizontal_flag indicating the division direction, and cu_sbt_pos_flag indicating sub-blocks including non-zero transform coefficients.

[0149] The prediction image generation unit 308 is configured to include an inter-frame prediction image generation unit 309 and an intra-frame prediction image generation unit 310.

[0150] The prediction parameter derivation unit 320 is configured to include an inter-frame prediction parameter derivation unit 303 and an intra-frame prediction parameter derivation unit 304.

[0151] Furthermore, examples of using CTU and CU as processing units are described below, but the process is not limited to these examples; processing can also be performed on a sub-CU basis. Alternatively, CTU and CU can be replaced with blocks, and sub-CUs can be replaced with sub-blocks, allowing processing to be performed on a block or sub-block basis.

[0152] The entropy decoding unit 301 performs entropy decoding on the externally input encoded stream Te, decoding each code (syntactic element). Entropy encoding can be performed in two ways: using a context (probability model) appropriately selected based on the type of syntactic element and its surrounding conditions to perform variable-length encoding of syntactic elements; and using a predetermined table or formula to perform variable-length encoding of syntactic elements. The former, CABAC (Context Adaptive Binary Arithmetic Coding), stores the CABAC state of the context (the category of the dominant symbol (0 or 1) and the probability state index pStateIdx with the specified probability) in memory. The entropy decoding unit 301 initializes all CABAC states at the beginning of each segment (tile, CTU line, slice). The entropy decoding unit 301 transforms the syntactic elements into binary strings and decodes each bit of the binary string. When using a context, it derives the context index ctxInc for each bit of the syntactic element, uses the context to decode the bits, and updates the CABAC state of the used context.

[0153] Bits not using context are decoded with equal probability (EP, bypass), omitting ctxInc export and CABAC state. The decoded syntax elements contain prediction information for generating the predicted image and prediction errors for generating the difference image.

[0154] The entropy decoding unit 301 outputs the decoded code to the parameter decoding unit 302. The decoded code may refer to, for example, the prediction mode `predMode`, `merge_flag`, `merge_idx`, `inter_pred_idc`, `refIdxLX`, `mvp_LX_idx`, `mvdLX`, `amvr_mode`, etc. The parameter decoding unit 302 controls which code to decode.

[0155] (Basic process)

[0156] Figure 7 This is a flowchart explaining the general operation of the motion picture decoding device 31.

[0157] (S1100: Parameter Set Information Decoding) The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS from the encoded data.

[0158] (S1200: Slice Information Decoding) The header decoding unit 3020 decodes the slice header (slice information) from the encoded data.

[0159] Hereinafter, the motion picture decoding device 31 derives the decoded image of each CTU by repeatedly performing S1300 to S5000 processing on each CTU included in the object picture.

[0160] (S1300: CTU Information Decoding) The CT information decoding unit 3021 decodes the CTU from the encoded data.

[0161] (S1400: CT Information Decoding) The CT information decoding unit 3021 decodes the CT from the encoded data.

[0162] (S1500: CU Decoding) The CU decoding unit 3022 implements S1510 and S1520 to decode the CU from the encoded data.

[0163] (S1510: CU Information Decoding) The CU decoding unit 3022 decodes CU information, prediction information, TU segmentation flag split_transform_flag, CU residual flags cbf_cb, cbf_cr, cbf_luma, etc. from the encoded data.

[0164] (S1520: TU Information Decoding) When the TU includes prediction error, the TU decoding unit 3024 decodes the QP update information, quantization prediction error, and transform index mts_idx from the encoded data. It should be noted that the QP update information is the difference between the quantization parameter prediction value qPpred and the prediction value of the quantization parameter QP.

[0165] (S2000: Predictive Image Generation) The predictive image generation unit 308 generates a predictive image for each block included in the object CU based on the prediction information.

[0166] (S3000: Inverse quantization / inverse transformation) The inverse quantization / inverse transformation unit 311 performs inverse quantization / inverse transformation processing on each TU included in the target CU.

[0167] (S4000: Decoded Image Generation) The addition unit 312 generates a decoded image of the object CU by adding the predicted image provided by the predicted image generation unit 308 to the prediction error provided by the inverse quantization / inverse transform unit 311.

[0168] (S5000: Loop Filter) The loop filter 305 applies deblocking filtering, SAO (Sample Adaptive Offset), ALF (Adaptive Loop Filter) and other loop filters to the decoded image to generate the decoded image.

[0169] (The structure of the inter-frame prediction parameter derivation section)

[0170] Figure 9 The diagram shows a schematic representation of the configuration of the inter-frame prediction parameter derivation unit 303 in this embodiment. The inter-frame prediction parameter derivation unit 303 derives inter-frame prediction parameters based on the syntax elements input from the parameter decoding unit 302 and referring to prediction parameters stored in the prediction parameter memory 307. Furthermore, the inter-frame prediction parameters are output to the inter-frame prediction image generation unit 309 and the prediction parameter memory 307. The inter-frame prediction parameter derivation unit 303 and its internal components, including the AMVP prediction parameter derivation unit 3032, the merged prediction parameter derivation unit 3036, the affine prediction unit 30372, the MMVD prediction unit 30373, the GPM prediction unit 30377, the DMVR unit 30537, and the MV (Motion Vector) addition unit 3038, are common units in moving image coding devices and moving image decoding devices; therefore, they can also be collectively referred to as motion vector derivation units (motion vector derivation devices).

[0171] The scale parameter derivation unit 30378 derives the horizontal scaling ratio RefPicScale[i][j][0], the vertical scaling ratio RefPicScale[i][j][1], and RefPicIsScaled[i][j], which indicates whether the reference image has been scaled. Here, i represents whether the reference image list is an L0 list or an L1 list, and j is set to the value of the L0 or L1 reference image list, as shown in the derivation below.

[0172] RefPicScale[i][j][0]=

[0173] ((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL

[0174] RefPicScale[i][j][1]=

[0175] ((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL

[0176] RefPicIsScaled[i][j]=

[0177] (RefPicScale[i][j][0]!= (1<<14)) || (RefPicScale[i][j][1]!= (1<<14))

[0178] Here, the variable `PicOutputWidthL` is the value used to calculate the horizontal scaling ratio when referencing the encoded image. It is obtained by subtracting the left and right offset values ​​from the horizontal pixel count of the encoded image's brightness. The variable `PicOutputHeightL` is the value used to calculate the vertical scaling ratio when referencing the encoded image. It is obtained by subtracting the vertical offset values ​​from the vertical pixel count of the encoded image's brightness. The variable `fRefWidth` is set to the value of `PicOutputWidthL` of the reference image in list i, and the variable `fRefHight` is set to the value of `PicOutputHeightL` of the reference image in list i, where `reference image j` is the reference image.

[0179] When affine_flag is 1, indicating affine prediction mode, the affine prediction unit 30372 derives inter-frame prediction parameters in units of sub-blocks.

[0180] When mmvd_flag is 1, indicating MMVD prediction mode, the MMVD prediction unit 30373 derives the inter-frame prediction parameters from the merging candidate and difference vector derived by the merging prediction parameter derivation unit 3036.

[0181] When GPM Flag is 1, which indicates GPM (Geometric Partitioning Mode) prediction mode, GPM prediction unit 30377 derives GPM prediction parameters.

[0182] When merge_flag is 1, indicating the merge prediction mode, merge_idx is derived and output to the merge prediction parameter derivation section 3036.

[0183] When merge_flag is 0, indicating AMVP prediction mode, the AMVP prediction parameter derivation unit 3032 derives mvpLX from inter_pred_idc, refIdxLX, or mvp_LX_idx.

[0184] (MV Addition Department)

[0185] In the MV addition section 3038, the derived mvpLX and mvdLX are added together to derive mvLX.

[0186] (Affine Prediction Department)

[0187] In the affine prediction unit 30372, 1) the motion vectors of two control points CP0, CP1 or three control points CP0, CP1, CP2 of the object block are derived, 2) the affine prediction parameters of the object block are derived, and 3) the motion vectors of each sub-block are derived from the affine prediction parameters.

[0188] In the case of merged affine prediction, the motion vectors cpMvLX[] of each control point CP0, CP1, and CP2 are derived from the motion vectors of the adjacent blocks of the object block. In the case of inter-frame affine prediction, cpMvLX[] of each control point is derived from the sum of the prediction vectors of each control point CP0, CP1, and CP2 and the difference vector mvdCpLX[] derived from the coded data.

[0189] (Combined Forecast)

[0190] Figure 10 The diagram shows a schematic representation of the configuration of the merge prediction parameter derivation unit 3036 in this embodiment. The merge prediction parameter derivation unit 3036 includes a merge candidate derivation unit 30361 and a merge candidate selection unit 30362. It should be noted that the merge candidates are configured to include prediction parameters (predFlagLX, mvLX, refIdxLX), which are stored in a merge candidate list. Indexes are assigned to merge candidates stored in the merge candidate list according to predetermined rules.

[0191] The merge candidate derivation unit 30361 directly derives merge candidates using the motion vectors and refIdxLX of the decoded adjacent blocks. In addition, the merge candidate derivation unit 30361 can apply spatial merge candidate derivation processing, temporal merge candidate derivation processing, paired merge candidate derivation processing, and zero merge candidate derivation processing, which will be described later.

[0192] As part of the spatial merging candidate derivation process, the merging candidate derivation unit 30361 reads the prediction parameters stored in the prediction parameter memory 307 according to a predetermined rule and sets them as merging candidates. In the specified method referring to the image, these are, for example, the prediction parameters of each adjacent block within a predetermined distance from the object block (e.g., all or part of the blocks that are connected to the left A1, right B1, upper right B0, lower left A0, and upper left B2 of the object block, respectively). Each merging candidate is referred to as A1, B1, B0, A0, and B2.

[0193] Here, A1, B1, B0, A0, and B2 are motion information derived from the block containing the following coordinates. Figure 8 The positions of A1, B1, B0, A0, and B2 are shown in the configuration of the merge candidates in the object image.

[0194] A1: (xCb-1, yCb+cbHeight-1)

[0195] B1: (xCb+cbWidth-1, yCb-1)

[0196] B0: (xCb+cbWidth, yCb-1)

[0197] A0: (xCb-1, yCb+cbHeight)

[0198] B2: (xCb-1, yCb-1)

[0199] Set the top-left coordinates of the object block to (xCb, yCb), width to cbWidth, and height to cbHeight.

[0200] As a time merging derivation process, such as Figure 8 As shown in the corresponding image, the merge candidate derivation unit 30361 reads the prediction parameters of the lower right CBR of the object block or the block C in the reference image including the center coordinates from the prediction parameter memory 307 and sets them as merge candidate Col, and stores them in the merge candidate list mergeCandList[].

[0201] Generally, block CBRs are preferentially added to mergeCandList[]. In cases where the CBR does not have motion vectors (e.g., intra-frame prediction blocks) or the CBR is located outside the image, the motion vector of block C is added to the prediction vector candidate. By adding motion vectors of co-position blocks with high probability of different motions as prediction candidates, the options for prediction vectors are increased, thus improving coding efficiency.

[0202] When ph_temporal_mvp_enabled_flag is 0, or when cbWidth*cbHeight is less than 32, set the parimetric motion vector mvLXCol of the object block to 0, and set the availableFlagLXCol of the parimetric block to 0.

[0203] In other cases (where SliceTemporalMvpEnabledFlag is 1), the following processing is performed.

[0204] For example, the position of C (xColCtr, yColCtr) and the position of CBR (xColCBr, yColCBr) can be derived by merging candidate derivation part 30361 through the following formula.

[0205] xColCtr = xCb + (cbWidth >> 1)

[0206] yColCtr = yCb + (cbHeight >> 1)

[0207] xColCBr=xCb+cbWidth

[0208] yColCBr=yCb+cbHeight

[0209] If the CBR is available, the candidate COL for merging is derived using the CBR's motion vector. If the CBR is unavailable, the COL is derived using C. Then, availableFlagLXCol is set to 1. It should be noted that the reference image can also be the collocated_ref_idx notified in the slice header.

[0210] The pairwise candidate derivation part derives the pairwise candidate avgK from the average of the two merge candidates (p0Cand, p1Cand) already stored in mergeCandList, and stores it in mergeCandList[].

[0211] mvLXavgK[0]=(mvLXp0Cand[0]+mvLXp1Cand[0]) / 2

[0212] mvLXavgK[1]=(mvLXp0Cand[1]+mvLXp1Cand[1]) / 2

[0213] The candidate merging derivation unit 30361 derives zero merging candidates Z0, ..., ZM for refIdxLX of 0...M and mvLX whose X and Y components are both 0, and stores them in the candidate merging list.

[0214] The order in which mergeCandList[] is stored is, for example, spatial merge candidates (A1, B1, B0, A0, B2), temporal merge candidates Col, paired candidates avgK, and zero merge candidates ZK. It should be noted that unusable reference blocks (such as intra-frame prediction blocks) are not stored in the merge candidate list.

[0215] i=0

[0216] if (availableFlagA1)

[0217] mergeCandList[i++]=A1

[0218] if (availableFlagB1)

[0219] mergeCandList[i++]=B1

[0220] if (availableFlagB0)

[0221] mergeCandList[i++]=B0

[0222] if (availableFlagA0)

[0223] mergeCandList[i++]=A0

[0224] if (availableFlagB2)

[0225] mergeCandList[i++]=B2

[0226] if (availableFlagCol)

[0227] mergeCandList[i++]=Col

[0228] if (availableFlagAvgK)

[0229] mergeCandList[i++]=avgK

[0230] if (i <MaxNumMergeCand)

[0231] mergeCandList[i++]=ZK

[0232] The merge candidate selection unit 30362 selects the merge candidate N shown as merge_idx from the merge candidates included in the merge candidate list according to the following formula.

[0233] N = mergeCandList[merge_idx]

[0234] Here, N represents the label of the merging candidate, which can be A1, B1, B0, A0, B2, Col, avgK, ZK, etc. The motion information of the merging candidate shown by label N is represented by (mvLXN[0], mvLXN[0]), predFlagLXN, refIdxLXN.

[0235] The selected (mvLXN[0], mvLXN[0]), predFlagLXN, and refIdxLXN are chosen as the inter-frame prediction parameters for the target block. The merging candidate selection unit 30362 stores the selected inter-frame prediction parameters of the merging candidates in the prediction parameter memory 307 and outputs them to the inter-frame prediction image generation unit 309.

[0236] (DMVR)

[0237] Next, the DMVR (Decoder-side Motion Vector Refinement) processing performed by the DMVR unit 30375 will be described. When the merge_flag is 1 or the skip_flag is 1 for an object CU, the DMVR unit 30375 uses reference images to correct the mvLX of the object CU derived by the merging prediction unit 30374. Specifically, when the prediction parameters derived by the merging prediction unit 30374 are bidirectional predictions, the motion vector is corrected using a prediction image derived from the motion vectors corresponding to the two reference images. The corrected mvLX is then supplied to the inter-frame prediction image generation unit 309.

[0238] Furthermore, in the derivation of the flag dmvrFlag that specifies whether DMVR processing is performed, one of the several conditions for setting dmvrFlag to 1 includes the values ​​of RefPicIsScaled[0][refIdxL0] being 0 and RefPicIsScaled[1][refIdxL1] being 0. When the value of dmvrFlag is set to 1, DMVR processing performed by the DMVR unit 30375 is executed.

[0239] Furthermore, in the derivation of the flag dmvrFlag that specifies whether DMVR processing is performed, one of the several conditions for setting dmvrFlag to 1 is that ciip_flag is 0, which means that intra-inter frame synthesis processing is not applied.

[0240] Furthermore, in the derivation of the flag dmvrFlag specifying whether DMVR processing is performed, one of the several conditions for setting dmvrFlag to 1 includes that luma_weight_l0_flag[i], which indicates whether there is weighted prediction coefficient information for L0 luminance prediction (described later), is 0, and that luma_weight_l1_flag[i], which indicates whether there is weighted prediction coefficient information for L1 luminance prediction, is 0. When the value of dmvrFlag is set to 1, DMVR processing performed by the DMVR unit 30375 is executed.

[0241] It should be noted that in the derivation of the flag dmvrFlag specifying whether DMVR processing is performed, one of the conditions for setting dmvrFlag to 1 may include luma_weight_l0_flag[i] being 0, luma_weight_l1_flag[i] being 0, chroma_weight_l0_flag[i] being 0 (indicating the presence of L0 prediction weighted prediction coefficients for color difference as described later), and chroma_weight_l1_flag[i] being 0 (indicating the presence of L1 prediction weighted prediction coefficients for color difference as described later). When the value of dmvrFlag is set to 1, DMVR processing performed by the DMVR unit 30375 is executed.

[0242] (Prof)

[0243] Furthermore, if the value of RefPicIsScaled[0][refIdxLX] is 1 or the value of RefPicIsScaled[1][refIdxLX] is 1, then the value of cbProfFlagLX is set to FALSE. Here, cbProfFlagLX is a flag that specifies whether to perform prediction refinement (PROF) for affine prediction.

[0244] (AMVP Prediction)

[0245] Figure 10 The diagram shows a schematic representation of the configuration of the AMVP prediction parameter derivation unit 3032 in this embodiment. The AMVP prediction parameter derivation unit 3032 includes a vector candidate derivation unit 3033 and a vector candidate selection unit 3034. The vector candidate derivation unit 3033 derives prediction vector candidates based on refIdxLX and the motion vectors of the decoded adjacent blocks stored in the prediction parameter memory 307, and stores them in the prediction vector candidate list mvpListLX[].

[0246] The vector candidate selection unit 3034 selects the motion vector mvpListLX[mvp_LX_idx] shown in the predicted vector candidate mvpListLX[] as mvpLX. The vector candidate selection unit 3034 outputs the selected mvpLX to the MV addition unit 3038.

[0247] (MV Addition Department)

[0248] The MV addition unit 3038 adds the mvpLX input from the AMVP prediction parameter derivation unit 3032 and the decoded mvdLX to calculate mvLX. The addition unit 3038 outputs the calculated mvLX to the inter-frame prediction image generation unit 309 and the prediction parameter memory 307.

[0249] mvLX[0]=mvpLX[0]+mvdLX[0]

[0250] mvLX[1]=mvpLX[1]+mvdLX[1]

[0251] (Detailed classification of sub-block merging)

[0252] The types of prediction processing associated with sub-block merging are summarized. As mentioned above, they can be broadly categorized into merge prediction and AMVP prediction.

[0253] The combined forecasts are further categorized as follows.

[0254] • Regular merge forecast (block-based merge forecast)

[0255] • Sub-block merging prediction

[0256] Sub-block merging predictions are further categorized as follows.

[0257] • Sub-block prediction (ATMVP)

[0258] • Affine prediction

[0259] • Inferred affine prediction

[0260] • Constructed affine prediction

[0261] On the other hand, the AMVP prediction classification is as follows.

[0262] • AMVP (Translation)

[0263] • MVD Affine Prediction

[0264] MVD affine prediction is further classified as follows.

[0265] • 4-parameter MVD affine prediction

[0266] • 6-parameter MVD affine prediction

[0267] It should be noted that MVD affine prediction refers to affine prediction that decodes and uses difference vectors.

[0268] In sub-block prediction, the availability of the corresponding sub-block COL of the target sub-block is determined in the same way as in the time merging derivation process. If it is available, the prediction parameters are derived. At least when SliceTemporalMvpEnabledFlag is 0, availableFlagSbCol is set to 0.

[0269] MMVD prediction (Merge with Motion Vector Difference) can be classified as either merge prediction or AMVP prediction. For the former, when merge_flag=1, the mmvd_flag and the associated MMVD syntax elements are decoded; for the latter, when merge_flag=0, the mmvd_flag and the associated MMVD syntax elements are decoded.

[0270] The loop filter 305 is a filter located within the encoding loop, used to remove block distortion and ringing distortion to improve image quality. The loop filter 305 performs deblocking filtering, sample adaptive offset (SAO), and adaptive loop filtering (ALF) on the decoded image of the CU generated by the adder 312.

[0271] The image storage 306 stores the decoded images of the CU in predetermined locations for each object image and object CU.

[0272] The prediction parameter memory 307 stores the prediction parameters in a predetermined location for each CTU or CU. Specifically, the prediction parameter memory 307 stores parameters decoded by the parameter decoding unit 302 and parameters derived by the prediction parameter derivation unit 320, etc.

[0273] The prediction image generation unit 308 is input with parameters derived by the prediction parameter derivation unit 320. Furthermore, the prediction image generation unit 308 reads a reference image from the reference image memory 306. In the prediction mode indicated by predMode, the prediction image generation unit 308 uses the parameters and the reference image (reference image block) to generate a prediction image of a block or sub-block. Here, a reference image block refers to a set of pixels (usually rectangular, hence called a block) on the reference image, which is the area referenced for generating the prediction image.

[0274] (Inter-frame prediction image generation unit 309)

[0275] When predMode indicates the inter-frame prediction mode, the inter-frame prediction image generation unit 309 uses the inter-frame prediction parameters input from the inter-frame prediction parameter derivation unit 303 and a reference image to generate a predicted image of a block or sub-block through inter-frame prediction.

[0276] Figure 11 This is a schematic diagram showing the configuration of the inter-frame prediction image generation unit 309 included in the prediction image generation unit 308 of this embodiment. The inter-frame prediction image generation unit 309 is configured to include a motion compensation unit (prediction image generation device) 3091 and a synthesis unit 3095. The synthesis unit 3095 is configured to include an intra-frame / inter-frame synthesis unit 30951, a GPM synthesis unit 30952, a BDOF (Bi-Directional Optical Flow) unit 30954, and a weighted prediction unit 3094.

[0277] (Motion compensation)

[0278] The motion compensation unit 3091 (interpolation image generation unit 3091) generates an interpolated image (motion-compensated image) by reading a reference block from the reference image memory 306 based on the inter-frame prediction parameters (predFlagLX, refIdxLX, mvLX) input from the inter-frame prediction parameter derivation unit 303. The reference block is a block on the reference image RefPicLX specified by refIdxLX that has been shifted by mvLX from the position of the target block. Here, if mvLX is not of integer precision, a filter called motion compensation filtering is implemented to generate pixels at fractional positions, thus generating the interpolated image.

[0279] The motion compensation unit 3091 first derives the integer position (xInt, yInt) and phase (xFrac, yFrac) corresponding to the coordinates (x, y) in the prediction block using the following formula.

[0280] xInt=xPb+(mvLX[0]>>(log2(MVPREC)))+x

[0281] xFrac=mvLX[0]&(MVPREC-1)

[0282] yInt=yPb+(mvLX[1]>>(log2(MVPREC)))+y

[0283] yFrac=mvLX[1]&(MVPREC-1)

[0284] Here, (xPb, yPb) are the top-left coordinates of a block of size bW*bH, x=0……bW-1, y=0……bH-1, and MVPREC represents the precision of mvLX (1 / MVPREC pixel precision). For example, MVPREC=16.

[0285] The motion compensation unit 3091 uses an interpolation filter to perform horizontal interpolation on the reference image refImg, thereby deriving a temporary image temp[][]. The following Σ is the sum of k related to k=0..NTAP-1, shift1 is the normalization parameter for the interval of adjustment values, offset1=1<<(shift1-1).

[0286] temp[x][y]=(ΣmcFilter[xFrac][k]*refImg[xInt+k-NTAP / 2+1][yInt]+offset1)>>shift1

[0287] Next, the motion compensation unit 3091 derives the interpolated image Pred[][] by performing vertical interpolation on the temporary image temp[][]. The following Σ is the sum related to k from k=0..NTAP-1, shift2 is the normalization parameter for the interval of adjustment values, and offset2=1<<(shift2-1).

[0288] Pred[x][y]=(ΣmcFilter[yFrac][k]*temp[x][y+k-NTAP / 2+1]+offset2)>>shift2

[0289] It should be noted that in the case of bidirectional prediction, the above Pred[][] (referred to as the interpolated images PredL0[][] and PredL1[][]) are derived according to each L0 list and L1 list, and the interpolated image Pred[][] is generated based on PredL0[][] and PredL1[][].

[0290] It should be noted that the motion compensation unit 3091 has the function of scaling the interpolated image based on the horizontal scaling ratio RefPicScale[i][j][0] and the vertical scaling ratio RefPicScale[i][j][1] of the reference image derived by the scale parameter derivation unit 30378.

[0291] The compositing unit 3095 includes: an intra-frame / inter-frame compositing unit 30951, a GPM compositing unit 30952, a weighted prediction unit 3094, and a BDOF unit 30954.

[0292] (Interpolation filter processing)

[0293] The interpolation filter processing performed by the predictive image generation unit 308, specifically the interpolation filter processing where the size of the reference image varies within a single sequence due to the aforementioned resampling, will be described below. It should be noted that this processing can, for example, be performed by the motion compensation unit 3091.

[0294] When the value of RefPicIsScaled[i][j] input from the inter-frame prediction parameter derivation unit 303 indicates that the reference image is scaled, the prediction image generation unit 308 switches multiple filter coefficients and performs interpolation filter processing.

[0295] (Intra-frame and inter-frame compositing)

[0296] The intra-frame / inter-frame synthesis unit 30951 generates a prediction image by weighted sum of the inter-frame prediction image and the intra-frame prediction image.

[0297] For the pixel value predSamplesComb[x][y] of the predicted image, if the flag ciip_flag indicating whether intra-frame inter-frame synthesis processing is applied is 1, the derivation is as follows.

[0298] predSamplesComb[x][y]=(w*predSamplesIntra[x][y]

[0299] +(4-w)*predSamplesInter[x][y]+2)>>2

[0300] Here, predSamplesIntra[x][y] represents the intra-frame predicted image, not limited to planar prediction. predSamplesInter[x][y] represents the reconstructed inter-frame predicted image.

[0301] The weight w is derived as shown below.

[0302] When both the bottommost block adjacent to the left of the object coding block and the rightmost block adjacent to the top of the object coding block are within the same frame, w is set to 3.

[0303] In other cases, w is set to 1 if neither the bottommost block adjacent to the left of the object coding block nor the rightmost block adjacent to the top of the object coding block is within the same frame.

[0304] In all other cases, w is set to 2.

[0305] (GPM synthesis treatment)

[0306] The GPM synthesis unit 30952 generates a predicted image using the aforementioned GPM prediction.

[0307] (BDOF prediction)

[0308] Next, the details of BDOF prediction (Bi-Directional Optical Flow, BDOF processing) performed by BDOF unit 30954 will be explained. In bi-directional prediction mode, BDOF unit 30954 generates a prediction image by referring to two prediction images (a first prediction image and a second prediction image) and a gradient correction term.

[0309] (Weighted Prediction)

[0310] The weighted prediction unit 3094 generates prediction images pbSamples for the blocks from the interpolated image predSamplesLX.

[0311] First, the variable `weightedPredFlag`, which indicates whether weighted prediction processing is performed, is derived as follows: When `slice_type` equals P, `weightedPredFlag` is set to be equal to `pps_weighted_pred_flag` defined by PPS. Otherwise, when `slice_type` equals B, `weightedPredFlag` is set to be equal to `pps_weighted_bipred_flag` && (!dmvrFlag) defined by PPS.

[0312] Here, bcw_idx is the weight index for bidirectional prediction with weights in CU units. It is set to bcw_idx=0 if no notification is given. bcwIdx is set to bcwIdxN for nearby blocks in merge prediction mode and to bcw_idx for the target block in AMVP prediction mode.

[0313] If the value of the variable weightedPredFlag is equal to 0 or the value of the variable bcwIdx is 0, the predicted images pbSamples are derived as follows for the usual predicted image processing.

[0314] In the case where one of the flags used in the prediction list (predFlagL0 or predFlagL1) is 1 (one-way prediction) (without using weighted prediction), the following formula is performed to match predSamplesLX (LX is L0 or L1) with the pixel bit depth bitDepth.

[0315] pbSamples[x][y]=Clip3(0,(1<<bitDepth)-1,(predSamplesLX[x][y]+offset1)> >shift1)

[0316] Here, shift1 = 14-bit Depth, offset1 = 1 << (shift1 - 1). PredLX is the interpolated image predicted by L0 or L1.

[0317] Furthermore, when the prediction list uses flags (predFlagL0 and predFlagL1) set to 1 (bidirectional prediction PRED_BI) and does not use weighted prediction, the following formula is used to average predSamplesL0 and predSamplesL1 and match the pixel bit depth.

[0318] pbSamples[x][y]=Clip3(0,(1<<bitDepth)-1,(predSamplesL0[x][y]+predSamplesL1[x][y]+offset2)> >shift2)

[0319] Here, shift2 = 15-bitDepth, offset2 = 1 << (shift2 - 1).

[0320] If the value of the variable weightedPredFlag is equal to 1 and the value of the variable bcwIdx is equal to 0, the predicted images pbSamples are derived as follows for weighted prediction processing.

[0321] Let the variable shift1 be equal to Max(2, 14-bitDepth). The variables log2Wd, o0, o1, w0, and w1 are derived as follows.

[0322] If cIdx is 0 and represents the brightness, apply the following formula.

[0323] log2Wd=luma_log2_weight_denom+shift1

[0324] w0=LumaWeightL0[refIdxL0]

[0325] w1 = LumaWeightL1[refIdxL1]

[0326] o0=luma_offset_l0[refIdxL0]<<(bitDepth-8)

[0327] o1=luma_offset_l1[refIdxL1]<<(bitDepth-8)

[0328] In other cases (color difference where cIdx is not equal to 0), the following formula shall be applied.

[0329] log2Wd=ChromaLog2WeightDenom+shift1

[0330] w0=ChromaWeightL0[refIdxL0][cIdx-1]

[0331] w1=ChromaWeightL1[refIdxL1][cIdx-1]

[0332] o0=ChromaOffsetL0[refIdxL0][cIdx-1]<<(bitDepth-8)

[0333] o1=ChromaOffsetL1[refIdxL1][cIdx-1]<<(bitDepth-8)

[0334] The pixel values ​​pbSamples[x][y] of the predicted images for x=0..nCbW-1 and y=0..nCbH-1 are derived as follows.

[0335] Next, if predFlagL0 equals 1 and predFlagL1 equals 0, the pixel values ​​pbSamples[x][y] of the predicted image are derived as follows.

[0336] if (log2Wd>=1)

[0337] pbSamples[x][y]=Clip3(0,(1< <bitDepth)-1,

[0338] ((predSamplesL0[x][y]*w0+2^(log2Wd-1))>>log2Wd)+o0)

[0339] else

[0340] pbSamples[x][y]=Clip3(0,(1< <bitDepth)-1,predSamplesL0[x][y]*w0+o0)

[0341] In addition, if predFlagL0 is 0 and predFlagL1 is 1, the pixel values ​​pbSamples[x][y] of the predicted image are derived as follows.

[0342] if (log2Wd>=1)

[0343] pbSamples[x][y]=Clip3(0,(1< <bitDepth)-1,

[0344] ((predSamplesL1[x][y]*w1+2^(log2Wd-1))>>log2Wd)+o1)

[0345] else

[0346] pbSamples[x][y]=Clip3(0,(1< <bitDepth)-1,predSamplesL1[x][y]*w1+o1)

[0347] In addition, if predFlagL0 equals 1 and predFlagL1 equals 1, the pixel values ​​pbSamples[x][y] of the predicted image are derived as follows.

[0348] pbSamples[x][y]=Clip3(0,(1< <bitDepth)-1,

[0349] (predSamplesL0[x][y]*w0+predSamplesL1[x][y]*w1+

[0350] ((o0+o1+1)<<log2Wd))> >(log2Wd+1))

[0351] (BCW Prediction)

[0352] BCW (Bi-prediction with CU-level Weights) is a prediction method that allows switching between pre-determined weighting coefficients based on CU level.

[0353] The input consists of two variables nCbW and nCbH specifying the width and height of the current coding block, two permutations of (nCbW) x (nCbH) predSamplesL0 and predSamplesL1, flags predFlagL0 and predFlagL1 indicating whether to use the prediction list, reference image indices refIdxL0 and refIdxL1, BCW prediction index bcw_idx, and variable cIdx specifying the indices of the luminance and chrominance components. BCW prediction is performed, and the output is the pixel values ​​of the predicted image of the permutation pbSamples of (nCbW) x (nCbH).

[0354] If the `sps_bcw_enabled_flag` indicating whether to use the prediction according to the SPS level is true (TURE), the variable `weightedPredFlag` is 0, and there are no weighted prediction coefficients in the reference images shown by the two reference image indices `refIdxL0` and `refIdxL1`, and the coding block size is below a certain value, then the CU-level grammar's `bcw_idx` is explicitly notified to substitute this value into the variable `bcwIdx`. If `bcw_idx` does not exist, then 0 is substituted into the variable `bcwIdx`.

[0355] With the variable bcwIdx equal to 0, the pixel values ​​of the predicted image are derived as follows.

[0356] pbSamples[x][y]=Clip3(0,(1<<bitDepth)-1,

[0357] (predSamplesL0[x][y]+ predSamplesL1[x][y]+ offset2) >> shift2)

[0358] In other cases (where bcwIdx is not equal to 0), the following formula shall be applied.

[0359] Set the variable w1 to be equal to bcwWLut[bcwIdx]. bcwWLut[k] = {4, 5, 3, 10, -2}.

[0360] The variable w0 is set to (8-w1). Furthermore, the pixel values ​​of the predicted image are derived as follows.

[0361] pbSamples[x][y]=Clip3(0,(1< <bitDepth)-1,

[0362] (w0*predSamplesL0[x][y]+

[0363] w1*predSamplesL1[x][y]+offset3)>>(shift2+3))

[0364] When using BCW prediction in AMVP prediction mode, the inter-frame prediction parameter decoding unit 303 decodes bcw_idx and sends it to the BCW unit 30955. Furthermore, when using BCW prediction in merge prediction mode, the inter-frame prediction parameter decoding unit 303 decodes the merge index merge_idx, and the merge candidate derivation unit 30361 derives the bcwIdx of each merge candidate. Specifically, the merge candidate derivation unit 30361 uses the weighting coefficients of adjacent blocks used for deriving merge candidates as weighting coefficients for merge candidates of the target block. That is, in merge mode, previously used weighting coefficients are inherited as weighting coefficients of the target block.

[0365] (Intra-frame prediction image generation unit 310)

[0366] When predMode indicates intra-prediction mode, the intra-prediction image generation unit 310 uses the intra-prediction parameters input from the intra-prediction parameter derivation unit 304 and the reference pixels read from the reference image memory 306 to perform intra-prediction.

[0367] The inverse quantization / inverse transform unit 311 inverse quantizes the quantization transform coefficients input from the parameter decoding unit 302 to obtain the transform coefficients.

[0368] The adder 312 adds the predicted image of the block input from the predicted image generation unit 308 to the prediction error input from the inverse quantization / inverse transform unit 311 for each pixel to generate the decoded image of the block. The adder 312 stores the decoded image of the block in the reference image memory 306 and outputs it to the loop filter 305.

[0369] The inverse quantization / inverse transform unit 311 inverse quantizes the quantization transform coefficients input from the parameter decoding unit 302 to obtain the transform coefficients.

[0370] The adder 312 adds the predicted image of the block input from the predicted image generation unit 308 to the prediction error input from the inverse quantization / inverse transform unit 311 for each pixel to generate the decoded image of the block. The adder 312 stores the decoded image of the block in the reference image memory 306 and outputs it to the loop filter 305.

[0371] (Composition of a motion picture encoding device)

[0372] Next, the configuration of the motion image encoding device 11 in this embodiment will be described. Figure 12This is a block diagram illustrating the configuration of the motion picture encoding apparatus 11 according to this embodiment. The motion picture encoding apparatus 11 is configured to include: a prediction image generation unit 101, a subtraction unit 102, a transform / quantization unit 103, an inverse quantization / inverse transform unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference image memory (reference image storage unit, frame memory) 109, an encoding parameter determination unit 110, a parameter encoding unit 111, a prediction parameter derivation unit 120, and an entropy encoding unit 104.

[0373] The prediction image generation unit 101 generates a prediction image for each CU. The prediction image generation unit 101 includes the inter-frame prediction image generation unit 309 and the intra-frame prediction image generation unit 310, which have already been described, and their descriptions are omitted.

[0374] The subtraction unit 102 subtracts the pixel values ​​of the predicted image of the block input from the prediction image generation unit 101 from the pixel values ​​of the image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the transform / quantization unit 103.

[0375] The transform / quantization unit 103 calculates the transform coefficients from the prediction error input from the subtraction unit 102 through frequency transformation, and derives the quantization transform coefficients through quantization. The transform / quantization unit 103 outputs the quantization transform coefficients to the parameter encoding unit 111 and the inverse quantization / inverse transform unit 105.

[0376] Inverse quantization / inverse transform unit 105 and inverse quantization / inverse transform unit 311 in motion image decoding device 31 Figure 6 The same applies, and its description is omitted. The calculated prediction error is output to the adder 106.

[0377] The parameter encoding unit 111 includes a header encoding unit 1110, a CT information encoding unit 1111, and a CU encoding unit 1112 (predictive mode encoding unit). The CU encoding unit 1112 also includes a TU encoding unit 1114. The general operation of each module will be explained below.

[0378] The header encoding unit 1110 performs encoding processing on parameters such as header information, segmentation information, prediction information, and quantization transformation coefficients.

[0379] The CT information encoding unit 1111 encodes QT, MT (BT, TT) segmentation information, etc.

[0380] The CU encoding unit 1112 encodes CU information, prediction information, segmentation information, etc.

[0381] The TU encoding unit 1114 encodes the QP update information and the quantization prediction error when the prediction error is included in the TU.

[0382] The CT information coding unit 1111 and the CU coding unit 1112 supply inter-frame prediction parameters (predMode, merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX), intra-frame prediction parameters (intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_reminder, intra_chroma_pred_mode), quantization transform coefficients and other syntax elements to the parameter coding unit 111.

[0383] The entropy coding unit 104 receives quantization transform coefficients and coding parameters (segmentation information, prediction parameters) from the parameter coding unit 111. The entropy coding unit 104 performs entropy coding on them, generating and outputting the coded stream Te.

[0384] The prediction parameter derivation unit 120 is a unit that includes an inter-frame prediction parameter coding unit 112 and an intra-frame prediction parameter coding unit 113. It derives intra-frame prediction parameters and intra-frame prediction parameters based on the parameters input from the coding parameter determination unit 110. The derived intra-frame prediction parameters and intra-frame prediction parameters are output to the parameter coding unit 111.

[0385] (The structure of the inter-frame prediction parameter coding unit)

[0386] like Figure 13 As shown, the inter-frame prediction parameter coding unit 112 is configured to include a parameter coding control unit 1121 and an inter-frame prediction parameter derivation unit 303. The inter-frame prediction parameter derivation unit 303 is a common configuration with the moving image decoding device. The parameter coding control unit 1121 includes a merge index derivation unit 11211 and a vector candidate index derivation unit 11212.

[0387] The merge index derivation unit 11211 derives merge candidates, etc., and outputs them to the inter-frame prediction parameter derivation unit 303. The vector candidate index derivation unit 11212 derives prediction vector candidates, etc., and outputs them to the inter-frame prediction parameter derivation unit 303 and the parameter encoding unit 111.

[0388] (The structure of the intra-frame prediction parameter coding unit 113)

[0389] like Figure 14 As shown, the intra-prediction parameter coding unit 113 includes a parameter coding control unit 1131 and an intra-prediction parameter derivation unit 304. The intra-prediction parameter derivation unit 304 is a common configuration with the moving image decoding device.

[0390] The parameter encoding control unit 1131 derives IntraPredModeY and IntraPredModeC. Then, it determines intra_luma_mpm_flag by referring to mpmCandList[]. These prediction parameters are output to the intra-prediction parameter derivation unit 304 and the parameter encoding unit 111.

[0391] However, unlike the motion picture decoding device, the encoding parameter determination unit 110 and the prediction parameter memory 108 are input to the inter-frame prediction parameter derivation unit 303 and the intra-frame prediction parameter derivation unit 304, and output to the parameter encoding unit 111.

[0392] The addition unit 106 adds the pixel values ​​of the prediction block input from the prediction image generation unit 101 and the prediction error input from the inverse quantization / inverse transform unit 105 to generate a decoded image by adding each pixel. The addition unit 106 stores the generated decoded image in the reference image memory 109.

[0393] The loop filter 107 applies a deblocking filter, SAO, and ALF to the decoded image generated by the adder 106. It should be noted that the loop filter 107 does not necessarily include the above three filters; for example, it may only include a deblocking filter.

[0394] The prediction parameter memory 108 stores the prediction parameters generated by the encoding parameter determination unit 110 in a predetermined location for each object image and CU.

[0395] The image memory 109 stores the decoded images generated by the loop filter 107 in predetermined locations according to each object image and CU.

[0396] The encoding parameter determination unit 110 selects one set from a plurality of sets of encoding parameters. The encoding parameters refer to the QT, BT, or TT segmentation information, prediction parameters, or parameters generated in association with them as encoding objects, as described above. The prediction image generation unit 101 uses these encoding parameters to generate a prediction image.

[0397] The encoding parameter determination unit 110 calculates the RD cost value, representing the information content and encoding error, for each of the multiple sets. The RD cost value is, for example, the sum of the code size and the squared error multiplied by a coefficient λ. The code size is the information content of the encoded stream Te obtained by entropy encoding the quantization error and the encoding parameters. The squared error is the sum of squares of the prediction errors calculated in the subtraction unit 102. The coefficient λ is a real number greater than a predetermined zero. The encoding parameter determination unit 110 selects the set of encoding parameters with the minimum calculated cost value. The encoding parameter determination unit 110 outputs the determined encoding parameters to the parameter encoding unit 111 and the prediction parameter derivation unit 120.

[0398] It should be noted that a portion of the motion picture encoding device 11 and motion picture decoding device 31 described above, such as the entropy decoding unit 301, parameter decoding unit 302, loop filter 305, prediction image generation unit 308, inverse quantization / inverse transform unit 311, addition unit 312, prediction parameter derivation unit 320, prediction image generation unit 101, subtraction unit 102, transform / quantization unit 103, entropy encoding unit 104, inverse quantization / inverse transform unit 105, loop filter 107, encoding parameter determination unit 110, parameter encoding unit 111, and prediction parameter derivation unit 120, can be implemented by a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the computer system can read and execute the program recorded on the recording medium. It should be noted that the "computer system" mentioned here refers to a computer system built into either the motion picture encoding device 11 or the motion picture decoding device 31, employing hardware including an operating system and peripheral devices. Furthermore, "computer-readable recording medium" refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. Moreover, "computer-readable recording medium" can also include: recording media that dynamically stores programs for a short period of time, such as communication lines used to transmit programs via networks like the Internet or telephone lines; and recording media that store programs for a fixed period of time, such as volatile memory within a computer system serving as a server or client in such cases. Furthermore, the aforementioned program can be a program used to implement the above-mentioned functions, or a program that can implement the above-mentioned functions by combining with programs already recorded in the computer system.

[0399] Furthermore, some or all of the motion picture encoding device 11 and motion picture decoding device 31 in the above embodiments can be implemented as integrated circuits such as LSI (Large Scale Integration). Each functional block of the motion picture encoding device 11 and motion picture decoding device 31 can be processorized individually, or some or all can be integrated for processorization. Moreover, the method of integrated circuit implementation is not limited to LSI; it can also be implemented using dedicated circuits or general-purpose processors. Furthermore, if advancements in semiconductor technology lead to integrated circuit technologies that replace LSI, integrated circuits based on such technologies can also be used.

[0400] The above description, with reference to the accompanying drawings, details one embodiment of the invention. However, the specific configuration is not limited to the above embodiment, and various design changes can be made without departing from the spirit of the invention.

[0401] (grammar)

[0402] Figure 15 (a) shows a portion of the syntax of the Sequence Parameter Set (SPS) of Non-Patent Document 1.

[0403] `sps_weighted_pred_flag` is a flag indicating whether weighted prediction can be applied to the P-slice of the reference SPS. `sps_weighted_pred_flag` equal to 0 indicates that weighted prediction is applied to the P-slice of the reference SPS. `sps_weighted_pred_flag` equal to 0 indicates that weighted prediction is not applied to the P-slice of the reference SPS.

[0404] `sps_weighted_bipred_flag` is a flag indicating whether weighted prediction can be applied to the B-slice of the reference SPS. `sps_weighted_bipred_flag` equal to 0 indicates that weighted prediction is applied to the B-slice of the reference SPS. `sps_weighted_bipred_flag` equal to 0 indicates that weighted prediction is not applied to the B-slice of the reference SPS.

[0405] The long_term_ref_pics_flag flag indicates whether to use long-term images.

[0406] inter_layer_ref_pics_present_flag is a flag indicating whether inter-layer prediction is used.

[0407] sps_idr_rpl_present_flag is the slice header of an IDR (Instantaneous Decoding Refresh Picture) and is a flag indicating whether a list of reference pictures is defined.

[0408] When rpl1_same_as_rpl0_flag is 1, it means that there is no information for reference image list 1, and num_ref_pic_lists_in_sps[0] is the same as ref_pic_list_struct(0, rplsIdx).

[0409] Figure 15 (b) shows a portion of the syntax of the Picture Parameter Set (PPS) in Non-Patent Document 1.

[0410] For num_ref_idx_default_active_minus1[i]+1, when i is 0, it represents the value of the variable NumRefIdxActive[0] for slice P or B when num_ref_idx_active_override_flag is 0. When i is 1, it represents the value of the variable NumRefIdxActive[1] for slice B when num_ref_idx_active_override_flag is equal to 0. The value of num_ref_idx_default_active_minus1[i] must be in the range of 0 to 14.

[0411] `pps_weighted_pred_flag` is a flag indicating whether weighted prediction is applied to the P-slice of the reference PPS. `pps_weighted_pred_flag` equal to 0 indicates that weighted prediction is not applied to the P-slice of the reference PPS. `pps_weighted_pred_flag` equal to 1 indicates that weighted prediction is applied to the P-slice of the reference PPS. When `pps_weighted_pred_flag` equals 0, the weighted prediction unit 3094 sets the value of `pps_weighted_pred_flag` to 0. If `pps_weighted_pred_flag` does not exist, its value is set to 0.

[0412] `pps_weighted_bipred_flag` is a flag indicating whether weighted prediction is applied to the B-slice of the reference PPS. A value of 0 for `pps_weighted_bipred_flag` indicates that no weighted prediction is applied to the B-slice of the reference PPS. A value of 1 for `pps_weighted_bipred_flag` indicates that weighted prediction is applied to the B-slice of the reference PPS. When `pps_weighted_bipred_flag` is 0, the weighted prediction unit 3094 sets the value of `pps_weighted_bipred_flag` to 0. If `pps_weighted_bipred_flag` is not present, the value is set to 0.

[0413] `rpl_info_in_ph_flag` is a flag indicating whether the reference image list information exists in the image header. `rpl_info_in_ph_flag` equal to 1 indicates that the reference image list information exists in the image header. `rpl_info_in_ph_flag` equal to 0 indicates that the reference image list information does not exist in the image header, but may exist in the slice header.

[0414] The `wp_info_in_ph_flag` exists when `pps_weighted_pred_flag`, `pps_weighted_bipred_flag`, or `rpl_info_in_ph_flag` are all equal to 1. `wp_info_in_ph_flag` equal to 1 indicates that the weighted prediction information `pred_weight_table` exists in the image header but not in the tile header. `wp_info_in_ph_flag` equal to 0 indicates that the weighted prediction information `pred_weight_table` does not exist in the image header but may exist in the tile header. If `wp_info_in_ph_flag` does not exist, its value is set to 0.

[0415] Figure 16 A portion of the syntax of the image header PH of non-patent document 1 is shown.

[0416] When ph_inter_slice_allowed_flag is 0, it means that all slices in the image have a slice_type of 2 (I Slice). When ph_inter_slice_allowed_flag is 1, it means that the slices included in the image have a slice_type of at least one 0 (B Slice) or 1 (P Slice).

[0417] When rpl_info_in_ph_flag is 1, the ref_pic_lists() function that defines the reference image list is paged, and the reference image list is selected.

[0418] `ph_temporal_mvp_enabled_flag` is a flag indicating whether temporal motion vector prediction is used in inter-frame prediction of slices associated with a PH. When `ph_temporal_mvp_enabled_flag` is 0, temporal motion vector prediction cannot be used in slices associated with a PH. When this is not the case (when `ph_temporal_mvp_enabled_flag` equals 1), temporal motion vector prediction can be used in slices associated with a PH. If it does not exist, it is presumed that the value of `ph_temporal_mvp_enabled_flag` is 0. The value of `ph_temporal_mvp_enabled_flag` is 0 if the reference image within the DPB does not have the same spatial resolution as the current image. When `ph_collocated_from_l0_flag` is 1, it indicates that the reference image used for temporal motion vector prediction is specified using reference image list 0. When `ph_collocated_from_l0_flag` is 0, it indicates that the reference image used for temporal motion vector prediction is specified using reference image list 1. `ph_collocated_ref_idx` represents the index value of the reference image used for temporal motion vector prediction. When `ph_collocated_from_l0_flag` is 1, `ph_collocated_ref_idx` references reference image list 0, and its value must be in the range of 0 to `num_ref_entries[0][RplsIdx[0]]-1`. Furthermore, when `ph_collocated_from_l0_flag` is 0, `ph_collocated_ref_idx` references reference image list 1, and its value must be in the range of 0 to `num_ref_entries[1][RplsIdx[1]]-1`. If it does not exist, it is assumed that `ph_collocated_ref_idx` is equal to 0.

[0419] If ph_inter_slice_allowed_flag is not 0, and pps_weighted_pred_flag is equal to 1, or pps_weighted_bipred_flag is equal to 1, or wp_info_in_ph_flag is equal to 1, then weighted prediction information pred_weight_table exists.

[0420] Figure 17(a) shows a portion of the syntax of the slice header of Non-Patent Document 1. This syntax is decoded, for example, by the parameter decoding unit 302.

[0421] When num_ref_idx_active_override_flag is 1, it means that the syntax element num_ref_idx_active_minus1[0] exists in slices P and B, and the syntax element num_ref_idx_active_minus1[1] exists in slice B. When num_ref_idx_active_override_flag is 0, it means that the syntax element num_ref_idx_active_minus1[0] does not exist in slices P and B. If it does not exist, it is assumed that the value of num_ref_idx_active_override_flag is equal to 1.

[0422] `num_ref_idx_active_minus1[i]` is used to deduce the actual number of reference images used in reference image list `i`. The variable `NumRefIdxActive[i]`, representing the actual number of reference images used, is obtained through... Figure 17 The method shown in (b) is used to derive the value. The value of num_ref_idx_active_minus1[i] must be a value between 0 and 14. In the case that the slice is B and num_ref_idx_active_override_flag is 1 and num_ref_idx_active_minus1[i] does not exist, it is speculated that num_ref_idx_active_minus1[i] is equal to 0.

[0423] When `ph_temporal_mvp_enabled_flag` is 1 and `rpl_info_in_ph_flag` is not 1, the slice header contains information related to temporal motion vector prediction. In this case, if the slice_type is equal to B, `slice_collocated_from_l0_flag` is specified. `rpl_info_in_ph_flag` is a flag indicating that information related to the list of reference images exists in the image header.

[0424] When `slice_collocated_from_l0_flag` is 1, it indicates that the reference image used for temporal motion vector prediction is derived from reference image list 0. When `slice_collocated_from_l0_flag` is 0, it indicates that the reference image used for temporal motion vector prediction is derived from reference image list 1. If `slice_type` equals B or P, `ph_temporal_mvp_enabled_flag` equals 1, and `slice_collocated_from_l0_flag` does not exist, the following applies: If `rpl_info_in_ph_flag` is not 1, it is assumed that `slice_collocated_from_l0_flag` is equal to `ph_collocated_from_l0_flag`. In other cases (where `rpl_info_in_ph_flag` equals 0 and `slice_type` equals P), it is assumed that `slice_collocated_from_l0_flag` is equal to 1.

[0425] `slice_collocated_ref_idx` specifies the index of the reference image used for temporal motion vector prediction. When `slice_type` is P or `slice_type` is B and `slice_collocated_from_l0_flag` is 1, `slice_collocated_ref_idx` references reference image list 0, and the value of `slice_collocated_ref_idx` must be above 0 and below `NumRefIdxActive[0]-1`. When `slice_type` is B and `slice_collocated_from_l0_flag` is 0, `slice_collocated_ref_idx` references reference image list 1, and the value of `slice_collocated_ref_idx` must be above 0 and below `NumRefIdxActive[1]-1`. If `slice_collocated_ref_idx` does not exist, the following applies. When `rpl_info_in_ph_flag` is 1, it is presumed that the value of `slice_collocated_ref_idx` is equal to `ph_collocated_ref_idx`. In all other cases (where rpl_info_in_ph_flag equals 0), it is presumed that slice_collocated_ref_idx is equal to 0. Furthermore, the reference image represented by slice_collocated_ref_idx must be identical across all slices within the image. The values ​​of pic_width_in_luma_samples and pic_height_in_luma_samples of the reference image represented by slice_collocated_ref_idx must be equal to the values ​​of pic_width_in_luma_samples and pic_height_in_luma_samples of the current image, and RprConstraintsActive[slice_collocated_from_l0_flag?0:1][slice_collocated_ref_idx] must be equal to 0.

[0426] Page the pred_weight_table if wp_info_in_ph_flag is not 1 and pps_weighted_pred_flag is 1 and slice_type is 1 (P Slice), or if pps_weighted_bipred_flag is 1 and slice_type is 0 (B Slice).

[0427] Figure 17 (b) shows the derivation method of the variable NumRefIdxActive[i] in Non-Patent Document 1 based on the prediction parameter derivation unit 320. In the case where the reference image list i (=0,1) is reference image list 0 in slice B or slice P, if num_ref_idx_active_override_flag is equal to 1, the variable NumRefIdxActive[i] is substituted with the value of num_ref_idx_active_minus1[i] plus 1. In cases other than those (where reference image list 0 exists in slice B or slice P, i.e., num_ref_idx_active_override_flag equals 0), if the value of num_ref_entries[i][RplsIdx[i]] is greater than or equal to the value obtained by adding 1 to num_ref_idx_default_active_minus1[i], then the variable NumRefIdxActive[i] is substituted with the value obtained by adding 1 to num_ref_idx_default_active_minus1[i]. Otherwise, the variable NumRefIdxActive[i] is substituted with the value of num_ref_entries[i][RplsIdx[i]]. num_ref_idx_default_active_minus1[i] is the default value of the variable NumRefIdxActive[i] defined by PPS. In the case of slice I, or in the case of reference image list 1 in slice P, the variable NumRefIdxActive[i] is substituted with 0.

[0428] Figure 18 The syntax of the weighted prediction information pred_weight_table for non-patent literature 1 is shown.

[0429] Here, num_l0_weights represents the number of weights signaled to the entries of reference image list 0 when wp_info_in_ph_flag equals 1. The value of num_l0_weights is in the range of min(15, num_ref_entries[0][RplsIdx[0]]) above 0. When wp_info_in_ph_flag equals 1, the variable NumWeightsL0 is set to equal num_l0_weights. In other cases (when wp_info_in_ph_flag equals 0), NumWeightsL0 is set to NumRefIdxActive[0]. Here, num_ref_entries[i][RplsIdx[i]] represents the number of reference images in reference image list i. The variable RplsIdx[i] is the index value of the list representing multiple reference image lists i.

[0430] num_l1_weights specifies the number of weights signaled to the entries of reference image list 1 when both pps_weighted_bipred_flag and wp_info_in_ph_flag are equal to 1. The value of num_l1_weights is set to be in the range of min(15, num_ref_entries[1][RplsIdx[1]]) above 0.

[0431] If pps_weighted_bipred_flag is 0, then set the variable NumWeightsL1 to 0; otherwise, if wp_info_in_ph_flag is 1, then substitute the value of num_l1_weights into the variable NumWeightsL1; otherwise, substitute NumRefIdxActive[1] into the variable NumWeightsL1.

[0432] luma_log2_weight_denom is the base-2 logarithm of the denominator of the weighting coefficient for all luma. The value of luma_log2_weight_denom must be within the range of 0 to 7. delta_chroma_log2_weight_denom is the difference of the base-2 logarithms of the denominators of all chroma weighting coefficients. In the absence of delta_chroma_log2_weight_denom, it is assumed to be equal to 0. It is derived that the variable ChromaLog2WeightDenom is equal to luma_log2_weight_denom + delta_chroma_log2_weight_denom, and the value must be within the range of 0 to 7.

[0433] When luma_weight_l0_flag[i] is 1, it indicates the presence of the weighting coefficient for the luma component predicted by L0. When luma_weight_l0_flag[i] is 0, it indicates the absence of the weighting coefficient for the luma component predicted by L0. In the absence of luma_weight_l0_flag[i], it is assumed that the weighted prediction unit 3094 is equal to 0. When chroma_weight_l0_flag[i] is 1, it indicates the presence of the weighting coefficient for the chroma prediction value predicted by L0. When chroma_weight_l0_flag[i] is 0, it indicates the absence of the weighting coefficient for the chroma prediction value predicted by L0. In the absence of chroma_weight_l0_flag[i], it is assumed that the weighted prediction unit 3094 is equal to 0.

[0434] delta_luma_weight_l0[i] is the difference of the weighting coefficients applied to the luma prediction value predicted by L0 using RefPicList[0][i]. It is derived that the variable LumaWeightL0[i] is equal to (1 << luma_log2_weight_denom) + delta_luma_weight_l0[i]. When luma_weight_l0_flag[i] is equal to 1, the value of delta_luma_weight_l0[i] must be within the range of -128 to 127. When luma_weight_l0_flag[i] is equal to 0, it is assumed by the weighted prediction unit 3094 that LumaWeightL0[i] is equal to the power of 2 to the luma_log2_weight_denom (2 ^ luma_log2_weight_denom).

[0435] luma_offset_l0[i] is an offset value for the predicted value of the luminance used for L0 prediction using RefPicList[0][i]. The value of luma_offset_l0[i] must be in the range of -128 to 127. When luma_weight_l0_flag[i] is equal to 0, the weighted prediction unit 3094 assumes that luma_offset_l0[i] is equal to 0.

[0436] delta_chroma_weight_l0[i][j] is the difference in the weighting coefficient for the predicted value of the chrominance used for L0 prediction using RefPicList0[i] where j is 0 for Cb and j is 1 for Cr. It is derived that the variable ChromaWeightL0[i][j] is equal to (1<<ChromaLog2WeightDenom) + delta_chroma_weight_l0[i][j]. When chroma_weight_l0_flag[i] is equal to 1, the value of delta_chroma_weight_l0[i][j] must be in the range of -128 to 127. When chroma_weight_l0_flag[i] is 0, the weighted prediction unit 3094 assumes that ChromaWeightL0[i][j] is equal to the power of 2 to the ChromaLog2WeightDenom (2^ChromaLog2WeightDenom). delta_chroma_offset_l0[i][j] is the difference in the offset value for the predicted value of the chrominance used for L0 prediction using RefPicList0[i] where j is 0 for Cb and j is 1 for Cr. The variable ChromaOffsetL0[i][j] is derived as follows.

[0437] ChromaOffsetL0[i][j]=Clip3(-128, 127,

[0438] (128 + delta_chroma_offset_l0[i][j] -

[0439] ((128 * ChromaWeightL0[i][j]) >> ChromaLog2WeightDenom)))

[0440] The value of delta_chroma_offset_l0[i][j] must be in the range of -4*128 to 4*127. When chroma_weight_l0_flag[i] equals 0, the weighted prediction unit 3094 speculates that ChromaOffsetL0[i][j] equals 0.

[0441] It should be noted that the following are interpretations: luma_weight_l1_flag[i], chroma_weight_l1_flag[i], delta_luma_weight_l1[i], luma_offset_l1[i], delta_chroma_weight_l1[i][j], and delta_chroma_offset_l1[i][j] are replaced with luma_weight_l0_flag[i], chroma_weight_l0_flag[i], delta_luma_weight_l0[i], luma_offset_l0[i], delta_chroma_weight_l0[i][j], and delta_chroma_offset_l0[i][j], respectively. Similarly, l0, L0, list0, and List0 are replaced with l1, l1, list1, and List1, respectively.

[0442] Figure 19 (a) shows the syntax of ref_pic_lists() for the definition of the picture list in Non-Patent Document 1. ref_pic_lists() sometimes exists in the picture header or slice header. If rpl_sps_flag[i] is 1, it means that the picture list i of ref_pic_lists() is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) of SPS. Here listIdx is equal to i (=0,1), and rplsIdx = rpl_idx[i].

[0443] When rpl_sps_flag[i] is 0, it means that the reference image list i is derived based on ref_pic_list_struct(listIdx, rplsIdx). Here, listIdx is equal to i directly included in ref_pic_lists(). If rpl_sps_flag[i] does not exist, the following applies. When num_ref_pic_lists_in_sps[i] is 0, the value of rpl_sps_flag[i] is presumed to be 0. When num_ref_pic_lists_in_sps[i] is greater than 0, if rpl1_idx_present_flag is equal to 0 and i is equal to 1, then the value of rpl_sps_flag[1] is presumed to be equal to rpl_sps_flag[0].

[0444] If rpl_sps_flag[i] is 1, decode rpl_idx[i]. rpl_idx[i] is used for the derivation of RplsIdx[i] (described later), where RplsIdx[i] represents the index rplsIdx of ref_pic_list_struct(listIdx, rplsIdx). ref_pic_list_struct(listIdx, rplsIdx) is used for the derivation of reference image i. Here, listIdx equals i. If it does not exist, it is presumed that the value of rpl_idx[i] is 0. The value of rpl_idx[i] is in the range of 0 or higher and below num_ref_pic_lists_in_sps[i] - 1. If rpl_sps_flag[i] is 1 and num_ref_pic_lists_in_sps[i] is 1, it is presumed that the value of rpl_idx[i] is 0. Given that rpl_sps_flag[i] is 1 and rpl1_idx_present_flag is 0, it is inferred that the value of rpl_idx[1] is equal to rpl_idx[0]. The derivation of variable RplsIdx[i] is as follows.

[0445] RplsIdx[i]=(rpl_sps_flag[i])?rpl_idx[i]:num_ref_pic_lists_in_sps[i]

[0446] Figure 19 (b) shows the syntax of the non-patent document 1, which refers to the picture list structure ref_pic_list_struct(listIdx, rplsIdx).

[0447] Sometimes `ref_pic_list_struct(listIdx, rplsIdx)` exists in the SPS, picture header, or slice header. Depending on whether it's included in the SPS, picture header, or slice header, the following applies: If it exists in the picture or slice header, `ref_pic_list_struct(listIdx, rplsIdx)` represents the list of reference images `listIdx` for the current picture (including sliced ​​pictures). If it exists in the SPS, `ref_pic_list_struct(listIdx, rplsIdx)` represents a candidate for the list of reference images `listIdx`. Furthermore, the current picture can be referenced by index from the picture header or slice header to the list of `ref_pic_list_struct(listIdx, rplsIdx)` included in the SPS.

[0448] Here, `num_ref_entries[listIdx][rplsIdx]` represents the number of `ref_pic_list_struct(listIdx, rplsIdx)`. The value of `num_ref_entries[listIdx][rplsIdx]` is a value greater than 0 and less than `MaxDpbSize + 13`. `MaxDpbSize` is the number of images to be decoded, determined based on the configuration file level.

[0449] ltrp_in_header_flag[listIdx][rplsIdx] is a flag indicating whether a long-term reference image exists in ref_pic_list_struct(listIdx, rplsIdx).

[0450] inter_layer_ref_pic_flag[listIdx][rplsIdx][i] is a flag indicating whether the i-th image in the list of reference images of ref_pic_list_struct(listIdx, rplsIdx) is an inter-layer prediction.

[0451] st_ref_pic_flag[listIdx][rplsIdx][i] is a flag indicating whether the i-th element in the list of reference images of ref_pic_list_struct(listIdx, rplsIdx) is a short-term reference image.

[0452] abs_delta_poc_st[listIdx][rplsIdx][i] is a syntactic element used to derive the absolute difference of the POC for a short-term reference image.

[0453] strp_entry_sign_flag[listIdx][rplsIdx][i] is a flag used to deduce the sign of positive or negative signs.

[0454] rpls_poc_lsb_lt[listIdx][rplsIdx][i] is a syntactic element used to deduce the POC of the i-th long-term reference image in the list of reference images for ref_pic_list_struct(listIdx, rplsIdx).

[0455] ilrp_idx[listIdx][rplsIdx][i] is a syntactic element used to derive the layer information of the i-th interlayer predicted reference image from the list of reference images in ref_pic_list_struct(listIdx, rplsIdx).

[0456] As for the problem with the method described in Non-Patent Document 1, such as Figure 19 As shown in (b), the following aspects exist: It is possible to specify 0 as the value of num_ref_entries[listIdx][rplsIdx] for the reference image list structure ref_pic_list_struct(listIdx, rplsIdx). 0 indicates that the number of reference images in the reference image list listIdx of the pic_list_struct represented by rplsIdx is 0. num_ref_entries can be specified regardless of slice_type. In the case of 0 reference images for slice P and slice B, the number of reference images in the reference image list is at least 1. If 0 is specified, there are no reference images to refer to, therefore, the reference images become uncertain.

[0457] Therefore, in this embodiment, as Figure 20 As shown, the notified syntax element is not `num_ref_entries[listIdx][rplsIdx]`, but rather `num_ref_entries_minus1[listIdx][rplsIdx]`, ensuring that `num_ref_entries_minus1[listIdx][rplsIdx]` takes a value greater than 0 and less than `MaxDpbSize+14`. By doing this, and by arbitrarily prohibiting the number of reference images to be 0, the reference images can be kept from becoming uncertain.

[0458] Furthermore, other problems with the method described in Non-Patent Document 1 can be listed as follows: Figure 18In the `pred_weight_table`, the weights of reference image list 0 and reference image list 1 are explicitly described using the syntax `num_l0_weights` and `num_l1_weights`. In the image header, at the time of paging the `pred_weight_table`, the number of reference images in reference image list i is already defined by `ref_pic_list_struct(listIdx, rplsIdx)`, and its syntax is redundant. Furthermore, in the slice header, at the time of paging the `pred_weight_table`, the number of reference images in reference image list i is already defined by `NumRefIdxActive[i]`, and its syntax is redundant. Therefore, in this embodiment, as... Figure 21 As shown in (a), in the image header, before the paging pred_weight_table, the variable NumWeightsL0 is substituted with the value of num_ref_entries_minus1[0][RplsIdx[0]]+1. Furthermore, for the variable NumWeightsL1, if pps_weighted_bipred_flag is 1, then the value of num_ref_entries_minus1[1][RplsIdx[1]]+1 is substituted; otherwise, 0 is substituted. This is because if pps_weighted_bipred_flag is 0, there is no weighted prediction for bidirectional prediction. Additionally, as... Figure 21 As shown in (b), in the slice header, before paging pred_weight_table, the value of variable NumRefIdxActive[0] is substituted into variable NumWeightsL0, and the value of variable NumRefIdxActive[1] is substituted into variable NumWeightsL1. Incidentally, the value of variable NumRefIdxActive[1] is 0 when slicing P. And, as Figure 22 As shown, pred_weight_table does not explicitly describe num_l0_weights and num_l1_weights as syntax. Instead, it sets the variable NumWeightsL0 to the number of weights in reference image list 0 and sets the variable NumWeightsL1 to reference image list 1. This eliminates redundancy.

[0459] Figure 23This is an example of another implementation of this embodiment. In this example, variables NumWeightsL0 and NumWeightsL1 are defined in pred_wight_table. When wp_info_in_ph_flag is equal to 1, the value of num_ref_entries[0][PicRplsIdx[0]] is substituted into variable NumWeightsL0; otherwise, the value of variable NumRefIdxActive[0] is substituted. wp_info_in_ph_flag is a flag indicating that weighted prediction information exists in the image header.

[0460] Furthermore, when wp_info_in_ph_flag equals 1 and pps_weighted_bipred_flag equals 1, the value of num_ref_entries[1][PicRplsIdx[1]] is substituted into the variable NumWeightsL1. pps_weighted_bipred_flag is a flag indicating bidirectional weighted prediction. When wp_info_in_ph_flag equals 1 and pps_weighted_bipred_flag is 0, 0 is substituted into the variable NumWeightsL1. When wp_info_in_ph_flag equals 0, the value of the variable NumRefIdxActive[1] is substituted into NumWeightsL1. By doing this, instead of explicitly describing num_l0_weights and num_l1_weights through syntax, setting the variable NumWeightsL0 to the number of weights of reference image list 0 and the variable NumWeightsL1 to the number of weights of reference image list 1, redundancy can be eliminated.

[0461] As another problem with the method described in Non-Patent Document 1, the number of active reference images is defined in the slice header, but not in the image header.

[0462] Therefore, in other embodiments of this implementation, such as Figure 24As shown in (a), the number of active reference images can be defined even in the image header. The following syntax is decoded, for example, by the parameter decoding unit 302. When ph_inter_slice_allowed_flag equals 1 and rpl_info_in_ph_flag equals 1, the number of active reference images is defined. When ph_inter_slice_allowed_flag is 1, it indicates that a P-slice or a B-slice may exist within the image. When rpl_info_in_ph_flag equals 1, it indicates that the reference image list information exists in the image header.

[0463] ph_num_ref_idx_active_override_flag is a flag indicating whether ph_num_ref_idx_active_minus1[0] and ph_num_ref_idx_active_minus1[1] exist.

[0464] ph_num_ref_idx_active_minus1[i] is a syntactic element used to deduce the variable NumRefIdxActive[i] relative to the list of reference images i, and is a value greater than 0 and less than 14.

[0465] `ph_collocated_ref_idx` represents the index of the reference image used for time-based motion vector prediction. When `ph_collocated_from_l0_flag` is 1, `ph_collocated_ref_idx` references reference image list 0, and its value is 0 or higher than `NumRefIdxActive[0]-1`. When `ph_collocated_from_l0_flag` is 0, `ph_collocated_ref_idx` references an entry in reference image list 1, and its value is 0 or higher than `NumRefIdxActive[1]-1`. If it does not exist, it is presumed that `ph_collocated_ref_idx` is equal to 0.

[0466] Figure 24(b) illustrates the derivation method of the variable NumRefIdxActive[i] based on the prediction parameter derivation unit 320. For the reference image list i (=0,1), when ph_num_ref_idx_active_override_flag is 1, if num_ref_entries_minus1[i][RplsIdx[i]] is greater than 0, then the variable NumRefIdxActive[i] is substituted with the value obtained by adding 1 to the value of ph_num_ref_idx_active_minus1[i]. Otherwise, 1 is substituted. On the other hand, if ph_num_ref_idx_active_override_flag is not 1, and the value of num_ref_entries_minus1[i][RplsIdx[i]] is greater than or equal to the value obtained by adding 1 to num_ref_idx_default_active_minus1[i], then the variable NumRefIdxActive[i] is substituted with the value obtained by adding 1 to num_ref_idx_default_active_minus1[i]. Otherwise, the variable NumRefIdxActive[i] is substituted with the value obtained by adding 1 to num_ref_entries_minus1[i][RplsIdx[i]]. num_ref_idx_default_active_minus1[i] is the default value of the variable NumRefIdxActive[i] defined by PPS.

[0467] Figure 25 (a) is the syntax of the slice header. The following syntax is decoded, for example, by parameter decoding unit 302. In the slice header, if the rpl_info_in_ph_flag indicating that the reference image list information exists in the image header is not 1 and it is a P slice or a B slice, the number of active reference images is defined.

[0468] Figure 25(b) shows the derivation method of the variable NumRefIdxActive[i] based on the prediction parameter derivation unit 320. For the reference image list i (=0,1), the variable NumRefIdxActive[i] is overridden when rpl_info_in_ph_flag is not 1 and i is 0 in slice B or slice P. If num_ref_idx_active_override_flag is 1 and num_ref_entries_minus1[i][RplsIdx[i]] is greater than 0, then the variable NumRefIdxActive[i] is substituted with the value obtained by adding 1 to the value of num_ref_idx_active_minus1[i]; otherwise, 1 is substituted. If num_ref_idx_active_override_flag is not 1, and the value of num_ref_entries_minus1[i][RplsIdx[i]] is greater than or equal to the value obtained by adding 1 to num_ref_idx_default_active_minus1[i], then the variable NumRefIdxActive[i] is substituted with the value obtained by adding 1 to num_ref_idx_default_active_minus1[i]. Otherwise, the variable NumRefIdxActive[i] is substituted with the value obtained by adding 1 to num_ref_entries_minus1[i][RplsIdx[i]]. When i is 1 in I-slice or P-slice, 0 is substituted into the variable NumRefIdxActive[i] regardless of the value of rpl_info_in_ph_flag. rpl_info_in_ph_flag is a flag indicating that the reference image list information exists in the image header. num_ref_idx_default_active_minus1[i] is the value of the default variable NumRefIdxActive[i] defined by PPS.

[0469] like Figure 26 As shown, pred_weight_table does not explicitly describe num_l0_weights and num_l1_weights as syntax. Instead, it sets the variable NumRefIdxActive[0] to the number of weights of reference image list 0 and sets the variable NumRefIdxActive[1] to the number of weights of reference image list 1. This eliminates redundancy.

[0470] As mentioned earlier, a problem with the method described in Non-Patent Document 1 is that it's possible to specify 0 as the value of num_ref_entries[listIdx][rplsIdx] of the reference image list structure ref_pic_list_struct(listIdx, rplsIdx). However, the number of reference images in the reference image list 0 of slice P and the reference image list of slice B must be at least 1. Therefore, the following problem exists: if 0 is specified, no reference image exists, and the reference image becomes uncertain.

[0471] Therefore, firstly, based on nal_unit_type and slice_type, which represent the types of NAL (Network Abstraction Layer) units, the following restrictions are imposed on the value of num_ref_entries, which represents the number of reference images in the list of reference images.

[0472] When `nal_unit_type` is either `IDR_W_RADL` or `IDR_N_LP` (i.e., for IDR images), `num_ref_entries[i][RplsIdx[i]]` must be 0 if `i` is either 0 or 1. `i` represents the index `listIdx` of the list of reference images. It's important to note that `nal_unit_type` for IDR images can be either `IDR_W_RADL` or `IDR_N_LP`. `IDR_W_RADL` indicates the possibility of having a RADL (Random Access Decodable Leading) image output before the IDR image, while `IDR_N_LP` indicates the possibility of not having one.

[0473] When slice_type is P or B, num_ref_entries[0][RplsIdx[0]] must be a value greater than 0.

[0474] When slice_type is B, num_ref_entries[1][RplsIdx[1]] must be a value greater than 0.

[0475] Based on this, in this embodiment, Figure 27 The syntax of the image header in (a) is set as an example. The following syntax is decoded, for example, by the parameter decoding unit 302.

[0476] When rpl_info_in_ph_flag is 1, such as Figure 27As shown in (a), in the image header, page ref_pic_lists() and select the reference image list.

[0477] However, in the case of IDR images, num_ref_entries[i][RplsIdx[i]] must be the list of reference images with a value of 0, either i is 0 or 1.

[0478] In addition, if there are P-slices or B-slices in the image, a list of reference images with a value greater than 0 must be selected as num_ref_entries[0][RplsIdx[0]].

[0479] In addition, if there is a B slice in the image, a list of reference images with a value greater than 0 must be selected as num_ref_entries[1][RplsIdx[1]].

[0480] Based on this, the number of activated reference images is used to deduce the variable NumRefIdxActive[i]. ph_num_ref_idx_active_override_flag is a flag indicating whether ph_num_ref_idx_active_minus1[0] and ph_num_ref_idx_active_minus1[1] exist. When ph_num_ref_idx_active_override_flag is 1, ph_num_ref_idx_active_minus1[0] and ph_num_ref_idx_active_minus1[1] exist; when ph_num_ref_idx_active_override_flag is 0, ph_num_ref_idx_active_minus1[0] and ph_num_ref_idx_active_minus1[1] do not exist. When ph_num_ref_idx_active_override_flag does not exist, the value is presumed to be 0.

[0481] ph_num_ref_idx_active_minus1[i], i=0, 1 is a syntactic element used to derive the variable NumRefIdxActive[i] relative to the list of reference images i, with a value between 0 and 14.

[0482] In addition, ph_num_ref_idx_active_minus1[i]+1 must be a value below num_ref_entries[i][RplsIdx[i]].

[0483] If num_ref_entries[i][RplsIdx[i]] is greater than 1, the value is explicitly notified. If ph_num_ref_idx_active_minus1[i] does not exist, the value is presumed to be 0.

[0484] Figure 27 (b) shows the derivation method of the variable NumRefIdxActive[i] based on the prediction parameter derivation unit 320. For the reference image list i (=0,1), when ph_num_ref_idx_active_override_flag is 1, the variable NumRefIdxActive[i] is substituted with the value of ph_num_ref_idx_active_minus1[i] plus 1. On the other hand, when ph_num_ref_idx_active_override_flag is not 1, if the value of num_ref_entries[i][RplsIdx[i]] is greater than or equal to the value of num_ref_idx_default_active_minus1[i] plus 1, then the variable NumRefIdxActive[i] is substituted with the value of num_ref_idx_default_active_minus1[i] plus 1. Otherwise, substitute num_ref_entries[i][RplsIdx[i]] into the variable NumRefIdxActive[i]. num_ref_idx_default_active_minus1[i] is the default value of the variable NumRefIdxActive[i] defined by PPS.

[0485] Figure 28 (a) is the syntax for the slice header. The following syntax is decoded, for example, by the parameter decoding unit 302.

[0486] In the slice header, if rpl_info_in_ph_flag is not 1 and nal_unit_type is neither IDR_W_RADL nor IDR_N_LP or if sps_idr_rpl_present_flag is 1, page ref_pic_lists() and select the reference image list.

[0487] However, in the case of IDR images, num_ref_entries[i][RplsIdx[i]] must be the list of reference images with a value of 0, either i is 0 or 1.

[0488] In addition, in the case of P-slice or B-slice, a list of reference images with a value greater than 0 must be selected as num_ref_entries[0][RplsIdx[0]].

[0489] In addition, in the case of B slices, a list of reference images with values ​​greater than 0 must be selected as num_ref_entries[1][RplsIdx[1]].

[0490] Based on this, the number of activated reference images is used to deduce the variable NumRefIdxActive[i]. When slice_type is not I and num_ref_entries[0][RplsIdx[0]] is greater than 1, or when slice_type is B and num_ref_entries[1][RplsIdx[1]] is greater than 1, num_ref_idx_active_override_flag exists. When num_ref_idx_active_override_flag is 1, num_ref_idx_active_minus1[0] and num_ref_idx_active_minus1[1] exist. When num_ref_idx_active_override_flag is 0, num_ref_idx_active_minus1[0] and num_ref_idx_active_minus1[1] do not exist. When num_ref_idx_active_override_flag does not exist, the value is presumed to be 0.

[0491] num_ref_idx_active_minus1[i] is a syntactic element used to deduce the variable NumRefIdxActive[i] relative to the list of reference images i, and is a value between 0 and 14.

[0492] In addition, num_ref_idx_active_minus1[i]+1 must be a value below num_ref_entries[i][RplsIdx[i]].

[0493] When slice_type is B, there is a syntax element num_ref_idx_active_minus1[i] with i being 0 and 1. When slice_type is not B (when slice_type is P), there is only a syntax element num_ref_idx_active_minus1[0].

[0494] If num_ref_entries[i][RplsIdx[i]] is greater than 1, the value is explicitly represented. If num_ref_idx_active_minus1[i] does not exist, the value is presumed to be 0.

[0495] Figure 28 (b) shows the derivation method of the variable NumRefIdxActive[i] based on the prediction parameter derivation unit 320.

[0496] If rpl_info_in_ph_flag is not 1 and it is a B slice or a P slice with i=0, the following processing is performed. If num_ref_idx_active_override_flag is 1, substitute the value obtained by adding 1 to the value of num_ref_idx_active_minus1[i] into the variable NumRefIdxActive[i]. If num_ref_idx_active_override_flag is not 1, and the value of num_ref_entries[i][RplsIdx[i]] is greater than the value obtained by adding 1 to num_ref_idx_default_active_minus1[i], then substitute the value obtained by adding 1 to num_ref_idx_default_active_minus1[i] into the variable NumRefIdxActive[i]. Otherwise, substitute num_ref_entries[i][RplsIdx[i] into the variable NumRefIdxActive[i].

[0497] That's not the case. In the case of I slices or P slices with i=1, the variable NumRefIdxActive[i] is substituted with 0 regardless of the value of rpl_info_in_ph_flag. rpl_info_in_ph_flag is a flag indicating that the reference image list information exists in the image header. num_ref_idx_default_active_minus1[i] is the default value of the variable NumRefIdxActive[i] defined by PPS.

[0498] [Application Example]

[0499] The aforementioned motion picture encoding device 11 and motion picture decoding device 31 can be mounted on various devices for transmitting, receiving, recording, and reproducing motion pictures. It should be noted that motion pictures can be natural motion pictures captured by cameras or the like, or artificial motion pictures generated by computers or the like (including CG (Computer Graphics) and GUI (Graphical User Interface)).

[0500] First, refer to Figure 2 The following describes the situation where the above-described motion picture encoding device 11 and motion picture decoding device 31 can be used for the transmission and reception of motion pictures.

[0501] Figure 2 PROD_A is a block diagram representing the configuration of the transmitting device PROD_A equipped with the motion picture encoding device 11. For example... Figure 2 As shown, the transmitting device PROD_A includes: an encoding unit PROD_A1 that obtains encoded data by encoding a moving image; a modulation unit PROD_A2 that obtains a modulated signal by modulating a carrier wave using the encoded data obtained by the encoding unit PROD_A1; and a transmitting unit PROD_A3 that transmits the modulated signal obtained by the modulation unit PROD_A2. The moving image encoding device 11 described above is used as the encoding unit PROD_A1.

[0502] As a source of motion images input to the encoding unit PROD_A1, the transmitting device PROD_A may further include: a camera PROD_A4 for capturing motion images, a recording medium PROD_A5 for recording motion images, an input terminal PROD_A6 for inputting motion images from an external source, and an image processing unit A7 for generating or processing images. Figure 2 The example shows that the sending device PROD_A has all of these components, but some can be omitted.

[0503] It should be noted that the recording medium PROD_A5 can be a medium that records unencoded motion images, or a medium that records motion images encoded using a recording encoding method different from the encoding method used for transmission. In the latter case, it is preferable that the decoding unit (not shown) that decodes the encoded data read from the recording medium PROD_A5 according to the recording encoding method is located between the recording medium PROD_A5 and the encoding unit PROD_A1.

[0504] Figure 2 PROD_B is a block diagram representing the configuration of the receiving device PROD_B equipped with the motion picture decoding device 31. For example... Figure 2As shown, the receiving device PROD_B includes: a receiving unit PROD_B1 for receiving a modulated signal, a demodulation unit PROD_B2 for obtaining coded data by demodulating the modulated signal received by the receiving unit PROD_B1, and a decoding unit PROD_B3 for obtaining a moving image by decoding the coded data obtained by the demodulation unit PROD_B2. The aforementioned moving image decoding device 31 is used as the decoding unit PROD_B3.

[0505] The receiving device PROD_B, serving as the destination for the moving images output by the decoding unit PROD_B3, may further include a display PROD_B4 for displaying the moving images, a recording medium PROD_B5 for recording the moving images, and an output terminal PROD_B6 for outputting the moving images to an external device. Figure 2 The example shows that the receiving device PROD_B has all of these components, but some can be omitted.

[0506] It should be noted that the recording medium PROD_B5 can be a medium for recording unencoded motion images, or it can be a medium encoded with a recording encoding method different from the encoding method used for transmission. In the latter case, it is preferable that the encoding unit (not shown) that encodes the motion images acquired from the decoding unit PROD_B3 according to the recording encoding method is located between the decoding unit PROD_B3 and the recording medium PROD_B5.

[0507] It should be noted that the transmission medium for modulated signals can be wireless or wired. Furthermore, the transmission scheme for modulated signals can be broadcast (here, a transmission scheme where the destination is not predetermined) or communication (here, a transmission scheme where the destination is predetermined). That is, the transmission of modulated signals can be achieved through any of the following: wireless broadcasting, wired broadcasting, wireless communication, and wired communication.

[0508] For example, a terrestrial digital broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals via wireless broadcasting. Similarly, a cable television broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals via cable broadcasting.

[0509] Furthermore, servers (workstations, etc.) and clients (TV receivers, personal computers, smartphones, etc.) using internet-based VOD (Video On Demand) services, moving image sharing services, etc., are examples of transmitting devices PROD_A and receiving devices PROD_B that transmit and receive modulated signals via communication (typically, either wireless or wired is used as the transmission medium in a LAN, and wired is used in a WAN). Here, personal computers include desktop PCs, laptop PCs, and tablet PCs. Additionally, smartphones also include multi-functional portable telephone terminals.

[0510] It should be noted that, in addition to decoding and displaying the encoded data downloaded from the server, the client of the motion picture sharing service also has the function of encoding and uploading motion pictures captured by a camera to the server. That is, the client of the motion picture sharing service performs the functions of both the sending device PROD_A and the receiving device PROD_B.

[0511] Next, refer to Figure 3 The following describes the situation where the above-mentioned motion picture encoding device 11 and motion picture decoding device 31 can be used for recording and reproducing motion pictures.

[0512] Figure 3 PROD_C is a block diagram representing the configuration of the recording device PROD_C equipped with the aforementioned motion picture encoding device 11. For example... Figure 3 As shown, the recording device PROD_C includes an encoding unit PROD_C1 that obtains encoded data by encoding moving images, and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 to the recording medium PROD_M. The moving image encoding device 11 described above is used as the encoding unit PROD_C1.

[0513] It should be noted that the recording medium PROD_M can be (1) a type of recording medium built into the recording device PROD_C, such as HDD (Hard Disk Drive) or SSD (Solid State Drive), or (2) a type of recording medium connected to the recording device PROD_C, such as SD (Secure Digital) memory card or USB (Universal Serial Bus) flash memory, or (3) a recording medium loaded into a drive device (not shown) built into the recording device PROD_C, such as DVD (Digital Versatile Disc) or BD (Blu-ray Disc).

[0514] Furthermore, as a source of motion images input to the encoding unit PROD_C1, the recording device PROD_C may further include: a camera PROD_C3 for capturing motion images, an input terminal PROD_C4 for inputting motion images from the outside, a receiving unit PROD_C5 for receiving motion images, and an image processing unit PROD_C6 for generating or processing images. Figure 3 The example shows that the recording device PROD_C has all of these components, but some can be omitted.

[0515] It should be noted that the receiving unit PROD_C5 can receive unencoded motion images, or it can receive encoded data encoded using a transmission encoding method different from the encoding method used for recording. In the latter case, it is preferable to place the transmission decoding unit (not shown) that decodes the encoded data encoded using the transmission encoding method between the receiving unit PROD_C5 and the encoding unit PROD_C1.

[0516] Examples of such recording devices PROD_C include DVD recorders, BD recorders, and HDD (Hard Disk Drive) recorders (in which case the input terminal PROD_C4 or the receiving unit PROD_C5 is the main source of moving images). Furthermore, portable camcorders (in which case the camera PROD_C3 is the main source of moving images), personal computers (in which case the receiving unit PROD_C5 or the image processing unit C6 is the main source of moving images), and smartphones (in which case the camera PROD_C3 or the receiving unit PROD_C5 is the main source of moving images) are also examples of such recording devices PROD_C.

[0517] Figure 3PROD_D is a block representing the configuration of the playback device PROD_D equipped with the aforementioned motion picture decoding device 31. For example... Figure 3 As shown, the playback device PROD_D includes a readout unit PROD_D1 that reads encoded data written to the recording medium PROD_M, and a decoding unit PROD_D2 that obtains a moving image by decoding the encoded data read out by the readout unit PROD_D1. The aforementioned moving image decoding device 31 is used as the decoding unit PROD_D2.

[0518] It should be noted that the recording medium PROD_M can be (1) a recording medium built into the playback device PROD_D, such as HDD or SSD, or (2) a recording medium connected to the playback device PROD_D, such as SD memory card or USB flash drive, or (3) a recording medium loaded into a drive device (not shown) built into the playback device PROD_D, such as DVD or BD.

[0519] Furthermore, as the destination for the motion images output by the decoding unit PROD_D2, the playback device PROD_D may further include: a display PROD_D3 for displaying motion images, an output terminal PROD_D4 for outputting motion images to the outside, and a transmitting unit PROD_D5 for transmitting motion images. Figure 3 The example shows that the reproduction device PROD_D has all of these components, but some can be omitted.

[0520] It should be noted that the transmitting unit PROD_D5 can transmit unencoded motion images, or it can transmit encoded data encoded using a transmission encoding method different from the encoding method used for recording. In the latter case, it is preferable to place the encoding unit (not shown) that encodes the motion images using the transmission encoding method between the decoding unit PROD_D2 and the transmitting unit PROD_D5.

[0521] Examples of such playback devices PROD_D include DVD players, BD players, HDD players, etc. (in which case, the output terminal PROD_D4 connected to a TV receiver, etc., is the main destination for the moving images). Other examples include TV receivers (in which the display PROD_D3 is the main destination for the moving images), digital signage (also called electronic billboards, electronic bulletin boards, etc., where the display PROD_D3 or the transmitter PROD_D5 is the main destination for the moving images), desktop PCs (in which the output terminal PROD_D4 or the transmitter PROD_D5 is the main destination for the moving images), laptop or tablet PCs (in which the display PROD_D3 or the transmitter PROD_D5 is the main destination for the moving images), and smartphones (in which the display PROD_D3 or the transmitter PROD_D5 is the main destination for the moving images).

[0522] (Hardware implementation and software implementation)

[0523] Furthermore, each of the aforementioned motion picture decoding device 31 and motion picture encoding device 11 can be implemented in hardware using logic circuits formed on an integrated circuit (IC chip), or in software using a CPU (Central Processing Unit).

[0524] In the latter case, the aforementioned devices include: a CPU that executes commands for programs that perform various functions; a ROM (Read Only Memory) that stores the programs; a RAM (Random Access Memory) that expands the programs; and a memory that stores the programs and various data, etc., and other storage devices (recording media). Furthermore, the objective of embodiments of the present invention is to achieve this by supplying the aforementioned devices with a recording medium containing program code (executable program, intermediate code program, source program) of the software that performs the aforementioned functions, i.e., the control program of the aforementioned devices, in a computer-readable manner; the computer (or CPU, MPU (Microprocessor Unit)) reads the program code recorded on the recording medium and executes it.

[0525] As recording media, the following can be used: tapes, cassette tapes, etc.; disks including floppy disks (registered trademark) / hard disks, CD-ROMs (Compact Disc Read-Only Memory), MO discs (Magneto-Optical Disc), MD discs (Mini Disc), DVDs (Digital Versatile Disc), CD-R discs (CD Recordable), Blu-ray discs (Blu-ray Disc); cards (including memory cards) / optical cards; semiconductor memory such as mask ROMs / EPROMs (Erasable Programmable Read-Only Memory) / EEPROMs (Electrically Erasable and Programmable Read-Only Memory, registered trademark) / flash memory ROMs; or logic circuits such as PLDs (Programmable Logic Devices) and FPGAs (Field Programmable Gate Arrays).

[0526] Furthermore, the aforementioned devices can be configured to connect to a communication network and supply the program code via the communication network. This communication network need only be capable of transmitting program code and is not particularly limited. For example, the Internet, intranet, extranet, LAN (Local Area Network), ISDN (Integrated Services Digital Network), VAN (Value-Added Network), CATV (Community Antenna television / Cable Television) communication network, Virtual Private Network, telephone line network, mobile communication network, satellite communication network, etc., can be used. Moreover, the transmission medium constituting this communication network need only be a medium capable of transmitting program code and is not limited to a specific configuration or type. For example, it can be used in wired networks such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line transmission, cable TV lines, telephone lines, and ADSL (Asymmetric Digital Subscriber Line) lines, as well as in wireless networks such as IrDA (Infrared Data Association), infrared (like remote controls), Bluetooth, IEEE 802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance), portable telephone networks, satellite lines, and terrestrial digital broadcasting networks. It should be noted that embodiments of the present invention can also be implemented as computer data signals embedded in a carrier wave, which electronically transmit the aforementioned program code.

[0527] The embodiments of the present invention are not limited to the embodiments described above, and various modifications can be made within the scope of the claims. That is, embodiments obtained by combining technical solutions with appropriate modifications within the scope of the claims are also included within the technical scope of the present invention.

[0528] Industrial availability

[0529] The embodiments of the present invention are preferably applied to a moving image decoding apparatus for decoding encoded data obtained by encoding image data, and to a moving image encoding apparatus for generating encoded data obtained by encoding image data. Furthermore, they are preferably applied to a data structure of encoded data generated by the moving image encoding apparatus and referenced by the moving image decoding apparatus.

Claims

1. A motion picture decoding apparatus for decoding encoded data, the motion picture decoding apparatus comprising: The prediction parameter derivation unit is configured to determine whether the first syntax element of a specified NAL unit type indicates instantaneous decoding refresh of the image; and The parameter decoding unit is configured to decode the following items: The second syntactic element in the sequence parameter set, where, When the first syntax element indicates instantaneous decoding and image refresh, the second syntax element specifies whether the list of reference images exists in the slice header. The third syntax element in the image parameter set specifies whether the reference image list information exists in the image header or the slice header, and The fourth syntax element, wherein the fourth syntax element is the number of items in the reference image list structure, in, When the value of the third syntax element is equal to 0 and the first syntax element does not indicate instantaneous decoding refresh of the image, or when the value of the third syntax element is equal to 0 and the value of the second syntax element is equal to 1, the list of reference images exists in the slice header, and When the value of the third syntax element is equal to 0 and the first syntax element indicates instantaneous decoding refresh of the image, the value of the fourth syntax element in the reference image list is inferred to be equal to 0, wherein the fourth syntax element is defined by a variable i that is equal to 0 or 1.

2. A motion picture encoding apparatus for encoding data, the motion picture encoding apparatus comprising: The prediction parameter derivation unit is configured to determine whether the first syntax element of a specified NAL unit type indicates instantaneous decoding refresh of the image; and The parameter decoding unit is configured to encode the following items: The second syntactic element in the sequence parameter set, where, When the first syntax element indicates instantaneous decoding and image refresh, the second syntax element specifies whether the list of reference images exists in the slice header. The third syntax element in the image parameter set specifies whether the reference image list information exists in the image header or the slice header, and The fourth syntax element, wherein the fourth syntax element is the number of items in the reference image list structure, in, When the value of the third syntax element is equal to 0 and the first syntax element does not indicate instantaneous decoding refresh of the image, or when the value of the third syntax element is equal to 0 and the value of the second syntax element is equal to 1, the list of reference images exists in the slice header, and When the value of the third syntax element is equal to 0 and the first syntax element indicates instantaneous decoding refresh of the image, the value of the fourth syntax element in the reference image list is inferred to be equal to 0, wherein the fourth syntax element is defined by a variable i that is equal to 0 or 1.

3. A motion picture coding method for encoding data, the motion picture coding method comprising: Determine whether the first syntax element of the specified NAL unit type indicates instantaneous decoding refresh of the image; as well as The second syntax element in the sequence parameter set is encoded, wherein, in the case where the first syntax element indicates instantaneous decoding and image refresh, the second syntax element specifies whether the list of reference images exists in the slice header. The third syntax element in the image parameter set is encoded, wherein the third syntax element specifies whether the reference image list information exists in the image header or the slice header, and The fourth syntax element is encoded, wherein the fourth syntax element is the number of items in the reference image list structure. in, When the value of the third syntax element is equal to 0 and the first syntax element does not indicate instantaneous decoding refresh of the image, or when the value of the third syntax element is equal to 0 and the value of the second syntax element is equal to 1, the list of reference images exists in the slice header, and When the value of the third syntax element is equal to 0 and the first syntax element indicates instantaneous decoding refresh of the image, the value of the fourth syntax element in the reference image list is inferred to be equal to 0, wherein the fourth syntax element is defined by a variable i that is equal to 0 or 1.

4. A non-transitory computer-readable medium storing a bitstream generated by encoding moving images, the non-transitory computer-readable medium comprising: Store the bitstream, wherein the bitstream comprises: The second syntax element in the sequence parameter set, wherein, in the case where the first syntax element indicates instantaneous decoding and image refresh, the second syntax element specifies whether the list of reference images exists in the slice header of the slice. The third syntax element in the image parameter set specifies whether the reference image list information exists in the image header or the slice header, and The fourth syntax element, wherein the fourth syntax element is the number of items in the reference image list structure, in, When the value of the third syntax element is equal to 0 and the first syntax element does not indicate instantaneous decoding refresh of the image, or when the value of the third syntax element is equal to 0 and the value of the second syntax element is equal to 1, the list of reference images exists in the slice header, and When the value of the third syntax element is equal to 0 and the first syntax element indicates instantaneous decoding refresh of the image, the value of the fourth syntax element in the reference image list is inferred to be equal to 0, wherein the fourth syntax element is defined by a variable i that is equal to 0 or 1.