Moving picture decoding apparatus and moving picture decoding method
By selecting a reference image list structure for the image or slice unit, and allowing the L1 prediction motion vector difference to be set to zero, the problem of reduced coding efficiency in B slices is solved, achieving efficient encoding and decoding.
Patent Information
- Application Number
- CN202180024998.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-02
- Filing Date
- 2021-03-26
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2041-03-26
AI Technical Summary
In motion vector encoding and decoding of B slices, the existing technology of setting the motion vector difference value of L1 prediction to zero will lead to a decrease in encoding efficiency, especially when there are multiple slices in an image, the encoding efficiency may deteriorate significantly.
The mode employs bidirectional prediction with image unit switching and L1 prediction motion vector difference set to zero. It selects a reference image list structure for image or slice unit selection, allowing the mode to be applied when both reference image lists contain only previous or future images.
Even if an image contains multiple slices, it can be encoded and decoded efficiently, improving encoding efficiency.
Smart Images

Figure CN115398917B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to a moving image encoding apparatus, a moving image decoding apparatus, and a prediction image generating apparatus. BACKGROUND
[0002] In order to efficiently transmit or record a moving image, a moving image encoding apparatus that generates encoded data by encoding a moving image, and a moving image decoding apparatus that generates a decoded image by decoding the encoded data are used.
[0003] As a specific moving image encoding method, for example, H.264 / AVC (Advanced Video Coding) or H.265 / HEVC (High-Efficiency Video Coding) can be cited.
[0004] In the above moving image encoding method, an image (picture) constituting a moving image is managed by a hierarchical structure including a slice, a coding tree unit (CTU), a coding unit (also referred to as a coding unit (CU)), and a transform unit (TU) obtained by dividing the image, the coding tree unit obtained by dividing the slice, the coding unit obtained by dividing the coding tree unit, and the transform unit obtained by dividing the coding unit, and is encoded / decoded per CU.
[0005] Further, in the above moving image encoding method, a prediction image is generally generated on the basis of a local decoded image obtained by encoding / decoding an input image, and a prediction error (also referred to as a "difference image" or a "residual image") obtained by subtracting the prediction image from the input image (original image) is encoded. As a method of generating a prediction image, inter-picture prediction (inter prediction) and intra-picture prediction (intra prediction) can be cited.
[0006] Further, as a technique of moving image encoding and decoding in recent years, Non-Patent Literature 1 can be cited.
[0007] In Non-Patent Literature 1, the following method is adopted: in encoding and decoding of a motion vector of a B slice, a mode in which a difference value of a motion vector of L1 prediction is set to zero is defined in a picture header.
[0008] PRIOR ART DOCUMENT
[0009] NON-PATENT LITERATURE
[0010] Non-Patent Literature 1: "Versatile Video Coding (Draft 8)", JVET-P2001-vE, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 2020-03-12 SUMMARY
[0011] PROBLEMS TO BE SOLVED BY THE INVENTION
[0012] However, in the method described in Non-Patent Literature 1, in encoding and decoding of the motion vector of the B slice, there is defined in the picture header a mode in which the difference value of the motion vector of the L1 prediction is set to zero. However, if this mode is set, the symmetrical motion vector difference mode is not run regardless of the reference picture list structure. Therefore, when there are multiple slices in one picture, there is a problem that the coding efficiency can be significantly degraded depending on the selected reference picture.
[0013] SOLUTION TO THE PROBLEM
[0014] The motion image decoding apparatus of one aspect of the present application is characterized in that:
[0015] has a mode in which the difference of the motion vector of the L1 prediction, which can be switched in a picture unit, of the bidirectional prediction is set to zero,
[0016] when all the short-term reference pictures referable in the two reference picture lists are either previous or future pictures, the mode in which the difference of the motion vector of the L1 prediction is set to zero can be applied.
[0017] By so configuring, even if there are multiple slices in one picture, efficient encoding and decoding are possible.
[0018] The motion image decoding apparatus of one aspect of the present application is characterized in that:
[0019] has a prediction section that decodes a reference picture list structure including multiple reference picture lists, and selects a reference picture list from the reference picture list structure in a picture unit or a slice unit,
[0020] when the prediction section selects a reference picture list in a picture unit, a mode in which the difference of the motion vector of the L1 prediction of the bidirectional prediction is set to zero can be applied in a picture unit,
[0021] when the prediction section selects a reference picture list in a slice unit, the mode in which the difference of the motion vector of the L1 prediction is set to zero can be applied in a slice unit.
[0022] This configuration allows for efficient encoding and decoding even if an image contains multiple slices.
[0023] Invention Effects
[0024] According to one aspect of the present invention, the above-mentioned problems can be solved. Attached Figure Description
[0025] Figure 1 This is a schematic diagram showing the configuration of the image transmission system of this embodiment.
[0026] Figure 2 This diagram illustrates the configuration of a transmitting device equipped with the motion picture encoding apparatus of this embodiment and a receiving device equipped with the motion picture decoding apparatus. PROD_A represents the transmitting device equipped with the motion picture encoding apparatus, and PROD_B represents the receiving device equipped with the motion picture decoding apparatus.
[0027] Figure 3 This diagram illustrates the configuration of a recording apparatus equipped with a motion picture encoding device according to this embodiment, and a playback apparatus equipped with a motion picture decoding device. PROD_C represents the recording apparatus equipped with the motion picture encoding device, and PROD_D represents the playback apparatus equipped with the motion picture decoding device.
[0028] Figure 4 It is a diagram representing the hierarchical structure of the encoded stream data.
[0029] Figure 5 This is a concept diagram representing an example of a reference image and a list of reference images.
[0030] Figure 6 This is a schematic diagram showing the configuration of a motion picture decoding device.
[0031] Figure 7 This is a flowchart illustrating the general operation of a motion picture decoding device.
[0032] Figure 8 This is a diagram illustrating the configuration of the merge candidates.
[0033] Figure 9 This is a schematic diagram showing the structure of the inter-frame prediction parameter derivation unit.
[0034] Figure 10 This is a schematic diagram showing the structure of the merged prediction parameter derivation unit and the AMVP prediction parameter derivation unit.
[0035] Figure 11 This is a schematic diagram showing the structure of the inter-frame prediction image generation unit.
[0036] Figure 12 This is a block diagram illustrating the structure of a motion picture encoding device.
[0037] Figure 13 This is a schematic diagram showing the structure of the inter-frame prediction parameter coding unit.
[0038] Figure 14 This is a schematic diagram showing the structure of the intra-frame prediction parameter coding unit.
[0039] Figure 15 It is a diagram representing part of the syntax of Sequence Parameter Set (SPS) and Picture Parameter Set (PPS).
[0040] Figure 16 This is a diagram representing part of the syntax for the image header PH.
[0041] Figure 17 This is a diagram representing part of the syntax of the slice header.
[0042] Figure 18 This is a diagram showing the syntax for defining a list of reference images using `ref_pic_lists()` and defining a structure for a list of reference images using `ref_pic_list_struct(listIdx, rplsIdx)`.
[0043] Figure 19 This is a diagram representing a part of the syntax of the coding unit (CU).
[0044] Figure 20 This is a diagram illustrating the syntax of the encoding unit CU in this embodiment.
[0045] Figure 21 This is a diagram illustrating the syntax of the encoding unit CU in this embodiment.
[0046] Figure 22 This is a diagram illustrating a portion of the syntax for the image header PH and slice header in this embodiment. Detailed Implementation
[0047] (First Implementation)
[0048] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.
[0049] Figure 1 This is a schematic diagram showing the configuration of the image transmission system 1 of this embodiment.
[0050] Image transmission system 1 is a system that transmits encoded streams of images with different resolutions that have been converted, decodes the transmitted encoded streams to reverse convert the images back to their original resolutions, and then displays them. Image transmission system 1 is configured to include a resolution conversion device (resolution conversion unit) 51, a moving image encoding device (image encoding device) 11, a network 21, a moving image decoding device (image decoding device) 31, a resolution inverse conversion device (resolution inverse conversion unit) 61, and a moving image display device (image display device) 41.
[0051] The resolution conversion device 51 converts the resolution of the image T contained in the moving image, and supplies the variable resolution moving image signal containing images with different resolutions to the image encoding device 11. Furthermore, the resolution conversion device 51 supplies information indicating whether the image has undergone resolution conversion to the moving image encoding device 11. When this information indicates resolution conversion, the moving image encoding device sets the resolution conversion information ref_pic_resampling_enabled_flag (described later) to 1, and includes it in the Sequence Parameter Set (SPS) of the encoded data for encoding.
[0052] The motion picture encoding device 11 inputs an image T whose resolution has been converted.
[0053] Network 21 transmits the encoded stream Te generated by the motion picture encoding device 11 to the motion picture decoding device 31. Network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. Network 21 is not necessarily limited to a two-way communication network; it can also be a one-way communication network that transmits broadcast waves, such as digital terrestrial broadcasting or satellite broadcasting. In addition, network 21 can also be replaced by a storage medium that records the encoded stream Te, such as DVD (Digital Versatile Disc: registered trademark) or BD (Blu-ray Disc: registered trademark).
[0054] The motion picture decoding device 31 decodes each encoded stream Te transmitted by the network 21, generates a variable resolution decoded image signal, and supplies it to the resolution inverse conversion device 61.
[0055] When the resolution conversion information contained in the variable resolution decoded image signal indicates a resolution conversion, the resolution inverse conversion device 61 performs an inverse conversion on the resolution-converted image to generate a decoded image signal of the original size.
[0056] The moving image display device 41 displays all or part of one or more decoded images Td represented by the decoded image signal input from the resolution inverter. The moving image display device 41 includes display devices such as liquid crystal displays (LCDs) and organic EL (electroluminescence) displays. Examples of display types include fixed types, mobile types, and HMDs (head-mounted displays). Furthermore, when the moving image decoding device 31 has high processing power, it displays high-quality images; when it has lower processing power, it displays images that do not require high processing power but have high display capabilities.
[0057] <operator>
[0058] The following describes the operators used in this specification.
[0059] >> is for right bit shift, << is for left bit shift, & is for bitwise AND, | is for bitwise OR, |= is the OR substitution operator, and || represents logical OR.
[0060] x? y: z is a ternary operator that takes y when x is true (other than 0) and takes z when x is false (0).
[0061] Clip3(a, b, c) is a function that truncates c to a value greater than a and less than b. It returns a if c < a, b if c > b, and c otherwise (where a <= b).
[0062] abs(a) is a function that returns the absolute value of a.
[0063] Int(a) is a function that returns the integer value of a.
[0064] floor(a) is a function that returns the largest integer less than or equal to a.
[0065] ceil(a) is a function that returns the smallest integer greater than or equal to a.
[0066] a / d represents a division operation where d is divided by a (the decimal point is discarded).
[0067] min(a, b) represents the smaller of a and b.
[0068] <Structure of encoded stream Te>
[0069] Before describing in detail the motion picture encoding device 11 and motion picture decoding device 31 of this embodiment, the data structure of the encoded stream Te generated by the motion picture encoding device 11 and decoded by the motion picture decoding device 31 will be described.
[0070] Figure 4 This is a diagram representing the hierarchical structure of the data in the encoded stream Te. The encoded stream Te typically contains a sequence, as well as multiple images that make up the sequence. Figure 4 The diagram shows an encoded video sequence representing a defined sequence SEQ, an encoded picture representing a defined picture PICT, an encoded slice representing a defined slice S, encoded slice data representing defined slice data, encoded tree units contained in the encoded slice data, and encoded units contained in the encoded tree units.
[0071] (Encoded video sequence)
[0072] In the encoded video sequence, a set of data is specified for reference by the moving image decoding device 31 to decode the sequence SEQ of the object being processed. For example... Figure 4 As shown, the sequence SEQ includes the video parameter set (VPS), the sequence parameter set (SPS), the picture parameter set (PPS), the adaptation parameter set (APS), the picture (PICT), and the supplemental enhancement information (SEI).
[0073] In a motion picture composed of multiple layers, the Video Parameter Set (VPS) specifies a set of common coding parameters for multiple motion pictures, as well as a set of coding parameters associated with the multiple layers contained in the motion picture.
[0074] The Sequence Parameter Set (SPS) specifies a set of encoding parameters that the moving image decoding device 31 refers to in order to decode the object sequence. For example, it specifies the width and height of the image. Alternatively, there can be multiple SPSs. In this case, any one of the multiple SPSs is selected from the PPS.
[0075] (Encoded image)
[0076] In the encoded image, a set of data is specified for reference by the motion picture decoding device 31 to decode the image PICT of the object being processed. The image PICT is as follows: Figure 4 As shown, it includes the image header PH, slice 0 to slice NS-1 (NS is the total number of slices contained in the image PICT).
[0077] Hereinafter, when it is not necessary to distinguish between individual slices 0 to NS-1, ellipses will sometimes be used to indicate subscripts. Furthermore, the same applies to the data contained in the encoded stream Te described below, i.e., other data with subscripts.
[0078] (Encoded slice)
[0079] In the encoded slice, a set of data is specified for the motion picture decoding device 31 to reference in order to decode the slice S of the object being processed. For example... Figure 4 As shown, a slice contains a slice header and slice data.
[0080] The slice header contains a group of encoded parameters referenced by the moving image decoding device 31 to determine the decoding method for the object slice. The slice type specification information (slice_type) is an example of the encoded parameters contained in the slice header.
[0081] As slice types that can be specified through slice type specification information, examples include (1) I slices that use only intra-frame prediction during encoding, (2) P slices that use single prediction (L0 prediction) or intra-frame prediction during encoding, and (3) B slices that use single prediction (using only L0 prediction of reference image list 0 or only L1 prediction of reference image list 1), double prediction, or intra-frame prediction during encoding. Furthermore, inter-frame prediction is not limited to single or double prediction; more reference images can be used to generate the prediction image. Hereinafter, when referred to as P or B slices, it means a slice containing blocks that can use inter-frame prediction.
[0082] Additionally, the slice header can also contain a reference to the picture parameter set (PPS) (pic_parameter_set_id).
[0083] (Encoded slice data)
[0084] In the encoded slice data, a set of data is specified for reference by the motion picture decoding device 31 to decode the slice data of the object being processed. Slice data such as... Figure 4 The encoded slice header, as shown, contains CTUs. A CTU is a fixed-size (e.g., 64×64) block that makes up a slice, sometimes also called the Largest Coding Unit (LCU).
[0085] (Coding Tree Unit)
[0086] Figure 4The specification defines a set of data for the motion picture decoding device 31 to reference in decoding the CTU of the processing object. The CTU is divided into basic units of encoding processing, namely coding units (CUs), through recursive quadtree (QT) segmentation, binary tree (BT) segmentation, or ternary tree (TT) segmentation. BT segmentation and TT segmentation are collectively referred to as multi-tree (MT) segmentation. The nodes of the tree structure obtained by recursive quadtree segmentation are called coding nodes. The intermediate nodes of quadtrees, binary trees, and ternary trees are coding nodes, and the CTU itself is defined as the top-level coding node.
[0087] In CT, CT information includes a CU segmentation flag (split_cu_flag) indicating whether CT segmentation is performed, a QT segmentation flag (qt_split_cu_flag) indicating whether QT segmentation is performed, an MT segmentation direction flag (mtt_split_cu_vertical_flag) indicating the segmentation direction of MT segmentation, and an MT segmentation type flag (mtt_split_cu_binary_flag) indicating the segmentation type of MT segmentation. `split_cu_flag`, `qt_split_cu_flag`, `mtt_split_cu_vertical_flag`, and `mtt_split_cu_binary_flag` are transmitted for each encoding node.
[0088] Luminance and chromatic aberration can also use different trees. The type of tree is indicated by `treeType`. For example, when luminance (Y, cIdx = 0) and chromatic aberration (Cb / Cr, cIdx = 1, 2) use a common tree, the common single tree is represented by `treeType = SINGLE_TREE`. When luminance and chromatic aberration use two different trees (DUAL trees), the luminance tree is represented by `treeType = DUAL_TREE_LUMA`, and the chromatic aberration tree is represented by `treeType = DUAL_TREE_CHROMA`.
[0089] (Encoding unit)
[0090] Figure 4 The document specifies a set of data for the motion picture decoding device 31 to reference in order to decode the encoding unit of the object being processed. Specifically, the CU includes a CU header (CUH), prediction parameters, transformation parameters, quantization transformation coefficients, etc. The CU header specifies the prediction mode, etc.
[0091] Predictive processing can be performed at the CU (Computer Unit) level or at the sub-CU level, where the CU is further divided into sub-CUs. When the size of the CU and the sub-CUs are equal, there is one sub-CU within the CU. When the size of the CU is larger than the size of the sub-CUs, the CU is divided into sub-CUs. For example, when the CU is 8×8 and the sub-CUs are 4×4, the CU is divided into four sub-CUs: one horizontally and one vertically.
[0092] There are two types of prediction (prediction modes): intra-frame prediction and inter-frame prediction. Intra-frame prediction is prediction within the same image, while inter-frame prediction refers to prediction processing between different images (such as between display times or between layers).
[0093] The conversion / quantization process is performed in units of CU, but the quantization conversion coefficients can also be entropy encoded in sub-blocks such as 4×4.
[0094] (Prediction parameters)
[0095] The predicted image is derived from the prediction parameters attached to the block. These prediction parameters include those for intra-frame prediction and inter-frame prediction.
[0096] The prediction parameters for inter-frame prediction will be explained below. The inter-frame prediction parameters are composed of the prediction list using flags predFlagL0 and predFlagL1, reference image indices refIdxL0 and refIdxL1, and motion vectors mvL0 and mvL1. predFlagL0 and predFlagL1 are flags indicating whether the reference image list (L0 list and L1 list) is used; a value of 1 indicates that the corresponding reference image list is used. Additionally, in this specification, when denoted as "flag indicating whether it is XX," a flag that is not 0 (e.g., 1) is considered to be XX, and a flag that is 0 is considered not to be XX. In logical negation, logical product, etc., 1 is treated as true, and 0 is treated as false (the same applies below). However, in actual devices and methods, other values may be used for true and false values.
[0097] Syntax elements used to derive inter-frame prediction parameters include, for example, the affine flag (affine_flag), merge flag (merge_fag), merge index (merge_idx), MMVD flag (mmvd_flag) used in merge mode, the inter-frame prediction identifier (inter_pred_idc) used to select the reference image used in AMVP mode, the reference image index (refIdxLX), the prediction vector index (mvp_LX_idx) used to derive motion vectors, the difference vector (mvdLX), and the motion vector precision mode (amvr_mode).
[0098] (Refer to the image list)
[0099] The reference image list is a list containing the reference images stored in the reference image memory 306. Figure 5 This is a concept diagram representing an example of a reference image and a list of reference images. In Figure 5 In a conceptual diagram representing a reference image, rectangles represent images, arrows indicate the reference relationships between images, the horizontal axis represents time, and I, P, and B within the rectangles represent intra-frame images, single-prediction images, and double-prediction images, respectively. The numbers within the rectangles indicate the decoding order. As shown in the figure, the decoding order of the images is I0, P1, B2, B3, B4, and the display order is I0, B3, B2, B4, P1. Figure 5 The diagram shows an example of a reference image list for image B3 (the object image). A reference image list is a list of candidates for reference images; an image (slice) can have more than one reference image list. In the example, object image B3 has reference image lists L0 (RefPicList0) and L1 (RefPicList1). In each CU, refIdxLX specifies which image in the reference image list RefPicListX (X = 0 or 1) is actually referenced. The diagram shows an example where refIdxL0 = 2 and refIdxL1 = 0. Furthermore, LX is the notation used when not distinguishing between L0 and L1 predictions. Later, LX will be replaced with L0 and L1 to distinguish the parameters corresponding to the L0 and L1 lists.
[0100] (Merge forecast and AMVP forecast)
[0101] There are two methods for decoding (encoding) prediction parameters: merge prediction mode and AMVP (Advanced Motion Vector Prediction) mode. The `merge_flag` is a flag used for their identification. In merge prediction mode, the prediction list is not included in the encoded data using the flag `predFlagLX`, the reference image index `refIdxLX`, and the motion vector `mvLX`, but is derived from the prediction parameters of processed neighboring blocks. In AMVP mode, `inter_pred_idc`, `refIdxLX`, and `mvLX` are included in the encoded data. Additionally, `mvLX` is encoded as `mvp_LX_idx`, which identifies the prediction vector `mvpLX`, and the difference vector `mvdLX`. Besides merge prediction mode, there are also affine prediction mode and MMVD prediction mode.
[0102] `inter_pred_idc` is a value representing the type and number of reference images, taking any one of `PRED_L0`, `PRED_L1`, or `PRED_BI`. `PRED_L0` and `PRED_L1` represent single prediction using one reference image managed in the L0 list and L1 list, respectively. `PRED_BI` represents double prediction using two reference images managed in the L0 list and L1 list.
[0103] merge_idx is an index indicating whether any of the prediction parameters from the prediction parameter candidates (merge candidates) derived from the processed block should be used as prediction parameters for the object block.
[0104] (Motion Vector)
[0105] mvLX represents the shift amount between blocks on two different images. The prediction vector and difference vector associated with mvLX are called mvpLX and mvdLX, respectively.
[0106] (The inter-frame prediction identifier inter_pred_idc and the prediction list utilize the predFlagLX)
[0107] The relationship between inter_pred_idc and predFlagL0 and predFlagL1 is as follows: they can be converted to each other.
[0108] inter_pred_idc=(predFlagL1<<1)+predFlagL0
[0109] predFlagL0 = inter_pred_idc & 1
[0110] predFlagL1=inter_pred_idc>>1
[0111] Furthermore, inter-frame prediction parameters can use either the prediction list utilization flag or the inter-frame prediction identifier. Additionally, the decision using the prediction list utilization flag can be replaced by the decision using the inter-frame prediction identifier. Conversely, the decision using the inter-frame prediction identifier can also be replaced by the decision using the prediction list utilization flag.
[0112] (Determination of biPred prediction)
[0113] The flag biPred indicating whether a prediction is dual-prediction can be derived from whether both prediction lists are 1. For example, it can be derived from the following formula.
[0114] biPred=(predFlagL0==1&&predFlagL1==1)
[0115] Alternatively, biPred can also be derived based on whether the inter-frame prediction identifier is a value indicating the use of two prediction lists (see image). For example, it can be derived using the following formula.
[0116] biPred=(inter_pred_idc==PRED_BI)? 1:0
[0117] (Composition of a motion picture decoding device)
[0118] The motion picture decoding device 31 of this embodiment ( Figure 6 The composition of ) will be explained.
[0119] The motion picture decoding device 31 is configured to include an entropy decoding unit 301, a parameter decoding unit (predictive image decoding device) 302, a loop filter 305, a reference image memory 306, a prediction parameter memory 307, a prediction image generation unit (predictive image generation device) 308, an inverse quantization / inverse conversion unit 311, an adder 312, and a prediction parameter derivation unit 320. Alternatively, there is a configuration that, in conjunction with the motion picture encoding device 11 described later, excludes the loop filter 305 from the motion picture decoding device 31.
[0120] The parameter decoding unit 302 also includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit). The CU decoding unit 3022 also includes a TU decoding unit 3024. These can be collectively referred to as decoding modules. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, PPS, and APS, and the slice header (slice information) from the encoded data. The CT information decoding unit 3021 decodes the CT from the encoded data. The CU decoding unit 3022 decodes the CU from the encoded data. When the TU contains prediction errors, the TU decoding unit 3024 decodes the QP (Quantization Parameter) update information (quantization correction value) and the quantization prediction error (residual_coding) from the encoded data.
[0121] When the mode is other than skip mode (skip_mode == 0), the TU decoding unit 3024 decodes the QP update information and quantization prediction error from the encoded data. More specifically, when skip_mode == 0, the TU decoding unit 3024 decodes the flag cu_cbp, which indicates whether the target block contains quantization prediction error; when cu_cbp is 1, it decodes the quantization prediction error. When cu_cbp is not present in the encoded data, it is set to 0.
[0122] The TU decoding unit 3024 decodes the index mts idx representing the conversion basis from the encoded data. Furthermore, the TU decoding unit 3024 decodes the index stIdx representing the utilization of the secondary conversion and the conversion basis from the encoded data. When stIdx is 0, it indicates that the secondary conversion is not applied; when stIdx is 1, it indicates one conversion in a set (pair) of secondary conversion bases; and when stIdx is 2, it indicates the other conversion in the aforementioned pair.
[0123] Furthermore, the TU decoding unit 3024 can also decode the sub-block conversion flag cu_sbt_flag. When cu_sbt_flag is 1, the CU is divided into multiple sub-blocks, and residual decoding is performed only on a specific sub-block. Moreover, the TU decoding unit 3024 can also decode the flag cu_sbt_quad_flag indicating whether the number of sub-blocks is 4 or 2, the flag cu_sbt_horizontal_flag indicating the division direction, and the flag cu_sbt_pos_flag indicating sub-blocks containing non-zero conversion coefficients.
[0124] The prediction image generation unit 308 is configured to include an inter-frame prediction image generation unit 309 and an intra-frame prediction image generation unit 310.
[0125] The prediction parameter derivation unit 320 is configured to include an inter-frame prediction parameter derivation unit 303 and an intra-frame prediction parameter derivation unit 304.
[0126] Furthermore, the following describes examples of using CTU and CU as processing units, but it is not limited to these examples; processing can also be performed in sub-CU units. Alternatively, CTU and CU can be renamed blocks, and sub-CUs can be renamed sub-blocks, resulting in processing in blocks or sub-blocks.
[0127] The entropy decoding unit 301 performs entropy decoding on the externally input encoded stream Te, decoding each code (syntax element). Entropy encoding can be performed in two ways: using a context (probability model) adaptively selected based on the type of syntax element or surrounding conditions to perform variable-length encoding of syntax elements; and using a predefined table or formula to perform variable-length encoding of syntax elements. The former, CABAC (Context Adaptive Binary Arithmetic Coding), stores the CABAC state of the specified context (the type of the dominant symbol (0 or 1) and the probability state index pStateIdx) in memory. The entropy decoding unit 301 initializes all CABAC states using the beginning of segments (tiles, CTU lines, slices). The entropy decoding unit 301 converts the syntax elements into binary strings and decodes each bit of the binary string. When using a context, the context index ctxInc is derived for each bit of the syntax element, the bits are decoded using the context, and the CABAC state of the used context is updated. Bits without context are decoded with equal probability (EP, bypass), omitting the ctxInc export and CABAC state. The decoded syntax elements contain prediction information for generating the predicted image and prediction errors for generating the difference image.
[0128] The entropy decoding unit 301 outputs the decoded code to the parameter decoding unit 302. The decoded code refers to, for example, prediction modes such as predMode, merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX, and amvr_mode. The control over which code to decode is based on the instruction from the parameter decoding unit 302.
[0129] (Basic Process)
[0130] Figure 7 This is a flowchart illustrating the general operation of the motion picture decoding device 31.
[0131] (S1100: Parameter Set Information Decoding) The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS from the encoded data.
[0132] (S1200: Slice Information Decoding) The header decoding unit 3020 decodes the slice header (slice information) from the encoded data.
[0133] Hereinafter, the motion picture decoding device 31 extracts the decoded image of each CTU by repeatedly performing S1300 to S5000 processing on each CTU contained in the object picture.
[0134] (S1300: CTU Information Decoding) The CT information decoding unit 3021 decodes the CTU from the encoded data.
[0135] (S1400: CT Information Decoding) The CT information decoding unit 3021 decodes the CT data from the encoded data.
[0136] (S1500: CU Decoding) The CU decoding unit 3022 performs S1510 and S1520 to decode the CU from the encoded data.
[0137] (S1510: CU Information Decoding) The CU decoding unit 3022 decodes CU information, prediction information, TU segmentation flag split_transform_flag, CU residual flags cbf_cb, cbf_cr, cbf_luma, etc. from the encoded data.
[0138] (S1520: TU Information Decoding) When the TU contains prediction error, the TU decoding unit 3024 decodes the QP update information, quantization prediction error, and transformation index mts_idx from the encoded data. Furthermore, the QP update information is the difference between the predicted value of the quantization parameter QP and the predicted value of the quantization parameter qPpred.
[0139] (S2000: Predictive Image Generation) The predictive image generation unit 308 generates a predictive image for each block contained in the object CU based on the prediction information.
[0140] (S3000: Inverse quantization / inverse conversion) The inverse quantization / inverse conversion unit 311 performs inverse quantization / inverse conversion processing on each TU contained in the object CU.
[0141] (S4000: Decoded Image Generation) The addition unit 312 generates a decoded image of the target CU by adding the predicted image supplied by the predicted image generation unit 308 to the prediction error supplied by the inverse quantization / inverse conversion unit 311.
[0142] (S5000: Loop Filter) The loop filter 305 performs deblocking filtering, SAO (Sample Adaptive Offset), ALF (Adaptive Loop Filter) and other loop filtering on the decoded image to generate the decoded image.
[0143] (The structure of the inter-frame prediction parameter derivation section)
[0144] Figure 9A schematic diagram illustrating the configuration of the inter-frame prediction parameter derivation unit 303 according to this embodiment is shown. The inter-frame prediction parameter derivation unit 303 derives inter-frame prediction parameters based on the syntax elements input from the parameter decoding unit 302 and referring to the prediction parameters stored in the prediction parameter memory 307. Furthermore, the inter-frame prediction parameters are output to the inter-frame prediction image generation unit 309 and the prediction parameter memory 307. The inter-frame prediction parameter derivation unit 303 and its internal elements, including the AMVP prediction parameter derivation unit 3032, the merged prediction parameter derivation unit 3036, the affine prediction unit 30372, the MMVD prediction unit 30373, the GPM prediction unit 30377, the DMVR unit 30537, and the MV (Motion Vector) addition unit 3038, are common units in moving image coding devices and moving image decoding devices; therefore, they can also be collectively referred to as motion vector derivation units (motion vector derivation devices).
[0145] The scale parameter derivation unit 30378 derives the horizontal scaling ratio RefPicScale[i][j][0] and the vertical scaling ratio RefPicScale[i][j][1] of the reference image, as well as RefPicIsScaled[i][j] indicating whether the reference image is scaled. Here, i indicates whether the list of reference images is an L0 list or an L1 list, and j is derived as the value of the L0 reference image list or the L1 reference image list in the following manner.
[0146] RefPicScale[i][j][0] =
[0147] ((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthLRefPicScale[i][j][1]=
[0148] ((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightLRefPicIsScaled[i][j]=
[0149] (RefPicScale[i][j][0]!=(1<<14))||(RefPicScale[i][j][1]!=(1<<14))
[0150] Here, the variable `PicOutputWidthL` is the value used to calculate the horizontal scaling ratio when referencing the encoded image. It is obtained by subtracting the left and right offset values from the horizontal pixel count of the encoded image's brightness. The variable `PicOutputHeightL` is the value used to calculate the vertical scaling ratio when referencing the encoded image. It is obtained by subtracting the vertical offset values from the vertical pixel count of the encoded image's brightness. The variable `fRefWidth` serves as the value of `PicOutputWidthL` for the reference image in list value j of list i, and the variable `fRefHight` serves as the value of `PicOutputHeightL` for the reference image in list value j of list i.
[0151] When affine_flag is 1, which indicates affine prediction mode, the affine prediction unit 30372 derives the inter-frame prediction parameters for sub-block units.
[0152] When mmvd_flag is 1, which indicates MMVD prediction mode, the MMVD prediction unit 30373 derives inter-frame prediction parameters from the merged candidate and difference vector derived by the merged prediction parameter derivation unit 3036.
[0153] When GPMFlag is 1, which indicates GPM (Geometric Partitioning Mode) prediction mode, the GPM prediction unit 30377 derives the GPM prediction parameters.
[0154] When merge_flag is 1, indicating the merge prediction mode, export merge_idx and output it to the merge prediction parameter export section 3036.
[0155] When merge_flag is 0, which indicates AMVP prediction mode, the AMVP prediction parameter export unit 3032 exports mvpLX from inter_pred_idc, refIdxLX, or mvp_LX_idx.
[0156] (MV Addition Department)
[0157] In the MV addition section 3038, the exported mvpLX and mvdLX are added together to export mvLX.
[0158] (Affine Prediction Department)
[0159] As an affine prediction unit 30372, 1) it derives the motion vectors of two control points CP0, CP1 or three control points CP0, CP1, CP2 of the object block; 2) it derives the affine prediction parameters of the object block; and 3) it derives the motion vectors of each sub-block from the affine prediction parameters.
[0160] When performing merged affine prediction, the motion vectors cpMvLX[] of each control point CP0, CP1, and CP2 are derived from the motion vectors of adjacent blocks of the object block. When performing inter-frame affine prediction, cpMvLX[] of each control point is derived from the sum of the prediction vectors of each control point CP0, CP1, and CP2 and the difference vector mvdCpLX[] derived from the coded data.
[0161] (Consolidated Forecast)
[0162] Figure 10 A schematic diagram illustrating the configuration of the merge prediction parameter derivation unit 3036 in this embodiment is shown. The merge prediction parameter derivation unit 3036 includes a merge candidate derivation unit 30361 and a merge candidate selection unit 30362. Furthermore, the merge candidates are configured to include prediction parameters (predFlagLX, mvLX, refIdxLX) and are stored in a merge candidate list. An index is assigned to the merge candidates stored in the merge candidate list according to specific rules.
[0163] The merge candidate derivation unit 30361 directly derives merge candidates using the motion vectors and refIdxLX of the adjacent blocks after decoding. In addition, the merge candidate derivation unit 30361 can also employ spatial merge candidate derivation processing, temporal merge candidate derivation processing, pairwise merge candidate derivation processing, and zero merge candidate derivation processing, as described later.
[0164] As part of the spatial merging candidate derivation process, the merging candidate derivation unit 30361 reads the prediction parameters stored in the prediction parameter memory 307 according to specific rules and sets them as merging candidates. The method of specifying the reference image is, for example, the prediction parameters associated with each of the adjacent blocks (e.g., all or part of the blocks that are connected to the left A1, right B1, upper right B0, lower left A0, and upper left B2 of the object block) within a predefined range. Each merging candidate is referred to as A1, B1, B0, A0, and B2. Here, A1, B1, B0, A0, and B2 are motion information derived from the blocks containing the following coordinates. Figure 8 In the object image, the configuration of the merge candidates is shown as positions A1, B1, B0, A0, and B2.
[0165] A1: (xCb-1, yCb+cbHeight-1)
[0166] B1: (xCb + cbWidth-1, yCb-1)
[0167] B0: (xCb+cbWidth.yCb-1)
[0168] A0: (xCb-1, yCb+cbHeight)
[0169] B2: (xCb-1, yCb-1)
[0170] Set the top-left coordinate of the object block to (xCb, yCb), the width to cbWidth, and the height to cbHeight.
[0171] As part of the time-merging export process, merge candidate export section 30361 is as follows: Figure 8 As shown in the corresponding image, the prediction parameters of the lower right CBR of the object block or the block C in the reference image containing the central coordinate are read from the prediction parameter memory 307 as merging candidate Col and stored in the merging candidate list mergeCandList[].
[0172] Generally, block CBRs are preferentially added to mergeCandList[]. When a CBR does not have a motion vector (e.g., an intra-predicted block) or when the CBR is located outside the image, the motion vector of block C is added to the prediction vector candidate. By adding the motion vectors of co-position blocks with high probability of different motions as prediction candidates, the selection of prediction vectors increases, and coding efficiency is improved.
[0173] When ph_temporal_mvp_enabled_flag is 0, or cbWidth*cbHeight is 32 or less, the parimetric motion vector mvLXCol of the object block is set to 0, and the availability flag availableFlagLXCol of the parimetric block is set to 0.
[0174] In other cases (where SliceTemporalMvpEnabledFlag is 1), the following processing is performed.
[0175] For example, the position of C (xColCtr, yColCtr) and the position of CBR (xColCBr, yColCBr) can also be derived using the following formula by merging candidate derived parts 30361.
[0176] xColCtr=xCb+(cbWidth>>1)
[0177] yColCtr=yCb+(cbHeight>>1)
[0178] xColCBr=xCb+cbWidth
[0179] yColCBr=yCb+cbHeight
[0180] If CBR is available, the candidate COL for merging is derived using the motion vector of CBR. If CBR is not available, the COL is derived using C. Furthermore, availableFlagLXCol is set to 1. Alternatively, the collocated_ref_idx notified in the slice header can also be referenced in the image.
[0181] The paired candidate derivation part derives the paired candidate avgK from the average of the two merged candidates (p0Cand, p1Cand) that have been stored in mergeCandList and stores them in mergeCandList[].
[0182] mvLXavgK[0]=(mvLXp0Cand[0]+mvLXp1Cand[0]) / 2
[0183] mvLXavgK[1]=(mvLXp0Cand[1]+mvLXp1Cand[1]) / 2
[0184] The candidate merging derivation unit 30361 derives zero merging candidates Z0, ..., ZM for refIdxLX with X components and Y components of 0 and mvLX, and stores them in the merging candidate list.
[0185] The order in which mergeCandList[] is stored is, for example, spatial merge candidates (A1, B1, B0, A0, B2), temporal merge candidates Col, paired candidates avgK, and zero merge candidates ZK. Additionally, unusable reference blocks (such as intra-frame prediction blocks) are not stored in the merge candidate list.
[0186] i=0
[0187] if(availableFlagA1)
[0188] mergeCandList[i++] = A1
[0189] if(availableFlagB1)
[0190] mergeCandList[i++] = B1
[0191] if(availableFlagB0)
[0192] mergeCandList[i++] = B0
[0193] if(availableFlagA0)
[0194] mergeCandList[i++] = A0
[0195] if(availableFlagB2)
[0196] mergeCandList[i++] = B2
[0197] if(availableFlagCol)
[0198] mergeCandList[i++] = Col
[0199] if(availableFlagAvgK)
[0200] mergeCandList[i++] = avgK
[0201] if (i < MaxNumMergeCand)
[0202] mergeCandList[i++] = ZK
[0203] Merge candidate selection unit 30362 selects merge candidate N, represented by merge_idx, from the merge candidates included in the merge candidate list according to the following formula.
[0204] N = mergeCandList[merge_idx]
[0205] Here, N represents the label of the merging candidate, which can be A1, B1, B0, A0, B2, Col, avgK, ZK, etc. The motion information of the merging candidate shown by label N is represented by (mvLXN[0], mvLXN[0]), predFlagLXN, refIdxLXN.
[0206] The selected (mvLXN[0], mvLXN[0]), predFlagLXN, and refIdxLXN are selected as inter-frame prediction parameters for the target block. The merging candidate selection unit 30362 stores the inter-frame prediction parameters of the selected merging candidates in the prediction parameter memory 307 and outputs them to the inter-frame prediction image generation unit 309.
[0207] (DMVR)
[0208] Next, the DMVR (Decoder-side Motion Vector Refinement) processing performed by the DMVR unit 30375 will be explained. For an object CU, when merge_flag is 1 or skip_flag is 1, the DMVR unit 30375 corrects the mvLX of that object CU derived by the merging prediction unit 30374 using a reference image. Specifically, when the prediction parameters derived by the merging prediction unit 30374 are dual predictions, if they correspond to two reference images, the motion vector is corrected using the prediction image derived from the motion vector. The corrected mvLX is then supplied to the inter-frame prediction image generation unit 309.
[0209] Furthermore, in the derivation of the flag dmvrFlag that specifies whether DMVR processing is performed, one of the several conditions for setting dmvrFlag to 1 includes the values of RefPicIsScaled[0][refIdxL0] being 0 and RefPicIsScaled[1][refIdxL1] being 0. When the value of dmvrFlag is set to 1, DMVR processing is performed by the DMVR unit 30375.
[0210] Furthermore, in the export of the flag dmvrFlag that specifies whether DMVR processing is performed, one of the multiple conditions for setting dmvrFlag to 1 is that ciip_flag is 0, meaning that intra-inter frame composition processing is not applicable.
[0211] Furthermore, in the derivation of the flag dmvrFlag specifying whether DMVR processing is performed, one of the several conditions for setting dmvrFlag to 1 includes a flag indicating the existence of coefficient information for weighted prediction of L0 luminance (described later), i.e., luma_weight_l0_flag[i], being 0, and a flag indicating the existence of coefficient information for weighted prediction of L1 luminance, i.e., luma_weight_l1_flag[i], being 0. When the value of dmvrFlag is set to 1, DMVR processing is performed by the DMVR unit 30375.
[0212] Furthermore, in the derivation of the flag dmvrFlag specifying whether DMVR processing is performed, one of the multiple conditions for setting dmvrFlag to 1 may include: luma_weight_l0_flag[i] being 0, luma_weight_l1_flag[i] being 0, and the flag indicating the existence of L0 prediction weighted coefficient information for color difference (described later) being chroma_weight_l0_flag[i] being 0, and the flag indicating the existence of L1 prediction weighted coefficient information for color difference (described later) being chroma_weight_l1_flag[i] being 0. When the value of dmvrFlag is set to 1, DMVR processing is performed by the DMVR unit 30375.
[0213] (Prof)
[0214] Furthermore, if the value of RefPicIsScaled[0][refIdxLX] is 1, or the value of RefPicIsScaled[1][refIdxLX] is 1, then the value of cbProfFlagLX is set to FALSE (=0). Here, cbProfFlagLX is a flag that specifies whether to perform prediction refinement (PROF) for affine prediction.
[0215] (AMVP Prediction)
[0216] Figure 10 A schematic diagram showing the configuration of the AMVP prediction parameter derivation unit 3032 in this embodiment is shown. The AMVP prediction parameter derivation unit 3032 has a vector candidate derivation unit 3033 and a vector candidate selection unit 3034. The vector candidate derivation unit 3033 derives prediction vector candidates based on the motion vectors of adjacent blocks where decoding has ended, stored in the prediction parameter memory 307, and stores them in the prediction vector candidate list mvpListLX[].
[0217] The vector candidate selection unit 3034 selects the motion vector mvpListLX[mvp_LX_idx] represented by mvp_LX_idx from the predicted vector candidates of mvpListLX[]. The vector candidate selection unit 3034 outputs the selected mvpLX to the MV addition unit 3038.
[0218] (MV Addition Department)
[0219] The MV addition unit 3038 adds the mvpLX input from the AMVP prediction parameter derivation unit 3032 to the decoded mvdLX to calculate mvLX. The addition unit 3038 outputs the calculated mvLX to the inter-frame prediction image generation unit 309 and the prediction parameter memory 307.
[0220] mvLX[0]=mvpLX[0]+mvdLX[0]
[0221] mvLX[1]=mvpLX[1]+mvdLX[1]
[0222] (Detailed classification of sub-block merging)
[0223] This paper summarizes the types of prediction processing related to sub-block merging. As mentioned above, they can be broadly categorized into merge prediction and AMVP prediction.
[0224] The combined forecasts are further categorized as follows.
[0225] • Normal merge prediction (block-based merge prediction)
[0226] • Sub-block merging prediction
[0227] Sub-block merging predictions are further categorized as follows.
[0228] • Sub-block prediction (ATMVP)
[0229] Affine prediction
[0230] • Inferred affine prediction
[0231] • Constructed affine prediction
[0232] On the other hand, AMVP predictions are divided into the following categories.
[0233] • AMVP (Parallel)
[0234] ·MVD Affine Prediction
[0235] MVD affine prediction is further divided into the following categories.
[0236] ·4-parameter MVD affine prediction
[0237] ·6-parameter MVD affine prediction
[0238] In addition, MVD affine prediction refers to affine prediction used for decoding difference vectors.
[0239] In sub-block prediction, similar to the time-merging export process, the availability of the corresponding sub-block COL of the target sub-block is determined using availableFlagSbCol. If available, the prediction parameters are exported. AvailableFlagSbCol is set to 0 when at least the aforementioned SliceTemporalMvpEnabledFlag is 0.
[0240] MMVD prediction (Merge with Motion Vector Difference) can be classified as either merge prediction or AMVP prediction. In the former case, if merge_flag = 1, then mmvd_flag and MMVD-related syntax elements are decoded; in the latter case, if merge_flag = 0, then mmvd_flag and MMVD-related syntax elements are decoded.
[0241] The loop filter 305 is a filter set within the encoding loop, and is used to improve image quality by removing block distortion or ringing distortion. The loop filter 305 performs deblocking filtering, sample adaptive offset (SAO), and adaptive loop filtering (ALF) on the decoded image of the CU generated by the adder 312.
[0242] The image memory 306 stores the decoded images of the CU in predefined locations for each object image and each object CU.
[0243] The prediction parameter memory 307 stores the prediction parameters in a pre-defined location for each CTU or CU. Specifically, the prediction parameter memory 307 stores parameters decoded by the parameter decoding unit 302 and parameters derived by the prediction parameter derivation unit 320, etc.
[0244] The parameters derived by the prediction parameter derivation unit 320 are input to the prediction image generation unit 308. Furthermore, the prediction image generation unit 308 reads a reference image from the reference image memory 306. In the prediction mode indicated by predMode, the prediction image generation unit 308 generates a prediction image of a block or sub-block using the parameters and the reference image (reference image block). Here, a reference image block refers to a set of pixels on the reference image (usually rectangular, hence called a block), and is the region referenced for generating the prediction image. (Inter-frame prediction image generation unit 309)
[0245] When predMode indicates inter-frame prediction mode, the inter-frame prediction image generation unit 309 uses the inter-frame prediction parameters input from the inter-frame prediction parameter derivation unit 303 and the reference image to generate a predicted image of a block or sub-block through inter-frame prediction.
[0246] Figure 11This is a schematic diagram showing the configuration of the inter-frame prediction image generation unit 309 included in the prediction image generation unit 308 of this embodiment. The inter-frame prediction image generation unit 309 is configured to include a motion compensation unit (prediction image generation device) 3091 and a compositing unit 3095. The compositing unit 3095 is configured to include an intra-frame / inter-frame compositing unit 30951, a GPM compositing unit 30952, a BDOF unit 30954, and a weighted prediction unit 3094.
[0247] (Motion compensation)
[0248] The motion compensation unit 3091 (interpolated image generation unit 3091) reads a reference block from the reference image memory 306 based on the inter-frame prediction parameters (predFlagLX, refIdxLX, mvLX) input from the inter-frame prediction parameter derivation unit 303, thereby generating an interpolated image (motion-compensated image). The reference block is the block on the reference image RefPicLX specified by refIdxLX at a position shifted by mvLX relative to the target block. Here, when mvLX is not of integer precision, a filter called motion compensation filtering is performed to generate pixels at fractional positions, thus generating the interpolated image.
[0249] The motion compensation unit 3091 first derives the integer position (xInt, yInt) and phase (xFrac, yFrac) corresponding to the coordinates (x, y) within the prediction block using the following formula.
[0250] xInt=xPb+(mvLX[0]>>(log2(MVPREC)))+x
[0251] xFrac = mvLX[0] & (MVPREC-1)
[0252] yInt=yPb+(mvLX[1]>>(log2(MVPREC)))+y
[0253] yFrac=mvLX[1]&(MVPREC-1)
[0254] Here, (xPb, yPb) are the top-left coordinates of a block of size bW*bH, where x = 0…bW-1 and y = 0…bH-1. MVPREC represents the precision of mvLX (1 / MVPREC pixel precision). For example, MVPREC = 16.
[0255] The motion compensation unit 3091 performs horizontal interpolation on the reference image refImg using interpolation filtering to derive a temporary image temp[][]. The following ∑ is the sum of k for k = 0..NTAP-1, shift1 is a normalization parameter that adjusts the range of values, and offset1 = 1 << (shift1-1).
[0256] temp[x][y]=(∑mcFilter[xFrac][k]*refImg[xInt+k-NTAP / 2+1][yInt]+offsetl)>>shift1
[0257] Next, the motion compensation unit 3091 derives the interpolated image Pred[][] by performing vertical interpolation on the temporary image temp[][]. Here, ∑ is the sum of k for k = 0..NTAP-1, shift2 is a normalization parameter that adjusts the range of values, and offset2 = 1 << (shift2-1).
[0258] Pred[x][y]=(∑mcFilter[yFrac][k]*temp[x][y+k-NTAP / 2+1]+offset2)>>shift2
[0259] In addition, when it is a double prediction, the above Pred[][] (called the interpolated images PredL0[][] and PredL1[][]) are derived according to each L0 list and L1 list, and the interpolated image Pred[][] is generated based on PredL0[][] and PredL1[][].
[0260] In addition, the motion compensation unit 3091 has the function of scaling and interpolating the image based on the horizontal scaling ratio RefPicScale[i][j][0] and the vertical scaling ratio RefPicScale[i][j][1] of the reference image derived by the scale parameter derivation unit 30378.
[0261] The compositing unit 3095 includes an intra-frame / inter-frame compositing unit 30951, a GPM compositing unit 30952, a weight prediction unit 3094, and a BDOF unit 30954.
[0262] (Interpolation filter processing)
[0263] The following describes the interpolation filtering process performed by the prediction image generation unit 308, and the interpolation filtering process when the size of the reference image changes in a single sequence for the purpose of applying the above-mentioned resampling. Alternatively, this process can also be performed, for example, by the motion compensation unit 3091.
[0264] When the value of RefPicIsScaled[i][j] input from the inter-frame prediction parameter derivation unit 303 indicates that the reference image has been scaled, the prediction image generation unit 308 switches multiple filtering coefficients and performs interpolation filtering.
[0265] (Intra-frame and inter-frame compositing)
[0266] The intra-frame / inter-frame synthesis unit 30951 generates a prediction image by weighted sum of the inter-frame prediction image and the intra-frame prediction image.
[0267] The pixel values of the predicted image, predSamplesComb[x][y], can be derived as follows when the flag ciip_flag indicating whether intra-frame / inter-frame synthesis processing is applicable is 1.
[0268] predSamplesComb[x][y]=(w*predSamplesIntra[x][y]
[0269] +(4-w)*predSamplesInter[x][y]+2)>>2
[0270] Here, predSamplesIntra[x][y] is the intra-predicted image, defined by planar prediction. predSamplesInter[x][y] is the reconstructed inter-predicted image.
[0271] The weight w is derived as follows.
[0272] When the bottommost block to the left of the object coding block and the rightmost block to the top of the object coding block are both intra-frame, w is set to 3.
[0273] In all other cases, i.e. when the bottommost block to the left of the object coding block and the rightmost block to the top of the object coding block are not within the same frame, w is set to 1.
[0274] In all other cases, w is set to 2.
[0275] (GPM synthesis treatment)
[0276] The GPM synthesis unit 30952 generates a predicted image using the aforementioned GPM prediction.
[0277] (BDOF Prediction)
[0278] Next, the details of the BDOF prediction (Bi-Directional Optical Flow, BDOF processing) performed by the BDOF unit 30954 will be explained. In the dual prediction mode, the BDOF unit 30954 generates a prediction image by referring to two prediction images (a first prediction image and a second prediction image) and a gradient correction term.
[0279] (Weight Prediction)
[0280] The weight prediction unit 3094 generates a prediction image pbSamples for the block based on the interpolated image predSamplesLX.
[0281] First, the variable `weightedPredFlag`, indicating whether weighted prediction processing is performed, is derived as follows: When `slice_type` equals `P`, `weightedPredFlag` is set to be equal to `pps_weighted_pred_flag` defined in PPS. Otherwise, when `slice_type` equals `B`, `weightedPredFlag` is set to be equal to `pps_weighted_bipred_flag && (!dmvrFlag)` defined in PPS.
[0282] Hereinafter, bcw_idx is the weight index of the double prediction with weights in units of CU. When bcw_idx is not notified, it is set to bcw_idx = 0. bcwIdx is set to bcwldxN of the nearby block in merge prediction mode and to bcw_idx of the target block in AMVP prediction mode.
[0283] If the value of the variable weightedPredFlag is equal to 0, or the value of the variable bcwIdx is 0, then the predicted image pbSamples is derived as follows, as in the usual predicted image processing.
[0284] When the prediction list uses one of the flags (predFlagL0 or predFlagL1) as 1 (single prediction) (no weighted prediction), the following process is performed to match predSamplesLX (LX is L0 or L1) with the number of pixel bits bitDepth.
[0285] pbSamples[x][y]=Clip3(0, (1<<bitDepth)-1, (predSamplesLX[x][y]+offset1)>>shift1)
[0286] Here, shift1 = 14-bit Depth, offset1 = 1 << (shift1 - 1). PredLX is the interpolated image predicted by L0 or L1.
[0287] Furthermore, when the prediction list uses flags (predFlagL0 and predFlagL1) set to 1 (double prediction PRED_BI) and does not use weighted prediction, the following process is performed to average predSamplesL0 and predSamplesL1 and match them with the number of pixel bits.
[0288] pbSamp1es[x][y]=Clip3(0, (1<<bitDepth)-1, (predSamplesL0[x][y]+predSamplesL1[x][y]+offset2)>>shift2)
[0289] Here, shift2 = 15-bitDepth, offset2 = 1 << (shift2-1).
[0290] If the value of the variable weightedPredFlag is equal to 1 and the value of the variable bcwIdx is equal to 0, then the prediction is processed as a weighted prediction, and the predicted image pbSamples is exported in the following manner.
[0291] The variable shift1 is set to equal Max(2, 14-bitDepth). The variables log2Wd, o0, o1, w0, and w1 are derived as follows.
[0292] If the brightness is 0 when cIdx is 0, the following method applies.
[0293] log2Wd=luma_log2_weight_denom+shift1
[0294] w0 = LumaWeightL0[refIdxL0]
[0295] w1 = LumaWeightL1[refIdxL1]
[0296] o0=luma_offset_10[refIdxL0]<<(bitDepth-8)
[0297] o1=luma_offset_11[refIdxL1]<<(bitDepth-8)
[0298] In cases other than (color difference where cIdx is not 0), the following method applies.
[0299] log2Wd=ChromaLog2WeightDenom+shift1
[0300] w0=ChromaWeightL0[refIdxL0][cIdx-1]
[0301] w1=ChromaWeightL1[refIdxL1][cIdx-1]
[0302] o0=ChromaOffsetL0[refldxL0][cIdx-1]<<<(bitDepth-8)
[0303] o1=ChromaOffsetL1[refIdxL1][cIdx-1]<<(bitDepth-8)
[0304] The pixel values pbSamples[x][y] of the predicted images for x = 0..nCbW-1 and y = 0..nCbH-1 are derived as follows.
[0305] Next, if predFlagL0 equals 1 and predFlagL1 equals 0, the pixel values pbSamples[x][y] of the predicted image are derived as follows.
[0306] if (log2Wd>=1)
[0307] pbSamples[x][y]=Clip3(0, (1<<bitDepth)-1,
[0308] ((predSamplesL0[x][y]*w0+2^(log2Wd-1))>>log2Wd)+o0)
[0309] else
[0310] pbSamples[x][y]=Clip3(0, (1<<bitDepth)-1, predSamplesL0[x][y]*w0+o0)
[0311] In addition, if predFlagL0 is 0 and predFlagL1 is 1, the pixel values pbSamples[x][y] of the predicted image are derived as follows.
[0312] if (log2Wd>=1)
[0313] pbSamples[x][y]=Clip3(0, (1<<bitDepth)-1,
[0314] ((predSamplesL1[x][y]*w1+2^(log2Wd-1))>>log2Wd)+o1)
[0315] else
[0316] pbSamples[x][y]=Clip3(0, (1<<bitDepth)-1, predSamplesL1[x][y]*w1+o1)
[0317] In addition, if predFlagL0 equals 1 and predFlagL1 equals 1, the pixel values pbSamples[x][y] of the predicted image are derived as follows.
[0318] pbSamples[x][y]=Clip3(0, (1<<bitDepth)-1,
[0319] (predSamplesL0[x][y]*w0+predSamplesL1[x][y]*w1+
[0320] ((o0+o1+1)<<log2Wd))>>(1og2Wd+1))
[0321] (BCW Prediction)
[0322] BCW (Bi-prediction with CU-level Weights) is a prediction method that can switch between pre-determined weighting coefficients based on CU level.
[0323] Input two variables nCbW and nCbH specifying the width and height of the current coding block, two permutations of (nCbW)x(nCbH) predSamplesL0 and predSamplesL1, flags predFlagL0 and predFlagL1 indicating whether to use the prediction list, reference image indexes refIdxL0 and refIdxL1, BCW prediction index bcw_idx, and variable cIdx specifying the index of the luminance and chromaticity components. Perform BCW prediction processing and output the pixel values of the predicted image of the permutation pbSamples of (nCbW)x(nCbH).
[0324] When the `sps_bcw_enabled_flag` indicating whether to use the prediction at the SPS level is `TURE`, the variable `weightedPredFlag` is 0, the reference images shown by the two reference image indices `refIdxL0` and `refIdxL1` have no weighted prediction coefficients, and the coding block size is below a certain size, the CU-level grammar's `bcw_idx` is explicitly notified to substitute this value into the variable `bcwIdx`. If `bcw_idx` does not exist, then 0 is substituted into the variable `bcwIdx`.
[0325] When the variable bcwIdx is 0, the pixel values of the predicted image are derived as follows.
[0326] pbSamples[x][y]=Clip3(0,(1< <bitDepth)-1,
[0327] (predSamplesL0[x][y]+predSamplesL1[x][y]+offset2)>>shift2) In other cases (where bcwIdx is not 0), the following method applies.
[0328] The variable w1 is set to be equal to bcwWLut[bcwIdx]. bcwWLut[k] = {4, 5, 3, 10, -2}.
[0329] The variable w0 is set to (8-w1). Furthermore, the pixel values of the predicted image are derived as follows: pbSamples[x][y] = Clip3(0, (1 << bitDepth) - 1, ...
[0330] (w0*predSamplesL0[x][y]+
[0331] w1*predSamplesL1[x][y]+offset3)>>(shift2+3))
[0332] When BCW prediction is used in AMVP prediction mode, the inter-frame prediction parameter decoding unit 303 decodes the bcw_idx and sends it to the BCW unit 30955. Furthermore, when BCW prediction is used in merge prediction mode, the inter-frame prediction parameter decoding unit 303 decodes the merge index merge_idx, and the merge candidate derivation unit 30361 derives the bcwIdx of each merge candidate. Specifically, the merge candidate derivation unit 30361 uses the weight coefficients of adjacent blocks used in the derivation of the merge candidates as the weight coefficients of the merge candidates used in the object block. That is, in merge mode, the previously used weight coefficients are inherited as the weight coefficients of the object block.
[0333] (Intra-frame prediction image generation unit 310)
[0334] When predMode indicates intra-prediction mode, the intra-prediction image generation unit 310 performs intra-prediction using the intra-prediction parameters input from the intra-prediction parameter derivation unit 304 and the reference pixels read from the reference image memory 306.
[0335] The inverse quantization / inverse conversion unit 311 inverse quantizes the quantization conversion coefficients input from the parameter decoding unit 302 to obtain the conversion coefficients.
[0336] The adder 312 adds the predicted image of the block input from the predicted image generation unit 308 to the prediction error input from the inverse quantization / inverse conversion unit 311, pixel by pixel, to generate the decoded image of the block. The adder 312 stores the decoded image of the block in the reference image memory 306, and outputs it to the loop filter 305.
[0337] The inverse quantization / inverse conversion unit 311 inverse quantizes the quantization conversion coefficients input from the parameter decoding unit 302 to obtain the conversion coefficients.
[0338] The adder 312 adds the predicted image of the block input from the predicted image generation unit 308 to the prediction error input from the inverse quantization / inverse conversion unit 311, pixel by pixel, to generate the decoded image of the block. The adder 312 stores the decoded image of the block in the reference image memory 306, and outputs it to the loop filter 305.
[0339] (Composition of a motion picture encoding device)
[0340] Next, the configuration of the motion image encoding device 11 in this embodiment will be described. Figure 12 This is a block diagram illustrating the configuration of the motion picture encoding apparatus 11 according to this embodiment. The motion picture encoding apparatus 11 is configured to include a prediction image generation unit 101, a subtraction unit 102, a conversion / quantization unit 103, an inverse quantization / inverse conversion unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference image memory (reference image storage unit, frame memory) 109, an encoding parameter determination unit 110, a parameter encoding unit 111, a prediction parameter derivation unit 120, and an entropy encoding unit 104.
[0341] The prediction image generation unit 101 generates a prediction image for each CU. The prediction image generation unit 101 includes the inter-frame prediction image generation unit 309 and the intra-frame prediction image generation unit 310, which have already been described, but their descriptions are omitted.
[0342] The subtraction unit 102 subtracts the pixel values of the predicted image of the block input from the prediction image generation unit 101 from the pixel values of the image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the conversion / quantization unit 103.
[0343] The conversion / quantization unit 103 calculates the conversion coefficients by frequency conversion based on the prediction error input from the subtraction unit 102, and derives the quantization conversion coefficients by quantization. The conversion / quantization unit 103 outputs the quantization conversion coefficients to the parameter encoding unit 111 and the inverse quantization / inverse conversion unit 105.
[0344] Inverse quantization / inverse conversion unit 105 and inverse quantization / inverse conversion unit 311 in motion image decoding device 31 Figure 6The same applies, but the explanation is omitted. The calculated prediction error is output to the adder 106.
[0345] The parameter encoding unit 111 includes a header encoding unit 1110, a CT information encoding unit 1111, and a CU encoding unit 1112 (prediction mode encoding unit). The CU encoding unit 1112 also has a TU encoding unit 1114. The general operation of each module will be described below.
[0346] The header encoding unit 1110 encodes parameters such as header information, segmentation information, prediction information, and quantization conversion coefficients.
[0347] The CT information encoding unit 1111 encodes QT and MT (BT, TT) segmentation information, etc. The CU encoding unit 1112 encodes CU information, prediction information, segmentation information, etc.
[0348] When the TU encoding unit 1114 contains prediction error, it encodes QP update information and quantization prediction error.
[0349] The CT information coding unit 1111 and the CU coding unit 1112 supply inter-frame prediction parameters (predMode, merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX), intra-frame prediction parameters (intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_reminder, intra_chroma_pred_mode), quantization conversion coefficients and other syntax elements to the parameter coding unit 111.
[0350] The quantization conversion coefficients and encoding parameters (segmentation information, prediction parameters) are input to the entropy encoding unit 104 by the parameter encoding unit 111. The entropy encoding unit 104 entropy encodes them to generate an encoded stream Te and outputs it.
[0351] The prediction parameter derivation unit 120 is a unit that includes an inter-frame prediction parameter coding unit 112 and an intra-frame prediction parameter coding unit 113, and derives intra-frame prediction parameters and intra-frame prediction parameters from the parameters input from the coding parameter determination unit 110. The derived intra-frame prediction parameters and intra-frame prediction parameters are output to the parameter coding unit 111.
[0352] (The structure of the inter-frame prediction parameter coding unit)
[0353] Inter-frame prediction parameter coding unit 112 as follows Figure 13As shown, the configuration includes a parameter encoding control unit 1121 and an inter-frame prediction parameter derivation unit 303. The configuration of the inter-frame prediction parameter derivation unit 303 is common to that of a moving image decoding apparatus. The parameter encoding control unit 1121 includes a merge index derivation unit 11211 and a vector candidate index derivation unit 11212.
[0354] The merge index derivation unit 11211 derives merge candidates and outputs them to the inter-frame prediction parameter derivation unit 303. The vector candidate index derivation unit 11212 derives prediction vector candidates and outputs them to the inter-frame prediction parameter derivation unit 303 and the parameter encoding unit 111.
[0355] (The structure of the intra-frame prediction parameter coding unit 113)
[0356] Intra-prediction parameter coding unit 113 as follows Figure 14 As shown, it includes a parameter encoding control unit 1131 and an intra-frame prediction parameter derivation unit 304. The configuration of the intra-frame prediction parameter derivation unit 304 is the same as that of a moving image decoding device.
[0357] The parameter encoding control unit 1131 derives IntraPredModeY and IntraPredModeC. Then, it determines intra_luma_mpm_flag by referring to mpmCandList[]. These prediction parameters are output to the intra-prediction parameter derivation unit 304 and the parameter encoding unit 111.
[0358] However, unlike the motion picture decoding device, the encoding parameter determination unit 110 and the prediction parameter memory 108 input the inter-frame prediction parameter derivation unit 303 and the intra-frame prediction parameter derivation unit 304 and output them to the parameter encoding unit 111.
[0359] The addition unit 106 adds the pixel value of the prediction block input from the prediction image generation unit 101 to the prediction error input from the inverse quantization / inverse conversion unit 105 for each pixel to generate a decoded image. The addition unit 106 stores the generated decoded image in the reference image memory 109.
[0360] The loop filter 107 performs deblocking filtering, SAO, and ALF on the decoded image generated by the adder 106. Alternatively, the loop filter 107 may not necessarily include all three filters mentioned above; for example, it may be configured to contain only a deblocking filter.
[0361] The prediction parameter memory 108 stores the prediction parameters generated by the encoding parameter determination unit 110 in a predefined location for each object image and each CU.
[0362] The image memory 109 stores the decoded image generated by the loop filter 107 in a predefined location for each object image and each CU.
[0363] The encoding parameter determination unit 110 selects one set from a plurality of sets of encoding parameters. The encoding parameters refer to the aforementioned QT, BT, or TT segmentation information, prediction parameters, or parameters generated in relation to them as encoding objects. The prediction image generation unit 101 uses these encoding parameters to generate a prediction image.
[0364] The encoding parameter determination unit 110 calculates the amount of information and the RD cost value representing the encoding error for each of the multiple sets. The RD cost value is, for example, the sum of the code amount and the mean square error multiplied by a coefficient λ. The code amount is the amount of information in the encoded stream Te obtained by entropy encoding the quantization error and the encoding parameters. The mean square error is the mean square sum of the prediction errors calculated in the subtraction unit 102. The coefficient λ is a real number greater than a predetermined zero. The encoding parameter determination unit 110 selects the set of encoding parameters with the smallest calculated cost value. The encoding parameter determination unit 110 outputs the determined encoding parameters to the parameter encoding unit 111 and the prediction parameter derivation unit 120.
[0365] Alternatively, a portion of the motion picture encoding device 11 and motion picture decoding device 31 described above, such as the entropy decoding unit 301, parameter decoding unit 302, loop filter 305, prediction image generation unit 308, inverse quantization / inverse conversion unit 311, addition unit 312, prediction parameter derivation unit 320, prediction image generation unit 101, subtraction unit 102, conversion / quantization unit 103, entropy encoding unit 104, inverse quantization / inverse conversion unit 105, loop filter 107, encoding parameter determination unit 110, parameter encoding unit 111, and prediction parameter derivation unit 120, can be implemented by a computer. In this case, it can also be implemented by recording the program for implementing this control function on a computer-readable recording medium, reading the program recorded on the recording medium into a computer system, and executing it. Furthermore, the "computer system" mentioned here refers to a computer system built into either the motion picture encoding device 11 or the motion picture decoding device 31, including hardware such as an operating system (OS) and peripheral devices. Furthermore, "computer-readable recording medium" refers to removable media such as floppy disks, optical disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. Moreover, "computer-readable recording medium" can also refer to a recording medium that dynamically retains a program for a short period, such as a communication line used to transmit a program via a network like the Internet or a telephone line, or a recording medium that retains a program for a certain period, such as volatile memory within a computer system that acts as a server or client in such cases. Furthermore, the aforementioned program can be a program used to implement the above-mentioned functions, or it can be a program that can achieve the above-mentioned functions by combining with programs already recorded in the computer system.
[0366] Furthermore, the motion picture encoding device 11 and motion picture decoding device 31 described above can be implemented, either partially or entirely, in the form of an integrated circuit such as LSI (Large Scale Integration). Each functional block of the motion picture encoding device 11 and motion picture decoding device 31 can be individually processorized, or partially or entirely integrated and processorized. Moreover, the method of integrated circuit implementation is not limited to LSI; it can also be implemented using dedicated circuits or general-purpose processors. Furthermore, when advancements in semiconductor technology have led to the emergence of integrated circuit technologies that replace LSI, these integrated circuits can also be used.
[0367] The above description, with reference to the accompanying drawings, details one embodiment of the present invention. However, the specific configuration is not limited to the above description, and various design changes can be made without departing from the spirit of the present invention.
[0368] (grammar)
[0369] Figure 15 (a) represents a portion of the syntax of the Sequence Parameter Set (SPS) of Non-Patent Document 1.
[0370] The long_term_ref_pics_flag flag indicates whether to use long-term images.
[0371] inter_layer_ref_pics_present_flag is a flag indicating whether inter-layer prediction is used.
[0372] sps_idr_rp1_present_flag is a flag indicating whether the slice header of an IDR image contains a syntax element that refers to a list of images.
[0373] rpl1_same_as_rpl0_flag is a flag indicating whether information for reference image list 1 exists. When rpl1_same_as_rpl0_flag is 1, no information for reference image list 1 exists, indicating that it is the same as num_ref_pic_lists_in_sps[0] and ref_pic_list_struct(0, rplsIdx).
[0374] `sps_smvd_enabled_flag` indicates whether Symmetric Motion Vector Differential Mode (SMVD) is applied to the encoding and decoding of motion vectors. When `sps_smvd_enabled_flag` is 1, it means that Symmetric Motion Vector Differential Mode is applicable. When `sps_smvd_enabled_flag` is 0, it means that Symmetric Motion Vector Differential Mode is not applicable.
[0375] Figure 15 (b) represents a portion of the syntax of the Picture Parameter Set (PPS) in Non-Patent Document 1.
[0376] A value of 1 for `rpl_info_in_ph_flag` indicates whether a list of reference images exists in the image header. `rpl_info_in_ph_flag` equal to 1 indicates the presence of a list of reference images in the image header. A value of 0 for `rpl_info_in_ph_flag` indicates the absence of a list of reference images in the image header, but this might be present in the slice header.
[0377] Figure 16 This is part of the syntax representing the image header PH of non-patent document 1.
[0378] `ph_inter_slice_allowed_flag` is a flag indicating whether a slice within an image is an inter-frame slice. When `ph_inter_slice_allowed_flag` is 0, it means that all slices in the image have a slice type of 2 (I Slice). When `ph_inter_slice_allowed_flag` is 1, it means that at least one slice in the image has a slice type of 0 (B Slice) or 1 (P Slice).
[0379] `mvd_l1_zero_flag` is a flag indicating whether a mode that makes the motion vector difference zero is applicable in L1 prediction for bidirectional prediction. When `mvd_l1_zero_flag` is 1, `mvd_coding()` is not called, and the variables `MvdL1[x0][y0][compIdx]` and `MvdCpL1[x0][y0][cpIdx][compIdx]` representing the difference information of the motion vectors are set to 0. `mvd_coding()` is a syntax structure that informs the difference information of the motion vectors for reference image list 1. When `mvd_l1_zero_flag` is 0, `mvd_coding` is called to encode and decode the required difference information of the motion vectors.
[0380] Figure 17(a) represents a portion of the syntax of the slice header of Non-Patent Document 1. This syntax is decoded, for example, by the parameter decoding unit 302.
[0381] When num_ref_idx_active_override_flag is 1, it means that the syntax element num_ref_idx_active_minus1[0] exists in slices P and B, and the syntax element num_ref_idx_active_minus1[1] exists in slice B. When num_ref_idx_active_override_flag is 0, it means that the syntax element num_ref_idx_active_minus1[i] does not exist in slices P and B. If it does not exist, it is inferred that the value of num_ref_idx_active_override_flag is equal to 1.
[0382] `num_ref_idx_active_minus1[i]` is used to derive the actual number of usable reference images in the reference image list `i`. The actual number of usable reference images, i.e., the variable `NumRefIdxActive[i]`, can be derived from... Figure 17 The method shown in (b) is used for export. The value of hum_ref_idx_active_minus1[i] must be a value between 0 and 14. When the slice is a B slice, and num_ref_idx_active_override_flag is 1, and num_ref_idx_active_minus1[i] does not exist, it is speculated that num_ref_idx_active_minus1[i] is equal to 0.
[0383] Figure 17(b) represents the method for deriving the variable NumRefIdxActive[i] from Non-Patent Document 1 based on the prediction parameter derivation unit 320. The following processing is performed on the reference picture list i (=0,1). When it is a B slice, or a P slice and i=0, if num_ref_idx_active_override_flag is equal to 1, then the value of num_ref_idx_active_minus1[i] plus 1 is substituted into the variable NumRefIdxActive[i]. In other cases (when it is a B slice, or a P slice and i=0, and num_ref_idx_active_override_flag is equal to 0), if the value of num_ref_entries[i][RplsIdx[i]] is greater than or equal to the value of num_ref_idx_default_active_minus1[i] plus 1, then the value of num_ref_idx_default_active_minus1[i] plus 1 is substituted into the variable NumRefIdxActive[i]. In all other cases (when it is a B slice, or a P slice and i = 0, and num_ref_idx_active_override_flag is not equal to 0), the value of num_ref_entries[i][RplsIdx[i]] is substituted into the variable NumRefIdxActive[i]. num_ref_idx_default_active_minus1[i] is the default value of the variable NumRefIdxActive[i] defined in PPS. When it is an I slice or a P slice and i = 1, 0 is substituted into the variable NumRefIdxActive[i].
[0384] Figure 18 (a) indicates the syntax of ref_pic_lists() for the definition of the picture list in Non-Patent Document 1. ref_pic_lists() sometimes exists in the picture header or slice header. If rpl_sps_flag[i] is 1, it means that the picture list i of ref_pic_lists() is derived from one of the ref_pic_list_struct(listIdx, rplsIdx) of SPS. Here, listIdx is equal to i.
[0385] When rpl_sps_flag[i] is 0, it means that the reference image list i is derived based on ref_pic_list_struct(listIdx, rplsIdx). Here, listIdx is equal to i directly contained in ref_pic_lists(). When rpl_sps_flag[i] does not exist, the following method applies. When num_ref_pic_lists_in_sps[i] is 0, the value of rpl_sps_flag[i] is presumed to be 0. In other cases (when num_ref_pic_lists_in_sps[i] is greater than 0), if rpl1_idx_present_flag is equal to 0 and i is equal to 1, then the value of rpl_sps_flag[1] is presumed to be equal to rpl_sps_flag[0].
[0386] `rpl_idx[i]` represents the index of `ref_pic_list_struct(listIdx, rplsIdx)`. `ref_pic_list_struct(listIdx, rplsIdx)` is used to derive the reference image `i`. Here, `listIdx` equals `i`. If it does not exist, the value of `rpl_idx[i]` is presumed to be 0. The value of `rpl_idx[i]` is in the range above 0 and below `num_ref_pic_lists_in_sps[i] - 1`. When `rpl_sps_flag[i]` is 1 and `num_ref_pic_lists_in_sps[i]` is 1, the value of `rpl_idx[i]` is presumed to be 0. When `rpl_sps_flag[i]` is 1 and `rpl1_idx_present_flag` is 0, the value of `rpl_idx[1]` is presumed to be equal to `rpl_idx[0]`. The variable `RplsIdx[i]` is derived as follows.
[0387] Rplsldx[i]=(rpl_sps_flag[i])? rpl_idx[i]:num_ref_pic_lists_in_sps[i]
[0388] Figure 18 (b) represents the syntax for defining the reference picture list structure ref_pic_list_struct(listIdx, rplsIdx) of non-patent document 1.
[0389] The `ref_pic_list_struct(listIdx, rplsIdx)` sometimes exists in the SPS, the picture header, or the slice header. Depending on whether it's included in the SPS, the picture header, or the slice header, the following applies: When it exists in the picture or slice header, `ref_pic_list_struct(listIdx, rplsIdx)` represents the list of reference images `listIdx` for the current picture (including slices). When it exists in the SPS, `ref_pic_list_struct(listIdx, rplsIdx)` represents a candidate list of reference images `listIdx`. Furthermore, the current picture references the list of `ref_pic_list_struct(listIdx, rplsIdx)` contained in the SPS using index values from the picture header or slice header.
[0390] Here, `num_ref_entries[listIdx][rplsIdx]` represents the number of `ref_pic_list_struct`(listIdx, rplsIdx). The value of `num_ref_entries[listIdx][rplsIdx]` is greater than 0 and less than `MaxDpbSize+13`. `MaxDpbSize` is the number of decoded images determined by the profile level.
[0391] ltrp_in_header_flag[listIdx][rplsIdx] is a flag indicating whether a long-term reference image exists in ref_pic_list_struct(listIdx, rplsIdx).
[0392] inter_layer_ref_pic_flag[listIdx][rplsIdx][i] is a flag indicating whether the i-th entry in the reference image list of ref_pic_list_struct(listIdx, rplsIdx) is an inter-layer prediction.
[0393] st_ref_pic_flag[listIdx][rplsIdx][i] is a flag indicating whether the i-th entry in the list of reference pictures of ref_pic_list_struct(listldx, rplsIdx) is a short-term reference picture.
[0394] abs_delta_poc_st[listIdx][rplsIdx][i] is a syntax element used to derive the absolute difference of a POC from a short-term reference image.
[0395] strp_entry_sign_flag[listIdx][rplsIdx][i] is a flag used to derive the sign of a sign.
[0396] rpls_poc_lsb_lt[listIdx][rplsIdx][i] is a syntax element used to derive the POC of the i-th long-term reference image from the list of reference images for ref_pic_list_struct(listIdx, rplsIdx).
[0397] ilrp_idx[listIdx][rplsIdx][i] is a syntax element used to derive the layer information of the i-th layer of the reference image from the list of reference images for ref_pic_list_struct(listIdx, rplsIdx).
[0398] Figure 19 This represents a portion of the syntax of the CU in non-patent document 1. This syntax is decoded, for example, by the parameter decoding unit 302.
[0399] As shown in IF_SYMMVD1, when sps_smvd_enabled_flag is 1, mvd_l1_zero_flag is FALSE, inter_pred_idc[x0][y0] is bidirectional prediction (PRED_BI), inter_affine_flag is FALSE, and the variables RefldxSymL0 and RefIdxSymL1 are greater than -1, then sym_mvd_flag[x0][y0] exists in this CU. sps_smvd_enabled_flag is a flag indicating whether the symmetric motion vector differential mode is applied to the encoding and decoding of motion vectors. mvd_l1_zero_flag is a flag indicating whether the mode that makes the difference of motion vectors zero is applied in the L1 prediction of bidirectional prediction. inter_pred_idc[x0][y0] is the inter-frame prediction identifier. sym_mvd_flag[x0][y0] is a flag indicating whether the symmetric motion vector differential mode is applied. If symmvd_flag[x0][y0] does not exist, it is assumed to be 0. Here, the indexes x0 and y0 represent the pixel position (x0, y0) of the brightness of the upper left corner of the CU relative to the upper left corner of the image.
[0400] The variable RefIdxSymL0 is the reference index value of reference image list 0 for the symmetric motion vector difference mode, and the variable RefIdxSymL1 is the reference index value of reference image list 1 for the symmetric motion vector difference mode.
[0401] When it is a bidirectional prediction with a relationship where the current image is sandwiched between two reference images, the reference index value with the smallest POC difference between the current image and the reference image in reference image list 0 is set as variable RefIdxSymL0, and the reference index value with the smallest POC difference between the current image and the reference image in reference image list 1 is set as variable RefIdxSymL1. If no index value meets the conditions, then -1 is substituted.
[0402] inter_affine_flag[x0][y0] is a flag indicating whether the predicted pixels of the current CU are generated based on motion compensation of the affine model when decoding P or B slices.
[0403] Next, if inter_pred_idc[x0][y0] is not PRED_L1, that is, when using unidirectional or bidirectional prediction with reference image list 0, the motion vector information used for L0 prediction is encoded and decoded. Otherwise, 0 is substituted into the variables MvdL0[x0][y0][0] and MvdL0[x0][y0][1]. In the differential information of the motion vector used for L0 prediction, the variable MvdL0[x0][y0][0] represents the value in the horizontal direction, and the variable MvdL0[x0][y0][1] represents the value in the vertical direction.
[0404] When encoding and decoding motion vector information used for L0 prediction, if NumRefIdxActive[0] is greater than 1 and sym_mvd_flag[x0][y0] is FALSE, then ref_idx_10[x0][y0] exists.
[0405] `ref_idx_10[x0][y0]` represents the reference image index of reference image list 0 in the current CU. When `ref_idx_10[x0][y0]` does not exist, and `sym_mvd_flag[x0][y0]` is estimated to be 1 as follows, `ref_idx_l0[x0][y0]` is set to the value of `RefIdxSymL0`. Otherwise (when `sym_mvd_flag[x0][y0]` is 0), `ref_idx_l0[x0][y0]` is set to 0.
[0406] Next, when inter_pred_idc[x0][y0] is not PRED_L0, that is, when using the one-way or two-way prediction of reference image list 1, the motion vector information used for L1 prediction is encoded and decoded. Otherwise, 0 is substituted into variables MvdL1[x0][y0][0] and MvdL1[x0][y0][1].
[0407] When encoding and decoding motion vector information used for L1 prediction, if NumRefIdxActive[1] is greater than 1 and sym_mvd_flag[x0][y0] is FALSE, then ref_idx_l1[x0][y0] exists.
[0408] `ref_idx_l1[x0][y0]` represents the reference image index of reference image list 0 in the current CU. When `ref_idx_l1[x0][y0]` does not exist, and `sym_mvd_flag[x0][y0]` is estimated to be 1 as follows, `ref_idx_l1[x0][y0]` is set to the value of `RefIdxSymL1`. Otherwise (when `sym_mvd_flag[x0][y0]` is 0), `ref_idx_l1[x0][y0]` is set to 0.
[0409] The variable `MotionModelIdc[x0][y0]` represents the motion compensation model of the CU, where 0 represents typical block motion compensation, 1 represents 4-parameter affine motion compensation, and 2 represents 6-parameter affine motion compensation. Based on the value of `MotionModelIdc[x0][y0]`, the function `mvd_coding(x0, y0, refList, cPIdx)` encodes and decodes the difference information of the motion vectors. Here, the independent variable `refList` provides the values of the list of reference images, and the independent variable `cpIdx` provides the value of the variable `MotionModelIdc[x0][y0]`.
[0410] `mvp_l0_flag[x0][y0]` represents the predicted vector index of reference image list 0. When `mvp_l0_flag[x0][y0]` does not exist, the prediction is 0.
[0411] As shown in IF_SYMMVD2, when mvd_l1_zero_flag is 1 and inter_pred_idc[x0][y0] is PRED_BI (bidirectional prediction), the mode in which the difference information of the motion vectors used for L1 prediction is zero is applied. In this case, 0 is substituted into the variables MvdL1[x0][y0][0] and MvdL1[x0][y0][1]. In addition, 0 is substituted into the difference information of the six motion vectors used for affine prediction: MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1].
[0412] In other cases, the following processing is performed. If sym_mvd_flag[x0][y0] is 1, then -MvdL0[x0][y0][0] is substituted into the variable MvdL1[x0][y0][0], and -MvdL0[x0][y0][1] is substituted into the variable MvdL1[x0][y0][1]. The differential information of the motion vectors predicted by L1 is not encoded or decoded. When sym_mvd_flag[x0][y0] is FALSE, the differential information of the motion vectors used for L1 prediction is encoded and decoded using the function mvd_coding.
[0413] Next, based on the value of MotionModelIdc[x0][y0], the function mvd_coding is used to encode and decode the differential information of the motion vector used for L1 prediction during affine prediction.
[0414] mvp_l1_flag[x0][y0] represents the predicted vector index of reference image list 1. When mvp_l1_flag[x0][y0] does not exist, the prediction is 0.
[0415] A problem with the method described in Non-Patent Document 1 is the definition of `mvd_l1_zero_flag` in the image header. In Non-Patent Document 1, multiple slices exist within an image, and each slice can select a different list of reference images. The encoding efficiency of setting `mvd_l1_zero_flag` to 1 depends on the list of reference images. Therefore, when multiple slices exist within an image, the encoding efficiency may significantly deteriorate depending on the selected reference image.
[0416] Therefore, in this embodiment, a variable `IdenticalDirecitionFlag` is defined to indicate that two reference images are in the same direction as the current image (both are previous images, or both are future images). Furthermore, it is added as one of the conditions for encoding and decoding `mvd_l1_zero_flag`. That is, in this embodiment, the reference image list does not adopt a structure such as the current image sandwiched between two previous and future reference images.
[0417] Specifically, in this embodiment, when as Figure 20As shown in IF_SYMMVD2_A, when mvd_11_zero_flag is 1, the variable IdenticalDirecitionFlag is 1, and inter_pred_idc[x0][y0] is PRED_BI (bidirectional prediction), the differential information for the motion vector used for L1 prediction is 0. In this case, 0 is substituted into the variables MvdL1[x0][y0][0] and MvdL1[x0][y0][1]. Furthermore, 0 is substituted into the differential information MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1]. These syntaxes are encoded, for example, by the prediction parameter derivation unit 120 or the parameter encoding unit 111, and decoded by the parameter decoding unit 302 or the prediction parameter derivation unit 320.
[0418] The variable IdenticalDirecitionFlag is set after the slice header of the P or B image is encoded or decoded and the list of reference images for the slice is made, but before the encoding or decoding of the CU.
[0419] The variable IdenticalDirecitionFlag is exported as follows.
[0420] If the difference DiffPicOrderCnt(aPic,CurrPic) between the POC of all short-term reference images aPic in the reference image list 0 and reference image list 1 of the current slice and the current image CurrPic is less than 0, then IdenticalDirecitionFlag is set to 1.
[0421] In other cases, if the difference between the POCs of CurrPic and aPic, DiffPicOrderCnt(CurrPic, aPic), is less than 0, then IdenticalDirecitionFlag is set to 1.
[0422] In all other cases, set IdenticalDirecitionFlag to 0.
[0423] Here, the variable PicOrderCntVal represents the POC (Picture Order Count) describing the output order from the DPB associated with each image. PicOrderCnt(picX) is a function representing the PicOrderCntVal of image picX, and the function DiffPicOrderCnt(picA, picB) is shown below.
[0424] DiffPicOrderCnt(picA,picB)=PicOrderCnt(picA)-PicOrderCnt(picB)
[0425] If the difference in POC between aPic and CurrPic, DiffPicOrderCnt(aPic, CurrPic), is less than 0, then all short-term reference images aPic become the previous images relative to the current image CurrPic.
[0426] Furthermore, if the difference in POC between CurrPic and aPic, DiffPicOrderCnt(CurrPic, aPic), is less than 0, then all short-term reference images aPic become future images relative to the current image CurrPic.
[0427] Another way to derive the variable `IdenticalDirecitionFlag` is to define it as setting `IdenticalDirecitionFlag` to 1 only if both reference images are previous images relative to the current image. That is, the list of reference images is not structured as if the current image is sandwiched between two previous and future reference images. In this case, it is derived as follows.
[0428] If the difference DiffPicOrderCnt(aPic,CurrPic) between the POC of all reference images aPic in the reference image list 0 and reference image list 1 of the current slice and the current image CurrPic is less than 0, then IdenticalDirecitionFlag is set to 1.
[0429] In all other cases, set IdenticalDirecitionFlag to 0.
[0430] In addition, this mark can also replace the mark previously used as the variable NoBackwadPredFlag in Non-Patent Document 1.
[0431] Alternatively, as another implementation, the variable IdenticalDirecitionFlag can be set after ref_idx_l0[x0][y0] and ref_idX_l1[x0][y0]. In this case, the variable IdenticalDirecitionFlag is derived as follows.
[0432] If the difference DiffPicOrderCnt(aPic,CurrPic) between the two short-term reference images aPic indicated by ref_idx_10[x0][y0] of reference image list 0 and ref_idx_11[x0][y0] of reference image list 1 and the current image CurrPic is less than 0, then IdenticalDirecitionFlag is set to 1.
[0433] In other cases, if the difference between the POCs of CurrPic and aPic, DiffPicOrderCnt(CurrPic, aPic), is less than 0, then IdenticalDirecitionFlag is set to 1.
[0434] In all other cases, set IdenticalDirecitionFlag to 0.
[0435] In addition, as another implementation, the variable IdenticalDirecitionFlag is derived as follows.
[0436] If the difference DiffPicOrderCnt(aPic,CurrPic) between the two short-term reference images aPic indicated by ref_idx_l0[x0][y0] of reference image list 0 and ref_idx_l1[x0][y0] of reference image list 1 and the current image CurrPic is less than 0, then IdenticalDirecitionFlag is set to 1.
[0437] In all other cases, set IdenticalDirecitionFlag to 0.
[0438] Other problems with the method described in Non-Patent Document 1 can be listed as follows: Figure 19As shown, when `mvd_l1_zero_flag` in the image header is 1, even if `sps_smvd_enabled_flag` is 1, the points in the symmetric motion vector difference mode will never run, regardless of the reference image list structure. In Non-Patent Document 1, multiple slices can exist within an image, and each slice can select a different reference image list. Therefore, when multiple slices exist within an image, the coding efficiency may significantly deteriorate depending on the selected reference image.
[0439] Therefore, in this embodiment, as Figure 21 As shown in IF_SYMMVD1_A, the condition that mvd_l1_zero_flag is 1 is removed from the applicable conditions of the symmetric motion vector difference mode, and replaced with the following condition.
[0440] if(sps_smvd_enabled_flag&&
[0441] inter_pred_idc[x0][y0]==PRED_BI&&
[0442] ! inter_affine_flag[x0][y0]&&
[0443] RefIdxSymL0>-1&&RefIdxSymL1>-1)
[0444] That is, when mvd_l1_zero_flag is 0, in this embodiment, sym_mvd_flag[x0][y0] is encoded by the prediction parameter derivation unit 120 or the parameter encoding unit 111 based on the above-mentioned conditional expression. Moreover, sym_mvd_flag[x0][y0] is decoded by the parameter decoding unit 302 or the prediction parameter derivation unit 320.
[0445] Alternatively, the mvd_l1_zero_flag check can be left unremoved, and instead, the check in this implementation can be modified as follows: Figure 21 As shown in IF_SYMMVD2_A, the condition for making the difference of motion vectors zero in L1 prediction applicable to bidirectional prediction is as follows, with the condition that the variable IdenticalDirecitionFlag is 1 added. IdenticalDirecitionFlag is a flag indicating whether two reference images are in the same direction as the current image (both are previous images, or both are future images).
[0446] if(mvd_l1_zero_flag&&IdenticalDirectionFlag&&
[0447] inter_pred_idc[x0][y0]==PRED_BI){
[0448] That is, when two reference images are in the same direction as the current image (both are previous images or both are future images) (IdenticalDirectionFlag is 1), the difference between the motion vectors is made zero in L1 prediction. In this case, 0 is substituted into the variables MvdL1[x0][y0][0] and MvdL1[x0][y0][1]. In addition, 0 is substituted into the difference information of the six motion vectors used for affine prediction: MvdCpL1[x0][y0][0][0], MvdCpL1[x0][y0][0][1], MvdCpL1[x0][y0][1][0], MvdCpL1[x0][y0][1][1], MvdCpL1[x0][y0][2][0], and MvdCpL1[x0][y0][2][1].
[0449] Furthermore, as a condition for the mode that makes the difference of motion vectors zero in L1 prediction (inter_pred_idc[x0][y0]!=PRED_L0) applicable to bidirectional prediction, it can also be shown as follows (IF_SYMMVD2_B).
[0450] if(mvd_l1_zero_flag&&
[0451] inter_pred_idc[x0][y0]==PRED_BI&&
[0452] ! (RefIdxSymL0>-1&&RefIdxSymL1>-1)){
[0453] By setting it this way, the following problem can be solved: when mvd_l1_zero_flag is set to 1 in the image header, even if sps_smvd_enabled_flag is set to 1, the symmetric motion vector difference mode will not run regardless of the reference image list structure.
[0454] In addition, as another implementation of the method for deriving the variable IdenticalDirecitionFlag, the reference index value of the reference image list employing the symmetric motion vector difference mode can also be used in the following formula.
[0455] IdenticalDirecitionFlag=(RefIdxSymL0>-1&&RefIdxSymL1>-1)? 0:1
[0456] Here, the variable RefIdxSymL0 is the reference index value of reference image list 0 for the symmetric motion vector difference mode, and the variable RefIdxSymL1 is the reference index value of reference image list 1 for the symmetric motion vector difference mode.
[0457] Furthermore, another implementation of the method for deriving the variable IdenticalDirecitionFlag can be as follows: If the difference DiffPicOrderCnt(aPic[i], CurrPic) between the active short-term reference image aPic[i] (i = 0, 1) and the current image CurrPic in the reference image list 0 and reference image list 1 of the current slice is less than 0, then IdenticalDirecitionFlag is set to 1.
[0458] In other cases, if DiffPicOrderCnt(CurrPic, aPic[i]) is less than 0, then IdenticalDirecitionFlag is set to 1.
[0459] In all other cases, set IdenticalDirecitionFlag to 0.
[0460] aPic[i] (i = 0, 1) is the actual active short-term reference image defined by the variables NumRefIdxActive[0] and NumRefIdxActive[1] in the reference image list i and reference image list 1 of the current slice.
[0461] As another derived method of the variable IdenticalDirecitionFlag, it can also be defined as setting the variable IdenticalDirecitionFlag to 1 only when the two reference images indicated by ref_idx_l0[x0][y0] and ref_idx_11[x0][y0] are both previous images relative to the current image.
[0462] If the difference between DiffPicOrderCnt(aPic,CurrPic) and the POC of the actual usable active short-term reference image aPic defined by variables NumRefIdxActive[0] and NumRefIdxActive[1] and the current image CurrPic is less than 0 in the reference image list 0 and reference image list 1 of the current slice, then IdenticalDirecitionFlag is set to 1.
[0463] In all other cases, set IdenticalDirecitionFlag to 0.
[0464] In addition, this mark can also replace the mark previously used as the variable NoBackwadPredFlag in Non-Patent Document 1.
[0465] Figure 22 This diagram illustrates the syntax of the image header PH and slice header used in other embodiments of the method described in Non-Patent Document 1 to solve the problem. These syntaxes are encoded, for example, by the prediction parameter derivation unit 120 or the parameter encoding unit 111, and decoded by the parameter decoding unit 302 or the prediction parameter derivation unit 320.
[0466] Figure 22 In the image header PH of (a), when ph_inter_slice_allowed_flag is 1 and rpl_info_in_ph_flag is 1, mvd_l1zero_flag is encoded and decoded. ph_inter_slice_allowed_flag is a flag indicating whether a slice within the image is an inter-frame slice. rpl_info_in_ph_flag is a flag indicating whether a list of reference images exists in the image header.
[0467] Figure 22 In the slice header of (b), when ph_inter_slice_allowed_flag is 0 and slice_type is B, mvd_l1_zero_flag is encoded and decoded. That is, when the image header PH does not contain reference image list information, the slice header contains reference image list information, and it is a B slice, mvd_l1_zero_flag is encoded and decoded.
[0468] By configuring it in this way, mvd_l1_zero_flag can be set according to the timing of changes to the reference image list. Therefore, the following problem can be solved: even if sps_smvd_enabled_flag is set to 1, the symmetric motion vector difference mode will not run regardless of the reference image list structure.
[0469] [Application Example]
[0470] The aforementioned motion picture encoding device 11 and motion picture decoding device 31 can be mounted on various devices for transmitting, receiving, recording, and playing motion pictures. Furthermore, motion pictures can be natural motion pictures captured by cameras or the like, or artificial motion pictures generated by computers or the like (including CG (Computer Graphics) and GUI (Graphical User Interface)).
[0471] First, refer to Figure 2 The following describes the situation where the above-mentioned motion picture encoding device 11 and motion picture decoding device 31 can be used for the transmission and reception of motion pictures.
[0472] Figure 2 PROD_A is a block diagram representing the configuration of the transmitting device PROD_A equipped with the motion picture encoding device 11. For example... Figure 2 As shown, the transmitting device PROD_A includes: an encoding unit PROD_A1, which obtains encoded data by encoding a moving image; a modulation unit PROD_A2, which obtains a modulated signal by modulating a carrier wave using the encoded data obtained by the encoding unit PROD_A1; and a transmitting unit PROD_A3, which transmits the modulated signal obtained by the modulation unit PROD_A2. The moving image encoding device 11 described above serves as the encoding unit PROD_A1.
[0473] The transmitting device PROD_A may also include a camera PROD_A4 for capturing moving images as a source of moving images input to the encoding unit PROD_A1, a recording medium PROD_A5 for recording moving images, an input terminal PROD_A6 for inputting moving images from an external source, and an image processing unit A7 for generating or processing images. Figure 2 The example shows the configuration of the transmitting device PROD_A with all these components, but some parts may be omitted.
[0474] Furthermore, the recording medium PROD_A5 can record unencoded motion images or motion images encoded using a recording encoding method different from the encoding method used for transmission. In the latter case, it is preferable to insert a decoding unit (not shown) between the recording medium PROD_A5 and the encoding unit PROD_A1 to decode the encoded data read from the recording medium PROD_A5 according to the recording encoding method.
[0475] Figure 2 PROD_B is a block diagram representing the configuration of the receiving device PROD_B equipped with the motion picture decoding device 31. For example... Figure 2As shown, the receiving device PROD_B includes: a receiving unit PROD_B1 that receives a modulated signal; a demodulation unit PROD_B2 that obtains encoded data by demodulating the modulated signal received by the receiving unit PROD_B1; and a decoding unit PROD_B3 that obtains a motion image by decoding the encoded data obtained by the demodulation unit PROD_B2. The motion image decoding device 31 described above is used as the decoding unit PROD_B3.
[0476] The receiving device PROD_B may also include a display PROD_B4 that serves as the destination for supplying the moving images output by the decoding unit PROD_B3, a recording medium PROD_B5 for recording the moving images, and an output terminal PROD_B6 for outputting the moving images to an external device. The figure illustrates the configuration of the receiving device PROD_B with all these components, but some may be omitted.
[0477] Furthermore, the recording medium PROD_B5 can be used to record unencoded motion images, or it can be encoded using a recording encoding method different from the encoding method used for transmission. When the latter is used, it is preferable to insert an encoding unit (not shown) between the decoding unit PROD_B3 and the recording medium PROD_B5, which encodes the motion images obtained from the decoding unit PROD_B3 according to the recording encoding method.
[0478] Furthermore, the transmission medium for transmitting modulated signals can be wireless or wired. Additionally, the transmission method for modulated signals can be broadcast (meaning transmission without a pre-defined destination) or communication (meaning transmission with a pre-defined destination). That is, the transmission of modulated signals can be achieved through any of the following: wireless broadcasting, wired broadcasting, wireless communication, and wired communication.
[0479] For example, a digital terrestrial broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) is an example of a wireless broadcasting transceiver modulated signal transmitting device PROD_A / receiving device PROD_B. Similarly, a cable television broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) is an example of a cable broadcasting transceiver modulated signal transmitting device PROD_A / receiving device PROD_B.
[0480] Furthermore, servers (workstations, etc.) / clients (TV receivers, personal computers, smartphones, etc.) using Internet-based VOD (Video On Demand) services or moving image sharing services are examples of transmitting devices PROD_A and receiving devices PROD_B that utilize communication transceiver modulated signals (typically, LANs use either wireless or wired transmission media, while WANs use wired transmission media). Here, personal computers include desktop PCs, laptop PCs, and tablet PCs. Additionally, smartphones also include multi-functional mobile phone terminals.
[0481] In addition to decoding and displaying the encoded data downloaded from the server, the client of the motion picture sharing service also has the function of encoding and uploading motion pictures captured by the camera to the server. That is, the client of the motion picture sharing service functions as both a sending device PROD_A and a receiving device PROD_B.
[0482] Next, refer to Figure 3 The following describes the situation where the above-mentioned motion picture encoding device 11 and motion picture decoding device 31 can be used for recording and playing motion pictures.
[0483] Figure 3 PROD_C is a block diagram representing the configuration of the recording device PROD_C equipped with the aforementioned motion picture encoding device 11. For example... Figure 3 As shown, the recording device PROD_C includes: an encoding unit PROD_C1, which obtains encoded data by encoding a moving image; and a writing unit PROD_C2, which writes the encoded data obtained by the encoding unit PROD_C1 to the recording medium PROD_M. The moving image encoding device 11 described above serves as the encoding unit PROD_C1.
[0484] In addition, the recording medium PROD_M can be (1) a type built into the recording device PROD_C, such as HDD (Hard Disk Drive) or SSD (Solid State Drive), or (2) a type connected to the recording device PROD_C, such as SD (Secure Digital) memory card or USB (Universal Serial Bus) flash memory, or (3) a drive device (not shown) installed in the recording device PROD_C, such as DVD (Digital Versatile Disc) or BD (Blu-ray Disc).
[0485] Furthermore, the recording device PROD_C may also include a camera PROD_C3 for capturing moving images as a source of moving images input to the encoding unit PROD_C1, an input terminal PROD_C4 for inputting moving images from the outside, a receiving unit PROD_C5 for receiving moving images, and an image processing unit PROD_C6 for generating or processing images. Figure 3 The example shows the configuration of the recording device PROD_C with all these components, but some parts may be omitted.
[0486] In addition, the receiving unit PROD_C5 can receive unencoded motion images, or it can receive encoded data encoded using a transmission encoding method different from the encoding method used for recording. When the latter is the case, it is preferable to insert a transmission decoding unit (not shown) between the receiving unit PROD_C5 and the encoding unit PROD_C1 to decode the encoded data encoded using the transmission encoding method.
[0487] Examples of such recording devices PROD_C include DVD recorders, BD recorders, and HDD (Hard Disk Drive) recorders (in which case the input terminal PROD_C4 or the receiving unit PROD_C5 becomes the main source of moving images). Furthermore, portable camcorders (in which case the camera PROD_C3 becomes the main source of moving images), personal computers (in which case the receiving unit PROD_C5 or the image processing unit C6 becomes the main source of moving images), and smartphones (in which case the camera PROD_C3 or the receiving unit PROD_C5 becomes the main source of moving images) are also examples of such recording devices PROD_C.
[0488] Figure 3 PROD_D is a block representing the configuration of the playback device PROD_D equipped with the aforementioned motion picture decoding device 31. For example... Figure 3 As shown, the playback device PROD_D includes: a readout unit PROD_D1, which reads out the encoded data written to the recording medium PROD_M; and a decoding unit PROD_D2, which obtains a motion picture by decoding the encoded data read out by the readout unit PROD_D1. The motion picture decoding device 31 described above is used as the decoding unit PROD_D2.
[0489] In addition, the recording medium PROD_M can be (1) a type built into the playback device PROD_D, such as HDD or SSD, or (2) a type connected to the playback device PROD_D, such as SD memory card or USB flash drive, or (3) a drive device (not shown) built into the playback device PROD_D, such as DVD or BD.
[0490] Furthermore, the playback device PROD_D may also include a display PROD_D3 for displaying moving images, serving as the destination for the moving images output by the decoding unit PROD_D2; an output terminal PROD_D4 for outputting the moving images to an external device; and a transmission unit PROD_D5 for transmitting the moving images. Figure 3 Figure 3 The example shows that the playback device PROD_D has all these components, but some can be omitted.
[0491] Furthermore, the transmitting unit PROD_D5 can transmit unencoded motion images or encoded data encoded using a transmission encoding method different from the encoding method used for recording. In the latter case, it is preferable to insert an encoding unit (not shown) that encodes the motion images using the transmission encoding method between the decoding unit PROD_D2 and the transmitting unit PROD_D5.
[0492] Examples of such playback devices PROD_D include DVD players, BD players, and HDD players (in which case, the output terminal PROD_D4 connected to a TV receiver or the like becomes the main destination for the moving images). Other examples include TV receivers (in which case the display PROD_D3 becomes the main destination for the moving images), digital signage (also known as electronic billboards or electronic signs, where the display PROD_D3 or the transmitter PROD_D5 becomes the main destination for the moving images), desktop PCs (in which case the output terminal PROD_D4 or the transmitter PROD_D5 becomes the main destination for the moving images), laptop or tablet PCs (in which case the display PROD_D3 or the transmitter PROD_D5 becomes the main destination for the moving images), and smartphones (in which case the display PROD_D3 or the transmitter PROD_D5 becomes the main destination for the moving images).
[0493] (Hardware implementation and software implementation)
[0494] Furthermore, each block of the aforementioned motion picture decoding device 31 and motion picture encoding device 11 can be implemented in hardware using logic circuits formed on an integrated circuit (IC chip), or in software using a CPU (Central Processing Unit).
[0495] When the latter is used, each of the aforementioned devices includes a CPU that executes commands for programs that perform various functions, a ROM (Read Only Memory) that stores the programs, a RAM (Random Access Memory) that expands the programs, and a memory that stores the programs and various data, etc., and other storage devices (recording media). Furthermore, the objective of embodiments of the present invention can also be achieved by supplying a recording medium containing, in a computer-readable format, program code (executable form program, intermediate code program, source program) of the software implementing the aforementioned functions, i.e., the control program of each of the aforementioned devices, to each of the aforementioned devices, and having the computer (or CPU or MPU (Microprocessor Unit)) read and execute the program code recorded in the recording medium.
[0496] As the aforementioned recording media, magnetic tapes such as magnetic tapes or cassette tapes can be used; disks such as floppy disks (registered trademark) / hard disks or optical discs such as CD-ROM (Compact Disc Read-Only Memory), MO (Magneto-Optical Disc), MD (Mini Disc), DVD (Digital Versatile Disc), CD-R (CD Recordable), Blu-ray Disc; cards such as IC cards (including memory cards) / optical cards; semiconductor memory such as ROMs with optical masks, EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable and Programmable Read-Only Memory), and flash memory ROMs; or logic circuits such as PLDs (Programmable Logic Devices) and FPGAs (Field Programmable Gate Arrays).
[0497] Furthermore, the aforementioned devices can be configured to connect to a communication network, via which the program code can be supplied. This communication network is not particularly limited as long as it can transmit program code. Examples include the Internet, intranet, extranet, LAN (Local Area Network), ISDN (Integrated Services Digital Network), VAN (Value-Added Network), CATV (Community Antenna Television / Cable Television) communication network, Virtual Private Network, telephone line network, mobile communication network, satellite communication network, etc. Moreover, the transmission medium constituting this communication network is also only required to be a medium capable of transmitting program code, and is not limited to a specific configuration or type. For example, wired connections such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line transmission, cable TV, telephone lines, and ADSL (Asymmetric Digital Subscriber Line) lines can be used. Wireless connections such as IrDA (Infrared Data Association), infrared (e.g., remote controls), Bluetooth (registered trademark), IEEE 802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance, registered trademark), portable telephone networks, satellite lines, and digital terrestrial broadcast networks can also be used. Furthermore, embodiments of the present invention can also be implemented as computer data signals that embody the aforementioned program code and embed it into a carrier wave using electronic transmission.
[0498] The embodiments of the present invention are not limited to the above-described embodiments, and various modifications can be made within the scope of the claims. That is, embodiments obtained by combining technical means that are appropriately modified within the scope of the claims are also included within the technical scope of the present invention.
[0499] Industrial availability
[0500] The embodiments of the present invention are preferably applied to a moving image decoding apparatus that decodes encoded data of encoded image data, and a moving image encoding apparatus that generates encoded data of encoded image data. Furthermore, they are preferably applied to a data structure of encoded data generated by the moving image encoding apparatus and referenced by the moving image decoding apparatus.
[0501] (Mutual references between related applications)
[0502] This application is based on the interest claimed by Japanese Patent Application No. 2020-066614, filed on April 2, 2020, the entire contents of which are incorporated herein by reference.
[0503] Explanation of reference numerals in the attached figures
[0504] 31: Image decoding device
[0505] 301: Entropy Decoding Department
[0506] 302: Parameter Decoding Unit
[0507] 303: Inter-frame Prediction Parameter Derivation Section
[0508] 304: Intra-frame prediction parameter derivation section
[0509] 305, 107: Loop filters
[0510] 306, 109: Reference image storage
[0511] 307, 108: Prediction parameter memory
[0512] 308, 101: Predictive Image Generation Unit
[0513] 309: Inter-frame Prediction Image Generation Unit
[0514] 310: Intra-frame prediction image generation unit
[0515] 311, 105: Inverse quantization / inverse conversion section
[0516] 312, 106: Addition Department
[0517] 320: Prediction Parameter Derivation Section
[0518] 11: Image encoding device
[0519] 102: Subtraction Section
[0520] 103: Conversion / Quantization Department
[0521] 104: Entropy Coding Department
[0522] 110: Encoding Parameter Determination Unit
[0523] 111: Parameter Encoding Section
[0524] 112: Inter-frame prediction parameter coding unit
[0525] 113: Intra-frame prediction parameter coding unit
[0526] 120: Prediction Parameter Derivation Section
Claims
1. A motion image decoding device, characterized in that, have: The decoding unit decodes the reference image list structure for each slice; and The prediction unit derives a reference image list based on the aforementioned reference image list structure. When it is a bidirectional prediction and has a relationship where the current image is sandwiched between two reference images,... The decoding unit sets the reference index of the first reference image list in the first reference index for the symmetrical motion vector difference mode, that is, the value of the reference index whose difference between the POC of the reference image in the first reference image list and the current image is the smallest. The decoding unit sets the reference index of the second reference image list in the second reference index used for the symmetrical motion vector difference mode, that is, the value of the reference index that minimizes the difference between the POC of the reference image in the second reference image list and the current image. The prediction unit derives co-location blocks and flag variables. The motion vector of the object block is derived using a merge candidate list containing the co-position blocks. If the difference between the POC of the first or second active reference image and the current image in the first and second reference image lists of the current slice is less than a first threshold, the value of the flag variable is set to the first value; otherwise, the value of the flag variable is set to the second value.
2. A method for decoding moving images, characterized in that, Decode each slice according to the reference image list structure. Derive the reference image list based on the aforementioned reference image list structure. When it is a bidirectional prediction with a relationship of two reference images sandwiching the current image, in the first reference index used for the symmetric motion vector difference mode, the reference index of the first reference image list is set to the value of the reference index that minimizes the difference between the POC of the reference images in the first reference image list and the current image. In the second reference index used for the symmetric motion vector difference mode, the reference index of the second reference image list is set to the value of the reference index that minimizes the difference between the POC of the reference images in the second reference image list and the current image. The motion vector of the object block is derived using a merge candidate list containing co-position blocks. If the difference between the POC of the first or second active reference image and the current image in the first and second reference image lists of the current slice is less than a first threshold, the value of the flag variable is set to the first value; otherwise, the value of the flag variable is set to the second value.
Citation Information
Patent Citations
Oral composition
JP2020066614A