Image decoding apparatus

By applying a combination of separable transform and secondary transform in an image decoding device, the performance loss problem of implicit MTS under secondary transform is solved, and the efficiency and applicability of image decoding are improved.

CN113892265BActive Publication Date: 2025-10-24SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080039017.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-30
Filing Date
2020-05-29
Publication Date
2025-10-24
Estimated Expiration
2040-05-29

AI Technical Summary

Technical Problem

In existing image coding techniques, when combining secondary transform and multiple transform selection (MTS), the performance loss of implicit MTS is a problem, especially in the case of secondary transform, the performance is poor.

Method used

An image decoding device is adopted, which modifies the transform coefficients through a second transform unit and applies a separable transform when the secondary transform is valid, including vertical and horizontal transforms; when the implicit transform is enabled and the sub-block transform or intra-frame sub-division mode is not used, the vertical and horizontal transform types are derived according to the width and height of the transform unit.

Benefits of technology

The performance of image decoding is improved, especially in the case of implicit MTS and secondary transform combination, which improves the appropriateness and efficiency of the transform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113892265B_ABST
    Figure CN113892265B_ABST
Patent Text Reader

Abstract

The present application provides an image decoding apparatus capable of more appropriately applying a transform realized by MTS and a secondary transform. The image decoding apparatus includes: a second transform section that, in a case where the secondary transform is effective, applies a transform using a transform matrix to a transform coefficient to correct the transform coefficient; a first transform section that applies a separation type transform constituted by a vertical transform and a horizontal transform to the transform coefficient; and an implicit transform setting section that, in a case where the secondary transform is effective, no intra-sub partition mode is used, and no sub-block transform is used, sets an implicit transform to be off, and in a case where the implicit transform is enabled, derives a horizontal transform type in accordance with a width of a target TU, derives a vertical transform type in accordance with a height of the target TU, and the first transform section performs a transform corresponding to the vertical transform type and a transform corresponding to the horizontal transform type.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to an image decoding apparatus and an image encoding apparatus. BACKGROUND

[0002] In order to efficiently transmit or record an image, an image encoding apparatus that generates encoded data by encoding an image and an image decoding apparatus that generates a decoded image by decoding the encoded data are used.

[0003] As a specific image encoding method, for example, H.264 / AVC, HEVC (High- Efficiency Video Coding), and the like can be cited.

[0004] In such an image encoding method, an image (picture) constituting an image is managed by a hierarchical structure including a slice obtained by dividing an image, a coding tree unit (CTU: Coding Tree Unit) obtained by dividing a slice, a coding unit (sometimes also referred to as a coding unit (CU)) obtained by dividing a coding tree unit, and a transform unit (TU: Transform Unit) obtained by dividing a coding unit, and is encoded / decoded per CU.

[0005] Further, in such an image encoding method, generally, a prediction image is generated on the basis of a local decoded image obtained by encoding / decoding an input image, and a prediction error (sometimes also referred to as a "difference image" or a "residual image") obtained by subtracting the prediction image from an input image (original image) is encoded. As a method of generating a prediction image, inter-picture prediction (inter-frame prediction) and intra-picture prediction (intra-frame prediction) can be cited.

[0006] Further, as a recent image encoding and decoding technique, Non-Patent Literature 1, Non-Patent Literature 2 can be cited. In Non-Patent Literature 1, a technique called Multiple Transform Selection (MTS) is disclosed, which switches a transform matrix according to an explicit syntax in encoded data or an implicit block size. In Non-Patent Literature 2, an image encoding apparatus is disclosed, which derives a transform coefficient by RST (Reduced Secondary Transform) transform, that is, a secondary transform, by transforming each coefficient of a transformed prediction error per transform unit. Further, in Non-Patent Literature 2, an image decoding apparatus is disclosed, which inversely transforms a transform coefficient by a secondary transform per transform unit.

[0007] Prior Art Documents

[0008] Non-Patent Literature

[0009] Non-Patent Literature 1: "Versatile Video Coding (Draft 5)", JVET-N1001-v6, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG11, 2019-05-23

[0010] Non-Patent Literature 2: "CE12: Mapping functions (test CE12-1 and CE12-2)", JVET-M0427-v2, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 13th Meeting: Marrakech, MA, 9-18 Jan. 2019 SUMMARY

[0011] Problems to be Solved by the Invention

[0012] In the secondary transform and the technology related to the secondary transform as in Non-Patent Literature 1, there is a problem that the performance is not enough in the case where the secondary transform and the transform realized by the MTS are combined. In particular, there is a problem that the performance of the implicit MTS is lost in the case where the secondary transform is combined.

[0013] An object of the present application is to provide an image decoding apparatus capable of more appropriately applying a transform realized by the MTS and a secondary transform and its related technology.

[0014] Technical Solution

[0015] The moving image decoding apparatus of one aspect of the present application is an image decoding apparatus that transforms transform coefficients on a per transform unit basis, characterized by comprising: a second transform section that, in a case where secondary transform is effective, applies transform using a transform matrix to the transform coefficients, and corrects the transform coefficients; a first transform section that applies a separation type transform composed of vertical transform and horizontal transform to the transform coefficients; and an implicit transform setting section that, in a case where the secondary transform is effective, no intra-sub partition mode is used, and no sub-block transform is used, sets implicit transform to off, and in a case where the implicit transform is on, derives a horizontal transform type in accordance with a width of a target TU, derives a vertical transform type in accordance with a height of the target TU, and the first transform section performs transform corresponding to the vertical transform type and transform corresponding to the horizontal transform type. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 is a diagram showing the configuration of an image transmission system of the present embodiment.

[0017] Figure 2 is a diagram showing the configuration of a transmission apparatus equipped with the moving image encoding apparatus of the present embodiment and a reception apparatus equipped with the moving image decoding apparatus. PROD_A denotes a transmission apparatus equipped with the moving image encoding apparatus, and PROD_B denotes a reception apparatus equipped with the moving image decoding apparatus.

[0018] Figure 3 is a diagram showing the configuration of a recording apparatus equipped with the moving image encoding apparatus of the present embodiment and a reproducing apparatus equipped with the moving image decoding apparatus. PROD_C denotes a recording apparatus equipped with the moving image encoding apparatus, and PROD_D denotes a reproducing apparatus equipped with the moving image decoding apparatus.

[0019] Figure 4 is a diagram showing the hierarchical structure of data of an encoded stream.

[0020] Figure 5 is a diagram showing a partitioning example of a CTU.

[0021] Figure 6 is a diagram showing the kinds (mode numbers) of intra prediction modes.

[0022] Figure 7 is a diagram showing the configuration of a moving image decoding apparatus.

[0023] Figure 8 is a flowchart showing the outline of the operation of a moving image decoding apparatus.

[0024] Figure 9 is a diagram showing the configuration of an intra prediction parameter decoding section.

[0025] Figure 10 is a diagram showing a reference region for intra prediction.

[0026] Figure 11 is a diagram showing the configuration of the intra prediction image generation section.

[0027] Figure 12 is a functional block diagram showing a configuration example of the inverse quantization / inverse transform section.

[0028] Figure 13 is a diagram explaining the transform range of the secondary transform.

[0029] Figure 14 is a diagram explaining the operation of implicit MTS in the case of using an intra sub- partition mode (intra sub- partition prediction).

[0030] Figure 15 is a diagram explaining the operation of implicit MTS in the case of using a sub-block transform.

[0031] Figure 16 is a block diagram showing the configuration of a moving image encoding apparatus.

[0032] Figure 17 is a diagram showing the configuration of the intra prediction parameter encoding section.

[0033] Figure 18 is a block diagram explaining the kernel transform section 1521.

[0034] Figure 19 is a block diagram explaining the secondary transform and the kernel transform.

[0035] Figure 20 is a flowchart explaining the operation of the MTS setting section 15211 of the embodiment.

[0036] Figure 21 is a flowchart explaining the operation of the MTS setting section 15211 of the embodiment.

[0037] Figure 22 is a flowchart explaining the operation of the MTS setting section 15211 of the embodiment.

[0038] Figure 23 is a block diagram showing the relationship between the TU decoding section and the inverse transform section.

[0039] Figure 24 is a flowchart showing the processing of the secondary transform. DETAILED DESCRIPTION

[0040] [Embodiment 1]

[0041] An embodiment of the present application will be described below with reference to the accompanying drawings.

[0042] Figure 1 is a diagram showing the configuration of the image transmission system 1 of the present embodiment.

[0043] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding an encoding target image, decodes the transmitted encoded stream, and displays an image. The image transmission system 1 is configured to include a moving image encoding device (image encoding device) 11, a network 21, a moving image decoding device (image decoding device) 31, and an image display device (image display device) 41.

[0044] The moving image encoding device 11 is input with an image T.

[0045] The network 21 transmits the encoded stream Te generated by the moving image encoding device 11 to the moving image decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network 21 is not necessarily limited to a bidirectional communication network, and can be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Furthermore, the network 21 can be replaced with a storage medium such as a DVD (Digital Versatile Disc, registered trademark), a BD (Blue-ray Disc, registered trademark), or the like, in which the encoded stream Te is recorded.

[0046] The moving image decoding device 31 decodes the encoded stream Te transmitted by the network 21, and generates one or a plurality of decoded images Td.

[0047] The image display device 41 displays all or a part of the one or a plurality of decoded images Td generated by the moving image decoding device 31. The image display device 41 is provided with, for example, a liquid crystal display, an organic EL (Electro-luminescence) display, or the like. As a form of the display, a stationary type, a mobile type, an HMD, or the like can be cited. Furthermore, in a case where the moving image decoding device 31 has a high processing capability, an image with high display quality is displayed, and in a case where only a low processing capability is possessed, an image that does not require a high processing capability and a high display capability is displayed.

[0048] <Operator>

[0049] The operator used in the present specification will be described below.

[0050] >> for right shift, << for left shift, & for bitwise AND, | for bitwise OR, |= for OR assignment operator, || for logical AND.

[0051] x? y : z is a ternary operator that takes y if x is true (other than 0) and z if x is false (0).

[0052] Clip3(a, b, c) is a function that clips the value of c to be between a and b, and is a function that returns a if c < a, b if c > b, and c if otherwise (where a <= b).

[0053] abs(a) is a function that returns the absolute value of a.

[0054] Int(a) is a function that returns the integer value of a.

[0055] floor(a) is a function that returns the largest integer less than a.

[0056] ceil(a) is a function that returns the smallest integer greater than or equal to a.

[0057] a / d means a divided by d (rounding down).

[0058] <Construction of coded stream Te>

[0059] Before the moving image encoding apparatus 11 and the moving image decoding apparatus 31 of the present embodiment are described in detail, the data structure of a coded stream Te generated by the moving image encoding apparatus 11 and decoded by the moving image decoding apparatus 31 is described.

[0060] Figure 4 A diagram showing the hierarchical structure of data in the coded stream Te. The coded stream Te exemplarily includes a sequence and a plurality of pictures constituting the sequence. Figure 4 A diagram showing a coded video sequence representing a predetermined sequence SEQ, a coded picture representing a predetermined picture PICT, a coded slice representing a predetermined slice S, coded slice data representing a predetermined slice data, a coding tree unit included in the coded slice data, and a coding unit included in the coding tree unit.

[0061] (Coded video sequence)

[0062] In the coded video sequence, a set of data for the moving image decoding apparatus 31 to refer to in order to decode the sequence SEQ that is a processing target is defined. As shown in FIG. 6, the coded video sequence includes a sequence header SH and a plurality of pictures PICT. Figure 4The coded video sequence shown in FIG. 1 includes a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a picture PICT, and supplemental enhancement information (SEI).

[0063] The video parameter set (VPS) defines a set of coding parameters common to a plurality of pictures and a plurality of layers included in the pictures and a set of coding parameters associated with each layer.

[0064] In the sequence parameter set (SPS), a set of coding parameters for reference by the moving picture decoding device 31 to decode the target sequence is defined. For example, the width and height of a picture are defined. Note that a plurality of SPSs can exist. In this case, any one of the plurality of SPSs is selected from the PPS.

[0065] In the picture parameter set (PPS), a set of coding parameters for reference by the moving picture decoding device 31 to decode each picture within the target sequence is defined. For example, a reference value of the quantization width for decoding of a picture (pic_init_qp_minus26), a flag indicating the application of weighted prediction (weighted_pred_flag), and a scaling list (quantization matrix) are included. Note that a plurality of PPSs can exist. In this case, any one of the plurality of PPSs is selected from each picture within the target sequence.

[0066] (Coded picture)

[0067] In the coded picture, a set of data for reference by the moving picture decoding device 31 to decode the picture PICT that is the processing target is defined. As shown in the coded picture of FIG. 1, the picture PICT includes slices 0 to NS-1 (NS is the total number of slices included in the picture PICT). Figure 4

[0068] Note that hereinafter, the index of the code will be omitted in some cases in the description without distinguishing each slice 0 to NS-1. The same applies to other data included in the coded stream Te described hereinafter, that is, data labeled with an index.

[0069] (Coded slice)

[0070] In the coded slice, a set of data for reference by the moving picture decoding device 31 to decode the slice S that is the processing target is defined. As shown in the coded slice of FIG. 1, the slice S includes a slice header and a slice data. Figure 4 ​As shown in the coded slice, the slice includes a slice header and slice data.

[0071] The slice header includes a coding parameter group that the moving picture decoding device 31 refers to in order to determine a decoding method for a target slice. Slice type designation information (slice_type) that designates a slice type is one example of a coding parameter included in the slice header.

[0072] Examples of slice types that can be specified by the slice type specification information include: (1) I slices that use only intra-frame prediction during encoding, (2) P slices that use unidirectional prediction or intra-frame prediction during encoding, and (3) B slices that use unidirectional prediction, bidirectional prediction, or intra-frame prediction during encoding. It should be noted that inter-frame prediction is not limited to unidirectional prediction and bidirectional prediction, and more reference pictures can be used to generate predicted images. Hereinafter, P and B slices refer to slices that include blocks that can use inter-frame prediction.

[0073] It should be noted that the slice header may also include a reference to the picture parameter set PPS (pic_parameter_set_id).

[0074] (Encoded slice data)

[0075] The coded slice data specifies a set of data that the moving image decoding device 31 refers to in order to decode the slice data to be processed. Figure 4 As shown in the coded slice header, the slice data includes CTUs. A CTU is a fixed-size (e.g., 64×64) block that constitutes a slice and is also called the Largest Coding Unit (LCU).

[0076] (Coding Tree Unit)

[0077] exist Figure 4 In the coding tree unit, a set of data is specified for reference by the motion picture decoding device 31 in order to decode the CTU of the processing object. The CTU is divided into coding units CU, which are the basic units of the encoding process, by recursive quadtree partitioning (QT (Quad Tree) partitioning), binary tree partitioning (BT (Binary Tree) partitioning), or ternary tree partitioning (TT (Ternary Tree) partitioning). BT partitioning and TT partitioning are collectively referred to as multi-tree partitioning (MT (Multi Tree) partitioning). The nodes of the tree structure obtained by recursive quadtree partitioning are called coding nodes. The intermediate nodes of the quadtree, binary tree, and ternary tree are coding nodes, and the CTU itself is also specified as the top-level coding node.

[0078] The CT includes, as the CT information, a QT split flag (cu_split_flag) indicating whether or not QT split is performed, an MT split flag (split_mt_flag) indicating the presence or absence of MT split, an MT split direction (split_mt_dir) indicating the split direction of MT split, and an MT split type (split_mt_type) indicating the split type of MT split. The cu_split_flag, the split_mt_flag, the split_mt_dir, and the split_mt_type are transmitted per coding node.

[0079] In the case where the cu_split_flag is 1, the coding node is split into 4 coding nodes (QT). Figure 5 In the case where the cu_split_flag is 0, in the case where the split_mt_flag is 0, the coding node is not split, and 1 CU is maintained as the node (no split of QT).

[0080] In the case where the cu_split_flag is 0, in the case where the split_mt_flag is 0, the coding node is not split, and 1 CU is maintained as the node (no split of QT). Figure 5 The CU is a terminal node of the coding node, and is not further split. The CU is a basic unit of coding processing.

[0081] In the case where the split_mt_flag is 1, the coding node is MT split as shown below. In the case where the split_mt_type is 0, in the case where the split_mt_dir is 1, the coding node is horizontally split into 2 coding nodes (BT (horizontal split) of QT), and in the case where the split_mt_dir is 0, the coding node is vertically split into 2 coding nodes (BT (vertical split) of QT). Further, in the case where the split_mt_type is 1, in the case where the split_mt_dir is 1, the coding node is horizontally split into 3 coding nodes (TT (horizontal split) of QT), and in the case where the split_mt_dir is 0, the coding node is vertically split into 3 coding nodes (TT (vertical split) of QT). These are shown in the CT information of QT. Figure 5 Figure 5 Figure 5 Figure 5 Figure 5

[0082] ​​​​​Further, in a case where the size of the CTU is 64x64 pixels, the size of the CU can take any one of 64x64 pixels, 64x32 pixels, 32x64 pixels, 32x32 pixels, 64x16 pixels, 16x64 pixels, 32x16 pixels, 16x32 pixels, 16x16 pixels, 64x8 pixels, 8x64 pixels, 32x8 pixels, 8x32 pixels, 16x8 pixels, 8x16 pixels, 8x8 pixels, 64x4 pixels, 4x64 pixels, 32x4 pixels, 4x32 pixels, 16x4 pixels, 4x16 pixels, 8x4 pixels, 4x8 pixels, and 4x4 pixels.

[0083] (encoding unit)

[0084] As shown in the encoding unit of FIG. 1, a set of data for reference by the motion picture decoding device 31 to decode the encoding unit that is a processing target is defined. Specifically, the CU is constituted by a CU header CUH, prediction parameters, transform parameters, quantized transform coefficients, and the like. The prediction mode and the like are defined in the CU header. Figure 4

[0085] The prediction processing is performed in units of CUs and in units of sub-CUs obtained by further dividing the CUs. In a case where the size of the CU is equal to that of the sub-CU, the sub-CU in the CU is one. In a case where the size of the CU is larger than that of the sub-CU, the CU is divided into sub-CUs. For example, in a case where the CU is 8x8 and the sub-CU is 4x4, the CU is divided into four sub-CUs constituted by two parts divided horizontally and two parts divided vertically.

[0086] The kind of prediction (prediction mode) is either intra prediction or inter prediction. The intra prediction is prediction within the same picture, and the inter prediction is prediction processing between mutually different pictures (e.g., between display times, between layer pictures).

[0087] The transform / quantization section is processed in units of CUs, but the quantized transform coefficients can also be entropy-encoded in units of sub-blocks such as 4x4.

[0088] (prediction parameters)

[0089] The prediction image is derived from the prediction parameters attached to the block. The prediction parameters include prediction parameters for intra prediction and inter prediction.

[0090] Hereinafter, the prediction parameters for intra prediction will be described. The intra prediction parameters are constituted by a luminance prediction mode IntraPredModeY and a color difference prediction mode IntraPredModeC. Figure 6 is a schematic diagram showing the kind of intra prediction mode (mode number). As shown in FIG. 2, the intra prediction mode is constituted by 33 modes. The mode number is 0 to 32. Figure 6 ​As shown, there are, for example, 67 kinds of intra prediction modes (0 to 66). For example, there are planar prediction (0), DC prediction (1), Angular prediction (2 to 66). Moreover, LM modes (67 to 72) can be added in the color difference.

[0091] Among the syntax elements for deriving the intra prediction parameters, there are, for example, intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_remainder, and the like.

[0092] (MPM)

[0093] intra_luma_mpm_flag is a flag indicating whether or not the IntraPredModeY of the target block coincides with the MPM (Most Probable Mode). The MPM is a prediction mode included in the MPM candidate list mpmCandList[]. The MPM candidate list is a list storing candidates estimated to have high probability of being applied to the target block from the intra prediction modes of the neighboring blocks and the prescribed intra prediction modes. In the case where intra_luma_mpm_flag is 1, the IntraPredModeY of the target block is derived using the MPM candidate list and the index intra_luma_mpm_idx.

[0094] IntraPredModeY = mpmCandList[intra_luma_mpm_idx]

[0095] (REM)

[0096] In the case where intra_luma_mpm_flag is 0, the intra prediction mode is selected from RemIntraPredMode which is the mode remaining after the intra prediction modes included in the MPM candidate list are removed from all the intra prediction modes. As RemIntraPredMode, the intra prediction mode which can be selected is referred to as "non-MPM" or "REM". RemIntraPredMode is derived using intra_luma_mpm_remainder.

[0097] (Configuration of Motion Picture Decoding Apparatus)

[0098] The configuration of the motion picture decoding apparatus 31 of the present embodiment will be described. Figure 7

[0099] ​The moving image decoding apparatus 31 is configured to include an entropy decoding section 301, a parameter decoding section (prediction image decoding apparatus) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a prediction image generation section (prediction image generation apparatus) 308, an inverse quantization / inverse transform section 311, and an addition section 312. Note that, according to the moving image encoding apparatus 11 described later, there is also a configuration in which the loop filter 305 is not included in the moving image decoding apparatus 31.

[0100] The parameter decoding section 302 further includes a header decoding section 3020, a CT information decoding section 3021, and a CU decoding section 3022 (prediction mode decoding section), and the CU decoding section 3022 includes a TU decoding section 3024. These can be collectively referred to as a decoding module. The header decoding section 3020 decodes parameter set information such as VPS, SPS, and PPS, and a slice header (slice information) from the encoded data. The CT information decoding section 3021 decodes CT from the encoded data. The CU decoding section 3022 decodes a CU from the encoded data. The TU decoding section 3024 decodes QP update information (quantization correction value) and quantized prediction error (residual_coding) from the encoded data in the case where a prediction error is included in a TU.

[0101] Figure 23 is a block diagram showing the relationship between the TU decoding section 3024 and the inverse transform section 3112. The stIdx decoding section 131 of the TU decoding section 3024 decodes a value stIdx indicating the use of secondary transform and the transform basis from the encoded data, and outputs it to the secondary transform section 31121. The mts_idx decoding section 132 of the TU decoding section 3024 decodes a value mts_idx indicating the transform matrix of MTS from the encoded data, and outputs it to the kernel transform section 31123. Specifically, the TU decoding section 3024 decodes stIdx in the case where the width and height of a CU are 4 or more, the prediction mode is an intra mode, and the number of transform coefficients within the CU, numSigCoeff, is greater than a predetermined number THSt (for example, 2 in SINGLE_TREE, and 1 otherwise). Note that, in the case where stIdx is 0, no secondary transform is applied, in the case where stIdx is 1, a transform of one side of a set (pair) of secondary transform matrices is indicated, and in the case where stIdx is 2, a transform of the other side of the pair is indicated. Further, the secondary transform matrix secTransMatrix can be selected not only according to the value of stIdx, but also according to the intra prediction mode and the size of the transform.

[0102] Further, the parameter decoding section 302 is configured to include an inter prediction parameter decoding section 303 and an intra prediction parameter decoding section 304, which are not shown. The prediction image generating section 308 is configured to include an inter prediction image generating section 309 and an intra prediction image generating section 310.

[0103] Further, in the following, an example in which a CTU, a CU is used as a processing unit is described, but is not limited thereto, and processing can be performed in units of sub-CUs. Alternatively, it can be configured to replace the CTU, the CU with a block, and the sub-CU with a sub-block, and perform processing in units of blocks or sub-blocks.

[0104] The entropy decoding section 301 entropy-decodes the encoded stream Te input from the outside, separates each code (syntax element), and decodes it. In entropy coding, there are a method of variable-length encoding a syntax element using a context (probability model) appropriately selected according to the kind of the syntax element, the surrounding situation, and a method of variable-length encoding a syntax element using a predetermined table or a calculation formula. The former CABAC (Context Adaptive Binary Arithmetic Coding) stores a probability model updated for each encoded or decoded picture (slice) in a memory. Then, as an initial state of a context of a P picture or a B picture, a probability model of a picture using the same slice type, the same slice-level quantization parameter is set according to the probability model stored in the memory. This initial state is used for encoding and decoding processing. Among the separated codes, there are prediction information for generating a prediction image and prediction error for generating a difference image, and the like.

[0105] The entropy decoding section 301 outputs the separated codes to the parameter decoding section 302. The separated codes are, for example, a prediction mode predMode. Control of which code is decoded is performed based on an instruction of the parameter decoding section 302.

[0106] (Basic Flow)

[0107] Figure 8 is a flowchart illustrating the outline of the operation of the moving image decoding apparatus 31.

[0108] (S1100: Parameter Set Information Decoding) The header decoding section 3020 decodes parameter set information such as VPS, SPS, and PPS from the encoded data.

[0109] (S1200: Slice Information Decoding) The header decoding section 3020 decodes slice headers (slice information) from the encoded data.

[0110] The moving image decoding apparatus 31 derives a decoded image of each CTU by repeating the processes of S1300 to S5000 for each CTU included in the target picture.

[0111] (S1300: CTU information decoding) The CT information decoding section 3021 decodes a CTU from the encoded data.

[0112] (S1400: CT information decoding) The CT information decoding section 3021 decodes a CT from the encoded data.

[0113] (S1500: CU decoding) The CU decoding section 3022 implements S1510, S1520 to decode a CU from the encoded data.

[0114] (S1510: CU information decoding) The CU decoding section 3022 decodes CU information, prediction information, a TU split flag split_transform_flag, CU residual flags cbf_cb, cbf_cr, cbf_luma, and the like from the encoded data.

[0115] (S1520: TU information decoding) The TU decoding section 3024 decodes QP update information (quantization correction value) and quantized prediction error (residual_coding) from the encoded data in a case where a prediction error is included in a TU. Note that the QP update information is a difference value from a predicted value of a quantization parameter qPpred which is a predicted value of a quantization parameter QP.

[0116] (S2000: Prediction image generation) The prediction image generation section 308 generates a prediction image based on prediction information for each block included in a target CU.

[0117] (S3000: Inverse quantization / inverse transform) The inverse quantization / inverse transform section 311 performs inverse quantization / inverse transform processing for each TU included in a target CU.

[0118] (S4000: Decoded image generation) The addition section 312 generates a decoded image of a target CU by adding a prediction image provided by the prediction image generation section 308 and a prediction error provided by the inverse quantization / inverse transform section 311.

[0119] (S5000: Loop filtering) The loop filter 305 applies loop filtering such as deblocking filtering, SAO, ALF, and the like to a decoded image to generate a decoded image.

[0120] Further, the parameter decoding section 302 is configured to include an inter prediction parameter decoding section 303 and an intra prediction parameter decoding section 304 which are not illustrated. The prediction image generation section 308 is configured to include an inter prediction image generation section 309 and an intra prediction image generation section 310 which are not illustrated.

[0121] Configuration of Intra Prediction Parameter Decoding Section 304

[0122] The intra prediction parameter decoding section 304 decodes the intra prediction parameter, for example, IntraPredMode, based on the code input from the entropy decoding section 301, with reference to the prediction parameter stored in the prediction parameter memory 307. The intra prediction parameter decoding section 304 outputs the decoded intra prediction parameter to the prediction image generating section 308, and stores it in the prediction parameter memory 307. The intra prediction parameter decoding section 304 can also derive the intra prediction mode that is different between the luminance and the color difference.

[0123] Figure 9 is a schematic diagram showing the configuration of the intra prediction parameter decoding section 304 of the parameter decoding section 302. As shown in Figure 9 The intra prediction parameter decoding section 304 is configured to include a parameter decoding control section 3041, a luminance intra prediction parameter decoding section 3042, and a color difference intra prediction parameter decoding section 3043.

[0124] The parameter decoding control section 3041 instructs the decoding of the syntax elements to the entropy decoding section 301, and receives the syntax elements from the entropy decoding section 301. In a case where intra_luma_mpm_flag is 1, the parameter decoding control section 3041 outputs intra_luma_mpm_idx to the MPM parameter decoding section 30422 within the luminance intra prediction parameter decoding section 3042. In addition, in a case where intra_luma_mpm_flag is 0, the parameter decoding control section 3041 outputs intra_luma_mpm_remainder to the non-MPM parameter decoding section 30423 of the luminance intra prediction parameter decoding section 3042. Further, the parameter decoding control section 3041 outputs the syntax elements of the intra prediction parameter of the color difference to the color difference intra prediction parameter decoding section 3043.

[0125] The luminance intra prediction parameter decoding section 3042 is configured to include an MPM candidate list derivation section 30421, an MPM parameter decoding section 30422, and a non-MPM parameter decoding section 30423 (decoding section, derivation section).

[0126] The MPM parameter decoding section 30422 derives IntraPredModeY with reference to mpmCandList[] and intra_luma_mpm_idx derived by the MPM candidate list derivation section 30421, and outputs it to the intra prediction image generating section 310.

[0127] The non-MPM parameter decoding section 30423 derives RemIntraPredMode from mpmCandList[] and intra_luma_mpm_remainder, and outputs IntraPredModeY to the intra prediction image generating section 310.

[0128] The chroma intra prediction parameter decoding section 3043 derives IntraPredModeC from the syntax elements of the intra prediction parameter of the chroma, and outputs it to the intra prediction image generating section 310.

[0129] The luminance intra prediction parameter decoding section 3042 can also decode a flag intra_subpartitions_mode_flag indicating whether or not to perform intra sub-partitioning that performs intra prediction by dividing the CU into smaller sub-blocks. In the case where intra_subpartitions_mode_flag is 0, further, intra_subpartitions_split_flag is decoded. The intra sub-partitioning mode is derived by the following equation.

[0130] IntraSubPartSplitType = (intra_subpartitions_mode_flag == 0)? 0 : 1 + intra_subpartitions_split_flag In the case where IntraSubPartSplitType is 0 (ISP_NO_SPLIT), the CU is not further divided and intra prediction is performed. In the case where IntraSubPartSplitType is 1 (ISP_HOR_SPLIT: horizontal split), the CU is divided into four sub-blocks from two in the vertical direction, and intra prediction and transform coefficient decoding, inverse quantization / inverse transform are performed in units of sub-blocks. In the case where IntraSubPartSplitType is 2 (ISP_VER_SPLIT: vertical split), the CU is divided into four sub-blocks from two in the horizontal direction, and intra prediction and transform coefficient decoding, inverse quantization / inverse transform are performed in units of sub-blocks. The number of divisions of the sub-blocks NumIntraSubPart is derived by the following equation.

[0131] NumIntraSubPart = (cbWidth == 4 && cbHeight == 8) || (cbWidth == 8 && cbHeight == 4)? 2 : 4

[0132] The width nW and height nH of the sub-block, and the number of divisions numPartsX and numPartY in the horizontal direction and the vertical direction are derived by the following.

[0133] nW = (IntraSubPartSplitType == ISP_VER_SPLIT? ) nTbW / NumIntraSubPart : nTbW

[0134] nH = (IntraSubPartSplitType == ISP_HOR_SPLIT? ) nTbH / NumIntraSubPart : nTbH

[0135] numPartsX = (IntraSubPartSplitType == ISP_VER_SPLIT? ) NumIntraSubPart : 1

[0136] numPartsY = (IntraSubPartSplitType == ISP_HOR_SPLIT? ) NumIntraSubPart : 1

[0137] Here, nTbW and nTbH are the width and height of the CU (or TU).

[0138] The loop filter 305 is a filter provided in the encoding loop, and is a filter that removes block distortion and ringing distortion to improve image quality. The loop filter 305 performs filtering such as deblocking filtering, sample adaptive offset (SAO), adaptive loop filtering (ALF), and the like on the decoded image of the CU generated by the addition section 312.

[0139] The reference picture memory 306 stores the decoded image of the CU generated by the addition section 312 in a predetermined location for each object picture and object CU.

[0140] The prediction parameter memory 307 stores prediction parameters in a predetermined location for each decoded object CTU or CU. Specifically, the prediction parameter memory 307 stores the parameters decoded by the parameter decoding section 302 and predMode and the like separated by the entropy decoding section 301.

[0141] The prediction image generation section 308 is input with predMode, prediction parameters, and the like. In addition, the prediction image generation section 308 reads a reference picture from the reference picture memory 306. The prediction image generation section 308 generates a prediction image of a block or sub-block using the prediction parameters and the read reference picture (reference picture block) in the prediction mode indicated by predMode. Here, the reference picture block refers to a set of pixels (typically rectangular, and thus referred to as a block) on the reference picture, and is a region referred to for generating a prediction image.

[0142] (Intra prediction image generation section 310)

[0143] When predMode indicates an intra prediction mode, the intra prediction image generation unit 310 performs intra prediction using the intra prediction parameters input from the intra prediction parameter decoding unit 304 and the reference pixels read from the reference picture memory 306 .

[0144] Specifically, the intra-prediction image generator 310 reads adjacent blocks within a predetermined range from the target block in the target picture from the reference picture memory 306. The predetermined range is the adjacent blocks to the left, upper left, upper, and upper right of the target block, and the reference area varies depending on the intra-prediction mode.

[0145] The intra-frame prediction image generation unit 310 generates a prediction image for the target block by referring to the read decoded pixel value and the prediction mode indicated by IntraPredMode. The intra-frame prediction image generation unit 310 outputs the generated prediction image for the target block to the addition unit 312.

[0146] The following describes the generation of a predicted image based on an intra-frame prediction mode. In Planar prediction, DC prediction, and Angular prediction, a decoded surrounding area adjacent to (close to) the prediction target block is set as a reference area R. Then, the predicted image is generated by extrapolating the pixels in the reference area R in a specific direction. For example, the reference area R can be set to an L-shaped area (e.g., a region formed by the left and top (or further left top, right top, and bottom left) of the prediction target block. Figure 10 The reference area of ​​Example 1 is represented by the area marked with oblique lines and pixels in the circle).

[0147] (Details of the Prediction Image Generator)

[0148] Next, use Figure 11 The following describes the details of the configuration of the intra-frame prediction image generation unit 310. The intra-frame prediction image generation unit 310 includes a prediction target block setting unit 3101, an unfiltered reference image setting unit 3102 (first reference image setting unit), a filtered reference image setting unit 3103 (second reference image setting unit), an intra-frame prediction unit 3104, and a prediction image correction unit 3105 (prediction image correction unit, filter switching unit, and weighting coefficient changing unit).

[0149] The intra prediction unit 3104 generates a temporary predicted image (pre-corrected predicted image) for the prediction target block based on each reference pixel in the reference region R (unfiltered reference image), the filtered reference image generated by applying the reference pixel filter (first filter), and the intra prediction mode, and outputs the image to the predicted image correction unit 3105. The predicted image correction unit 3105 corrects the temporary predicted image according to the intra prediction mode, generates a predicted image (corrected predicted image), and outputs the image.

[0150] The following describes each unit included in the intra prediction image generation unit 310.

[0151] (Prediction target block setting unit 3101)

[0152] The prediction target block setting unit 3101 sets the target CU as a prediction target block, and outputs information about the prediction target block (prediction target block information). The prediction target block information includes at least the size, position, and index indicating the luma or chroma of the prediction target block.

[0153] (Unfiltered reference image setting unit 3102)

[0154] The unfiltered reference image setting unit 3102 sets the neighboring peripheral region of the prediction target block as the reference region R based on the size and position of the prediction target block. Next, each decoded pixel value of the pixels in the reference region R (unfiltered reference image, boundary pixel) is set to the corresponding position on the reference picture memory 306. Figure 10 The row r[x][-1] of decoded pixels adjacent to the top edge of the prediction target block and the column r[-1][v] of decoded pixels adjacent to the left edge of the prediction target block in Example 1 of the reference region are unfiltered reference images.

[0155] (Filterd reference image setting unit 3103)

[0156] The filtered reference image setting unit 3103 applies a reference pixel filter (first filter) to the unfiltered reference image according to the intra prediction mode, and derives the filtered reference image s[x][y] at each position (x, y) on the reference region R. Specifically, a low-pass filter is applied to the unfiltered reference image at the position (x, y) and its periphery, and the filtered reference image s[x][y] is derived. Figure 10 Example 2 of the reference region). Note that the low-pass filter is not necessarily applied to all intra prediction modes, and can be applied to a part of the intra prediction modes. Note that the filter applied to the unfiltered reference image on the reference region R in the filtered reference image setting unit 3103 is referred to as a "reference pixel filter (first filter)", and in contrast, the filter used to correct the temporary prediction image in the prediction image correction unit 3105 described later is referred to as a "boundary filter (second filter)".

[0157] (Configuration of intra prediction unit 3104)

[0158] The intra prediction section 3104 generates a temporary prediction image (temporary prediction pixel value, pre-correction prediction image) of the prediction target block based on the intra prediction mode, the unfiltered reference image, and the filtered reference pixel value, and outputs it to the prediction image correction section 3105. The intra prediction section 3104 has, inside, a Planar prediction section 31041, a DC prediction section 31042, an Angular prediction section 31043, and an LM prediction section 31044. The intra prediction section 3104 selects a specific prediction section according to the intra prediction mode, and inputs the unfiltered reference image and the filtered reference image. The relationship between the intra prediction mode and the corresponding prediction section is as follows.

[0159] • Planar prediction • Planar prediction section 31041

[0160] • DC prediction • DC prediction section 31042

[0161] • Angular prediction • Angular prediction section 31043

[0162] • LM prediction • LM prediction section 31044

[0163] (Planar prediction)

[0164] The Planar prediction section 31041 generates a temporary prediction image by linearly adding a plurality of filtered reference images according to the distance between the prediction target pixel position and the reference pixel position, and outputs it to the prediction image correction section 3105.

[0165] (DC prediction)

[0166] The DC prediction section 31042 derives a DC prediction value corresponding to the average value of the filtered reference image s[x][y], and outputs a temporary prediction image q[x][y] in which the DC prediction value is taken as the pixel value.

[0167] (Angular prediction)

[0168] The Angular prediction section 31043 generates a temporary prediction image q[x][y] using the filtered reference image s[x][y] of the prediction direction (reference direction) indicated by the intra prediction mode, and outputs it to the prediction image correction section 3105.

[0169] (LM prediction)

[0170] The LM prediction unit 31044 predicts the pixel values of the color difference based on the pixel values of the luminance. Specifically, it is a method of generating a predicted image of the color difference image (Cb, Cr) using a linear model based on the decoded luminance image. CCLM (Cross-Component Linear Model prediction), which is one of the LM predictions, is a prediction method that uses a linear model for predicting the color difference from the luminance for one block.

[0171] (Configuration of the prediction image correction unit 3105)

[0172] The prediction image correction unit 3105 corrects the temporary prediction image output from the intra prediction unit 3104 according to the intra prediction mode. Specifically, the prediction image correction unit 3105 performs weighted addition (weighted averaging) of the unfiltered reference image and the temporary prediction image for each pixel of the temporary prediction image, according to the distance between the reference region R and the target prediction pixel, thereby deriving the prediction image (corrected prediction image) Pred that is the temporary prediction image corrected. Note that in the partial intra prediction mode, the temporary prediction image is not corrected by the prediction image correction unit 3105, and the output of the intra prediction unit 3104 can be directly used as the prediction image.

[0173] (Inverse quantization / inverse transform unit 311)

[0174] The inverse quantization / inverse transform unit 311 inverse quantizes the quantized transform coefficient qd[][] input from the entropy decoding unit 301 to obtain a transform coefficient d[][]. This quantized transform coefficient qd[][] is a coefficient obtained by performing frequency transform such as DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), or the like on the prediction error and quantizing it in the encoding process. The inverse quantization / inverse transform unit 311 performs inverse frequency transform such as inverse DCT, inverse DST, or the like on the obtained transform coefficient, and calculates the prediction error. The inverse quantization / inverse transform unit 311 outputs the prediction error to the addition unit 312.

[0175] Hereinafter, a configuration example of the inverse quantization / inverse transform unit 311 will be described with reference to Figure 12 Figure 12 is a functional block diagram that shows a configuration example of the inverse quantization / inverse transform unit 311. As shown in Figure 12 , the inverse quantization / inverse transform unit 311 includes an inverse quantization unit 3111 and a transform unit 3112. The inverse quantization unit 3111 inverse quantizes the quantized transform coefficient qd[][] decoded in the TU decoding unit 3024 to obtain a transform coefficient d[][]. The inverse quantization unit 3111 outputs the obtained transform coefficient d[][] to the transform unit 3112. ​

[0176] The transform unit 3112 inversely transforms the received transform coefficients d[][] on a per-TU basis to reproduce the prediction error r[][]. The transform unit 3112 outputs the reproduced prediction error r[][] to the addition unit 312.

[0177] Note that in the present specification, the process of transforming a difference image into transform coefficients in an image encoding apparatus is referred to as forward transform, and the process of transforming transform coefficients into a difference image in an image decoding apparatus is referred to as inverse transform, but can be referred to as transform and inverse transform, respectively. Note that the forward transform (transform) and the inverse transform (inverse transform) are not different in processing other than the value of the transform matrix as a transform base. Therefore, in the following description, with respect to the transform processing in the transform unit 3112, the term "inverse transform" can be used instead of "transform".

[0178] The transform unit 3112 includes a secondary transform unit (second transform unit) 31121 and a core transform unit (first transform unit) 31123.

[0179] The TU decoding unit 3024 can decode a sub-block transform flag cu_sbt_flag indicating that the transform coefficients are decoded, inverse-quantized, and inverse-transformed only for one of the plurality of sub-blocks by further dividing the CU into the plurality of sub-blocks. In a case where the cu_sbt_flag is 1, a flag cu_sbt_quad_flag indicating whether or not the division into four sub-blocks is performed can also be decoded. In a case where the cu_sbt_quad_flag is 0, the number of sub-blocks is two. In a case where the cu_sbt_quad_flag is 1, the number of sub-blocks is four. Further, a cu_sbt_horizontal_flag indicating whether the division is performed horizontally or vertically is decoded. Further, a cu_sbt_pos_flag indicating which sub-block includes the transform coefficients is decoded.

[0180] (Scale unit 31112)

[0181] The scale unit 31112 performs scaling using a weighting in a coefficient unit with respect to the transform coefficients decoded by the TU decoding unit.

[0182] The scale unit 31112 performs scaling by the following equation in a case where the transform skip is valid (transform_skip == 1).

[0183] r[x][y] = d[x][y] << tsShift

[0184] Here, tsShift = 5 + ((log2(nTbW) + log2(nTbH)) / 2).

[0185] In the case other than the above, the quantization matrix m[x][y] and the scaling factor ls[x][y] are derived by the following formula.

[0186] ls[x][y] = (m[x][y] * levelScale[(qP + 1) % 6]) « (qP / 6)

[0187] Or can be derived by the following formula.

[0188] ls[x][y] = (m[x][y] * levelScale[qP % 6]) « (qP / 6)

[0189] Here, levelScale[] = {40, 45, 51, 57, 64, 72}.

[0190] Note that the value of the quantization matrix m[x][y] can be decoded from the encoded data, and m[x][y] = 16 can be used as uniform quantization.

[0191] The scaling unit 31112 derives dnc[][] from the product of the scaling factor ls[][] and the decoded transform coefficient TransCoeffLevel, and inverse quantizes it.

[0192] dnc[x][y] = (TransCoeffLevel[xTbY][yTbY][cIdx][x][y] * ls[x][y] * rectNorm + bdOffset) » bdShift

[0193] Finally, the scaling unit 31112 clips the inverse quantized transform coefficient and derives d[x][y].

[0194] d[x][y] = Clip3(CoeffMin, CoeffMax, dnc[x][y])

[0195] d[x][y] is transmitted to the kernel transform unit 31123 or the secondary transform unit 31121. The secondary transform unit (second transform unit) 31121 applies a secondary transform to the transform coefficient d[][] after inverse quantization and before kernel transform.

[0196] (Secondary transform and kernel transform)

[0197] The secondary transform unit 31121 applies a transform using a transform matrix to a part or all of the transform coefficients d[][] received from the inverse quantization unit 3111, thereby restoring the modified transform coefficients (transform coefficients after the transform by the second transform unit) d[][]. The secondary transform unit 31121 applies the secondary transform to the transform coefficients d[][] of the specified unit per transform unit TU. The secondary transform is applied only in the intra CU, and the transform basis is determined with reference to the intra prediction mode IntraPredMode. The selection of the transform basis is described later. The secondary transform unit 31121 outputs the restored modified transform coefficients d[][] to the core transform unit 31123.

[0198] The core transform unit 31123 acquires the transform coefficients d[][] or the modified transform coefficients d[][] restored by the secondary transform unit 31121, performs a transform, and derives the prediction error r[][]. The core transform unit 31123 outputs the prediction error r[][] to the addition unit 312.

[0199] (Secondary transform)

[0200] In the moving image encoding apparatus 11, a transform (positive secondary transform) is further applied to the transform coefficients after the core transform (DCT2 and DST7, etc.) of the difference image, the correlation remaining in the transform coefficients is removed, and the energy is concentrated in a part of the transform coefficients. In the moving image encoding apparatus 11, the positive transform unit 1032 included in the transform / quantization unit 103 and the inverse transform unit 152 included in the inverse transform / inverse quantization unit 105 are shown in FIG. 11. In the moving image decoding apparatus 3, on the contrary, the secondary transform is applied to the transform coefficients of a part or all of the decoded TUs, and the core transform (DCT2 and DST7, etc.) is applied to the transform coefficients after the secondary transform. Figure 19

[0201] In the secondary transform, the following processing is performed in accordance with the size of the TU and the intra prediction mode. Hereinafter, the processing of the secondary transform is described sequentially. Figure 13 is a diagram for explaining the secondary transform. In Figure 13 In the case of the 8x8 TU, the following processing is shown in FIG. 12: the transform coefficients d[][] of the 4x4 region are stored in the one-dimensional array u[] by the processing of S2, the one-dimensional array v[] is transformed from the one-dimensional array u[] by the processing of S3, and finally, d[][] is stored again by the processing of S4.

[0202] Figure 24 is a flowchart showing the processing of the secondary transform.

[0203] (S1: Setting of transform size and input / output size)

[0204] ​The secondary transform unit 31121 derives the size of the secondary transform (4 x 4 or 8 x 8), the number of output transform coefficients (nStOutSize), the number of applied transform coefficients (input transform coefficients) (nonZeroSize), and the number of sub-blocks to which the secondary transform is applied (numStX, numStY) according to the size (width nTbW, height nTbH) of the TU. The size of the secondary transform of 4 x 4 or 8 x 8 is represented by nStSize = 4, 8. In addition, the size of the secondary transform of 4 x 4 or 8 x 8 can be referred to as RST4x4, RST8x8, respectively.

[0205] The secondary transform unit 31121 outputs 48 transform coefficients by the secondary transform of RST8x8 in the case where the TU is equal to or larger than a predetermined size. In other cases, 16 transform coefficients are output by the secondary transform of RST4x4. In the case where the TU is 4 x 4, 16 transform coefficients are derived from 8 transform coefficients using RST4x4, and in the case where the TU is 8 x 8, 48 transform coefficients are derived from 8 transform coefficients using RST8x8. In other cases, 16 or 48 transform coefficients are output from 16 transform coefficients according to the size of the TU.

[0206] In the case where both nTbW and nTbH are equal to or larger than 8, log2StSize = 3, nStOutSize = 48

[0207] In other cases, log2StSize = 2, nStOutSize = 16

[0208] nStSize = 1 << log2StSize

[0209] In the case where both nTbW and nTbH are equal to or larger than 4 or 8 x 8, nonZeroSize = 8

[0210] In other cases, nonZeroSize = 16

[0211] numStX = (nTbH == 4 && nTbW > 8)? 2 : 1

[0212] numStY = (nTbW == 4 && nTbH > 8)? 2 : 1

[0213] (S2: Rearranged into one-dimensional signal)

[0214] The secondary transform unit 31121 processes the part of the TU's transform coefficients d[][] again by rearranging them into a one-dimensional array u[]. Specifically, in the secondary transform, u[] is derived from the two-dimensional transform coefficients d[][] of the object TU by referring to the transform coefficients for x = 0..nonZeroSize - 1. xC, yC are positions on the TU, which are derived from the arrangement DiagScanOrder representing the scan order and the position x of the transform coefficient in the sub-block.

[0215]

[0216] (S3: Application of the transform process)

[0217] The secondary transform unit 31121 performs a transform using a first transform basis (matrix) T on u[] (vector F') of length nonZeroSize to derive a one-dimensional array v'[] (vector V') of length nStOutSize as output.

[0218] This transform can be expressed in matrix operations by the following equation.

[0219] V' = T x F'

[0220] Here, the transform basis for the case of a transform size of 4 x 4 (RST 4 x 4) is referred to as the first transform basis T1. The transform basis for the case of a transform size of 8 x 8 (RST 8 x 8) is referred to as the second transform basis T2. T1 is a 16 x 16 (16 rows x 16 columns) matrix, i.e., the transform derives a 16 x 1 (16 rows x 1 column) vector V', i.e., a one-dimensional array v' [] of length 16, as the product of the 16 x 16 matrix T1 and the 16 x 1 (16 rows x 1 column) vector F'. T2 is a 48 x 16 (48 rows x 16 columns) matrix, i.e., the transform derives a 48 x 1 (48 rows x 1 column, length 48) vector V', i.e., a one-dimensional array v' [] of length 48, as the product of the 48 x 16 matrix T2 and the 16 x 1 vector F'.

[0221] Specifically, the secondary transform unit 31121 derives, depending on the secondary transform size nStSize (nTrS), the set number (stTrSetld) of the secondary transform derived from the intra prediction mode IntraPredMode, the stIdx indicating the transform basis of the secondary transform decoded from the coded data, and the corresponding transform matrix secTranMatrix [][] (transform basis T1 or T2). Furthermore, the secondary transform unit 31121 performs a product-sum operation of the transform matrix and the one-dimensional array u[] as shown in the following equation.

[0222] v' [i] = Clip3 (CoeffMin, CoeffMax,∑secTransMatrix [j] [i] * u [j])

[0223] Here, ∑ is a sum for j = 0..nonZeroSize - 1. Further, i is processed for 0..nStSize - 1. CoeffMin, CoeffMax indicate a range of values of the transform coefficient.

[0224] (S4: Two-dimensional arrangement of one-dimensional signal after transform processing)

[0225] The secondary transform unit 31121 arranges the coefficients v' [] of the one-dimensional array after the transform again at a designated position within the TU.

[0226] In the processing S4, the secondary transform unit 31121 arranges the coefficients v' [] of the length nStOutSize obtained through the above-described processing S3 in the region at the upper left of the arrangement d[][] of the transform coefficients.

[0227] The secondary transform unit 31121 performs the following processing for x = 0..nStSize - 1, y = 0..nStSize - 1. Specifically, the secondary transform unit 31121 applies the following formula in the case of IntraPredMode <= 34 or INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM.

[0228] d[(xSbIdx « log2StSize) + x][(ySbIdx « log2StSize) + y] = (y < 4)? v[x + (y « log2StSize)] : ((x < 4)? v[32 + x + ((y - 4) « 2)] :

[0229] d[(xSbIdx « log2StSize) + x][(ySbIdx « log2StSize) + y]

[0230] In the case other than this, the secondary transform unit 31121 applies the following formula.

[0231] d[(xSbIdx « log2StSize) + x][(ySbIdx « log2StSize) + y] = (y < 4)? v[x + (y « log2StSize)] : ((x < 4)? v[32 + x + ((y - 4) « 2)] : d[(xSbIdx « log2StSize) + x][(ySbIdx « log2StSize) + y] (Kernel transform unit 31123)

[0232] <Kernel transform>

[0233] A method that can switch the transform adaptively and a transform that is switched by an explicit flag, index, and prediction mode, and the like are referred to as a transform (first transform, core transform). The transform (core transform) used in the core transform is a separable transform constituted of a vertical transform and a horizontal transform. Further, a transform that separates a two-dimensional signal in the horizontal direction and the vertical direction can also be defined as the first transform. Further, in the image decoding device, a transform applied after the second transform (secondary transform) can also be defined as the first transform. The transform basis (transform matrix) of the core transform is DCT2, DST7, DCT8. In the core transform, the transform basis is switched independently for the vertical transform and the horizontal transform. Note that the transform to be selected is not limited to the above, and other transforms (transform basis) can also be used. Note that DCT2, DST7, DCT8, DST1, and DCT5 can be respectively denoted as DCT-II, DST-VII, DCT-VIII, DST-I, and DCT-V. Further, as a mode that explicitly skips the core transform, there can also be a transform skip.

[0234] In the core transform, there is explicit MTS and implicit MTS. In the case of explicit MTS, mts_idx is decoded from the encoded data, and the transform matrix is switched. In the case of implicit MTS, mts_idx is derived from the intra prediction mode, the block size.

[0235] Note that in the present embodiment, an example in which mts_idx is decoded in the CU unit or the TU unit is described, but the unit of decoding (switching) is not limited thereto.

[0236] mts_idx is a switching index for selecting the transform basis of the core transform. mts_idx has any one of 0, 1, 2, 3, 4, and derives the transform type trTypeHor in the horizontal direction and the transform type trTypeVer in the vertical direction.

[0237] Using Figure 18 The core transform described in the above description will be specifically described. Figure 18 The core transform section 1521 of Figure 12 The core transform section 31123 of Figure 19 The core transform section 1521 is one example of Figure 18The core transformation unit 1521 is configured with an MTS setting unit 15211 that sets the type of transformation used from among a plurality of transformation bases, a coefficient transformation processing unit 15212 that calculates the prediction residual r[][] from the (modified) transformation coefficient d[][] using the derived transformation, and a matrix transformation processing unit 15213 that performs the actual transformation. In the case where no secondary transformation is performed, the modified transformation coefficient is equal to the transformation coefficient. In the case where secondary transformation is performed, the modified transformation coefficient takes a different value from the transformation coefficient. The MTS setting unit 15211 is configured with an MTS setting unit 152111 that determines the derivation method of the index mts_idx of the transformation used, and an implicit MTS setting unit 152112 that derives mts_idx implicitly.

[0238] The MTS setting unit 152111 selects whether to perform explicit MTS, whether to perform implicit MTS, or whether to perform no MTS.

[0239] The MTS setting unit 152111 uses explicit MTS in the case where explicit MTS is effective (in the case where sps_explicit_mts_flag is 1), and uses mts_idx decoded from the encoded data in the subsequent processing. The flag explicitMtsEnabled indicating whether explicit MTS is effective can be set separately for the intra mode and the inter mode. In this case, it can also be that, in the case where the prediction mode PredMode is the inter mode (except for MODE_INTRA) and sps_explicit_mts_inter_enabled_flag is 1, or in the case where PredMode is the intra mode (MODE_INTRA) and sps_explicit_mts_intra_enabled_flag is 1, it is determined that explicit MTS is effective, and mts_idx is decoded from the encoded data. Furthermore, it can also be that, in the case where both the width and the height of the TU are 32 or less (nTbW <= 32 && nTbH <= 32), mts_idx is decoded. (implicitMtsEnabled flag setting)

[0240] The MTS setting unit 152111 sets the implicit MTS flag (implicitMtsEnabled) to 1 in the case where the MTS flag is effective (sps_mts_enabled_flag == 1) and the explicit MTS flag does not indicate that it is effective (explicitMtsEnabled == 0). More specifically, the MTS setting unit 152111 sets implicitMtsEnabled = 1 in the case where any one of the following conditions is satisfied, and sets implicitMtsEnabled = 0 in other cases.

[0241] • Intra sub-partition is enabled (IntraSubPartSplitType!= ISP_NO_SPLIT)

[0242] • CU sub-transform is enabled and TU is smaller than a specified size (cu_sbt_flag == 1 and Max(nTbW, nTbH) < 32)

[0243] • Explicit MTS is off (both sps_explicit_mts_intra_enabled_flag and sps_explicit_mts_inter_enabled_flag are 0) and PredMode is MODE_INTRA

[0244] The MTS setting section 152111 sets mts_idx = 0 in other cases.

[0245] (Explicit MTS)

[0246] The TU decoding section 3024 decodes mts_idx from the encoded data in the case where explicit MTS is effective (in the case where sps_explicit_mts_flag is 1).

[0247] (Limitation of Explicit MTS)

[0248] The TU decoding section 3024 can limit the range (category) of the transform matrix selected by the transform section depending on whether secondary transform is effective or not. For example, the TU decoding section 3024 decodes mts_idx of the maximum value cMaxSt1 in the case where explicit MTS is effective and secondary transform is effective (stIdx!= 0). In other cases, the TU decoding section 3024 decodes mts_idx of the maximum value cMaxSt0 in the case where secondary transform is not effective (stIdx == 0). Here, cMaxSt0 > cMaxSt1 is assumed.

[0249] (Limitation Example 1)

[0250] TU decoding section 3024 sets mts_idx = 0 (trTypeHor = trTypeVer = 0 = DCT2) in the case where explicit MTS is effective and secondary transform is effective (stIdx ≠ 0). At this time, the maximum value cMax of mts_idx = 0. In other cases, decode any one of 0, 1, 2, 3, 4 in mts_idx = 0, 1, 2, 3, 4. At this time, the maximum value cMax of mts_idx = 4. As described later, mts_idx = 1, 2, 3, 4 can be DST7 and DCT8 combination, DCT8 and DST7 combination, or DCT8 and DCT8 combination as trTypeHor, trTypeVer, respectively. Note that, instead of DST7, DST1, DCT4, or a combination of DCT2 and pre- and post-processing can be performed.

[0251] (Limiting example 2)

[0252] TU decoding section 3024 decodes mts_idx in the case where explicit MTS is effective and secondary transform is effective (stIdx ≠ 0). mtx_idx is 0 (trTypeHor = trTypeVer = 0 = DCT2) or 1 (trTypeHor = trTypeVer = 1 = DST7). At this time, the maximum value cMax of mts_idx = 1. In other cases, decode any one of 0, 1, 2, 3, 4 as mts_idx. At this time, the maximum value cMax of mts_idx = 4. Note that, as described later, mts_idx = 2, 3, 4 can be DST7 and DCT8 combination or DCT8 and DST7 combination, DCT8 and DCT8 combination as trTypeHor, trTypeVer, respectively.

[0253] (Limiting example 3)

[0254] TU decoding section 3024 decodes any one of 0, 1, 2 as mts_idx in the case where explicit MTS is effective and secondary transform is enabled (stIdx ≠ 0). At this time, the maximum value cMax of mts_idx = 2. In other cases, decode any one of 0, 1, 2, 3, 4 as mts_idx. At this time, the maximum value cMax of mts_idx = 4.

[0255] According to the above configuration, in the case of secondary transform, the effective range of MTS can be limited, and thus the effect of simplifying encoding is achieved. For example, in limiting example 2, in the case of DCT8 where secondary transform and the effect overlap, secondary transform is not performed, and thus the effect of reducing the overhead generated by mts_idx and improving the encoding efficiency is achieved.

[0256] <Summary of explicit MTS>

[0257] An image decoding apparatus includes a transform unit that transforms transform coefficients on a per-TU basis, the transform unit including a second transform unit that applies a transform using a transform matrix to input transform coefficients when secondary transform is effective, and a first transform unit that applies a transform using one transform matrix indicated by mtx_idx selected from two or more transform matrices to the transform coefficients, and a TU decoding unit that decodes mts_idx, the TU decoding unit decoding a value of a first range as mts_idx when secondary transform is effective (stIdx!= 0), and decoding a value of a second range when secondary transform is not effective (stIdx == 0), the second range including the first range. Further, the TU decoding unit decodes mts_idx. Here, mts_idx has a configuration in which a maximum value is cMaxSt1 when secondary transform is effective (stIdx!= 0), and a maximum value is cMaxSt0 when secondary transform is not effective (stIdx == 0), and cMaxSt1 < cMaxSt0.

[0258] (Implicit MTS)

[0259] The implicit MTS setting unit 152112 performs the following processing in the case of implicit MTS.

[0260] (SM001) The implicit MTS setting unit 152112 sets either of 0 (DCT2) or 1 (DST7) as the transform types tyTypeHor and tyTypeVer according to the intra prediction mode IntraPredMode and the TU size in the case of using an intra sub-partition mode (IntraSubPartSplitType!= ISP_NO_SPLIT) other than the above. Figure 14

[0261] (SM002) The implicit MTS setting unit 152112 sets either of 1 (DST7) or 2 (DCT8) as tyTypeHor and tyTypeVer according to cu_sbt_horizontal_flag and cu_sbt_pos_flag in the case of using a sub-block transform (cu_sbt_flag == 1) other than the above. Figure 15

[0262] (SM003) The implicit MTS setting unit 152112 sets either of 0 (DCT2) or 1 (DST7) as tyTypeHor and tyTypeVer according to the TU size (width nTbW, height nTbH) in the case of using a default implicit MTS other than the above. Specifically, as shown in​​Figure 21 As shown, in a case where the width nTbW is in a prescribed range as the horizontal transform type trTypeHor (S1301), 1 (DCT1) is set (S1302), and in other cases, 0 (DCT2) is set (S1303). Similarly, in a case where the height nTbH is in a prescribed range as the vertical transform type trTypeVer (S1304), 1 (DCT1) is set (S1305), and in other cases, 0 (DCT2) is set (S1306).

[0263] trTypeHor = (nTbW >= 4 && nTbW <= 16 && nTbW <= nTbH)? 1 : 0

[0264] trTypeVer = (nTbH >= 4 && nTbH <= 16 && nTbH <= nTbW)? 1 : 0

[0265] Note that the prescribed range is not limited to the above. For example, it can be as follows.

[0266] trTypeHor = (nTbW >= 4 && nTbW <= 8 && nTbW <= nTbH)? 1 : 0

[0267] trTypeVer = (nTbH >= 4 && nTbH <= 8 && nTbH <= nTbW)? 1 : 0

[0268] The above-described default implicit MTS is a mode of the most common implicit MTS.

[0269] (Embodiment 1 of implicit MTS)

[0270] The MTS setting section 15211 does not perform implicit MTS and sets implicitMtsEnabled to 0 in a case where secondary transform is enabled (stIdx!= 0). Specifically, as shown in Figure 20 In the above-described (implicit MTS flag setting), the MTS setting section 152111 sets implicitMtsEnabled = 1 (S1504) in a case where any one of the following conditions is satisfied and secondary transform is not enabled (stIdx!= 0) (S1500), and sets implicitMtsEnabled = 0 (S1505) in other cases.

[0271] • (S1501) In a case where intra sub-partition is enabled (IntraSubPartSplitType!= ISP_NO_SPLIT)

[0272] (S1502) When CU sub-transform is enabled and TU is smaller than the specified size (cu_sbt_flag == 1 and Max(nTbW, nTbH) < 32)

[0273] (S1503) Explicit MTS is disabled (both sps_explicit_mts_intra_enabled_flag and sps_explicit_mts_ihter_enabled_flag are 0) and PredMode is MODE_INTRA

[0274] It should be noted that, Figure 20 The determinations in S1501 to S1503 indicated by the dotted line may be different. For example, the determination of intra-frame sub-division may not be performed, or sub-block transformation may not be performed. In addition, other prediction and transformation determinations may be added.

[0275] According to the above configuration, even when MTS is valid, implicit MTS is not used when using the secondary transform (stIdx!=0). Thus, when using the secondary transform, DCT2 is used, thereby achieving an effect of improving coding efficiency.

[0276] (Implementation Method 2 of Implicit MTS)

[0277] When the MTS flag is valid (sps_mts_enabled_flag==1) and the explicit MTS flag does not indicate validity (explicitMtsEnabled==0), the MTS setting unit 152112 can derive trTypHor=trTypeVer=0. For example, SM000 can be performed before the above-mentioned SM001.

[0278] Figure 21 This is a diagram illustrating the operation of the implicit MTS setting unit 152112.

[0279] (SM000) When stIdx!=0, the implicit MTS setting unit 152112 derives trTypeHor=trTypeVer=0.

[0280] It should be noted that if Figure 21 As shown in SM003, the implicit MTS setting unit 152112 can derive the transform type by deriving the default implicit MTS (SM003) as described above for the case of stIdx == 0. Alternatively, the transform type can be derived through SM001 and SM002.

[0281] According to the above configuration, in a case where implicit MTS is effective and secondary transform is effective, DCT2 is used as MTS, and an effect of improving coding efficiency is achieved.

[0282] (Embodiment 3 of implicit MTS)

[0283] Also, the implicit MTS setting section 152112 does not use implicit MTS (e.g., sets implicitMtsEnabled to 0) in a case where secondary transform is enabled (stIdx!= 0) and neither MTS obtained by intra sub-partition mode (SM001, IntraSubPartSplitType!= ISP_NO_SPLIT) nor MTS obtained by sub-block transform (SM002, cu_sbt_flag == 1) is used.

[0284] Further, in a case where secondary transform is enabled (stIdx!= 0) and neither MTS obtained by intra sub-partition mode (SM001, IntraSubPartSplitType!= ISP_NO_SPLIT) nor MTS obtained by sub-block transform (SM002, cu_sbt_flag == 1) is used, trTypeHor = trTypeVer = 0 is derived. For example, as shown in SM003', it is possible to replace the above-described SM003. Figure 22

[0285] (SM003') The implicit MTS setting section 152112 sets either of 0 (DCT2) or 1 (DST7) as tyTypeHor, tyTypeVer according to secondary transform and TU size (width nTbW, height nTbH) in a case other than the above (default implicit MTS). For example, the implicit MTS setting section 152112 selects 1 (DCT1) as the horizontal transform type trTypeHor in a case where stIdx == 0 and the width nTbW is in a prescribed range (S1301'), and sets 0 (DCT2) in a case other than this (S1303). Similarly, 1 (DCT1) is set as the vertical transform type trTypeVer in a case where stIdx == 0 and the height nTbH is in a prescribed range (S1304'), and 0 (DCT2) is set in a case other than this (S1306).

[0286] trTypeHor = (stIdx == 0 && nTbW >= 4 && nTbW <= 16 && nTbW <= nTbH)? 1 : 0

[0287] ​trTypeVer = (stIdx == 0 && nTbH >= 4 && nTbH <= 16 && nTbH <= nTbW)? 1 : 0

[0288] According to the above configuration, in a case where implicit MTS is valid and secondary transform is valid, DCT2 is used as a default MTS, and thus an effect of improving coding efficiency is achieved.

[0289] The MTS setting section 15211 derives an index trType of a transform set used by the following equation, and outputs it to the coefficient transform processing section 15212. The coefficient transform processing section 15212 outputs the input trType to the transform matrix deriving section 152131. The MTS setting section 152111 derives a value indicating the MTS used by the following equation.

[0290] In a case where mts_idx == 0, trTypeHor = 0 trTypeVer = 0

[0291] In a case where mts_idx == 1, trTypeHor = 1 trTypeVer = 1

[0292] In a case where mts_idx == 2, trTypeHor = 2 trTypeVer = 1

[0293] In a case where mts_idx == 3, trTypeHor = 1 trTypeVer = 2

[0294] In a case where mts_idx == 4, trTypeHor = 2 trTypeVer = 2

[0295] Note that the transform bases corresponding to the cases where tyType (trTypeHor or trTypeVer) is 0, 1, 2 can be DCT2, DST7, DCT8.

[0296] The coefficient transform processing section 15212 is configured of a vertical transform section 152121 that performs vertical transform on the modified transform coefficients d [][], and a horizontal transform section 152123 that performs horizontal transform.

[0297] The vertical transform section 152121 (coefficient transform processing section 15212) performs the following processing.

[0298] e [x] [y] = å (transMatrix [y] [j] x d [x] [j]) (j = 0..nTbS-1)

[0299] Here, transMatrix [ ][ ] (= transMatrixV [ ] [ ] ) is a transform basis represented by an nTbS x nTbS matrix derived using trTypeVer. nTbS is the height nTbH of the TU. In the case of DCT2 of 4 x 4 transform (nTbS = 4) with trType = 0, for example, transMatrix = { { 29, 55, 74, 84}, { 74, 74, 0, -74}, { 84, -29, -74, 55}, { 55, -84, 74, -29}} is used. The symbol ∑ means a process of adding up the indices j = 0..nTbS - 1 of the product of the matrix transMatrix [y] [j] and the transform coefficient d [x] [j]. That is, e [x] [y] is obtained by arranging the column obtained by multiplying the vector x [j] (j = 0..nTbS - 1) composed of d [x] [j] (j = 0..nTbS - 1) which is each column of d [x] [y] by the element transMatrix [y] [j] of the matrix.

[0300] The intermediate clipping section 152122 derives the intermediate value g [ ] [ ] by clipping the intermediate value e [ ] [ ] and transmits it to the horizontal transform section 152123.

[0301] g [x] [y] = Clip3 (coeffMin, coeffMax, (e [x] [y] + 64) » 7)

[0302] The 64, 7 in the above expression are values determined in accordance with the bit depth of the transform basis, and the transform basis is assumed to be 7 bits in the above expression. Further, coeffMin and coeffMax are the minimum and maximum values of the clipping.

[0303] The horizontal transform section 152123 (coefficient transform processing section 15212) performs the following processing. transMatrix [ ] [ ] (= transMatrixH [ ] [ ] ) is a transform basis represented by an nTbS x nTbS matrix derived using trTypeHor. nTbS is the width nTbW of the TU. The horizontal transform section 152123 transforms the intermediate value g [x] [y] into the prediction residual r [x] [y] by horizontal one-dimensional transform.

[0304] r [x] [y] = ∑ transMatrix [x] [j] x g [j] [y] (j = 0..nTbS - 1)

[0305] The above symbol ∑ means a process of adding the index j for j = 0..nTbS - 1 to the product of the matrix transMatrix[x][j] and g[j][y]. That is, r[x][y] is obtained by arranging the rows obtained by the product of g[j][y] (j = 0..nTbS - 1) which is each row of g[x][y] and the matrix transMatrix.

[0306] The prediction residual r[][] is transferred from the horizontal transform section 152123 to the adder 312.

[0307] The vertical transform section 152121 and the horizontal transform section 152123 are transformed by the matrix transform processing section 15213. The matrix transform processing section 15213 is constituted by a transform matrix derivation section 152131 and a transform processing section 152132.

[0308] The transform matrix derivation section 152131 derives the transform matrix transMatrix[][] from the length (nTbW, nTbH) of the TU and the index tyType (trTypeHor, trTypeVer) of the core transform.

[0309] The matrix transform processing section 15213 transforms the one-dimensional array xx[j] input using the derived transform matrix transMatrix[][] into the one-dimensional array yy[i], and performs the vertical transform and the horizontal transform. In the vertical transform, the transform coefficient d[x][j] of the x column is input as the one-dimensional transform coefficient xx[j] to perform the transform. In the horizontal transform, the intermediate coefficient g[j][y] of the y row is input as xx[j] to perform the transform.

[0310] yy[i] = ∑ (transMatrix[i][j] x xx[j]) (j = 0..nTbS - 1)

[0311] <Secondary transform>

[0312] The TU decoding section 3024 can limit the range of the value of stIdx to be decoded depending on whether the value of mts_idx is valid in the case of decoding the secondary transform stIdx.

[0313] (Limiting example 1)

[0314] The TU decoding section 3024 sets the maximum value cMax of stIdx to be decoded to 1 in the case where mts_idx is 0, and decodes stIdx = 0 to 2. In the case other than this, that is, in the case where mts_idx = 1, 2, 3, 4, stIdx is not decoded from the coded data, but stIdx = 0 is derived.

[0315] (Restriction example 2)

[0316] The TU decoding section 3024 sets the maximum value cMax of stIdx to be decoded to 1 in the case where mts_idx is 0...1, and decodes stIdx = 0...2. In the case where mts_idx is 2, 3, 4, stIdx is not decoded from the encoded data, but is derived as stIdx = 0.

[0317] (Restriction example 3)

[0318] The TU decoding section 3024 sets the maximum value cMax of stIdx to be decoded to 1 in the case where mts_idx is 0...2, and decodes stIdx = 0...2. In the case where mts_idx is 3, 4, stIdx is not decoded from the encoded data, but is derived as stIdx = 0.

[0319] According to the above configuration, the variable stIdx indicating the type of secondary transform is decoded only in the case where the transform in the range specified by MTS is performed, so the effective range of secondary transform can be limited, and thus the effect of simplifying the encoding is achieved. Further, for example, in restriction example 2, in the case where DCT8 is overlapped with the secondary transform and the effect, the secondary transform is not performed, and thus the effect of reducing the overhead generated by stIdx and improving the encoding efficiency is achieved.

[0320] The addition section 312 adds the prediction image of the block input from the prediction image generation section 308 to the prediction error input from the inverse quantization / inverse transform section 311 on a pixel-by-pixel basis, and generates a decoded image of the block. The addition section 312 stores the decoded image of the block in the reference picture storage 306, and outputs to the loop filter 305.

[0321] (Configuration of a moving image encoding apparatus)

[0322] Next, the configuration of the moving image encoding apparatus 11 of the present embodiment will be described. Figure 16 is a block diagram showing the configuration of the moving image encoding apparatus 11 of the present embodiment. The moving image encoding apparatus 11 is configured to include a prediction image generation section 101, a subtraction section 102, a transform / quantization section 103, an inverse quantization / inverse transform section 105, an addition section 106, a loop filter 107, a prediction parameter storage (prediction parameter storage section, frame memory) 108, a reference picture storage (reference image storage section, frame memory) 109, an encoding parameter determination section 110, a parameter encoding section 111, and an entropy encoding section 104.

[0323] The predicted image generation unit 101 generates a predicted image for each CU, which is a region obtained by dividing each picture of each image T. The predicted image generation unit 101 performs the same operation as the predicted image generation unit 308 already described, and the description thereof is omitted here.

[0324] The subtraction unit 102 generates a prediction error by subtracting the pixel value of the predicted image of the block input from the predicted image generation unit 101 from the pixel value of the image T. The subtraction unit 102 outputs the prediction error to the transformation and quantization unit 103 .

[0325] The transform / quantization unit 103 calculates transform coefficients by frequency transforming the prediction error input from the subtraction unit 102 and derives quantized transform coefficients by quantization. The transform / quantization unit 103 outputs the quantized transform coefficients to the entropy coding unit 104 and the inverse quantization / inverse transform unit 105.

[0326] like Figure 19 As shown, the transform / quantization unit 103 includes a positive kernel transform 10321 (first transform unit) and a positive quadratic transform unit 10322 (second transform unit).

[0327] The forward quadratic transform applied in the moving picture encoding device 11 performs substantially the same processing as that of the quadratic transform applied in the moving picture decoding device 31 except that the processes S1 to S4 of the quadratic transform applied in the moving picture decoding device 31 are applied in the reverse order of processes S1, S4, S3, and S2.

[0328] In process S1, the forward quadratic transform unit 10322 performs the same process as the quadratic transform unit 31121, except that the input and output of the quadratic transform are of length nStOutSize and nonZeroSize, respectively.

[0329] In processing S4, the quadratic transform unit 10322 derives a one-dimensional array v[] of nStOutSize (or nStSize*nStSize) from the transform coefficients d[][] at the specified position within the TU.

[0330] In processing S3, the quadratic transformation unit 10322 obtains the one-dimensional array u[] (vector F) of nonZeroSize from the one-dimensional array v[] (vector V) of nStOutSize and the transformation matrix T[][] through the following transformation.

[0331] F=trans(T)×V

[0332] Here, trans(T) is the transposed matrix of T. The secondary transformation unit can also derive a one-dimensional array u[] (vector F) by the following formula.

[0333] F=Tinv×V

[0334] Here, Tinv is an inverse matrix of T. T is constituted by the first transform basis T1 and the second transform basis T2. Note that it is also possible that the secondary transform unit uses an orthogonal matrix with respect to T, and thus sets trans(T) of T as Tinv.

[0335] Note that in actual processing, T is a matrix of integer values, and thus it is not the case that T x Tinv = I (an identity matrix), but rather a constant multiple of the identity matrix (T x Tinv = K2 x I, K2 being a constant). In this case, it is also possible that the secondary transform unit uses a matrix that is a constant multiple of the inverse matrix as Tinv, and directly uses the inverse matrix with respect to the transposed matrix.

[0336] In the process S2, the primary secondary transform unit 10322 rearranges the one-dimensional array u[] of nonZeroSize into a two-dimensional arrangement, and derives the transform coefficients d[][].

[0337]

[0338] The inverse quantization / inverse transform unit 105 is the same as the inverse quantization / inverse transform unit 311 in the moving image decoding apparatus 31, and thus the description is omitted. The calculated prediction error is input to the addition unit 106. Figure 15

[0339] In the entropy encoding unit 104, the quantized transform coefficients are input from the transform / quantization unit 103, and the encoding parameters are input from the parameter encoding unit 111. The encoding parameters are, for example, predMode.

[0340] The entropy encoding unit 104 entropy-encodes the partition information, the prediction parameters, the quantized transform coefficients, and the like, and generates and outputs the encoded stream Te.

[0341] The parameter encoding unit 111 includes an unillustrated header encoding unit 1110, a CT information encoding unit 1111, a CU encoding unit 1112 (prediction mode encoding unit), an inter prediction parameter encoding unit 112, and an intra prediction parameter encoding unit 113. The CU encoding unit 1112 further includes a TU encoding unit 1114.

[0342] The following describes the outline of the operation of each module. The parameter encoding unit 111 performs encoding processing of the header information, the partition information, the prediction information, the quantized transform coefficients, and the like.

[0343] The CT information encoding unit 1111 encodes the QT, MT (BT, TT) partition information, and the like, in accordance with the encoding data.

[0344] The CU encoding unit 1112 encodes the CU information, the prediction information, the TU partition flag, the CU residual flag, and the like.

[0345] ​The TU coding section 1114 encodes QP update information (quantization correction value) and quantized prediction error (residual_coding) in the case where the prediction error is included in the TU.

[0346] The CT information coding section 1111 and the CU coding section 1112 supply syntax elements such as inter prediction parameters, intra prediction parameters (intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_remainder), and quantized transform coefficients to the entropy coding section 104.

[0347] (Configuration of the intra prediction parameter coding section 113)

[0348] The intra prediction parameter coding section 113 derives forms (e.g., intra_luma_mpm_idx, intra_luma_mpm_remainder, etc.) for coding in accordance with IntraPredMode input from the coding parameter determination section 110. The intra prediction parameter coding section 113 includes the same configuration as that of the intra prediction parameter derivation section 304 for deriving the intra prediction parameters.

[0349] Figure 17 is a schematic diagram showing the configuration of the intra prediction parameter coding section 113 of the parameter coding section 111. The intra prediction parameter coding section 113 is configured to include a parameter coding control section 1131, a luma intra prediction parameter derivation section 1132, and a chroma intra prediction parameter derivation section 1133.

[0350] IntraPredModeY and IntraPredModeC are input from the coding parameter determination section 110 to the parameter coding control section 1131. The parameter coding control section 1131 determines intra_luma_mpm_flag with reference to mpmCandList[] of the MPM candidate list derivation section 30421. Then, intra_luma_mpm_flag and IntraPredModeY are output to the luma intra prediction parameter derivation section 1132. In addition, IntraPredModeC is output to the chroma intra prediction parameter derivation section 1133.

[0351] The luma intra prediction parameter derivation section 1132 is configured to include the MPM candidate list derivation section 30421 (candidate list derivation section), the MPM parameter derivation section 11322 (parameter derivation section), and the non-MPM parameter derivation section 11323 (coding section, derivation section).

[0352] The MPM candidate list derivation section 30421 derives the mpmCandList[] with reference to the intra prediction mode of the neighboring block stored in the prediction parameter storage 108. The MPM parameter derivation section 11322 derives the intra_luma_mpm_idx from the IntraPredModeY and the mpmCandList[] in the case where the intra_luma_mpm_flag is 1, and outputs it to the entropy coding section 104. The non-MPM parameter derivation section 11323 derives the RemIntraPredMode from the IntraPredModeY and the mpmCandList[] in the case where the intra_luma_mpm_flag is 0, and outputs the intra_luma_mpm_remainder to the entropy coding section 104.

[0353] The chroma intra prediction parameter derivation section 1133 derives the intra_chroma_pred_mode from the IntraPredModeY and the IntraPredModeC and outputs it.

[0354] The addition section 106 adds the pixel value of the block prediction image input from the prediction image generation section 101 and the prediction error input from the inverse quantization / inverse transform section 105 per pixel to generate a decoded image. The addition section 106 stores the generated decoded image in the reference picture storage 109.

[0355] The loop filter 107 performs deblocking filtering, SAO, and ALF on the decoded image generated by the addition section 106. Note that the loop filter 107 does not necessarily include all of the three filters described above, and may, for example, include only a deblocking filter.

[0356] The prediction parameter storage 108 stores the prediction parameters generated by the encoding parameter determination section 110 in a predetermined position per target picture and CU.

[0357] The reference picture storage 109 stores the decoded image generated by the loop filter 107 in a predetermined position per target picture and CU.

[0358] The encoding parameter determination section 110 selects one of a plurality of sets of encoding parameters. The encoding parameters refer to the QT, BT, or TT split information described above, the prediction parameters, or parameters generated in association therewith as encoding targets. The prediction image generation section 101 generates a prediction image using these encoding parameters.

[0359] The encoding parameter determination section 110 calculates the RD cost value indicating the amount of information and the encoding error for each of the plurality of sets. The encoding parameter determination section 110 selects the set of encoding parameters for which the calculated cost value is the smallest. Thus, the entropy encoding section 104 outputs the selected set of encoding parameters as the encoded stream Te. The encoding parameter determination section 110 stores the determined encoding parameters in the prediction parameter storage 108.

[0360] Note that a part of the moving image encoding apparatus 11, the moving image decoding apparatus 31 in the above-described embodiments, for example, the entropy decoding section 301, the parameter decoding section 302, the loop filter 305, the prediction image generation section 308, the inverse quantization / inverse transform section 311, the addition section 312, the prediction image generation section 101, the subtraction section 102, the transform / quantization section 103, the entropy encoding section 104, the inverse quantization / inverse transform section 105, the loop filter 107, the encoding parameter determination section 110, and the parameter encoding section 111 can be implemented by a computer. In this case, the program for realizing the control function can be recorded in a computer-readable recording medium, and a computer system can be caused to read the program recorded in the recording medium and execute it. Note that the "computer system" referred to here means a computer system built into either of the moving image encoding apparatus 11 and the moving image decoding apparatus 31, and a computer system including hardware such as an OS and a peripheral device. Further, the "computer-readable recording medium" means a removable medium such as a floppy disk, a magneto-optical disk, a ROM, a CD-ROM, and a storage device such as a hard disk built into a computer system. Furthermore, the "computer-readable recording medium" can include a medium that dynamically stores a program for a short period of time, such as a communication line in the case of transmitting a program via a network such as the Internet or a communication line such as a telephone line, and a medium that stores a program for a fixed period of time, such as a volatile memory inside a computer system that is a server or a client in this case. Further, the program can be a program for realizing a part of the above-described functions, and can also be a program that realizes the above-described functions by being combined with a program already recorded in a computer system.

[0361] Further, a part or all of the moving image encoding apparatus 11 and the moving image decoding apparatus 31 in the above-described embodiments can also be implemented as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the moving image encoding apparatus 11 and the moving image decoding apparatus 31 can be processed individually by a processor, or a part or all of them can be integrated and processed by a processor. Further, the method of integration is not limited to an LSI, and can be realized by a dedicated circuit or a general-purpose processor. Furthermore, in the case where an integrated circuit technology replacing LSIs emerges as a result of advances in semiconductor technology, an integrated circuit based on this technology can also be used.

[0362] An embodiment of the present invention has been described above in detail with reference to the drawings. However, the specific configuration is not limited to the above embodiment, and various design changes can be made without departing from the gist of the present invention.

[0363] [Application Examples]

[0364] The moving image encoding device 11 and the moving image decoding device 31 can be installed in various devices that transmit, receive, record, and reproduce moving images. It should be noted that the moving images can be natural moving images captured by a camera or the like, or artificial moving images (including CG and GUI) generated by a computer or the like.

[0365] First, refer to Figure 2 A case where the above-described moving picture encoding device 11 and moving picture decoding device 31 can be used for transmission and reception of moving pictures will be described.

[0366] Figure 2 , a block diagram showing the structure of a transmitting device PROD_A equipped with a motion picture encoding device 11 is shown in FIG. Figure 2 As shown, the transmitting device PROD_A includes an encoding unit PROD_A1 that encodes a moving image to obtain encoded data, a modulating unit PROD_A2 that modulates a carrier wave using the encoded data obtained by the encoding unit PROD_A1 to obtain a modulated signal, and a transmitting unit PROD_A3 that transmits the modulated signal obtained by the modulating unit PROD_A2. The moving image encoding device 11 described above is used as the encoding unit PROD_A1.

[0367] As a source of motion images input to the encoding unit PROD_A1, the transmitting device PROD_A may further include: a camera PROD_A4 for shooting motion images, a recording medium PROD_A5 for recording motion images, an input terminal PROD_A6 for inputting motion images from the outside, and an image processing unit A7 for generating or processing images. Figure 2 The example in which the transmitting device PROD_A includes all of these configurations is shown, but some of them may be omitted.

[0368] It should be noted that the recording medium PROD_A5 may be a medium recording unencoded moving images, or a medium recording moving images encoded using a recording encoding method different from the transmission encoding method. In the latter case, a decoding unit (not shown) that decodes the encoded data read from the recording medium PROD_A5 using the recording encoding method is preferably interposed between the recording medium PROD_A5 and the encoding unit PROD_A1.

[0369] Further, in Figure 2 a block diagram showing the configuration of a reception apparatus PROD_B mounting the moving picture decoding apparatus 31 is shown. As Figure 2 shown, the reception apparatus PROD_B is provided with a reception section PROD_B1 that receives a modulated signal, a demodulation section PROD_B2 that obtains encoded data by demodulating the modulated signal received by the reception section PROD_B1, and a decoding section PROD_B3 that obtains a moving picture by decoding the encoded data obtained by the demodulation section PROD_B2. The moving picture decoding apparatus 31 described above is used as this decoding section PROD_B3.

[0370] The reception apparatus PROD_B can further be provided with a display PROD_B4 that displays the moving picture, a recording medium PROD_B5 that records the moving picture, and an output terminal PROD_B6 that outputs the moving picture to the outside, as a supply destination of the moving picture output by the decoding section PROD_B3. In Figure 2 the reception apparatus PROD_B is exemplified as being provided with all of these, but a part of them can be omitted.

[0371] Note that the recording medium PROD_B5 can be a medium for recording a moving picture that has not been encoded, or a medium that has been encoded in an encoding scheme for recording that is different from an encoding scheme for transmission. In the latter case, it is preferable that an encoding section (not shown) that encodes the moving picture obtained from the decoding section PROD_B3 in the encoding scheme for recording be interposed between the decoding section PROD_B3 and the recording medium PROD_B5.

[0372] Note that the transmission medium that transmits the modulated signal can be wireless or wired. Further, the transmission scheme that transmits the modulated signal can be broadcasting (here, a transmission scheme in which the transmission destination is not determined in advance), or communication (here, a transmission scheme in which the transmission destination is determined in advance). That is, the transmission of the modulated signal can be achieved by any one of wireless broadcasting, wired broadcasting, wireless communication, and wired communication.

[0373] For example, a broadcasting station (broadcasting device, etc.) / receiving station (television receiver, etc.) of terrestrial digital broadcasting is one example of a transmission apparatus PROD_A / reception apparatus PROD_B that transceives a modulated signal by wireless broadcasting. Further, a broadcasting station (broadcasting device, etc.) / receiving station (television receiver, etc.) of cable television broadcasting is one example of a transmission apparatus PROD_A / reception apparatus PROD_B that transceives a modulated signal by wired broadcasting.

[0374] Moreover, a server (workstation or the like) / client (television receiver, personal computer, smartphone or the like) using a VOD (Video On Demand) service, a moving image sharing service or the like is an example of a transmission apparatus PROD_A / reception apparatus PROD_B that transmits / receives a modulated signal (typically, either wireless or wired is used as a transmission medium in a LAN, and wired is used as a transmission medium in a WAN). Here, the personal computer includes a desktop PC, a laptop PC, and a tablet PC. Moreover, the smartphone also includes a multifunctional portable telephone terminal.

[0375] Note that the client of the moving image sharing service has a function of encoding a moving image captured by a camera and uploading it to the server in addition to a function of decoding encoded data downloaded from the server and displaying it on a display. That is, the client of the moving image sharing service functions as both the transmission apparatus PROD_A and the reception apparatus PROD_B.

[0376] Next, a case in which the moving image encoding apparatus 11 and the moving image decoding apparatus 31 described above are used for recording and reproduction of a moving image will be described with reference to Figure 3

[0377] Figure 3 A block diagram showing the configuration of a recording apparatus PROD_C equipped with the moving image encoding apparatus 11 described above is shown in FIG. 1. As shown in FIG. 1, the recording apparatus PROD_C is provided with an encoding section PROD_C1 that obtains encoded data by encoding a moving image and a writing section PROD_C2 that writes the encoded data obtained by the encoding section PROD_C1 to a recording medium PROD_M. The moving image encoding apparatus 11 described above is used as the encoding section PROD_C1. Figure 3 Note that the recording medium PROD_M can be a recording medium of a type that is built in the recording apparatus PROD_C such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or the like (1), a recording medium of a type that is connected to the recording apparatus PROD_C such as an SD memory card, a USB (Universal Serial Bus) flash memory, or the like (2), or a recording medium that is loaded into a drive apparatus (not shown) built in the recording apparatus PROD_C such as a DVD (Digital Versatile Disc, registered trademark), a BD (Blu-ray Disc, registered trademark), or the like (3).

[0378]

[0379] ​​Further, as a supply source of the moving image input to the encoding section PROD_C1, the recording apparatus PROD_C can further include a camera PROD_C3 that captures a moving image, an input terminal PROD_C4 for externally inputting a moving image, a reception section PROD_C5 for receiving a moving image, and an image processing section PROD_C6 that generates or processes an image. In Figure 3 The recording apparatus PROD_C is configured to include all of these, but a part of them can be omitted.

[0380] Note that the reception section PROD_C5 can receive a moving image that is not encoded, or can receive encoded data that is encoded in a transmission encoding scheme different from an encoding scheme for recording. In the latter case, it is preferable to interpose a transmission decoding section (not shown) that decodes encoded data encoded in a transmission encoding scheme between the reception section PROD_C5 and the encoding section PROD_C1.

[0381] As such a recording apparatus PROD_C, for example, a DVD recorder, a BD recorder, an HDD (Hard Disk Drive) recorder, or the like can be given (in this case, the input terminal PROD_C4 or the reception section PROD_C5 is a main supply source of a moving image). Further, a camcorder (in this case, the camera PROD_C3 is a main supply source of a moving image), a personal computer (in this case, the reception section PROD_C5 or the image processing section C6 is a main supply source of a moving image), a smartphone (in this case, the camera PROD_C3 or the reception section PROD_C5 is a main supply source of a moving image), or the like are also examples of such a recording apparatus PROD_C.

[0382] Further, a block diagram showing a configuration of a reproduction apparatus PROD_D that mounts the moving image decoding apparatus 31 described above is shown in Figure 3 As shown in Figure 3 The reproduction apparatus PROD_D includes a readout section PROD_D1 that reads out encoded data written in a recording medium PROD_M, and a decoding section PROD_D2 that obtains a moving image by decoding the encoded data read out by the readout section PROD_D1. The moving image decoding apparatus 31 described above is used as the decoding section PROD_D2.

[0383] Note that the recording medium PROD_M can be a recording medium of a type (1) that is built in the reproduction apparatus PROD_D like an HDD, an SSD, or the like, a recording medium of a type (2) that is connected to the reproduction apparatus PROD_D like an SD memory card, a USB flash memory, or the like, or a recording medium of a type (3) that is loaded into a drive device (not shown) built in the reproduction apparatus PROD_D like a DVD, a BD, or the like.

[0384] Further, as a supply destination of the moving image output by the decoding section PROD_D2, the reproduction apparatus PROD_D can further include a display PROD_D3 that displays the moving image, an output terminal PROD_D4 that outputs the moving image to the outside, and a transmission section PROD_D5 that transmits the moving image. In Figure 3 The reproduction apparatus PROD_D is exemplified in the above-mentioned embodiment as including all of these, but a part of them can be omitted.

[0385] Note that the transmission section PROD_D5 can transmit the moving image that is not encoded, or can transmit encoded data that is encoded in a transmission-use encoding mode different from an encoding mode used for recording. In the latter case, it is preferable that an encoding section (not shown) that encodes the moving image in the transmission-use encoding mode be interposed between the decoding section PROD_D2 and the transmission section PROD_D5.

[0386] As such a reproduction apparatus PROD_D, for example, a DVD player, a BD player, an HDD player, or the like can be exemplified (in this case, the output terminal PROD_D4 connected to a television receiver or the like is the main supply destination of the moving image). Further, a television receiver (in this case, the display PROD_D3 is the main supply destination of the moving image), a digital signage (also called an electronic billboard, an electronic bulletin board, or the like, the display PROD_D3 or the transmission section PROD_D5 is the main supply destination of the moving image), a desktop PC (in this case, the output terminal PROD_D4 or the transmission section PROD_D5 is the main supply destination of the moving image), a laptop or tablet PC (in this case, the display PROD_D3 or the transmission section PROD_D5 is the main supply destination of the moving image), a smartphone (in this case, the display PROD_D3 or the transmission section PROD_D5 is the main supply destination of the moving image), or the like are also one example of such a reproduction apparatus PROD_D.

[0387] (Hardware implementation and software implementation)

[0388] Moreover, each block of the moving image decoding apparatus 31 and the moving image encoding apparatus 11 described above can be realized in hardware by a logic circuit formed on an integrated circuit (IC chip), or in software by a CPU (Central Processing Unit).

[0389] In the latter case, each apparatus described above is provided with a CPU that executes a command of a program that realizes each function, a ROM (Read Only Memory) that stores the program, a RAM (Random Access Memory) that expands the program, a storage device (recording medium) such as a memory that stores the program and various data, and the like. Then, an object of the embodiments of the present application is also achieved by supplying a recording medium that records a program code (an execution form program, an intermediate code program, a source program) of a control program of each apparatus described above, which is software that realizes the functions described above, in a manner that is readable by a computer (or a CPU, an MPU), to each apparatus described above, and the computer (or the CPU, the MPU) reads out the program code recorded on the recording medium and executes it.

[0390] As the recording medium described above, for example, a tape such as a magnetic tape, a cartridge tape, or the like; a disk such as a floppy disk (registered trademark) / hard disk, a CD-ROM (Compact Disc Read-Only Memory) / MO disk (Magneto-Optical disc) / MD (Mini Disc) / DVD (Digital Versatile Disc, registered trademark) / CD-R (CD Recordable) / Blu-ray Disc (registered trademark), or the like; a card such as an IC card (including a memory card) / optical card, or the like; a semiconductor memory such as a mask ROM / EPROM (Erasable Programmable Read-Only Memory) / EEPROM (Electrically Erasable and Programmable Read-Only Memory) / flash ROM, or the like; or a logic circuit such as a PLD (Programmable logic device) / FPGA (Field Programmable Gate Array), or the like can be used.

[0391] Further, each of the above-described devices can be configured to be connectable to a communication network, and supply the above-described program codes via the communication network. The communication network can transmit the program codes, and is not particularly limited. For example, the Internet, an intranet, an extranet, a LAN (Local Area Network), an ISDN (Integrated Services Digital Network), a VAN (Value-Added Network), a CATV (Community Antenna television / Cable Television) communication network, a virtual private network, a telephone line network, a mobile communication network, a satellite communication network, and the like can be used. Further, a transmission medium configuring the communication network is also a medium capable of transmitting the program codes, and is not limited to a particular configuration or type. For example, either in a wire such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, a power line, a wired TV line, a telephone line, an ADSL (Asymmetric Digital Subscriber Line) line, or the like, or in wireless such as IrDA (Infrared Data Association), infrared of a remote controller, Bluetooth (registered trademark), IEEE 802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance, registered trademark), a portable telephone network, a satellite line, a terrestrial digital broadcast network, and the like can be used. Note that the embodiments of the present application can be realized even in the form of a computer data signal of an embedded carrier that embodies the above-described program codes by electronic transmission.

[0392] The embodiments of the present application are not limited to the above-described embodiments, and various modifications can be made within the scope of the claims. That is, embodiments obtained by combining technical means appropriately modified within the scope of the claims are also included in the technical scope of the present application.

[0393] Industrial Applicability

[0394] Embodiments of the present application can be preferably applied to a moving image decoding apparatus that decodes encoded data obtained by encoding image data, and a moving image encoding apparatus that generates encoded data obtained by encoding image data. Furthermore, can be preferably applied to a data structure of encoded data generated by the moving image encoding apparatus and referred to by the moving image decoding apparatus.

[0395] (CROSS-REFERENCE TO RELATED APPLICATIONS)

[0396] This application claims the benefit of priority to Japanese Patent Application No. 2019-101179 filed on May 30, 2019, and is hereby incorporated by reference in its entirety into the present specification.

[0397] EXPLANATION OF REFERENCE NUMERALS

[0398] 31 Moving image decoding apparatus

[0399] 301 Entropy decoding section

[0400] 302 Parameter decoding section

[0401] 3020 Header decoding section

[0402] 303 Inter prediction parameter decoding section

[0403] 304 Intra prediction parameter decoding section

[0404] 308 Prediction image generating section

[0405] 309 Inter prediction image generating section

[0406] 310 Intra prediction image generating section

[0407] 311 Inverse quantization / inverse transform section

[0408] 312 Addition section

[0409] 11 Moving image encoding apparatus

[0410] 101 Prediction image generating section

[0411] 102 Subtraction section

[0412] 103 Transform / quantization section

[0413] 104 Entropy encoding section

[0414] 105 Inverse quantization / inverse transform section

[0415] 107 Loop filter

[0416] 110 Encoding parameter determining section

[0417] 111 parameter encoding section

[0418] 112 inter prediction parameter encoding section

[0419] 113 intra prediction parameter encoding section

[0420] 1110 header encoding section

[0421] 1111 CT information encoding section

[0422] 1112 CU encoding section (prediction mode encoding section)

[0423] 1114 TU encoding section

[0424] 3111 inverse quantization section

[0425] 3112 inverse transform section

[0426] 31121 secondary transform section

[0427] 31112 scaling section

[0428] 31123 kernel transform section

[0429] 10322 forward secondary transform section

[0430] 10323 forward kernel transform section

Claims

1. An image decoding apparatus that performs a transform on transform coefficients on a per transform unit basis, characterized by the image decoding apparatus comprising: a decoding section that decodes a secondary index indicating whether or not a secondary transform is used and a transform basis, and a multi transform selection (MTS) index that is a switching index of a transform basis used for selecting a core transform, derives a secondary transform matrix based on the secondary index and a transform size, a secondary transform section that, in a case where the secondary transform is used, applies the secondary transform using the secondary transform matrix to the transform coefficients, restores a modified transform coefficient; and a core transform section that applies a separable transform consisting of a vertical transform and a horizontal transform to the transform coefficients or the modified transform coefficients, the core transform section including an MTS setting section and an implicit MTS setting section, the MTS setting section derives a horizontal transform type variable and a vertical transform type variable in accordance with the MTS index in a case where explicit MTS is valid, the implicit MTS setting section sets an implicit transform to be off in a case where (i) a value of the secondary index is not equal to 0, an intra-sub partition mode is not used in MTS, and a sub-block transform is not used in MTS, (ii) the value of the secondary index is equal to 0, a width of a transform unit is 4 or more and 16 or less, and the width of the transform unit is equal to or less than a height of the transform unit, sets a value of the horizontal transform type variable to be equal to 1, and in other cases, sets the value of the horizontal transform type variable to be equal to 0, and (iii) the value of the secondary index is equal to 0, the height of the transform unit is 4 or more and 16 or less, and the height of the transform unit is equal to or less than the width of the transform unit, sets a value of the vertical transform type variable to be equal to 1, and in other cases, sets the value of the vertical transform type variable to be equal to 0, the core transform section performs the vertical transform corresponding to the vertical transform type variable and the horizontal transform corresponding to the horizontal transform type variable.

2. The image decoding apparatus according to claim 1, wherein the core transform section includes a horizontal transform section and a vertical transform section, the horizontal transform section applies the horizontal transform to the transform coefficients or the modified transform coefficients in accordance with the horizontal transform type variable, the vertical transform section applies the vertical transform to the transform coefficients or the modified transform coefficients in accordance with the vertical transform type variable, and the implicit MTS setting section sets the value of the horizontal transform type variable to be equal to 1 in a case where the value of the secondary index is equal to 0, the width of the transform unit is 4 or more and 16 or less, and the width of the transform unit is equal to or less than the height of the transform unit, and sets the value of the horizontal transform type variable to be equal to 0 in other cases, and sets the value of the vertical transform type variable to be equal to 1 in a case where the value of the secondary index is equal to 0, the height of the transform unit is 4 or more and 16 or less, and the height of the transform unit is equal to or less than the width of the transform unit, and sets the value of the vertical transform type variable to be equal to 0 in other cases.

Citation Information

Patent Citations

  • Fixation device and image formation device

    JP2019101179A