Image decoding device and recording medium

The image decoding apparatus addresses performance issues in combining secondary conversion with MTS by using a secondary index and multi-conversion selection index, resulting in improved efficiency and image quality.

JP2025094136AActive Publication Date: 2025-06-24SHARP KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025045944
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-05-30
Filing Date
2025-03-19
Publication Date
2025-06-24
Estimated Expiration
2040-05-29

AI Technical Summary

Technical Problem

Existing image encoding and decoding technologies face performance issues when combining secondary conversion with Multiple Transform Selection (MTS), particularly with implicit MTS, leading to suboptimal results.

Method used

An image decoding apparatus that includes a secondary index for indicating the use of inverse secondary conversion, a multi-conversion selection index for switching between conversion bases, and a core conversion unit that applies inverse core conversion using an MTS setting unit and an implicit MTS setting unit, thereby optimizing conversion processes.

Benefits of technology

The proposed solution enhances the performance of image decoding by effectively applying secondary conversion and MTS, improving the overall efficiency and quality of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094136000001_ABST
    Figure 2025094136000001_ABST
Patent Text Reader

Abstract

To provide an image decoding device and a related technology thereof which more suitably apply transform through Multiple Transform Selection (MTS) and secondary transform.SOLUTION: A moving image decoding device 31 includes: a transform unit decoding part 3024; and an inverse quantization / inverse transform part that includes a secondary transform section, a core transform section, and an inverse quantization section. The core transform section includes an MTS setting portion and an implicit MTS setting portion. When explicit MTS is enabled, the MTS setting portion derives horizontal transform type and vertical transform type. When a value of a secondary index is not equal to 0, the implicit MTS setting portion sets the horizontal transform type and the vertical transform type to 0. The core transform section performs inverse transform on the basis of the vertical transform type and also performs inverse transform on the basis of the horizontal transform type.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to an image decoding apparatus and an image encoding apparatus.

Background Art

[0002] In order to efficiently transmit or record an image, an image encoding apparatus that generates encoded data by encoding an image, and an image decoding apparatus that generates a decoded image by decoding the encoded data are used.

[0003] Specific image encoding methods include, for example, H.264 / AVC and HEVC (High-Efficiency Video Coding).

[0004] In such an image encoding method, an image (picture) constituting an image is managed by a hierarchical structure including a slice obtained by dividing the image, a coding tree unit (CTU) obtained by dividing the slice, a coding unit (sometimes called a coding unit (CU)) obtained by dividing the coding tree unit, and a transform unit (TU) obtained by dividing the coding unit, and is encoded / decoded for each CU.

[0005] Also, in such an image encoding method, usually, a predicted image is generated based on a local decoded image obtained by encoding / decoding an input image, and a prediction error (sometimes called a "difference image" or "residual image") obtained by subtracting the predicted image from the input image (original image) is encoded. Examples of the method for generating a predicted image include inter-picture prediction (inter prediction) and intra-picture prediction (intra prediction).

[0006] In recent years, Non-Patent Document 1 and Non-Patent Document 2 can be cited as image encoding and decoding technologies. Non-Patent Document 1 discloses a technology called Multiple Transform Selection (MTS) that switches the conversion matrix according to the explicit syntax or implicit block size in the encoded data. Non-Patent Document 2 discloses an image encoding device that, for each conversion unit, converts each coefficient of the prediction error after conversion by RST (Reduced Secondary Transform), that is, secondary transform, to derive conversion coefficients. Also, Non-Patent Document 2 discloses an image decoding device that inverse-transforms the conversion coefficients by secondary transform for each conversion unit.

Prior Art Documents

Non-Patent Documents

[0007]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0008] In technologies related to secondary conversion and secondary conversion such as Non-Patent Document 1, there is a problem that the performance when combining secondary conversion and conversion by MTS is not sufficient. In particular, there is a problem that the performance of implicit MTS becomes a loss when combined with secondary conversion.

[0009] An object of the present invention is to provide an image decoding apparatus and related technologies that can more suitably apply conversion by MTS and secondary conversion.

Means for Solving the Problems

[0010] An image decoding apparatus according to an aspect of the present invention is an image decoding apparatus that converts conversion coefficients for each conversion unit, and includes: (i) a secondary index indicating the presence or absence of use of inverse secondary conversion and a conversion basis from encoded data; and (ii) a conversion unit decoding unit that decodes a multi-conversion selection index that is a switching index for selecting a conversion basis for inverse core conversion; an inverse quantization unit that inverse quantizes quantization conversion coefficients and calculates conversion coefficients; and a secondary conversion unit that derives corrected conversion coefficients by applying the inverse secondary conversion to the conversion coefficients using a conversion matrix when the inverse secondary conversion is effective. A core conversion unit that applies the inverse core conversion including vertical conversion and horizontal conversion to the conversion coefficients or the corrected conversion coefficients, the core conversion unit includes an MTS setting unit and an implicit MTS setting unit, and the MTS setting unit is based on the multi-conversion selection index when explicit MTS is effective. The implicit MTS setting unit derives a horizontal conversion type and a vertical conversion type, and when the value of the secondary index is not equal to 0, sets the horizontal conversion type and the vertical conversion type to 0, and the core conversion unit performs inverse conversion based on the vertical conversion type and performs inverse conversion based on the horizontal conversion type.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Best Mode for Carrying Out the Invention

[0012] 〔Embodiment 1〕 Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0013] FIG. 1 is a schematic diagram showing the configuration of the image transmission system 1 according to the present embodiment.

[0014] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding an image to be encoded and decodes the transmitted encoded stream to display an image. The image transmission system 1 includes a moving image encoding device (image encoding device) 11, a network 21, a moving image decoding device (image decoding device) 31, and an image display device (image display device) 41.

[0015] An image T is input to the moving image encoding device 11.

[0016] The network 21 transmits the encoded stream Te generated by the moving image encoding device 11 to the moving image decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network 21 is not necessarily limited to a bidirectional communication network, and may be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Further, the network 21 may be replaced by a storage medium on which the encoded stream Te such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark) is recorded.

[0017] The moving image decoding device 31 decodes each of the encoded streams Te transmitted by the network 21 and generates one or more decoded images Td that have been decoded.

[0018] The image display device 41 displays all or part of the one or more decoded images Td generated by the moving image decoding device 31. The image display device 41 includes, for example, a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Examples of the form of the display include a stationary type, a mobile type, and an HMD. Further, when the moving image decoding device 31 has high processing power, an image with high image quality is displayed, and when it has only low processing power, an image that does not require high processing power and display power is displayed.

[0019] <Operator> The operators used in this specification are described below.

[0020] [[ID=18>> is a right bit shift, << is a left bit shift, & is a bitwise AND, | is a bitwise OR, |= is an OR assignment operator, and || indicates a logical OR.

[0021] x?y:z is a ternary operator that takes y when x is true (other than 0) and z when x is false (0).

[0022] Clip3(a, b, c) is a function that clips c to a value between a and b (inclusive). If c < a, it returns a; if c > b, it returns b; otherwise, it returns c (where a <= b).

[0023] abs(a) is a function that returns the absolute value of a.

[0024] Int(a) is a function that returns the integer value of a.

[0025] floor(a) is a function that returns the largest integer less than or equal to a.

[0026] ceil(a) is a function that returns the smallest integer greater than or equal to a.

[0027] a / d represents the division of a by d, rounded down to the nearest integer.

[0028] <Structure of the Encoded Stream Te>[[]]END]] Prior to the detailed description of the moving image encoding device 11 and the moving image decoding device 31 according to this embodiment, the data structure of the encoded stream Te generated by the moving image encoding device 11 and decoded by the moving image decoding device 31 will be described.

[0029] Figure 4 is a diagram showing the hierarchical structure of the data in the encoded stream Te. The encoded stream Te illustratively includes a sequence and a plurality of pictures that make up the sequence. Figure 4 shows, respectively, an encoded video sequence that defines a sequence SEQ, an encoded picture that defines a picture PICT, an encoded slice that defines a slice S, encoded slice data that defines the slice data, an encoded tree unit included in the encoded slice data, and an encoded unit included in the encoded tree unit.

[0030] (Encoded Video Sequence) In a coded video sequence, a set of data that the moving image decoding device 31 refers to in order to decode the sequence SEQ to be processed is defined. As shown in the coded video sequence of FIG. 4, the sequence SEQ includes a Video Parameter Set, a Sequence Parameter Set SPS, a Picture Parameter Set PPS, a picture PICT, and Supplemental Enhancement Information SEI.

[0031] The Video Parameter Set VPS defines a set of coding parameters common to a plurality of images, a plurality of layers included in the images, and a set of coding parameters related to each individual layer in an image composed of a plurality of layers.

[0032] The Sequence Parameter Set SPS defines a set of coding parameters that the moving image decoding device 31 refers to in order to decode the target sequence. For example, the width and height of a picture are defined. Note that there may be a plurality of SPSs. In that case, one of the plurality of SPSs is selected from the PPS.

[0033] The Picture Parameter Set PPS defines a set of coding parameters that the moving image decoding device 31 refers to in order to decode each picture in the target sequence. For example, it includes a reference value of the quantization width (pic_init_qp_minus26) used for decoding a picture, a flag (weighted_pred_flag) indicating the application of weighted prediction, and a scaling list (quantization matrix). Note that there may be a plurality of PPSs. In that case, one of the plurality of PPSs is selected from each picture in the target sequence.

[0034] (Coded Picture) In the case of a coded picture, a set of data that the moving picture decoding apparatus 31 refers to in order to decode the picture PICT to be processed is defined. As shown in the coded picture of FIG. 4, the picture PICT includes slices 0 to NS - 1 (NS is the total number of slices included in the picture PICT).

[0035] Incidentally, hereinafter, when it is not necessary to distinguish each of slices 0 to NS - 1, the subscript of the code may be omitted in the description. The same applies to the data included in the coded stream Te described below and other data with subscripts attached.

[0036] (Coded slice) In the case of a coded slice, a set of data that the moving picture decoding apparatus 31 refers to in order to decode the slice S to be processed is defined. As shown in the coded slice of FIG. 4, the slice includes a slice header and slice data.

[0037] The slice header includes a set of coding parameters that the moving picture decoding apparatus 31 refers to in order to determine the decoding method of the target slice. The slice type designation information (slice_type) for designating the slice type is an example of the coding parameters included in the slice header.

[0038] Examples of the slice types that can be specified by the slice type designation information include (1) an I slice that uses only intra prediction during coding, (2) a P slice that uses either unidirectional prediction or intra prediction during coding, and (3) a B slice that uses unidirectional prediction, bidirectional prediction, or intra prediction during coding. Note that the inter prediction is not limited to unidirectional prediction and bidirectional prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referring to P and B slices, it refers to a slice including a block that can use inter prediction.

[0039] Note that the slice header may include a reference (pic_parameter_set_id) to the picture parameter set PPS.

[0040] (Symbolized slice data) In the symbolized slice data, a set of data that the moving image decoding device 31 refers to in order to decode the slice data to be processed is defined. The slice data includes CTUs as shown in the symbolized slice header of FIG. 4. A CTU is a block of a fixed size (for example, 64x64) that constitutes a slice, and may also be called the largest coding unit (LCU).

[0041] (Coding tree unit) In the coding tree unit of FIG. 4, a set of data that the moving image decoding device 31 refers to in order to decode the CTU to be processed is defined. The CTU is divided into coding units CU, which are the basic units of the encoding process, by recursive quadtree splitting (QT (Quad Tree) splitting), binary tree splitting (BT (Binary Tree) splitting), or ternary tree splitting (TT (Ternary Tree) splitting). The combination of BT splitting and TT splitting is called multi-tree splitting (MT (Multi Tree) splitting). A node of the tree structure obtained by recursive quadtree splitting is called a coding node. Intermediate nodes of the quadtree, binary tree, and ternary tree are coding nodes, and the CTU itself is also defined as the topmost coding node.

[0042] CT, as CT information, includes a QT split flag (cu_split_flag) indicating whether to perform QT splitting, an MT split flag (split_mt_flag) indicating the presence or absence of MT splitting, an MT split direction (split_mt_dir) indicating the split direction of MT splitting, and an MT split type (split_mt_type) indicating the split type of MT splitting. The cu_split_flag, split_mt_flag, split_mt_dir, and split_mt_type are transmitted for each coding node.

[0043] When the cu_split_flag is 1, the coding node is divided into four coding nodes (QT in FIG. 5).

[0044] When cu_split_flag is 0 and split_mt_flag is 0, the encoding node is not split and has one CU as a node (no split in Fig. 5). The CU is the terminal node of the encoding node and is not split further. The CU is the basic unit of the encoding process.

[0045] When split_mt_flag is 1, the encoding node is MT split as follows. When split_mt_type is 0 and split_mt_dir is 1, the encoding node is horizontally split into two encoding nodes (BT (horizontal split) in Fig. 5), and when split_mt_dir is 0, the encoding node is vertically split into two encoding nodes (BT (vertical split) in Fig. 5). Also, when split_mt_type is 1 and split_mt_dir is 1, the encoding node is horizontally split into three encoding nodes (TT (horizontal split) in Fig. 5), and when split_mt_dir is 0, the encoding node is vertically split into three encoding nodes (TT (vertical split) in Fig. 5). These are shown in the CT information of Fig. 5.

[0046] Also, when the size of the CTU is 64x64 pixels, the size of the CU can be any of 64x64 pixels, 64x32 pixels, 32x64 pixels, 32x32 pixels, 64x16 pixels, 16x64 pixels, 32x16 pixels, 16x32 pixels, 16x16 pixels, 64x8 pixels, 8x64 pixels, 32x8 pixels, 8x32 pixels, 16x8 pixels, 8x16 pixels, 8x8 pixels, 64x4 pixels, 4x64 pixels, 32x4 pixels, 4x32 pixels, 16x4 pixels, 4x16 pixels, 8x4 pixels, 4x8 pixels, and 4x4 pixels.

[0047] (Encoding Unit) As shown in the encoding unit of Fig. 4, a set of data that the moving image decoding device 31 refers to in order to decode the encoding unit to be processed is defined. Specifically, the CU is composed of a CU header CUH, prediction parameters, transform parameters, quantized transform coefficients, etc. The prediction mode, etc. are defined in the CU header.

[0048] The prediction process may be performed in units of CUs or in units of sub-CUs obtained by further dividing a CU. When the sizes of a CU and a sub-CU are equal, there is one sub-CU in the CU. When a CU is larger than the size of a sub-CU, the CU is divided into sub-CUs. For example, when a CU is 8x8 and a sub-CU is 4x4, the CU is divided into four sub-CUs, which consist of two horizontal divisions and two vertical divisions.

[0049] There are two types of prediction (prediction mode), namely intra prediction and inter prediction. Intra prediction is prediction within the same picture, and inter prediction refers to the prediction process performed between different pictures (for example, between display times, between layer images).

[0050] The transformation and quantization process is performed in units of CUs, but the quantized transform coefficients may be entropy coded in units of sub-blocks such as 4x4.

[0051] (Prediction parameters) The predicted image is derived from the prediction parameters associated with the block. The prediction parameters include intra prediction parameters and inter prediction parameters.

[0052] Hereinafter, the intra prediction parameters will be described. The intra prediction parameters are composed of a luminance prediction mode IntraPredModeY and a chrominance prediction mode IntraPredModeC. FIG. 6 is a schematic diagram showing the types (mode numbers) of intra prediction modes. As shown in FIG. 6, there are, for example, 67 types (0 to 66) of intra prediction modes. For example, planar prediction (0), DC prediction (1), and angular prediction (2 to 66). Further, an LM mode (67 to 72) may be added for chrominance.

[0053] The syntax elements for deriving the intra prediction parameters include, for example, intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_remainder, etc.

[0054] (MPM) The intra_luma_mpm_flag is a flag indicating whether the IntraPredModeY of the target block matches the MPM (Most Probable Mode). The MPM is a prediction mode included in the MPM candidate list mpmCandList[]. The MPM candidate list is a list that stores candidates estimated to have a high probability of being applied to the target block from the intra prediction modes of adjacent blocks and predetermined intra prediction modes. When intra_luma_mpm_flag is 1, the IntraPredModeY of the target block is derived using the MPM candidate list and the index intra_luma_mpm_idx.

[0055] IntraPredModeY = mpmCandList[intra_luma_mpm_idx] (REM) When intra_luma_mpm_flag is 0, the intra prediction mode is selected from the remaining modes RemIntraPredMode obtained by excluding the intra prediction modes included in the MPM candidate list from all intra prediction modes. The intra prediction modes that can be selected as RemIntraPredMode are called "non-MPM" or "REM". RemIntraPredMode is derived using intra_luma_mpm_remainder.

[0056] (Configuration of Video Decoder) The configuration of the video decoder 31 (Fig. 7) according to this embodiment will be described.

[0057] The video decoder 31 includes an entropy decoder 301, a parameter decoder (predicted image decoder) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a predicted image generation unit (predicted image generation device) 308, an inverse quantization and inverse transformation unit 311, and an addition unit 312. Note that, in accordance with the video encoder 11 described later, there is also a configuration in which the video decoder 31 does not include the loop filter 305.

[0058] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit). The CU decoding unit 3022 includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS, and a slice header (slice information) from the encoded data. The CT information decoding unit 3021 decodes CT from the encoded data. The CU decoding unit 3022 decodes CU from the encoded data. The TU decoding unit 3024 decodes QP update information (quantization correction value) and quantized prediction error (residual_coding) from the encoded data when the TU contains a prediction error.

[0059] Figure 23 is a block diagram showing the relationship between the TU decoding unit 3024 and the inverse transformation unit 3112. The stIdx decoding unit 131 of the TU decoding unit 3024 decodes a value stIdx indicating the presence or absence of use of secondary transformation and the transformation basis from the encoded data, and outputs it to the secondary transformation unit 31121. The mts_idx decoding unit 132 of the TU decoding unit 3024 decodes a value mts_idx indicating the transformation matrix of MTS from the encoded data, and outputs it to the core transformation unit 31123. Specifically, the TU decoding unit 3024 decodes stIdx when the width and height of the CU are 4 or more, the prediction mode is the intra mode, and the number numSigCoeff of transform coefficients in the CU is greater than a predetermined number THSt (for example, 2 in SINGLE_TREE and 1 otherwise). When stIdx is 0, no secondary transformation is applied. When stIdx is 1, it indicates one of the transformations of a set (pair) of secondary transformation matrices. When stIdx is 2, it indicates the other transformation of the above pair. Also, the secondary transformation matrix secTransMatrix may be selected according to not only the value of stIdx but also the intra prediction mode and the size of the transformation.

[0060] The parameter decoding unit 302 includes an inter-prediction parameter decoding unit 303 and an intra-prediction parameter decoding unit 304 (not shown). The prediction image generation unit 308 includes an inter-prediction image generation unit 309 and an intra-prediction image generation unit 310.

[0061] In the following, an example using CTUs and CUs as processing units will be described. However, the present invention is not limited to this example, and processing may be performed in units of sub-CUs. Alternatively, CTUs and CUs may be read as blocks, and sub-CUs may be read as sub-blocks, and processing may be performed in units of blocks or sub-blocks.

[0062] The entropy decoding unit 301 performs entropy decoding on the encoded stream Te input from the outside to separate and decode individual codes (syntax elements). For entropy encoding, there are a method of variably encoding syntax elements using a context (probability model) adaptively selected according to the type of syntax element and the surrounding situation, and a method of variably encoding syntax elements using a predetermined table or calculation formula. The former CABAC (Context Adaptive Binary Arithmetic Coding) stores the probability model updated for each encoded or decoded picture (slice) in memory. Then, as the initial state of the context of the P picture or B picture, a probability model of a picture using the quantization parameter of the same slice type and the same slice level is set from the probability models stored in the memory. This initial state is used for encoding and decoding processes. The separated codes include prediction information for generating a prediction image, a prediction error for generating a difference image, and the like.

[0063] The entropy decoding unit 301 outputs the separated codes to the parameter decoding unit 302. The separated codes are, for example, the prediction mode predMode. The control of which code to decode is performed based on the instruction of the parameter decoding unit 302.

[0064] (Basic Flow) FIG. 8 is a flowchart for explaining the schematic operation of the moving image decoding apparatus 31.

[0065] (S1100: Parameter set information decoding) The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS from the encoded data.

[0066] (S1200: Slice information decoding) The header decoding unit 3020 decodes a slice header (slice information) from the encoded data.

[0067] Hereinafter, the moving image decoding apparatus 31 derives a decoded image of each CTU by repeating the processes from S1300 to S5000 for each CTU included in the target picture.

[0068] (S1300: CTU information decoding) The CT information decoding unit 3021 decodes a CTU from the encoded data.

[0069] (S1400: CT information decoding) The CT information decoding unit 3021 decodes a CT from the encoded data.

[0070] (S1500: CU decoding) The CU decoding unit 3022 performs S1510 and S1520 to decode a CU from the encoded data.

[0071] (S1510: CU information decoding) The CU decoding unit 3022 decodes CU information, prediction information, TU split flag split_transform_flag, CU residual flags cbf_cb, cbf_cr, cbf_luma, etc. from the encoded data.

[0072] (S1520: TU information decoding) When the TU contains a prediction error, the TU decoding unit 3024 decodes QP update information (quantization correction value) and quantized prediction error (residual_coding) from the encoded data. Note that the QP update information is a difference value from the quantization parameter prediction value qPpred, which is a predicted value of the quantization parameter QP.

[0073] (S2000: Prediction Image Generation) The prediction image generation unit 308 generates a prediction image for each block included in the target CU based on the prediction information.

[0074] (S3000: Inverse Quantization and Inverse Transformation) The inverse quantization and inverse transformation unit 311 performs inverse quantization and inverse transformation processing for each TU included in the target CU.

[0075] (S4000: Decoded Image Generation) The addition unit 312 generates a decoded image of the target CU by adding the prediction image supplied from the prediction image generation unit 308 and the prediction error supplied from the inverse quantization and inverse transformation unit 311.

[0076] (S5000: Loop Filter) The loop filter 305 applies loop filters such as a deblocking filter, SAO, and ALF to the decoded image to generate a decoded image.

[0077] Also, the parameter decoding unit 302 is configured to include an inter prediction parameter decoding unit 303 and an intra prediction parameter decoding unit 304 (not shown). The prediction image generation unit 308 is configured to include an inter prediction image generation unit 309 and an intra prediction image generation unit 310 (not shown).

[0078] (Configuration of Intra Prediction Parameter Decoding Unit 304) The intra prediction parameter decoding unit 304 decodes intra prediction parameters, for example, an intra prediction mode IntraPredMode, by referring to the prediction parameters stored in the prediction parameter memory 307 based on the code input from the entropy decoding unit 301. The intra prediction parameter decoding unit 304 outputs the decoded intra prediction parameters to the prediction image generation unit 308 and stores them in the prediction parameter memory 307. The intra prediction parameter decoding unit 304 may derive different intra prediction modes for luminance and color difference.

[0079] FIG. 9 is a schematic diagram showing the configuration of the intra prediction parameter decoding section 304 of the parameter decoding section 302. As shown in FIG. 9, the intra prediction parameter decoding section 304 includes a parameter decoding control section 3041, a luminance intra prediction parameter decoding section 3042, and a chrominance intra prediction parameter decoding section 3043.

[0080] The parameter decoding control section 3041 instructs the entropy decoding section 301 to decode the syntax elements, and receives the syntax elements from the entropy decoding section 301. When intra_luma_mpm_flag therein is 1, the parameter decoding control section 3041 outputs intra_luma_mpm_idx to the MPM parameter decoding section 30422 in the luminance intra prediction parameter decoding section 3042. When intra_luma_mpm_flag is 0, the parameter decoding control section 3041 outputs intra_luma_mpm_remainder to the non-MPM parameter decoding section 30423 of the luminance intra prediction parameter decoding section 3042. Also, the parameter decoding control section 3041 outputs the syntax elements of the chrominance intra prediction parameters to the chrominance intra prediction parameter decoding section 3043.

[0081] The luminance intra prediction parameter decoding section 3042 includes an MPM candidate list derivation section 30421, an MPM parameter decoding section 30422, and a non-MPM parameter decoding section 30423 (decoding section, derivation section).

[0082] The MPM parameter decoding section 30422 refers to mpmCandList[] derived by the MPM candidate list derivation section 30421 and intra_luma_mpm_idx, derives IntraPredModeY, and outputs it to the intra prediction image generation section 310.

[0083] The non-MPM parameter decoding section 30423 derives RemIntraPredMode from mpmCandList[] and intra_luma_mpm_remainder, and outputs IntraPredModeY to the intra prediction image generation section 310.

[0084] The chrominance intra prediction parameter decoding unit 3043 derives IntraPredModeC from the syntax elements of the chrominance intra prediction parameter and outputs it to the intra prediction image generation unit 310.

[0085] The luma intra prediction parameter decoding unit 3042 may further decode a flag intra_subpartitions_mode_flag indicating whether to perform intra sub-division for performing intra prediction by further dividing a CU into smaller sub-blocks. When intra_subpartitions_mode_flag is other than 0, intra_subpartitions_split_flag is further decoded. The intra sub-division mode is derived by the following formula.

[0086] IntraSubPartSplitType = (intra_subpartitions_mode_flag == 0)? 0 : 1 + intra_subpartitions_split_flag When IntraSubPartSplitType is 0 (ISP_NO_SPLIT), intra prediction is performed without further dividing the CU. When IntraSubPartSplitType is 1 (ISP_HOR_SPLIT: horizontal split), the CU is divided into 2 to 4 sub-blocks in the vertical direction, and intra prediction, transform coefficient decoding, inverse quantization, and inverse transformation are performed for each sub-block. When IntraSubPartSplitType is 2 (ISP_VER_SPLIT: vertical split), the CU is divided into 2 to 4 sub-blocks in the horizontal direction, and intra prediction, transform coefficient decoding, inverse quantization, and inverse transformation are performed for each sub-block. The number of sub-blocks divided NumIntraSubPart is derived by the following formula.

[0087] NumIntraSubPart = (cbWidth == 4 && cbHeight == 8) || (cbWidth == 8 && cbHeight == 4)? 2 : 4 The width nW and height nH of the sub-block, and the number of divisions numPartsX and numPartY in the horizontal and vertical directions are derived as follows.

[0088] nW = (IntraSubPartSplitType == ISP_VER_SPLIT?) nTbW / NumIntraSubPart : nTbW nH = (IntraSubPartSplitType == ISP_HOR_SPLIT?) nTbH / NumIntraSubPart : nTbH numPartsX = (IntraSubPartSplitType == ISP_VER_SPLIT?) NumIntraSubPart : 1 numPartsY = (IntraSubPartSplitType == ISP_HOR_SPLIT?) NumIntraSubPart : 1 Here, nTbW and nTbH are the width and height of the CU (or TU).

[0089] The loop filter 305 is a filter provided within the encoding loop, which removes block distortion and ringing distortion and improves the image quality. The loop filter 305 applies filters such as a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF) to the decoded image of the CU generated by the adder 312.

[0090] The reference picture memory 306 stores the decoded image of the CU generated by the adder 312 at a predetermined position for each target picture and target CU.

[0091] The prediction parameter memory 307 stores prediction parameters at a predetermined position for each CTU or CU to be decoded. Specifically, the prediction parameter memory 307 stores the parameters decoded by the parameter decoder 302 and predMode and the like separated by the entropy decoder 301.

[0092] The prediction image generation unit 308 is input with predMode, prediction parameters, etc. Further, the prediction image generation unit 308 reads a reference picture from the reference picture memory 306. The prediction image generation unit 308 generates a prediction image of a block or a sub-block in the prediction mode indicated by predMode, using the prediction parameters and the read reference picture (reference picture block). Here, the reference picture block is a set of pixels on the reference picture (usually a rectangle, so it is called a block), and is the area referred to for generating the prediction image.

[0093] (Intra prediction image generation unit 310) When predMode indicates the intra prediction mode, the intra prediction image generation unit 310 performs intra prediction using the intra prediction parameters input from the intra prediction parameter decoding unit 304 and the reference pixels read from the reference picture memory 306.

[0094] Specifically, the intra prediction image generation unit 310 reads adjacent blocks within a predetermined range from the target block on the target picture from the reference picture memory 306. The predetermined range refers to the adjacent blocks to the left, upper left, upper, and upper right of the target block, and the area referred to varies depending on the intra prediction mode.

[0095] The intra prediction image generation unit 310 generates a prediction image of the target block with reference to the read decoded pixel values and the prediction mode indicated by IntraPredMode. The intra prediction image generation unit 310 outputs the generated prediction image of the block to the addition unit 312.

[0096] The generation of a predicted image based on the intra prediction mode will be described below. In planar prediction, DC prediction, and angular prediction, a decoded peripheral area adjacent (proximate) to the block to be predicted is set as the reference area R. Then, a predicted image is generated by extrapolating the pixels on the reference area R in a specific direction. For example, the reference area R may be set as an L-shaped area (e.g., the area indicated by the hatched round-marked pixels in Example 1 of the reference area in FIG. 10) including the left and top (or, further, the upper left, upper right, and lower left) of the block to be predicted.

[0097] (Details of the Predicted Image Generation Unit) Next, the details of the configuration of the intra predicted image generation unit 310 will be described with reference to FIG. 11. The intra predicted image generation unit 310 includes a block to be predicted setting unit 3101, an unfiltered reference image setting unit 3102 (first reference image setting unit), a filtered reference image setting unit 3103 (second reference image setting unit), an intra prediction unit 3104, and a predicted image correction unit 3105 (predicted image correction unit, filter switching unit, weight coefficient changing unit).

[0098] Based on each reference pixel (unfiltered reference image) on the reference area R, the filtered reference image generated by applying a reference pixel filter (first filter), and the intra prediction mode, the intra prediction unit 3104 generates a temporary predicted image (predicted image before correction) of the block to be predicted and outputs it to the predicted image correction unit 3105. The predicted image correction unit 3105 corrects the temporary predicted image according to the intra prediction mode, generates a predicted image (predicted image after correction), and outputs it.

[0099] Hereinafter, each unit included in the intra predicted image generation unit 310 will be described.

[0100] (Block to be Predicted Setting Unit 3101) The block to be predicted setting unit 3101 sets the target CU as the block to be predicted and outputs information regarding the block to be predicted (block to be predicted information). The block to be predicted information includes at least the size, position, and an index indicating whether it is luminance or color difference of the block to be predicted.

[0101] (Non-filtered reference image setting unit 3102) The non-filtered reference image setting unit 3102 sets the adjacent peripheral region of the prediction target block as the reference region R based on the size and position of the prediction target block. Subsequently, each decoded pixel value at the corresponding position on the reference picture memory 306 is set to each pixel value (non-filtered reference image, boundary pixel) within the reference region R. The line r[x][-1] of decoded pixels adjacent to the upper side of the prediction target block and the column r[-1][y] of decoded pixels adjacent to the left side of the prediction target block shown in Example 1 of the reference region in FIG. 10 are the non-filtered reference images.

[0102] (Filtered reference image setting unit 3103) The filtered reference image setting unit 3103 applies a reference pixel filter (first filter) to the non-filtered reference image according to the intra prediction mode to derive the filtered reference image s[x][y] at each position (x, y) on the reference region R. Specifically, a low-pass filter is applied to the position (x, y) and its surrounding non-filtered reference image to derive the filtered reference image (Example 2 of the reference region in FIG. 10). Note that it is not necessarily required to apply a low-pass filter to all intra prediction modes, and a low-pass filter may be applied to some intra prediction modes. The filter applied to the non-filtered reference image on the reference region R in the filtered reference image setting unit 3103 is referred to as a "reference pixel filter (first filter)", while the filter that corrects the temporary prediction image in the prediction image correction unit 3105 described later is referred to as a "boundary filter (second filter)".

[0103] (Configuration of intra prediction unit 3104) The intra prediction unit 3104 generates a temporary prediction image (temporary prediction pixel values, pre-correction prediction image) of the block to be predicted based on the intra prediction mode, the unfiltered reference image, and the filtered reference pixel values, and outputs it to the prediction image correction unit 3105. The intra prediction unit 3104 includes an internal Planar prediction unit 31041, a DC prediction unit 31042, an Angular prediction unit 31043, and an LM prediction unit 31044. The intra prediction unit 3104 selects a specific prediction unit according to the intra prediction mode and inputs the unfiltered reference image and the filtered reference image. The relationship between the intra prediction mode and the corresponding prediction unit is as follows. · Planar prediction ··· Planar prediction unit 31041 · DC prediction ··· DC prediction unit 31042 · Angular prediction ··· Angular prediction unit 31043 · LM prediction ··· LM prediction unit 31044 (Planar prediction) The Planar prediction unit 31041 linearly adds a plurality of filtered reference images according to the distance between the pixel position to be predicted and the reference pixel position to generate a temporary prediction image, and outputs it to the prediction image correction unit 3105.

[0104] (DC prediction) The DC prediction unit 31042 derives a DC prediction value corresponding to the average value of the filtered reference image s[x][y], and outputs a temporary prediction image q[x][y] with the DC prediction value as the pixel value.

[0105] (Angular prediction) The Angular prediction unit 31043 generates a temporary prediction image q[x][y] using the filtered reference image s[x][y] in the prediction direction (reference direction) indicated by the intra prediction mode, and outputs it to the prediction image correction unit 3105.

[0106] (LM prediction) The LM prediction unit 31044 predicts the pixel values of the color difference based on the pixel values of the luminance. Specifically, based on the decoded luminance image, a linear model is used to generate a predicted image of the color difference image (Cb, Cr). One of the LM predictions, the CCLM (Cross-Component Linear Model prediction), is a prediction method that uses a linear model to predict the color difference from the luminance for one block.

[0107] (Configuration of the predicted image correction unit 3105) The predicted image correction unit 3105 corrects the temporary predicted image output from the intra prediction unit 3104 according to the intra prediction mode. Specifically, for each pixel of the temporary predicted image, the predicted image correction unit 3105 weights and adds (weighted average) the unfiltered reference image and the temporary predicted image according to the distance between the reference region R and the target predicted pixel, thereby deriving a predicted image (corrected predicted image) Pred obtained by correcting the temporary predicted image. Note that in some intra prediction modes, the predicted image correction unit 3105 may not correct the temporary predicted image and use the output of the intra prediction unit 3104 as the predicted image as it is.

[0108] (Inverse quantization and inverse transformation unit 311) The inverse quantization and inverse transformation unit 311 inverse quantizes the quantized transformation coefficients qd[ ][ ] input from the entropy decoding unit 301 to obtain the transformation coefficients d[ ][ ]. These quantized transformation coefficients qd[ ][ ] are coefficients obtained by performing frequency transformation such as DCT (Discrete Cosine Transform) and DST (Discrete Sine Transform) on the prediction error and then quantizing it in the encoding process. The inverse quantization and inverse transformation unit 311 performs inverse frequency transformation such as inverse DCT and inverse DST on the obtained transformation coefficients to calculate the prediction error. The inverse quantization and inverse transformation unit 311 outputs the prediction error to the addition unit 312.

[0109] Next, a configuration example of the inverse quantization and inverse transformation unit 311 will be described with reference to FIG. 12. FIG. 12 is a functional block diagram showing a configuration example of the inverse quantization and inverse transformation unit 311. As shown in FIG. 12, the inverse quantization and inverse transformation unit 311 includes an inverse quantization unit 3111 and a transformation unit 3112. The inverse quantization unit 3111 inverse quantizes the quantized transformation coefficients qd[ ][ ] decoded in the TU decoder 3024 to derive the transformation coefficients d[ ][ ]. The inverse quantization unit 3111 outputs the derived transformation coefficients d[ ][ ] to the transformation unit 3112.

[0110] The transformation unit 3112 inverse-transforms the received transformation coefficients d[ ][ ] for each transformation unit TU to restore the prediction error r[ ][ ]. The transformation unit 3112 outputs the restored prediction error r[ ][ ] to the addition unit 312.

[0111] In this specification, the process of converting the differential image in the image encoding device into transformation coefficients is called forward transformation, and the process of converting from the transformation coefficients in the image decoding device into a differential image is called transformation. However, they may also be simply called transformation and inverse transformation, respectively. Note that there is no difference in the processes other than the values of the transformation matrix serving as the transformation basis between the forward transformation (transformation) and the transformation (inverse transformation). Therefore, in the following description, the term "inverse transformation" may be used instead of "transformation" for the transformation process in the transformation unit 3112.

[0112] The transformation unit 3112 includes a secondary transformation unit (second transformation unit) 31121 and a core transformation unit (first transformation unit) 31123.

[0113] The TU decoding unit 3024 may further divide the CU into a plurality of sub-blocks, and decode only one sub-block among the plurality of sub-blocks for the inverse quantization and inverse transform of the transform coefficients, and decode the sub-block transform flag cu_sbt_flag. When cu_sbt_flag is 1, the flag cu_sbt_quad_flag indicating whether to further divide into four sub-blocks may be decoded. When cu_sbt_quad_flag is 0, the number of sub-blocks is 2. When cu_sbt_quad_flag is 1, the number of sub-blocks is 4. Also, the cu_sbt_horizontal_flag indicating whether to divide horizontally or vertically is decoded. Also, the cu_sbt_pos_flag indicating which sub-block contains the transform coefficients is decoded.

[0114] (Scaling unit 31112) The scaling unit 31112 scales the transform coefficients decoded by the TU decoding unit using the weight of the coefficient unit.

[0115] When the transform skip is valid (transform_skip == 1), the scaling unit 31112 performs scaling according to the following formula.

[0116] r[x][y] = d[x][y] << tsShift Here, tsShift = 5 + ((log2(nTbW) + log2(nTbH)) / 2).

[0117] In other cases, the quantization matrix m[x][y] and the scaling factor ls[x][y] are derived according to the following formula.

[0118] ls[x][y] = (m[x][y] * levelScale[(qP+1)%6]) << (qP / 6) Or it may be derived according to the following formula.

[0119] ls[x][y] = (m[x][y] * levelScale[qP%6]) << (qP / 6) Here, levelScale[] = { 40, 45, 51, 57, 64, 72}.

[0120] Note that the value of the quantization matrix m[x][y] may be decoded from the encoded data, or m[x][y] = 16 may be used as uniform quantization.

[0121] The scaling unit 31112 derives dnc[][] from the product of the scaling factor ls[][] and the decoded transform coefficient TransCoeffLevel, and performs inverse quantization.

[0122] dnc[x][y] = ( TransCoeffLevel[xTbY][yTbY][cIdx][x][y] * ls[x][y] * rectNorm +bdOffset ) >> bdShift Finally, the scaling unit 31112 clips the inverse quantized transform coefficient to derive d[x][y].

[0123] d[x][y] = Clip3( CoeffMin, CoeffMax, dnc[x][y] ) d[x][y] is transmitted to the core transform unit 31123 or the secondary transform unit 31121. The secondary transform unit (the second transform unit) 31121 applies a secondary transform to the transform coefficient d[ ][ ] after inverse quantization and before core transform.

[0124] (Secondary transform and core transform) The secondary conversion unit 31121 restores the modified conversion coefficients (conversion coefficients after conversion by the second conversion unit) d[ ][ ] by applying a conversion using a conversion matrix to some or all of the conversion coefficients d[ ][ ] received from the inverse quantization unit 3111. The secondary conversion unit 31121 applies secondary conversion to the conversion coefficients d[ ][ ] in a predetermined unit for each conversion unit TU. The secondary conversion is applied only in the intra CU, and the conversion basis is determined with reference to the intra prediction mode IntraPredMode. The selection of the conversion basis will be described later. The secondary conversion unit 31121 outputs the restored modified conversion coefficients d[ ][ ] to the core conversion unit 31123.

[0125] The core conversion unit 31123 obtains the conversion coefficients d[ ][ ] or the modified conversion coefficients d[ ][ ] restored by the secondary conversion unit 31121, performs conversion, and derives the prediction error r[][]. The core conversion unit 31123 outputs the prediction error r[][ ] to the addition unit 312.

[0126] (Secondary Conversion) In the moving image encoding device 11, further conversion (forward secondary conversion) is applied to the conversion coefficients after core conversion (such as DCT2 and DST7) of the difference image to remove the remaining correlation in the conversion coefficients and concentrate the energy on some of the conversion coefficients. The forward conversion unit 1032 included in the conversion and quantization unit 103 of the moving image encoding device 11 and the inverse conversion unit 152 included in the inverse conversion and inverse quantization unit 105 are shown in FIG. 19. In the moving image decoding device 3, conversely, secondary conversion is applied to the conversion coefficients of some or all regions of the decoded TU, and core conversion (such as DCT2 and DST7) is applied to the conversion coefficients after secondary conversion.

[0127] In the secondary transformation, the following processing is performed according to the size of the TU and the intra prediction mode. Hereinafter, the processing of the secondary transformation will be described in order. FIG. 13 is a diagram for explaining the secondary transformation. In the figure, for an 8x8 TU, in the process of S2, the transformation coefficients d[][] of a 4x4 region are stored in a one-dimensional array u[] of nonZeroSize, and in the process of S3, they are transformed from the one-dimensional array u[] to a one-dimensional array v[], and finally in the process of S4, they are stored again in d[][].

[0128] FIG. 24 is a flowchart showing the processing of the secondary transformation.

[0129] (S1: Setting the transformation size and the input / output size) The secondary transformation unit 31121 derives the size of the secondary transformation (4x4 or 8x8), the number of output transformation coefficients (nStOutSize), the number of transformation coefficients to be applied (the number of input transformation coefficients) nonZeroSize, and the number of sub-blocks (numStX, numStY) to which the secondary transformation is applied according to the size of the TU (width nTbW, height nTbH). The sizes of the 4x4 and 8x8 secondary transformations are indicated by nStSize = 4 and 8. Also, the sizes of the 4x4 and 8x8 secondary transformations may be referred to as RST4x4 and RST8x8, respectively.

[0130] When the TU is of a predetermined size or more, the secondary transformation unit 31121 outputs 48 transformation coefficients by the RST8x8 secondary transformation. Otherwise, it outputs 16 transformation coefficients by the RST4x4 secondary transformation. When the TU is 4x4, 16 transformation coefficients are derived from 8 transformation coefficients using RST4x4, and when the TU is 8x8, 48 transformation coefficients are derived from 8 transformation coefficients using RST8x8. Otherwise, 16 or 48 transformation coefficients are output according to the size of the TU from 16 transformation coefficients.

[0131] When both nTbW and nTbH are 8 or more, log2StSize = 3, nStOutSize = 48 In other cases, log2StSize = 2, nStOutSize = 16 nStSize = 1 << log2StSize When both nTbW and nTbH are 4, or in the case of 8x8, nonZeroSize = 8 In other cases, nonZeroSize = 16 numStX = (nTbH == 4 && nTbW > 8)? 2 : 1 numStY = (nTbW == 4 && nTbH > 8)? 2 : 1 (S2: Rearrange into a one-dimensional signal) The secondary conversion unit 31121 rearranges and processes some of the transform coefficients d[][] of the TU into a one-dimensional array u[] once. Specifically, in the secondary conversion, from the two-dimensional transform coefficients d[][] of the target TU, the transform coefficients for x = 0..nonZeroSize - 1 are referred to derive u[]. xC and yC are positions on the TU and are derived from the array DiagScanOrder indicating the scan order and the position x of the transform coefficients in the sub-block.

[0132] for (x = 0; x < nonZeroSize; x++) { xC = (xSbIdx << log2StSize) + DiagScanOrder[log2StSize][log2StSize][x][0] yC = (ySbIdx << log2StSize) + DiagScanOrder[log2StSize][log2StSize][x][1] u[x] = d[xC][yC] } (S3: Apply the conversion process) The secondary conversion unit 31121 performs a conversion on u[] (vector F') with a length of nonZeroSize using the first type of conversion basis (matrix) T and derives a one-dimensional array v'[] (vector V') with a length of nStOutSize as the output.

[0133] This conversion can be represented by the following equation in matrix operations.

[0134] V' = T × F' Here, when the conversion size is 4x4 (RST4x4), the conversion basis is called the first type of conversion basis T1. When the conversion size is 8x8 (RST8x8), the conversion basis is called the second type of conversion basis T2. T1 is a 16×16 (16 rows and 16 columns) matrix, and the conversion derives a 16×1 (16 rows and 1 column) vector V', that is, a one-dimensional array v'[] of length 16, as the product of a 16x16 matrix T and a 16x1 (16 rows and 1 column) vector F'. T2 is a 48×16 (48 rows and 16 columns) matrix, and the conversion derives a 48×1 (48 rows and 1 column, length 48) vector V', that is, a one-dimensional array v'[] of length 48, as the product of a 48x16 matrix T and a 16x1 vector F'.

[0135] Specifically, the secondary conversion unit 31121 derives a set number (stTrSetId) of secondary conversions derived from the intra prediction mode IntraPredMode, an stIdx indicating the conversion basis of the secondary conversion decoded from the encoded data, and a secondary conversion size nStSize (nTrS) to obtain a corresponding conversion matrix secTranMatrix[][](conversion basis T1 or T2). Further, the secondary conversion unit 31121 performs a sum-of-products operation of the conversion matrix and the one-dimensional array u[] as shown in the following formula. v'[i] = Clip3( CoeffMin, CoeffMax, ΣsecTransMatrix[j][i]*u[j]) Here, Σ is the sum from j = 0..nonZeroSize - 1. Also, i is processed for 0..nStSize - 1. CoeffMin and CoeffMax indicate the value range of the conversion coefficients.

[0136] (S4: Two-dimensional arrangement of the one-dimensional signal after conversion processing) The secondary conversion unit 31121 arranges the coefficients v'[] of the converted one-dimensional array at a predetermined position within TU again.

[0137] In process S4, the secondary conversion unit 31121 places the coefficient v'[] of length nStOutSize obtained by the above-described process S3 in the upper left region of the array d[][] of conversion coefficients.

[0138] The secondary conversion unit 31121 performs the following process for x = 0..nStSize - 1, y = 0..nStSize - 1. Specifically, when IntraPredMode <= 34 or INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM, the secondary conversion unit 31121 applies the following formula.

[0139] d[(xSbIdx<<log2StSize)+x][(ySbIdx<<log2StSize)+y] = (y < 4)? v[x+(y<<log2StSize)] : ((x < 4)? v[32 + x + ((y - 4) << 2)] : d[(xSbIdx<<log2StSize)+x][(ySbIdx<<log2StSize)+y]) In other cases, the secondary conversion unit 31121 applies the following formula. d[(xSbIdx<<log2StSize)+x][(ySbIdx<<log2StSize)+y] = (y < 4)? v[y+(x<<log2StSize)] : ((x < 4)? v[32 + (y - 4) + (x << 2)] : d[(xSbIdx<<log2StSize)+x][(ySbIdx<<log2StSize)+y]) (Core conversion unit 31123) <Core conversion> The method of conversion can be adaptively switched, and the conversion that can be switched by an explicit flag, index, prediction mode, etc. is called conversion (first conversion, core conversion). The conversion (core conversion) used in the core conversion is a separable conversion composed of a vertical conversion and a horizontal conversion. Also, a conversion that separates a two-dimensional signal into a horizontal direction and a vertical direction may be defined as the first conversion. Further, in an image decoding apparatus, a conversion applied after a second conversion (secondary conversion) may be defined as the first conversion. The conversion basis (conversion matrix) of the core conversion is DCT2, DST7, DCT8. In the core conversion, the conversion basis is switched independently for each of the vertical conversion and the horizontal conversion. Note that the selectable conversion is not limited to the above, and another conversion (conversion basis) can also be used. Note that DCT2, DST7, DCT8, DST1, and DCT5 may be represented as DCT-II, DST-VII, DCT-VIII, DST-I, and DCT-V, respectively. Also, there may be a conversion skip as a mode for explicitly skipping the core conversion.

[0140] The core conversion includes an explicit MTS and an implicit MTS. In the case of the explicit MTS, the mts_idx is decoded from the encoded data, and the conversion matrix is switched. In the case of the implicit MTS, the mts_idx is derived according to the intra prediction mode and the block size.

[0141] Note that in this embodiment, an example of decoding the mts_idx in CU units or TU units is described, but the unit of decoding (switching) is not limited to this.

[0142] The mts_idx is a switching index for selecting the conversion basis of the core conversion. The mts_idx has a value of any one of 0, 1, 2, 3, 4, and derives the horizontal conversion type trTypeHor and the vertical conversion type trTypeVer.

[0143] The core conversion described above will be specifically described with reference to FIG. 18. The core conversion unit 1521 in FIG. 18 is an example of the core conversion unit 31123 in FIG. 12 and the core conversion unit 1521 in FIG. 19. The core conversion unit 1521 in FIG. 18 includes an MTS setting unit 15211 that sets the type of conversion to be used from a plurality of conversion bases, a coefficient conversion processing unit 15212 that calculates a prediction residual r[ ][ ] from a (corrected) conversion coefficient d[ ][ ] using the derived conversion, and a matrix conversion processing unit 15213 that performs an actual conversion. When the secondary conversion is not performed, the corrected conversion coefficient is equal to the conversion coefficient. When the secondary conversion is performed, the corrected conversion coefficient takes a value different from the conversion coefficient. The MTS setting unit 15211 includes an MTS setting unit 152111 that determines a method for deriving the index mts_ids of the conversion to be used and an implicit MTS setting unit 152112 that implicitly derives mts_idx.

[0144] The MTS setting unit 152111 selects whether to perform explicit MTS, implicit MTS, or no MTS.

[0145] When explicit MTS is valid (when sps_explicit_mts_flag is 1), the MTS setting unit 152111 uses explicit MTS and uses the mts_idx decoded from the encoded data in subsequent processing. The flag explicitMtsEnabled indicating whether explicit MTS is valid may be set separately for the intra mode and the inter mode. In this case, when the prediction mode PredMode is the inter mode (other than MODE_INTRA) and sps_explicit_mts_inter_enabled_flag is 1, or when PredMode is the intra mode (MODE_INTRA) and sps_explicit_mts_intra_enabled_flag is 1, it is determined that explicit MTS is valid, and mts_idx may be decoded from the encoded data. Further, mts_idx may be decoded limited to the case where both the width and height of the TU are 32 or less (nTbW <= 32 && nTbH <= 32).

[0146] (implicit MTS flag setting) The MTS setting unit 152111 sets the implicit MTS flag (implicitMtsEnabled) to 1 when the MTS flag is valid (sps_mts_enabled_flag == 1) and the explicit MTS flag does not indicate validity (explicitMtsEnabled == 0). More specifically, the MTS setting unit 152111 sets implicitMtsEnabled = 1 when any of the following conditions are met, and implicitMtsEnabled = 0 otherwise.

[0147] · When Intra Sub Part Split is on (IntraSubPartSplitType != ISP_NO_SPLIT) · When CU sub-conversion is on and the TU is less than a predetermined size (cu_sbt_flag == 1 and Max(nTbW, nTbH ) < 32) · When explicit MTS is off (both sps_explicit_mts_intra_enabled_flag and sps_explicit_mts_inter_enabled_flag are 0) and PredMode is MODE_INTRA In other cases, the MTS setting unit 152111 sets mts_idx = 0.

[0148] (Explicit MTS) When the explicit MTS indicates validity (when sps_explicit_mts_flag is 1), the TU decoding unit 3024 decodes mts_idx from the encoded data.

[0149] (Explicit MTS Restriction) The TU decoding unit 3024 may limit the range (type) of the transformation matrix selected by the transformation unit according to whether the secondary transformation is effective or not. For example, when the explicit MTS is effective and the secondary transformation is effective (stIdx!= 0), the TU decoding unit 3024 decodes the mts_idx of the maximum value cMaxSt1. When the secondary transformation is not effective otherwise (stIdx == 0), the TU decoding unit 3024 decodes the mts_idx of the maximum value cMaxSt0. Assume that cMaxSt0 > cMaxSt1 here.

[0150] (Restriction example 1) When the explicit MTS is effective and the secondary transformation is effective (stIdx!= 0), the TU decoding unit 3024 sets mts_idx = 0 (trTypeHor = trTypeVer = 0 = DCT2). At this time, the maximum value cMax of mts_idx is 0. In other cases, mts_idx = 0, 1, 2, 3, or 4 is decoded. At this time, the maximum value cMax of mts_idx is 4. As described later, mts_idx = 1, 2, 3, 4 may be a combination of DST7 and DCT8, a combination of DCT8 and DST7, or a combination of DCT8 and DCT8 as trTypeHor and trTypeVer, respectively. Instead of DST7, a transformation combining DST1, DCT4, or pre / post-processing and DCT2 may be performed.

[0151] (Restriction example 2) When the explicit MTS is valid and the secondary transformation is valid (stIdx!= 0), the TU decoder 3024 decodes mts_idx. mts_idx is 0 (trTypeHor = trTypeVer = 0 = DCT2) or 1 (trTypeHor = trTypeVer = 1 = DST7). At this time, the maximum value cMax of mts_idx is 1. Otherwise, it decodes any one of 0, 1, 2, 3, 4 as mts_idx. At this time, the maximum value cMax of mts_idx is 4. Note that as described later, mts_idx = 2, 3, 4 may be a combination of DST7 and DCT8, or a combination of DCT8 and DST7, or a combination of DCT8 and DCT8 for trTypeHor and trTypeVer respectively.

[0152] (Restriction Example 3) When the explicit MTS is valid and the secondary transformation is on (stIdx!= 0), the TU decoder 3024 decodes any one of 0, 1, 2 as mts_idx. At this time, the maximum value cMax of mts_idx is 2. Otherwise, it decodes any one of 0, 1, 2, 3, 4 as mts_idx. At this time, the maximum value cMax of mts_idx is 4.

[0153] According to the above configuration, in the case of secondary transformation, the effective range of MTS can be limited, so that the effect of simplifying encoding is achieved. For example, in Restriction Example 2, since the secondary transformation is not performed in the case of DCT8 where the effect overlaps with the secondary transformation, the overhead due to mts_idx is reduced and the encoding efficiency is improved.

[0154] <Summary of Explicit MTS> An image decoding apparatus including a conversion unit that converts a conversion coefficient for each TU, the conversion unit including: a second conversion unit that applies a conversion using a conversion matrix to an input conversion coefficient when secondary conversion is valid; and a first conversion unit that selects one conversion matrix indicated by mtx_idx from two or more conversion matrices and applies a conversion to the conversion coefficient. The TU decoding unit that decodes mts_idx decodes a value in a first range as mts_idx when secondary conversion is valid (stIdx!= 0), and decodes a value in a second range when secondary conversion is not valid (stIdx == 0), where the second range includes the first range. The TU decoding unit also decodes mts_idx. Here, mts_idx has a configuration where, when secondary conversion is valid (stIdx!= 0), the maximum value is cMaxSt1, when secondary conversion is not valid (stIdx == 0), the maximum value is cMaxSt0, and cMaxSt1 < cMaxSt0.

[0155] (Implicit MTS) The implicit MTS setting unit 152112 performs the following processing in the case of implicit MTS.

[0156] (SM001) When using the intra sub-partition mode (IntraSubPartSplitType!= ISP_NO_SPLIT), the implicit MTS setting unit 152112 sets either 0 (DCT2) or 1 (DST7) as the conversion types tyTypeHor and tyTypeVer according to the intra prediction mode IntraPredMode and the TU size as shown in FIG. 14.

[0157] (SM002) In other cases and when sub-block conversion is on (cu_sbt_flag == 1), the implicit MTS setting unit 152112 sets either 1 (DST7) or 2 (DCT8) as tyTypeHor and tyTypeVer according to cu_sbt_horizontal_flag and cu_sbt_pos_flag as shown in FIG. 15.

[0158] (SM003) The implicit MTS setting unit 152112 sets either 0 (DCT2) or 1 (DST7) as tyTypeHor and tyTypeVer according to the TU size (width nTbW, height nTbH) in cases other than the above (default implicit MTS). Specifically, as shown in FIG. 21, when the width nTbW is within a predetermined range as the horizontal conversion type trTypeHor (S1301), it is set to 1 (DCT1) (S1302), and in other cases, it is set to 0 (DCT2) (S1303). Similarly, when the height nTbH is within a predetermined range as the vertical conversion type trTypeVer (S1304), 1 (DCT1) is set (S1305), and in other cases, 0 (DCT2) is set (S1306).

[0159] trTypeHor = ( nTbW >= 4 && nTbW <= 16 && nTbW <= nTbH )? 1 : 0 trTypeVer = ( nTbH >= 4 && nTbH <= 16 && nTbH <= nTbW )? 1 : 0 Note that the predetermined range is not limited to the above. For example, the following may also be used.

[0160] trTypeHor = ( nTbW >= 4 && nTbW <= 8 && nTbW <= nTbH )? 1 : 0 trTypeVer = ( nTbH >= 4 && nTbH <= 8 && nTbH <= nTbW )? 1 : 0 The above default implicit MTS is the most common mode of implicit MTS.

[0161] (Embodiment 1 of Implicit MTS) When the secondary conversion is on (stIdx!=0), the MTS setting unit 15211 does not perform implicit MTS and sets implicitMtsEnabled to 0. Specifically, as shown in FIG. 20, in the above (implicit MTS flag setting), when the MTS setting unit 15211 satisfies any of the following conditions and the secondary conversion is not on (other than stIdx!=0) (S1500), it sets implicitMtsEnabled = 1 (S1504), and in other cases, it sets implicitMtsEnabled = 0 (S1505).

[0162] · (S1501) When the intra sub-part split is on (IntraSubPartSplitType!= ISP_NO_SPLIT) · (S1502) When the CU sub-conversion is on and the TU is less than a predetermined size (cu_sbt_flag == 1 and Max(nTbW, nTbH) < 32) · (S1503) When the explicit MTS is off (both sps_explicit_mts_intra_enabled_flag and sps_explicit_mts_inter_enabled_flag are 0) and PredMode is MODE_INTRA Note that the determinations of S1501 to S1503 shown in the boxes indicated by the dotted lines in the figure may be different. For example, when the intra sub-part split determination is not performed, when the sub-block conversion is not performed, or when other prediction or conversion determinations are added, it is also possible.

[0163] According to the above configuration, even when MTS is effective, when using secondary conversion (stIdx!=0), implicit MTS is not used. As a result, when using secondary conversion, by using DCT2 as MTS, the effect of improving the coding efficiency is achieved.

[0164] (Embodiment 2 of Implicit MTS) The implicit MTS setting unit 152112 may derive trTypHor = trTypeVer = 0 when the MTS flag is valid (sps_mts_enabled_flag == 1) and the explicit MTS flag does not indicate validity (explicitMtsEnabled == 0). For example, SM000 may be performed before the above SM001.

[0165] Figure 21 is a diagram for explaining the operation of the implicit MTS setting unit 152112.

[0166] (SM000) When stIdx!= 0, the implicit MTS setting unit 152112 derives trTypeHor = trTypeVer = 0.

[0167] As shown in SM003 of Figure 21, when stIdx == 0, the implicit MTS setting unit 152112 may derive the conversion type by the default implicit MTS derivation (SM003) already described. Also, the conversion type may be derived by SM001 and SM002.

[0168] According to the above configuration, when the implicit MTS is valid and the secondary conversion is valid, by using DCT2 as the MTS, the effect of improving the encoding efficiency is achieved.

[0169] (Embodiment 3 of Implicit MTS) When the secondary conversion is on (stIdx!= 0) and neither the MTS by the intra sub-part split mode (SM001, IntraSubPartSplitType!= ISP_NO_SPLIT) nor the MTS by the sub-block conversion (SM002, cu_sbt_flag == 1) is used, the implicit MTS setting unit 152112 may not use the implicit MTS (for example, set implicitMtsEnabled to 0).

[0170] Also, when the secondary conversion is on (stIdx!= 0), and neither the MTS by the intra sub - partition mode (SM001, IntraSubPartSplitType!= ISP_NO_SPLIT) nor the MTS by the sub - block conversion (SM002, cu_sbt_flag == 1) is used, trTypeHor = trTypeVer = 0 is derived. For example, as shown in FIG. 22, SM003´ may be performed instead of the above - mentioned SM003.

[0171] (SM003´) The implicit MTS setting unit 152112 sets either 0 (DCT2) or 1 (DST7) as tyTypeHor and tyTypeVer according to the secondary conversion and the TU size (width nTbW, height nTbH) in cases other than the above (default implicit MTS). For example, the implicit MTS setting unit 152112 selects 1 (DCT1) as the horizontal conversion type trTypeHor when stIdx == 0 and the width nTbW is within a predetermined range (S1301´) (S1302), and sets 0 (DCT2) otherwise (S1303). Similarly, as the vertical conversion type trTypeVer, 1 (DCT1) is set when stIdx == 0 and the height nTbH is within a predetermined range (S1304´) (S1305), and 0 (DCT2) is set otherwise (S1306).

[0172] trTypeHor = (stIdx == 0 && nTbW >= 4 && nTbW <= 16 && nTbW <= nTbH)? 1 : 0 trTypeVer = (stIdx == 0 && nTbH >= 4 && nTbH <= 16 && nTbH <= nTbW)? 1 : 0 According to the above configuration, when the implicit MTS is effective and the secondary conversion is effective, by using DCT2 as the default MTS, the effect of improving the coding efficiency is achieved.

[0173] The MTS setting unit 15211 derives the index trType of the conversion set to be used according to the following formula and outputs it to the coefficient conversion processing unit 15212. The coefficient conversion processing unit 15212 outputs the input trType to the conversion matrix derivation unit 152131. The MTS setting unit 152111 derives the value indicating the MTS to be used according to the following formula.

[0174] When mts_idx == 0, trTypeHor = 0 trTypeVer = 0 When mts_idx == 1, trTypeHor = 1 trTypeVer = 1 When mts_idx == 2, trTypeHor = 2 trTypeVer = 1 When mts_idx == 3, trTypeHor = 1 trTypeVer = 2 When mts_idx == 4, trTypeHor = 2 trTypeVer = 2 Note that the conversion bases corresponding to the cases where tyType (trTypeHor or trTypeVer) is 0, 1, 2 may be DCT2, DST7, DCT8.

[0175] The coefficient conversion processing unit 15212 is composed of a vertical conversion unit 152121 that performs vertical conversion on the correction conversion coefficient d[ ][ ] and a horizontal conversion unit 152123 that performs horizontal conversion.

[0176] The vertical conversion unit 152121 (coefficient conversion processing unit 15212) performs the following processing.

[0177] e[ x ][ y ] = Σ (transMatrix[ y ][ j ]×d[ x ][ j ]) (j = 0..nTbS - 1) Here, transMatrix[ ][ ](=transMatrixV[ ][ ]) is a conversion basis represented by an nTbS × nTbS matrix derived using trTypeVer. nTbS is the height nTbH of the TU. In the case of the 4×4 conversion (nTbS = 4) of DCT2 with trType == 0, for example, transMatrix = {{29, 55, 74, 84}, {74, 74, 0, -74}, {84, -29, -74, 55}, {55, -84, 74, -29}} is used. The symbol Σ means the process of adding the products of the matrix transMatrix[ y ][ j ] and the conversion coefficient d[ x ][ j ] for the subscript j from j = 0..nTbS-1. That is, e[ x ][ y ] is the result of arranging the columns obtained from the product of the vector x[ j ] (j = 0..nTbS-1) consisting of d[ x ][ j ] (j = 0..nTbS-1) which is each column of d[ x ][ y ] and the element transMatrix[ y ][ j ] of the matrix.

[0178] The intermediate clip unit 152122 derives the intermediate value g[ ][ ] by clipping the intermediate value e[ ][ ] and sends it to the horizontal conversion unit 152123.

[0179] g[ x ][ y ] = Clip3( coeffMin, coeffMax, ( e[ x ][ y ] + 64 ) >> 7 ) The 64 and 7 in the above formula are numerical values determined from the bit depth of the conversion basis. In the above formula, the conversion basis is assumed to be 7 bits. Also, coeffMin and coeffMax are the minimum and maximum values for clipping.

[0180] The horizontal conversion unit 152123 (coefficient conversion processing unit 15212) performs the following processing. transMatrix[ ][ ] (=transMatrixH[ ][ ]) is a conversion basis represented by an nTbS × nTbS matrix derived using trTypeHor. nTbS is the width nTbW of the TU. The horizontal conversion unit 152123 converts the intermediate value g[ x ][ y ] into the prediction residual r[ x ][ y ] by one-dimensional horizontal conversion.

[0181] r[ x ][ y ] =Σ transMatrix[ x ][ j ]×g[ j ][ y ] (j = 0..nTbS-1) The symbol Σ above means the process of adding the product of the matrix transMatrix[ x ][j] and g[j][ y ] for the subscript j from j = 0 to nTbS-1. That is, r[ x ][ y ] is the result of arranging the rows obtained from the product of g[ j ][ y ] (j = 0..nTbS-1), which are the rows of g[ x ][ y ], and the matrix transMatrix.

[0182] The prediction residual r[ ][ ] is sent from the horizontal conversion unit 152123 to the adder 312.

[0183] The vertical conversion unit 152121 and the horizontal conversion unit 152123 perform the conversion by the matrix conversion processing unit 15213. The matrix conversion processing unit 15213 is composed of a conversion matrix derivation unit 152131 and a conversion processing unit 152132.

[0184] The conversion matrix derivation unit 152131 derives the conversion matrix transMatrix[ ][ ] according to the length of the TU (nTbW, nTbH) and the index tyType (trTypeHor, trTypeVer) of the core conversion.

[0185] The matrix transformation processing unit 15213 uses the derived transformation matrix transMatrix[ ][ ] to transform the input one-dimensional array xx[ j ] into a one-dimensional array yy[ i ], performing vertical transformation and horizontal transformation. In the vertical transformation, the transformation coefficient d[ x ][j] of column x is input as the one-dimensional transformation coefficient xx[j] for transformation. In the horizontal transformation, the intermediate coefficient g[ j ][ y ] of row y is input as xx[ j ] for transformation.

[0186] yy[i] = Σ (transMatrix[ i ][ j ] × xx[ j ]) (j = 0..nTbS-1) <Secondary transformation> When the TU decoding unit 3024 decodes the secondary transformation stIdx, it may limit the range of the value of stIdx to be decoded according to whether the value of mts_idx is valid.

[0187] (Restriction example 1) When mts_idx is 0, the TU decoding unit 3024 sets the maximum value cMax of stIdx to be decoded to 1 and decodes stIdx = 0~2. In other cases, that is, when mts_idx = 1, 2, 3, 4, stIdx = 0 is derived without decoding stIdx from the encoded data.

[0188] (Restriction example 2) When mts_idx is 0..1, the TU decoding unit 3024 sets the maximum value cMax of stIdx to be decoded to 1 and decodes stIdx = 0~2. In other cases, that is, when mts_idx = 2, 3, 4, stIdx = 0 is derived without decoding stIdx from the encoded data.

[0189] (Restriction example 3) When mts_idx is 0, 1, 2, the TU decoding unit 3024 sets the maximum value cMax of stIdx to be decoded to 1 and decodes stIdx = 0~2. In other cases, that is, when mts_idx = 3, 4, stIdx = 0 is derived without decoding stIdx from the encoded data.

[0190] According to the above configuration, since the variable stIdx indicating the type of secondary conversion is decoded only when the MTS performs conversion within a predetermined range, the effective range of the secondary conversion is limited, resulting in the effect of simplified encoding. Also, for example, in Restriction Example 2, since the secondary conversion is not performed in the case of DCT8 where the effects overlap with the secondary conversion, the overhead due to stIdx is reduced, resulting in the effect of improved encoding efficiency.

[0191] The addition unit 312 adds the predicted image of the block input from the predicted image generation unit 308 and the prediction error input from the inverse quantization and inverse transformation unit 311 for each pixel to generate the decoded image of the block. The addition unit 312 stores the decoded image of the block in the reference picture memory 306 and also outputs it to the loop filter 305.

[0192] (Configuration of the moving image encoding device) Next, the configuration of the moving image encoding device 11 according to the present embodiment will be described. FIG. 16 is a block diagram showing the configuration of the moving image encoding device 11 according to the present embodiment. The moving image encoding device 11 includes a predicted image generation unit 101, a subtraction unit 102, a conversion and quantization unit 103, an inverse quantization and inverse transformation unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, an encoding parameter determination unit 110, a parameter encoding unit 111, and an entropy encoding unit 104.

[0193] The predicted image generation unit 101 generates a predicted image for each CU, which is a region obtained by dividing each picture of the image T. The predicted image generation unit 101 performs the same operation as the predicted image generation unit 308 described above, and the description thereof is omitted.

[0194] The subtraction unit 102 subtracts the pixel value of the predicted image of the block input from the predicted image generation unit 101 from the pixel value of the image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the conversion and quantization unit 103.

[0195] The conversion and quantization unit 103 calculates conversion coefficients by frequency conversion and derives quantized conversion coefficients by quantization for the prediction error input from the subtraction unit 102. The conversion and quantization unit 103 outputs the quantized conversion coefficients to the entropy encoding unit 104 and the inverse quantization and inverse conversion unit 105.

[0196] As shown in FIG. 19, the conversion and quantization unit 103 includes a forward core conversion 10321 (first conversion unit) and a forward secondary conversion unit 10322 (second conversion unit).

[0197] In the forward secondary conversion applied to the moving image encoding device 11, substantially the same processing is performed as that of the secondary conversion applied to the moving image decoding device 31, except that the processing S1 - S4 of the secondary conversion is applied in the reverse order of S1, S4, S3, S2.

[0198] In processing S1, the forward secondary conversion unit 10322 performs the same processing as the secondary conversion unit 31121, except that the input and output of the secondary conversion are the length nStOutSize and nonZeroSize, respectively.

[0199] In processing S4, the forward secondary conversion unit 10322 derives a one - dimensional array v[] of nStOutSize (or nStSize * nStSize) from the conversion coefficients d[][] at a predetermined position within the TU.

[0200] In processing S3, the forward secondary conversion unit 10322 obtains a one - dimensional array u[] (vector F) of nonZeroSize by the following conversion from a one - dimensional array v[] of nStOutSize (vector V) and a conversion matrix T[][].

[0201] F = trans(T) × V Here, trans(T) is the transpose matrix of T. The secondary conversion unit may derive the one - dimensional array u[] (vector F) by the following formula.

[0202] F = Tinv × V Here, Tinv is the inverse matrix of T. T is composed of a first type of transformation basis T1 and a second type of transformation basis T2. Note that the secondary conversion unit may use an orthogonal matrix for T to set trans(T) of T as Tinv.

[0203] Note that in actual processing, since T is a matrix of integer values, instead of T×Tinv = I (identity matrix), it becomes a constant multiple of the identity matrix (T×Tinv = K2×I, where K2 is a constant). In this case, the secondary conversion unit uses a matrix that is a constant multiple of the inverse matrix as Tinv, but for the transposed matrix, the inverse matrix may be used as it is.

[0204] In process S2, the forward secondary conversion unit 10322 rearranges the one-dimensional array u[] of nonZeroSize into a two-dimensional array to derive the conversion coefficient d[][].

[0205] for (x = 0; x < nonZeroSize; x++) { xC = (xSbIdx << log2StSize) + DiagScanOrder[log2StSize][log2StSize][x][0] yC = (ySbIdx << log2StSize) + DiagScanOrder[log2StSize][log2StSize][x][1] d[xC][yC] = u[x] } The inverse quantization and inverse transformation unit 105 is the same as the inverse quantization and inverse transformation unit 311 (Fig. 15) in the moving image decoding device 31, and the description is omitted. The calculated prediction error is output to the addition unit 106.

[0206] The entropy encoding unit 104 receives the quantized conversion coefficients from the conversion and quantization unit 103 and the encoding parameters from the parameter encoding unit 111. The encoding parameters are, for example, predMode.

[0207] The entropy encoding unit 104 entropy-encodes split information, prediction parameters, quantized transform coefficients, etc., to generate and output an encoded stream Te.

[0208] The parameter encoding unit 111 includes a header encoding unit 1110 (not shown), a CT information encoding unit 1111, a CU encoding unit 1112 (prediction mode encoding unit), and an inter prediction parameter encoding unit 112 and an intra prediction parameter encoding unit 113. The CU encoding unit 1112 further includes a TU encoding unit 1114.

[0209] The following is an explanation of the general operation of each module. The parameter encoding unit 111 performs encoding processing of parameters such as header information, split information, prediction information, and quantized transform coefficients.

[0210] The CT information encoding unit 1111 encodes QT, MT (BT, TT) split information, etc., from the encoded data.

[0211] The CU encoding unit 1112 encodes CU information, prediction information, TU split flag, CU residual flag, etc.

[0212] When the TU contains prediction error, the TU encoding unit 1114 encodes QP update information (quantization correction value) and quantized prediction error (residual_coding).

[0213] The CT information encoding unit 1111 and the CU encoding unit 1112 supply syntax elements such as inter prediction parameters, intra prediction parameters (intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_remainder), and quantized transform coefficients to the entropy encoding unit 104.

[0214] (Configuration of the intra prediction parameter encoding unit 113) The Intra Prediction Parameter Encoding Unit 113 derives a format for encoding (such as intra_luma_mpm_idx, intra_luma_mpm_remainder, etc.) from the IntraPredMode input from the Encoding Parameter Determination Unit 110. The Intra Prediction Parameter Encoding Unit 113 includes a configuration that is partially the same as the configuration in which the Intra Prediction Parameter Decoding Unit 304 derives the intra prediction parameters.

[0215] FIG. 17 is a schematic diagram showing the configuration of the Intra Prediction Parameter Encoding Unit 113 of the Parameter Encoding Unit 111. The Intra Prediction Parameter Encoding Unit 113 includes a Parameter Encoding Control Unit 1131, a Luminance Intra Prediction Parameter Derivation Unit 1132, and a Chrominance Intra Prediction Parameter Derivation Unit 1133.

[0216] The Parameter Encoding Control Unit 1131 receives IntraPredModeY and IntraPredModeC from the Encoding Parameter Determination Unit 110. The Parameter Encoding Control Unit 1131 determines intra_luma_mpm_flag by referring to mpmCandList[] of the MPM Candidate List Derivation Unit 30421. Then, intra_luma_mpm_flag and IntraPredModeY are output to the Luminance Intra Prediction Parameter Derivation Unit 1132. Also, IntraPredModeC is output to the Chrominance Intra Prediction Parameter Derivation Unit 1133.

[0217] The Luminance Intra Prediction Parameter Derivation Unit 1132 includes an MPM Candidate List Derivation Unit 30421 (candidate list derivation unit), an MPM Parameter Derivation Unit 11322, and a non-MPM Parameter Derivation Unit 11323 (encoding unit, derivation unit).

[0218] The MPM candidate list derivation unit 30421 derives mpmCandList[] by referring to the intra prediction mode of adjacent blocks stored in the prediction parameter memory 108. When intra_luma_mpm_flag is 1, the MPM parameter derivation unit 11322 derives intra_luma_mpm_idx from IntraPredModeY and mpmCandList[] and outputs it to the entropy encoding unit 104. When intra_luma_mpm_flag is 0, the non-MPM parameter derivation unit 11323 derives RemIntraPredMode from IntraPredModeY and mpmCandList[] and outputs intra_luma_mpm_remainder to the entropy encoding unit 104.

[0219] The chroma intra prediction parameter derivation unit 1133 derives and outputs intra_chroma_pred_mode from IntraPredModeY and IntraPredModeC.

[0220] The addition unit 106 adds the pixel values of the predicted image of the block input from the prediction image generation unit 101 and the prediction error input from the inverse quantization / inverse transformation unit 105 for each pixel to generate a decoded image. The addition unit 106 stores the generated decoded image in the reference picture memory 109.

[0221] The loop filter 107 performs a deblocking filter, SAO, and ALF on the decoded image generated by the addition unit 106. Note that the loop filter 107 does not necessarily include the above three types of filters, and for example, it may have a configuration including only the deblocking filter.

[0222] The prediction parameter memory 108 stores the prediction parameters generated by the encoding parameter determination unit 110 at predetermined positions for each target picture and CU.

[0223] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at predetermined positions for each target picture and CU.

[0224] The parameter determination unit 110 for encoding selects one set from a plurality of sets of encoding parameters. The encoding parameters are the QT, BT, or TT segmentation information, prediction parameters, or parameters to be encoded generated in relation to these, as described above. The prediction image generation unit 101 generates a prediction image using these encoding parameters.

[0225] The parameter determination unit 110 for encoding calculates an RD cost value indicating the amount of information and the encoding error for each of the plurality of sets. The parameter determination unit 110 for encoding selects the set of encoding parameters for which the calculated cost value is the minimum. As a result, the entropy encoding unit 104 outputs the selected set of encoding parameters as an encoded stream Te. The parameter determination unit 110 for encoding stores the determined encoding parameters in the prediction parameter memory 108.

[0226] Note that, a part of the moving image encoding device 11 and the moving image decoding device 31 in the above-described embodiment, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generation unit 308, the inverse quantization / inverse transform unit 311, the addition unit 312, the predicted image generation unit 101, the subtraction unit 102, the transform / quantization unit 103, the entropy encoding unit 104, the inverse quantization / inverse transform unit 105, the loop filter 107, the encoding parameter determination unit 110, and the parameter encoding unit 111 may be implemented by a computer. In that case, a program for realizing this control function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to realize it. Here, the "computer system" refers to a computer system built in either the moving image encoding device 11 or the moving image decoding device 31, including hardware such as an OS and peripheral devices. Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, etc., and a storage device such as a hard disk built in a computer system. Furthermore, the "computer-readable recording medium" also includes, like a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, something that holds a program dynamically for a short time, and something that holds a program for a certain time, like a volatile memory inside a computer system that becomes a server or a client in that case. Also, the above program may be for realizing a part of the aforementioned functions, and may further be something that can be realized in combination with a program already recorded in the computer system for the aforementioned functions.

[0227] Further, part or all of the moving image encoding device 11 and the moving image decoding device 31 in the above-described embodiments may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the moving image encoding device 11 and the moving image decoding device 31 may be individually processed by a processor, or part or all of them may be integrated and processed by a processor. Further, the method of integrating into an integrated circuit is not limited to an LSI, and may be realized by a dedicated circuit or a general-purpose processor. Also, when a technology for integrating into an integrated circuit that replaces an LSI appears due to the progress of semiconductor technology, an integrated circuit using such technology may be used.

[0228] As described above, an embodiment of the present invention has been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist of the present invention.

[0229] 〔Application Example〕 The above-described moving image encoding device 11 and moving image decoding device 31 can be mounted and used in various devices that transmit, receive, record, and play back moving images. Note that the moving image may be a natural moving image captured by a camera or the like, or an artificial moving image (including CG and GUI) generated by a computer or the like.

[0230] First, the fact that the above-described moving image encoding device 11 and moving image decoding device 31 can be used for transmitting and receiving moving images will be described with reference to FIG. 2.

[0231] FIG. 2 shows a block diagram showing the configuration of a transmission device PROD_A equipped with the moving image encoding device 11. As shown in the figure, the transmission device PROD_A includes an encoding unit PROD_A1 that obtains encoded data by encoding a moving image, a modulation unit PROD_A2 that obtains a modulation signal by modulating a carrier wave with the encoded data obtained by the encoding unit PROD_A1, and a transmission unit PROD_A3 that transmits the modulation signal obtained by the modulation unit PROD_A2. The above-described moving image encoding device 11 is used as this encoding unit PROD_A1.

[0232] The transmitting device PROD_A may further include a camera PROD_A4 for capturing a moving image, a recording medium PROD_A5 for recording the moving image, an input terminal PROD_A6 for externally inputting the moving image, and an image processing unit A7 for generating or processing an image, as a source of the moving image input to the encoding unit PROD_A1. In the figure, a configuration in which the transmitting device PROD_A includes all of these is illustrated, but a part of them may be omitted.

[0233] Note that the recording medium PROD_A5 may record an unencoded moving image, or may record a moving image encoded by a recording encoding method different from the transmission encoding method. In the latter case, a decoding unit (not shown) for decoding the encoded data read from the recording medium PROD_A5 according to the recording encoding method may be interposed between the recording medium PROD_A5 and the encoding unit PROD_A1.

[0234] Also, FIG. 2 shows a block diagram showing the configuration of the receiving device PROD_B equipped with the moving image decoding device 31. As shown in the figure, the receiving device PROD_B includes a receiving unit PROD_B1 for receiving a modulation signal, a demodulation unit PROD_B2 for obtaining encoded data by demodulating the modulation signal received by the receiving unit PROD_B1, and a decoding unit PROD_B3 for obtaining a moving image by decoding the encoded data obtained by the demodulation unit PROD_B2. The above-described moving image decoding device 31 is used as this decoding unit PROD_B3.

[0235] The receiving device PROD_B may further include a display PROD_B4 for displaying a moving image, a recording medium PROD_B5 for recording the moving image, and an output terminal PROD_B6 for externally outputting the moving image, as a destination of the moving image output from the decoding unit PROD_B3. In the figure, a configuration in which the receiving device PROD_B includes all of these is illustrated, but a part of them may be omitted.

[0236] Note that the recording medium PROD_B5 may be for recording unencoded moving images, or may be encoded using an encoding method for recording different from the encoding method for transmission. In the latter case, an encoding unit (not shown) for encoding the moving image obtained from the decoding unit PROD_B3 according to the encoding method for recording may be interposed between the decoding unit PROD_B3 and the recording medium PROD_B5.

[0237] Note that the transmission medium for transmitting the modulation signal may be wireless or wired. Also, the transmission mode for transmitting the modulation signal may be broadcast (here, referring to a transmission mode where the transmission destination is not specified in advance), or may be communication (here, referring to a transmission mode where the transmission destination is specified in advance). That is, the transmission of the modulation signal may be realized by any of wireless broadcast, wired broadcast, wireless communication, and wired communication.

[0238] For example, a broadcast station (broadcast equipment, etc.) / reception station (television receiver, etc.) for terrestrial digital broadcast is an example of the transmission device PROD_A / reception device PROD_B that transmits and receives the modulation signal by wireless broadcast. Also, a broadcast station (broadcast equipment, etc.) / reception station (television receiver, etc.) for cable television broadcast is an example of the transmission device PROD_A / reception device PROD_B that transmits and receives the modulation signal by wired broadcast.

[0239] Also, a server (workstation, etc.) / client (television receiver, personal computer, smartphone, etc.) for VOD (Video On Demand) service or video sharing service using the Internet is an example of the transmission device PROD_A / reception device PROD_B that transmits and receives the modulation signal by communication (usually, either wireless or wired is used as the transmission medium in a LAN, and wired is used as the transmission medium in a WAN). Here, personal computers include desktop PCs, laptop PCs, and tablet PCs. Also, smartphones include multifunctional mobile phone terminals.

[0240] In addition to the function of decoding the encoded data downloaded from the server and displaying it on the display, the client of the video sharing service has a function of encoding the moving images captured by the camera and uploading them to the server. That is, the client of the video sharing service functions as both the transmission device PROD_A and the reception device PROD_B.

[0241] Next, it will be described with reference to FIG. 3 that the above-described moving image encoding device 11 and moving image decoding device 31 can be used for recording and playing back moving images.

[0242] FIG. 3 shows a block diagram showing the configuration of the recording device PROD_C equipped with the above-described moving image encoding device 11. As shown in the figure, the recording device PROD_C includes an encoding unit PROD_C1 that obtains encoded data by encoding a moving image, and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 to the recording medium PROD_M. The above-described moving image encoding device 11 is used as this encoding unit PROD_C1.

[0243] Note that the recording medium PROD_M may be of a type built into the recording device PROD_C, such as (1) an HDD (Hard Disk Drive) or an SSD (Solid State Drive), or may be of a type connected to the recording device PROD_C, such as (2) an SD memory card or a USB (Universal Serial Bus) flash memory, or may be loaded into a drive device (not shown) built into the recording device PROD_C, such as (3) a DVD (Digital Versatile Disc: registered trademark) or a BD (Blu-ray Disc: registered trademark).

[0244] In addition, the recording device PROD_C may further include a camera PROD_C3 that captures a moving image, an input terminal PROD_C4 for inputting a moving image from the outside, a receiving unit PROD_C5 for receiving a moving image, and an image processing unit PROD_C6 for generating or processing an image, as a supply source of the moving image input to the encoding unit PROD_C1. In the figure, a configuration in which the recording device PROD_C includes all of these is illustrated, but a part of them may be omitted.

[0245] Note that the receiving unit PROD_C5 may receive an unencoded moving image, or may receive encoded data encoded by a transmission encoding method different from the recording encoding method. In the latter case, a transmission decoding unit (not shown) for decoding the encoded data encoded by the transmission encoding method may be interposed between the receiving unit PROD_C5 and the encoding unit PROD_C1.

[0246] Examples of such a recording device PROD_C include a DVD recorder, a BD recorder, an HDD (Hard Disk Drive) recorder, etc. (in this case, the input terminal PROD_C4 or the receiving unit PROD_C5 serves as the main supply source of the moving image). Also, a camcorder (in this case, the camera PROD_C3 serves as the main supply source of the moving image), a personal computer (in this case, the receiving unit PROD_C5 or the image processing unit C6 serves as the main supply source of the moving image), a smartphone (in this case, the camera PROD_C3 or the receiving unit PROD_C5 serves as the main supply source of the moving image), etc. are also examples of such a recording device PROD_C.

[0247] Also, FIG. 3 shows a block diagram showing the configuration of a playback device PROD_D equipped with the above-described moving image decoding device 31. As shown in the figure, the playback device PROD_D includes a reading unit PROD_D1 that reads the encoded data written on the recording medium PROD_M, and a decoding unit PROD_D2 that obtains a moving image by decoding the encoded data read by the reading unit PROD_D1. The above-described moving image decoding device 31 is used as this decoding unit PROD_D2.

[0248] Note that the recording medium PROD_M may be of a type built into the playback device PROD_D, such as an HDD or SSD, etc., or may be of a type connected to the playback device PROD_D, such as an SD memory card or USB flash memory, etc., or may be loaded into a drive device (not shown) built into the playback device PROD_D, such as a DVD or BD, etc.

[0249] In addition, the playback device PROD_D may further include a display PROD_D3 for displaying a moving image, an output terminal PROD_D4 for outputting the moving image externally, and a transmission unit PROD_D5 for transmitting the moving image as destinations for the moving image output by the decoding unit PROD_D2. In the figure, a configuration in which the playback device PROD_D includes all of these is illustrated, but a part of them may be omitted.

[0250] Note that the transmission unit PROD_D5 may transmit an unencoded moving image or may transmit encoded data encoded by a transmission encoding method different from the recording encoding method. In the latter case, an encoding unit (not shown) for encoding the moving image by the transmission encoding method may be interposed between the decoding unit PROD_D2 and the transmission unit PROD_D5.

[0251] Examples of such a playback device PROD_D include a DVD player, a BD player, an HDD player, etc. (in this case, the output terminal PROD_D4 to which a television receiver or the like is connected serves as the main supply destination of the moving image). Also, a television receiver (in this case, the display PROD_D3 serves as the main supply destination of the moving image), digital signage (also referred to as an electronic signboard or an electronic bulletin board, etc., and the display PROD_D3 or the transmission unit PROD_D5 serves as the main supply destination of the moving image), a desktop PC (in this case, the output terminal PROD_D4 or the transmission unit PROD_D5 serves as the main supply destination of the moving image), a laptop or tablet PC (in this case, the display PROD_D3 or the transmission unit PROD_D5 serves as the main supply destination of the moving image), a smartphone (in this case, the display PROD_D3 or the transmission unit PROD_D5 serves as the main supply destination of the moving image), etc. are also examples of such a playback device PROD_D.

[0252] (Hardware Implementation and Software Implementation) In addition, each block of the above-described moving image decoding device 31 and moving image encoding device 11 may be implemented hardware-wise by a logic circuit formed on an integrated circuit (IC chip), or may be implemented software-wise using a CPU (Central Processing Unit).

[0253] In the latter case, each of the above devices includes a CPU that executes instructions of a program for realizing each function, a ROM (Read Only Memory) that stores the above program, a RAM (Random Access Memory) that expands the above program, a storage device (recording medium) such as a memory that stores the above program and various data, etc. And the object of the embodiment of the present invention can also be achieved by supplying a recording medium in which program codes (executable format programs, intermediate code programs, source programs) of control programs of each of the above devices, which are software for realizing the above-described functions, are recorded in a computer-readable manner to each of the above devices, and having the computer (or CPU or MPU) read and execute the program codes recorded in the recording medium.

[0254] As the above-mentioned recording medium, for example, tapes such as magnetic tapes and cassette tapes, magnetic disks such as floppy (registered trademark) disks / hard disks, and optical disks including optical disks such as CD-ROM (Compact Disc Read-Only Memory) / MO disk (Magneto-Optical disc) / MD (Mini Disc) / DVD (Digital Versatile Disc: registered trademark) / CD-R (CD Recordable) / Blu-ray Disc (Blu-ray Disc: registered trademark), cards such as IC cards (including memory cards) / optical cards, semiconductor memories such as mask ROM / EPROM (Erasable Programmable Read-Only Memory) / EEPROM (Electrically Erasable and Programmable Read-Only Memory: registered trademark) / flash ROM, or logic circuits such as PLD (Programmable logic device) and FPGA (Field Programmable Gate Array) can be used.

[0255] In addition, each of the above devices may be configured to be connectable to a communication network, and the above program code may be supplied via the communication network. This communication network only needs to be capable of transmitting the program code and is not particularly limited. For example, the Internet, intranet, extranet, LAN (Local Area Network), ISDN (Integrated Services Digital Network), VAN (Value-Added Network), CATV (Community Antenna television / Cable Television) communication network, virtual private network, telephone line network, mobile communication network, satellite communication network, etc. can be used. Also, the transmission medium constituting this communication network only needs to be a medium capable of transmitting the program code and is not limited to a specific configuration or type. For example, it can be wired such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line carrier, cable TV line, telephone line, ADSL (Asymmetric Digital Subscriber Line) line, etc., or wireless such as infrared rays like IrDA (Infrared Data Association) and remote controls, Bluetooth (registered trademark), IEEE802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance: registered trademark), mobile phone network, satellite line, terrestrial digital broadcast network, etc. Note that the embodiments of the present invention can also be realized in the form of a computer data signal embedded in a carrier wave, in which the above program code is embodied by electronic transmission.

[0256] The embodiments of the present invention are not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. That is, embodiments obtained by combining technical means appropriately modified within the scope shown in the claims are also included in the technical scope of the present invention.

[0257] Summary A moving image decoding apparatus according to an aspect of the present invention is an image decoding apparatus that converts conversion coefficients for each conversion unit. When secondary conversion is valid, a second conversion unit that corrects the conversion coefficients by applying a conversion using a conversion matrix to the conversion coefficients, a first conversion unit that applies a separable conversion composed of a vertical conversion and a horizontal conversion to the conversion coefficients, and when the secondary conversion is valid, the intra sub-partition mode is not used, and when sub-block conversion is not used, the implicit conversion is turned off. When the implicit conversion is on, an implicit conversion setting unit that derives a horizontal conversion type according to the width of the target TU and derives a vertical conversion type according to the height of the target TU, wherein the first conversion unit performs a conversion according to the vertical conversion type and a conversion according to the horizontal conversion type. An image decoding apparatus according to an aspect of the present invention is an image decoding apparatus that converts conversion coefficients for each conversion unit, a decoding unit that decodes an index indicating not to use secondary conversion when the value is 0, and when the value of the index is other than 0, a second conversion unit that applies the secondary conversion to the conversion coefficients and outputs corrected conversion coefficients, a first conversion unit that applies a separable conversion composed of a vertical conversion and a horizontal conversion to the conversion coefficients or the corrected conversion coefficients, and based on whether the value of the index is 0 and whether the width of the conversion unit is within a predetermined range, sets the value of the horizontal conversion type variable, and based on whether the value of the index is 0 and whether the height of the conversion unit is within a predetermined range, an implicit conversion setting unit that sets the value of the vertical conversion type variable, wherein the first conversion unit performs the vertical conversion according to the vertical conversion type variable and performs the horizontal conversion according to the horizontal conversion type variable. Industrial Applicability

[0258] Embodiments of the present invention can be suitably applied to a moving image decoding apparatus that decodes encoded data in which image data is encoded, and a moving image encoding apparatus that generates encoded data in which image data is encoded. Further, it can be suitably applied to the data structure of the encoded data generated by the moving image encoding apparatus and referred to by the moving image decoding apparatus.

[0259] (Cross-reference to related applications) This application claims the benefit of priority to Japanese Patent Application: Japanese Patent Application No. 2019-101179, filed on May 30, 2019, and by reference thereto, the entire contents thereof are incorporated herein.

Explanation of symbols

[0260] 31 Moving image decoding apparatus 301 Entropy decoding unit 302 Parameter decoding unit 3020 Header decoding unit 303 Inter prediction parameter decoding unit 304 Intra prediction parameter decoding unit 308 Predicted image generation unit 309 Inter predicted image generation unit 310 Intra predicted image generation unit 311 Inverse quantization and inverse transformation unit 312 Addition unit 11 Moving image encoding apparatus 101 Predicted image generation unit 102 Subtraction unit 103 Transformation and quantization unit 104 Entropy encoding unit 105 Inverse quantization and inverse transformation unit 107 Loop filter 110 Encoded parameter determination unit 111 Parameter encoding unit 112 Inter prediction parameter encoding unit 113 Intra prediction parameter encoding unit 1110 Header encoding unit 1111 CT information encoding unit 1112 CU Symbolization Unit (Prediction Mode Symbolization Unit) 1114 TU Symbolization Unit 3111 Inverse Quantization Unit 3112 Inverse Transformation Unit 31121 Secondary Transformation Unit 31112 Scaling Unit 31123 Core Transformation Unit 10322 Forward Secondary Transformation Unit 10323 Forward Core Transformation Unit

Claims

1. An image decoding device that transforms transform coefficients for each transform unit, comprising: A transform unit decoder that decodes from the encoded data (i) a secondary index indicating whether or not an inverse secondary transform is used and a transform base, and (ii) a multi-transform selection index that is a switching index for selecting a transform base of the inverse core transform; an inverse quantization unit for inverse quantizing the quantized transform coefficients to calculate transform coefficients; a secondary transform unit configured to derive modified transform coefficients by applying the inverse secondary transform to the transform coefficients using a transform matrix if the inverse secondary transform is enabled; a core transform unit that applies the inverse core transform including a vertical transform and a horizontal transform to the transform coefficients or the modified transform coefficients; The core conversion unit includes an MTS setting unit and an implicit MTS setting unit, The MTS setting unit derives a horizontal transform type and a vertical transform type based on the multi-transform selection index when an explicit MTS is enabled; the implicit MTS setting unit sets the horizontal transform type and the vertical transform type to 0 if the value of the secondary index is not equal to 0; The image decoding device, wherein the core transform unit performs an inverse transform based on the vertical transform type, and performs an inverse transform based on the horizontal transform type.

2. The image decoding device according to claim 1, characterized in that, when an intra sub-division mode is used, the implicit MTS setting unit sets the horizontal transform type and the vertical transform type to 0 or 1 based on the intra prediction mode and the size of the transform unit.

3. A computer-readable recording medium for recording a program for causing a computer to convert a conversion coefficient for each conversion unit, The program causes the computer to A step of decoding a secondary index indicating whether an inverse secondary transform is used and a transform base from the encoded data; A step of decoding a multi-transform selection index, which is a switching index for selecting a transform basis of an inverse core transform, from the encoded data; dequantizing the quantized transform coefficients to calculate transform coefficients; if the inverse secondary transform is valid, deriving modified transform coefficients by applying the inverse secondary transform to the transform coefficients using a transform matrix; if explicit MTS is enabled, deriving a horizontal transform type and a vertical transform type based on the multi-transform selection index; if the value of the secondary index is not equal to 0, setting the horizontal transform type and the vertical transform type to 0; applying a vertical transform to the transform coefficients or the modified transform coefficients based on the vertical transform type and a horizontal transform to the transform coefficients or the modified transform coefficients based on the horizontal transform type; A computer-readable recording medium that causes a computer to execute the above-mentioned steps.

Citation Information

Patent Citations

  • Method and apparatus for improved implicit transform selection

    WO2020247306A1

  • Method and apparatus for video coding

    WO2020251743A1