Image decoding device, image encoding device, and recording medium
By applying transformation matrix and separable transformation in the image decoding device, the performance loss problem when combining implicit MTS with quadratic transformation is solved, thereby improving the efficiency and effect of image decoding.
Patent Information
- Application Number
- CN202511394247.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-30
- Filing Date
- 2020-05-29
- Publication Date
- 2025-11-14
AI Technical Summary
In existing image coding methods, implicit MTS performs poorly when combined with quadratic transform and multiple transform selection (MTS), especially when combined with quadratic transform, resulting in performance loss.
An image decoding device is employed to perform transformation by applying a transformation matrix when the quadratic transformation is valid, and when the implicit transformation is enabled, the horizontal and vertical transformation types are derived based on the width and height of the transformation unit, and a split transformation is used for correction.
It improves the performance of image decoding, especially in the case of implicit MTS and quadratic transform combination, enhancing the effectiveness and efficiency of the transform.
Smart Images

Figure CN120956897A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on May 29, 2020 (entering the Chinese national phase on November 25, 2021), with application number 202080039017.1 and title "Image Decoding Device". Technical Field
[0002] The embodiments of the present invention relate to an image decoding device and an image encoding device. Background Technology
[0003] In order to efficiently transmit or record images, an image encoding device is used to generate encoded data by encoding the image, and an image decoding device is used to generate a decoded image by decoding the encoded data.
[0004] Specific image coding methods include H.264 / AVC and HEVC (High-Efficiency Video Coding).
[0005] In this image encoding method, the images (pictures) that constitute the image are managed through a hierarchical structure and encoded / decoded by each CU. The hierarchical structure includes slices obtained by segmenting the image, coding tree units (CTUs) obtained by segmenting the slices, coding units (sometimes also called coding units (CUs)) obtained by segmenting the coding tree units, and transformation units (TUs) obtained by segmenting the coding units.
[0006] Furthermore, in such image encoding methods, a prediction image is typically generated based on a locally decoded image obtained by encoding / decoding the input image, and the prediction error (sometimes called a "difference image" or "residual image") obtained by subtracting the prediction image from the input image (original image) is encoded. Methods for generating prediction images include inter-frame prediction and intra-frame prediction.
[0007] Furthermore, non-patent documents 1 and 2 can be cited as examples of image encoding and decoding technologies in recent years. Non-patent document 1 discloses a technique called Multiple Transform Selection (MTS), which switches the transform matrix based on the explicit syntax or implicit block size in the encoded data. Non-patent document 2 discloses an image encoding apparatus that derives transform coefficients by performing an RST (Reduced Secondary Transform) transformation on each transform unit, i.e., a quadratic transformation, on the transformed coefficients of the prediction error. Furthermore, non-patent document 2 discloses an image decoding apparatus that performs an inverse quadratic transformation on each transform unit on the transform coefficients.
[0008] Existing technical documents
[0009] Non-patent literature
[0010] Non-patent document 1: "Versatile Video Coding (Draft 5)", JVET-N1001-v6, Joint Video Exploration Team (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 2019-05-23
[0011] Non-Patent Document 2: “CE12: Mapping functions (test CE12-1 and CE12-2)”, JVET-M0427-v2, Joint Video Experts Team (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC29 / WG 11 13th Meeting: Marrakech, MA, 9-18 Jan. 2019 Summary of the Invention
[0012] The problem the invention aims to solve
[0013] In techniques such as non-patent document 1, including quadratic transformations and related technologies, there is a problem that the performance of combinations of quadratic transformations and transformations implemented by MTS is insufficient. In particular, there is a problem that the performance of implicit MTS becomes a loss when combined with quadratic transformations.
[0014] The purpose of this invention is to provide an image decoding device and related technologies that can more appropriately apply the transformation and quadratic transformation implemented by MTS.
[0015] Technical solution
[0016] One aspect of the present invention is a motion picture decoding apparatus that transforms transform coefficients per transform unit, characterized by comprising: a second transform unit that, when the secondary transform is valid, applies a transform using a transform matrix to the transform coefficients to correct them; a first transform unit that applies a split-type transform consisting of a vertical transform and a horizontal transform to the transform coefficients; and an implicit transform setting unit that, when the secondary transform is valid, does not utilize intra-frame sub-segmentation mode and does not utilize sub-block transform, sets the implicit transform to be off, and when the implicit transform is enabled, derives a horizontal transform type based on the width of the object TU and derives a vertical transform type based on the height of the object TU, wherein the first transform unit performs a transform corresponding to the vertical transform type and a transform corresponding to the horizontal transform type. Attached Figure Description
[0017] Figure 1 This is a schematic diagram showing the configuration of the image transmission system of this embodiment.
[0018] Figure 2 This diagram illustrates the configuration of a transmitting device equipped with a motion picture encoding apparatus according to this embodiment and a receiving device equipped with a motion picture decoding apparatus. PROD_A represents the transmitting device equipped with the motion picture encoding apparatus, and PROD_B represents the receiving device equipped with the motion picture decoding apparatus.
[0019] Figure 3 This diagram illustrates the configuration of a recording apparatus equipped with a motion picture encoding device according to this embodiment and a playback apparatus equipped with a motion picture decoding device. PROD_C represents the recording apparatus equipped with the motion picture encoding device, and PROD_D represents the playback apparatus equipped with the motion picture decoding device.
[0020] Figure 4 It is a diagram representing the hierarchical structure of the encoded stream data.
[0021] Figure 5 This is a diagram representing a segmentation example of CTU.
[0022] Figure 6 It is a schematic diagram representing the types (mode numbers) of intra-frame prediction modes.
[0023] Figure 7 This is a schematic diagram showing the configuration of a motion picture decoding device.
[0024] Figure 8 This is a flowchart illustrating the general operation of a motion picture decoding device.
[0025] Figure 9 This is a schematic diagram showing the structure of the intra-frame prediction parameter decoding unit.
[0026] Figure 10 It is a map representing the reference region used for intra-frame prediction.
[0027] Figure 11 This is a diagram showing the structure of the intra-frame prediction image generation unit.
[0028] Figure 12 This is a functional block diagram representing an example of the configuration of the inverse quantization / inverse transform unit.
[0029] Figure 13 This diagram illustrates the range of the quadratic transformation.
[0030] Figure 14 This diagram illustrates the operation of implicit MTS when using intra-sub-segmentation mode (intra-sub-segmentation prediction).
[0031] Figure 15 This diagram illustrates the operation of implicit MTS when using subblock transformations.
[0032] Figure 16 This is a block diagram illustrating the structure of a motion picture encoding device.
[0033] Figure 17 This is a schematic diagram showing the structure of the intra-frame prediction parameter coding unit.
[0034] Figure 18 This is a block diagram illustrating the nuclear conversion unit 1521.
[0035] Figure 19 This is a flowchart illustrating the quadratic transformation and the kernel transformation.
[0036] Figure 20 This is a flowchart explaining the operation of the MTS setting unit 15211 in the embodiment.
[0037] Figure 21 This is a flowchart explaining the operation of the MTS setting unit 15211 in the embodiment.
[0038] Figure 22 This is a flowchart explaining the operation of the MTS setting unit 15211 in the embodiment.
[0039] Figure 23 This is a block diagram showing the relationship between the TU decoding unit and the inverse transform unit.
[0040] Figure 24 This is a flowchart illustrating the process of a quadratic transformation. Detailed Implementation
[0041] [Implementation Method 1]
[0042] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.
[0043] Figure 1 This is a schematic diagram showing the configuration of the image transmission system 1 of this embodiment.
[0044] Image transmission system 1 is a system for transmitting an encoded stream obtained by encoding an image of an encoded object, decoding the transmitted encoded stream, and displaying the image. Image transmission system 1 is configured to include: a moving picture encoding device (image encoding device) 11, a network 21, a moving picture decoding device (image decoding device) 31, and an image display device (image display device) 41.
[0045] The motion picture encoding device 11 is input to the image T.
[0046] Network 21 transmits the encoded stream Te generated by the motion picture encoding device 11 to the motion picture decoding device 31. Network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. Network 21 is not necessarily limited to a two-way communication network; it can also be a one-way communication network transmitting broadcast waves such as terrestrial digital broadcasts and satellite broadcasts. Furthermore, network 21 can also be replaced by a storage medium containing the encoded stream Te, such as a DVD (Digital Versatile Disc) or a Blu-ray Disc (Blu-ray Disc).
[0047] The motion picture decoding device 31 decodes the encoded stream Te transmitted by the network 21 to generate one or more decoded images Td.
[0048] The image display device 41 displays all or part of one or more decoded images Td generated by the motion picture decoding device 31. The image display device 41 may include, for example, a liquid crystal display (LCD) or an organic EL (electroluminescence) display. Examples of display types include fixed, mobile, and HMD (Head-Down Display). Furthermore, when the motion picture decoding device 31 has high processing power, it displays high-quality images; when it has only low processing power, it displays images that do not require high processing power or high display capabilities.
[0049] <operator>
[0050] The operators used in this specification are described below.
[0051] >> is for right shift, << is for left shift, & is for bitwise AND, | is for bitwise OR, |= is for OR assignment operator, and || represents logical OR.
[0052] x? y : z is a ternary operator that takes y when x is true (other than 0) and takes z when x is false (0).
[0053] Clip3(a, b, c) is a function that clips c to a value between a and b (where a <= b). It returns a when c < a, b when c > b, and c otherwise.
[0054] abs(a) is a function that returns the absolute value of a.
[0055] Int(a) is a function that returns the integer value of a.
[0056] floor(a) is a function that returns the largest integer less than or equal to a.
[0057] ceil(a) is a function that returns the smallest integer greater than or equal to a.
[0058] a / d means a divided by d (truncating the decimal part).
[0059] <Structure of the coded stream Te>
[0060] Before explaining the motion picture encoding device 11 and the motion picture decoding device 31 of this embodiment in detail, the data structure of the coded stream Te generated by the motion picture encoding device 11 and decoded by the motion picture decoding device 31 is explained.
[0061] Figure 4 A diagram showing the hierarchical structure of the data in the coded stream Te. The coded stream Te exemplarily includes a sequence and a plurality of pictures constituting the sequence. Figure 4 Diagrams respectively showing the coded video sequence representing a given sequence SEQ, the coded picture representing a specified picture PICT, the coded slice representing a specified slice S, the coded slice data representing the specified slice data, the coding tree units included in the coded slice data, and the coding units included in the coding tree units are shown.
[0062] (Coded video sequence)
[0063] In the coded video sequence, a set of data for the motion picture decoding device 31 to refer to for decoding the sequence SEQ to be processed is specified. As Figure 4As shown in the encoded video sequence, the sequence SEQ includes: Video Parameter Set, Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture PCT, and Supplemental Enhancement Information (SEI).
[0064] The Video Parameter Set (VPS) specifies a set of encoding parameters shared by multiple images in an image composed of multiple layers, as well as a set of encoding parameters associated with the multiple layers included in the image.
[0065] The Sequence Parameter Set (SPS) specifies a set of encoding parameters that the moving image decoding device 31 refers to in order to decode the object sequence. For example, it specifies the width and height of the image. It should be noted that multiple SPSs can exist. In this case, any one of the multiple SPSs is selected from the PPS.
[0066] The Picture Parameter Set (PPS) specifies a set of encoding parameters that the moving image decoding device 31 refers to for decoding each picture in the object sequence. These parameters include, for example, a reference value for the quantization width used for picture decoding (pic_init__qp_minus26), a flag indicating the application of weighted prediction (weighted_pred_flag), and a scaling list (quantization matrix). It should be noted that multiple PPSs can exist. In this case, any one of the multiple PPSs is selected from the pictures in the object sequence.
[0067] (Encoded image)
[0068] In the encoded image, a set of data is specified for the motion picture decoding device 31 to refer to in order to decode the image PICT of the object being processed. For example... Figure 4 As shown in the encoded image, the image PICT includes slices 0 to NS-1 (NS is the total number of slices included in the image PICT).
[0069] It should be noted that, in the following descriptions, where there is no need to distinguish between slices 0 to NS-1, the code indices may sometimes be omitted. Furthermore, the same applies to the data included in the encoded stream Te described below, i.e., other data marked with indices.
[0070] (Encoded slice)
[0071] In the encoded slice, a set of data is specified for the motion picture decoding device 31 to refer to in order to decode the slice S of the object being processed. For example... Figure 4As shown in the encoded slice, the slice includes a slice header and slice data.
[0072] The slice header includes a set of encoded parameters for the moving image decoding device 31 to refer to in order to determine the decoding method for the object slice. The slice type specification information (slice_type) is an example of the encoded parameters included in the slice header.
[0073] As slice types that can be specified by the slice type specification information, the following can be listed: (1) I slices that use only intra-frame prediction during encoding, (2) P slices that use unidirectional prediction or intra-frame prediction during encoding, and (3) B slices that use unidirectional prediction, bidirectional prediction, or intra-frame prediction during encoding. It should be noted that inter-frame prediction is not limited to unidirectional or bidirectional prediction, and more reference images can be used to generate the prediction image. Hereinafter, the cases referred to as P and B slices refer to slices that include blocks that can use inter-frame prediction.
[0074] It should be noted that the slice header may also include a reference to the image parameter set (PPS) (pic_parameter_set_id).
[0075] (Encoded slice data)
[0076] In the encoded slice data, a set of data is specified for the motion picture decoding device 31 to refer to in order to decode the slice data of the object being processed. For example... Figure 4 As shown in the encoded slice header, the slice data includes CTUs. A CTU is a fixed-size (e.g., 64×64) block that makes up a slice, also known as the Largest Coding Unit (LCU).
[0077] (Coding Tree Unit)
[0078] exist Figure 4 Within the coding tree unit, a set of data is defined for the motion picture decoding device 31 to refer to in order to decode the CTU of the processing object. The CTU is divided into coding units CU, which serve as the basic unit of coding processing, through recursive quadtree partitioning (QT (QuadTree) partitioning), binary tree partitioning (BT (Binary Tree) partitioning), or ternary tree partitioning (TT (Ternary Tree) partitioning). BT partitioning and TT partitioning are collectively referred to as multi-tree partitioning (MT (Multi Tree) partitioning). The nodes of the tree structure obtained through recursive quadtree partitioning are called coding nodes. The intermediate nodes of quadtrees, binary trees, and ternary trees are coding nodes, and the CTU itself is defined as the top-level coding node.
[0079] CT includes the following information as CT information: a QT segmentation flag (cu_split_flag) indicating whether QT segmentation is performed, an MT segmentation flag (split_mt_flag) indicating whether MT segmentation is performed, an MT segmentation direction (split_mt_dir) indicating the segmentation direction of MT segmentation, and an MT segmentation type (split_mt_type) indicating the segmentation type of MT segmentation. cu_split_flag, split_mt_flag, split_mt_dir, and split_mt_type are transmitted per encoding node.
[0080] When cu_split_flag is 1, the encoding node is split into 4 encoding nodes. Figure 5 (QT).
[0081] When cu_split_flag is 0, and split_mt_flag is also 0, the encoding nodes are not split, and one CU is retained as a node. Figure 5 (Without further segmentation). CU is the terminal node of the encoding node and is not further segmented. CU is the basic unit of encoding processing.
[0082] With split_mt_flag set to 1, the encoded node is split into two MT nodes as shown below. With split_mt_type set to 0 and split_mt_dir set to 1, the encoded node is horizontally split into two encoded nodes. Figure 5 The BT (horizontal split) method, when split_mt_dir is 0, vertically splits the encoding node into 2 encoding nodes. Figure 5 BT (vertical splitting). Furthermore, when split_mt_type is 1 and split_mt_dir is 1, the encoding node is horizontally split into 3 encoding nodes. Figure 5 The TT (horizontal split) method, when split_mt_dir is 0, vertically splits the encoding node into 3 encoding nodes. Figure 5 The TT (vertical segmentation)). They are in Figure 5 The CT information is shown.
[0083] Furthermore, when the CTU size is 64×64 pixels, the CU size can be any of the following: 64×64 pixels, 64×32 pixels, 32×64 pixels, 32×32 pixels, 64×16 pixels, 16×64 pixels, 32×16 pixels, 16×32 pixels, 16×16 pixels, 64×8 pixels, 8×64 pixels, 32×8 pixels, 8×32 pixels, 16×8 pixels, 8×16 pixels, 8×8 pixels, 64×4 pixels, 4×64 pixels, 32×4 pixels, 4×32 pixels, 16×4 pixels, 4×16 pixels, 8×4 pixels, 4×8 pixels, and 4×4 pixels.
[0084] (Encoding unit)
[0085] like Figure 4 As shown in the encoding unit, a set of data is specified for the motion picture decoding device 31 to refer to in order to decode the encoding unit of the object being processed. Specifically, the CU consists of a CU header CUH, prediction parameters, transform parameters, quantization transform coefficients, etc. The prediction mode, etc., are specified in the CU header.
[0086] Predictive processing can be performed on a per-CU (Combined Unit) basis or on a per-sub-CU basis, where the CU is further divided into sub-CUs. When the size of the CU and its sub-CUs are equal, there is one sub-CU within the CU. When the size of the CU is larger than the size of its sub-CUs, the CU is divided into sub-CUs. For example, if the CU is 8×8 and the sub-CUs are 4×4, the CU is divided into four sub-CUs, each consisting of two horizontally divided parts and two vertically divided parts.
[0087] There are two types of prediction (prediction modes): intra-frame prediction and inter-frame prediction. Intra-frame prediction is prediction within the same image, while inter-frame prediction refers to prediction processing performed between different images (such as between display times or between layers).
[0088] Transform / quantization processing is performed on a unit basis (CU), but quantization transform coefficients can also be entropy encoded on a unit basis (4×4 sub-blocks).
[0089] (Prediction parameters)
[0090] The predicted image is derived from the prediction parameters appended to the block. These prediction parameters include those for intra-frame and inter-frame prediction.
[0091] The following explains the prediction parameters for intra-frame prediction. The intra-frame prediction parameters consist of the luminance prediction mode (IntraPredModeY) and the chrominance prediction mode (IntraPredModeC). Figure 6 This is a schematic diagram representing the types (mode numbers) of intra-frame prediction modes. For example... Figure 6As shown, there are, for example, 67 intra-frame prediction modes (0-66). These include, for example, planar prediction (0), DC prediction (1), and Angular prediction (2-66). Furthermore, LM modes (67-72) can be added to the chromatic aberration.
[0092] Syntax elements used to derive intra-prediction parameters include, for example, intra_luma_mpm_flag, intra_luma_mpm_idx, and intra_luma_mpm_remainder.
[0093] (MPM)
[0094] `intra_luma_mpm_flag` is a flag indicating whether the IntraPredModeY of the object block matches the MPM (Most ProbableMode). MPM is the prediction mode included in the MPM candidate list `mpmCandList[]`. The MPM candidate list stores a list of candidates with high probabilities of being applied to the object block based on the intra prediction modes of adjacent blocks and the specified intra prediction mode. When `intra_luma_mpm_flag` is 1, the IntraPredModeY of the object block is derived using the MPM candidate list and the index `intra_luma_mpm_idx`.
[0095] IntraPredModeY=mpmCandList [intra_luma_mpm_idx](REM)
[0096] With `intra_luma_mpm_flag` set to 0, the intra-prediction mode is selected from the remaining modes `RemIntraPredMode` after removing the intra-prediction modes included in the MPM candidate list from all intra-prediction modes. The selectable intra-prediction mode is referred to as "Non-MPM" or "REM". `RemIntraPredMode` is exported using `intra_luma_mpm_remainder`.
[0097] (Composition of a motion picture decoding device)
[0098] The motion picture decoding device 31 of this embodiment ( Figure 7 The composition of ) will be explained.
[0099] The motion picture decoding device 31 is configured to include: an entropy decoding unit 301, a parameter decoding unit (predictive image decoding device) 302, a loop filter 305, a reference image memory 306, a prediction parameter memory 307, a prediction image generation unit (predictive image generation device) 308, an inverse quantization / inverse transform unit 311, and an adder unit 312. It should be noted that, according to the motion picture encoding device 11 described later, there is also a configuration where the motion picture decoding device 31 does not include the loop filter 305.
[0100] The parameter decoding unit 302 also includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit). The CU decoding unit 3022 includes a TU decoding unit 3024. These can also be collectively referred to as decoding modules. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS, and the slice header (slice information) from the encoded data. The CT information decoding unit 3021 decodes the CT from the encoded data. The CU decoding unit 3022 decodes the CU from the encoded data. The TU decoding unit 3024, when the TU includes prediction errors, decodes QP update information (quantization correction value) and quantization prediction errors (residual_coding) from the encoded data.
[0101] Figure 23 This is a block diagram illustrating the relationship between the TU decoding unit 3024 and the inverse transform unit 3112. The stIdx decoding unit 131 of the TU decoding unit 3024 decodes the value stIdx from the encoded data, indicating whether a secondary transform is used and the transform basis, and outputs it to the secondary transform unit 31121. The mts_idx decoding unit 132 of the TU decoding unit 3024 decodes the value mts_idx from the encoded data, indicating the transform matrix of the MTS, and outputs it to the kernel transform unit 31123. Specifically, the TU decoding unit 3024 decodes stIdx when the width and height of the CU are 4 or more, the prediction mode is intra-frame mode, and the number of transform coefficients numSigCoeff within the CU is greater than a predetermined number THSt (e.g., 2 in SINGLE_TREE, and 1 otherwise). It should be noted that when stIdx is 0, no quadratic transformation is applied; when stIdx is 1, it represents the transformation of one side of the set (pair) of quadratic transformation matrices; and when stIdx is 2, it represents the transformation of the other side of the pair. Furthermore, the quadratic transformation matrix secTransMatrix can be selected not only based on the value of stIdx, but also based on the intra-frame prediction mode and the size of the transformation.
[0102] Furthermore, the parameter decoding unit 302 is configured to include an inter-frame prediction parameter decoding unit 303 (not shown) and an intra-frame prediction parameter decoding unit 304. The prediction image generation unit 308 is configured to include an inter-frame prediction image generation unit 309 and an intra-frame prediction image generation unit 310.
[0103] Furthermore, examples of using CTU and CU as processing units are described below, but this is not the only approach; processing can also be performed on a sub-CU basis. Alternatively, processing can be performed on a block or sub-block basis by replacing CTU and CU with blocks and sub-CUs with sub-blocks.
[0104] The entropy decoding unit 301 performs entropy decoding on the externally input encoded stream Te, separating and decoding each code (syntactic element). Entropy coding can be performed in two ways: using a context (probability model) appropriately selected based on the type of syntactic element and its surrounding conditions to perform variable-length encoding of syntactic elements; and using a predetermined table or formula to perform variable-length encoding of syntactic elements. The former, CABAC (Context Adaptive Binary Arithmetic Coding), stores a probability model updated for each encoded or decoded image (slice) in memory. Then, as the initial state of the context for image P or image B, a probability model for images using the same slice type and the same slice level quantization parameters is set based on the probability model stored in memory. This initial state is used for encoding and decoding processes. The separated code contains prediction information for generating the predicted image and prediction errors for generating the difference image.
[0105] The entropy decoding unit 301 outputs the separated code to the parameter decoding unit 302. The separated code, for example, refers to the prediction mode predMode. The parameter decoding unit 302 controls which code to decode.
[0106] (Basic Process)
[0107] Figure 8 This is a flowchart explaining the general operation of the motion picture decoding device 31.
[0108] (S1100: Parameter Set Information Decoding) The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS from the encoded data.
[0109] (S1200: Slice Information Decoding) The header decoding unit 3020 decodes the slice header (slice information) from the encoded data.
[0110] Hereinafter, the motion picture decoding device 31 exports the decoded image of each CTU by repeatedly performing the processing steps S1300 to S5000 on each CTU included in the object picture.
[0111] (S1300: CTU Information Decoding) The CT information decoding unit 3021 decodes the CTU from the encoded data.
[0112] (S1400: CT Information Decoding) The CT information decoding unit 3021 decodes the CT from the encoded data.
[0113] (S1500: CU Decoding) The CU decoding unit 3022 implements S1510 and S1520 to decode the CU from the encoded data.
[0114] (S1510: CU Information Decoding) The CU decoding unit 3022 decodes CU information, prediction information, TU segmentation flag split_transform_flag, CU residual flags cbf_cb, cbf_cr, cbf_luma, etc. from the encoded data.
[0115] (S1520: TU Information Decoding) When the TU includes prediction error, the TU decoding unit 3024 decodes the QP update information (quantization correction value) and the quantization prediction error (residual_coding) from the encoded data. It should be noted that the QP update information is the difference between the quantization parameter prediction value qPpred and the prediction value of the quantization parameter QP.
[0116] (S2000: Predictive Image Generation) The predictive image generation unit 308 generates a predictive image for each block included in the object CU based on the prediction information.
[0117] (S3000: Inverse quantization / inverse transformation) The inverse quantization / inverse transformation unit 311 performs inverse quantization / inverse transformation processing on each TU included in the target CU.
[0118] (S4000: Decoded Image Generation) The addition unit 312 generates a decoded image of the object CU by adding the predicted image provided by the predicted image generation unit 308 to the prediction error provided by the inverse quantization / inverse transform unit 311.
[0119] (S5000: Loop Filter) The loop filter 305 applies loop filters such as deblocking filter, SAO, and ALF to the decoded image to generate the decoded image.
[0120] Furthermore, the parameter decoding unit 302 is configured to include an inter-frame prediction parameter decoding unit 303 (not shown) and an intra-frame prediction parameter decoding unit 304. The prediction image generation unit 308 is configured to include an inter-frame prediction image generation unit 309 (not shown) and an intra-frame prediction image generation unit 310.
[0121] (Configuration of the intra-frame prediction parameter decoding unit 304)
[0122] The intra-prediction parameter decoding unit 304 decodes the intra-prediction parameters based on the code input from the entropy decoding unit 301 and with reference to the prediction parameters stored in the prediction parameter memory 307, for example, decoding the intra-prediction mode IntraPredMode. The intra-prediction parameter decoding unit 304 outputs the decoded intra-prediction parameters to the prediction image generation unit 308 and then stores them in the prediction parameter memory 307. The intra-prediction parameter decoding unit 304 can also derive intra-prediction modes that differ in brightness and chromatic difference.
[0123] Figure 9 This is a schematic diagram showing the configuration of the intra-frame prediction parameter decoding unit 304 of the parameter decoding unit 302. (See diagram below.) Figure 9 As shown, the intra-frame prediction parameter decoding unit 304 is configured to include: a parameter decoding control unit 3041, a luminance intra-frame prediction parameter decoding unit 3042, and a chrominance intra-frame prediction parameter decoding unit 3043.
[0124] The parameter decoding control unit 3041 instructs the entropy decoding unit 301 to decode the syntax elements and receives the syntax elements from the entropy decoding unit 301. When intra_luma_mpm_flag is 1, the parameter decoding control unit 3041 outputs intra_luma_mpm_idx to the MPM parameter decoding unit 30422 within the luma intra-frame prediction parameter decoding unit 3042. Furthermore, when intra_luma_mpm_flag is 0, the parameter decoding control unit 3041 outputs intra_luma_mpm_remainder to the non-MPM parameter decoding unit 30423 of the luma intra-frame prediction parameter decoding unit 3042. Additionally, the parameter decoding control unit 3041 outputs the syntax elements of the intra-frame prediction parameters for chroma to the chroma intra-frame prediction parameter decoding unit 3043.
[0125] The luminance intra-frame prediction parameter decoding unit 3042 is configured to include: an MPM candidate list derivation unit 30421, an MPM parameter decoding unit 30422, and a non-MPM parameter decoding unit 30423 (decoding unit and derivation unit).
[0126] The MPM parameter decoding unit 30422 refers to the mpmCandList[] and intra_luma_mpm_idx exported by the MPM candidate list exporting unit 30421 to export IntraPredModeY, and outputs it to the intra-frame prediction image generation unit 310.
[0127] The non-MPM parameter decoding unit 30423 derives RemIntraPredMode from mpmCandList[] and intra_luma_mpm_remainder, and outputs IntraPredModeY to the intra-frame prediction image generation unit 310.
[0128] The chromatic difference intra-prediction parameter decoding unit 3043 derives IntraPredModeC based on the syntax elements of the chromatic difference intra-prediction parameters and outputs it to the intra-prediction image generation unit 310.
[0129] The lumen intra-prediction parameter decoding unit 3042 can also decode the intra_subpartitions_mode_flag flag, which indicates whether intra-prediction is performed by dividing the CU into smaller sub-blocks. If intra_subpartitions_mode_flag is not 0, intra_subpartitions_split_flag is further decoded. The intra-subpartition mode is derived using the following formula.
[0130] = IntraSubPartSplitType = (intra_subpartitions_mode_flag == 0) ? 0: 1 + intra_subpartitions_split_flag. When IntraSubPartSplitType is 0 (ISP_NO_SPLIT), intra-frame prediction is performed without further CU segmentation. When IntraSubPartSplitType is 1 (ISP_HOR_SPLIT: horizontal segmentation), the CU is vertically segmented from two blocks into four sub-blocks, and intra-frame prediction, transform coefficient decoding, and inverse quantization / inverse transform are performed on a sub-block basis. When IntraSubPartSplitType is 2 (ISP_VER_SPLIT: vertical segmentation), the CU is horizontally segmented from two blocks into four sub-blocks, and intra-frame prediction, transform coefficient decoding, and inverse quantization / inverse transform are performed on a sub-block basis. The number of sub-blocks, NumIntraSubPart, is derived using the following formula.
[0131] NumIntraSubPart==(cbWidth==4&&cbHeight==8)||(cbWidth==8&&cbHeight==4)? 2:4
[0132] The width nW and height nH of the sub-block, the number of horizontal and vertical divisions numPartsX, and numPartY are derived as follows.
[0133] nW=(IntraSubPartSplitType==ISP_VER_SPLIT?)nTbW / NumIntraSubPart:nTbW
[0134] nH=(IntraSubPartSplitType==ISP_HOR_SPLIT?)nTbH / NumIntraSubPart:nTbH
[0135] numPartsX=(IntraSubPartSplitType==ISP_VER_SPLIT?)NumIntraSubPart:1
[0136] numPartsY=(IntraSubPartSplitType==ISP_HOR_SPLIT?)NumIntraSubPart:1
[0137] Here, nTbW and nTbH are the width and height of CU (or TU).
[0138] The loop filter 305 is a filter located within the encoding loop, used to remove block distortion and ringing distortion to improve image quality. The loop filter 305 performs deblocking filtering, sample adaptive offset (SAO), and adaptive loop filtering (ALF) on the decoded image of the CU generated by the adder 312.
[0139] The image memory 306 stores the decoded images of the CU generated by the addition unit 312 in predetermined locations according to each object image and object CU.
[0140] The prediction parameter memory 307 stores the prediction parameters in predetermined locations according to the CTU or CU of each decoded object. Specifically, the prediction parameter memory 307 stores the parameters decoded by the parameter decoding unit 302 and the predMode separated by the entropy decoding unit 301, etc.
[0141] The prediction image generation unit 308 receives inputs such as predMode and prediction parameters. Furthermore, the prediction image generation unit 308 reads a reference image from the reference image memory 306. Under the prediction mode indicated by predMode, the prediction image generation unit 308 uses the prediction parameters and the read-out reference image (reference image block) to generate a prediction image of a block or sub-block. Here, a reference image block refers to a set of pixels (usually rectangular, hence called a block) on the reference image, which is the area referenced for generating the prediction image.
[0142] (Intra-frame prediction image generation unit 310)
[0143] When predMode indicates intra-prediction mode, the intra-prediction image generation unit 310 uses intra-prediction parameters input from the intra-prediction parameter decoding unit 304 and reference pixels read from the reference image memory 306 to perform intra-prediction.
[0144] Specifically, the intra-frame prediction image generation unit 310 reads adjacent blocks on the object image from the reference image memory 306, which are within a predetermined range from the object block. The predetermined range consists of adjacent blocks to the left, upper left, upper, and upper right of the object block, and varies depending on the region referenced by the intra-frame prediction mode.
[0145] The intra-prediction image generation unit 310 generates a prediction image of the target block by referring to the read-out decoded pixel values and the prediction mode represented by IntraPredMode. The intra-prediction image generation unit 310 outputs the generated prediction image of the block to the addition unit 312.
[0146] The following describes the generation of predicted images based on intra-frame prediction modes. In Planar prediction, DC prediction, and Angular prediction, the decoded surrounding region adjacent to (close to) the predicted object block is set as a reference region R. Then, the predicted image is generated by extrapolating pixels on the reference region R in a specific direction. For example, the reference region R can be set as an L-shaped region including the left and top (or further top-left, top-right, and bottom-left) of the predicted object block (e.g., by...). Figure 10 (The area represented by the pixel of the circular marker with a diagonal line in Example 1 of the reference area).
[0147] (Details of the predictive image generation department)
[0148] Next, use Figure 11 The detailed configuration of the intra-frame prediction image generation unit 310 will be described below. The intra-frame prediction image generation unit 310 includes: a prediction target block setting unit 3101, an unfiltered reference image setting unit 3102 (first reference image setting unit), a filtered reference image setting unit 3103 (second reference image setting unit), an intra-frame prediction unit 3104, and a prediction image correction unit 3105 (prediction image correction unit, filter switching unit, weighting coefficient changing unit).
[0149] The intra-frame prediction unit 3104 generates a temporary prediction image (pre-correction prediction image) of the prediction target block based on each reference pixel (unfiltered reference image) on the application reference region R, the filtered reference image generated by the reference pixel filter (first filter), and the intra-frame prediction mode, and outputs it to the prediction image correction unit 3105. The prediction image correction unit 3105 corrects the temporary prediction image according to the intra-frame prediction mode, generates a prediction image (corrected prediction image), and outputs it.
[0150] The following describes the components included in the intra-frame prediction image generation unit 310.
[0151] (Prediction Object Block Setting Unit 3101)
[0152] The prediction object block setting unit 3101 sets the object CU as the prediction object block and outputs information related to the prediction object block (prediction object block information). The prediction object block information includes at least the size, position, and index representing the brightness or color difference of the prediction object block.
[0153] (Unfiltered reference image setting unit 3102)
[0154] The unfiltered reference image setting unit 3102 sets the adjacent surrounding area of the predicted object block as a reference region R based on the size and position of the predicted object block. Then, it sets the corresponding decoded pixel values (unfiltered reference image, boundary pixels) in the reference image memory 306 for each pixel value in the reference region R. Figure 10 In Example 1 of the reference region, the row r[x][-1] of the decoded pixels adjacent to the top of the predicted object block and the column r[-1][v] of the decoded pixels adjacent to the left of the predicted object block are unfiltered reference images.
[0155] (Filtered reference image setting unit 3103)
[0156] The filtered reference image setting unit 3103 applies a reference pixel filter (first filter) to the unfiltered reference image according to the intra-frame prediction mode, and derives filtered reference images s[x][y] for each position (x, y) on the reference region R. Specifically, a low-pass filter is applied to the unfiltered reference image at position (x, y) and its surrounding unfiltered reference image to derive filtered reference images (s[x][y]). Figure 10 Example 2 (reference area). It should be noted that a low-pass filter is not necessarily applied to all intra-frame prediction modes; it can be applied to a portion of the intra-frame prediction modes. It should be noted that the filter applied to the unfiltered reference image on the reference area R in the filtered reference image setting unit 3103 is called the "reference pixel filter (first filter)". In contrast, the filter used to correct the temporary prediction image in the prediction image correction unit 3105, which will be described later, is called the "boundary filter (second filter)".
[0157] (Configuration of the intra-frame prediction unit 3104)
[0158] The intra-prediction unit 3104 generates a temporary prediction image (temporary prediction pixel values, pre-correction prediction image) of the prediction target block based on the intra-prediction mode, the unfiltered reference image, and the filtered reference pixel values, and outputs it to the prediction image correction unit 3105. The intra-prediction unit 3104 internally includes a Planar prediction unit 31041, a DC prediction unit 31042, an Angular prediction unit 31043, and an LM prediction unit 31044. The intra-prediction unit 3104 selects a specific prediction unit according to the intra-prediction mode and inputs the unfiltered reference image and the filtered reference image. The relationship between the intra-prediction mode and the corresponding prediction unit is shown below.
[0159] •Planar Prediction•··Planar Prediction Department 31041
[0160] DC Forecasting... DC Forecasting Department 31042
[0161] ·Angular Forecasting Department···Angular Forecasting Department
[0162] Test 31043
[0163] ·LM Forecast···LM Forecast Department 31044
[0164] (Planar prediction)
[0165] The Planar prediction unit 31041 generates a temporary prediction image by linearly adding multiple filtered reference images based on the distance between the predicted object pixel position and the reference pixel position, and outputs it to the prediction image correction unit 3105.
[0166] (DC forecast)
[0167] The DC prediction unit 31042 derives a DC prediction value equivalent to the average value of the filtered reference image s[x][v], and outputs a temporary prediction image q[x][y] with the DC prediction value as the pixel value.
[0168] (Angular prediction)
[0169] The Angular prediction unit 31043 generates a temporary prediction image q[x][y] using the filtered reference image s[x][y] in the prediction direction (reference direction) indicated by the intra-frame prediction mode, and outputs it to the prediction image correction unit 3105.
[0170] (LM Prediction)
[0171] The LM prediction unit 31044 predicts the pixel values of the color difference based on the pixel values of the brightness. Specifically, it is a method of generating a predicted image of the color difference (Cb, Cr) using a linear model based on the decoded brightness image. CCLM (Cross-Component Linear Model prediction), one type of LM prediction, is a prediction method that uses a linear model to predict the color difference based on the brightness for a block.
[0172] (Configuration of the predictive image correction unit 3105)
[0173] The prediction image correction unit 3105 corrects the temporary prediction image output from the intra-prediction unit 3104 according to the intra-prediction mode. Specifically, the prediction image correction unit 3105 performs a weighted sum (weighted average) of the unfiltered reference image and the temporary prediction image for each pixel of the temporary prediction image based on the distance between the reference region R and the target prediction pixel, thereby deriving the corrected prediction image (corrected prediction image) Pred. It should be noted that in some intra-prediction modes, the temporary prediction image can be directly used as the prediction image without correcting it through the prediction image correction unit 3105.
[0174] (Inverse quantization / inverse transform unit 311)
[0175] The inverse quantization / inverse transform unit 311 inversely quantizes the quantization transform coefficients qd[][] input from the entropy decoding unit 301 to obtain the transform coefficients d[][]. These quantization transform coefficients qd[][] are obtained by performing frequency transformations such as DCT (Discrete Cosine Transform) and DST (Discrete Sine Transform) on the prediction error during the encoding process and then quantizing it. The inverse quantization / inverse transform unit 311 performs inverse DCT, inverse DST, and other inverse frequency transformations on the obtained transform coefficients to calculate the prediction error. The inverse quantization / inverse transform unit 311 outputs the prediction error to the adder unit 312.
[0176] The following is for reference Figure 12 An example of the configuration of the inverse quantization / inverse transformation unit 311 will be explained. Figure 12 This is a functional block diagram illustrating an example of the configuration of the inverse quantization / inverse transform unit 311. For example... Figure 12 As shown, the quantization / inverse transform unit 311 includes an inverse quantization unit 3111 and a transform unit 3112. The inverse quantization unit 3111 performs inverse quantization on the quantization transform coefficients qd[][] obtained by decoding in the TU decoding unit 3024 to derive transform coefficients d[][]. The inverse quantization unit 3111 outputs the derived transform coefficients d[][] to the transform unit 3112.
[0177] The transformation unit 3112 performs an inverse transformation on the received transformation coefficients d[][] for each transformation unit TU to restore the prediction error r[][]. The transformation unit 3112 outputs the restored prediction error r[][] to the addition unit 312.
[0178] It should be noted that in this specification, the process of transforming the differential image in the image encoding device into transform coefficients is called forward transform, and the process of transforming the transform coefficients in the image decoding device into a differential image is called transform. However, they can also be referred to as transform and inverse transform, respectively. It should be noted that there is no difference between forward transform (transform) and transform (inverse transform) in terms of processing other than the value of the transform matrix that serves as the transform basis. Therefore, in the following description, the term "inverse transform" can be used instead of "transform" for the transform processing in the transform unit 3112.
[0179] The transformation unit 3112 includes a secondary transformation unit (second transformation unit) 31121 and a core transformation unit (first transformation unit) 31123.
[0180] The TU decoding unit 3024 can decode the sub-block transform flag cu_sbt_flag, which indicates whether the transform coefficients of only one sub-block are decoded, inverse quantized, and inverse transformed by further dividing the CU into multiple sub-blocks. When cu_sbt_flag is 1, the flag cu_sbt_quad_flag, indicating whether the CU is divided into four sub-blocks, can also be decoded. When cu_sbt_quad_flag is 0, the number of sub-blocks is 2. When cu_sbt_quad_flag is 1, the number of sub-blocks is 4. Furthermore, the flag cu_sbt_horizontal_flag, indicating whether the CU is divided horizontally or vertically, is decoded. Additionally, the flag cu_sbt_pos_flag, indicating which sub-block includes the transform coefficients, is decoded.
[0181] (Scaling section 31112)
[0182] The scaling unit 31112 scales the transform coefficients decoded by the TU decoding unit using a weighted average of coefficient units.
[0183] The scaling unit 31112 scales using the following formula when transformation skip is valid (transform_skip == 1).
[0184] r[x][y] = d[x][y] << tsShift
[0185] Here, tsShift = 5 + ((log2(nTbW) + log2(nTbH)) / 2).
[0186] In cases other than those described above, the quantization matrix m[x][y] and scaling factor ls[x][v] are derived using the following formulas.
[0187] ls[x][y]=(m[x][y]*levelScale[(qP+1)%6])<<(qP / 6)
[0188] Alternatively, it can be derived using the following formula.
[0189] 1s[x][y]=(m[x][y]*1evelScale[qP%6])<<(qP / 6)
[0190] Here, levelScale[] = {40, 45, 51, 57, 64, 72}.
[0191] It should be noted that the value of the quantization matrix m[x][y] can be decoded from the encoded data, and m[x][y] = 16 can be used as uniform quantization.
[0192] The scaling unit 31112 derives dnc[][] based on the product of the scaling factor 1s[][] and the decoded transform coefficients TransCoeffLevel, and then inversely quantizes it.
[0193] dnc[x][y]=(TransCoeffLevel[xTbY][yTbY][cIdx][x][y]*ls[x][y]*rectNorm+bdOffset)>>bdShift
[0194] Finally, the scaling part 31112 truncates the inverse-quantized transformation coefficients and derives d[x][y].
[0195] d[x][y]=Clip3(CoeffMin, CoeffMax, dnc[x][y])
[0196] d[x][y] is transmitted to the kernel transformation unit 31123 or the second transformation unit 31121. The second transformation unit (second transformation unit) 31121 applies a second transformation to the transformation coefficients d[][] after inverse quantization and before kernel transformation.
[0197] (Quadratic transformation and kernel transformation)
[0198] The secondary transformation unit 31121 applies a transformation matrix to a portion or all of the transformation coefficients d[][] received from the inverse quantization unit 3111, thereby restoring the corrected transformation coefficients (the transformed coefficients after the transformation performed by the second transformation unit) d[][]. The secondary transformation unit 31121 applies the secondary transformation to the transformation coefficients d[][] of a specified unit per transformation unit TU. The secondary transformation is applied only in the intra-frame CU, and the transformation basis is determined with reference to the intra-frame prediction mode IntraPredMode. The selection of the transformation basis will be described later. The secondary transformation unit 31121 outputs the restored corrected transformation coefficients d[][] to the kernel transformation unit 31123.
[0199] The kernel transformation unit 31123 acquires the transformation coefficients d[][] or the corrected transformation coefficients d[][] restored by the secondary transformation unit 31121, performs the transformation, and derives the prediction error r[][]. The kernel transformation unit 31123 outputs the prediction error r[][] to the adder unit 312.
[0200] (Quadratic transformation)
[0201] In the moving image coding apparatus 11, a further transformation (positive quadratic transformation) is applied to the transform coefficients after kernel transformations (DCT2 and DST7, etc.) of the differential image to remove residual correlations in the transform coefficients and concentrate energy on a portion of the transform coefficients. Figure 19 The diagram shows the forward transform unit 1032 included in the transform / quantization unit 103 of the motion picture encoding apparatus 11 and the inverse transform unit 152 included in the inverse transform / inverse quantization unit 105. In the motion picture decoding apparatus 3, conversely, a quadratic transform is applied to the transform coefficients of a portion or all regions of the decoded TU, and a kernel transform (DCT2 and DST7, etc.) is applied to the transform coefficients after the quadratic transform.
[0202] In the second transformation, the following processing is performed based on the TU size and intra-frame prediction mode. The processing of the second transformation will be explained in turn below. Figure 13 This is a diagram illustrating a quadratic transformation. In Figure 13 In the example, for an 8×8 TU, the following process is shown: through the process of S2, the transformation coefficients d[][] of the 4×4 region are stored in the one-dimensional array u[] of nonZeroSize. Through the process of S3, the one-dimensional array u[] is transformed into the one-dimensional array v[]. Finally, through the process of S4, they are stored in d[][] again.
[0203] Figure 24 This is a flowchart illustrating the process of a quadratic transformation.
[0204] (S1: Setting the transformation size and input / output size)
[0205] The secondary transformation unit 31121 derives the dimensions of the secondary transformation (4×4 or 8×8), the number of output transformation coefficients (nStOutSize), the number of applied transformation coefficients (input transformation coefficients) nonZeroSize, and the number of sub-blocks to which the secondary transformation is applied (numStX, numStY) based on the dimensions of the TU (width nTbW, height nTbH). The dimensions of the 4×4 and 8×8 secondary transformations are represented by nStSize = 4 and 8, respectively. Alternatively, the dimensions of the 4×4 and 8×8 secondary transformations can be referred to as RST4×4 and RST8×8, respectively.
[0206] When the TU is a specified size or larger, the secondary transformation unit 31121 outputs 48 transformation coefficients through an RST8×8 secondary transformation. Otherwise, it outputs 16 transformation coefficients through an RST4×4 secondary transformation. When the TU is 4×4, 16 transformation coefficients are derived from 8 transformation coefficients using RST4×4; when the TU is 8×8, 48 transformation coefficients are derived from 8 transformation coefficients using RST8×8. Otherwise, depending on the size of the TU, 16 or 48 transformation coefficients are output from 16 transformation coefficients.
[0207] When both nTbW and nTbH are greater than 8, log2StSize = 3 and nStOutSize = 48.
[0208] In cases other than those mentioned above, log2StSize = 2, nStOutSize = 16
[0209] nStSize = 1 <log2StSize
[0210] When both nTbW and nTbH are 4 or 8×8, nonZeroSize = 8. Otherwise, nonZeroSize = 16.
[0211] numStX=(nTbH==4&&nTbW>8)? 2:1
[0212] numStY=(nTbW==4&&nTbH>8)? 2:1
[0213] (S2: Rearranged into a one-dimensional signal)
[0214] The secondary transformation unit 31121 rearranges a portion of the transformation coefficients d[][] of TU into a one-dimensional array u[] for processing. Specifically, in the secondary transformation, u[] is derived based on the two-dimensional transformation coefficients d[][] of the object TU, with reference to the transformation coefficients of x = 0..nonZeroSize-1. xC and yC are the positions on TU, derived based on the arrangement DiagScanOrder representing the scan order and the position x of the transformation coefficients in the sub-block.
[0215]
[0216] (S3: Application of Transformation Processing)
[0217] The second transformation unit 31121 performs a transformation on u[] (vector F′) of length nonZeroSize using the first transformation basis (matrix) T, and derives a one-dimensional array v′[] (vector V′) of length nStOutSize as the output.
[0218] This transformation can be expressed by the following formula in matrix operations.
[0219] V′=T×F′
[0220] Here, the transformation basis for the case with a transformation size of 4×4 (RST4×4) is called the first transformation basis T1. The transformation basis for the case with a transformation size of 8×8 (RST8×8) is called the second transformation basis T2. T1 is a 16×16 (16 rows and 16 columns) matrix, that is, the transformation derives a 16×1 (16 rows and 1 column) vector V′, that is, a one-dimensional array v′[] of length 16, which is used as the product of the 16×16 matrix T1 and the 16×1 (16 rows and 1 column) vector F′. T2 is a 48×16 (48 rows and 16 columns) matrix, that is, the transformation derives a 48×1 (48 rows and 1 column, length 48) vector V′, that is, a one-dimensional array v′[] of length 48, which is used as the product of the 48×16 matrix T2 and the 16×1 vector F′.
[0221] Specifically, the secondary transform unit 31121 derives the set number (stTrSetId) of the secondary transform derived from the IntraPredMode based on the secondary transform size nStSize(nTrS), the transform basis stIdx representing the secondary transform decoded from the encoded data, and the corresponding transform matrix secTranMatrix[][] (transform basis T1 or T2). Furthermore, the secondary transform unit 31121 performs a product summation operation between the transform matrix and the one-dimensional array u[] as shown in the following formula.
[0222] v′[i]=C1ip3(CoeffMin, CoeffMax, ∑secTransMatrix[j][i]*u[j])
[0223] Here, ∑ is the sum from j = 0 to nonZeroSize - 1. In addition, i is processed for 0 to nStSize - 1. CoeffMin and CoeffMax represent the range of values of the transform coefficients.
[0224] (S4: Two-dimensional configuration of the one-dimensional signal after transformation processing)
[0225] The secondary transformation unit 31121 re-arranges the coefficients v′[] of the transformed one-dimensional array at the specified positions within the TU.
[0226] In process S4, the secondary transformation unit 31121 arranges the coefficients v′[] of length nStOutSize obtained through the above process S3 in the upper left region of the transform coefficient arrangement d[][].
[0227] The secondary transformation unit 31121 performs the following process for x = 0 to nStSize - 1 and y = 0 to nStSize - 1. Specifically, when IntraPredMode <= 34 or INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM, the secondary transformation unit 31121 applies the following formula.
[0228] d[(xSbIdx << log2StSize) + x][(ySbIdx << log2StSize) + y] =
[0229] (y < 4)? v[x + (y << log2StSize)] : ((x < 4)? v[32 + X + ((y - 4) << 2)] :
[0230] d[(xSbIdx << log2StSize) + x][(ySbIdx << log2StSize) + y])
[0231] In other cases, the secondary transformation unit 31121 applies the following formula.
[0232] d[(xSbIdx << log2StSize) + x][(ySbIdx << log2StSize) + y] =
[0233] (y < 4)? v[y + (x << log2StSize)] : ((x < 4)? v[32 + (y - 4) + (x << 2)] : d[(xSbIdx << log2StSize) + x][(ySbIdx << log2StSize) + y]) (Kernel transformation unit 31123)
[0234] Nuclear Transformation
[0235] A transformation that adaptively switches between transformation methods and switches through explicit flags, indices, and prediction modes is called a transformation (first transformation, kernel transformation). The transformation used in kernel transformation is a separate transformation consisting of a vertical transformation and a horizontal transformation. Alternatively, a transformation that separates two-dimensional signals in the horizontal and vertical directions can also be defined as a first transformation. Furthermore, in an image decoding device, a transformation applied after a second transformation can also be defined as a first transformation. The transformation basis (transformation matrix) of the kernel transformation is DCT2, DST7, and DCT8. In the kernel transformation, the transformation basis is switched independently for the vertical and horizontal transformations. It should be noted that the selectable transformations are not limited to those mentioned above, and other transformations (transformation bases) can also be used. It should be noted that DCT2, DST7, DCT8, DST1, and DCT5 will be represented as DCT-II, DST-VII, DCT-VIII, DST-I, and DCT-V, respectively. Furthermore, transformation skipping can also exist as a mode that explicitly skips the kernel transformation.
[0236] In kernel transform, there are explicit MTS and implicit MTS. In the case of explicit MTS, mts_idx is decoded from the encoded data, and the transform matrix is switched. In the case of implicit MTS, mts_idx is derived based on the intra-frame prediction mode and block size.
[0237] It should be noted that this embodiment describes an example of decoding mts_idx in CU or TU units, but the decoding (switching) unit is not limited to these.
[0238] `mts_idx` is the switching index used to select the transform basis for the kernel transform. `mts_idx` has any value of 0, 1, 2, 3, or 4, which derives the horizontal transform type `trTypeHor` and the vertical transform type `trTypeVer`.
[0239] use Figure 18 The nuclear transformation described above will be explained in detail. Figure 18 The nuclear transformation unit 1521 is Figure 12 Nuclear transformation unit 31123, Figure 19 An example of the nuclear transformation unit 1521. Figure 18The kernel transformation unit 1521 comprises: an MTS setting unit 15211, which sets the type of transformation to be used based on multiple transformation bases; a coefficient transformation processing unit 15212, which calculates the prediction residual r[][] based on the (corrected) transformation coefficients d[][] using the derived transformation; and a matrix transformation processing unit 15213, which performs the actual transformation. When no secondary transformation is performed, the corrected transformation coefficients are equal to the transformation coefficients. When a secondary transformation is performed, the corrected transformation coefficients take values different from the transformation coefficients. The MTS setting unit 15211 comprises an MTS setting unit 152111 that determines the derivation method of the index mts_ids of the transformation to be used, and an implicit MTS setting unit 152112 that implicitly derives mts_idx.
[0240] The MTS setting unit 152111 selects whether to perform explicit MTS, implicit MTS, or no MTS.
[0241] When the explicit MTS setting unit 152111 is enabled (when sps_explicit_mts_flag is 1), it uses explicit MTS and uses the mts_idx decoded from the encoded data in subsequent processing. The flag indicating whether explicit MTS is enabled (explicitMtsEnabled) can be set separately for intra-frame mode and inter-frame mode. In this case, explicit MTS can be determined to be enabled and mts_idx decoded from the encoded data when the prediction mode (PredMode) is inter-frame mode (other than MODE_INTRA) and sps_explicit_mts_inter_enabled_flag is 1, or when the prediction mode (PredMode) is intra-frame mode (MODE_INTRA) and sps_explicit_mts_intra_enabled_flag is 1. Furthermore, mts_idx can be decoded when both the width and height of the TU are 32 or less (nTbW <= 32 && nTbH <= 32). (implicitMTS flag setting)
[0242] When the MTS setting unit 152111 is active (sps_mts_enabled_flag == 1) and the explicit MTS flag is not active (explicitMtsEnabled == 0), the implicit MTS flag (implicitMtsEnabled) is set to 1. More specifically, the MTS setting unit 152111 sets implicitMtsEnabled to 1 if any of the following conditions are met, and sets implicitMtsEnabled to 0 otherwise.
[0243] • When intra-frame sub-segmentation is enabled (IntraSubPartSplitType != ISP_NO_SPLIT)
[0244] • When the CU subtransform is enabled and the TU is less than the specified size (cu_sbt_flag == 1 and Max(nTbW, nTbH) < 32)
[0245] • When explicit MTS is disabled (both sps_explicit_mts_intra_enabled_flag and sps_explicit_mts_inter_enabled_flag are 0) and PredMode is MODE_INTRA
[0246] In cases other than those described above, the MTS setting unit 152111 is set to mts_idx = 0.
[0247] (Explicit MTS)
[0248] The TU decoding unit 3024 decodes mts_idx from the encoded data when the explicit MTS representation is valid (when sps_explicit_mts_flag is 1).
[0249] (Explicit MTS Limitation)
[0250] The TU decoding unit 3024 can limit the range (category) of the transformation matrix selected by the transformation unit based on whether the quadratic transformation is valid. For example, when both the explicit MTS and the quadratic transformation are valid (stIdx != 0), the TU decoding unit 3024 decodes the mts_idx of the maximum value cMaxSt1. Otherwise, when the quadratic transformation is not valid (stIdx == 0), it decodes the mts_idx of the maximum value cMaxSt0. Here, cMaxSt0 is set to > cMaxSt1.
[0251] (Restriction Example 1)
[0252] When explicit MTS is valid and the second transformation is valid (stIdx != 0), the TU decoding unit 3024 sets mts_idx = 0 (trTypeHor = trTypeVer = 0 = DCT2). In this case, the maximum value of mts_idx, cMax, is 0. Otherwise, it decodes mts_idx as any one of 0, 1, 2, 3, or 4. In this case, the maximum value of mts_idx, cMax, is 4. As described later, mts_idx = 1, 2, 3, or 4 can be combinations of DST7 and DCT8, DCT8 and DST7, or DCT8 and DCT8, respectively. It should be noted that DST7 can also be substituted for DST1, DCT4, or a transformation combining DCT2 with pre- and post-processing.
[0253] (Restriction Example 2)
[0254] When the explicit MTS is valid and the second transformation is valid (stIdx != 0), the TU decoding unit 3024 decodes mts_idx. mts_idx is 0 (trTypeHor = trTypeVer = 0 = DCT2) or 1 (trTypeHor = trTypeVer = 1 = DST7). In this case, the maximum value of mts_idx, cMax = 1. Otherwise, it decodes mts_idx as any one of 0, 1, 2, 3, or 4. In this case, the maximum value of mts_idx, cMax = 4. It should be noted that, as described later, mts_idx = 2, 3, or 4, acting as trTypeHor and trTypeVer respectively, can be combinations of DST7 and DCT8, DCT8 and DST7, or DCT8 and DCT8.
[0255] (Restriction Example 3)
[0256] When explicit MTS is enabled and secondary transformation is activated (stIdx != 0), the TU decoding unit 3024 decodes any one of mts_idx 0, 1, or 2. In this case, the maximum value of mts_idx, cMax, is 2. Otherwise, it decodes any one of mts_idx 0, 1, 2, 3, or 4. In this case, the maximum value of mts_idx, cMax, is 4.
[0257] Based on the above configuration, the effective range of MTS can be limited in the case of a quadratic transformation, thus achieving the effect of simplified coding. For example, in the limitation example 2, no quadratic transformation is performed in the case of DCT8 where the effects overlap, thus achieving the effect of reducing the overhead generated by mts_idx and improving coding efficiency.
[0258] Summary of Explicit MTS
[0259] An image decoding apparatus includes a transformation unit that transforms transform coefficients per transform unit (TU). The transformation unit includes: a second transformation unit that applies a transformation matrix to the input transform coefficients when a quadratic transformation is valid; and a first transformation unit that selects a transformation matrix represented by `mtx_idx` from two or more transformation matrices for each transform coefficient and applies the transformation. The TU decoding unit decodes `mtx_idx` by decoding values in a first range as `mts_idx` when the quadratic transformation is valid (`stIdx != 0`), and by decoding values in a second range, where the second range includes the first range, when the quadratic transformation is not valid (`stIdx == 0`). Furthermore, the TU decoding unit decodes `mts_idx`. Here, `mts_idx` has the following configuration: when the quadratic transformation is valid (`stIdx != 0`), its maximum value is `cMaxSt1`; when the quadratic transformation is not valid (`stIdx == 0`), its maximum value is `cMaxSt0`, where `cMaxSt1` < `cMaxSt0`.
[0260] (Implicit MTS)
[0261] The implicit MTS setting unit 152112 performs the following processing when the implicit MTS is enabled.
[0262] (SM001) Implicit MTS setting unit 152112, when using intra-subpart splitting mode (IntraSubPartSplitType != ISP_NO_SPLIT), such as Figure 14 As shown, the transform type tyTypeHor and tyTypeVer are determined by either IntraPredMode (IntraPredMode) or TU size setting 0 (DCT2) or 1 (DST7).
[0263] (SM002) Implicit MTS setting unit 152112, in addition to the above, and when the sub-block is changed to enabled (cu_sbt_flag == 1), such as Figure 15 As shown, either 1 (DST7) or 2 (DCT8) is set according to cu_sbt_horizontal_flag and cu_sbt_pos_flag to serve as tyTypeHor and tyTypeVer.
[0264] (SM003) Implicit MTS setting unit 152112, in cases other than those described above (default implicit MTS), sets either 0 (DCT2) or 1 (DST7) as tyTypeHor and tyTypeVer based on the TU dimensions (width nTbW, height nTbH). Specifically, as... Figure 21 As shown, when the horizontal transformation type trTypeHor has a width nTbW within the specified range (S1301), it is set to 1 (DCT1) (S1302); otherwise, it is set to 0 (DCT2) (S1303). Similarly, when the vertical transformation type trTypeVer has a height nTbH within the specified range (S1304), it is set to 1 (DCT1) (S1305); otherwise, it is set to 0 (DCT2) (S1306).
[0265] trTypeHor=(nTbW>=4&&nTbW<=16&&nTbW<=nTbH)? 1:0
[0266] trTypeVer=(nTbH>=4&&nTbH<=16&&nTbH<=nTbW)? 1:0
[0267] It should be noted that the scope of the regulations is not limited to the above. For example, it could also be as follows.
[0268] trTypeHor=(nTbW>=4&&nTbW<=8&&nTbW<=nTbH)? 1:0
[0269] trTypeVer=(nTbH>=4&&nTbH<=8&&nTbH<=nTbW)? 1:0
[0270] The default implicit MTS mentioned above is the most common implicit MTS mode.
[0271] (Implementation method 1 of implicit MTS)
[0272] When the secondary transformation is enabled (stIdx != 0), the MTS setting unit 15211 does not perform implicit MTS, but sets implicitMtsEnabled to 0. Specifically, as follows: Figure 20 As shown, in the above (implicitMTS flag setting), the MTS setting unit 152111 sets implicitMtsEnabled=1 (S1504) when any of the following conditions are met and the secondary transformation is not enabled (except stIdx!=0), and sets implicitMtsEnabled=0 (S1505) otherwise.
[0273] (S1501) Intra-subpartSplit is enabled (IntraSubPartSplitType != ISP_NO_SPLIT)
[0274] (S1502) The case where the CU subtransform is enabled and the TU is less than the specified size (cu_sbt_flag == 1 and Max(nTbW, nTbH) < 32)
[0275] • (S1503) Explicit MTS is off (both sps_explicit_mts_intra_enabled_flag and sps_explicit_mts_inter_enabled_flag are 0) and PredMode is MODE_INTRA
[0276] It should be noted that, by use Figure 20 The decisions for S1501 to S1503, indicated by the dashed boxes, can also be different. For example, it could be a case where no intra-frame sub-segmentation decision is made, or a case where no sub-block transformation is performed. In addition, other prediction and transformation decisions can be added.
[0277] Based on the above configuration, even when MTS is valid, implicit MTS is not utilized when using the quadratic transform (stIdx != 0). Therefore, by using DCT2 as the quadratic transform, the coding efficiency is improved.
[0278] (Implementation method 2 of implicit MTS)
[0279] When the MTS setting unit 152112 is in a situation where the MTS flag is enabled (sps_mts_enabled_flag == 1) and the explicit MTS flag is not enabled (explicitMtsEnabled == 0), it can derive trTypHor = trTypeVer = 0. For example, SM000 can be performed before SM001 mentioned above.
[0280] Figure 21 This diagram illustrates the operation of the implicit MTS setting unit 152112.
[0281] (SM000) Implicit MTS setting unit 152112 outputs trTypeHor=trTypeVer=0 when stIdx!=0.
[0282] It should be noted that, as Figure 21 As shown in SM003, the implicit MTS setting unit 152112 can derive the transformation type through the export of the default implicit MTS (SM003), which has already been explained when stIdx == 0. Furthermore, the transformation type can also be exported through SM001 and SM002.
[0283] Based on the above configuration, when the implicit MTS is valid and the quadratic transformation is valid, DCT2 is used as the MTS to improve coding efficiency.
[0284] (Implementation method 3 of implicit MTS)
[0285] Alternatively, the implicit MTS setting unit 152112 may not utilize implicit MTS (for example, setting implicitMtsEnabled to 0) when the secondary transformation is enabled (stIdx != 0) and neither the MTS obtained from the intra-sub-segmentation mode (SM001, IntraSubPartSplitType != ISP_NO_SPLIT) nor the MTS obtained from the sub-block transformation (SM002, cu_sbt_flag == 1) is utilized.
[0286] Furthermore, when the secondary transform is enabled (stIdx != 0) and neither the MTS obtained from the intra-sub-segmentation mode (SM001, IntraSubPartSplitType != ISP_NO_SPLIT) nor the MTS obtained from the sub-block transform (SM002, cu_sbt_flag == 1) is used, trTypeHor = trTypeVer = 0 is derived. For example, as... Figure 22 As shown, SM003′ can also be used instead of SM003 as described above.
[0287] (SM003') The implicit MTS setting unit 152112, in cases other than those described above (default implicit MTS), sets either 0 (DCT2) or 1 (DST7) as tyTypeHor and tyTypeVer based on the secondary transformation and TU dimensions (width nTbW, height nTbH). For example, the implicit MTS setting unit 152112 selects 1 (DCT1) (S1302) when the horizontal transformation type trTypeHor is used, stIdx == 0, and the width nTbW is within the specified range (S1301'), and sets 0 (DCT2) in other cases (S1303). Similarly, when the vertical transformation type trTypeVer is used, stIdx == 0, and the height nTbH is within the specified range (S1304'), it sets 1 (DCT1) (S1305), and sets 0 (DCT2) in other cases (S1306).
[0288] trTypeHor=(stIdx==0&&nTbW>=4&&nTbW<=16&&nTbW<=nTbH)? 1:0
[0289] trTypeVer=(stIdx==0&&nTbH>=4&&nTbH<=16&&nTbH<=nTbW)? 1:0
[0290] Based on the above structure, when the implicit MTS is valid and the quadratic transformation is valid, DCT2 is used as the default MTS, thereby improving coding efficiency.
[0291] The MTS setting unit 15211 derives the index trType of the transformation set used by the following formula and outputs it to the coefficient transformation processing unit 15212. The coefficient transformation processing unit 15212 outputs the input trType to the transformation matrix export unit 152131. The MTS setting unit 152111 derives the value representing the MTS used by the following formula.
[0292] When mts_idx == 0, trTypeHor = 0 and trTypeVer = 0.
[0293] With mts_idx == 1, trTypeHor = 1 and trTypeVer = 1
[0294] With mts_idx == 2, trTypeHor = 2 and trTypeVer = 1
[0295] With mts_idx == 3, trTypeHor = 1 and trTypeVer = 2
[0296] With mts_idx == 4, trTypeHor = 2 and trTypeVer = 2
[0297] It should be noted that the transformation basis corresponding to the cases where tyType(trTypeHor or trTypeVer) is 0, 1, or 2 can be DCT2, DST7, or DCT8.
[0298] The coefficient transformation processing unit 15212 is composed of a vertical transformation unit 152121 that performs vertical transformation on the modified transformation coefficients d[][] and a horizontal transformation unit 152123 that performs horizontal transformation.
[0299] The vertical transformation unit 152121 (coefficient transformation processing unit 15212) performs the following processing.
[0300] e[x][y]=∑(transMatrix[y][j]×d[x][j])(j=0..nTbS-1)
[0301] Here, `transMatrix[][]` (=transMatrixV[][]) is the transformation basis represented by an nTbS×nTbS matrix derived using `trTypeVer`. `nTbS` is the height of `TU`, `nTbH`. In the case of a 4×4 transformation of DCT2 with `trType==0` (nTbS=4), for example, `transMatrix={{29, 55, 74, 84}{74, 74, 0, -74}{84, -29, -74, 55}{55, -84, 74, -29}}` is used. The symbol `∑` refers to the addition of the product of the matrix `transMatrix[y][j]` and the transformation coefficients `d[x][j]` to the subscript `j` up to `j=0..nTbS-1`. That is, e[x][y] is obtained by arranging the columns obtained by the product of the vector x[j] (j = 0..nTbS-1) consisting of the columns of d[x][j] (j = 0..nTbS-1) and the elements of the matrix transMatrix[y][j].
[0302] The intermediate truncation unit 152122 derives the intermediate value g[][] by truncating the intermediate value e[][] and transmits it to the horizontal transformation unit 152123.
[0303] g[x][y]=Clip3(coeffMin, coeffMax, (e[x][y]+64)>>7)
[0304] In the above formula, 64 and 7 are values determined based on the bit depth of the transform basis, which is assumed to be 7 bits. Furthermore, coeffMin and coeffMax are the minimum and maximum values truncated.
[0305] The horizontal transformation unit 152123 (coefficient transformation processing unit 15212) performs the following processing. transMatrix[][] (=transMatrixH[][]) is the transformation basis represented by an nTbS×nTbS matrix derived using trTypeHor. nTbS is the width nTbW of TU. The horizontal transformation unit 152123 transforms the intermediate value g[x][y] into the prediction residual r[x][y] through a one-dimensional transformation in the horizontal direction.
[0306] r[x][y]=∑transMatrix[x][j]×g[j][y](j=0..nTbS-1)
[0307] The symbol ∑ above refers to the process of adding the product of matrix transMatrix[x][j] and g[j][y] to the index j up to j = 0..nTbS-1. That is, r[x][y] is obtained by arranging the rows obtained by combining g[j][y] (j = 0..nTbS-1) which are the rows of g[x][y] with the matrix transMatrix.
[0308] The predicted residual r[][] is transmitted from the horizontal transformation unit 152123 to the adder 312.
[0309] The vertical transformation unit 152121 and the horizontal transformation unit 152123 are transformed by the matrix transformation processing unit 15213. The matrix transformation processing unit 15213 is composed of the transformation matrix derivation unit 152131 and the transformation processing unit 152132.
[0310] The transformation matrix derivation unit 152131 derives the transformation matrix transMatrix[][] based on the length of TU (nTbW, nTbH) and the index tyType (trTypeHor, trTypeVer) of the kernel transformation.
[0311] The matrix transformation processing unit 15213 transforms the one-dimensional array xx[j] input using the derived transformation matrix transMatrix[][] into a one-dimensional array yy[i], performing both vertical and horizontal transformations. In the vertical transformation, the transformation coefficients d[x][j] of column x are used as the one-dimensional transformation coefficients xx[j] input. In the horizontal transformation, the middle coefficients g[j][y] of row y are used as the input xx[j].
[0312] yy[i]=∑(transMatrix[i][j]×xx[j])(j=0..nTbS-1)
[0313] <Quadratic Transformation>
[0314] When decoding the secondary transform stIdx, the TU decoding unit 3024 can limit the range of stIdx values to be decoded based on whether the value of mts_idx is valid.
[0315] (Restriction Example 1)
[0316] When mts_idx is 0, the TU decoding unit 3024 sets the maximum value cMax of stIdx to be decoded to 1, and decodes stIdx = 0 to 2. In other cases, that is, when mts_idx = 1, 2, 3, or 4, stIdx is not decoded from the encoded data, and stIdx = 0 is derived.
[0317] (Restriction Example 2)
[0318] When mts_idx is 0…1, the TU decoding unit 3024 sets the maximum value cMax of stIdx to be decoded to 1, and decodes stIdx = 0 to 2. In other cases, that is, when mts_idx = 2, 3, or 4, stIdx is not decoded from the encoded data, and stIdx = 0 is derived instead.
[0319] (Restriction Example 3)
[0320] When mts_idx is 0, 1, or 2, the TU decoding unit 3024 sets the maximum value cMax of stIdx to be decoded to 1, and decodes stIdx = 0 to 2. In other cases, that is, when mts_idx = 3 or 4, stIdx is not decoded from the encoded data, and stIdx = 0 is derived.
[0321] Based on the above configuration, the variable stIdx, representing the category of the quadratic transform, is decoded only when a transformation within a specified range is performed via MTS. Therefore, the effective range of the quadratic transform can be limited, thus achieving the effect of simplified encoding. Furthermore, for example, in Limitation Example 2, the quadratic transform is not performed in the case of DCT8 where the quadratic transform and its effect overlap. This reduces the overhead generated by stIdx and improves encoding efficiency.
[0322] The adder 312 adds the predicted image of the block input from the predicted image generation unit 308 to the prediction error input from the inverse quantization / inverse transform unit 311 for each pixel to generate the decoded image of the block. The adder 312 stores the decoded image of the block in the reference image memory 306 and outputs it to the loop filter 305.
[0323] (Composition of a motion picture encoding device)
[0324] Next, the configuration of the motion image encoding device 11 in this embodiment will be described. Figure 16 This is a block diagram illustrating the configuration of the motion picture encoding apparatus 11 according to this embodiment. The motion picture encoding apparatus 11 is configured to include: a prediction image generation unit 101, a subtraction unit 102, a transform / quantization unit 103, an inverse quantization / inverse transform unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference image memory (reference image storage unit, frame memory) 109, an encoding parameter determination unit 110, a parameter encoding unit 111, and an entropy encoding unit 104.
[0325] The prediction image generation unit 101 generates a prediction image based on the regions (CU) formed by dividing each image T into its constituent parts. The prediction image generation unit 101 performs the same operation as the prediction image generation unit 308 described previously, and its description is omitted here.
[0326] The subtraction unit 102 subtracts the pixel values of the predicted image of the block input from the prediction image generation unit 101 from the pixel values of the image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the transform / quantization unit 103.
[0327] The transform / quantization unit 103 calculates the transform coefficients from the prediction error input from the subtraction unit 102 through frequency transformation, and derives the quantized transform coefficients through quantization. The transform / quantization unit 103 outputs the quantized transform coefficients to the entropy encoding unit 104 and the inverse quantization / inverse transform unit 105.
[0328] like Figure 19 As shown, the transformation / quantization unit 103 includes a positive core transformation 10321 (first transformation unit) and a positive quadratic transformation unit 10322 (second transformation unit).
[0329] In the positive quadratic transform applied in the motion picture coding device 11, the processing is roughly the same as that of the quadratic transform applied in the motion picture decoding device 31, except that the processing S1-S4 are applied in reverse order of processing S1, S4, S3, S2.
[0330] In processing S1, except that the input and output of the quadratic transformation are lengths nStOutSize and nonZeroSize respectively, the positive quadratic transformation unit 10322 performs the same processing as the quadratic transformation unit 31121.
[0331] In processing S4, the positive quadratic transformation unit 10322 derives a one-dimensional array v[] of nStOutSize (or nStSize*nStSize) from the transformation coefficients d[][] at a specified position in TU.
[0332] In processing S3, the positive quadratic transformation unit 10322 obtains the nonZeroSize one-dimensional array u[] (vector V) from the one-dimensional array v[] (vector V) of nStOutSize and the transformation matrix T[][] through the following transformation.
[0333] F = trans(T) × V
[0334] Here, trans(T) is the transpose of T. The quadratic transformation part can also be derived from the one-dimensional array u[] (vector F) by the following formula.
[0335] F = Tinv × V
[0336] Here, Tinv is the inverse matrix of T. T is composed of the first transformation basis T1 and the second transformation basis T2. It should be noted that, alternatively, the quadratic transformation part can use an orthogonal matrix with respect to T, thereby setting trans(T) of T as Tinv.
[0337] It should be noted that in actual processing, T is a matrix with integer values, so it is not T×Tinv=I (the identity matrix), but a constant multiple of the identity matrix (T×Tinv=K2×I, where K2 is a constant). In this case, the quadratic transformation part can also use a matrix that is a constant multiple of the inverse matrix as Tinv, while the transpose matrix is used directly.
[0338] In processing S2, the positive quadratic transformation unit 10322 rearranges the one-dimensional array u[] of nonZeroSize into a two-dimensional arrangement and derives the transformation coefficients d[][].
[0339]
[0340] Inverse quantization / inverse transform unit 105 and inverse quantization / inverse transform unit 311 in motion image decoding device 31 Figure 15 (Same as above, explanation omitted.) The calculated prediction error is input into the adder 106.
[0341] In the entropy coding unit 104, quantization transformation coefficients are input from the transform / quantization unit 103, and coding parameters are input from the parameter coding unit 111. The coding parameter is, for example, predMode.
[0342] The entropy coding unit 104 entropy codes the segmentation information, prediction parameters, quantization transformation coefficients, etc. to generate a coded stream Te and outputs it.
[0343] The parameter coding unit 111 includes: a header coding unit 1110 (not shown), a CT information coding unit 1111, a CU coding unit 1112 (prediction mode coding unit), an inter-frame prediction parameter coding unit 112, and an intra-frame prediction parameter coding unit 113. The CU coding unit 1112 also includes a TU coding unit 1114.
[0344] The following is a brief description of the general operation of each module. The parameter encoding unit 111 performs encoding processing of parameters such as header information, segmentation information, prediction information, and quantization transformation coefficients.
[0345] The CT information encoding unit 1111 encodes QT, MT (BT, TT) segmentation information, etc., based on the encoding data.
[0346] The CU encoding unit 1112 encodes CU information, prediction information, TU segmentation flag, CU residual flag, etc.
[0347] When the prediction error is included in the TU, the TU encoding unit 1114 encodes the QP update information (quantization correction value) and the quantization prediction error (residual_coding).
[0348] The CT information coding unit 1111 and the CU coding unit 1112 provide inter-frame prediction parameters, intra-frame prediction parameters (intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_remainder), quantization transform coefficients and other syntax elements to the entropy coding unit 104.
[0349] (The structure of the intra-frame prediction parameter coding unit 113)
[0350] The intra-prediction parameter encoding unit 113 derives the encoding format (e.g., intra_luma_mpm_idx, intra_luma_mpm_remainder, etc.) based on the IntraPredMode input from the encoding parameter determination unit 110. The intra-prediction parameter encoding unit 113 includes the same configuration as the configuration portion of the intra-prediction parameters derived by the intra-prediction parameter decoding unit 304.
[0351] Figure 17 This is a schematic diagram showing the configuration of the intra-prediction parameter coding unit 113 of the parameter coding unit 111. The intra-prediction parameter coding unit 113 is configured to include: a parameter coding control unit 1131, a luma intra-prediction parameter derivation unit 1132, and a chromatic intra-prediction parameter derivation unit 1133.
[0352] The encoding parameter determination unit 110 inputs IntraPredModeY and IntraPredModeC to the parameter encoding control unit 1131. The parameter encoding control unit 1131 determines intra_luma_mpm_flag by referring to the mpmCandList[] of the MPM candidate list derivation unit 30421. Then, intra_luma_mpm_flag and IntraPredModeY are output to the luma intra-prediction parameter derivation unit 1132. In addition, IntraPredModeC is output to the chroma intra-prediction parameter derivation unit 1133.
[0353] The luminance intra-frame prediction parameter derivation unit 1132 is configured to include: an MPM candidate list derivation unit 30421 (candidate list derivation unit), an MPM parameter derivation unit 11322 (parameter derivation unit), and a non-MPM parameter derivation unit 11323 (coding unit, derivation unit).
[0354] The MPM candidate list derivation unit 30421 derives `mpmCandList[]` by referring to the intra-prediction modes of adjacent blocks stored in the prediction parameter memory 108. The MPM parameter derivation unit 11322, when `intra_luma_mpm_flag` is 1, derives `intra_luma_mpm_idx` from `IntraPredModeY` and `mpmCandList[]`, and outputs it to the entropy coding unit 104. The non-MPM parameter derivation unit 11323, when `intra_luma_mpm_flag` is 0, derives `RemIntraPredMode` from `IntraPredModeY` and `mpmCandList[]`, and outputs `intra_luma_mpm_remainder` to the entropy coding unit 104.
[0355] The chroma intra-pred prediction parameter export unit 1133 exports intra_chroma_pred_mode from IntraPredModeY and IntraPredModeC and outputs it.
[0356] The addition unit 106 adds the pixel values of the block prediction image input from the prediction image generation unit 101 and the prediction error input from the inverse quantization / inverse transform unit 105 to generate a decoded image by adding each pixel. The addition unit 106 stores the generated decoded image in the reference image memory 109.
[0357] The loop filter 107 performs deblocking filtering, SAO, and ALF on the decoded image generated by the adder 106. It should be noted that the loop filter 107 does not necessarily include the above three filters; for example, it may only include a deblocking filter.
[0358] The prediction parameter memory 108 stores the prediction parameters generated by the encoding parameter determination unit 110 in a predetermined location for each object image and CU.
[0359] The image memory 109 stores the decoded images generated by the loop filter 107 in predetermined locations for each object image and each CU.
[0360] The encoding parameter determination unit 110 selects one set from a plurality of sets of encoding parameters. The encoding parameters refer to the QT, BT, or TT segmentation information, prediction parameters, or parameters generated in association with them as encoding objects, as described above. The prediction image generation unit 101 uses these encoding parameters to generate a prediction image.
[0361] The encoding parameter determination unit 110 calculates the RD cost value, representing the information content and encoding error, for each of the multiple sets. The encoding parameter determination unit 110 selects the set of encoding parameters with the smallest calculated cost value. Therefore, the entropy encoding unit 104 outputs the selected set of encoding parameters as the encoded stream Te. The encoding parameter determination unit 110 stores the determined encoding parameters in the prediction parameter memory 108.
[0362] It should be noted that a portion of the motion picture encoding device 11 and motion picture decoding device 31 described above can be implemented using a computer. For example, this includes the entropy decoding unit 301, parameter decoding unit 302, loop filter 305, prediction image generation unit 308, inverse quantization / inverse transform unit 311, addition unit 312, prediction image generation unit 101, subtraction unit 102, transform / quantization unit 103, entropy encoding unit 104, inverse quantization / inverse transform unit 105, loop filter 107, encoding parameter determination unit 110, and parameter encoding unit 111. In this case, the program for implementing this control function can be recorded on a computer-readable recording medium, and the computer system can read and execute the program recorded on the recording medium. It should be noted that the "computer system" mentioned here refers to a computer system built into either the motion picture encoding device 11 or the motion picture decoding device 31, employing hardware including an operating system and peripheral devices. Furthermore, "computer-readable recording medium" refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. Moreover, "computer-readable recording medium" can include: media that dynamically stores programs for a short period of time, like a communication line used to transmit programs via a network such as the Internet or a communication line such as a telephone line; and media that stores programs for a fixed period of time, like volatile memory within a computer system serving as a server or client in this case. Furthermore, the program can be a program used to implement the functions described above, or it can be a program that can implement the functions described above by combining with programs already recorded in the computer system.
[0363] Furthermore, the motion picture encoding device 11 and motion picture decoding device 31 in the above embodiments can also be implemented as integrated circuits such as LSI (Large Scale Integration). Each functional block of the motion picture encoding device 11 and motion picture decoding device 31 can be processorized individually, or some or all can be integrated for processorization. Moreover, the method of integrated circuit implementation is not limited to LSI; it can also be implemented using dedicated circuits or general-purpose processors. Furthermore, if advancements in semiconductor technology lead to integrated circuit technologies that replace LSI, integrated circuits based on such technologies can also be used.
[0364] The above description, with reference to the accompanying drawings, details one embodiment of the invention. However, the specific configuration is not limited to the above embodiment, and various design changes can be made without departing from the spirit of the invention.
[0365] (Application Example)
[0366] The aforementioned moving image encoding device 11 and moving image decoding device 31 can be mounted on various devices for transmitting, receiving, recording, and reproducing moving images. It should be noted that the moving images can be natural moving images captured by a camera or the like, or artificial moving images (including CG and GUI) generated by a computer or the like.
[0367] First, refer to Figure 2 The following describes the situation where the above-described motion picture encoding device 11 and motion picture decoding device 31 can be used for the transmission and reception of motion pictures.
[0368] Figure 2 The diagram shows the configuration of the transmitting device PROD_A, which is equipped with the motion picture encoding device 11. Figure 2 As shown, the transmitting device PROD_A includes: an encoding unit PROD_A1 that obtains encoded data by encoding a moving image; a modulation unit PROD_A2 that obtains a modulated signal by modulating a carrier wave using the encoded data obtained by the encoding unit PROD_A1; and a transmitting unit PROD_A3 that transmits the modulated signal obtained by the modulation unit PROD_A2. The moving image encoding device 11 described above is used as the encoding unit PROD_A1.
[0369] As a source of motion images input to the encoding unit PROD_A1, the transmitting device PROD_A may further include: a camera PROD_A4 for capturing motion images, a recording medium PROD_A5 for recording motion images, an input terminal PROD_A6 for inputting motion images from an external source, and an image processing unit A7 for generating or processing images. Figure 2 The example shows that the sending device PROD_A has all of these components, but some can be omitted.
[0370] It should be noted that the recording medium PROD_A5 can be a medium that records unencoded motion images, or a medium that records motion images encoded using a recording encoding method different from the encoding method used for transmission. In the latter case, it is preferable that the decoding unit (not shown) that decodes the encoded data read from the recording medium PROD_A5 according to the recording encoding method is located between the recording medium PROD_A5 and the encoding unit PROD_A1.
[0371] In addition, Figure 2 The diagram shows the configuration of the receiving device PROD_B equipped with the motion picture decoding device 31. Figure 2 As shown, the receiving device PROD_B includes: a receiving unit PROD_B1 for receiving a modulated signal, a demodulation unit PROD_B2 for obtaining coded data by demodulating the modulated signal received by the receiving unit PROD_B1, and a decoding unit PROD_B3 for obtaining a moving image by decoding the coded data obtained by the demodulation unit PROD_B2. The aforementioned moving image decoding device 31 is used as the decoding unit PROD_B3.
[0372] The receiving device PROD_B, serving as the destination for the moving images output by the decoding unit PROD_B3, may further include a display PROD_B4 for displaying the moving images, a recording medium PROD_B5 for recording the moving images, and an output terminal PROD_B6 for outputting the moving images to an external device. Figure 2 The example shows that the receiving device PROD_B has all of these components, but some can be omitted.
[0373] It should be noted that the recording medium PROD_B5 can be a medium for recording unencoded motion images, or it can be a medium encoded with a recording encoding method different from the encoding method used for transmission. In the latter case, it is preferable that the encoding unit (not shown) that encodes the motion images acquired from the decoding unit PROD_B3 according to the recording encoding method is located between the decoding unit PROD_B3 and the recording medium PROD_B5.
[0374] It should be noted that the transmission medium for modulated signals can be wireless or wired. Furthermore, the transmission scheme for modulated signals can be broadcast (here, a transmission scheme where the destination is not predetermined) or communication (here, a transmission scheme where the destination is predetermined). That is, the transmission of modulated signals can be achieved through any of the following: wireless broadcasting, wired broadcasting, wireless communication, and wired communication.
[0375] For example, a terrestrial digital broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals via wireless broadcasting. Similarly, a cable television broadcasting station (broadcasting equipment, etc.) / receiving station (television receiver, etc.) is an example of a transmitting device PROD_A / receiving device PROD_B that transmits and receives modulated signals via cable broadcasting.
[0376] Furthermore, servers (workstations, etc.) and clients (TV receivers, personal computers, smartphones, etc.) using internet-based VOD (Video On Demand) services, moving image sharing services, etc., are examples of transmitting devices PROD_A and receiving devices PROD_B that transmit and receive modulated signals via communication (typically, either wireless or wired is used as the transmission medium in a LAN, and wired is used in a WAN). Here, personal computers include desktop PCs, laptop PCs, and tablet PCs. Additionally, smartphones also include multi-functional portable telephone terminals.
[0377] It should be noted that, in addition to decoding the encoded data downloaded from the server and displaying it on the screen, the client of the motion picture sharing service also has the function of encoding motion pictures captured by a camera and uploading them to the server. That is, the client of the motion picture sharing service performs the functions of both the sending device PROD_A and the receiving device PROD_B.
[0378] Next, refer to Figure 3 The following describes the situation where the above-mentioned motion picture encoding device 11 and motion picture decoding device 31 can be used for recording and reproducing motion pictures.
[0379] Figure 3 The diagram shows a block diagram illustrating the configuration of a recording device PROD_C equipped with the aforementioned motion picture encoding device 11. Figure 3 As shown, the recording device PROD_C includes an encoding unit PROD_C1 that obtains encoded data by encoding moving images, and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 to the recording medium PROD_M. The moving image encoding device 11 described above is used as the encoding unit PROD_C1.
[0380] It should be noted that the recording medium PROD_M can be (1) a type of recording medium built into the recording device PROD_C, such as HDD (Hard Disk Drive) or SSD (Solid State Drive), or (2) a type of recording medium connected to the recording device PROD_C, such as SD memory card or USB (Universal Serial Bus) flash memory, or (3) a recording medium loaded into a drive (not shown) built into the recording device PROD_C, such as DVD (Digital Versatile Disc) or BD (Blu-ray Disc).
[0381] Furthermore, as a source of motion images input to the encoding unit PROD_C1, the recording device PROD_C may further include: a camera PROD_C3 for capturing motion images, an input terminal PROD_C4 for inputting motion images from the outside, a receiving unit PROD_C5 for receiving motion images, and an image processing unit PROD_C6 for generating or processing images. Figure 3 The example shows that the recording device PROD_C has all of these components, but some can be omitted.
[0382] It should be noted that the receiving unit PROD_C5 can receive unencoded motion images, or it can receive encoded data encoded using a transmission encoding method different from the encoding method used for recording. In the latter case, it is preferable to place the transmission decoding unit (not shown) that decodes the encoded data encoded using the transmission encoding method between the receiving unit PROD_C5 and the encoding unit PROD_C1.
[0383] Examples of such recording devices PROD_C include DVD recorders, BD recorders, and HDD (Hard Disk Drive) recorders (in which case the input terminal PROD_C4 or the receiving unit PROD_C5 is the main source of moving images). Furthermore, portable camcorders (in which case the camera PROD_C3 is the main source of moving images), personal computers (in which case the receiving unit PROD_C5 or the image processing unit C6 is the main source of moving images), and smartphones (in which case the camera PROD_C3 or the receiving unit PROD_C5 is the main source of moving images) are also examples of such recording devices PROD_C.
[0384] In addition, Figure 3 The diagram shows the configuration of the playback device PROD_D equipped with the aforementioned motion picture decoding device 31. Figure 3 As shown, the playback device PROD_D includes a readout unit PROD_D1 that reads encoded data written to the recording medium PROD_M, and a decoding unit PROD_D2 that obtains a moving image by decoding the encoded data read out by the readout unit PROD_D1. The aforementioned moving image decoding device 31 is used as the decoding unit PROD_D2.
[0385] It should be noted that the recording medium PROD_M can be (1) a recording medium built into the playback device PROD_D, such as HDD or SSD, or (2) a recording medium connected to the playback device PROD_D, such as SD memory card or USB flash drive, or (3) a recording medium loaded into a drive device (not shown) built into the playback device PROD_D, such as DVD or BD.
[0386] Furthermore, as the destination for the motion images output by the decoding unit PROD_D2, the playback device PROD_D may further include: a display PROD_D3 for displaying motion images, an output terminal PROD_D4 for outputting motion images to the outside, and a transmitting unit PROD_D5 for transmitting motion images. Figure 3 The example shows that the reproduction device PROD_D has all of these components, but some can be omitted.
[0387] It should be noted that the transmitting unit PROD_D5 can transmit unencoded motion images, or it can transmit encoded data encoded using a transmission encoding method different from the encoding method used for recording. In the latter case, it is preferable to place the encoding unit (not shown) that encodes the motion images using the transmission encoding method between the decoding unit PROD_D2 and the transmitting unit PROD_D5.
[0388] Examples of such playback devices PROD_D include DVD players, BD players, HDD players, etc. (in which case, the output terminal PROD_D4 connected to a TV receiver, etc., is the main destination for the moving images). Other examples include TV receivers (in which the display PROD_D3 is the main destination for the moving images), digital signage (also called electronic billboards, electronic bulletin boards, etc., where the display PROD_D3 or the transmitter PROD_D5 is the main destination for the moving images), desktop PCs (in which the output terminal PROD_D4 or the transmitter PROD_D5 is the main destination for the moving images), laptop or tablet PCs (in which the display PROD_D3 or the transmitter PROD_D5 is the main destination for the moving images), and smartphones (in which the display PROD_D3 or the transmitter PROD_D5 is the main destination for the moving images).
[0389] (Hardware implementation and software implementation)
[0390] Furthermore, each of the aforementioned motion picture decoding device 31 and motion picture encoding device 11 can be implemented in hardware using logic circuits formed on an integrated circuit (IC chip), or in software using a CPU (Central Processing Unit).
[0391] In the latter case, the aforementioned devices include: a CPU that executes commands for programs that perform various functions; a ROM (Read Only Memory) that stores the programs; a RAM (Random Access Memory) that expands the programs; and a memory that stores the programs and various data, etc., such as a storage device (recording medium). Furthermore, the objective of embodiments of the present invention is to achieve this by supplying a recording medium containing program code (executable form program, intermediate code program, source program) of the software implementing the aforementioned functions, i.e., the control program of the aforementioned devices, in a computer-readable manner to the aforementioned devices, wherein the computer (or CPU, MPU) reads the program code recorded on the recording medium and executes it.
[0392] As recording media, the following can be used: tapes, cassette tapes, etc.; disks including floppy disks (registered trademark) / hard disks, CD-ROMs (Compact Disc Read-Only Memory), MO discs (Magneto-Optical Disc), MD discs (Mini Disc), DVDs (Digital Versatile Disc), CD-Rs (CD Recordable), Blu-ray discs (registered trademark), etc.; cards (including memory cards) / optical cards, etc.; semiconductor memory types such as mask ROMs, EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable and Programmable Read-Only Memory), and flash memory ROMs; or logic circuits such as PLDs (Programmable Logic Devices) and FPGAs (Field Programmable Gate Arrays), etc.
[0393] Furthermore, the aforementioned devices can be configured to connect to a communication network and supply the program code via the communication network. This communication network need only be capable of transmitting program code and is not particularly limited. For example, the Internet, intranet, extranet, LAN (Local Area Network), ISDN (Integrated Services Digital Network), VAN (Value-Added Network), CATV (Community Antenna television / Cable Television) communication network, Virtual Private Network, telephone line network, mobile communication network, satellite communication network, etc., can be used. Furthermore, the transmission medium constituting this communication network need only be a medium capable of transmitting program code and is not limited to a specific configuration or type. For example, it can be used in wired networks such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line transmission, cable TV lines, telephone lines, and ADSL (Asymmetric Digital Subscriber Line) lines, as well as in wireless networks such as IrDA (Infrared Data Association), infrared (like remote controls), Bluetooth (registered trademark), IEEE 802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance, registered trademark), mobile phone networks, satellite lines, and terrestrial digital broadcasting networks. It should be noted that embodiments of the present invention can also be implemented as computer data signals with embedded carriers that embody the above-mentioned program code via electronic transmission.
[0394] The embodiments of the present invention are not limited to the embodiments described above, and various modifications can be made within the scope of the claims. That is, embodiments obtained by combining technical solutions with appropriate modifications within the scope of the claims are also included within the technical scope of the present invention.
[0395] Industrial availability
[0396] The embodiments of the present invention are preferably applied to a moving image decoding apparatus for decoding encoded data obtained by encoding image data, and to a moving image encoding apparatus for generating encoded data obtained by encoding image data. Furthermore, they are preferably applied to a data structure of encoded data generated by the moving image encoding apparatus and referenced by the moving image decoding apparatus.
[0397] (Mutual references between related applications)
[0398] This application claims priority to Japanese Patent Application No. 2019-101179, filed on May 30, 2019, and incorporates the entire contents of which hereby refer to.
[0399] Explanation of reference numerals in the attached figures
[0400] 31 Motion Image Decoding Device
[0401] 301 Entropy Decoding Department
[0402] 302 Parameter Decoding Unit
[0403] 3020 Header Decoding Department
[0404] 303 Inter-frame Prediction Parameter Decoding Unit
[0405] 304-frame intra-prediction parameter decoding unit
[0406] 308 Predictive Image Generation Unit
[0407] 309 Inter-frame Prediction Image Generation Unit
[0408] 310-frame intra-prediction image generation unit
[0409] 311 Inverse Quantization / Inverse Transformation Unit
[0410] 312 Addition Department
[0411] 11. Motion Picture Coding Device
[0412] 101 Predictive Image Generation Unit
[0413] 102 Subtraction Department
[0414] 103 Transformation / Quantization Section
[0415] 104 Entropy Coding Department
[0416] 105 Inverse Quantization / Inverse Transformation Unit
[0417] 107 loop filter
[0418] 110 Coding Parameter Determination Unit
[0419] 111 Parameter Encoding Section
[0420] 112 Inter-frame Prediction Parameter Coding Unit
[0421] 113 Intra-frame Prediction Parameter Coding Unit
[0422] 1110 Header Code Department
[0423] 1111CT Information Coding Department
[0424] 1112CU Coding Unit (Predictive Mode Coding Unit)
[0425] 1114TU Coding Department
[0426] 3111 Inverse Quantization Department
[0427] 3112 Inverse Transformer
[0428] 31121 Secondary Transformation Unit
[0429] 31112 Scaling section
[0430] 31123 Nuclear Transformation Unit
[0431] 10322 positive quadratic transform unit
[0432] 10323 Positron Transformation Unit.
Claims
1. An image decoding apparatus, wherein transforming transform coefficients according to each transform unit, characterized in that, The transformation is performed by at least one of a kernel transformation and a quadratic transformation. The kernel transform is a split transform consisting of a vertical transform and a horizontal transform. The second transformation is a transformation applied to the transformation coefficients of a portion or all of the transformation unit before the kernel transformation. In the case of using implicit multiple transform selection (MTS) in the kernel transform, parameters specifying the kernel transform type are derived based on multiple encoding parameters. In the absence of the implicit MTS in the kernel transformation, parameters specifying the kernel transformation type are derived based on an encoding parameter. Without performing the aforementioned secondary transformation, based on the first flag, the second flag, the block size of the object block, the third flag, and the prediction mode of the object block, it is determined whether to use the implicit MTS. The first flag indicates whether intra-frame sub-segmentation is performed, the second flag indicates whether sub-block transformation is performed, and the third flag indicates whether there are parameters specifying the kernel transformation type in the vertical / horizontal direction.
2. The image decoding device according to claim 1, characterized in that, Without performing the aforementioned secondary transformation The first flag indicates a division, or The second flag indicates that the object block is segmented and the block size is smaller than a specified size, or If the third flag is false and the prediction mode of the object block is intra-frame prediction, then the implicit MTS is used.
3. The image decoding apparatus according to claim 1, characterized in that, Without performing the aforementioned quadratic transformation and using the implicit MTS. If the second flag is true, the parameter specifying the kernel transformation type is set to 1 or 2 using a table. If the second flag is false, the parameter specifying the kernel transformation type is set to 0 or 1 using the block size of the object block.
4. The image decoding apparatus according to claim 1 or 2, characterized in that, When performing the second transformation, the parameter specifying the kernel transformation type is set to 0.
5. An image encoding apparatus, wherein transform coefficients are transformed according to each transform unit, characterized in that, The transformation is performed by at least one of a kernel transformation and a quadratic transformation. The kernel transform is a split transform consisting of a vertical transform and a horizontal transform. The second transformation is a transformation applied to the transformation coefficients of a portion or all of the transformation unit before the kernel transformation. In the case where implicit multiple transform selection (MTS) is used in the kernel transform, parameters specifying the kernel transform type are derived based on multiple encoding parameters. In the absence of the implicit MTS in the kernel transformation, parameters specifying the kernel transformation type are derived based on an encoding parameter. Without performing the aforementioned secondary transformation, based on the first flag, the second flag, the block size of the object block, the third flag, and the prediction mode of the object block, it is determined whether to use the implicit MTS. The first flag indicates whether intra-frame sub-segmentation is performed, the second flag indicates whether sub-block transformation is performed, and the third flag indicates whether there are parameters specifying the kernel transformation type in the vertical / horizontal direction.
6. A computer-readable recording medium containing a program that causes a computer to transform the transformation coefficients at each transformation unit, characterized in that, The transformation is performed by at least one of a kernel transformation and a quadratic transformation. The kernel transform is a split transform consisting of a vertical transform and a horizontal transform. The second transformation is a transformation applied to the transformation coefficients of a portion or all of the transformation unit before the kernel transformation. The program causes the computer to perform the following steps: In the case of using implicit multitransform selection (MTS) in the kernel transform, parameters for specifying the kernel transform type are derived based on multiple encoding parameters; In the absence of the implicit MTS in the kernel transformation, parameters specifying the kernel transformation type are derived based on an encoding parameter; and Without performing the aforementioned secondary transformation, based on the first flag, the second flag, the block size of the object block, the third flag, and the prediction mode of the object block, it is determined whether to use the implicit MTS. The first flag indicates whether intra-frame sub-segmentation is performed, the second flag indicates whether sub-block transformation is performed, and the third flag indicates whether there are parameters specifying the kernel transformation type in the vertical / horizontal direction.
Citation Information
Patent Citations
Fixation device and image formation device
JP2019101179A